Methods for predicting active compounds

Molecular dynamics simulations and machine learning models predict absolute ligand binding free energy and off-rates, addressing inefficiencies in existing methods by simulating unbinding events and ranking ligands, enabling efficient high-throughput drug discovery.

WO2025226957A1PCT designated stage Publication Date: 2025-10-30ALIVEXIS INC +3
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/026222
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-26
Filing Date
2025-04-24
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Current methods for predicting ligand off-rates (koff) and binding free energy in drug discovery are limited by insubstantial validation, reliance on relative free energy changes, and require structural similarity between ligands, leading to inefficiencies and inaccuracies in high-throughput virtual screening.

Method used

A method involving molecular dynamics simulations to predict absolute ligand binding free energy by simulating unbinding events, calculating temporal information and dissociation free energy, and using machine learning models to rank candidate ligands based on empirical binding free energy predictions.

Benefits of technology

Enables accurate and efficient prediction of ligand binding free energy and off-rates, suitable for high-throughput screening without requiring structural similarity, significantly faster than existing methods, supporting hit-to-lead and lead optimization in drug discovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025026222_30102025_PF_FP_ABST
    Figure US2025026222_30102025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure is related generally to molecular dynamics (MD) simulation methods, systems, and devices, including computer-assisted methods for predicting active pharmaceutical compounds. The present disclosure relates to simulation-based prediction of ligand off-rates, and more particularly, to simulation-based prediction of ligand binding free energy, which is useful for predicting drug activity.
Need to check novelty before this filing date? Find Prior Art

Description

Atty Ref.: AVX0003PCT METHODS FOR PREDICTING ACTIVE COMPOUNDS CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority to U.S. Provisional Patent Application Serial No.63 / 639,473 filed: 26 April 2024, and titled: Methods for Predicting Active Compounds, which is incorporated herein in its entirety. BACKGROUND

[0002] 1. Field. The present disclosure is related generally to molecular dynamics (MD) simulation, pharmaceutical drug discovery, and computational chemistry, and more particularly, to computer-assisted methods for predicting active compounds. The present disclosure relates to simulation-based prediction of ligand off-rates, and more particularly, to simulation-based prediction of ligand binding free energy.

[0003] 2. Background Information. In rational drug discovery approaches, both the free energy of binding and binding half-life (koff) are important factors in determining the efficacy of drugs. Computational methods have been developed to predict the free energy of binding and binding half-life (koff) and may rely on molecular dynamics (MD) simulations.

[0004] Whilemethods have been described as having the ability to drive daily synthesis decisions in a commercial drug discovery setting, the prediction of koff has had limited validation, and the deployment of predictive methods in drug discovery settings are desired. For example, current free energy models rely on relative free energy changes, and therefore, are limited to optimization of a congeneric series. As such, there is a need in the art for methods that allow for absolute prediction of koff and / or absolute prediction of free energy of ligand binding, which require no structural similarity between ligands to facilitate high throughput virtual screening of diverse ligands.

[0005] Binding half-life may drive pharmacodynamic effects in vivo after drug concentration drops, and longer residence times can improve both the selectivity and in vivo efficacy of drugs. Although binding affinity and free energy of binding are properties that may often be the main consideration in lead discovery and optimization, any drug discovery program that focuses solely on the optimization of free energy of binding is inadequate, asAtty Ref.: AVX0003PCT pharmacodynamic effects may greatly affect drug efficacy due to changes in drug concentration over time. In the area of free energy prediction, significant academic and commercial works have demonstrated that with reasonable structural information, some methods can predict the relative free energy differences between a congeneric series of ligands with good accuracy using validation sets of hundreds of compounds that bind targets across a broad range of target classes. However, even among some of the most effective free energy methods, the accuracy associated with the methods may generally be less than one order of magnitude and, in some cases, close to half an order of magnitude. In the best case, the accuracy may support excellent predictivity to help drive synthetic decisions in the hit-to- lead and lead optimization phases of drug discovery. In contrast to binding affinity, the prediction of kinetic on and off rates (kon and koff) has received considerably less attention from the computational chemistry field.

[0006] Although several methods to predict koffhave been published, the validation of the approaches has been insubstantial, and few demonstrations of commercial drug discovery impact have been published. The lack of development of koff prediction tools may be attributable to several factors such as, for example, relative slow processing time or high complexity associated with computational workflows, relative low availability of public datasets of konand koffcompared to binding affinities, and in practice, the binding half-life of compounds is measured much less routinely than binding or inhibition constants (which can be directly related to relative free energy through the Cheng-Prussoff equation). SUMMARY

[0007] A method is disclosed including: simulating, by one or more processors, an unbinding event in which a candidate ligand dissociates from a biomolecular target to which the candidate ligand is bound, wherein the candidate ligand dissociating from the biomolecular target includes the candidate ligand traveling from a first set of coordinates at which the candidate ligand is bound to the biomolecular target to a second set of coordinates at which the candidate ligand is not bound to the biomolecular target; calculating, by the one or more processors based on a result of simulating the unbinding event, temporal information associated with the candidate ligand traveling from the first set of coordinates to the second set of coordinates; calculating, by the one or more processors, a dissociation free energy associated with the candidate ligand and the biomolecular target based on at least one of: theAtty Ref.: AVX0003PCT temporal information, a distance between the first set of coordinates and the second set of coordinates, or a combination thereof; and providing the dissociation free energy as a prediction of empirically derived binding free energy associated with the candidate ligand.

[0008] In any one or combination of the embodiments disclosed herein, the method may include screening the candidate ligand in association with one or more pharmaceutical applications based on the prediction of the empirically derived binding free energy.

[0009] In any one or combination of the embodiments disclosed herein, simulating the unbinding event includes accelerating the unbinding event by performing a scaled molecular dynamics run; and the temporal information and the distance are based on a result of accelerating the unbinding event.

[0010] In any one or combination of the embodiments disclosed herein, the scaledmolecular dynamics run is based on the equation ^^^^⃗ = ^∗^^⃗^^ / ^, where P(^⃗) is population,P*(^⃗) is population of the scaled molecular dynamics run, and λ is a scaling factor.

[0011] In any one or combination of the embodiments disclosed herein, simulating the unbinding event includes accelerating the unbinding event by performing a molecular dynamics run according to a temperature range associated with accelerating the unbinding event; and the temporal information and the distance are based on a result of accelerating the unbinding event.

[0012] In any one or combination of the embodiments disclosed herein, the method may include providing the temporal information to one or more prediction models; and predicting, by the one or more prediction models in response to processing the temporal information associated with the candidate ligand, respective free energy differences of a set of candidate ligands.

[0013] In any one or combination of the embodiments disclosed herein, the method may include generating ranking information associated with the set of candidate ligands based on the respective free energy differences; and screening the set of candidate ligands in association with one or more pharmaceutical application based on the ranking information.

[0014] In any one or combination of the embodiments disclosed herein, the method may include simulating, by the one or more processors, the candidate ligand in a solution; calculating, by the one or more processors based on a result of simulating the candidate ligand in the solution, a free energy associated with the candidate ligand in the solution; andAtty Ref.: AVX0003PCT calculating, by the one or more processors, a binding free energy associated with the candidate ligand based on a difference between the dissociation free energy and the free energy associated with the candidate ligand in the solution.

[0015] In any one or combination of the embodiments disclosed herein, calculating thebinding free energy is based on the equation ∆^^ = ∆^^^^^^ − ∆^^^^^^^^ =−^^^^^^∗^^^⃗^^^^^^⁄ ^ ^ + ^^^^^^∗^^^⃗^^^^^^^^⁄ ^ ^, where: ∆^^ is binding free energy,∆^^^^^^is the dissociation free energy, ∆^^^^^^^^is the free energy, R is a gas constant, T is temperature, ^∗^^^⃗^^^^^ is population associated with simulating the unbinding event, ^∗^^^⃗^^^^^^^ is population associated with simulating the candidate ligand in the solution, and λ is a scaling factor.

[0016] In any one or combination of the embodiments disclosed herein, the method may include training one or more prediction models based at least one of the binding free energy, the dissociation free energy, and the free energy; selecting a compound from a library of candidate compounds, wherein the compound includes a second candidate ligand and a second biomolecular target; and predicting, by the one or more prediction models, a binding free energy associated with the second candidate ligand.

[0017] In any one or combination of the embodiments disclosed herein, the method may include screening the second candidate ligand in association with one or more pharmaceutical applications based on the predicted binding free energy.

[0018] In any one or combination of the embodiments disclosed herein, the one or more prediction models include at least one of: a machine learning model; a statistical model; a quantitative structure-activity relationship (QSAR) model; or a combination thereof.

[0019] In any one or combination of the embodiments disclosed herein, the method may include predicting pose information associated with a bound state of the candidate ligand to the biomolecular target, wherein simulating the unbinding event is based on the pose information.

[0020] In any one or combination of the embodiments disclosed herein, predicting the pose information is performed using at least one of x-ray crystallography, docking, flexible docking, molecular dynamics or a combination thereof.

[0021] In any one or combination of the embodiments disclosed herein, the method may include predicting a solubility of a compound including the candidate ligand and theAtty Ref.: AVX0003PCT biomolecular target based on at least one of: the temporal information; the dissociation free energy; or a combination thereof.

[0022] In any one or combination of the embodiments disclosed herein, the method may include determining a correlation amount between the temporal information or dissociation free energy and the binding free energy; predicting, based on the correlation amount, at least one of: an empirical half maximal inhibitory concentration (IC50) associated with the candidate ligand and a pharmaceutical application; and an inhibitory constant (Ki) associated with the candidate ligand and the pharmaceutical application.

[0023] In some aspects, a method is disclosed including: simulating, by one or more processors, an unbinding event in which the candidate ligand dissociates from a biomolecular target to which the candidate ligand is bound, wherein the candidate ligand dissociating from the biomolecular target includes the candidate ligand traveling from a first set of coordinates at which the candidate ligand is bound to the biomolecular target to a second set of coordinates at which the candidate ligand is not bound to the biomolecular target; calculating, by the one or more processors based on a result of simulating the unbinding event, temporal information associated with the candidate ligand traveling from the first set of coordinates to the second set of coordinates; calculating, by the one or more processors, a dissociation free energy associated with the candidate ligand and the biomolecular target based on at least one of: the temporal information, a distance between the first set of coordinates and the second set of coordinates, or a combination thereof; simulating, by the one or more processors, the candidate ligand in a solution; calculating, by the one or more processors based on a result of simulating the candidate ligand in the solution, a free energy associated with the candidate ligand in the solution; and calculating, by the one or more processors, a binding free energy associated with the candidate ligand based on a difference between the dissociation free energy and the free energy associated with the candidate ligand in the solution.

[0024] In some aspects, the techniques described herein relate to a system including: a processor; and a memory storing instructions thereon that, when executed by the processor, cause the processor to: simulate an unbinding event in which a candidate ligand dissociates from a biomolecular target to which the candidate ligand is bound, wherein the candidate ligand dissociating from the biomolecular target includes the candidate ligand traveling from a first set of coordinates at which the candidate ligand is bound to the biomolecular target to a second set of coordinates at which the candidate ligand is not bound to the biomolecularAtty Ref.: AVX0003PCT target; calculate, based on a result of simulating the unbinding event, temporal information associated with the candidate ligand traveling from the first set of coordinates to the second set of coordinates; calculate a dissociation free energy associated with the candidate ligand and the biomolecular target based on at least one of: the temporal information, a distance between the first set of coordinates and the second set of coordinates, or a combination thereof; and provide the dissociation free energy as a prediction of empirically derived binding free energy associated with the candidate ligand.

[0025] In any one or combination of the embodiments disclosed herein, the instructions, when executed by the processor, further cause the processor to perform operations including: screening the candidate ligand in association with one or more pharmaceutical applications based on the prediction of the empirically derived binding free energy.

[0026] In any one or combination of the embodiments disclosed herein, the instructions causing the processor to simulate the unbinding event cause the processor to accelerate the unbinding event by at least one of: performing a scaled molecular dynamics run; and performing a molecular dynamics run according to a temperature range associated with accelerating the unbinding event, wherein the temporal information and the distance are based on a result of accelerating the unbinding event.

[0027] In any one or combination of the embodiments disclosed herein, the instructions, when executed by the processor, further cause the processor to perform operations including: simulate the candidate ligand in a solution; calculate, based on a result of simulating the candidate ligand in the solution, a free energy associated with the candidate ligand in the solution; and calculate a binding free energy associated with the candidate ligand based on a difference between the dissociation free energy and the free energy associated with the candidate ligand in the solution.

[0028] In some aspects, the techniques described herein relate to a method of screening a candidate ligand including: simulating, by one or more processors, an unbinding event in which the candidate ligand dissociates from a biomolecular target to which the candidate ligand is bound, wherein the candidate ligand dissociating from the biomolecular target includes the candidate ligand traveling from a first set of coordinates at which the candidate ligand is bound to the biomolecular target to a second set of coordinates at which the candidate ligand is not bound to the biomolecular target; calculating, by the one or more processors based on a result of simulating the unbinding event, temporal informationAtty Ref.: AVX0003PCT associated with the candidate ligand traveling from the first set of coordinates to the second set of coordinates; calculating, by the one or more processors, a dissociation free energy associated with the candidate ligand and the biomolecular target based on at least one of: the temporal information, a distance between the first set of coordinates and the second set of coordinates, or a combination thereof; providing the dissociation free energy as a prediction of empirically derived binding free energy associated with the candidate ligand; and screening the candidate ligand in association with one or more pharmaceutical applications based on the prediction of the empirically derived binding free energy associated with the candidate ligand.

[0029] In some aspects, the techniques described herein relate to a method of screening a candidate ligand including: simulating, by one or more processors, an unbinding event in which the candidate ligand dissociates from a biomolecular target to which the candidate ligand is bound, wherein the candidate ligand dissociating from the biomolecular target includes the candidate ligand traveling from a first set of coordinates at which the candidate ligand is bound to the biomolecular target to a second set of coordinates at which the candidate ligand is not bound to the biomolecular target; calculating, by the one or more processors based on a result of simulating the unbinding event, temporal information associated with the candidate ligand traveling from the first set of coordinates to the second set of coordinates; calculating, by the one or more processors, a dissociation free energy associated with the candidate ligand and the biomolecular target based on at least one of: the temporal information, a distance between the first set of coordinates and the second set of coordinates, or a combination thereof; simulating, by the one or more processors, the candidate ligand in a solution; calculating, by the one or more processors based on a result of simulating the candidate ligand in the solution, a free energy associated with the candidate ligand in the solution; calculating, by the one or more processors, a binding free energy associated with the candidate ligand based on a difference between the dissociation free energy and the free energy associated with the candidate ligand in the solution; and screening the candidate ligand in association with one or more pharmaceutical applications based on the binding free energy associated with the candidate ligand.

[0030] The preceding general areas of utility are given by way of example only and are not intended to be limiting on the scope of the present disclosure and appended claims. Additional objects and advantages associated with the compositions, methods, and processes of the present invention will be appreciated by one of ordinary skill in the art in light of theAtty Ref.: AVX0003PCT instant claims, description, and examples. For example, the various aspects and embodiments of the invention may be utilized in numerous combinations, all of which are expressly contemplated by the present description. These additional advantages objects and embodiments are expressly included within the scope of the present invention. The publications and other materials used herein to illuminate the background of the invention, and in particular cases, to provide additional details respecting the practice, are incorporated by reference. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The accompanying drawings, which are incorporated into and form a part of the specification, illustrate several embodiments of the present invention and, together with the description, serve to explain the principles of the invention. The drawings are only for the purpose of illustrating an embodiment of the invention and are not to be construed as limiting the invention. Further objects, features and advantages of the invention will become apparent from the following detailed description taken in conjunction with the accompanying figures showing illustrative embodiments of the invention.

[0032] FIG.1 illustrates an example of a system in accordance with aspects of the present disclosure.

[0033] FIG.2 illustrates an example flowchart of a method in accordance with one or more embodiments of the present disclosure.

[0034] FIG.3 illustrates an example flowchart of a method in accordance with one or more embodiments of the present disclosure.

[0035] FIGS.4A through 4C illustrate examples of a protein and ligand, bound and unbound simulations, and plots of free energy determined by the simulations in accordance with example aspects of the present disclosure.

[0036] FIG.5 illustrates an example plot representing kinetic on rate (kon), kinetic off rate (koff), bound energy, and unbound energy determined in accordance with example aspects of the present disclosure.

[0037] FIGS.6A, 6B, 6C, and 6D illustrate example plots in accordance with the free energy prediction techniques described herein compared to other techniques.Atty Ref.: AVX0003PCT

[0038] FIG.7 illustrates an example flowchart of a method in accordance with one or more embodiments of the present disclosure.

[0039] FIG.8 illustrates example plots in accordance with example aspects of the present disclosure.

[0040] FIG.9 illustrates example plots in accordance with example aspects of the present disclosure.

[0041] FIGS.10A, 10B, 10C, 10D, 10E and 10F illustrate example plots in accordance with example aspects of the present disclosure.

[0042] FIG.11 illustrates example plots in accordance with example aspects of the present disclosure.

[0043] FIG.12 illustrates example plots in accordance with example aspects of the present disclosure.

[0044] FIG.13 illustrates an example graph in accordance with example aspects of the present disclosure.

[0045] FIG.14 illustrates an example plot in accordance with example aspects of the present disclosure.

[0046] FIG.15 illustrates an example iteration process in accordance with one or more embodiments of the present disclosure.

[0047] FIG.16 illustrates a table of data indicating that the free energy prediction techniques described herein can be used to predict the best pose from a list of poses generated by docking software or other means.

[0048] FIG.17 illustrates example molecules and respective solubilities determined using a comparative free energy perturbation approach described in Hong, R.S., Rojas, A.V., Bhardwaj, R.M., Wang, L., Mattei, A., Abraham, N.S., Cusack, K.P., Pierce, M.O., Mondal, S., Mehio, N. and Bordawekar, S., 2023.

[0049] FIG 18 illustrates example plots solubility determined using the molecular simulation methods described herein as applied to the molecules indicated at FIG.17.Atty Ref.: AVX0003PCT DETAILED DESCRIPTION

[0050] A detailed description of one or more embodiments supported by aspects of the present disclosure are presented herein by way of exemplification and not limitation with reference to the Figures.

[0051] The present description provides novel and optimized methods of koffprediction supported by validation with a robust data set. In certain embodiments, the approach employs high-temperature molecular dynamics (MD) simulations to drive ligand dissociation that would otherwise occur on timescales (ms–hr) far beyond those accessible with equilibrium MD. Additionally, the prediction workflow is significantly more efficient than current free energy-based protocols, making the described methods suitable for high throughput hit finding while being accurate enough for hit-to-lead and lead optimization. Unlike most practical implementations of free energy perturbation, the methods as described herein are an absolute koff predictor that does not require a reference compound with known koff. The necessary inputs and setup are relatively simple, requiring only the three- dimensional coordinates of the biological target and a reasonable model for the ligand binding mode, both of which are basic prerequisites for any structure-based drug discovery campaign.

[0052] While methods as described herein can predict the rank-ordering of binding half- life, additional embodiments allow the direct prediction of the free energy of binding by modeling a two-state system. As discussed above, the described methods are advantageous because they are not based on relative free energies or constrained by the use of structural similarity as is true of conventional methods. Additionally, the methods as described herein performs greater than 100 times faster than most available MD simulation-based free energy or koffmethods, allowing for widespread use by the molecular modeling community.

[0053] Fig.1 illustrates an example of a system 100 in accordance with aspects of the present disclosure. The system 100 may include a device 105, a server 110, a database 115, a communication network 120, and a simulation environment 170. The devices 105, the server 110, the database 115, the communications network 120, and the simulation environment 170 may implement aspects of the present disclosure described herein. The device 105 may implement aspects of the techniques and features described herein.

[0054] In various aspects, settings of any of the device 105, the server 110, database 115, and the network 120 may be configured and modified by any user and / or administrator of theAtty Ref.: AVX0003PCT system 100. Settings may include thresholds or parameters described herein, as well as settings related to how data is managed. Settings may be configured to be personalized for one or more devices 105, users of the devices 105, and / or other groups of entities, and may be referred to herein as profile settings, user settings, or organization settings. In some aspects, rules and settings may be used in addition to, or instead of, parameters or thresholds described herein. In some examples, the rules and / or settings may be personalized by a user and / or administrator for any variable, threshold, user (user profile), device 105, entity (e.g., patient), or groups thereof.

[0055] A device 105 may include a processor 130, a network interface 135, a memory 140, and a user interface 145. In some examples, components of the device 105 (e.g., processor 130, network interface 135, memory 140, user interface 145) may communicate over a system bus (e.g., control busses, address busses, data busses) included in the device 105. In some cases, the device 105 may be referred to as a computing resource.

[0056] In some cases, the device 105 may transmit or receive packets to one or more other devices (e.g., another device 105, the server 110, the database 115, the simulation environment 170) via the communication network 120, using the network interface 135. The network interface 135 may include, for example, any combination of network interface cards (NICs), network ports, associated drivers, or the like. Communications between components (e.g., processor 130, memory 140) of the device 105 and one or more other devices (e.g., another device 105, the database 115) connected to the communication network 120 may, for example, flow through the network interface 135.

[0057] The processor 130 may correspond to one or many computer processing devices. For example, the processor 130 may include a silicon chip, such as a FPGA, an ASIC, any other type of IC chip, a collection of IC chips, or the like. In some aspects, the processors may include a microprocessor, CPU, a GPU, or plurality of microprocessors configured to execute the instructions sets stored in a corresponding memory (e.g., memory 140 of the device 105). For example, upon executing the instruction sets stored in memory 140, the processor 130 may enable or perform one or more functions of the device 105.

[0058] The processor 130 may utilize data stored in the memory 140 as a neural network (also referred to herein as a machine learning network). The neural network may include a machine learning architecture. In some aspects, the neural network may be or include an artificial neural network (ANN). In some other aspects, the neural network may be or includeAtty Ref.: AVX0003PCT any machine learning network such as, for example, a deep learning network, a convolutional neural network, or the like. Some elements stored in memory 140 may be described as or referred to as instructions or instruction sets, and some functions of the device 105 may be implemented using machine learning techniques.

[0059] The memory 140 may include one or multiple computer memory devices. The memory 140 may include, for example, Random Access Memory (RAM) devices, Read Only Memory (ROM) devices, flash memory devices, magnetic disk storage media, optical storage media, solid state storage devices, core memory, buffer memory devices, combinations thereof, and the like. The memory 140, in some examples, may correspond to a computer readable storage media. In some aspects, the memory 140 may be internal or external to the device 105.

[0060] The memory 140 may be configured to store instruction sets, neural networks, and other data structures (e.g., depicted herein) in addition to temporarily storing data for the processor 130 to execute various types of routines or functions. For example, the memory 140 may be configured to store program instructions (instruction sets) that are executable by the processor 130 and provide functionality of a machine learning engine 141 and a prediction engine 147 described herein. The memory 140 may also be configured to store data or information that is useable or capable of being called by the instructions stored in memory 140. One example of data that may be stored in memory 140 for use by components thereof is a data model(s) 142 (also referred to herein as a neural network model), training data 143 (also referred to herein as a training data and feedback), statistical models 146, and / or other prediction models 148.

[0061] The machine learning engine 141 may include a single or multiple engines. The device 105 (e.g., the machine learning engine 141) may utilize one or more data models 142 for recognizing and processing information obtained from other devices 105, the server 110, the database 115, and the simulation environment 170. In some aspects, the device 105 (e.g., the machine learning engine 141) may update one or more data models 142 based on learned information included in the training data 143. In some aspects, the machine learning engine 141 and the data models 142 may support forward learning based on the training data 143. The machine learning engine 141 and data models 142 may support reinforcement learning (e.g., deep reinforcement learning) and imitation learning described herein. The machine learning engine 141 may have access to and use one or more data models 142. For example,Atty Ref.: AVX0003PCT the data model(s) 142 may be built and updated by the machine learning engine 141 based on the training data 143. The data model(s) 142 may be provided in any number of formats or forms. Non limiting examples of the data model(s) 142 include Decision Trees, Support Vector Machines (SVMs), Nearest Neighbor, and / or Bayesian classifiers.

[0062] The engines described herein (e.g., machine learning engine 141, prediction engine 147) may create, select, and execute processing decisions as described herein. Processing decisions may be handled automatically by the engines (e.g., machine learning engine 141, prediction engine 147), with or without human input.

[0063] The engines (e.g., machine learning engine 141, prediction engine 147) may store, in the memory 140 (e.g., in a database included in the memory 140), historical information. Data within the database of the memory 140 may be updated, revised, edited, or deleted by the engines described herein. In some aspects, the engines described herein may support continuous, periodic, and / or batch fetching of data and data aggregation.

[0064] The device 105 may render a presentation (e.g., visually, audibly, using haptic feedback, etc.) of an application 144 (e.g., a browser application 144-a, an application 144-b). The application 144-b may be an application associated with executing, controlling, and / or monitoring the simulation environment 170 described herein. For example, the application 144-b may enable control of the device 105 or the simulation environment 170.

[0065] In an example, the device 105 may render the presentation via the user interface 145. The user interface 145 may include, for example, a display (e.g., a touchscreen display), an audio output device (e.g., a speaker, a headphone connector), or any combination thereof. In some aspects, the applications 144 may be stored on the memory 140. In some cases, the applications 144 may include cloud based applications or server based applications (e.g., supported and / or hosted by the database 115 or the server 110). Settings of the user interface 145 may be partially or entirely customizable and may be managed by one or more users, by automatic processing, and / or by artificial intelligence.

[0066] In an example, any of the applications 144 (e.g., browser application 144-a, application 144-b) may be configured to receive data in an electronic format and present content of data via the user interface 145. For example, the applications 144 may receive data from another device 105, the server 110, or the simulation environment 170 via the communications network 120, and the device 105 may display the content via the user interface 145.Atty Ref.: AVX0003PCT

[0067] The database 115 may include a relational database, a centralized database, a distributed database, an operational database, a hierarchical database, a network database, an object oriented database, a graph database, a NoSQL (non-relational) database, etc. In some aspects, the database 115 may store and provide access to, for example, any of the stored data described herein.

[0068] The server 110 may include a processor 150, a network interface 155, database interface instructions 160, and a memory 165. In some examples, components of the server 110 (e.g., processor 150, network interface 155, database interface 160, memory 165) may communicate over a system bus (e.g., control busses, address busses, data busses) included in the server 110. The processor 150, network interface 155, and memory 165 of the server 110 may include examples of aspects of the processor 130, network interface 135, and memory 140 of the device 105 described herein.

[0069] For example, the processor 150 may be configured to execute instruction sets stored in memory 165, upon which the processor 150 may enable or perform one or more functions of the server 110. In some aspects, the processor 150 may utilize data stored in the memory 165 as a neural network. In some examples, the server 110 may transmit or receive packets to one or more other devices (e.g., a device 105, the database 115, another server 110) via the communication network 120, using the network interface 155. Communications between components (e.g., processor 150, memory 165) of the server 110 and one or more other devices (e.g., a device 105, the database 115, the simulation environment 170, etc.) connected to the communication network 120 may, for example, flow through the network interface 155.

[0070] In some examples, the database interface instructions 160 (also referred to herein as database interface 160), when executed by the processor 150, may enable the server 110 to send data to and receive data from the database 115. For example, the database interface instructions 160, when executed by the processor 150, may enable the server 110 to generate database queries, provide one or more interfaces for system administrators to define database queries, transmit database queries to one or more databases (e.g., database 115), receive responses to database queries, access data associated with the database queries, and format responses received from the databases for processing by other components of the server 110.

[0071] The memory 165 may be configured to store instruction sets, neural networks, and other data structures (e.g., depicted herein) in addition to temporarily storing data for theAtty Ref.: AVX0003PCT processor 150 to execute various types of routines or functions. For example, the memory 165 may be configured to store program instructions (instruction sets) that are executable by the processor 150 and provide functionality of the machine learning engine 166 and a prediction engine 169 described herein. One example of data that may be stored in memory 165 for use by components thereof is a data model(s) 167 (also referred to herein as a neural network model) and / or training data 168. The data model(s) 167 and the training data 168 may include examples of aspects of the data model(s) 142 and the training data 143 described with reference to the device 105. For example, the server 110 (e.g., the machine learning engine 166) may utilize one or more data models 163 for recognizing and processing information obtained from devices 105, another server 110, the database 115, or the simulation environment 170. In some aspects, the server 110 (e.g., the machine learning engine 166) may update one or more data models 163 based on learned information included in the training data 168.

[0072] In some aspects, components of the machine learning engine 166 may be provided in a separate machine learning engine in communication with the server 110.

[0073] The prediction engine 147 and prediction engine 169 may support free energy prediction techniques described herein. The prediction engine 147 and prediction engine 169 may support techniques for identifying compounds as described herein. Example aspects of the techniques performable by the prediction engine 147 and the prediction engine 169 are described herein with reference to the following figures. The prediction engine 147 and the prediction engine 169 may be implemented using one or more models (e.g., model(s) 142, statistical models 146, prediction models 148, and the like) described herein.

[0074] According to one or more embodiments of the present disclosure, the system 100 supports molecular dynamics simulation-based koffpredictions. The free energy prediction techniques described herein demonstrate similar accuracy to some other free energy prediction methods, with advantages of performance speeds which are about 100 times faster than other molecular dynamics simulation-based free energy or koff methods, thus supporting widespread use by the molecular modeling community. While some other free energy methods rely on relative free energy changes and are primarily useful for optimization of a congeneric series, the free energy prediction techniques are independent of requirements for structural similarity between ligands as in other free energy methods. Accordingly, forAtty Ref.: AVX0003PCT example, the free energy prediction techniques described herein may serve as an absolute predictor of koff.

[0075] The system 100 provides a free energy prediction tool (e.g., implemented at device 105 and / or server 110) which can be used in virtual screening of diverse ligands, making the free energy prediction distinct from relative free energy methods.

[0076] As will be described herein, the system 100 supports features related to molecular dynamics simulation, drug discovery, and computational chemistry. The techniques described herein support the rank ordering and scoring of ligands based on either the binding half-life or the free energy of binding of the ligands. For example, both the binding half-life or the free energy of binding are important properties of ligands that may affect biological targets such as protein enzymes, protein-protein interactions (PPIs), nucleic acids, and other biomolecular targets. The system 100 (e.g., at simulation environment 170) provides a high- performance computing environment configured to simulate the binding of a ligand to a protein with high accuracy. In some aspects, the system 100 may be implemented in combination with molecular dynamics engines such as, for example, openMM, GROMACS, AMBER, Desmond, Tinker, NAMD, CHARMM, or the like.

[0077] In accordance with one or more embodiments of the present disclosure, the system 100 supports techniques for predicting ligand off-rates (e.g., kinetic off rates, koff), e.g., the absolute koff, and further, using the predicting ligand off-rates for the prediction of free energy, e.g., absolute free energy of binding, ranking by free energy, and / or ranking by empirical half maximal inhibitory concentration (IC50), inhibitory constant (Ki), or related measures of ligand efficacy.

[0078] The free energy prediction techniques supported by the system 100 may provide accurate predictions of ligand affinity and binding half-life. In some embodiments, the free energy prediction techniques supported by the system 100 apply a modified population-based reweighting theory, and specifically, for example, may use a two state model to predict the absolute binding energy.

[0079] In some embodiments, a method of koffprediction is provided herein, supported by validation with a robust data set. In some aspects, the approach may employ high-temperature molecular dynamics simulations to drive ligand dissociation that would otherwise occur on timescales (e.g., milliseconds (ms)-hour (hr)) far beyond those accessible with equilibrium molecular dynamics. The prediction workflow is significantly more efficient than current freeAtty Ref.: AVX0003PCT energy-based protocols and, accordingly, for example, supports high throughput hit finding while maintaining an accuracy level supportive of hit-to-lead and lead optimization. In some aspects, unlike some other implementations of free energy perturbation, the prediction techniques provided by the system 100 as described herein serve as an absolute koffpredictor which may be implemented without a reference compound with known koff. For example, the free energy prediction techniques described herein may be implemented with a relatively reduced number of inputs compared to some other prediction techniques. For example, the free energy prediction described herein may be implemented based on inputs including the three-dimensional coordinates of the biological target and a reasonable model for the ligand binding mode, both of which are basic prerequisites for any structure-based drug discovery campaign.

[0080] According to one or more additional or alternative embodiments of the present disclosure, the system 100 supports predicting the absolute free energy of ligand binding to protein targets by using a two-state model (bound and unbound) based on population-based reweighting of accelerated molecular dynamics simulations.

[0081] Aspects of the system 100 support predicting the rank ordering of binding half- life, and further, directly predicting the free energy of binding by modeling a two-state system. The techniques described herein for free energy prediction using a two-state model provide accurate binding predictions, both in terms of the actual free energy of binding and the rank-ordering of a ligand series, at a fraction of the simulation time associated with other prediction methodologies. The techniques described herein for free energy prediction using a two-state model differ from other free energy methods based on relative free energy changes. The techniques described herein for free energy prediction using a two-state model may be implemented with an improved accuracy level and reduced simulation time compared to some other techniques for determining absolute binding free energy, as the other techniques may provide reasonable results but with long simulation times and / or provide inaccurate absolute binding affinity predictions that require a correction factor.

[0082] The system 100 supports rapid simulation-based prediction of ligand off-rates and rapid simulation-based prediction of ligand binding free energy, example aspects of which are described herein with reference to the following diagrams and flowcharts.

[0083] In the descriptions of the flowcharts herein, the operations may be performed in a different order than the order shown, or the operations may be performed in different ordersAtty Ref.: AVX0003PCT or at different times. Certain operations may also be left out of the flowcharts, one or more operations may be repeated, or other operations may be added to the flowcharts.

[0084] FIG.2 illustrates an example flowchart of a method 200 in accordance with one or more embodiments of the present disclosure. The method 200 may be implemented by the example aspects of system 100 (e.g., performed by a device 105 and / or server 110) as described herein.

[0085] The method 200 supports predicting ligand off-rate (koff) using molecular simulations, in which the simulations are accelerated by high temperature or scaled potential.

[0086] The method 200 supports specific use cases capable of predicting free energy of binding or relative free energy of binding. The method 200 supports dynamic molecular simulation-based prediction of ligand efficacy.

[0087] At 205, the method 200 may include preparing a biomolecular structure for molecular modeling. In an example, the biomolecular structure may include a ligand and a protein (e.g., a target protein for binding with the ligand).

[0088] At 210, the method 200 may include generating design ideas associated with a compound. In some examples, the design ideas may include one or more candidate ligands and one or more candidate biomolecular targets. In some examples, the design ideas may include, but are not limited to, hand drawn ideas, medicinal chemist ideas, enumerated compound libraries, purchasable compound libraries, computer generated compounds, AI designed compounds, combinatorial chemistry compound libraries, fragment growing, fragment linking, recombination of known compounds, and R-group enumerations.

[0089] At 215, the method 200 may include positioning one or more compounds in a multidimensional space (e.g., a 3D space) of a binding pocket. The term binding pocket may refer to a cavity on the surface or in the interior of a protein that possesses suitable properties for binding a ligand. In an example, at 215, the method 200 may include positioning the ligand with respect to a binding pocket of the protein. In some aspects, at 215, the method 200 may include storing pose information (e.g., position information, orientation information, coordinates) of the ligand and / or the protein. In some examples, the pose information associated with 215 of the method 200 may be referred to as a first set of coordinates at which the ligand is bound to the protein.Atty Ref.: AVX0003PCT

[0090] At 220-a, the method 200 may include running a molecular simulation with acceleration from scaled potential. Additionally, or alternatively, at 220-a, the method 200 may include running the molecular simulation with acceleration from temperature. The simulation at 220-a may be a bound simulation as described herein. In some aspects, the bound simulation is an accelerated ligand unbinding event described herein.

[0091] At 220-b, the method 200 may include running a molecular simulation of ligand in solution. The simulation at 220-b may be an unbound simulation as described herein.

[0092] In some aspects in association with running the molecular simulations (e.g., at 220-a and 220-b), the method 200 may include tracking pose information of the ligand and / or the protein.

[0093] At 225, the method 200 may include calculating relative time to ligand dissociation from the bound state. For example, at 225, the method 200 may include calculating (e.g., from the start of the simulation of 220-a) the amount of time for a ligand to dissociate or separate from a protein (or other biomolecular target) to which the ligand is bound. In some embodiments, the method 200 may include calculating the dissociation rate in which the ligand dissociates or separates from the protein (or other biomolecular target). The amount of time to dissociate and the dissociation rate may be referred to as temporal information. In some examples, at 225, the method 200 may include calculating pose information associated with the disassociation or separation of the ligand from the protein (or other biomolecular target). In some aspects, the pose information calculated at 225 of the method 200 may be referred to as a second set of coordinates at which the ligand is not bound to the protein.

[0094] At 230-a, the method 200 may include calculating the dissociation free energy (DFE) for the ligand (also referred to herein as ligand dissociation free energy). In an example, the dissociation free energy quantifies the work associated with dissociating a complex (e.g., dissociating the ligand from the biomolecular target).

[0095] At 230-b, the method 200 may include calculating ligand in solution free energy. The ligand in solution free energy corresponds to the free energy of a ligand without its biomolecular target bound. The ligand in solution free energy may be related to the rate of diffusion of the ligand in solution.Atty Ref.: AVX0003PCT

[0096] At 235, the method 200 may include calculating ligand binding free energy (also referred to herein as “binding free energy” or “free energy difference of binding”). As described herein, the term ligand binding free energy may refer to the free energy difference between the bound state and the unbound state (free state) associated with the ligand. That is, for example, ligand binding free energy may refer to free energy of binding associated with the ligand and biomolecular targets associated with the ligand.

[0097] In some aspects, at 235, the method 200 may include calculating ligand binding free energy in accordance with Equations (5) through (8) later described herein.

[0098] At 240, the method 200 may include predicting ligand binding free energy. In some aspects, at 240, the method 200 may include evaluating the dissociation free energy of the ligand moving from the starting position or binding site as the prediction of the ligand binding free energy.

[0099] In some embodiments, at 240, the method 200 may include predicting the ligand binding free energy from the dissociation free energy (calculated at 230-a) minus the ligand free energy (calculated at 230-b). Some usual corrections could be applied such as normalizing the means or ranges to predict the ligand binding free energy of experiment.

[0100] In some other embodiments, at 240, the method 200 may include predicting the ligand binding free energy from the dissociation free energy (calculated at 230-a) directly, for example, in exchange for reduced accuracy in absolute terms between the dissociation free energy without using the ligand free energy (calculated at 230-b). That is, for example, the method 200 may include ignoring the ligand binding free energy.

[0101] At 245, the method 200 may include comparing the predicted ligand binding free energy to empirically determined ligand binding free energy (of 235). Expressed another way, at 245, the method 200 may include comparing the predicted free energy of binding associated with the ligand (as determined at 240) to an empirically determined free energy of binding associated with the ligand (as determined at 235).

[0102] Additionally, or alternatively, at 245, the method 200 may include comparing one or more other predicted properties of the ligand to one or more corresponding empirically determined properties of the ligand. In some aspects, the method 200 may include generating the one or more other predicted properties of the ligand at 240, and the method 200 may include generating the one or more empirically determined properties of the ligand at 235.Atty Ref.: AVX0003PCT

[0103] For example, the method 200 may include comparing a predicted IC50 associated with the ligand to an empirically determined IC50 associated with the ligand. IC50 is a measure of the potency of a substance in inhibiting a specific biological or biochemical function. IC50 is a quantitative measure which indicates the amount of a particular inhibitory substance which inhibits, in vitro, a given biological process or biological component by 50%.

[0104] In another example, the method 200 may include comparing a predicted Kiassociated with the ligand to an empirically determined Ki associated with the ligand. The Ki is a type of equilibrium dissociation constant (Kd) representing the equilibrium binding affinity for a ligand which reduces the activity of the binding partner of the ligand. Ki represents the concentration at which an inhibitor ligand occupies 50% of the receptor sites when no competing ligand is present.

[0105] As described herein with reference to FIG.2, aspects of the present disclosure support a molecular simulation method (e.g., method 200) which incorporates one or more simulation techniques in association with speeding up dynamic events of interest (e.g., unbinding of a ligand from a protein). Non-limiting examples of the simulation techniques supportive of speeding up the dynamic events of interest include 1) altered, high temperature molecular dynamics simulation, 2) scaled potential molecular dynamics simulation, and 3) other methods or simulation techniques that alter the Hamiltonian or potential energy surface to increase the speed of dynamic events. Aspects of the method 200 may include calculating a relative ranking of a ligand based on the length of the simulation time (also referred to herein as simulation duration).

[0106] As described herein with reference to FIG.2 (e.g., at 220-a), aspects of the present disclosure include running a molecular dynamics simulation with a ligand bound to a biological target of interest (e.g., a protein, an enzyme). Aspects of the present disclosure include accelerating the simulation by scaling the potential energy surface to ensure that the ligand will unbind during the simulation. In some aspects, the techniques described herein include accelerating the simulation because, for example, conventional molecular dynamics simulations are unable to provide simulation speeds fast enough to model unbinding of ligands with drug-like potency. Accordingly, for example, the techniques described herein provide improvements supportive of model unbinding of ligands with drug-like potency.Atty Ref.: AVX0003PCT

[0107] Aspects of the present disclosure include calculating a relative ranking of ligands. For example, the method 200 may include using the relative ranking in association with predicting (e.g., at 240) the potency of a ligand given other known references or for rank- ordering ligands from a virtual screen or medicinal chemistry design.

[0108] The use of scaled molecular dynamics and or similar koffpredictors in association with directly predicting free energy of binding (as described herein) is non-obvious because previous literature has characterized the relationship between koffand free energy as poorly correlated. In accordance with one or more embodiments of the present disclosure, the techniques described herein are based on an analysis of existing data showing that a correlation between free energy and unbinding kinetic rates does exist, except in rare instances where the protein has significantly changed (e.g., more than a threshold amount) conformation. The techniques described herein are based on a demonstrated reasonable correlation between koff and free energy, in which the accuracy of the correlation is sufficient to enhance the accuracy of virtual screening. Other approaches have failed to describe or implement the use of the koff predictor to predict the free energy of binding.

[0109] FIG.3 illustrates an example flowchart of a method 300 in accordance with one or more embodiments of the present disclosure. The method 300 may be implemented by the example aspects of system 100 as described herein. The method 300 illustrates an example of a general workflow for free energy prediction using a two-state model in accordance with one or more embodiments of the present disclosure. The method 300 includes aspects of method 200 described herein, and repeated descriptions of like elements are omitted for brevity.

[0110] Aspects of the method 300 are described herein with reference to FIGS.4A through 4C. FIGS.4A through 4C illustrate examples of a protein and ligand, bound and unbound simulations, and plots of free energy determined by the simulations in accordance with example aspects of the present disclosure and the method 300.

[0111] The method 300 supports the prediction of ligand binding free energy using molecular simulations that are accelerated by high temperature (e.g., above a threshold temperature value). In some aspects, the method 300 includes deriving the absolute free energy of binding by using a two-state system (e.g., ligand with a biomolecular target and ligand without a biomolecular target).

[0112] At 305, the method 300 may include identifying a ligand (e.g., a drug candidate ligand) and a pose (also referred to herein as pose information) associated with the ligand andAtty Ref.: AVX0003PCT a biomolecular target. In the examples described herein, the biomolecular target is described as a protein, but aspects of the present disclosure are not limited thereto. For example, the biomolecular target may be another biological target. In some aspects, the method 300 may include identifying the ligand and the pose via x-ray crystallography, docking, flexible docking, molecular dynamics, overlay, or any suitable modeling means supportive of the method 300. The pose (pose information) may be, for example, an orientation of the ligand and the biomolecular target according to a state in which the ligand and the biomolecular target are bound to each other and form a stable complex. In some examples, the pose information associated with 305 of the method 300 may be referred to as a first set of coordinates at which the ligand is bound to the biomolecular target. Accordingly, for example, at 305, the method 300 may include identifying a predicted 3D pose (also referred to herein as a predicted binding pose) before performing any simulation. That is, for example, the method 300 may include identifying the predicted 3D pose (predicted binding pose), before performing the bound simulation in 320-a.

[0113] At 320-a, the method 300 may include performing a bound simulation in which the ligand is bound to the biomolecular target. In some examples, the bound simulation at 320-a may include aspects of the simulation described with reference to 220-a. In some aspects, the bound simulation is an accelerated ligand unbinding event described herein. In some examples, at 320-a, the method 300 may include calculating pose information associated with the disassociation or separation (i.e., unbinding) of the ligand from the biomolecular target. In some aspects, the pose information calculated at 320-a of the method 300 may be referred to as a second set of coordinates at which the ligand is not bound to the biomolecular target.

[0114] At 320-b, the method 300 may include performing an unbound simulation in which the ligand is in a solution (e.g., water, a bulk solvent) and is not bound to a biomolecular target. In some examples, the unbound simulation at 320-b may include aspects of the simulation described with reference to 220-b.

[0115] In some embodiments, the method 300 may include performing the bound simulation (of 320-a) substantially simultaneous to performing the unbound simulation (of 320-b). Additionally, or alternatively, the method 300 may include performing the bound simulation at a different time point (or time duration) compared to performing the unbound simulation.Atty Ref.: AVX0003PCT

[0116] At 325, the method 300 may include counting the populations associated with each of the bound simulation and the unbound simulation along a reaction coordinate that defines unbinding from the protein.

[0117] With reference to FIG.4C, an example plot 405 of free energy 407 associated with the bound simulation and an example plot 410 of free energy 412 associated with the unbound simulation are illustrated. In the examples, the Y-axis corresponds to free energy, and the X-axis corresponds to reaction coordinate. With reference to the example plot 405 and the bound simulation, the free energy is about 11 kcal / mol at the reaction coordinate that defines unbinding of the ligand from the protein. With reference to the example plot 410 and the unbound simulation, the free energy is about 2 kcal / mol at the same reaction coordinate.

[0118] In some aspects, at 325, the method 300 may include counting populations in accordance with Equation (1) later described herein.

[0119] At 330, the method 300 may include reweighting the populations along the reaction coordinate. In some examples, the reweighting may be according to scaling factor lambda (^) and / or temperature T as described herein. In some aspects, at 330, the method 300 may include reweighting in accordance with Equation (3) and Equation (4) later described herein.

[0120] At 335, the method 300 may include calculating the free energy associated with the bound simulation and the free energy associated with the unbound simulation, based on the reweighting. In some examples, the calculation of free energy at 335 may include aspects of the calculating of free energy described with reference to 235.

[0121] In some aspects, at 335, the method 300 may include calculating ligand binding free energy in accordance with Equation (3) later described herein.

[0122] At 340, the method 300 may include calculating free energy difference. For example, at 340, the method 300 may include calculating the difference between the free energy associated with the bound simulation (e.g., 11 kcal / mol at the reaction coordinate that defines unbinding of the ligand from the protein, as described with reference to plot 405) and the free energy associated with the unbound simulation (e.g., 2 kcal / mol at the same reaction coordinate, as described with reference to plot 410) and using Equations (5) through (8).Atty Ref.: AVX0003PCT

[0123] In some aspects, the method 300 may further include predicting ligand binding free energy, which may include evaluating the dissociation free energy of the ligand moving from the starting position or binding site as the prediction of the ligand binding free energy.

[0124] Aspects of the present disclosure include extending molecular simulation methods described herein by creating a two-state system. Other approaches have failed to describe or implement a two-state system (e.g., ligand in solution and ligand-target complex) using population reweighting theory and acceleration via scaled potential. The techniques described herein employ an accelerated molecular dynamics simulation (e.g., as described with reference to 220-a and 320-a) to remove the ligand from the biological target binding site and a second molecular dynamics simulation (e.g., as described with reference to 220-b and 320- b) of the ligand in solution. The techniques described herein include accelerating simulations of the protein-ligand complex by a scaled potential.

[0125] In some embodiments, aspects of the present disclosure include applying the molecular simulation methods described herein in a single system. For example, aspects of the present disclosure support placing the ligand in the bound protein and in solution, in a single system instead of two discrete systems, as described by Azimi, S., et al., Relative Binding Free Energy Calculations for Ligands with Diverse Scaffolds with the Alchemical Transfer Method, Journal of Chemical Information and Modeling 202262 (2), 309-323.

[0126] The techniques described herein (e.g., using Equations (7) or (8)) support correcting the populations to reflect the probability distributions at any temperature of interest and calculating the free energy difference between the bound and unbound state (i.e., the binding free energy). The binding free energy is a high value parameter in drug discovery research and directly relates to biological efficacy. The predicted free energy determined using the techniques described herein may be advantageous in determining (e.g., by researchers) how to prioritize molecules in a medicinal chemistry campaign. In some cases, the predicted free energy may be advantageous for rank-ordering possible ligands such as, for example, from a virtual screening or medicinal chemistry design. Using a two-state system with accelerated off-rate predictors to generate the free energy of binding has not been previously described. The techniques described herein support generating accurate results both in terms of absolute free energy and of ligand ranking.Atty Ref.: AVX0003PCT

[0127] Accordingly, for example, aspects of the present disclosure may include selecting a ligand for binding to a biomolecular target, based on ranking of the ligand, and incorporating the ligand into a drug to be administered.

[0128] FIG.5 illustrates an example plot 500 representing kinetic on rate (kon), kinetic off rate (koff), bound energy, and unbound energy determined in accordance with example aspects of the present disclosure. FIG.5 illustrates an example plot 505 representing kinetic on rate (kon), kinetic off rate (koff), bound energy, and unbound energy, reweighted according to scaling factor lambda (^) of 0.25, 0.5, and 0.75 in accordance with example aspects of the present disclosure, for example, as described with reference to method 300.

[0129] The techniques described herein incorporate a two-state problem supportive of reduced computation durations, more effective computation (e.g., increased efficiency, reduced overhead) compared to other approaches which include computing the entire potential of mean force (PMF) reaction coordinate.

[0130] For the bound state, the techniques described herein include running a relatively short simulation (e.g., less than a threshold temporal duration) of a ligand bound to a biomolecular target (e.g., a protein) and determining, based on the simulation results, initial unbinding of the ligand from the biomolecular target.

[0131] For the unbound state, the techniques described herein include running a relatively short simulation (e.g., less than a threshold temporal duration) of the ligand in a solution (e.g., water, a bulk solvent) and measuring the same reaction coordinate. In some aspects, the difference in lambda (^) reweighted populations between the bound state and the unbound state is directly related to the binding free energy.

[0132] The free energy prediction techniques described herein incorporate a method for accelerated sampling and reweighting of a conformational landscape for a given molecular simulation. The method for accelerated sampling and reweighting has advantages in that the reweighting scheme is independent of the fluctuating energy of the system but based on the conformational state of the system, which is more interpretable and scalable to larger systems.

[0133] In some aspects, the free energy prediction techniques described herein are based on a population-based reweighting scheme and framework that provides a further framework to calculate koff, which, as further developed in accordance with one or more embodiments ofAtty Ref.: AVX0003PCT the present disclosure, significantly improves (e.g., by a threshold value, percentage, or the like) the speed and accuracy of the involved protocols.

[0134] For example, the systems and techniques described herein support running molecular dynamics simulations on a scaled potential energy surface, which is capable of accelerating the rate and scope of sampling conformational space and altering the free energy landscape. Accordingly, for example, the techniques described herein may include running molecular dynamics simulations in association with altering the correlation between the real distribution of sampled conformations and the partition function. Accordingly, for example, the population P at equilibrium for a configuration (or microstate), (^⃗), is based on the potential energy (V) at that point, as shown by Equation (1).^^^^⃗ = ^^^^^^⃗^ (1)

[0001] Similarly, for example, in the case of a simulation conducted on a scaled temperature potential, the P*(^⃗) of a scaled molecular dynamics run may be defined by Equation (2),^∗^^^⃗ = ^^^^∗^^⃗^ = ^^^^^^^⃗^ (2)= λ as a user-defined scaling factor. In some aspects, when λ is applied to a given system, an altered population of microstates is produced. In some non-limiting aspects, the systems and techniques described herein may include applying λ to the temperature or the potential energy surface. With reference to Equation (1) and Equation (2), a relation is seen between the population P*(^⃗) of the scaled molecular dynamics run and the population P(^⃗), corresponding to a Boltzmann distribution along the original potential, as shown by Equation (3).^^^⃗^ = ^∗^^⃗^^ / ^ (3)

[0136] The techniques described herein may include running a molecular simulation with acceleration from temperature and / or a scaled potential. In some aspects, by running simulations with a scaled potential, sampling is significantly enhanced. The techniques described herein may include applying the relationship in Equation (3) in association withAtty Ref.: AVX0003PCT recovering the Boltzmann distribution and energies of any microstate (^⃗) including the bioactive (or presumed bioactive) conformation.

[0137] The techniques described herein may include expanding Equation (3) in association with predicting dissociation rates of protein-ligand systems. For example, Equation (4) supports predicting the koff of ligands relative to the koff of other ligands. !"",$% ^^ !"",&% = ^^^'(⁄ )* = +^^'(⁄ )* , = - !"",$!"",&.(4)implementing Equation (4) with high temperatures T (e.g., above a threshold value). In some aspects, the techniques described herein may support conducting simulations on a scaled temperature potential, P*(^⃗). For example, the techniques described herein may use temperature as the scaling factor, and further, using a reweighting scheme described herein (e.g., population-based reweighting scheme) in association with recovering the correct distribution at the unscaled temperature.

[0139] The free energy prediction techniques described herein support rank-ordering of ligand kinetic off-rates and rank-ordering of ligand affinity, and in many cases, correlation between the ligand kinetic off-rates and the ligand affinity satisfies a target threshold value. In accordance with one or more embodiments of the present disclosure, the techniques described herein include predicting absolute values of the free energy of ligand binding, and further, directly relating the predicted absolute values to equilibrium dissociation constant (Kd) or affinity.

[0140] As described herein, embodiments of the present disclosure support the prediction of absolute values and directly relating the predicted absolute values to equilibrium dissociation constant (Kd) or affinity by running two states of simulations: 1) a simulation with a ligand bound to a protein and 2) a simulation with the ligand in a solution (e.g., water, a bulk solvent). The techniques described herein include determining, based on results of the two simulations, the effect that protein binding has on motion of the ligand away from the bound state. That is, for example, the techniques described herein support computational methods for predicting interactions between a ligand and a protein (or other biomolecular target).Atty Ref.: AVX0003PCT

[0141] In an example, Equation (3) supports techniques described herein for relating free energy differences along any coordinate by counting the populations and then scaling, by a factor related to lambda (^), the scaling factor of the potential energy surface or temperature. Aspects of the present disclosure include further relating the scaling factor to multiple states of a system.^^^^⃗ = ^∗^^⃗^^ / ^ (3)∆^^ = −^^^^^ / ^^ (5)

[0143] Kdis given by the concentration of the protein and ligand over the concentration of the protein ligand complex, as shown by Equation (6). / = 012032^ 0132 (6), , , substituting a two-state system (e.g., a system with the ligand bound, and a system with the ligand unbound), using a scaling factor lambda (^) to accelerate the simulation, as shown by Equation (7). ∗^ $⁄ %∆^ = −^^l ^ ^ 1 ^⃗9:;!9:<^^ n K7 = −^^^^ 8 %= (7)some same energy of the bound and unbound state, as shown by Equation (8), which is a reformulation of Equation (7).∆^ = ∆^ ^^^ − ∆^^^^^^^^ = −^^^^^ ∗^ ^^^^^^^⁄ ^ ∗^ ^^^^^^^^^⁄ ^^ ^^ ^ ^⃗ ^ + ^^^^^^ ^⃗ ^ (8)

[0146] Using the relationship in Equation (7) (and similarly, in Equation (8)), the techniques described herein support accurately calculating the free energy of binding at any scaled potential or any temperature. That is, for example, by performing two independent simulations which describe the populations of the endpoints (bound and unbound) along a collective variable (e.g., distance between the center of mass (COM) of a ligand and theAtty Ref.: AVX0003PCT target molecule or the root mean-squared deviation (RMSD) of a ligand from the starting position of the ligand), the techniques described herein support accurately calculating the free energy of binding at any scaled potential or any temperature. Root mean square deviation of atomic positions, or simply RMSD, is the measure of the average distance between the atoms (usually the backbone atoms) of superimposed molecules. In the study of globular protein conformations, some approaches measure the similarity in three-dimensional structure by the RMSD of the Cα atomic coordinates after optimal rigid body superposition.

[0147] FIG.6A illustrates an example plot 605 of predicted free energy perturbation (FEP) in accordance with FEP+ versus free energy prediction using a two-state model as described herein.

[0148] FIG.6B illustrates an example plot 610 of predicted FEP in accordance with other molecular dynamics engines (e.g., AMBER-TI) versus free energy prediction using a two- state model as described herein.

[0149] FIG.6C illustrates an example plot 615 of predicted FEP in accordance with the free energy prediction techniques described herein using a one-state model (i.e., model of binding kinetics for a ligand in a bound state to a target) of koff.

[0150] FIG.6D illustrates an example plot 620 of predicted off-rate versus free energy prediction using a two-state model as described herein.

[0151] The plots 605 through 620 include six FEP normalized targets (e.g., data is normalized). In the example plots 605 through 620, BACE and PTP1B are excluded. Referring to plots 605 through 620, the techniques described herein with reference to free energy prediction using a two-state model or one-state model and as described herein provide more effective computation (e.g., 10x to 100x faster computation) compared to other molecular dynamics engines.

[0152] FIG.7 illustrates an example flowchart of a method 700 in accordance with one or more embodiments of the present disclosure. The method 700 may be implemented by the example aspects of a system 100 as described herein. The method 700 supports predicting empirically derived binding free energy associated with a candidate ligand in accordance with one or more embodiments of the present disclosure. The method 700 includes aspects of the method 200 and the method 300 described herein, and repeated descriptions of like elements are omitted for brevity.Atty Ref.: AVX0003PCT

[0153] At 705, the method 700 includes simulating, by one or more processors, an unbinding event in which a candidate ligand dissociates from a biomolecular target to which the candidate ligand is bound, wherein the candidate ligand dissociating from the biomolecular target includes the candidate ligand traveling from a first set of coordinates at which the candidate ligand is bound to the biomolecular target to a second set of coordinates at which the candidate ligand is not bound to the biomolecular target.

[0154] At 710, the method 700 includes calculating, by the one or more processors based on a result of simulating the unbinding event, temporal information associated with the candidate ligand traveling from the first set of coordinates to the second set of coordinates.

[0155] At 715, the method 700 includes calculating, by the one or more processors, a dissociation free energy associated with the candidate ligand and the biomolecular target based on at least one of: the temporal information, a distance between the first set of coordinates and the second set of coordinates, or a combination thereof.

[0156] At 720, the method 700 includes providing the dissociation free energy as a prediction of empirically derived binding free energy associated with the candidate ligand.

[0157] In any one or combination of the embodiments disclosed herein, the method 700 may include screening the candidate ligand in association with one or more pharmaceutical applications based on the prediction of the empirically derived binding free energy. In certain embodiments, the screening comprises screening for molecules that may be tested in vitro, virtual screening, or validation screening of compounds already synthesized or a combination thereof. In certain additional embodiments, the method includes a step of prioritizing compounds to be synthesized.

[0158] In any one or combination of the embodiments disclosed herein, simulating the unbinding event includes accelerating the unbinding event by performing a scaled molecular dynamics run; and the temporal information and the distance are based on a result of accelerating the unbinding event.

[0159] In any one or combination of the embodiments disclosed herein, the scaledmolecular dynamics run is based on the equation ^^^⃗^ = ^∗^^⃗^^ / ^, where P(^⃗) is population,P*(^⃗) is population of the scaled molecular dynamics run, and λ is a scaling factor.

[0160] In any one or combination of the embodiments disclosed herein, simulating the unbinding event includes accelerating the unbinding event by performing a molecularAtty Ref.: AVX0003PCT dynamics run according to a temperature range associated with accelerating the unbinding event; and the temporal information and the distance are based on a result of accelerating the unbinding event.

[0161] In any one or combination of the embodiments disclosed herein, the method 700 may include providing the temporal information to one or more prediction models; and predicting, by the one or more prediction models in response to processing the temporal information associated with the candidate ligand, respective free energy differences of a set of candidate ligands. In some non-limiting examples, the one or more prediction models may support direct prediction via ligand unbinding rate estimation, from dissociation free energy estimates, and two-state model estimates.

[0162] In any one or combination of the embodiments disclosed herein, the method 700 may include generating ranking information associated with the set of candidate ligands based on the respective free energy differences; and screening the set of candidate ligands in association with one or more pharmaceutical application based on the ranking information.

[0163] In any one or combination of the embodiments disclosed herein, the method 700 may include simulating, by the one or more processors, the candidate ligand in a solution; calculating, by the one or more processors based on a result of simulating the candidate ligand in the solution, a free energy associated with the candidate ligand in the solution; and calculating a binding free energy associated with the ligand based on a difference between the dissociation free energy and the free energy associated with the candidate ligand in the solution.

[0164] In any one or combination of the embodiments disclosed herein, calculating thebinding free energy is based on the equation ∆^^ = ∆^^^^^^ − ∆^^^^^^^^ =−^^^^^^∗^^^⃗^^^^^^⁄ ^ ^ + ^^^^^^∗^^⃗^ ^, where: ∆^^ is binding free energy,∆^^^^^^is the dissociation free energy, ∆^^^^^^^^is the free energy, R is a gas constant, T is temperature, ^∗^^^⃗^^^^^ is population associated with simulating the unbinding event, ^∗^^^⃗^^^^^^^is population associated with simulating the candidate ligand in the solution, and λ is a scaling factor.

[0165] In any one or combination of the embodiments disclosed herein, the method 700 may include training one or more prediction models based at least one of the binding free energy, the dissociation free energy, and the free energy; selecting a compound from a libraryAtty Ref.: AVX0003PCT of candidate compounds, wherein the compound includes a second candidate ligand and a second biomolecular target; and predicting, by the one or more prediction models, a binding free energy associated with the second candidate ligand.

[0166] In any one or combination of the embodiments disclosed herein, the method 700 may include screening the second candidate ligand in association with one or more pharmaceutical applications based on the predicted binding free energy.

[0167] In any one or combination of the embodiments disclosed herein, the one or more prediction models include at least one of: a machine learning model; a statistical model; a quantitative structure-activity relationship (QSAR) model; or a combination thereof.

[0168] In any one or combination of the embodiments disclosed herein, the method 700 may include predicting pose information associated with a bound state of the candidate ligand to the biomolecular target, wherein simulating the unbinding event is based on the pose information.

[0169] In any one or combination of the embodiments disclosed herein, predicting the pose information is performed using at least one of x-ray crystallography, docking, flexible docking, molecular dynamics or a combination thereof.

[0170] In any one or combination of the embodiments disclosed herein, the method 700 may include predicting a solubility of a compound including the candidate ligand and the biomolecular target based on at least one of: the temporal information; the dissociation free energy; or a combination thereof.

[0171] In any one or combination of the embodiments disclosed herein, the method 700 may include determining a correlation amount between the temporal information or dissociation free energy and the binding free energy; predicting, based on the correlation amount, at least one of: an empirical half maximal inhibitory concentration (IC50) associated with the candidate ligand and a pharmaceutical application; and an inhibitory constant (Ki) associated with the candidate ligand and the pharmaceutical application.

[0172] Any of the steps, functions, and operations discussed herein can be performed continuously and automatically.

[0173] According to one or more embodiments of the present disclosure, a method is described which support evaluating the potency free energy, Ki, IC50, or similar measure ofAtty Ref.: AVX0003PCT binding efficacy of chemical compounds. The method includes using a computer to run an accelerated ligand unbinding event with either scaled potential molecular dynamics, high temperature molecular dynamics, and / or other methods or simulation techniques that alter the Hamiltonian or potential energy surface to increase the speed of dynamic events. The method includes evaluating the relative differences in the amount of time for a ligand associated with the accelerated ligand unbinding event to move from a starting position of the ligand. The method includes evaluating the relative differences in the amount of time for the ligand to ased on Equation (3) (i.e., ^ ^⃗ = ^∗ $move from the starting position b ^ ^ ^^⃗^%) described herein.The method includes evaluating the dissociation free energy of the ligand moving from the starting position or binding site (also referred to herein as ligand dissociation free energy) as a prediction of empirically derived free energy of binding, IC50, Ki, or related properties.

[0174] In some aspects, the method may include utilizing the predicted ligand dissociation rate or time to dissociate to predict the free energy differences of known ligands or rank the ligands by free energy.

[0175] In some aspects, the method may include calculating the ligand dissociation free energy by accelerating the unbinding of the ligand via scaled potential or high temperature. In some aspects, the method may include calculating the ligand dissociation free energy basedon the equation ^^^^⃗ = ^∗^^⃗^^ / ^.

[0176] In some aspects, the method may include performing a second simulation with a ligand in bulk solvent. Accordingly, for example, the method may include using a two-state system (e.g., ligand bound to biological target, and ligand without biological target) to calculate the absolute free energy difference of binding.

[0177] In some aspects, calculating the absolute free energy difference of binding using the two-state system (e.g., ligand bound to biological target, and ligand without biologicaltarget) may be based on the equation ∆^^ = ∆^^^^^^ − ∆^^^^^^^^ =−^^^^^^∗^^⃗ ^^⁄ ^ ^ + ^^^^ ∗ ^⁄ ^^^^^^ ^^ ^^^⃗^^^^^^^ ^.

[0178] In some aspects, the method may include evaluating compound libraries with machine learning, quantitative structure–activity relationship (QSAR) modeling, artificial intelligence (AI), or other similar methods, in which the data to train the prediction models 148 is generated by predictive simulations of the method.Atty Ref.: AVX0003PCT

[0179] In some aspects, the method may include predicting the ligand binding pose is accurate in 3D space after a number of hypothesized binding poses are predicted with other algorithms such as, but not limited to, docking, flexible docking, or molecular dynamics.

[0180] In some aspects, the method may include predicting the solubility of a compound or ranking the solubility of the compound based on the relative amount of time, dissociation free energy, or free energy taken to dissociate from a crystal or amorphous solid of the compound. In some examples, the method may include predicting or ranking the solubility by measuring “unbinding” from a glob of amorphous or crystalline ligands in solution, and comparing against the diffusion as a free ligand. In some examples, the method may include predicting or ranking the solubility by comparing diffusion rates in different solvents (e.g., octanol, water) to estimate partition coefficients.

[0181] Example aspects of system setup, simulation parameters, and analysis associated with techniques (e.g., method 200, method 300, method 700, iteration process 1500) supported by aspects of the present disclosure are further described herein.

[0182] System setup: All protein structures were obtained from the RCSB Protein Data Bank (RCSB PDB) unless otherwise specified. Protein structures were prepared using the Schrodinger protein preparation workflow to model in missing residues, determine tautomeric and pKastates, and perform a restrained minimization to optimize the structure before docking. If a known binding mode existed for a ligand to the protein target, 3D alignment and / or restrained docking was used. For targets without a known related ligand binding pose, docking was performed with ligands restricted to experimentally validated binding sites. Glide or Autodock Vina were used to generate all ligand-receptor poses unless otherwise noted. Ligand force field parameters were generated following the protocol implemented in OpenMM, which uses the openmmforcefields package to develop parameters consistent with the general Amber force field 2.1. Simulations of the ligand in solution for free energy prediction using a two-state model were set up in the exact same manner but the protein was removed.

[0183] Simulation parameters: Example high temperatures ranged from 400K - 1100K and were maintained with the Langevin thermostat. In some other example implementations, example high temperatures ranged from 450K - 1100K, ranged from 600K – 1000K, or ranged from 600K – 1100K. However, embodiments of the present disclosure are not limited thereto. Individual simulation times were based on the rate of ligand unbinding, typically 1 nsAtty Ref.: AVX0003PCT - 5 ns. Temperatures and simulation times were optimized by running simulations on a known inhibitor(s) and calculating the corresponding koffrates. Simulations were conducted in the NVT ensemble with a 2 fs timestep using OpenMM7. Backbone atoms of the protein were restrained (e.g., default value of σ = 3.0 Å, but adjusted depending on the system) except for residues within 6 Å of the protein binding site. Up to 32 replicate trajectories were run to generate sufficient statistics for reliable calculation of koffrates.

[0184] Analysis: Pytraj, a Python package binding to the CPPTRAJ program, was used for post-processing simulations. Trajectory frames were aligned to the first frame (protein-to- protein) for each respective trajectory. Time-dependent root-mean squared deviation (RMSD) was then calculated for each respective ligand, with a 5.0 Å cutoff to classify ligands in the bound / unbound state. The raw unbinding rate was obtained by calculating the median simulation time for unbinding for all trajectories of a respective ligand, producing a koff value at high temperature. Reweighting is carried out by using Eq. (4) with room temperature (T = 300 K), producing the relative dissociation rate (koff) per ligand.

[0185] Example results supportive of the free energy prediction techniques in accordance with one or more embodiments of the present disclosure are described herein.

[0186] Initial validation of the free energy prediction techniques approach: FAK, HSP90, and p38 MAP kinase. Validation of the free energy prediction approaches described herein required curation of an experimental data set measuring koff values for receptor-ligand complexes. The first set of compounds selected was from a study conducted by Boehringer Ingelheim on p38 MAP kinase.35Experimental koff rates for the p38 MAPK study had a range of five orders of magnitude (10-1–10-6s-1), which provides a good range to demonstrate that the free energy prediction techniques described herein have the ability to distinguish between ligands with fast and slow off-rates. Simulations with T = 1000 K and t = 2 ns were required to observe unbinding of all ligands. Good agreement with experiment were observed, with 13 of 16 data points less than 1 log unit deviations from exact predictions and the other three data points within 2 log unit deviations from exact predictions (FIG.8).

[0187] Validation included a second target, heat shock protein 90 (HSP90), to test the free energy prediction techniques described herein against a protein with multiple distinct structures and highly flexible loops. This was a larger dataset (66 ligands) with a similar range of koffrates (10-0–10-4s-1). The free energy prediction techniques settings of T = 900 K and t = 1 ns were sufficient to observe unbinding of all ligands. Similar agreement existedAtty Ref.: AVX0003PCT between simulation and experiment for the HSP90 data (FIG.8). Additionally, a third dataset was included for focal adhesion kinase (FAK) which contained 14 ligands. Again, the predicted results were well correlated between predicted koff and experimental values. In total, for the 96 ligands across three targets, the free energy prediction techniques achieve highly competitive performance relative to other industry-standard simulation-based predictions of ligand efficacy, with high correlation observed between simulation and experiment (R2= 0.61-0.65) with an RMSE = 0.61-0.76 log(koff) units and an MUE = 0.47-0.61 log(koff) units.

[0188] FIG.8 illustrates example plots 805 through 815 supportive of aspects of the present disclosure.

[0189] The free energy prediction techniques accurately predicts koffvalues for targets with a wide range of dissociation rates. Experimental log(koff) values versus log(koff) values predicted by the free energy prediction techniques for known ligands of FAK, HSP90, and p38 MAPK and Log(koff) values predicted by the free energy prediction techniques were normalized to allow for direct comparison to experiment.

[0190] Under certain conditions, scores provided by the free energy prediction techniques described herein are correlated with and may be applied toward prediction of experimental ligand efficacy. Although it has been demonstrated that the free energy prediction techniques described herein can reproduce koffvalues to a good accuracy (e.g., satisfying a target accuracy), the vast majority of small molecule experimental assays performed in drug discovery research focus on quantifying a measure of ligand efficacy such as binding affinity, either by the equilibrium dissociation constant (Kd) or by IC50 / EC50 (drug concentration producing half-maximal inhibition or effect). Kdis related to the free energy of binding (ΔG) by:

[0191] ∆^ = ^^ln^ / ^^ (5)

[0192] where R is the gas constant and T is temperature. Aspects of the present disclosure support extending the utility of the free energy prediction techniques described herein by developing the ability to extrapolate trends in koff to binding affinity data. Kd is related to kinetic on and off rates by the following equation: ^

[0193] =!:?< !""(6)Atty Ref.: AVX0003PCT

[0194] i.e., that konis inversely correlated with Kdand koffis directly correlated with Kd. Although this is true in theory, in some cases, empirically the strength of these correlative relationships can vary and are system-dependent. The most relevant previous study by Georgi et. al., on the relationship between Kd, kon, and kofffocused on type 1 and type 2 kinase inhibitors. Table 1 lists the ligand set statistics and results of the study. The overall conclusion from this original work was that konwas more correlated to Kdthan koffrates. Georgi et. al.32This Work Number of Ligands 3812 800 Number of Targets 78 37 Number of Distinct 490 295 ScaffoldsKon vs Kd Correlation 0.761 0.541 Koff vs Kd Correlation R=0.364 0.62 Table 1: Ligand statistics and results for Georgi et. al., and in accordance with example aspects of the present disclosure.

[0195] Further analysis was performed on this dataset to reveal that koff and kon are actually similarly correlated with Kd. Inclusion of all datapoints in the original work shows kon is indeed closely anti-correlated with Kd and that koff is poorly correlated with Kd (Fig.9). However, removal of (1) non-congeneric type 1 and type 2 kinase inhibitors for the same target, and (2) inhibitors that focus on transition state manipulation result in equally good kon vs Kdand koffvs Kdcorrelations. Table 1 lists the number of resulting ligands and targets used in this new analysis.

[0196] The improved correlation between Kd and koff is likely the result of focusing our analysis on congeneric series of compounds and excluding cases where a given protein target would likely adopt significant conformational changes across a series of non-congeneric ligands, e.g., ligands that bind to the active or inactive state of a kinase. Although strictly speaking, a correlation between koff and Kd is expected only if kon is constant across the ligands examined, here it can be empirically observed such a correlation across a large number and variety of congeneric series involving likely similar binding modes and protein conformations within each series.Atty Ref.: AVX0003PCT

[0197] FIG.9 illustrates example plots 905 through 920 in accordance with example aspects of the present disclosure. Koffshows good correlation with Kdfor generally related kinase inhibitor targets.

[0198] A) Left (plot 905): Relationship between log(Kd) and log(kon) for diverse (types 1 and 2) kinase inhibitors as well as data for InhA, which is optimized for correlation of Kdwith kon; right (plot 910): Relationship between log(Kd) and log(koff) for diverse (types 1 and 2) kinase inhibitors as well as data for InhA, which is optimized for correlation of Kdwith kon.

[0199] B) Left (plot 915): Relationship between log(Kd) and log(kon) for curated types 1 and 2 kinase inhibitors as well as data for InhA; right (plot 920): Relationship between log(Kd) and log(koff) for curated types 1 and 2 kinase inhibitors as well as data for InhA.

[0200] Based on the observation that koff is empirically correlated to Kd for a large and diverse set of ligands, the ability of the free energy prediction techniques described herein have been tested to directly predict Kd and compare its accuracy against the most widely used approach in industry for binding affinity prediction, namely relative free energy calculations. The following targets were chosen: TYK2, thrombin, MCL1, JNK1, CDK2, and p38 kinase. 2 targets (BACE and PTP1B) were excluded because they are known to be challenging to absolute efficacy methods due to loop rearrangement and high protein re-organizational energies upon binding (PTP1B) and titration of key residues upon binding (BACE).18

[0201] Temperature scans were performed to identify the optimal temperature to produce a dynamic range of koffvalues within a timeframe of ≤4 ns per replica. As a benchmark, the performance of the free energy prediction techniques described herein were compared against the results of a comparable academic free energy perturbation (FEP) study using the thermodynamic integration (TI) implementation in AMBER18 described in Song, L. et al., Using AMBER18 for Relative Free Energy Calculations, Journal of Chemical Information and Modeling 201959 (7), 3128-3135. The work with AMBER-TI also used the GAFF forcefield as did the free energy prediction techniques described herein. After normalizing scores from runs by both the AMBER TI and the free energy prediction techniques described herein to free energies (i.e., setting the equivalent free energy ranges for each protein target), we demonstrate that the free energy prediction techniques described herein exhibit comparable accuracy with AMBER TI (FIG.8). We next compared to the original FEP+ work published by Wang et. al, the correlation and error metrics of FEP+ were improvedAtty Ref.: AVX0003PCT compared to either the free energy prediction techniques or AMBER-TI. We expect that the forcefield used by FEP+ OPLS is superior to the GAFF forcefield and this is the reason we see improved performance from FEP+ with OPLS and comparable performance between the free energy prediction techniques and AMBER-TI which used GAFF. In some aspects, the free energy prediction techniques may be applied to the change in accuracy when using other forcefields, and we expect the method to benefit from the ongoing work to improve forcefields.

[0202] The prediction of small molecule binding energies in accordance with the free energy prediction techniques described herein is competitive with more expensive free energy approaches. A) Experimental versus computationally predicted free energies of binding for TYK2, thrombin, MCL1, JNK1, CDK2, and p38 kinase obtained using the FEP+ implementation in Desmond. B) Experimental versus computationally predicted free energies of binding for TYK2, thrombin, MCL1, JNK1, CDK2, and p38 kinase obtained using the thermodynamic integration (TI) implementation in AMBER18. C) Experimental versus computationally predicted free energies of binding for TYK2, thrombin, MCL1, JNK1, CDK2, and p38 kinase obtained using the free energy prediction techniques described herein. Experimental data for the six protein targets shown were taken from Wang et al. JACS 2015.

[0203] The free energy prediction techniques described herein can achieve a high degree of enrichment in small molecule virtual screening. One of the most promising applications of the free energy prediction techniques described herein in drug discovery is in structure-based virtual screening (VS). Although the advent of ultra-large commercially accessible compound libraries on the billion-compound scale make computational screening an attractive option for hit-finding, the traditionally low experimental hit rates from structure-based and ligand-based screens make it difficult to reduce the virtual hit lists down to a manageable number of compounds. The ability of the free energy prediction techniques described herein have been tested to increase enrichment of true positive hits using compound lists that were pre- prioritized via traditional virtual screening. We took the DUD-E dataset and selected the MAP38 kinase compounds which were docked to the MAP38 protein using AutoDock Vina. To mimic a typical VS campaign that would use docking results to inform compound selection for subsequent synthesis and testing, we took the top 100 scoring compounds (out of a total of 36,379 docked compounds), consisting of 20 active compounds and 80 decoy compounds, re-evaluated them using the free energy prediction techniques described herein, and compared the resulting enrichment to the docking results. In the example describedAtty Ref.: AVX0003PCT herein, the 20 active compounds were not spiked for MET or P38, about 20% of the top 100 compounds in each target were active, and 80% were inactive, which gave a reasonable range for testing enrichment and in which spiking was not required.

[0204] FIGS.10A through 10F illustrate example plots 1005 through 1030 in accordance with example aspects of the present disclosure and the re-evaluation described herein.

[0205] If we use docking scores from the entire 36,379-compound dataset, enrichment of active compounds over decoy compounds follows an expected trend, with an enrichment factor of 10 (Fig.10A). However, when we look at the top 100 docked results, differentiation between actives and decoys is negligible, leading to an enrichment factor of 2.1 (Fig.10B). In contrast, the free energy prediction techniques described herein vastly outperform docking, showing a 27-fold enrichment within the top 100 compounds by correctly identifying nearly 80% of the actives before ranking any decoys (Fig.10C). We performed the same analysis on the MET kinase dataset and saw a similar trend: docking improved enrichment of the entire compound library (Fig.10D), poor enrichment of the top 100 compounds (2.9) (Fig. 10E), and significantly improved enrichment with the free energy prediction techniques described herein (20) (Fig.10F). In numerous in-house studies on other protein targets, we find that the free energy prediction techniques described herein consistently and significantly outperform docking alone. Thus, the free energy prediction techniques described herein provide an invaluable tool to post-process VS results and increase both hit rate and efficiency of the hit-to-lead workflow.

[0206] The free energy prediction techniques described herein provide superior enrichment over docking, and have nearly perfect enrichment in a scenario exactly like using the free energy prediction techniques described herein as a final filter to the VS.

[0207] Early enrichment is improved more than 57 fold among top100 compounds.

[0208] The free energy prediction techniques described herein significantly improve enrichment to correctly identify active molecules from a virtual screening campaign. A) Enrichment curve for docking of the entire 36,379-molecule library to the MAP38 kinase. B) Enrichment curve for the top 100 docked compounds to MAP38 kinase. C) Enrichment curve for the free energy prediction techniques described herein, using the top 100 docked compounds to model dissociation of small molecules from MAP38 kinase. D) Enrichment curve for docking of the entire 11,398-molecule library to the MET kinase. E) Enrichment curve for the top 100 docked compounds to MET kinase. F) Enrichment curve for the freeAtty Ref.: AVX0003PCT energy prediction techniques described herein, using the top 100 docked compounds to model dissociation of small molecules from MET kinase. ROC: receiver operator characteristics; enrichment factor 1%: enrichment of actives over decoys using 1% of the decoy library, as given by EF = (a / n) / (A / N), where a is the number of actives found in sample size n, A is the total number of actives, and N is the total number of ligands (decoys and actives).

[0209] The free energy prediction techniques described herein in practice

[0210] The free energy prediction techniques described herein may be applied to poly- ADP ribose glycohydrolase (PARG). The free energy prediction techniques described herein have been used to rapidly and efficiently narrow down selection of compounds for synthesis. The free energy prediction techniques described herein have also been used prospectively to predict the free energy of compounds. The free energy prediction techniques described herein have been tested against FEP using a subset of PARG compounds. No changes were made to the three-dimensional structures of the PARG-ligand complexes used in the FEP calculations, allowing a direct comparison of the performance of the free energy prediction techniques described herein to FEP results. The free energy prediction techniques described herein predicted binding energies with comparable accuracy to the FEP calculations, showing moderate correlation with experimental results and acceptable levels of error (FIG.10). Note that the compounds were chosen prospectively and thus a systematic overprediction bias in both methods may be observed. However, in comparison to FEP, the computational turn- around time for calculations using the free energy prediction techniques described herein was approximately 100 times faster.

[0211] Multiple potential applications for the free energy prediction techniques described herein have been demonstrated in an industrial setting. For the molecular modeling practitioner, the free energy prediction techniques described herein provide a tool that when used properly can provide real benefits to a drug development campaign.

[0212] FIG.11 illustrates example plots 1105 and 1110 in accordance with example aspects of the present disclosure and the re-evaluation described herein.

[0213] The free energy prediction techniques described herein perform competitively with prospective free energy calculations. A) Prospective predictions of free energy of binding of PARG compounds via FEP calculations versus experimental free energy via IC50. B) Prospective predictions of free energy of binding of PARG compounds via calculationsAtty Ref.: AVX0003PCT using the free energy prediction techniques described herein versus experimental free energy via IC50. RMSE: root-mean squared error; MUE: mean unsigned error.

[0214] The techniques described herein for free energy prediction using a two-state model have been validated with multiple protein ligand systems. For each of the systems, two states were simulated: 1) the protein with the ligand, and 2) the ligand alone in solution. This gives us two end points along a reaction coordinate from which the free energy prediction techniques described herein can derive a free energy by subtraction of the two values. The free energy prediction techniques described herein may calculate the free energy by reweighting the populations at a specified distance, usually the distance from the starting positions as measured by the change in the center of mass (COM) of the ligand of interest.

[0215] Generally, a target favorable state (e.g., the most favorable state) is 1-2 Å and represents the distance with which we compare to the MD simulation of the ligand alone in solution. Since COM movement does not take an equivalent amount of time at all distances, we see some variation between the free energy at different points in the COM. For example, COM varies by x because the volume of each COM bin changes (e.g., larger) as the COM goes out further. Hence, it is important to standardize where we subtract the difference in free energy, and in some examples, the techniques described herein include implementing such standardization by using the same bin in each simulation. (FIG.4C). The free energy prediction techniques supported by aspects of the present disclosure may include using a ligand simulation as described herein. In some aspects, a reasonable estimate of the ligand free energy can also be derived from the Einstein-Smoluchowski diffusion equations and use the time of diffusion to estimate populations.

[0216] The techniques described herein for free energy prediction using a two-state model have been validated on focal adhesion kinase (FAK), and P38 kinase and observe good agreement of both the absolute values of the predicted free energy and the correlation of the free energies (FIG.2).

[0217] FIG.12 illustrate example plots 1205 and 1210 in accordance with example aspects of the present disclosure. FIG.12 illustrates predictions (at plot 1205) of the absolute binding free energy of focal adhesion kinase ligands FAK and (at plot 1210) of p38 ligands using a two-state model in accordance with example aspects of the present disclosure.

[0218] The free energy prediction techniques described herein provide an enhanced sampling method for the prediction of ligand off-rates. The methodology is firmly rooted inAtty Ref.: AVX0003PCT statistical mechanics and has been validated across eight targets and hundreds of ligands. The examples show how the free energy prediction techniques described herein has been successfully used prospectively in drug discovery campaigns. The free energy prediction techniques described herein support accelerating ligand dissociation and can accurately capture unbinding kinetics within a ½ log unit, all with a cumulative simulation time of tens of ns or less per ligand.

[0219] For a large variety of protein targets and compounds sets, the free energy prediction techniques described herein support indirectly predicting binding affinities or IC50 values. The free energy prediction techniques described herein are about 10-100X faster than the majority of free energy implementations but maintain the same degree of accuracy. A major advantage of the free energy prediction techniques described herein is that the techniques are capable of obtaining results without using alchemical changes, negating the requirement for a reference compound that is inherent in relative FEP approaches. This allows for de novo design of compounds and absolute comparisons between ligands in a virtual screening application. Although in some cases, koff rates are not always directly correlated with Kd (and thus free energies), in many cases, the free energy prediction techniques described herein support identifying potent and efficacious compounds, with the added potential benefit of an increased Koffwhich is often a desirable pharmacologically property in its own right. The free energy prediction techniques described herein may be implemented in almost any MD engine and, with the ease of access to GPU resources and the low computational cost of the free energy prediction techniques described herein, the free energy prediction techniques described herein provide an important tool in almost any computational drug discovery program moving forward.

[0220] Additionally a methodology in accordance with example aspects of the present disclosure is provided that can accurately predict absolute binding free energies both in range and ranking, using a two-state model as described herein. Another significant development is the relative speed of the calculations. Simulations can be completed in as little as a few nanoseconds of total simulation time compared with up to one thousand or more nanoseconds of simulation time for other known absolute free energy methods. The free energy prediction using a two-state model as described herein is also novel because it is the first known two- state model to use populations from two separate simulations to directly compute a free energy of binding. In some cases, populations from simulations are often more accurate in computing free energies than using the energy functions themselves due to the energeticAtty Ref.: AVX0003PCT noise problem in molecular simulations, and thus the free energy prediction using a two-state model provide technological improvements compared to other approaches.

[0221] Enrichment examples in accordance with example aspects of the present disclosure are described herein.

[0222] An example workflow in accordance with one or more embodiments of the present disclosure is described herein. The workflow may include docking a large or very large number of compounds (thousands to millions). The workflow may include identifying the top scoring docked compounds (both active and inactive). The workflow may include calculation of docking-based Receiver Operating Characteristic (ROC) (also referred to herein as a receiver operator curve). The workflow may include implementing the free energy prediction techniques described herein to determine whether further enrichment is possible on top-scoring docked compounds.

[0223] Table 2 below illustrates example results of docking-based ROC compared to ROC as determined by the free energy prediction techniques described herein. Table 2 illustrates that the free energy prediction techniques described herein provide superior performance to docking, which may translate to better VS hit rates.

[0224] An example implementation for identifying prospective virtual screen protein– protein interaction (PPI) targets in accordance with one or more embodiments of the present disclosure is described herein.

[0225] A system configured for performing the free energy prediction techniques and techniques for free energy prediction using a two-state model described herein may be applied to screening or analyzing a large library of idea compounds or commercially available compounds (e.g., about 10 million compounds). The system may rapidly score the compounds, reducing the quantity to about 25,000 compounds with good docking scoresAtty Ref.: AVX0003PCT (e.g., satisfying a score threshold). The system may perform the free energy prediction techniques and techniques for free energy prediction against the 25,000 compounds for final prioritization, identifying about 800 compounds based on the predictions. Accordingly, for example, an entity may purchase and experimentally test the 800 compounds. In an example, the system may determine a diverse set of 40 molecules with 16 unique novel scaffolds with measurable IC50.

[0226] Fig.13 illustrates an example graph 1305 illustrating compounds screened and scored (as determined by the free energy prediction techniques and techniques for free energy prediction using a two-state model) in accordance with example aspects of the present disclosure.

[0227] Fig.14 illustrates an example plot 1405 of representative hits from the FP Assay in accordance with example aspects of the present disclosure. Table 3 associated with the plot 1405 is provided herein.Table 3

[0228] Table 4 below illustrates an example of the impact of the free energy prediction techniques described herein on virtual screens in accordance with example aspects of the present disclosure.Atty Ref.: AVX0003PCT [s described herein and virtual screenings (VSs) according to other approaches. VSs involved 5M~1B compounds screened, 200-300 compounds prioritized and purchased. Hit rate is approximately 1.6 novel cores per VS prior to usage of the free energy prediction techniques described herein: total 11 novel cores found from 7 VSs (prior to the applying free energy prediction techniques). Use of the free energy prediction techniques described herein identified more novel cores (e.g., 16) than the 7 prior VSs combined (11)

[0230] FIG.15 illustrates an example iteration process 1500 in accordance with one or more embodiments of the present disclosure. The iteration process 1500 (at 1515) supports prediction of temporal information, dissociation free energy, or free energy in accordance with the techniques described herein. The iteration process 1500 (at 1520) supports further prioritization and synthesis. The iteration process 1500 may be performed by the system 100 (e.g., device 105, server 110) described herein.

[0231] The system 100 analyzes a lead compound having high potency / cellular activity but low solubility and adsorption, distribution, metabolism, elimination, and toxicity (ADMET) issue according to the iteration process 1500. The system 100 identifies new molecules with high potency / cellular activity, high solubility and no ADMET Issue. The compounds are proceeded to in vivo testing.

[0232] In some embodiments, the free energy prediction techniques described herein support features for improving accuracy through longer lower temperature simulations for lead-opt cost estimate. In some embodiments, the free energy prediction techniques describedAtty Ref.: AVX0003PCT herein support features for improving accuracy through the use of median value of simulation time rather than true summation of simulation time.

[0233] In some embodiments, the free energy prediction techniques described herein support features for improving accuracy through adding salt, for example, especially for charged ligands.

[0234] In some embodiments, the free energy prediction techniques described herein support features for improving accuracy through using polarizable force fields, for example, especially for charged ligands.

[0235] In some embodiments, the free energy prediction techniques described herein support features for improving accuracy through using OPLS / other force fields. Accordingly, for example, the free energy prediction techniques include features for adjusting accuracy according to a target accuracy level.

[0236] Further example results supportive of applying the free energy prediction techniques for solubility prediction in accordance with one or more embodiments of the present disclosure are described herein.

[0237] FIG.16 illustrates a table 1600 of data indicating that the free energy prediction techniques described herein can be used to predict the best pose from a list of poses generated by docking software or other means.

[0238] Particularly, the table 1600 lists root RMSD of atomic positions with respect to the protein structures and docked ligands, as determined using Induced Fit Docking (IFD) (see 1605) and Binding Pose Metadynamics (see 1610) as described in the publication Clark, A.J., Tiwary, P., Borrelli, K., Feng, S., Miller, E.B., Abel, R., Friesner, R.A. and Berne, B.J., 2016. Prediction of protein–ligand binding poses via a combination of induced fit docking and metadynamics simulations. Journal of chemical theory and computation, 12(6), pp.2990- 2998, and further, as determined using the free energy prediction techniques (see 1615) described herein.

[0239] Protein and ligand names included in the table 1600 are references to protein data bank (PDB) codes used by the Research Collaboratory for Structural Bioinformatics PDB.Atty Ref.: AVX0003PCT

[0240] 1PXJ (PDB: pdb_00001pxj) (2003. Structure 11: 399-410). HUMAN CYCLIN DEPENDENT KINASE 2 COMPLEXED WITH THE INHIBITOR 4-(2,4-Dimethyl-thiazol- 5-yl)-pyrimidin-2-ylamine.

[0241] 1WCC (PDB: pdb_00001wcc) (2005. J Med Chem 48: 403-413). Human CDK2 (298aa; MW: 34 kDa).

[0242] The table 1600 includes 12 Cyclin Dependent Kinase 2 (CDK2) ligands docked into non cognate crystal structures, as described at the above-mentioned publication by Clark, A.J., et al.

[0243] The table 1600 includes results generated by taking the top 5 poses from IFD and rescoring the top 5 poses, and applying the free energy prediction techniques described herein to select the top pose based on scores generated using the free energy prediction techniques. The table 1600 includes results generated by comparing the top pose by the 3 scoring methods (i.e., IFD, Metadynamics binding pose score, and score generated using the free energy prediction techniques described herein) and averaging the RMSD of each of those poses selected by the 3 methods to find which one has the best performance.

[0244] As seen by table 1600, the free energy prediction techniques described herein show the ability to select compounds with low RMSD to crystal structures in this CDK2 dataset.

[0245] An example implementation of the free energy prediction techniques described herein in accordance with one or more embodiments of the present disclosure as applied to solubility and a solubility workflow description is now described.

[0246] General aspects of solubility are first described herein. The relationship between a ligand's center of mass (COM) and the mass of a solute's COM is a fundamental concept in coordination chemistry and molecular dynamics. Ligands are molecules or ions that bind to a central metal atom or ion, forming a coordination complex. The COM of a ligand represents the average position of the constituent atoms of the ligand, weighted by their masses. The COM of the solute, which is the ligand in this context, is similarly determined by the masses and positions of the ligand's atoms. This COM may be crucial in understanding the complex's structure and dynamics, as the COM influences interactions with the solvent and other molecules.Atty Ref.: AVX0003PCT

[0247] A breakdown of some key aspects is provided herein. Ligand COM: The weighted average position of all the atoms in the ligand, taking into account their masses. Solute COM: In the context of a coordination complex, the solute is the ligand. The COM is the weighted average position of all atoms in the ligand.

[0248] Importance: The COM of the ligand influences molecular shape, solvent interactions, and molecular dynamics. As to molecular shape, the COM helps define the overall shape and orientation of the ligand within the complex. As to solvent interactions, the COM determines how the ligand interacts with solvent molecules. As to molecular dynamics, the COM plays a role in how the ligand moves and rotates within the complex.

[0249] Example: Consider a simple complex where a metal ion is bound to a ligand like water (H2O). The COM of the H2O ligand would be the average position of the two hydrogen atoms and the oxygen atom, each weighted by its respective mass. This COM influences how the water molecule interacts with the metal ion and the surrounding solvent.

[0250] In essence, understanding the COM of a ligand is crucial for comprehending the structure and behavior of coordination complexes, and the free energy prediction techniques described herein support effective understanding of the COM of ligands, thereby providing improved comprehension of the structure and behavior of coordination complexes.

[0251] Turning to an example solubility workflow in accordance with one or more embodiments of the present disclosure, the solubility workflow may include generating initial configuration with a packing optimization tool for a molecular dynamics simulations (e.g., PACKMOL) and 100 ligands in a box. The solubility workflow may include running a short equilibration at 300K to create packed solute. The solubility workflow may include selecting the ligand with a COM furthest from mass solute COM. The solubility workflow may include performing or executing a restrained run at high temperature with all but 1 ligand restrained, in which the 1 ligand which is not restrained has a highest ligand COM to mass solute COM. The solubility workflow may include calculating free energy as a ligand moves from solute to solvent with population reweighting

[0252] The solubility workflow may include estimating unbound state free energy via ligand simulation or a diffusion equation as follows:

[0253] ∆GABCDEFCFGH = ∆GEBDI7 − ∆GDIEBDI7 (7)Atty Ref.: AVX0003PCT

[0254] Applying the free energy prediction techniques described herein provides effective determination of the bound and unbound states, and thereby provides effective calculation or prediction of the solubility of pharmaceutically relevant compounds.

[0255] FIG.17 illustrates example molecules and respective solubilities determined using a comparative free energy perturbation approach described in Hong, R.S., Rojas, A.V., Bhardwaj, R.M., Wang, L., Mattei, A., Abraham, N.S., Cusack, K.P., Pierce, M.O., Mondal, S., Mehio, N. and Bordawekar, S., 2023. Free energy perturbation approach for accurate crystalline aqueous solubility predictions. Journal of Medicinal Chemistry, 66(23), pp.15883- 15893.

[0256] FIG.18 illustrates example plots 1805 and 1810 of solubility determined using the molecular simulation methods described herein as applied to the molecules indicated at FIG. 17, in comparison to the solubilities as determined using the comparative free energy perturbation approach.

[0257] Plot 1805 illustrates solubilities determined using a molecular simulation method described herein, using logS, as expressed by the following equation:

[0258] (8).

[0259] LogS is a common unit for measuring solubility. It is the 10-based logarithm of the solubility measured in mol / l unit, so logS = log (solubility measured in mol / l). The solubility predictor predicts solubility values at 25 °C.

[0260] Plot 1810 illustrates solubilities determined using a molecular simulation method described herein, with free energy prediction using a two-state model.

[0261] Referring to the results illustrated at plot 1805 and plot 1810, it can be seen that, the molecular simulation methods described herein provides accurate solubility predictions as compared to experimental determination of the solubilities.

[0262] Aspects of the present disclosure may take the form of an embodiment that is entirely hardware, an embodiment that is entirely software (including firmware, resident software, micro-code, etc.) or an embodiment combining software and hardware aspects that may all generally be referred to herein as a “circuit,” “module,” or “system.” AnyAtty Ref.: AVX0003PCT combination of one or more computer-readable medium(s) may be utilized. The computer- readable medium may be a computer-readable signal medium or a computer-readable storage medium.

[0263] A computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non- exhaustive list) of the computer-readable storage medium would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0264] The terms “determine,” “calculate,” “compute,” and variations thereof, as used herein, are used interchangeably and include any type of methodology, process, mathematical operation or technique.

[0265] While various embodiments of the present disclosure are described herein, it will be understood by those skilled in the art that such embodiments are provided by way of example only. It will be understood by those skilled in the art that numerous modifications and changes to, and variations and equivalent substitutions of, the embodiments described herein can be made without departing from the scope of the disclosure. It is understood that various alternatives to the embodiments described herein may be employed in practicing the disclosure, and modifications may be made to adapt a particular structure or material to the teachings of the disclosure. It is also understood that every embodiment of the disclosure may optionally be combined with any one or more of the other embodiments described herein which are consistent with that embodiment.

[0266] Where elements are presented in list format (e.g., in a Markush group), it is understood that each possible subgroup of the elements is also disclosed, and any one or more elements can be removed from the list or group.

[0267] It is also understood that, unless clearly indicated to the contrary, in any method described or claimed herein that includes more than one act or step, the order of the acts orAtty Ref.: AVX0003PCT steps of the method is not necessarily limited to the order in which the acts or steps of the method are recited, but the disclosure encompasses embodiments in which the order is so limited.

[0268] It is further understood that, in general, where an embodiment in the description or the claims is referred to as comprising one or more features, the disclosure also encompasses embodiments that consist of, or consist essentially of, such feature(s).

[0269] It is also understood that any embodiment of the disclosure, e.g., any embodiment found within the prior art, can be explicitly excluded from the claims, regardless of whether or not the specific exclusion is recited in the specification.

[0270] Headings are included herein for reference and to aid in locating certain sections. Headings are not intended to limit the scope of the embodiments and concepts described in the sections under those headings, and those embodiments and concepts may have applicability in other sections throughout the entire disclosure.

[0271] All patent literature and all non-patent literature cited herein are incorporated herein by reference in their entirety to the same extent as if each patent literature or non- patent literature were specifically and individually indicated to be incorporated herein by reference in its entirety.

[0272] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0273] Where a range of values is provided, it is understood that each intervening value between the upper and lower limit of that range and any other stated or intervening value in that stated range is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges is also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.

[0274] The articles "a" and "an" as used herein and in the appended claims are used herein to refer to one or to more than one (i.e., to at least one) of the grammatical object of the article unless the context clearly indicates otherwise. By way of example, "an element" means one element or more than one element.Atty Ref.: AVX0003PCT

[0275] The term “exemplary” as used herein means “serving as an example, instance or illustration”. Any embodiment or feature characterized herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments or features.

[0276] The phrase "and / or," as used herein in the specification and in the claims, should be understood to mean "either or both" of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with "and / or" should be construed in the same fashion, i.e., "one or more" of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the "and / or" clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to "A and / or B", when used in conjunction with open-ended language such as "comprising" can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.

[0277] As used herein in the specification and in the claims, "or" should be understood to have the same meaning as "and / or" as defined above. For example, when separating items in a list, "or" or "and / or" shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as "only one of' or "exactly one of," or, when used in the claims, "consisting of," will refer to the inclusion of exactly one element of a number or list of elements. In general, the term "or" as used herein shall only be interpreted as indicating exclusive alternatives (i.e., "one or the other but not both") when preceded by terms of exclusivity, such as "either," "one of," "only one of," or "exactly one of."

[0278] In the claims, as well as in the specification above, all transitional phrases such as "comprising," "including," "carrying," "having," "containing," "involving," "holding," "composed of," and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases "consisting of' and "consisting essentially of' shall be closed or semi-closed transitional phrases, respectively.

[0279] As used herein in the specification and in the claims, the phrase "at least one," in reference to a list of one or more elements, should be understood to mean at least one element selected from anyone or more of the elements in the list of elements, but not necessarilyAtty Ref.: AVX0003PCT including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase "at least one" refers, whether related or unrelated to those elements specifically identified. Thus, as a nonlimiting example, "at least one of A and B" (or, equivalently, "at least one of A or B," or, equivalently "at least one of A and / or B") can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.

[0280] It should also be understood that, in certain methods described herein that include more than one step or act, the order of the steps or acts of the method is not necessarily limited to the order in which the steps or acts of the method are recited unless the context indicates otherwise.

[0281] The term “about” or “approximately” means an acceptable error for a particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined. In certain embodiments, the term “about” or “approximately” means within one standard deviation. In some embodiments, when no particular margin of error (e.g., a standard deviation to a mean value given in a chart or table of data) is recited, the term “about” or “approximately” means that range which would encompass the recited value and the range which would be included by rounding up or down to the recited value as well, taking into account significant figures. In certain embodiments, the term “about” or “approximately” means within 10% or 5% of the specified value. Whenever the term “about” or “approximately” precedes the first numerical value in a series of two or more numerical values or in a series of two or more ranges of numerical values, the term “about” or “approximately” applies to each one of the numerical values in that series of numerical values or in that series of ranges of numerical values.

[0282] Whenever the term “at least” or “greater than” precedes the first numerical value in a series of two or more numerical values, the term “at least” or “greater than” applies to each one of the numerical values in that series of numerical values.Atty Ref.: AVX0003PCT

[0283] Whenever the term “no more than” or “less than” precedes the first numerical value in a series of two or more numerical values, the term “no more than” or “less than” applies to each one of the numerical values in that series of numerical values.

Claims

Atty Ref.: AVX0003PCT CLAIMS What is claimed is:

1. A method comprising: simulating, by one or more processors, an unbinding event in which a candidate ligand dissociates from a biomolecular target to which the candidate ligand is bound, wherein the candidate ligand dissociating from the biomolecular target comprises the candidate ligand traveling from a first set of coordinates at which the candidate ligand is bound to the biomolecular target to a second set of coordinates at which the candidate ligand is not bound to the biomolecular target; calculating, by the one or more processors based on a result of simulating the unbinding event, temporal information associated with the candidate ligand traveling from the first set of coordinates to the second set of coordinates; calculating, by the one or more processors, a dissociation free energy associated with the candidate ligand and the biomolecular target based on at least one of: the temporal information, a distance between the first set of coordinates and the second set of coordinates, or a combination thereof; and providing the dissociation free energy as a prediction of empirically derived binding free energy associated with the candidate ligand.

2. The method of claim 1, further comprising: screening the candidate ligand in association with one or more pharmaceutical applications based on the prediction of the empirically derived binding free energy.

3. The method of claim 1, wherein: simulating the unbinding event comprises accelerating the unbinding event by performing a scaled molecular dynamics run; and the temporal information and the distance are based on a result of accelerating the unbinding event.

4. The method of claim 3, wherein the scaled molecular dynamics run is based on the equationAtty Ref.: AVX0003PCT ^^^^⃗ = ^∗^^⃗^^ / ^where P(^⃗) is population, P*(^⃗) is of the scaled molecular dynamics run, and λ is a scaling factor.

5. The method of claim 1, wherein: simulating the unbinding event comprises accelerating the unbinding event by performing a molecular dynamics run according to a temperature range associated with accelerating the unbinding event; and the temporal information and the distance are based on a result of accelerating the unbinding event.

6. The method of claim 1, further comprising: providing the temporal information to one or more prediction models; and predicting, by the one or more prediction models in response to processing the temporal information associated with the candidate ligand, respective free energy differences of a set of candidate ligands.

7. The method of claim 6, further comprising: generating ranking information associated with the set of candidate ligands based on the respective free energy differences; and screening the set of candidate ligands in association with one or more pharmaceutical application based on the ranking information.

8. The method of claim 1, further comprising: simulating, by the one or more processors, the candidate ligand in a solution; calculating, by the one or more processors based on a result of simulating the candidate ligand in the solution, a free energy associated with the candidate ligand in the solution; and calculating, by the one or more processors, a binding free energy associated with the candidate ligand based on a difference between the dissociation free energy and the free energy associated with the candidate ligand in the solution.Atty Ref.: AVX0003PCT 9. The method of claim 8, wherein calculating the binding free energy is based on the equation∆^ ∗^ = ∆^^^^^^ − ∆^^^^^^^^ = −^^^^^^ ^^^⃗^^^^^^⁄ ^ ^ + ^^^^^^∗^^^⃗^^^^^^^^⁄ ^ ^wherein:∆^^^^^^is the dissociation free energy, ∆^^^^^^^^is the free energy, R is a gas constant, T is temperature, ^∗^^^⃗^^^^^ is population associated with simulating the unbinding event, ^∗^^^⃗^^^^^^^is population associated with simulating the candidate ligand in the solution, and λ is a scaling factor.

10. The method of claim 8, further comprising: training one or more prediction models based at least one of the binding free energy, the dissociation free energy, and the free energy; selecting a compound from a library of candidate compounds, wherein the compound comprises a second candidate ligand and a second biomolecular target; and predicting, by the one or more prediction models, a binding free energy associated with the second candidate ligand.

11. The method of claim 10, further comprising: screening the second candidate ligand in association with one or more pharmaceutical applications based on the predicted binding free energy.

12. The method of claim 10, wherein the one or more prediction models comprise at least one of: a machine learning model; a statistical model; a quantitative structure–activity relationship (QSAR) model; or a combination thereof.Atty Ref.: AVX0003PCT 13. The method of claim 1, further comprising: predicting pose information associated with a bound state of the candidate ligand to the biomolecular target, wherein simulating the unbinding event is based on the pose information.

14. The method of claim 13, wherein predicting the pose information is performed using at least one of x-ray crystallography, docking, flexible docking, molecular dynamics or a combination thereof.

15. The method of claim 1, further comprising: predicting a solubility of a compound comprising the candidate ligand and the biomolecular target based on at least one of: the temporal information; the dissociation free energy; or a combination thereof.

16. The method of claim 1, further comprising: determining a correlation amount between the temporal information or dissociation free energy and the binding free energy; predicting, based on the correlation amount, at least one of: an empirical half maximal inhibitory concentration (IC50) associated with the candidate ligand and a pharmaceutical application; and an inhibitory constant (Ki) associated with the candidate ligand and the pharmaceutical application.

17. A system comprising: a processor; and a memory storing instructions thereon that, when executed by the processor, cause the processor to: simulate an unbinding event in which a candidate ligand dissociates from a biomolecular target to which the candidate ligand is bound, wherein the candidate ligand dissociating from the biomolecular target comprises the candidate ligand traveling from a first set of coordinates at which the candidate ligand is bound to the biomolecular target to aAtty Ref.: AVX0003PCT second set of coordinates at which the candidate ligand is not bound to the biomolecular target; calculate, based on a result of simulating the unbinding event, temporal information associated with the candidate ligand traveling from the first set of coordinates to the second set of coordinates; calculate a dissociation free energy associated with the candidate ligand and the biomolecular target based on at least one of: the temporal information, a distance between the first set of coordinates and the second set of coordinates, or a combination thereof; and provide the dissociation free energy as a prediction of empirically derived binding free energy associated with the candidate ligand.

18. The system of claim 17, wherein the instructions, when executed by the processor, further cause the processor to perform operations comprising: screening the candidate ligand in association with one or more pharmaceutical applications based on the prediction of the empirically derived binding free energy.

19. The system of claim 17, wherein: the instructions causing the processor to simulate the unbinding event cause the processor to accelerate the unbinding event by at least one of: performing a scaled molecular dynamics run; and performing a molecular dynamics run according to a temperature range associated with accelerating the unbinding event, wherein the temporal information and the distance are based on a result of accelerating the unbinding event.

20. The system of claim 17, wherein the instructions, when executed by the processor, further cause the processor to perform operations comprising: simulate the candidate ligand in a solution; calculate, based on a result of simulating the candidate ligand in the solution, a free energy associated with the candidate ligand in the solution; andAtty Ref.: AVX0003PCT calculate a binding free energy associated with the candidate ligand based on a difference between the dissociation free energy and the free energy associated with the candidate ligand in the solution.

Citation Information

Patent Citations

  • Compositions And Methods For Predicting Inhibitors Of Protein Targets

    US20120142623A1

  • Physics-based computational methods for predicting compound solubility

    US20160321432A1

  • Methods and apparatuses for training prediction model

    US20230097667A1

  • Drug ranking method and system, comparison method for drug ranking method and new use of drug selected using the same

    US20230184749A1