Characterizing interactions between compounds and polymers using negative pose data and model conditioning
Patent Information
- Application Number
- JP2024519522
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-10-01
- Filing Date
- 2022-09-29
- Publication Date
- 2025-10-03
AI Technical Summary
Traditional structure-based virtual high-throughput screening (vHTS) machine learning methods lack pose sensitivity, leading to inaccurate characterization of interactions between compounds and target polymers, as they treat compounds and polymers separately and do not account for the correct orientation of molecular interactions.
Conditioning vHTS machine learning models to be pose sensitive by training them on both positive and negative poses of compound-polymer interactions using atomic coordinates from methods like nuclear magnetic resonance or cryo-electron microscopy, and adjusting model parameters through loss functions to improve accuracy.
Enhances the accuracy of characterizing compound-polymer interactions by making the models sensitive to pose, reducing errors and improving the reliability of interaction scoring.
Smart Images

Figure 00000069_0000 
Figure 00000069_0001 
Figure 00000069_0002
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 251,142, entitled “CHARACTERIZATION OF INTERACTIONS BETWEEN COMPOUNDS AND POLYMERS USING NEGATIVE POSE DATA AND MODEL CONDITIONING,” filed October 1, 2021, which is incorporated herein by reference.
[0002] The present application is directed to the use of models to characterize the interactions between test compounds and target polymers. [Background technology]
[0003] Essentially, biological systems function through the physical interaction of molecules such as compounds with target polymers. Structure-based virtual high-throughput screening (vHTS) machine learning methods are used to characterize the interaction between candidate (test) compounds and target polymers through machine learning approaches. Such characterization can report, for example, continuous or categorical activity indicators, PKa, or any other suitable metric to characterize the interaction between the candidate compound and the target polymer.
[0004] One drawback of vHTS machine learning methods is the way in which the machine learning model invoked in such methods interprets the pose between the compound and the binding site. The model represents the compound and the polymer separately, even though structural information about the two is provided. Therefore, any provided pose that allows for the discrimination of the polymer and the compound will give the same score. The model is insensitive to the pose. As shown in FIG. 19. This is exemplified in the Picasso problem, where a machine learning model such as a convolutional neural network may erroneously choose a pose that has all the correct components but is fundamentally wrong overall. As shown in FIG. 18. The left pose and the right pose both have the same parts, two eyes, two eyebrows, nose, lips, and the overall shape of the head. Thus, it proves difficult to teach a convolutional neural network that the left pose is correct. Thus, there is an inherent pose insensitivity in conventional vHTS machine learning methods. Such pose insensitivity can lead to incorrect or inaccurate characterization of the interaction between the test compound and the target polymer. For example, such pose insensitivity may cause a vHTS machine learning approach, which provides a categorical activity label for each compound in a screening library, to mislabel a certain percentage of compounds in a screening library.
[0005] Given the above background, what is needed in the art is a method for imparting pause sensitivity to vHTS machine learning methods. Summary of the Invention
[0006] The present disclosure addresses the problems identified in the Background Art by conditioning vHTS machine learning models to be pose-sensitive. Such models are trained on training compounds where the characterization of the interaction between each training compound and the target polymer is known. However, for each such training compound, the vHTS machine learning model is trained on both the positive pose of the training compound and the negative pose of the training compound, and such positive and negative poses are selected using an independent pose generation process. In this way, the vHTS machine learning model is trained to be pose-sensitive.
[0007] Thus, one aspect of the present disclosure is a computer system for providing characterization of an interaction between a test compound and a target polymer. The computer system comprises one or more processors and a memory addressable by the one or more processors. The memory stores at least one program for execution by the one or more processors. In some embodiments, the characterization of the interaction between the test compound and the target polymer is a binary activity score. In some embodiments, the target polymer is an assembly of proteins, polypeptides, polynucleic acids, polyribonucleic acids, polysaccharides, or any combination thereof. According to the present disclosure, a plurality of atomic coordinates of the target polymer is obtained. In some embodiments, the plurality of atomic coordinates includes atomic coordinates of at least 400 atoms.
[0008] In some embodiments, the plurality of atomic coordinates is a set of three-dimensional coordinates {x, ..., x} of a crystal structure of the target polymer resolved to a resolution of 2.5 Å or better, or a resolution of 3.3 Å or better. N}.
[0009] In some embodiments, the plurality of atomic coordinates of the target polymer comprises a collection of three-dimensional coordinates of the target polymer determined by nuclear magnetic resonance, neutron diffraction, or cryo-electron microscopy.
[0010] In some embodiments, the characterization of the interaction between the test compound and the target polymer is a binary score, and a first value of the binary score represents an IC of the test compound with respect to the target polymer that is above a first threshold. 50 , E.C. 50 , Kd, KI, or pKI, and the second value of the binary score represents the IC of the test compound with respect to the target polymer below the first threshold. 50 , E.C. 50 , Kd, KI, or pKI.
[0011] According to the present disclosure, a training data set is obtained that includes a respective electronic description of each training compound in a plurality of training compounds. In some embodiments, the plurality of training compounds includes at least 100 compounds. Each respective electronic description includes (i) a corresponding positive pose of the corresponding training compound with respect to a plurality of atomic space coordinates combined with a corresponding first positive interaction score, and (ii) a corresponding negative pose of the corresponding training compound with respect to a plurality of atomic space coordinates combined with a corresponding first negative interaction score.
[0012] In some embodiments, the corresponding positive scores of the corresponding positive poses of the corresponding training compounds for the target polymer are obtained by searching the corresponding positive voxel maps of the corresponding training compounds for the target polymer at the corresponding positive poses, expanding the corresponding positive voxel maps into corresponding positive vectors, and inputting the corresponding positive vectors into a neural network, thereby obtaining the corresponding positive scores of the corresponding positive poses. In some embodiments, the neural network includes more than 500 parameters. In some such embodiments, the corresponding positive vectors are first one-dimensional vectors.
[0013] In some embodiments, the corresponding negative scores of the corresponding negative poses of the corresponding training compound for the target polymer are obtained by searching the corresponding negative voxel maps of the corresponding training compound for the target polymer at the corresponding negative poses, expanding the corresponding negative voxel maps into corresponding negative vectors, and inputting the corresponding negative vectors into a neural network, thereby obtaining the corresponding negative scores of the corresponding negative poses of the corresponding training compound for the target polymer.
[0014] In some such embodiments, the corresponding negative vector is a second one-dimensional vector.
[0015] In some embodiments, the corresponding first positive interaction score and the corresponding first negative interaction score each represent a binding coefficient or an in silico quality score of the corresponding training compound for the target polymer.
[0016] In some embodiments, each training compound in the training dataset satisfies two or more, three or more, or all four of Lipinski's rule of five: (i) no more than five hydrogen bond donors, (ii) no more than ten hydrogen bond acceptors, (iii) molecular weight less than 500 daltons, and (iv) LogP less than 5.
[0017] In some embodiments, each training compound in the training dataset is an organic compound having a molecular weight of less than 500 Daltons, less than 1000 Daltons, less than 2000 Daltons, less than 4000 Daltons, less than 6000 Daltons, less than 8000 Daltons, less than 10000 Daltons, or less than 20000 Daltons.
[0018] According to the present disclosure, at least a first model is trained. The first model has a first plurality of parameters. In some embodiments, the first plurality of parameters includes more than 400 parameters. The training uses, for each corresponding training compound 46 in the plurality of training compounds, at least (i) the corresponding positive score of the corresponding positive pose of the corresponding training compound with respect to the target polymer as input to the first model against the corresponding first positive interaction score of the corresponding training compound with respect to the target polymer, and (ii) the corresponding negative score of the corresponding negative pose of the corresponding training compound with respect to the target polymer as input to the first model against the corresponding first negative interaction score of the corresponding training compound with respect to the target polymer, thereby adjusting the first plurality of parameters, and the output of at least the first model is used to provide, at least in part, a characterization of the interaction between the test compound and the target polymer.
[0019] In some embodiments, the first model is a first fully connected neural network.
[0020] In some embodiments, the training is a regression task in which the first plurality of parameters is tuned by backpropagation through an associated loss function. In such embodiments, the corresponding first positive interaction score is calculated using the formula B=N×A A is related to the corresponding first negative interaction score by A=B, where A is the corresponding positive interaction score, B is the corresponding negative interaction score, and N is a real number greater than zero and less than one (e.g., 0.90).
[0021] In some such embodiments, the associated loss function is a mean squared error loss function, a mean absolute error loss function, a Huber loss function, a Log-Cosh loss function, or a quantile loss function.
[0022] In some such embodiments, the corresponding first positive interaction score and the corresponding first negative interaction score each represent a binding coefficient, and the corresponding first positive interaction score is an in vitro measurement of the binding coefficient of the corresponding training compound to the target polymer.
[0023] In some such embodiments, the first positive interaction score is the IC 50 , E.C. 50 , Kd, KI, or pKI.
[0024] In some embodiments, each respective electronic description in the training data set further comprises a corresponding positive activity score of the corresponding positive pose of the corresponding training compound and a corresponding negative activity score of the corresponding negative pose of the corresponding training compound. In such an embodiment, training at least the first model further comprises jointly training a second model with the first model. The second model has a second plurality of parameters. Such training further uses, for each corresponding training compound in the plurality of training compounds, at least (iii) a corresponding positive score of the corresponding positive pose of the corresponding training compound with respect to the target polymer as an input to the second model relative to the corresponding positive activity score of the corresponding training compound, and (iv) a corresponding negative score of the corresponding negative pose of the corresponding training compound with respect to the target polymer as an input to the second model relative to the corresponding negative activity score of the corresponding training compound. In this manner, the second plurality of parameters is adjusted such that the second model provides, at least in part, an activity of the interaction between the test compound and the target polymer that is used together with the output of the first model to provide a characterization of the interaction between the test compound and the target polymer.
[0025] In some such embodiments, the second model is a second fully connected neural network.
[0026] In some embodiments, each respective electronic description in the training data set further comprises a corresponding positive activity score of the corresponding positive pose of the corresponding training compound and a corresponding negative activity score of the corresponding negative pose of the corresponding training compound. In such an embodiment, training at least the first model further comprises jointly training a second model with the first model, the second model having a second plurality of parameters. In such an embodiment, the training further uses, for each corresponding training compound in the plurality of training compounds, at least (iii) a corresponding positive score of the corresponding positive pose of the corresponding training compound with respect to the target polymer and the corresponding first positive interaction score as a binding input to the second model with respect to the corresponding positive activity score of the corresponding training compound, and (iv) a corresponding negative score of the corresponding negative pose of the corresponding training compound with respect to the target polymer and the corresponding first negative interaction score as a binding input to the second model with respect to the corresponding negative activity score of the corresponding training compound. In this way, the second plurality of parameters is adjusted so that the second model can be used, at least in part, with the output of the first model to provide a characterization of the interaction between the test compound and the target polymer.
[0027] In some such embodiments, the corresponding positive activity score is a first binary activity score and the corresponding negative activity score is a second binary activity score. In some embodiments, the corresponding first binary activity score is assigned a value of 1 and the corresponding second binary activity score is assigned a value of 0 based on the measured activity of the corresponding compound against the target polymer. In some such embodiments, training the first model is a regression task in which a first plurality of parameters are tuned by backpropagation through a first associated loss function, and training the second model is a classification task in which a second plurality of parameters are tuned by backpropagation through a second associated loss function. In some such embodiments, the corresponding first positive interaction score and the corresponding first negative interaction score each represent a binding coefficient or an in silico quality score of the corresponding training compound against the target polymer, and the corresponding positive activity score is a first binary activity score and the corresponding negative activity score is a second binary activity score. In some such embodiments, the first associated loss function is a mean squared error loss function, a mean absolute error loss function, a Huber loss function, a Log-Cosh loss function, or a quantile loss function, and the second associated loss function is a binary cross-entropy loss function, a hinge loss function, or a squared hinge loss function. In some such embodiments, the second model is a second fully connected neural network.
[0028] In some embodiments, each respective electronic description in the training data set further comprises a corresponding second positive interaction score for the corresponding positive pose of the corresponding training compound and a corresponding second negative interaction score for the corresponding negative pose of the corresponding training compound. In such embodiments, each respective electronic description in the training data set further comprises a corresponding positive activity score for the corresponding positive pose of the corresponding training compound and a corresponding negative activity score for the corresponding negative pose of the corresponding training compound. In such embodiments, training at least the first model further comprises jointly training a second model and a third model with the first model. The second model has a second plurality of parameters, and the third model has a third plurality of parameters. In such an embodiment, for each corresponding training compound in the plurality of training compounds, at least (iii) a corresponding positive score of a corresponding positive pose of the corresponding training compound with respect to the target polymer as input to the second model relative to a corresponding second positive interaction score of the corresponding training compound with respect to the target polymer, (iv) a corresponding negative score of a corresponding negative pose of the corresponding training compound with respect to the target polymer as input to the second model relative to a corresponding second negative interaction score of the corresponding training compound with respect to the target polymer, thereby adjusting the second plurality of parameters, and (v) a corresponding positive activity score of the corresponding training compound with respect to the target polymer, (vi) the output of the first model and the second model when the corresponding positive scores of the corresponding training compounds' corresponding positive poses on the target polymer and the corresponding positive scores of the corresponding training compounds' corresponding positive poses are input as binding inputs to the third model for the corresponding activity scores of the corresponding training compounds, and (vi) the output of the first model and the second model when the corresponding negative scores of the corresponding training compounds' corresponding negative poses on the target polymer and the corresponding negative scores of the corresponding training compounds' corresponding negative poses are input as binding inputs to the third model for the corresponding negative activity scores of the corresponding training compounds, thereby adjusting a third plurality of parameters of the third model. In such an embodiment, the output of the third model provides a characterization of the interaction between the test compound and the target polymer.In some such embodiments, the second model is a second fully connected neural network and the third model is a third fully connected neural network. In some such embodiments, the corresponding positive activity score is a first binary activity score and the corresponding negative activity score is a second binary activity score. In some embodiments, the corresponding first binary activity score is assigned a value of 1 and the corresponding second binary activity score is assigned a value of 0 based on the measured activity of the corresponding compound against the target polymer. In some such embodiments, training the first model is a first regression task where the first plurality of parameters are tuned by backpropagation through a first associated loss function, training the second model is a second regression task where the second plurality of parameters are tuned by backpropagation through a second associated loss function, and training the third model is a classification task where the third plurality of parameters are tuned by backpropagation through a third associated loss function. In some such embodiments, the corresponding first positive interaction score and the corresponding first negative interaction score each represent an in silico quality score of the corresponding training compound for the target polymer, the corresponding second positive interaction score and the corresponding second negative interaction score each represent a binding coefficient of the corresponding training compound for the target polymer, the corresponding positive activity score is a first binary activity score, and the corresponding negative activity score is a second binary activity score. In some such embodiments, the first associated loss function is a mean squared error loss function, a mean absolute error loss function, a Huber loss function, a Log-Cosh loss function, or a quantile loss function, the second associated loss function is a mean squared error loss function, a mean absolute error loss function, a Huber loss function, a Log-Cosh loss function, or a quantile loss function, and the third associated loss function is a binary cross-entropy loss function, a hinge loss function, a squared hinge loss function, or any other loss function described herein for use as the first or second associated loss function.
[0029] Another aspect of the disclosure provides a method for characterizing an interaction between a test compound and a target polymer, the method comprising: acquiring a plurality of atomic coordinates of the target polymer in a computer system comprising a memory. In some embodiments, the plurality of atomic coordinates comprises atomic coordinates of at least 400 atoms. A training data set is acquired. The training data set comprises a respective electronic description of each training compound in the plurality of training compounds. In some embodiments, the plurality of training compounds comprises at least 100 compounds. Each respective electronic description comprises: (i) a corresponding positive pose of the corresponding training compound for the plurality of atomic space coordinates combined with a corresponding first positive interaction score; and (ii) a corresponding negative pose of the corresponding training compound for the plurality of atomic space coordinates combined with a corresponding first negative interaction score. At least a first model is trained, the first model having a first plurality of parameters. In some embodiments, the first plurality of parameters comprises more than 400 parameters. For each corresponding training compound in the plurality of training compounds, the training uses at least (i) the corresponding positive score of the corresponding positive pose of the corresponding training compound with respect to the target polymer as input to the first model for the corresponding first positive interaction score of the corresponding training compound with respect to the target polymer, and (ii) the corresponding negative score of the corresponding negative pose of the corresponding training compound with respect to the target polymer as input to the first model for the corresponding first negative interaction score of the corresponding training compound with respect to the target polymer. In this manner, the first plurality of parameters are adjusted. After training, the output of at least the first model is used, at least in part, to provide a characterization of the interaction between the test compound and the target polymer.
[0030] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores instructions that, when executed by a computer system, cause the computer system to execute a method for characterizing an interaction between a test compound and a target polymer according to the method. The method includes obtaining a plurality of atomic coordinates of the target polymer. In some embodiments, the plurality of atomic coordinates includes atomic coordinates of at least 400 atoms. A training data set is obtained that includes a respective electronic description of each training compound in the plurality of training compounds. The plurality of training compounds includes at least 100 compounds. Each respective electronic description includes (i) a corresponding positive pose of the corresponding training compound for the plurality of atomic space coordinates combined with a corresponding first positive interaction score, and (ii) a corresponding negative pose of the corresponding training compound for the plurality of atomic space coordinates combined with a corresponding first negative interaction score. In the method, at least a first model is trained. The first model has a first plurality of parameters. In some embodiments, the first plurality of parameters includes more than 400 parameters. For each corresponding training compound in the plurality of training compounds, the training uses at least (i) the corresponding positive scores of the corresponding positive poses of the corresponding training compound with respect to the target polymer as inputs to the first model relative to the corresponding first positive interaction scores of the corresponding training compound with respect to the target polymer, and (ii) the corresponding negative scores of the corresponding negative poses of the corresponding training compound with respect to the target polymer as inputs to the first model relative to the corresponding first negative interaction scores of the corresponding training compound with respect to the target polymer, thereby adjusting the first plurality of parameters. The output of at least the first model is used, at least in part, to provide a characterization of the interaction between the test compound and the target polymer.
[0031] In the drawings, embodiments of the disclosed system and method are shown by way of example, and it is to be expressly understood that the description and drawings are for illustrative purposes only and are intended as an aid to understanding and are not intended as a definition of the limits of the disclosed system and method. [Brief description of the drawings]
[0032] [Figure 1] 1 illustrates a computer system according to some embodiments of the present disclosure. [Figure 2A] 1 illustrates a method for characterizing an interaction between a test compound and a target polymer according to some embodiments of the present disclosure. [Figure 2B] 1 illustrates a method for characterizing an interaction between a test compound and a target polymer according to some embodiments of the present disclosure. [Figure 2C] 1 illustrates a method for characterizing an interaction between a test compound and a target polymer according to some embodiments of the present disclosure. [Figure 2D] 1 illustrates a method for characterizing an interaction between a test compound and a target polymer according to some embodiments of the present disclosure. [Figure 2E] 1 illustrates a method for characterizing an interaction between a test compound and a target polymer according to some embodiments of the present disclosure. [Figure 2F] 1 illustrates a method for characterizing an interaction between a test compound and a target polymer according to some embodiments of the present disclosure. [Figure 2G] 1 illustrates a method for characterizing an interaction between a test compound and a target polymer according to some embodiments of the present disclosure. [Figure 2H] 1 illustrates a method for characterizing an interaction between a test compound and a target polymer according to some embodiments of the present disclosure. [Figure 2I] 1 illustrates a method for characterizing an interaction between a test compound and a target polymer according to some embodiments of the present disclosure. [Diagram 3] FIG. 1 is a schematic diagram of an exemplary training compound in pose relative to a target polymer, according to some embodiments of the present disclosure. [Figure 4] 1 is a schematic diagram of a geometric representation of input features in the form of a three-dimensional grid of voxels, according to some embodiments of the present disclosure. [Diagram 5]FIG. 2 is a diagram of a compound encoded onto a two-dimensional grid of voxels, according to some embodiments of the present disclosure. [Figure 6] FIG. 2 is a diagram of a compound encoded onto a two-dimensional grid of voxels, according to some embodiments of the present disclosure. [Figure 7] FIG. 7 is a diagram of the visualization of FIG. 6 with voxels numbered, according to some embodiments of the present disclosure. [Figure 8] 1A-1C are schematic diagrams of geometric representations of input features in the form of coordinate positions of atom centers, according to some embodiments of the present disclosure; [Figure 9A] According to one embodiment of the present disclosure, there is provided a system for characterizing an interaction between a test compound and a target polymer, wherein the characterization is a compound binding mode score, and the system is trained using combined positive and negative poses to train the compound. [Figure 9B] According to one embodiment of the present disclosure, there is provided a system for characterizing the interaction between a test compound and a target polymer, the characterization being activity and compound binding mode score, and the system is trained using combined positive and negative poses to train the compound. [Figure 9C] A system for characterizing interactions between a test compound and a target polymer according to one embodiment of the present disclosure, wherein the characterization is active, the system is trained using combined positive and negative poses to train the compound, and the final output model is conditioned against two different pose quality models. [Figure 10] According to one embodiment of the present disclosure, there is provided a system for characterizing the interaction between a test compound and a target polymer, wherein the characterization is (i) binary discrete activity and (ii) pKi, and the system is trained using combined positive and negative poses for training the compound. [Figure 11]According to one embodiment of the present disclosure, a system for characterizing the interaction between a test compound and a target polymer, wherein the characterization is pKi, the pKi is partially conditioned for activity, and the system is trained using combined positive and negative pauses to train the compound. [Figure 12] According to one embodiment of the present disclosure, a system for characterizing an interaction between a test compound and a target polymer, the characterization being activity, the activity being conditioned, in part, on both pKi and pause quality score, and the system being trained using combined positive and negative pauses to train the compound. [Figure 13] According to one embodiment of the present disclosure, a system for characterizing the interaction between a test compound and a target polymer, the characterization being activity, the activity being conditioned, in part, on both pKi and compound binding mode scores, and the system being trained using combined positive and negative poses to train the compound. [Figure 14] According to one embodiment of the present disclosure, there is provided a system for characterizing the interaction between a test compound and a target polymer, the characterization being activity and two different compound binding mode scores, and the system is trained using combined positive and negative poses to train the compound. [Figure 15] According to one embodiment of the present disclosure, a system for characterizing the interaction between a test compound and a target polymer, the characterization being activity, two different compound binding mode scores and pKi, and the system is trained using combined positive and negative poses to train the compound. [Figure 16A] According to one embodiment of the present disclosure, a system for characterizing the interaction between a test compound and a target polymer, the characterization being activity, the activity being conditioned, in part, on pKi and binding mode scores, and the system being trained using combined positive and negative poses to train the compound. [Figure 16B]According to one embodiment of the present disclosure, a system for characterizing the interaction between a test compound and a target polymer, the characterization being activity, the activity being conditioned, in part, on pKi and two different binding mode scores, and the system being trained using combined positive and negative poses to train the compound. [Figure 17] 1 is a depiction of applying multiple function computation elements (g1, g2, ...) to voxel inputs (x1, x2, ..., x100) and jointly composing the function computation element outputs using g(), in accordance with some embodiments of the present disclosure. [Figure 18] 1 illustrates the insensitivity faced by machine learning models in characterizing the pose of compounds with respect to target polymers according to the prior art. [Figure 19] Demonstrating the insensitivity of traditional machine learning models to the quality of the compound-polymer pose, as shown in the figure, the best possible pose receives the same score as a bad pose from the machine learning model, and an unlikely pose receives the same score as the best possible pose from the machine learning model. [Figure 20] Human ZAP70 protein with annotated ATP binding site (grey), allosteric site (red), and control binding site in the SH2 domain (blue). Used PDB ID: 2ozo. [Figure 21] 1 illustrates various benchmarks of receiver operating curve AUC performance according to embodiments of the present disclosure. [Figure 22] 1 shows a Picasso problem experiment in which 105 diverse compounds (labeled 0, non-binders) mixed with approximately 300 kinase inhibitors (labeled 1, binders) were docked and scored with three binding sites: i) the ATP binding site, ii) the allosteric binding site, and iii) the binding site in the SH2 domain, according to an embodiment of the disclosure. [Figure 23] 1 shows the median probability drop between good poses and bad (left panel) or impossible poses (right panel) according to an embodiment of the present disclosure. [Figure 24]1 illustrates an activity task conditioned on PoseRanker and Vina scores according to an embodiment of the present disclosure.
[0033] Like reference numbers refer to corresponding parts throughout the drawings. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0034] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.
[0035] It should also be understood that terms such as first, second, etc. may be used herein to describe various elements, but these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first object may be referred to as a second object, and similarly, a second object may be referred to as a first object, without departing from the scope of the present invention. A first object and a second object are both objects, but are not the same object.
[0036] The terms used in this disclosure are only for the purpose of describing particular embodiments and are not intended to limit the invention. When used in the description of the invention and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. The term "and / or" as used herein will also be understood to refer to and include any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms "comprises" and / or "comprising" as used herein specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0037] As used herein, the term "if" may be interpreted to mean "when" or "upon" or "upon determining" that a stated antecedent condition is true or "in accordance with determining" or "depending on the context," depending on the context. Similarly, the phrase "when determined" or "when [a stated condition or event] is detected" may be interpreted to mean "upon determining" or "in response to determining" or "upon detection of [a stated condition or event]," or "upon detection of [a stated condition or event]," depending on the context.
[0038] The present disclosure provides a system and method for characterizing interactions between test compounds and polymers using a polymer coordinate and a training dataset of compounds. Each respective training compound has a positive pose with respect to the target polymer coordinate with a positive interaction score. At least some of the respective training compounds in the training dataset of compounds also have a negative pose of the respective training compound with respect to the target polymer coordinate and a negative interaction score. A model is trained for each respective compound in the training set by applying at least (i) the positive score of the positive pose as an input to the model to the compound's positive interaction score, and (ii) in some cases, the negative score of the negative pose as an input to the model to the compound's negative interaction score, thereby adjusting the parameters of the model, and at least some of the compounds in the training set have both positive and negative poses. In some embodiments, at least 5%, 10%, 20%, 50%, or 70% of the compounds in the training set have both positive and negative poses, while the remaining compounds in the training set have only positive poses. In some embodiments, all of the compounds in the training set have both positive and negative poses.
[0039] In some embodiments, the positive scores for the positive poses are obtained by forming a corresponding positive voxel map for each training compound at each positive pose for the polymer. In some embodiments, the corresponding positive voxel map is vectorized and fed to the neural network. In some embodiments, the voxel map is input to the neural network without vectorization.
[0040] In some embodiments, the neural network is a convolutional neural network. In some such embodiments, the convolutional neural network includes an input layer, a plurality of individually weighted convolutional layers, and an output scorer. The convolutional layers include an initial layer and a final layer. In response to an input, the input layer provides a value to the initial convolutional layer. Each respective convolutional layer other than the final convolutional layer provides an intermediate value as a function of the respective convolutional layer weight and the respective convolutional layer input value to another one of the convolutional layers. The final convolutional layer provides a value to the scorer as a function of the final layer weight and the input value. In this way, the scorer scores the positive poses of each compound and arrives at a positive score for the positive poses of each compound.
[0041] In some embodiments, the negative scores of the negative poses are obtained by forming corresponding negative voxel maps of each training compound at each negative pose of the polymer. In some embodiments, the corresponding negative voxel maps are vectorized and fed into a neural network (e.g., a convolutional neural network) as described above. In some embodiments, the voxel maps are input into the neural network without vectorization. In this way, the scorer scores the negative poses of each compound to arrive at a negative score for the negative poses of each compound.
[0042] Once the model is trained on the training compound, it can be used to characterize the interaction between the test compound and the polymer. In some embodiments, in response to the positive poses of the test compound and the target polymer being input to the neural network, a score of the positive pose is provided by the neural network and the second (or third, fourth, ... xth) model. Upon conditioning through the buried layer, the score of the positive pose provided by the neural network serves as an input to the trained model, thereby providing a characterization of the interaction between the test compound and the polymer.
[0043] 1 shows a computer system 100 for characterizing interactions between test compounds and target polymers, which can be used, for example, as a binding affinity prediction system to generate accurate predictions regarding the binding affinity of one or more test compounds with a target polymer.
[0044] With reference to Figure 1, in a typical embodiment, computer system 100 comprises one or more computers. For purposes of illustration in Figure 1, computer system 100 is represented as a single computer that includes all of the disclosed functionality of computer system 100. However, the present disclosure is not so limited. The functionality of computer system 100 may be distributed across any number of networked computers and / or may reside on each of several networked computers and / or virtual machines. Those skilled in the art will appreciate that a variety of different computer topologies are possible for computer system 100, and all such topologies are within the scope of the present disclosure.
[0045] With the above in mind and referring to Figure 1, computer system 100 comprises one or more processing units (CPUs) 59, a network or other communication interface 84, a user interface 78 (e.g., including a display 82 and an optional keyboard 80 or other form of input device), memory 92 (e.g., random access memory), one or more magnetic disk storage and / or persistent devices 90 optionally accessed by one or more controllers 88, one or more communication buses 12 for interconnecting the aforementioned components, and a power supply 79 for powering the aforementioned components. Data in memory 92 may be seamlessly shared with non-volatile memory 90 using well-known computing techniques such as caching. Memory 92 and / or memory 90 may include mass storage located remotely relative to central processing unit(s) 59. In other words, memory 92 and / or some of the data stored in memory 90 may actually be hosted on a computer that is external to computer system 100, but can be electronically accessed by computer system 100 via the Internet, an intranet, or other form of network or electronic cable using network interface 84. In some embodiments, computer system 100 utilizes neural networks executed from memory 52 associated with one or more of the graphic processing units 50 to improve system speed and performance. In some alternative embodiments, computer system 100 utilizes neural networks executed from memory 92 rather than memory associated with the graphic processing units 50.
[0046] The memory 92 and / or optionally the memory 52 of the computer system 100 may include: an optional operating system 34 that includes procedures for handling various basic system services; and a spatial data evaluation module 36 for characterizing interactions between test compounds and target polymers; data 38 for the target polymer, including structural data (e.g., a plurality of atomic spatial coordinates 40 of the target polymer) and / or, optionally, active site information 42 of the target polymer; a training data set 44 including a respective electronic description 46 of each training compound in a plurality of training compounds, each respective electronic description in at least a subset of the training data set 44 including (i) a corresponding positive pose 48 of the corresponding training compound with respect to the plurality of atomic spatial coordinates 40 combined with a corresponding first positive interaction score 50, and (ii) a corresponding negative pose 60 of the corresponding training compound with respect to the plurality of atomic spatial coordinates 40 combined with a corresponding first negative interaction score 62; a first model 72 including a first plurality of parameters 73, the output of the first model being used, at least in part, to provide a characterization of the interaction between the test compound and the target polymer; An assessment module 20 for applying a neural network 24 to the spatial data (e.g., for applying the neural network to test or training compounds docked to the target polymer); one or more (optionally) vectorized representations of the voxel map; and A neural network 24, optionally including an input layer 26, optionally including one or more convolutional layers 28, and including a terminal scorer 30; a second model 74 including a second plurality of parameters 75, the output of the second model being used, at least in part, to (i) provide a characterization of the interaction between the test compound and the target polymer, and / or (ii) to condition the first model; and ● optionally a third model 76 including a third plurality of parameters 77, the output of the third model being used, at least in part, to (i) provide a characterization of the interaction between the test compound and the target polymer, and / or (ii) to condition the first model and / or the second model; and ●Optionally, storing any number of additional xth models, each such additional xth model including a corresponding plurality of parameters, and the output of the xth model being used, at least in part, to (i) provide a characterization of the interaction between the test compound and the target polymer, and / or (ii) an xth model used to condition any other single models and / or groups of models.
[0047] In some implementations, one or more of the above identified data elements or modules of computer system 100 are stored in one or more of the aforementioned memory devices and correspond to sets of instructions for performing the functions described above. The above identified data, modules or programs (e.g., sets of instructions) need not be implemented as separate software programs, procedures, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various implementations. In some implementations, memory 92 and / or 90 (and optionally 52) optionally stores a subset of the above identified modules and data structures. Additionally, in some embodiments, memory 92 and / or 90 (and optionally 52) stores additional modules and data structures not described above.
[0048] Disclosed herein is a system for characterizing interactions between test compounds and target polymers, and a method for carrying out such characterization is detailed with reference to FIG. 2 and discussed below.
[0049] Block 200. Referring to block 200 of FIG. 2A, a computer system 100 is disclosed that provides characterization of interactions between a test compound and a target polymer 38. As discussed above in conjunction with FIG. 1, the computer system comprises one or more processors 74 and memory 90 / 92 addressable by the one or more processors. The memory stores at least one program for execution by the one or more processors. The remainder of FIG. 2 details features of the at least one program, including training of the computer system and use of the trained computer system.
[0050] Blocks 202-204. Referring to block 202 of FIG. 2A, in some embodiments, once trained on the reference compound, the spatial data evaluation module 36 can characterize the interaction between the test compound and the target polymer 38. In some such embodiments, this characterization is a discrete (e.g., discrete binary) activity score. In other words, the characterization is categorical. For example, in some embodiments, the characterization is discrete binary, and the computer system provides one value, e.g., "1", when the test compound is determined to be active against the target polymer by in silico methods implemented in the spatial data evaluation module 36 and discussed in more detail below, and provides another value, e.g., "0", when the test compound is determined to be inactive against the target polymer.
[0051] In some embodiments, the characterization is a discrete scale other than binary. For example, in some embodiments, the characterization provides a first value, e.g., "0", when the test compound is determined by in silico methods implemented in spatial data evaluation module 36 and discussed in more detail below to have activity below a first threshold, a second value, e.g., "1", when the test compound is determined to have activity between the first and second thresholds, and a third value, e.g., "2", when the test compound is determined to have activity above the second threshold. In such embodiments, the first and second thresholds are predetermined and are selected to have values that are constant for a particular experiment (e.g., a particular evaluation of a particular database, set, or collection of test compounds against a particular target polymer) and that prove useful in identifying suitable test compounds from a particular database, set, or collection of test compounds for activity against the test polymer. For example, in some embodiments, any of the thresholds disclosed herein are designed to identify 0.1 percent or less, 0.5 percent or less, 1 percent or less, 2 percent or less, 5 percent or less, 10 percent or less, 20 percent or less, or 50 percent or less of a database of test compounds as active against a target polymer, and the database of test compounds may be greater than 100 compounds, greater than 1000 compounds, greater than 10,000 compounds, greater than 100,000 compounds, greater than 1×10 6 Compounds, 10x10 6 The compound may include one or more compounds.
[0052] In an alternative embodiment, once trained against the reference compound, the spatial data evaluation module 36 can characterize the interaction between the test compound and the target polymer 38 as an activity on a continuous scale. That is, the spatial data evaluation module 36 provides a continuous scale numerical value indicative of the activity of the test compound against the target polymer. The continuous scale activity value is useful, for example, for comparing the activity of each test compound in a database of test compounds against the target polymer to which it was assigned by the trained spatial data evaluation module 36.
[0053] Referring to block 204, the disclosed systems and methods are not limited to characterizing the interaction between the test compound and the target polymer 38 as a continuous or discrete scale activity. In an alternative embodiment, the spatial data evaluation module 36 may, in effect, characterize the interaction between the test compound and the target polymer, once trained on the reference compound, as the IC of the test compound against the target polymer on a continuous or discrete (categorical) scale. 50 , E.C. 50 , Kd, KI, or pKI.
[0054] Although a binary discrete scale and a discrete scale having three possible outcomes have been identified, the present disclosure is not limited to these two examples of discrete scales for characterizing the interaction between a test compound and a target polymer 38. Indeed, any discrete scale can be used for characterizing the interaction between a test compound and a target polymer 38, including, as non-limiting examples, discrete scales having 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 different outcomes.
[0055] Block 206. Referring to block 204 of FIG. 2A, in some embodiments, the target polymer 38 is an assembly of proteins, polypeptides, polynucleic acids, polyribonucleic acids, polysaccharides, metalloproteins, or any combination thereof. In some embodiments, the target polymer 38 is a macromolecule composed of repeating residues. In some embodiments, the target polymer 38 is a natural material. In some embodiments, the target polymer 38 is a synthetic material. In some embodiments, the target polymer 38 is an elastomer, shellac, amber, natural or synthetic rubber, cellulose, bakelite, nylon, polystyrene, polyethylene, polypropylene, polyacrylonitrile, polyethylene glycol, or a polysaccharide.
[0056] In some embodiments, the target polymer 38 is a heteropolymer (copolymer). A copolymer is a polymer derived from two (or more) monomer species, as opposed to a homopolymer, where only one monomer is used. Copolymerization refers to the method used to chemically synthesize a copolymer. Examples of copolymers include, but are not limited to, ABS plastic, SBR, nitrile rubber, styrene acrylonitrile, styrene-isoprene-styrene (SIS), and ethylene vinyl acetate. Because copolymers contain at least two types of building blocks (structural units, or particles as well), copolymers can be classified based on how these units are arranged along the chain. These include alternating copolymers, in which A and B units regularly alternate. See, for example, Jenkins, 1996, “Glossary of Basic Terms in Polymer Science,” Pure Appl. Chem. 68(12): 2287-2311, which is incorporated herein by reference in its entirety. Additional examples of copolymers include copolymers that contain repeating sequences (e.g., (ABABBAAAABBB) n ) are periodic copolymers with A and B units arranged in a cyclic manner. Additional examples of copolymers are statistical copolymers, in which the arrangement of monomer residues in the copolymer follows statistical laws. See, for example, Painter, 1997, Fundamentals of Polymer Science, CRC Press, 1997, p14, which is incorporated herein by reference in its entirety. Yet another example of a copolymer that can be evaluated using the disclosed systems and methods is a block copolymer, comprising two or more homopolymer subunits linked by covalent bonds. The linking of the homopolymer subunits may require an intermediate non-repeating subunit, known as a junction block. Block copolymers with two or three distinct blocks are called diblock and triblock copolymers, respectively.
[0057] In some embodiments, the target polymer 38 comprises 50 or more, 100 or more, 150 or more, 200 or more, 300 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more atoms.
[0058] In some embodiments, the target polymer 38 is actually a plurality of polymers (e.g., 2 or more, 3 or more, 10 or more, 100 or more, 1000 or more, or 5000 or more polymers), where each polymer in the plurality does not all have the same molecular weight. In some such embodiments, the target polymers 38 in the plurality of polymers share at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% sequence identity and fall into weight ranges with a corresponding chain length distribution. In some embodiments, the target polymer 38 is a branched polymer molecule that includes a backbone with one or more substituent side chains or branches. Types of branched polymers include, but are not limited to, star polymers, comb polymers, brush polymers, dendronized polymers, ladders, and dendrimers. See, e.g., Rubinstein et al., 2003, Polymer physics, Oxford; New York: Oxford University Press. p. 6, which is incorporated herein by reference in its entirety.
[0059] In some embodiments, the target polymer is a polypeptide. As used herein, the term "polypeptide" refers to two or more amino acids or residues linked by a peptide bond. The terms "polypeptide" and "protein" are used interchangeably herein and include oligopeptides and peptides. An "amino acid", "residue" or "peptide" refers to any of the 20 standard structural units of proteins, as known in the art, including imino acids such as proline and hydroxyproline. The names of amino acid isomers may include D, L, R, and S. The definition of amino acid includes unnatural amino acids. Thus, selenocysteine, pyrrolysine, lanthionine, 2-aminoisobutyric acid, gamma-aminobutyric acid, dehydroalanine, ornithine, citrulline, and homocysteine, as non-limiting examples, are all considered amino acids. Other variants or analogs of amino acids are known in the art. Thus, a polypeptide may include synthetic peptide analog structures, such as peptoids. See Simon et al., 1992, Proceedings of the National Academy of Sciences USA, 89, 9367, which is incorporated herein by reference in its entirety. See also Chin et al., 2003, Science 301, 964, and Chin et al., 2003, Chemistry & Biology 10, 511, each of which is incorporated herein by reference in its entirety.
[0060] The target polymers 38 evaluated according to some embodiments of the disclosed systems and methods may also have any number of post-translational modifications. Thus, target polymers 38 include those polymers that have been modified by acylation, alkylation, amidation, biotinylation, formylation, gamma-carboxylation, glutamylation, glycosylation, glycylation, hydroxylation, iodination, isoprenylation, lipoylation, cofactor addition (e.g., heme, flavin, metals, etc.), addition of nucleosides and their derivatives, oxidation, reduction, pegylation, phosphatidylinositol addition, phosphopantetheinylation, phosphorylation, pyroglutamate formation, racemization, addition of amino acids by tRNA (e.g., arginylation), sulfation, selenoylation, ISGylation, sumoylation, ubiquitination, chemical modifications (e.g., citrullination and deamidation), and processing by other enzymes (e.g., proteases, phosphatases, and kinases). Other types of post-translational modifications are known in the art and are within the scope of the target polymer 38 of the present disclosure.
[0061] In some embodiments, the target polymer 38 is a surfactant. A surfactant is a compound that reduces the surface tension of a liquid, the interfacial tension between two liquids, or the interfacial tension between a liquid and a solid. Surfactants can function as detergents, wetting agents, emulsifiers, foaming agents, and dispersants. Surfactants are usually organic compounds that are amphiphilic, meaning that they contain both hydrophobic groups (their tails) and hydrophilic groups (their heads). Thus, surfactant molecules contain both water-insoluble (or oil-soluble) and water-soluble components. Surfactant molecules diffuse into water and, when water is mixed with oil, adsorb to the air-water interface or the oil-water interface. The insoluble hydrophobic groups can extend from the bulk water phase into the air or oil phase, while the water-soluble head groups remain in the water phase. Such alignment of surfactant molecules at the surface modifies the surface properties of the water at the water / air or water / oil interface.
[0062] Examples of ionic surfactants include ionic surfactants such as anionic, cationic, or zwitterionic (ampotelic) surfactants, in some embodiments, the target entity 58 is a reverse micelle or a liposome.
[0063] In some embodiments, the target polymer 38 is a fullerene. A fullerene is any molecule composed entirely of carbon in the form of a hollow sphere, ellipsoid, or tube. Spherical fullerenes are also called buckyballs and resemble the balls used in soccer. Cylindrical ones are called carbon nanotubes or buckytubes. Fullerenes are similar in structure to graphite, which consists of stacked graphene sheets of interlocking hexagonal rings, but may also contain pentagonal (or sometimes heptagonal) rings.
[0064] Blocks 208-212. Referring to block 208 of FIG. 2A, a plurality of atomic coordinates 40 of the target polymer 38 are obtained. In some embodiments, the plurality of atomic coordinates includes atomic coordinates of at least 400 atoms of the target polymer. In some embodiments, the plurality of atomic coordinates includes atomic coordinates of at least 25 atoms, at least 50 atoms, at least 100 atoms, at least 200 atoms, at least 300 atoms, at least 400 atoms, at least 1000 atoms, at least 2000 atoms, or at least 5000 atoms of the target polymer. In some embodiments, only coordinates of active sites of the target polymer 38 to which the ligand is expected to bind the target polymer are obtained. Referring to block 210, in some embodiments, the plurality of atomic coordinates includes a set of three-dimensional coordinates {x1, ..., x2, ... N Referring to block 212, in some embodiments, the plurality of atomic coordinates of the target polymer comprises a collection of three-dimensional coordinates of the target polymer determined by nuclear magnetic resonance, neutron diffraction, or cryo-electron microscopy.
[0065] In some embodiments, the plurality of atomic coordinates is a set of three-dimensional coordinates {x1, ..., x2, ... N}.
[0066] In some embodiments, the plurality of atomic coordinates of the target polymer 38 is a collection of 10 or more, 20 or more, 30 or more, or more three-dimensional coordinates of the target polymer determined by nuclear magnetic resonance, the collection having a backbone root mean square deviation (RMSD) of 1.0 Å or more, 0.9 Å or more, 0.8 Å or more, 0.7 Å or more, 0.6 Å or more, 0.5 Å or more, 0.4 Å or more, 0.3 Å or more, or 0.2 Å or more. In some embodiments, the plurality of atomic coordinates are determined by neutron diffraction or cryo-electron microscopy.
[0067] In some embodiments, the target polymer 38 includes two different types of polymers, such as a nucleic acid bound to a polypeptide. In some embodiments, the natural target polymer includes two polypeptides bound to each other. In some embodiments, the natural target polymer under study includes one or more metal ions (e.g., a metalloproteinase having one or more zinc atoms). In such cases, the metal ions and / or small organic molecules may be included in the atomic coordinates 40 of the target polymer.
[0068] In some embodiments, target polymer 38 is a polymer having 10 or more, 20 or more, 30 or more, 50 or more, 100 or more, 100-1000, or less than 500 residues present in the target polymer.
[0069] In some embodiments, the atomic coordinates of the target polymer 38 are determined using modeling methods such as ab initio methods, density functional methods, semi-empirical and empirical methods, molecular mechanics, chemical mechanics, or molecular mechanics.
[0070] In some embodiments, the atomic coordinates 40 are represented by the Cartesian coordinates of the centers of atoms comprising the target polymer 38. In some alternative embodiments, the spatial coordinates 40 of the target polymer 38 are represented by the electron density of the target polymer as measured, for example, by X-ray crystallography. For example, in some embodiments, the atomic coordinates 40 are represented by the 2F observed -F calculated Includes electron density maps, F observed is the amplitude of the observed structure factor of the target polymer, and Fc is the amplitude of the structure factor calculated from the calculated atomic coordinates of the target polymer 38.
[0071] In various other embodiments, the atomic coordinates 40 of the target polymer 38 are obtained in accordance with block 206 from a wide variety of sources, including, but not limited to, co-complexes interpreted from structural ensembles generated by solution NMR, X-ray crystallography, neutron diffraction, cryo-electron microscopy, sampling from computational simulations, homology modeling, sampling from rotamer libraries, or any combination thereof.
[0072] Block 214. Referring to block 214 of FIG. 2B, a training data set 44 is obtained that includes a respective electronic description of each training compound 46 in the plurality of training compounds. In some embodiments, the plurality of training compounds includes at least 50, 100, 200, 1000, 5000, 10,000, 50,000, 100,000, 1×10 6 pieces, 1×10 7 Pieces or 1 x 10 8training compounds. Each electronic description of each training compound 46 in at least a subset of the training dataset includes (i) a corresponding positive pose 48 of the corresponding training compound 46 for the plurality of atomic spatial coordinates coupled with a corresponding first positive interaction score 50, and (ii) a corresponding negative pose 60 of the corresponding training compound for the plurality of atomic spatial coordinates coupled with a corresponding first negative interaction score 62. FIG. 3 shows the positive poses 48 of the training compounds 46 in the active site of the target polymer 38. In some embodiments, some of the training compounds 46 do not have a negative pose 60 and do not have a corresponding first negative interaction score 62. In some embodiments, some of the training compounds 46 do not have a positive pose 48 and do not have a corresponding first positive interaction score 50. In some embodiments, all of the training compounds 46 have both a positive pose and a negative pose, and both a corresponding first positive interaction score and a first negative interaction score.
[0073] In some embodiments, the target polymer 38 is a polymer with an active site, and the positive and negative poses are obtained by docking the training compound into the active site of the polymer. In some embodiments, the training compound is docked into the target polymer 38 multiple times to form multiple poses. In some embodiments, each training compound is docked into the target compound 38 two, three, four, five or more times, ten or more times, fifty or more times, one hundred or more times, or one thousand or more times. Each such docking represents a different pose of the training compound docked into the target polymer 38. In some embodiments, the target polymer 38 is a polymer with an active site, and each training compound is docked into the active site in each of a plurality of different ways, and each such way represents a different pose. Many of these poses are expected to be incorrect, meaning that such poses do not represent the true interaction between the training compounds and the naturally occurring target polymer.
[0074] In some embodiments, each pose of the training compounds is determined by AutoDock Vina. See Trott and Olson, “AutoDock Vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization and multithreading,” Journal of Computational Chemistry 31 (2010) 455-461. In such embodiments, for each training compound, the pose that received the best score by AutoDock Vina is assigned the positive pose 48, and the pose that received the worst score by AutoDock Vina is assigned the negative pose 60. In some embodiments, different docking programs are used to determine the positive pose 48 and the negative pose 60 for each training compound.For example, in some embodiments, Quick Vina 2 (Alhossary et al., 2015, “Fast, accurate, and reliable molecular docking with QuickVina,” Bioinformatics 31:13, pp. 2214-2216), VinaLC (Zhang et al., 2013, “Message Passing Interface and Multithreading Hybrid for Parallel Molecular Docking of Large Databases on Petascale High Performance Computing Machines,” J. Comput. Chem. DOI:10.1002 / jcc.23214), Smina (Koes et al., 2013, “Lessons learned in empirical scoring with smina from the CSAR 2011 benchmarking exercise,” Journal of chemical information and modeling 53:8, pp. 1893-1904), or Cuina (Morrison et al., “Efficient GPU Implementation of AutoDock Vina,”COMP poster 3432389) is used.
[0075] In some embodiments, the positive poses 48 are a positive ensemble of poses and the negative poses 60 are a negative ensemble of poses. For example, in some embodiments, the positive poses 48 are a corresponding first ensemble of 2-500 structurally similar poses and the negative poses 48 are a corresponding second ensemble of 2-500 structurally similar poses, where the corresponding first ensemble has a better overall docking score than the corresponding second ensemble. Methods for obtaining such ensembles are disclosed in Stafford et al., 2019, “Modeling protein flexibility with conformational sampling improves ligand pose and bioactivity prediction,” Abstracts of Papers of the American Chemical Society, Volume 258, which is incorporated herein by reference. In some embodiments, each corresponding first ensemble (collectively representing the positive poses 48) is 2-30, 2-20, 2-10, more than 100, 2-1000 structurally similar poses. In some embodiments, each corresponding second collection (collectively representing negative poses 48) is between 2 and 30, between 2 and 20, between 2 and 10, greater than 100, or between 2 and 1000 structurally similar poses.
[0076] In some embodiments, each pose (e.g., in a collection of poses) is scored against several (e.g., 2-100) different conformations of the target protein. In some embodiments, each pose (e.g., in a collection of poses) is scored against a fixed conformation of the target protein.
[0077] In some embodiments, the training compounds are docked to the target polymer 38 by either random pose generation techniques or biased pose generation. In some embodiments, the training compounds are docked to the target polymer 38 by Markov chain Monte Carlo sampling. In some embodiments, such sampling allows for sufficient flexibility of the training compounds in the docking calculations and a scoring function that is the sum of the interaction energy between the training compounds and the target polymer 38 and the conformational energy of the training (or test) objects. See, e.g., Liu and Wang, 1999, “MCDOCK: A Monte Carlo simulation approach to the molecular docking problem,” Journal of Computer-Aided Molecular Design 13, 435-451, which is incorporated herein by reference. In such embodiments, for a given training compound, the pose that received the best docking score is assigned the positive pose 48, and the pose that received the worst docking score is assigned the positive pose.
[0078] In some embodiments, algorithms such as DOCK (Shoichet, Bodian, and Kuntz, 1992, "Molecular docking using shape descriptors," Journal of Computational Chemistry 13(3), pp. 380-397, and Knegtel, Kuntz, and Oshiro, 1997 "Molecular docking to ensembles of protein structures," Journal of Molecular Biology 266, pp. 424-440, each of which is incorporated herein by reference) are used to find multiple poses for each of the training compounds relative to the target polymer 38. Such algorithms model the target polymer 38 and the training compounds as rigid bodies. The docked conformations are searched using complementary surfaces to find poses.
[0079] In some embodiments, AutoDOCK (Morris et al., 2009, “AutoDock4 and AutoDockTools4: Automated Docking with Selective Receptor Flexibility,” J. Comput. Chem. 30(16), pp. 2785-2791; Sotriffer et al., 2000, “Automated docking of ligands to antibodies: methods and applications,” Methods: A Companion to Methods in Enzymology 20, pp. 280-291, and Morris et al., 1998, “Automated Docking Using a Lamarckian Genetic Algorithm and Empirical Binding Free Energy Function,” Journal of Computational Chemistry 20, pp. 2771-2779, each of which is incorporated herein by reference) is used. 19:1639-1662) to find multiple poses for each of the training compounds against the target polymer 38. AutoDOCK uses a kinetic model of the ligand and supports Monte Carlo, simulated annealing, Lamarckian genetic algorithms, and genetic algorithms. Thus, in some embodiments, multiple different poses (for a given training compound) are obtained by Markov chain Monte Carlo sampling, simulated annealing, Lamarckian genetic algorithms, or genetic algorithms using a docking scoring function.
[0080] In some embodiments, an algorithm such as FlexX (Rarey et al., 1996, "A Fast Flexible Docking Method Using an Incremental Construction Algorithm," Journal of Molecular Biology 261, pp. 470-489, incorporated herein by reference) is used to find multiple poses for each training compound relative to the target polymer. FlexX uses a greedy algorithm to perform sequential construction of the training compound at the active site of the target polymer 38. Thus, in some embodiments, multiple different poses (for a given target compound) are obtained by the greedy algorithm.
[0081] In some embodiments, an algorithm such as GOLD (Jones et al., 1997, "Development and Validation of a Genetic Algorithm for flexible Docking," Journal Molecular Biology 267, pp. 727-748, incorporated herein by reference) is used to find multiple poses for each of the training compounds relative to the target polymer 38. GOLD stands for Genetic Optimization for Ligand Docking. GOLD builds a genetically optimized hydrogen bond network between the training compounds and the target polymer 38.
[0082] In some embodiments, molecular mechanics is performed on the target polymer (or a portion thereof, such as the active site of the target polymer) and each respective training compound to identify positive poses 48 and negative poses 60 for each respective training compound. During the molecular mechanics run, the atoms of the target polymer and training compounds are allowed to interact for a period of time to provide a view of the dynamic evolution of the system. The trajectories of the atoms in the target polymer and training compounds are determined by numerically solving Newton's equations of motion for the interacting particle system, and the forces between the particles and their respective potential energies are calculated using interatomic potentials or molecular mechanics force fields. See Alder and Wainwright, 1959, "Studies in Molecular Dynamics. I. General Method," J. Chem. Phys. 31(2):459, and Bibcode, 1959, J. Ch. Ph. 31, 459A, doi:10.1063 / 1.1730376, each of which is incorporated herein by reference in its entirety. Thus, in this manner, a molecular mechanics run generates a trajectory of the target polymer and each training compound over time, the trajectory including the trajectories of atoms within the target polymer and the training compound. In some embodiments, a subset of multiple different poses is obtained by snapshotting the trajectory over a period of time. In some embodiments, the poses are obtained from snapshotting several different trajectories, each trajectory including a different molecular mechanics run of the target polymer interacting with the training compound. In some embodiments, prior to the molecular mechanics run, the training compound is first docked into the active site of the target polymer using a docking technique.
[0083] In some embodiments, any pair from among multiple poses of each training compound against the target polymer (where one pose of the pose pair has a better docking score than the other pose of the pair) can serve as the positive pose 48 and negative pose 60 of each training compound, respectively.
[0084] Block 216. Several different non-limiting methods and programs for finding poses and determining in silico pose quality scores for such poses are disclosed above in conjunction with block 214 of FIG. 2B. In some embodiments, the first positive interaction score 50 of the positive pose 48 is an in silico pose quality score calculated for the positive pose 48 with respect to the target polymer 38 by any of these non-limiting methods and programs, or any combination thereof, or any equivalent or similar program. In some embodiments, the positive pose 48 is a collection of poses, as discussed above in block 214, and the first positive interaction score 50 of the positive pose 48 is an in silico pose quality score calculated for the positive pose 48 with respect to the target polymer 38 by any of these non-limiting methods and programs. Correspondingly, in some embodiments, the first negative interaction score 62 of the negative pose 60 is an in silico pose quality score calculated for the negative pose 60 with respect to the target polymer 38 by any of these non-limiting methods and programs. In some embodiments, the negative pose 60 is a collection of poses, as discussed above in block 214, and the first negative interaction score 62 of the negative pose 60 is an in silico pose quality score calculated for the negative pose 60 with respect to the target polymer 38 by any of these non-limiting methods and programs.
[0085] In some embodiments, rather than using an in silico quality score of the training compound, the first positive interaction score is determined by experimental means, based on the measured binding coefficient, IC, of the corresponding training compound 46 to the target polymer 38. 50 , E.C. 50 , Kd, KI, or pKI. IC 50 , E.C. 50Measured binding coefficients such as Kd, KI, and pKI are generally described in Huser ed., 2006, High-Throughput-Screening in Drug Discovery, Methods and Principles in Medicinal Chemistry35 and Chen ed., 2019, A Practical Guide to Assay Development and High-Throughput Screening in Drug Discovery, each of which is incorporated herein by reference in its entirety.
[0086] Block 218. Referring to block 218 of Figure 2B, in some embodiments, each training compound in the training dataset satisfies two or more, three or more, or all four of Lipinski's Rule of Five: (i) 5 or fewer hydrogen bond donors, (ii) 10 or fewer hydrogen bond acceptors, (iii) molecular weight less than 500 Daltons, and (iv) LogP less than 5. See Lipinski, 1997, Adv. Drug Del. Rev. 23, 3, which is incorporated herein by reference in its entirety.
[0087] In some embodiments, the training compound meets one or more criteria in addition to Lipinski's rule of five. For example, in some embodiments, the training compound has 5 or less aromatic rings, 4 or less aromatic rings, 3 or less aromatic rings, or 2 or less aromatic rings. In some embodiments, the training compound is any organic compound with a molecular weight of less than 2000 Daltons, less than 4000 Daltons, less than 6000 Daltons, less than 8000 Daltons, less than 10000 Daltons, or less than 20000 Daltons.
[0088] However, some embodiments of the disclosed systems and methods do not have a size limitation on the training compound, for example, in some embodiments, such training compounds are large polymers such as antibodies.
[0089] Referring to block 220, in some embodiments, each training compound in the training dataset is an organic compound having a molecular weight of less than 500 Daltons, less than 1000 Daltons, less than 2000 Daltons, less than 4000 Daltons, less than 6000 Daltons, less than 8000 Daltons, less than 10000 Daltons, or less than 20000 Daltons.
[0090] Blocks 224-226. Referring to block 224, the method trains at least a first model 72. The training uses, for each corresponding training compound 46 in at least a first subset of the plurality of training compounds, at least (i) a corresponding positive score of the corresponding positive pose 48 of the corresponding training compound 46 with respect to the target polymer 38 as input to the first model 72 relative to the corresponding first positive interaction score 50 of the corresponding training compound with respect to the target polymer, and (ii) a corresponding negative score of the corresponding negative pose 60 of the corresponding training compound with respect to the target polymer as input to the first model 72 relative to the corresponding first negative interaction score 62 of the corresponding training compound with respect to the target polymer, thereby adjusting a first plurality of parameters 73, and the output of at least the first model is used, at least in part, to provide a characterization of the interaction between the test compound and the target polymer. In some such embodiments, the training further uses, for each corresponding training compound 46 in the second subset of the plurality of training compounds, at least the corresponding positive score of the corresponding positive pose 48 of the corresponding training compound 46 with respect to the target polymer 38 as input to the first model 72 relative to the corresponding first positive interaction score 50 of the corresponding training compound with respect to the target polymer. In some embodiments, all of the training compounds have both positive and negative poses. In some embodiments, only a portion of the training compounds in the plurality of training compounds have both positive and negative poses, while other training compounds in the plurality of training compounds have positive poses but no negative poses. In some embodiments, only a portion of the training compounds in the plurality of training compounds have both positive and negative poses, while other training compounds in the plurality of training compounds either (i) have one or more positive poses but no negative poses, or (ii) have one or more negative poses but no positive poses.
[0091] Referring to block 234, in some embodiments, the first model 72 is a first fully connected neural network.
[0092] In Figure 9A, a first model 72 provides an estimate of the pose quality of a compound. For each corresponding training compound 46 in the plurality of training compounds, data in a training set 44 is used to train the model 72. For each training compound, the corresponding positive scores of the corresponding positive poses 48 of the corresponding training compound 46 are obtained with respect to the target polymer 38 as inputs to the first model 72.
[0093] According to the embodiment of FIG. 9A, the corresponding positive scores of the corresponding positive poses 48 are the output of the neural network 24 when the positive poses 48 are input to the neural network 24, as discussed in more detail below in block 228. As shown in FIG. 9A, in an exemplary embodiment, the positive scores are in the form of embedding from the embedding layer 96, which serves at least the purpose of sizing the positive scores to the dimensions necessary to serve as input to the first model. The output of the first model 72, upon input of the corresponding positive scores from the neural network 24, is compared against the corresponding first positive interaction scores 50 of the corresponding training compound with respect to the target polymer 38. The difference between the output of the first model 72 and the corresponding first positive interaction scores 50 is evaluated by a loss function to adjust the weights of the first model through a backpropagation technique 72.
[0094] Further, according to the embodiment of FIG. 9A, the corresponding negative scores of the corresponding negative poses 60 are the output of the neural network 24 when the negative poses 60 are input to the neural network 24, as discussed in more detail below in block 232, for those compounds in the training set that have negative poses. As shown in FIG. 9A, in an exemplary embodiment, the negative scores are in the form of embedding from the embedding layer 96, which serves at least the purpose of sizing the negative scores to the dimensions necessary to serve as input to the first model. The output of the first model 72, upon input of the corresponding negative scores from the neural network 24, is compared against the corresponding first negative interaction scores 62 of the corresponding training compound with respect to the target polymer 38. The difference between the output of the first model 72 and the corresponding first negative interaction scores 62 is also evaluated by a loss function to adjust the weights of the first model through a backpropagation technique.
[0095] The first model 72 has a first plurality of parameters 73. In some embodiments, the first plurality of parameters is 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 10,000, 50,000, 100,000, or 1×10 6 Contains more than one parameter.
[0096] Referring to block 226, in some embodiments, the first model 72 is a fully connected neural network, also known as a multilayer perceptron (MLP). In some embodiments, the MLP is a type of feedforward artificial neural network (ANN) that includes at least three layers: an input layer, a hidden layer, and an output layer of nodes. In such embodiments, except for the input nodes, each node is a neuron that uses a nonlinear activation function. In some embodiments, further disclosure regarding a suitable MLP that serves as the first model 72 can be found in Vang-mata ed., 2020, Multilayer Perceptrons: Theory and Applications, Nova Science Publishers, Hauppauge, New York, which is incorporated herein by reference.
[0097] Blocks 228 to 230. Referring to block 228 of FIG. 2B, in some embodiments, the corresponding positive scores of the corresponding positive poses 48 of the corresponding training compounds 46 with respect to the target polymer 38 are obtained by searching the corresponding positive voxel maps 52 of the corresponding training compounds 46 with respect to the target polymer 38 at the corresponding positive poses 48, expanding the corresponding positive voxel maps 52 into corresponding positive vectors 54, and inputting the corresponding positive vectors 54 into the neural network 24 in the form of a neural network (e.g., a convolutional neural network, a graph neural network, etc.). The graph neural network or the convolutional neural network 24 then provides the corresponding positive scores of the corresponding positive poses 48 at the output.
[0098] In some embodiments, neural network 24, with or without the use of voxel maps, may have more than 500 parameters, more than 1000 parameters, more than 2000 parameters, more than 5000 parameters, more than 10,000 parameters, more than 100,000 parameters, or more than 1×10 6Contains more than one parameter.
[0099] In some such embodiments, referring to block 230, the corresponding positive vector 54 referenced above is a first one-dimensional vector. In some embodiments, the corresponding positive vector 54 includes 10 or more elements, 20 or more elements, 100 or more elements, 500 or more elements, 1000 or more elements, or 10,000 or more elements.
[0100] In some embodiments, the neural network 24 is any of the convolutional neural networks 24 disclosed in Wallach et al., 2015, “AtomNet: A Deep Convolutional Neural Network for Bioactivity Prediction in Structure-based Drug Discovery,” arXiv:1510.02855v1, or U.S. Pat. Nos. 11,080,570, 10,546,237, 10,482,355, 10,002,312, or 9,373,059, each of which is incorporated herein by reference in its entirety. Further details regarding using a convolutional neural network to obtain corresponding positive scores for corresponding positive poses 48 of corresponding training compounds 46 with respect to a target polymer 38 are disclosed below in the section entitled “Using a Convolutional Neural Network to Obtain Pose Scores.”
[0101] In some embodiments, neural network 24 is an equivariant neural network. Non-limiting examples of equivariant convolutional neural networks include Thomas et al., 2018, “Tensor field networks: Rotation- and translation-equivariant neural networks for 3D point clouds,” arXiv:1802.08219; Anderson et al., 2019, “Cormorant: Covariant Molecular Neural Networks,” Neural Information Processing Systems; Johannes et al., 2020, “Directional Message Passing For Molecular Graphs,” International Conference on Learning Representations; Townshend et al., 2021, “ATOM3D: Tasks On Molecules in Three Dimensions,” International Conference on Learning Representations; Jing et al., 2009, “Learning from Protein Structure with Geometric Vector Perceptrons,” arXiv:2009.01411; and Satorras et al., 2021, “E(n) Equivariant Graph Neural Networks,” arXiv:2102.09844, each of which is incorporated herein by reference in its entirety.
[0102] In some embodiments, neural network 24 is a graph neural network (eg, a graph superposition neural network). Non-limiting examples of graph convolutional neural networks include Behler Parrinello, 2007, “Generalized Neural-Network Representation of High Dimensional Potential-Energy Surfaces,” Physical Review Letters 98, 146401; Chmiela et al., 2017, “Machine learning of accurate energy-conserving molecular force fields,” Science Advances 3(5): e1603015; Schuett et al., 2017, “SchNet: A continuous-filter convolutional neural network for modeling quantum interactions,” Advances in Neural Information Processing Systems 30, pp. 992-1002; Feinberg et al., 2018, “PotentialNet for Molecular Property Prediction,” ACS Cent. Sci. 4, 11, 1520-1530; and Stafford et al., “AtomNet PoseRanker: Enriching Ligand Pose Quality for Dynamic Proteins in Virtual High Throughput Screens,” https: / / chemrxiv.org / engage / chemrxiv / article-details / 614b905e39ef6a1c36268003, each of which is incorporated herein by reference in its entirety.
[0103] In some embodiments, the neural network 24 is any of the graph neural networks disclosed in U.S. Provisional Patent Application No. 63 / 336,841, entitled “Characterization of Interactions Between Compounds and Polymers Using Pose Ensembles,” filed May 10, 2022, which is incorporated by reference herein.
[0104] Blocks 232 to 234. Referring to block 232 of FIG. 2D, in some embodiments, the corresponding negative scores of the corresponding negative poses 60 of the corresponding training compounds 46 with respect to the target polymer 38 are obtained in the same manner as the corresponding positive scores of the corresponding positive poses 48 of the corresponding training compounds 46 with respect to the target polymer 38 are obtained. For example, in some embodiments, it is obtained by searching the corresponding negative voxel map of the corresponding training compounds with respect to the target polymer at the corresponding negative poses 60, expanding the corresponding negative voxel map into a corresponding negative vector 66, and inputting the corresponding negative vector into the neural network 24, thereby obtaining the corresponding negative scores of the corresponding negative poses 60 of the corresponding training compounds 46 with respect to the target polymer 38. Referring to block 234 of FIG. 2D, in some such embodiments, the corresponding negative vector 66 is a second one-dimensional vector.
[0105] Blocks 236-244. Referring to block 236 of FIG. 2D, in some embodiments, training the model 72 is a regression task in which a first plurality of parameters 73 of the first model 72 are tuned by backpropagation through an associated loss function, and a corresponding first positive interaction score 50 is calculated using the formula B=N×A where A is the corresponding positive interaction score, B is the corresponding negative interaction score, and N is a real number greater than zero and less than one. Training the model 72 as a regression task is suitable when the first positive interaction scores are measured properties of the respective training compounds from wet-lab (e.g., in vivo or in vitro) assays. An example of such a measured property of the respective training compounds is the IC of the respective training compounds with respect to the target polymer. 50 , E.C. 50 , Kd, KI, or pKI. In such an embodiment, it is reasonable to assign a first positive interaction score 50 to the measured property of the training compound. Then, for training purposes, the question becomes what to assign the first negative interaction score 62 of the training compound, given the measured property of the training compound. According to block 236 of FIG. 2D, in some embodiments, the negative interaction score 62 is assigned a fixed discounted value N of the measured property. By fixed, it is meant that the same value N is applied to each first positive interaction score 50 for each respective training compound to calculate the value of the corresponding first negative interaction score 62. Thus, if the value of N is 0.90, then for each respective training compound, the corresponding first negative interaction score 62 has a value that is 0.90 of the corresponding first positive interaction score 50. In some embodiments, N is a value between 0.10 and 0.99. In some embodiments, N is a value between 0.20 and 0.95. In some embodiments, N is a value between 0.30 and 0.90. In some embodiments, N is a value between 0.25 and 0.85. In some embodiments, N is a value between 0.60 and 0.95. In some alternative embodiments, the negative interaction score 62 is assigned the logarithm of the measured property. Thus, in such embodiments, for each respective training compound, the corresponding first negative interaction score is the logarithm of the corresponding first positive interaction score 50. The logarithm can be to any base, such as natural logarithm, base 10, etc.
[0106] Referring to block 238, in some embodiments, the associated loss function described above with respect to block 232 is any suitable regression task loss function. Examples of such loss functions include, but are not limited to, a mean squared error loss function, a mean absolute error loss function, a Huber loss function, a Log-Cosh loss function, or a quantile loss function. See Wang et al., 2020, “A Comprehensive Survey of Loss Functions in Machine Learning,” Annals of Data Science, https: / / doi.org / 10.1007 / s40745-020-00253-5, last accessed September 15, 2021, each of which is incorporated herein by reference.
[0107] Referring to block 240 of FIG. 2D , in some particular embodiments, the corresponding first positive interaction score 50 and the corresponding first negative interaction score 62 each represent a binding coefficient, and the corresponding first positive interaction score is an in vivo or in vitro measurement of the binding coefficient of the corresponding training compound 46 to the target polymer 38.
[0108] Referring to block 244 of FIG. 2E, in some embodiments, the first positive interaction score is determined based on the IC 50 , E.C. 50 , Kd, KI, or pKI. Measured binding coefficients are generally described in Huser ed., 2006, High-Throughput-Screening in Drug Discovery, Methods and Principles in Medicinal Chemistry 35 and Chen ed., 2019, A Practical Guide to Assay Development and High-Throughput Screening in Drug Discovery, each of which is incorporated herein by reference in its entirety.
[0109] Blocks 246-248. Referring to block 246 of FIG. 2E, in some embodiments, each respective electronic description 46 in at least a subset of electronic descriptions 46 in the training dataset 44 further includes a corresponding positive score 56 for a corresponding positive pose 48 of the corresponding training compound 46 and a corresponding negative activity score 58 for a corresponding negative pose 60 of the corresponding training compound. In some embodiments, at least some of the training compounds do not have a negative activity score 58. Referring to FIG. 9B, in some embodiments, training at least the first model 72 further includes jointly training a second model 74 with the first model.
[0110] Similar to the first model 72, the second model 74 has a plurality of parameters 75 (second plurality of parameters). In some embodiments, the second plurality of parameters is 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 10,000, 50,000, 100,000, or 1×10 6 Contains more than one parameter.
[0111] In the embodiment of Figure 9B, the second model 74 provides an estimate of the compound's pose quality. For each corresponding training compound 46 in the plurality of training compounds, the data in the training set 44 is used to train the second model 74. For each training compound, the corresponding positive scores of the corresponding positive poses 48 of the corresponding training compound 46 are obtained with respect to the target polymer 38 as input to the second model 74.
[0112] According to the embodiment of FIG. 9B, the corresponding positive score of the corresponding positive pose 48 is the output of the neural network 24 when the positive pose 48 is input to the neural network 24. As shown in FIG. 9B, in an exemplary embodiment, the positive score is a form of embedding from the embedding layer 96, which serves at least the purpose of sizing the positive score to a dimension necessary to serve as an input to the second model 74. The output of the second model 74 is compared against the corresponding first positive interaction score 50 of the corresponding training compound with respect to the target polymer 38 when the corresponding positive score is input from the neural network 24 to the second model 74, as shown by the edge 920. The difference between the output of the second model 74 and the corresponding first positive interaction score 50 is evaluated by a loss function to adjust the weights of the second model 74 through a backpropagation technique.
[0113] Further, according to the embodiment of FIG. 9B, the corresponding negative scores of the corresponding negative poses 60 are the output of the neural network 24 when the negative poses 60 are input to the neural network 24 for those compounds in the training set that have negative poses. As shown in FIG. 9B, in the exemplary embodiment, the negative scores are in the form of embedding from the embedding layer 96, which serves at least the purpose of sizing the negative scores to a dimension necessary to serve as an input to the second model. The output of the second model 74 is compared against the corresponding first negative interaction scores 62 of the corresponding training compounds with respect to the target polymer 38 upon inputting the corresponding negative scores from the neural network 24 to the second model 74, as shown by the edge 920. The difference between the output of the second model 74 and the corresponding first negative interaction scores 62 is also evaluated by a loss function to adjust the multiple parameters 75 of the second model through a backpropagation technique.
[0114] Further, in the embodiment shown in Figure 9B, for each corresponding training compound in the plurality of training compounds, the training of block 224 further uses, for at least a subset of the training compounds, (iii) corresponding positive scores of corresponding positive poses 48 of the corresponding training compound 46 with respect to the target polymer 38 as inputs to the first model 72 (shown by edge 930 in Figure 9B) relative to corresponding positive activity scores 56 of the corresponding training compound, and (iv) corresponding negative scores of corresponding negative poses 60 of the corresponding training compound 46 with respect to the target polymer 38 as inputs to the first model 72 (shown again by edge 930 in Figure 9B) relative to corresponding negative activity scores 68 of the corresponding training compound 46. In this manner, the first plurality of parameters 73 of the first model are adjusted during training.
[0115] 9B embodiment, the second model 74 is trained against the respective first positive interaction scores 50 and first negative interaction scores 62, while the first model 72 is trained against the positive activity scores 56 and negative activity scores 68. In some such embodiments, the first positive interaction scores 50 and first negative interaction scores 62 are docking scores, and the positive activity scores and negative activity scores are binary discrete activity values. For example, one of the two possible values of the binary discrete activity value would indicate that the corresponding training inhibits the activity of the target polymer, while the other of the two possible values of the binary discrete activity value would indicate that the corresponding training does not inhibit that activity of the target polymer.
[0116] As shown in Figure 9B, once trained, the poses of the test compound are propagated to neural network 24, which generates a score of the pose of the test compound relative to the target polymer. This score of the pose of the test compound with respect to the target polymer is input to both second model 74 (to provide a characterization of the interaction between the test compound and the target polymer in the form of a pose quality score) and first model 72 (to provide a characterization of the interaction between the test compound and the target polymer in the form of the activity of the interaction between the test compound and the target polymer 38). Thus, in the embodiment of Figure 9B, the characterization of the interaction between the test compound and the target polymer is both an activity score (e.g., a discrete binary score or a scalar score) and a pose quality score.
[0117] Referring to block 248, in some such embodiments, the first model 72 and the second model 74 are each fully connected neural networks, also known as multilayer perceptrons (MLPs). In some embodiments, the MLP is a type of feedforward artificial neural network (ANN) that includes at least three layers: an input layer, a hidden layer, and an output layer of nodes. In such embodiments, except for the input nodes, each node is a neuron that uses a nonlinear activation function. In some embodiments, further disclosure regarding suitable MLPs that function as the first model 72 can be found in Vang-mata ed., 2020, Multilayer Perceptrons: Theory and Applications, Nova Science Publishers, Hauppauge, New York, which is incorporated herein by reference.
[0118] Blocks 252-256. Referring to block 252 of FIG. 2F, in some embodiments, as shown in FIG. 9C, each respective electronic description 46 in at least a subset of the training dataset 44 further includes corresponding positive activity scores 56 for corresponding positive poses 48 of the corresponding training compound 46 and corresponding negative activity scores 58 for corresponding negative poses 60 of the corresponding training compound. In such embodiments, the training described above in block 224 (training at least the first model 72) further includes jointly training a second model 74 with the first model 72. The second model 74 has a second plurality of parameters 75.
[0119] In the embodiment of Figure 9C, the second model 74 provides an estimate of the compound's pose quality. For each corresponding training compound 46 in the plurality of training compounds, the data in the training set 44 is used to train the second model 74. For each training compound, the corresponding positive scores of the corresponding positive poses 48 of the corresponding training compound 46 are obtained with respect to the target polymer 38 as input to the second model 74.
[0120] According to the embodiment of FIG. 9C, the corresponding positive score of the corresponding positive pose 48 is the output of the neural network 24 when the positive pose 48 is input to the neural network 24. As shown in FIG. 9C, in an exemplary embodiment, the positive score is a form of embedding from the embedding layer 96, which serves the purpose of at least sizing the positive score to a dimension necessary to serve as an input to the first model and the second model. The output of the second model 74 is compared against the corresponding first positive interaction score 50 of the corresponding training compound with respect to the target polymer 38 when the corresponding positive score is input from the neural network 24 to the second model 74, as shown by the edge 940. The difference between the output of the second model 74 and the corresponding first positive interaction score 50 is evaluated by a loss function to adjust the weights of the second model 74 through a backpropagation technique.
[0121] Further, according to the embodiment of FIG. 9C, the corresponding negative scores of the corresponding negative poses 60 are the output of the neural network 24 when the negative poses 60 are input to the neural network 24 for those compounds in the training set that have negative poses. As shown in FIG. 9C, in the exemplary embodiment, the negative scores are in the form of embedding from the embedding layer 96, which serves at least the purpose of sizing the negative scores to a size necessary to serve as an input to both the first model and the second model. The output of the second model 74 is compared against the corresponding first negative interaction scores 62 of the corresponding training compounds with respect to the target polymer 38 upon inputting the corresponding negative scores from the neural network 24 to the second model 74, as shown by the edge 940. The difference between the output of the second model 74 and the corresponding first negative interaction scores 62 is also evaluated by a loss function to adjust the parameters 75 of the second model through a backpropagation technique.
[0122] 9C further uses, for each corresponding training compound 46 in at least a subset of the plurality of training compounds, the corresponding positive scores of the corresponding positive poses 48 of the corresponding training compound with respect to the target polymer 38 provided by both the model 24 (through the edge 950) and the second model 74 (through the edge 930) as combined inputs to the first model 72 for at least the corresponding positive activity scores 56 of the corresponding training compound, and the corresponding negative scores of the corresponding negative poses 60 of the corresponding training compound 46 with respect to the target polymer 38 provided by both the model 24 (again, through the edge 950) and the second model 74 (again, through the edge 930) for the corresponding negative activity scores 68 of the corresponding training compound. In this manner, the first plurality of parameters 73 of the first model 72 are adjusted (e.g., through backpropagation using a loss function).
[0123] The second model 74 is used, at least in part, with the output of the first model 72 to provide a characterization of the interaction between the test compound and the target polymer. For example, as shown in FIG. 9C, once trained, the pose of the test compound is passed to the neural network 24, which generates a score of the pose of the test compound relative to the target polymer 38. This score for the target polymer is input to both the first model 72 (through edge 950) and the second model 74 (through edge 940). Additionally, the output of the second model 74 for the test compound (which is a calculation of an interaction score, such as a pose quality score, pKA, etc.) is input to the first model 72 through edge 930. Thus, the first model 72 receives both the output of the second model and the output of model 24 in response to inputting the pose of the test compound to model 24. The first model 72 uses both of these inputs to determine a characterization of the interaction between the test compound and the target polymer. In some embodiments, this characterization is an activity score for the test compound. In some embodiments, the activity score is a discrete binary score, e.g., where "1" indicates that the test compound is active against the target polymer and "0" indicates that the test compound is inactive against the target polymer. In some embodiments, the activity score provided by the first model 72 is a scalar. Conditioning the (discrete binary) activity score of the first model 72 against both the output of model 24 and the second model 74 serves to improve the performance of the first model in characterizing the test compound.
[0124] Referring to block 254 of FIG. 2F, in some such embodiments, the corresponding positive activity score 56 is a first binary activity score and the corresponding negative activity score 68 is a second binary activity score. In some embodiments, the corresponding first binary activity score is assigned a value of 1 based on the measured activity of the corresponding compound against the target polymer based on meeting the activity criteria, and the corresponding second binary activity score is assigned a value of 0 based on not meeting the activity criteria. In some embodiments, these activity values of the training compounds are obtained by in vivo or in vitro assays. Such assays are generally described in Huser ed., 2006, High-Throughput-Screening in Drug Discovery, Methods and Principles in Medicinal Chemistry 35 and Chen ed., 2019, A Practical Guide to Assay Development and High-Throughput Screening in Drug Discovery, each of which is incorporated herein by reference in its entirety.
[0125] Referring to block 256 of FIG. 2F, in some embodiments, the training of the second model 74 is a regression task in which the second plurality of parameters 75 are tuned by backpropagation through a second associated loss function. Non-limiting examples of loss functions suitable for regression tasks include, but are not limited to, a mean squared error loss function, a mean absolute error loss function, a Huber loss function, a Log-Cosh loss function, or a quantile loss function. See Wang et al., 2020, “A Comprehensive Survey of Loss Functions in Machine Learning,” Annals of Data Science, https: / / doi.org / 10.1007 / s40745-020-00253-5, last accessed September 15, 2021, which is incorporated by reference in its entirety. Additionally, in some embodiments, the training of the first model 72 is a classification task in which the first plurality of parameters 73 are tuned by backpropagation through a first associated loss function. Non-limiting examples of loss functions suitable for classification tasks include, but are not limited to, a binary cross-entropy loss function, a hinge loss function, or a squared hinge loss function.
[0126] In some embodiments, the output of the first model is a discrete value other than binary. For example, a first output value of the second model (in response to inputting a pause into classifier 24 in the configuration shown in FIG. 9C) indicates poor activity of the test compound on the target polymer, a second output value indicates intermediate activity of the test compound on the target polymer, and a third output value indicates good activity of the test compound on the target polymer. In some such embodiments, the loss function used to train the first classifier can be a multiclass classification loss function, such as a multiclass cross-entropy loss function, a sparse multiclass cross-entropy loss function, or a Kullback-Leibler divergence loss function.
[0127] Block 260. Referring to block 260 of Figure 2G, in some embodiments, the corresponding first positive interaction score 50 and the corresponding first negative interaction score 62 each represent a binding coefficient or in silico quality score of the corresponding training compound for the target polymer, the corresponding positive activity score 56 is a first binary activity score, and the corresponding negative activity score 68 is a second binary activity score.
[0128] Block 262. Referring to block 262 of FIG. 2G, in some embodiments, the first associated loss function is a mean squared error loss function, a mean absolute error loss function, a Huber loss function, a Log-Cosh loss function, or a quantile loss function, and the second associated loss function is a binary cross-entropy loss function, a hinge loss function, or a squared hinge loss function.
[0129] Block 264. Referring to block 264 of FIG. 2G, in some embodiments, the second model 74 is a second fully connected neural network, also known as a multilayer perceptron (MLP). In some embodiments, the MLP is a type of feedforward artificial neural network (ANN) that includes at least three layers: an input layer, a hidden layer, and an output layer of nodes. In such embodiments, except for the input nodes, each node is a neuron that uses a nonlinear activation function. In some embodiments, further disclosure regarding a suitable MLP that serves as the first model 72 can be found in Vang-mata ed., 2020, Multilayer Perceptrons: Theory and Applications, Nova Science Publishers, Hauppauge, New York, which is incorporated herein by reference.
[0130] Blocks 268-276. Referring to block 268 of FIG. 2H, as shown in FIG. 16A, in some embodiments, each respective electronic description in the training data set further includes a corresponding second positive interaction score for the corresponding positive pose 48 of the corresponding training compound 46 and a corresponding second negative interaction score for the corresponding negative pose 60 of the corresponding training compound. Additionally, each respective electronic description in the training data set also includes a corresponding positive activity score 56 for the corresponding positive pose 48 of the corresponding training compound 46 and a corresponding negative activity score 68 for the corresponding negative pose 60 of the corresponding training compound.
[0131] In such an embodiment, the training of at least the first model 72, the second model 74, and the third model 76 are trained jointly.
[0132] The second model 74 has a second plurality of parameters 75. In some embodiments, the second plurality of parameters is 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 10,000, 50,000, 100,000, or 1×10 6 Contains more than one parameter.
[0133] The third model 76 has a third plurality of parameters 77. In some embodiments, the third plurality of parameters is 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 10,000, 50,000, 100,000, or 1×10 6 Contains more than one parameter.
[0134] Model co-training uses, for each corresponding training compound 46 in at least a subset of the multiple training compounds, at least (i) the corresponding positive score of the corresponding positive pose 48 of the corresponding training compound with respect to the target polymer 38 provided by model 24 (through edge 1610) as a combined input to the second model 74 relative to the corresponding first positive interaction score 50 of the corresponding training compound with respect to the target polymer 38, and (ii) the corresponding negative score of the corresponding negative pose 60 of the corresponding training compound 46 with respect to the target polymer 38 provided by model 24 (again, through edge 1610) as an input to the second model 74 relative to the corresponding first negative interaction score 62 of the corresponding training compound 46 with respect to the target polymer 38, thereby adjusting a second plurality of parameters of the second model.
[0135] The model co-training further uses, for each corresponding training compound 46 in at least a subset of the multiple training compounds, at least the corresponding positive score of the corresponding positive pose 48 of the corresponding training compound with respect to the target polymer 38 provided by the model 24 (through edge 1620) as input to the third model 76 for the corresponding second positive interaction score 58 of the corresponding training compound with respect to the target polymer 38, and the corresponding negative score of the corresponding negative pose 60 of the corresponding training compound 46 with respect to the target polymer 38 provided by the model 24 (again, through edge 1620) as input to the third model 76 for the corresponding second negative interaction score 70 of the corresponding training compound 46 with respect to the target polymer 38, thereby adjusting a third plurality of parameters 77 of the third model 76.
[0136] The model co-training includes, for each corresponding training compound 46 in at least a subset of the plurality of training compounds, outputting at least: (i) the corresponding positive scores of the corresponding positive poses 48 of the corresponding training compound with respect to the target polymer 38 provided by the model 24 (through the edge 1630) and (ii) the output of the second model 74 through the edge 1640 when inputting the corresponding positive scores of the corresponding positive poses 48 of the corresponding training compound with respect to the target polymer 38 provided by the model 24 into the second model 74; and (iii) the output of the third model 76 through the edge 1650 when inputting the corresponding positive scores of the corresponding positive poses 48 of the corresponding training compound with respect to the target polymer 38 provided by the model 24 into the third model 76 as a collectively input to the first model 72 for the corresponding positive activity scores of the corresponding training compound with respect to the target polymer 38. 76, and further using at least (i) the corresponding negative scores of the corresponding negative poses of the corresponding training compounds for the target polymer 38 provided by model 24 (through edge 1630), (ii) the output of the second model 74 through edge 1640 when the corresponding negative scores of the corresponding negative poses of the corresponding training compounds for the target polymer 38 provided by model 24 are input to the second model 74, and (iii) the output of the third model 76 through edge 1650 when the corresponding negative scores of the corresponding negative poses of the corresponding training compounds for the target polymer 38 provided by model 24 are input to the third model 76 as a lumped input to the first model 72 for the corresponding negative activity scores of the corresponding training compounds for the target polymer 38.
[0137] The first model 74 is used to provide a characterization of the interaction between the test compound and the target polymer. For example, as shown in FIG. 16A, once trained, the pose of the test compound is passed to the neural network 24, which generates a score of the pose of the test compound relative to the target polymer 38. This score for the target polymer is input to the first model 72 (through edge 1630), the second model 74 (through edge 1610), and the third model (through edge 1620). Additionally, the output of the second model 74 of the test compound (which is a calculation of an interaction score, such as a pose quality score) is input to the first model 72 through edge 1640. Additionally, the output of the third model 76 of the test compound (which is a calculation of an interaction score, such as pKA) is input to the first model 72 through edge 1650. Thus, the third model receives the outputs of the first model, the second model, and the model 24 in response to the input of the pose of the test compound to the model 24. The first model 72 uses each of these inputs to collectively determine a characterization of the interaction between the test compound and the target polymer. In some embodiments, this characterization is an activity score for the test compound. In some embodiments, this activity score is a discrete binary score, e.g., where "1" indicates that the test compound is active against the target polymer and "0" indicates that the test compound is inactive against the target polymer. In some embodiments, the activity score provided by the third model 74 is a scalar. Conditioning the discrete binary) activity score of the first model 72 against the outputs of model 24, second model 74, and third model 76 helps to improve the performance of the first model in characterizing the test compound by forcing this first model to take into account the binding mode when calculating activity, thus addressing the Picasso problem that arises in machine learning. Thus, the output of the first model provides a characterization of the interaction between the test compound and the target polymer.
[0138] In some embodiments, referring to FIG. 16A, the embeddings 96 generated by the neural network 24 are used to predict three outputs: activity (through the first model 72), CUina pose quality score (through the second model 74), and pKi score (through the third model 76). This is performed in two stages in the embodiment shown in FIG. 16A. First, the CUina and pKi score predictions are calculated by passing the test compound's pose score against the target polymer 38 from the neural network 24 through the second model 74 and the third model 76 (as embeddings 96). Second, a conditioned embedding 1690 is formed by concatenating (i) the input embeddings 96 (the test compound's pose score against the target polymer 38 from the neural network scores), (ii) the second model 74 score prediction resulting from the first stage, and (iii) the third model 76 score prediction from the first stage. This embedding 1690 is then passed to the first model 72, which is in the form of a multi-layer perceptron, to calculate an activity prediction for the test compound. In some embodiments, rather than simply concatenating (i) the input embedding 96 (the test compound's pose score for the target polymer 38 from the neural network scores), (ii) the second model 74 score prediction resulting from the first stage, and (iii) the third model 76 score prediction from the first stage, the embedding 1690 multiplies these three sources together and inputs the product of the multiplication into the third model as embedding 1690. In some embodiments, embedding 1690 in Figure 16A does not simply concatenate (i) the input embedding 96 (the scores of the poses of the test compounds relative to the target polymer 38 from the neural network scores), (ii) the score predictions of the second model 74 resulting from the first stage, and (iii) the score predictions of the third model 76 from the first stage, but rather multiplies these three sources together and inputs the product of the multiplication into the third model as embedding 1690. In some embodiments, rather than concatenating, embedding 1690 transforms each of the three sources in embedding 1690, and this transformation serves as an input to the first model 72.More generally, the embedding 1690 may perform any mathematical function on all or any portion of any of the inputs to the embedding 1690, including but not limited to multiplication, concatenation, linear or non-linear transformations, to form a conditioning embedding that is passed on to the first model 72.
[0139] With reference to Figure 16B, it is possible to condition the first model 72 against additional models as well. Thus, in Figure 16B, the first model 72 is conditioned against the output of the network 24 as well as the output of a second model 74, e.g., trained on the CUina scores of the training compounds, a third model 76, e.g., trained on the pKi scores of the training compounds, and a fourth model 990, e.g., trained on the PoseNet scores of the training compounds.
[0140] Referring to block 272 of FIG. 2I, in some such embodiments, the first model, the second model 74, the third model 76, and the fourth model 990 are each fully connected neural networks. Such fully connected neural networks are also known as multilayer perceptrons (MLPs). In some embodiments, the MLP is a type of feedforward artificial neural network (ANN) that includes at least three layers: an input layer, a hidden layer, and an output layer of nodes. In such embodiments, except for the input node, each node is a neuron that uses a nonlinear activation function. In some embodiments, further disclosure regarding suitable MLPs that function as the first model 72 can be found in Vang-mata ed., 2020, Multilayer Perceptrons: Theory and Applications, Nova Science Publishers, Hauppauge, New York, which is incorporated herein by reference.
[0141] 2I, in some embodiments, the corresponding positive activity score provided by the first model 72 is a first binary activity score and the corresponding negative activity score provided by the first model 72 is a second binary activity score. In some embodiments, the corresponding first binary activity score is assigned a value of "1" and the corresponding second binary activity score is assigned a value of "0" based on the measured activity of the corresponding training compound against the target polymer.
[0142] Referring to block 276 of FIG. 2I, in some embodiments, the training of the second model 74 is a regression task in which a second plurality of parameters associated with the second model are adjusted by backpropagation through a second associated loss function. Further, in some embodiments, the training of the third model 76 is a regression task in which a third plurality of parameters associated with the third model are adjusted by backpropagation through a third associated loss function. Further, in some embodiments, the training of the fourth model 990 is a regression task in which a fourth plurality of parameters associated with the fourth model 990 are adjusted by backpropagation through a fourth associated loss function. Non-limiting examples of loss functions suitable for these regression tasks include, but are not limited to, a mean squared error loss function, a mean absolute error loss function, a Huber loss function, a Log-Cosh loss function, or a quantile loss function. See Wang et al., 2020, “A Comprehensive Survey of Loss Functions in Machine Learning,” Annals of Data Science, https: / / doi.org / 10.1007 / s40745-020-00253-5, last accessed September 15, 2021, which is incorporated by reference in its entirety. Further, in some embodiments, the training of the first model 72 is a classification task in which a first plurality of parameters associated with the first model 72 are adjusted by backpropagation through a first associated loss function. Non-limiting examples of loss functions suitable for classification tasks include, but are not limited to, a binary cross-entropy loss function, a hinge loss function, or a squared hinge loss function.
[0143] In some such embodiments, the corresponding first positive interaction score and the corresponding first negative interaction score each represent an in silico quality score of the corresponding training compound for the target polymer, the corresponding second positive interaction score and the corresponding second negative interaction score each represent a binding coefficient of the corresponding training compound for the target polymer, the corresponding positive activity score is a first binary activity score, and the corresponding negative activity score is a second binary activity score. In some such embodiments, the second, third, and fourth associated loss functions are each independently a mean squared error loss function, a mean absolute error loss function, a Huber loss function, a Log-Cosh loss function, or a quantile loss function, while the first associated loss function is a binary cross-entropy loss function, a hinge loss function, or a squared hinge loss function.
[0144] 16B, the embeddings 96 generated by neural network 24 are used to predict four outputs: activity (through first model 72), CUina pose quality score (through second model 74), pKi score (through third model 76), and PoseNet score (through fourth model 990). This is performed in two stages in the embodiment shown in FIG. 16B. First, the CUina, pKi, and PoseNet score predictions are calculated by passing the pose scores of the test compound relative to the target polymer 38 from neural network 24 through the second model 74, the third model 76, and the fourth model 990 (as embeddings 96). Second, a conditioned embedding 1690 is formed by concatenating (i) the input embedding 96 (the score of the test compound's pose relative to the target polymer 38 from the neural network scores), (ii) the score prediction of the second model 74 resulting from the first stage, and (iii) the score prediction of the third model 76 from the first stage. This embedding 1690, along with the output of the fourth model, is then passed to the first model 72, which is in the form of a multi-layer perceptron, to calculate an activity prediction for the test compound. In some embodiments, the embedding 1690 of FIG. 16B does not simply concatenate (i) the input embedding 96 (the scores of the test compound's poses for the target polymer 38 from the neural network scores), (ii) the score predictions of the second model 74 resulting from the first stage, and (iii) the score predictions of the third model 76 from the first stage, but rather multiplies these three sources together and inputs the product of the multiplication into the third model as the embedding 1690. In some embodiments, rather than concatenating, the embedding 1690 transforms each of the three sources in the embedding 1690, and this transformation serves as an input to the first model 72. More generally, the embedding 1690 can perform any mathematical function on all or any portion of any of the inputs to the embedding 1690, including but not limited to multiplication, concatenation, linear or non-linear transformations, to form the conditioning embedding that is passed on to the first model 72.
[0145] FIG. 10 illustrates a system for characterizing the interaction between a test compound and a target polymer according to one embodiment of the present disclosure, where the characterization is (i) binary discrete activity and (ii) pKi, and the system is trained using combined positive and negative poses to train the compound. Although not shown in FIG. 10, a shared embedding layer receives the output from the neural network 24 upon input of the voxelated poses of the compound to the neural network 24. In the system of FIG. 10, the pKi model and the activity model are independent of each other. In some embodiments, the pKi model is trained as a regression task using a loss function such as mean squared error, while the activity model is trained as a classification task using a loss function such as binary cost entropy.
[0146] FIG. 11 is a system for characterizing the interaction between a test compound and a target polymer according to one embodiment of the present disclosure, where the characterization is pKi, where pKi is conditioned in part on activity, and the system is trained using combined positive and negative poses to train the compound. Although not shown in FIG. 11, a shared embedding layer receives output from the neural network 24 upon input of the voxelated poses of the compound to the neural network 24. In the system of FIG. 11, the pKi model is conditioned against the activity model. In some embodiments, the pKi model is trained as a regression task using a loss function such as mean squared error, while the activity model is trained as a classification task using a loss function such as binary cost entropy.
[0147] FIG. 12 is a system for characterizing the interaction between a test compound and a target polymer according to one embodiment of the present disclosure, where the characterization is activity, and the activity is conditioned, in part, on both pKi and pose quality scores, and the system is trained using combined positive and negative poses to train the compound. Although not shown in FIG. 12, a shared embedding layer receives output from the neural network 24 upon input of the voxelated poses of the compound to the neural network 24. In the system of FIG. 12, the activity model is conditioned on the pKi model. In some embodiments, the pKi model is trained as a regression task using a loss function such as mean squared error, while the activity model is trained as a classification task using a loss function such as binary cost entropy.
[0148] FIG. 13 is a system for characterizing the interaction between a test compound and a target polymer according to one embodiment of the present disclosure, where the characterization is activity, and the activity is conditioned, in part, on both pKi and compound binding mode scores, and the system is trained using combined positive and negative poses to train the compound. Although not shown in FIG. 13, a shared embedding layer receives the output from the neural network 24 upon input of the voxelated poses of the compound to the neural network 24. In the system of FIG. 13, the activity model is conditioned on both the pKi model and the posenet model. In some embodiments, the pKi model and the posenet model are trained as a regression task using a loss function such as mean squared error, while the activity model is trained as a classification task using a loss function such as binary cost entropy.
[0149] 14 is a system for characterizing the interaction between a test compound and a target polymer according to one embodiment of the present disclosure, where the characterization is activity and two different compound binding mode scores, and the system is trained using combined positive and negative poses to train the compound. In the system of FIG. 14, the activity model is conditioned against a pose quality score model. In some embodiments, the pose quality model is trained as a regression task using a loss function such as mean squared error, while the activity model is trained as a classification task using a loss function such as binary cost entropy.
[0150] 15 is a system for characterizing the interaction between a test compound and a target polymer according to one embodiment of the present disclosure, where the characterizations are activity, two different compound binding mode scores and pKi, and the system is trained using combined positive and negative poses to train the compound. In the system of FIG. 15, the activity model is conditioned against a pose quality score model. In some embodiments, the pose quality model is trained as a regression task using a loss function such as mean squared error, while the activity model is trained as a classification task using a loss function such as binary cost entropy.
[0151] Representative test compounds and training compounds. A significant difference between the test compounds and the training compounds is that the training compounds are labeled (e.g., by complementary binding data obtained from a wet-lab binding assay) and such labels are used to train the neural network 24 and other models of the present disclosure, whereas the test compounds are not labeled and the neural network 24 and other models of the present disclosure are used to classify the test compounds. In other words, the training compounds are already classified by labels and such classifications are used to train the neural network 24 and other models of the present disclosure, so that the models of the present disclosure can then classify the test compounds. The test compounds are typically not classified before application of the neural network 24 and other models of the present disclosure. In a typical embodiment, the classification associated with the training compounds is the binding data to the target polymer 38 obtained by a wet-lab binding assay.
[0152] Training a predictive model. In some embodiments where a deep neural network (e.g., neural network 24) is implemented, network 24 is trained to receive geometric data input and output a prediction (probability) of whether a given test compound will bind to a target polymer. For example, in some embodiments, training compounds with known binding data to the target polymer (for each associated binding data) are passed sequentially through neural network 24 and models of the present disclosure using the techniques discussed above in connection with FIG. 2, and neural network 24 provides a single value for each respective training compound.
[0153] In some such embodiments, the system of the present disclosure outputs one of two possible activity classes for each training object for a given target compound. For example, the single value provided by the system of the present disclosure for each respective training compound is in a first activity class (e.g., binder) if it is below a predetermined threshold, and in a second activity class (e.g., non-binder) if its number is above a predetermined threshold. The activity class assigned by the system of the present disclosure is compared to the actual activity class as represented by the training compound binding data. In an exemplary non-limiting embodiment, such training compound binding data is from an independent web lab binding assay. The error of the activity class assignment made by the system of the present disclosure is back-propagated through the weights of each model (e.g., 24, 72, 74, etc.) of the system of the present disclosure to be validated against the binding data, and then to train the system. For example, the filter weights of each filter in the optional convolutional layer 28 of the network are adjusted in such back-propagation. In an exemplary embodiment, the neural network 24 is trained by stochastic gradient descent with the AdaDelta adaptive learning method (Zeiler, 2012, “ADADELTA: an adaptive learning rate method,” 'CoRR, vol. abs / 1212.5701, incorporated herein by reference) taking into account the combined data for errors in class assignments made by the network, and the back-propagation algorithm as presented in Rumelhart et al., 1988, “Neurocomputing: Foundations of research,” ch. Learning Representations by Back-propagating Errors, pp. 696-699, Cambridge, MA, USA: MIT Press, incorporated herein by reference.In some such embodiments, the two possible activity classes are each determined by a binding constant greater than a given threshold amount (e.g., an IC of the training compound for the target polymer greater than 1 nanomolar, 10 nanomolar, 100 nanomolar, 1 micromolar, 10 micromolar, 100 micromolar, or 1 millimolar). 50 , E.C. 50 , or KI) and a binding constant below a given threshold amount (e.g., IC 50 , E.C. 50 , or KI).
[0154] In some embodiments, the system of the present disclosure outputs one of a plurality of possible activity classes (e.g., three or more activity classes, four or more activity classes, five or more activity classes) for each training compound for a given target polymer. For example, the single value provided by the system of the present disclosure for each respective training compound is in a first activity class if its number falls in a first range, in a second activity class if its number falls in a second range, in a third activity class if its number falls in a third range, etc. The activity class assigned by the system of the present disclosure is compared to the actual activity class as represented by the training compound binding data of other forms of training data. The error of the activity class assignment made by the system of the present disclosure is used to train the system of the present disclosure using the techniques discussed above, such that it is validated against the binding data (or other forms of measured or independently calculated data). In some embodiments, each respective classification in the plurality of classifications is determined by the IC of the training compound for the target polymer. 50 , E.C. 50 , pkA, or KI range.
[0155] In some embodiments, the classification of a plurality of training compounds by the system of the present disclosure is compared against training data (e.g., binding data or other independently measured data for the training compounds) using non-parametric techniques. For example, the system of the present disclosure is used to rank a plurality of training compounds for a given property (e.g., binding to a given target polymer), and this rank order is compared against a rank order provided by training data obtained by wet-lab binding assays of a plurality of training compounds. This gives rise to the ability to train the system of the present disclosure against errors in the calculated rank order using the system error correction techniques discussed above. In some embodiments, the error (difference) between the ranking of the training compounds by the system of the present disclosure and the ranking of the training compounds determined by the binding data (or other independently measured data for the training compounds) is calculated using the Wilcoxon Mann Whitney function (Wilcoxon signed rank test) or other non-parametric test, and such error is back-propagated through the system of the present disclosure (e.g., Model 72, Model 74, Model 24, etc.) to further train the system using the error correction techniques discussed above.
[0156] In embodiments in which deep learning techniques utilize neural network 24 as described above, training a system including network 24 may include modifying weights in filters in optional convolutional layers 28 and biases in network layers to improve the accuracy of its predictions. The weights and biases may be further constrained with various forms of regularization, such as L1, L2, weight decay, and dropout.
[0157] In an embodiment, the neural network 24 or any of the models disclosed herein may optionally have their respective parameters (e.g., weights) adjusted (to potentially minimize the error between the predicted binding affinity and / or categorization of the system and the reported binding affinity and / or categorization of the training data) when the training data is labeled (e.g., with binding data). Various methods may be used to minimize error functions such as gradient descent, which may include, but are not limited to, logarithmic loss, sum of squares error, hinge loss methods. These methods may include second-order methods or approximations such as momentum, Hessian-free estimation, Nesterov's accelerated gradient, Adagrad, etc. Unlabeled developmental pre-training and labeled discriminative training may be combined.
[0158] The input geometric data can be grouped into training examples. For example, a single set of molecules, cofactors, and proteins often has multiple geometric measurements, with each "snapshot" depicting alternative conformations and poses that the target polymer and training compounds can adopt. Similarly, if the target polymer is a protein, different tautomers of protein side chains, cofactors, and training compounds can also be sampled. Since all of these states contribute to the behavior of biological systems, a system that predicts binding affinities according to the Boltzmann distribution can be configured to consider these states together (e.g., by taking a weighted average of these samplings). Optionally, these training examples can be labeled with binding information. If quantitative binding information (e.g., binding data) is available, such a label can be a numerical binding affinity. Alternatively, the training examples may be assigned labels from a set of two or more ordered categories (e.g., two categories of binders and non-binders, or several potentially overlapping categories describing ligands as binders of sub-molar, sub-millimolar, sub-100 micromolar, sub-10 micromolar, sub-1 micromolar, sub-100 nanomolar, sub-10 nanomolar, sub-1 nanomolar potency). Training binding data may be derived or received from a variety of sources, such as experimental measurements, calculated estimates, expert insight, or guesswork (e.g., random pairs of molecules and proteins are highly unlikely to bind).
[0159] The neural network 24 is used to obtain scores for the poses. To score the poses using the neural network 24, in some embodiments, a voxel map is created for the poses (e.g., a positive voxel map 52 for positive poses and a negative voxel map 64 for negative poses 60). In some embodiments, the voxel map is created by (i) sampling the training compound at either the positive poses 48 (or collection thereof) or the negative poses (or collection thereof) and forming a corresponding three-dimensional uniform space-filling honeycomb including a corresponding plurality of space-filling (three-dimensional) polyhedral cells by sampling the target polymer 38 on a three-dimensional grid basis, and (ii) for each respective three-dimensional polyhedral cell in the corresponding plurality of three-dimensional cells, populating the respective voxel map with voxels (a discrete set of regularly spaced polyhedral cells) based on a property (e.g., a chemical property) of the respective three-dimensional polyhedral cell. Thus, for a particular training compound, two voxel maps are created: a positive voxel map 52 and a negative voxel map 65. Examples of space-filling honeycombs include a cubic honeycomb with parallelepiped cells, a hexagonal prism honeycomb with hexagonal prism cells, a rhombic dodecahedron with rhombic dodecahedral cells, an elongated dodecahedron with elongated dodecahedral cells, and a truncated octahedron with truncated octahedral cells.
[0160] In some embodiments, the space-filling honeycomb is a cubic honeycomb with cubic cells, the dimensions of such voxels determining their respective resolution. For example, a resolution of 1 Å may be selected, meaning that each voxel represents a corresponding cube of geometric data with 1 Å dimensions in such embodiments (e.g., 1 Å×1 Å×1 Å in the respective height, width, and depth of each cell). However, in some embodiments, a finer grid spacing (e.g., 0.1 Å, or even 0.01 Å) or a coarser grid spacing (e.g., 4 Å) is used, which spacing results in an integer number of voxels to cover the input geometric data. In some embodiments, sampling occurs at a resolution of 0.1 Å to 10 Å. By way of example, for a 40 Å input cube, at a resolution of 1 Å, such an arrangement would result in 40×40×40=64,000 input voxels.
[0161] In some embodiments, the atomic features generated in sampling (i) are arranged in a single voxel of the respective voxel map, with each voxel of the plurality of voxels representing at most one atomic feature. In some embodiments, the atomic features consist of an enumeration of atom types. As an example, some embodiments of the disclosed system and method are configured to represent the presence of all atoms in a given voxel of the voxel map 40 as different numbers of its entries, e.g., if carbon is present in the voxel, a value of 6 is assigned to the voxel since the atomic number of carbon is 6. However, such encoding may imply that atoms with close atomic numbers behave similarly, which may not be particularly useful in some applications. Furthermore, the behavior of elements may be more similar within a group (column of the periodic table), and thus such encoding imposes additional work for the neural network 24 to decode.
[0162] In some embodiments, atomic features are encoded in the voxels as binary categorical variables. In such embodiments, atom types are encoded in what is called "one-hot" encoding, where every atom type has a separate channel. Thus, in such embodiments, each voxel has multiple channels, where at least a subset of the multiple channels represent atom types. For example, one channel in each voxel may represent carbon, while another channel in each voxel may represent oxygen. If the three-dimensional grid element corresponding to a given voxel contains a given atom type, then the channel for that atom type in the given voxel is assigned a first value of the binary categorical variable, such as "1," and if the three-dimensional grid element corresponding to the given voxel does not contain that atom type, then the channel for that atom type in the given voxel is assigned a second value of the binary categorical variable, such as "0," in the given voxel.
[0163] There are over 100 elements, but most are not encountered in biology. However, representing even the most common biological elements (e.g., H, C, N, O, F, P, S, Cl, Br, I, Li, Na, Mg, K, Ca, Mn, Fe, Co, Zn) can result in 18 channels per voxel, or 10,483 x 18 = 188,694 inputs, to the receptor field. Thus, in some embodiments, each respective voxel in the voxel map includes multiple channels, each channel within the multiple channels representing a different property that may occur in the three-dimensional space-filling polyhedral cell corresponding to the respective voxel. The number of possible channels for a given voxel is even higher in embodiments where additional features of the atom (e.g., partial charges, presence of ligands for protein targets, electronegativity, or SYBYL atom type) are further presented as independent channels for each voxel, requiring more input channels to distinguish between otherwise equivalent atoms.
[0164] In some embodiments, each voxel has 5 or more input channels. In some embodiments, each voxel has 15 or more input channels. In some embodiments, each voxel has 20 or more input channels, 25 or more input channels, 30 or more input channels, 50 or more input channels, or 100 or more input channels. In some embodiments, each voxel has 5 or more input channels selected from the descriptors included in Table 1 below. For example, in some embodiments, each voxel has 5 or more channels, each encoded as a binary categorical variable, each such channel representing a SYBYL atom type selected from Table 1 below. For example, in some embodiments, each respective voxel in the voxel map includes a channel for a C.3 (sp3 carbon) atom type, meaning that if the grid in space of a given test object-target object (or training object-target object) complex represented by the respective voxel contains sp3 carbon, the channel adopts a first value (e.g., "1"), and otherwise is a second value (e.g., "0").
[0165] [Table 1-1]
[0166] [Table 1-2]
[0167] In some embodiments, each voxel includes 10 or more input channels, 15 or more input channels, or 20 or more input channels selected from the descriptors included in Table 1 above. In some embodiments, each voxel includes a channel for a halogen.
[0168] In some embodiments, a first Structural Protein-Ligand Interaction Fingerprint (SPLIF) score is generated for each training compound's positive pose 48 and a second SPLIF is generated for each training compound's negative pose 60. In such embodiments, these SPLIF scores are used as additional inputs to the underlying neural network or are encoded separately in the voxel map. For a description of SPLIF, see Da and Kireev, 2014, J. Chem. Inf. Model. 54, pp. 2555-2561, "Structural Protein-Ligand Interaction Fingerprints (SPLIF) for Structure-Based Virtual Screening: Method and Benchmark Study", which is incorporated herein by reference in its entirety. SPLIFs implicitly encode all possible interaction types (e.g., π-π, CH-π, etc.) that may occur between the interacting fragments of the training compound and the target polymer 38. In a first step, the intermolecular contacts of the training compound-target polymer 38 are examined. If the distance between two atoms is within a certain threshold (e.g., within 4.5 Å), they are considered to be in contact. For each such intermolecular atom pair, the respective training atom and the target polymer atom are expanded into a circular fragment, e.g., a fragment that includes the atom in question and each of its contiguous neighbors up to a certain distance. Each type of circular fragment is assigned an identifier. In some embodiments, such identifiers are encoded in individual channels within each voxel. In some embodiments, an extended connectivity fingerprint to the first nearest neighbor (ECFP2) can be used, as defined in Pipeline Pilot software. See Pipeline Pilot, ver. 8.5, Accelrys Software Inc., 2009, which is incorporated herein by reference in its entirety. ECFP holds information about all atoms / bond types and uses one unique integer identifier to represent one substructure (e.g., a circular fragment).The SPLIF fingerprint encodes all circular fragment identifiers that it contains. In some embodiments, the SPLIF fingerprint does not encode individual voxels, but serves as a separate and independent input in neural network 24, discussed below.
[0169] In some embodiments, rather than or in addition to SPLIFs, a Structural Interaction Fingerprint (SIFt) is calculated for each pose (positive pose 48 and negative pose 60) of a given training compound relative to the target polymer and provided separately as an input to the neural network 24 or encoded in the voxel map. For calculation of SIFt, see Deng et al., 2003, "Structural Interaction Fingerprint (SIFt): A Novel Method for Analyzing Three-Dimensional Protein-Ligand Binding Interactions," J. Med. Chem. 47 (2), pp. 337-344, which is incorporated herein by reference in its entirety.
[0170] In some embodiments, instead of or in addition to SPLIF and SIFT, atom pair-based interaction fragments (APIFs) are calculated for each pose (positive pose 48 and negative pose 60) of a given training compound relative to the target polymer 38 and provided separately as input to the neural network 24 or encoded separately in the voxel map. For calculation of APIFs, see Perez-Nueno et al., 2009, "APIF: a new interaction fingerprint based on atom pairs and its application to virtual screening," J. Chem. Inf. Model. 49(5), pp. 1245-1260, which is incorporated herein by reference in its entirety.
[0171] The data representation may be encoded in a manner that allows for expression of various structural relationships associated with the molecule / protein, for example. The geometric representation may be implemented in various ways and topographies according to various embodiments. The geometric representation is used for data visualization and analysis. For example, in an embodiment, the geometry may be represented using voxels laid out on various topographies, such as 2-D, 3-D Cartesian / Euclidean space, 3-D non-Euclidean space, manifolds, etc. For example, FIG. 4 shows a sample three-dimensional grid structure 400 including a series of subcontainers, according to an embodiment. Each subcontainer 402 may correspond to a voxel. A coordinate system may be defined for the grid such that each subcontainer has an identifier. In some embodiments of the disclosed systems and methods, the coordinate system is a Cartesian system in 3-D space, but in other embodiments of the system, the coordinate system may be any other type of coordinate system, such as an oblate spheroid, a cylindrical or spherical coordinate system, a polar coordinate system, other coordinate systems designed for various manifolds and vector spaces, among others. In some embodiments, the voxels may have a particular value associated with each, which may be represented, for example, by applying a label and / or determining the respective positioning, among other things.
[0172] Since neural networks require a fixed input size, some embodiments of the disclosed systems and methods trim the geometric data (target test or target training object complex) to fit within an appropriate bounding box. For example, a cube with 25-40 Å to a side may be used. In some embodiments where the target and / or test object is docketed to the active site of the target object 58, the center of the active site serves as the center of the cube.
[0173] In some embodiments, a square cube of fixed dimensions centered on the active site of the target polymer 38 is used to divide the space into a voxel grid, although the disclosed system is not so limited. In some embodiments, any of a variety of shapes are used to divide the space into a voxel grid. In some embodiments, a polyhedron, such as a rectangular prism, a polyhedral shape, etc., is used to divide the space.
[0174] In an embodiment, the grid structure may be configured to resemble an arrangement of voxels, for example, each substructure may be associated with a channel for each atom being analyzed, and an encoding method may be provided for numerically representing each atom.
[0175] In some embodiments, the voxel map takes into account the factor of time (eg, along with the training compound poses and the molecular dynamics runs of the target polymer) and therefore may be four-dimensional (X, Y, Z, and time).
[0176] In some embodiments, other implementations such as pixels, points, polygons, polyhedra, or any other type of shape of multiple dimensions (e.g., 3D, 4D, etc. shapes) may be used instead of voxels.
[0177] In some embodiments, the geometric data is normalized by selecting the origin of the X, Y, and Z coordinates to be the center of mass of the binding site of the target polymer 38, as determined by a cavity flooding algorithm. For representative details of such algorithms, see Ho and Marshall, 1990, "Cavity search: An algorithm for the isolation and display of cavity-like binding regions," Journal of Computer-Aided Molecular Design 4, pp. 337-354, and Hendlich et al., 1997, "Ligsite: automatic and efficient detection of potential small molecule-binding sites in proteins," J. Mol. Graph. Model 15:6, each of which is incorporated herein by reference in its entirety. Alternatively, in some embodiments, the origin of the voxel map is centered on the center of mass of the entire co-complex (of the training compound docked at its respective pose - positive pose 48 or negative pose 60 - bound to the target polymer). In some embodiments, the origin of the voxel map is centered on the center of mass of the training compound. In some embodiments, the origin of the voxel map is centered on the center of mass of the target polymer 38. The basis vectors can optionally be selected to be the principal moments of inertia of the entire co-complex, the target polymer only, or the training compound only. In some embodiments, the target polymer 38 has an active site, and the sampling samples the training compound at both positive poses 48 and negative poses 60, with the center of mass of the active site taken as the origin, and a corresponding three-dimensional uniform honeycomb for sampling the active site on a three-dimensional grid basis representing the portion of the polymer and training compound centered on the center of mass. In some embodiments, the uniform honeycomb is a regular cubic honeycomb, and the polymer and portion of the test object are cubes of predetermined fixed dimensions.In such embodiments, the use of a cube of predetermined fixed dimensions ensures that relevant portions of the geometric data are used to ensure that each voxel map is the same size. In some embodiments, the predetermined fixed dimensions of the cube are NÅ×NÅ×NÅ, where N is an integer or real number between 5 and 100, an integer between 8 and 50, or an integer between 15 and 40. In some embodiments, the uniform honeycomb is a rectangular prism honeycomb, and the polymer and training compound portions are rectangular prisms of predetermined fixed dimensions QÅ×RÅ×SÅ, where Q is a first integer between 5 and 100, R is a second integer between 5 and 100, and S is a third integer or real number between 5 and 100, and at least one number in the set {Q,R,S} is not equal to another value in the set {Q,R,S}.
[0178] In an embodiment, every voxel has one or more input channels, which may have various values associated with them, which in a simple implementation may be on / off, and may be configured to code for a certain type of atom. The atom type may indicate the element of the atom, or the atom type may be further refined to distinguish characteristics of other atoms. The atoms present may then be coded at each voxel. The various types of coding may be utilized using various techniques and / or methodologies. As an exemplary coding method, the atomic number of the atom may be utilized, resulting in one value per voxel from 1 for hydrogen to 118 for ununoctium (or any other element).
[0179] However, as discussed above, other encoding methods may be utilized, such as "one-hot encoding," where each voxel has many parallel input channels, each of which is either on or off and encodes for a certain type of atom. The atom types may indicate the elemental character of the atom, or the atom types may be further refined to distinguish other atomic characteristics. For example, the SYBYL atom type distinguishes a singly bonded carbon from a doubly bonded carbon, a triple bonded carbon, or an aromatic carbon. For a description of the SYBYL atom types, see Clark et al., 1989, "Validation of the General Purpose Tripos Force Field, 1989, J. Comput. Chem. 10, pp. 982-1012, which is incorporated herein by reference.
[0180] In some embodiments, each voxel further comprises one or more channels for distinguishing between atoms that are part of the target polymer 38 or cofactors to parts of the training compound. For example, in one embodiment, each voxel further comprises a first channel for the target polymer 38 and a second channel for the training compound. The first channel is set to a value such as "1" if the atoms in the portion of the space represented by the voxel are from the target polymer 38, and is zero otherwise (e.g., because the portion of the space represented by the voxel does not contain any atoms or contains one or more atoms from the training compound). Furthermore, the second channel is set to a value such as "1" if the atoms in the portion of the space represented by the voxel are from the training compound, and is zero otherwise (e.g., because the portion of the space represented by the voxel does not contain any atoms or contains one or more atoms from the target polymer 38). Similarly, the other channels may additionally (or alternatively) specify further information such as partial charges, polarizability, electronegativity, solvent accessible space, and electron density. For example, in some embodiments, the electron density map of the target object covers a set of three-dimensional coordinates, and the creation of the voxel map further samples the electron density map. Examples of suitable electron density maps include, but are not limited to, multiple isomorphous displacement maps, single isomorphous displacement with anomalous signal maps, single wavelength anomalous dispersion maps, multi-wavelength anomalous dispersion maps, and 2Fo-Fc maps (260). See McRee, 1993, Practical Protein Crystallography, Academic Press, which is incorporated herein by reference.
[0181] In some embodiments, voxel encoding according to the disclosed systems and methods may include additional optional encoding refinements. The following two are provided as examples:
[0182] In a first encoding refinement, the required memory can be reduced by reducing the set of atoms represented by a voxel (e.g., by reducing the number of channels represented by a voxel) based on the fact that most elements occur rarely in biological systems. Atoms can be mapped to share the same channel in a voxel either by combining rare atoms (thus, may have little impact on the performance of the system) or by combining atoms with similar properties (thus, may minimize inaccuracies from combinations). In some embodiments, 2, 3, 4, 5, 6, 7, 8, 9, or 10 different atoms share the same channel in a voxel.
[0183] An improvement to the encoding is to allow a voxel to represent the location of an atom by partially activating neighboring voxels. This results in partial activation of neighboring neurons in the subsequent neural network, moving from one-hot encoding to "several-warm" encoding. For example, 3 When the grid is aligned, the van der Waals diameter is 3.5 Å and therefore the volume is 22.4 Å. 3It may be exemplary to consider a chlorine atom where voxels within the chlorine atom are completely filled and voxels at the edge of the atom are only partially filled. Thus, channels representing chlorine in partially filled voxels are turned on in proportion to the amount of such voxels that fall within the chlorine atom. For example, if 50% of the voxel volume falls within the chlorine atom, then the channels in the voxel representing chlorine are activated 50%. This may result in a "smoothed" and more accurate representation relative to discrete one-hot coding. Thus, in some embodiments, the atomic features generated in the sampling are distributed to a subset of voxels in the voxel map, where this subset of voxels includes 2 or more voxels, 3 or more voxels, 5 or more voxels, 10 or more voxels, or 25 or more voxels. In some embodiments, the atomic features consist of an enumeration of atom types (e.g., one of the SYBYL atom types).
[0184] Thus, the voxelization (rasterization) of the encoded geometric data (docking of test or training objects to a target object) is based on various rules applied to the input data.
[0185] Figures 5 and 6 provide illustrations of two molecules 502 encoded on a two-dimensional grid 500 of voxels, according to some embodiments. Figure 5 provides two molecules superimposed on the two-dimensional grid. Figure 6 provides one-hot encoding, using different diagonal patterns to encode the presence of oxygen, nitrogen, carbon, and free space, respectively. As mentioned above, such encoding may be referred to as "one-hot" encoding. Figure 6 shows the grid 500 of Figure 5 with the molecules 502 omitted. Figure 7 provides an illustration of the two-dimensional grid of voxels of Figure 6 with the voxels numbered.
[0186] In some embodiments, feature shapes are represented in forms other than voxels. Figure 8 provides illustrations of various representations in which features (e.g., atomic centers) are represented as 0-D points (representation 802), 1-D points (representation 804), 2-D points (representation 806), or 3-D points (representation 808). Initially, the spacing between the points may be chosen randomly. However, as the predictive model is trained, the points may move closer together or further apart.
[0187] In some embodiments, the input representation may be in the form of a 1D array of features, including but not limited to three-dimensional coordinates.
[0188] In some embodiments, neural network 24 is a graph superposition neural network. Non-limiting examples of graph convolutional neural networks include Behler Parrinello, 2007, “Generalized Neural-Network Representation of High Dimensional Potential-Energy Surfaces,” Physical Review Letters 98, 146401; Chmiela et al., 2017, “Machine learning of accurate energy-conserving molecular force fields,” Science Advances 3(5): e1603015; Schuett et al., 2017, “SchNet: A continuous-filter convolutional neural network for modeling quantum interactions,” Advances in Neural Information Processing Systems 30, pp. 992-1002; Feinberg et al., 2018, “PotentialNet for Molecular Property Prediction,” ACS Cent. Sci. 4, 11, 1520-1530; and Stafford et al., “AtomNet PoseRanker: Enriching Ligand Pose Quality for Dynamic Proteins in Virtual High Throughput Screens,” https: / / chemrxiv.org / engage / chemrxiv / article-details / 614b905e39ef6a1c36268003, each of which is incorporated herein by reference in its entirety.
[0189] In some embodiments, the neural network is an equivariant neural network. Non-limiting examples of equivariant convolutional neural networks include Thomas et al., 2018, “Tensor field networks: Rotation- and translation-equivariant neural networks for 3D point clouds,” arXiv:1802.08219; Anderson et al., 2019, “Cormorant: Covariant Molecular Neural Networks,” Neural Information Processing Systems; Johannes et al., 2020, “Directional Message Passing For Molecular Graphs,” International Conference on Learning Representations; Townshend et al., 2021, “ATOM3D: Tasks On Molecules in Three Dimensions,” International Conference on Learning Representations; Jing et al., 2009, “Learning from Protein Structure with Geometric Vector Perceptrons,” arXiv:2009.01411; and Satorras et al., 2021, “E(n) Equivariant Graph Neural Networks,” arXiv:2102.09844, each of which is incorporated herein by reference in its entirety.
[0190] In some embodiments, the neural network 24 is any of the graph neural networks disclosed in U.S. Provisional Patent Application No. 63 / 336,841, entitled “Characterization of Interactions Between Compounds and Polymers Using Pose Ensembles,” filed May 10, 2022, which is incorporated by reference herein.
[0191] The voxel maps are expanded into corresponding vectors. Each voxel map (e.g., positive voxel map 52 and negative voxel map 64) is optionally expanded into a corresponding vector (e.g., positive vector 54 and negative vector 66 for each training compound in the training dataset 40). In some embodiments, each such vector is a one-dimensional vector. For example, in some embodiments, a 20 Å cube on each side is centered on the active site of the target polymer 38 and sampled with a three-dimensional fixed grid spacing of 1 Å to form corresponding voxels of the voxel map that are applied to each basic channel of voxel structural features such as atom type and, optionally, more complex training compound-target polymer descriptors, as discussed above. In some embodiments, the voxels of this three-dimensional voxel map are expanded into one-dimensional floating point vectors.
[0192] In some embodiments, the vectorized representation of the voxel map (e.g., the positive vector 54 and the negative vector 66 for each training compound in the training dataset 40) is provided to the neural network 24. In some embodiments, as shown in FIG. 1B, the vectorized representation of the voxel map is stored in the GPU memory 52 along with the assessment module 20 and the neural network 24. This provides the advantage of processing the vectorized representation of the voxel map at a faster rate through the neural network 24. However, in other embodiments, any or all of the vectorized representation of the voxel map (e.g., the positive vector 54 and the negative vector 66 for each training compound in the training dataset 40), the assessment module 20, and the neural network 24 are within the memory 92 of the system 100 or are simply addressable by the system 92 across a network. In some embodiments, any or all of the vectorized representation of the voxel map, the assessment module 20, and the neural network 24 are in a cloud computing environment.
[0193] In some embodiments, the vectors (e.g., positive vector 54 and negative vector 66 for each training compound of the training dataset 40) are provided to a graphic processing unit memory 52, which includes a network architecture including a neural network 24 including an input layer 26 for sequentially receiving a plurality of vectors, optionally a plurality of convolutional layers 28, and a scorer 30. In some embodiments, any of the plurality of convolutional layers includes an initial convolutional layer and a final convolutional layer. In some embodiments, the neural network 24 is not in a GPU memory, but in a general purpose memory of the system 100. In some embodiments, the voxel maps are not vectorized before being input to the network 24.
[0194] Making the Convolutional Layer 28 User In some embodiments, the convolutional layer 28 of the multiple convolutional layers comprises a set of learnable filters (also called kernels). Each filter has a fixed three-dimensional size that is convolved (stepped at a predetermined step rate) across the depth, height, and width of the input volume of the convolutional layer, computing a dot product (or other function) between the entries (weights, or more generally, parameters) of the filter and the input, thereby creating a multidimensional activation map for that filter. In some embodiments, the step rate of the filter is 1 element, 2 elements, 3 elements, 4 elements, 5 elements, 6 elements, 7 elements, 8 elements, 9 elements, 10 elements, or more than 10 elements of the input space. Thus, when a filter has a size of 5 3 In some embodiments, the filter computes the dot product (or other mathematical function) between consecutive cubes of input space with a depth of 5 elements, a width of 5 elements, and a height of 5 elements, for a total of 125 input space values per voxel channel.
[0195] The input space to the initial convolutional layer (e.g., output from the input layer 26) is formed from either the voxel map or a vectorized representation of the voxel map (e.g., the positive vector 54 and the negative vector 66 for each training compound in the training data set 40). In some embodiments, the vectorized representation of the voxel map is a one-dimensional vectorized representation of the voxel map that serves as the input space to the initial convolutional layer. Nonetheless, when the filter convolves its input space, which is a one-dimensional vectorized representation of the voxel map, the filter still obtains elements from the one-dimensional vectorized representation that represent a corresponding contiguous cube of the fixed space of the target polymer 38-training compound complex. In some embodiments, the filter uses bookkeeping techniques to select those elements from among the one-dimensional vectorized representation that form a corresponding contiguous cube of the fixed space of the target polymer 38-training compound complex. Thus, in some instances, this entails taking a non-contiguous subset of the elements in the one-dimensional vectorized representation to obtain the element values of the corresponding contiguous cube of the fixed space of the target polymer 38-training compound complex.
[0196] In some embodiments, the filter is initialized (e.g., to Gaussian noise) or trained to have 125 corresponding weights (per input channel) for taking a dot product (or some other form of mathematical operation, such as the function disclosed in FIG. 14) of the 125 input spatial values to calculate a first single value (or set of values) of the activation layer corresponding to the filter. In some embodiments, the values calculated by the filter are summed, weighted, and / or biased. To calculate the sum value of the activation layer corresponding to the filter, the filter is then stepped (convolved) in one of the three dimensions of the input volume by a step rate (stride) associated with the filter, where a dot product (or some other form of mathematical operation, such as the mathematical function disclosed in FIG. 17) of the filter weights and the 125 input spatial values (per channel) is taken at a new location within the input volume. This stepping (convolution) is repeated until the filter has sampled the entire input space according to the step rate. In some embodiments, the edges of the input space are zero-padded to control the spatial volume of the output space generated by the convolution layer. In typical embodiments, each of the filters of a convolutional layer thus canvasses the entire three-dimensional input volume, thereby forming a corresponding activation map. The collection of activation maps from the filters of a convolutional layer collectively forms the three-dimensional output volume of one convolutional layer, thereby serving as the three-dimensional (three spatial dimensions) input of a subsequent convolutional layer. Thus, every entry in the output volume can also be interpreted as the output of a single neuron (or set of neurons) that looks at a small region in the input space to the convolutional layer and shares parameters with neurons in the same activation map. Thus, in some embodiments, a convolutional layer of multiple convolutional layers has multiple filters, each filter of the multiple filters having a stride of N with stride Y (in the three spatial dimensions). 3 where N is an integer greater than or equal to 2 (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10) and Y is a positive integer (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10).
[0197] Each of the multiple convolutional layers is associated with a different set of weights, or more generally, a different set of parameters. More specifically, each of the multiple convolutional layers includes multiple filters, each filter including independent parameters (e.g., weights). In some embodiments, the convolutional layers are of dimension 5 or more. 3 For a given convolutional layer, the convolutional layer has 128 filters, i.e., 128×5×5×5 or 16,000 parameters (e.g., weights) per channel in the voxel map. Thus, if there are five channels in the voxel map, the convolutional layer has 16,000×5 parameters (e.g., weights), i.e., 80,000 parameters (e.g., weights). In some embodiments, some or all such parameters (and, optionally, biases) of all filters in a given convolutional layer may be tied together, e.g., constrained to be the same.
[0198] In response to input of a respective vector (e.g., positive vector 54 or negative vector 66), input layer 26 provides a first plurality of values to the initial convolutional layer as a first function of the values in the respective vector, the first function optionally being calculated using graphics processing unit 50. In some embodiments, computer system 100 has two or more graphics processing units 50.
[0199] Each respective convolutional layer 28 other than the final convolutional layer provides intermediate values to another of the multiple convolutional layers as a second function of (i) a different set of parameters (e.g., weights) associated with the respective convolutional layer and (ii) a respective input value received by the respective convolutional layer. In some embodiments, the second function is calculated using the graphics processing unit 50. For example, in some embodiments, each respective filter of each convolutional layer 28 canvases (in three spatial dimensions) the input volume to the convolutional layer according to the characteristic three-dimensional stride of the convolutional layer, and at each respective filter location, takes a dot product (or some other mathematical function) of the filter parameters (e.g., weights) of the respective filter and the values of the input volume at the respective filter location (a contiguous cube that is a subset of the total input space), thereby generating a calculated point (or set of points) on the activation layer that corresponds to the respective filter location. The activation layers of the filters of each convolutional layer collectively represent the intermediate values of the respective convolutional layer.
[0200] The final convolutional layer supplies a final value to the scorer as a function of (i) a set of distinct parameters (e.g., weights) associated with the final convolutional layer and (ii) optionally a third function of the input values received by the final convolutional layer calculated using the graphics processing unit 50. For example, each respective filter of the final convolutional layer 28 canvases (in three spatial dimensions) the input volume to the final convolutional layer according to the characteristic three-dimensional stride of the convolutional layer, and at each respective filter location, takes the dot product (or some other mathematical function) of the filter's filter weight and the value of the input volume at the respective filter location, thereby calculating a calculated point (or set of points) on the activation layer that corresponds to the respective filter location. The activation layers of the filters of the final convolutional layer collectively represent the final value that is supplied to the scorer 30.
[0201] In some embodiments, the convolutional neural network has one or more activation layers. In some embodiments, the activation layer is a layer of neurons that applies a non-saturating activation function f(x)=max(0,x), which increases the non-linearity of the decision function and the overall network without affecting the receptive field of the convolutional layer. In other embodiments, the activation layer applies other functions to increase the non-linearity, such as the saturated hyperbolic tangent function f(x)=tanh, f(x)=│tanh(x)│, and the sigmoid function f(x)=(1+e -x ) -1 In some embodiments for the neural network, non-limiting examples of other activation functions included in other activation layers may include, but are not limited to, logistic (or sigmoid), softmax, Gaussian, Boltzmann weighted averaging, absolute value, linear, rectified linear, bounded rectified linear, soft rectified linear, parameterized rectified linear, mean, maximum, minimum, any vector norm LP (for p=1, 2, 3, ..., ∞), sign, square, square root, biquadratic, inverse quadratic, inverse biquadratic, multiharmonic spline, and thin plate spline.
[0202] Network 24 learns filters in convolutional layers 28 that activate when they see some particular type of feature at some spatial location in the input. In some embodiments, the initial parameters (e.g., weights) of each filter in the convolutional layers are obtained by training the convolutional neural network on a compound training library. Thus, operation of the convolutional neural network 24 may yield more complex features than those historically used to perform binding affinity predictions. For example, a filter in a given convolutional layer of network 24 acting as a hydrogen bond detector may be able to not only recognize that hydrogen bond donors and acceptors are at a given distance and angle, but also recognize that the biochemical environment around the donor and acceptor strengthens or weakens the bond. Additionally, the filters in network 24 may be trained to effectively distinguish between binders and non-binders in the underlying data.
[0203] As described above, in some embodiments, neural network 24 is configured to form three-dimensional convolutional layers. The input region to the lowest level convolutional layer 28 may be a cube (or other contiguous region) of voxel channels from the receptive field. Higher convolutional layers 28 evaluate the outputs from lower convolutional layers, but still make each output a function of a bounded region of voxels that are close to each other (in 3-D Euclidean distance).
[0204] In one embodiment, network 24 is configured to apply regularization techniques to reduce the tendency of the model to overfit the training data.
[0205] Zero or more of the network layers in network 24 may consist of pooling layers. Similar to convolutional layers, pooling layers are a set of function computations that apply the same function to different spatially localized patches of the input. For pooling layers, the output is given by a pooling operator, e.g., some vector norm LP for p=1, 2, 3, ..., ∞, over several voxels. Pooling is typically done per channel, not across channels. Pooling divides the input space into a set of three-dimensional boxes, and for each such sub-region, outputs the maximum value. The pooling operation provides a form of translation invariance. The function of pooling layers is to progressively reduce the spatial size of the representation to reduce the amount of parameters and computations in the network, and thus also to control overfitting. In some embodiments, pooling layers are inserted between successive convolutional layers 28 in network 24. Such pooling layers operate independently on each depth slice of the input, varying the size spatially. In addition to max pooling, the pooling unit can also perform other functions such as average pooling or L2-norm pooling.
[0206] Zero or more layers in network 24 may consist of normalization layers, such as local response normalization or local contrast normalization, which may be applied across channels at the same location or to specific channels across several locations. These normalization layers may promote diversity in the response of several function calculations to the same input.
[0207] In some embodiments, the scorer 30 includes multiple fully-connected layers and an evaluation layer where the fully-connected layers of the multiple fully-connected layers feed into the evaluation layer. As in a regular neural network, the neurons of a fully-connected layer have full connections to all activations of the previous layer. Thus, their activations can be calculated by matrix multiplication followed by a bias offset. In some embodiments, each fully-connected layer has 512 hidden units, 1024 hidden units, or 2048 hidden units. In some embodiments, the scorer has no fully-connected layers, one fully-connected layer, two fully-connected layers, three fully-connected layers, four fully-connected layers, five fully-connected layers, six or more fully-connected layers, or ten or more fully-connected layers.
[0208] In some embodiments, the evaluation tier distinguishes between multiple activity classes, hi some embodiments, the evaluation tier includes a logistic regression cost tier spanning two activity classes, three activity classes, four activity classes, five activity classes, or six or more activity classes.
[0209] In some embodiments, the evaluation tier includes a logistic regression cost tier across multiple activity classes, hi some embodiments, the evaluation tier includes a logistic regression cost tier across two activity classes, three activity classes, four activity classes, five activity classes, or six or more activity classes.
[0210] In some embodiments, the evaluation layer distinguishes between two activity classes, the first activity class (first classification) being a training compound with respect to the target polymer that has an IC of greater than a first binding value. 50 , E.C. 50or KI, and the second activity class (second classification) is an IC of the training compound with respect to the target polymer that is less than the first binding value. 50 , E.C. 50 In some embodiments, the first binding value is 1 nanomolar, 10 nanomolar, 100 nanomolar, 1 micromolar, 10 micromolar, 100 micromolar, or 1 millimolar.
[0211] In some embodiments, the evaluation layer includes a logistic regression cost layer across two activity classes, where the first activity class (first classification) is the training compound's IC for the target polymer that exceeds a first binding value. 50 , E.C. 50 or KI, and the second activity class (second classification) is an IC of the training compound with respect to the target polymer that is less than the first binding value. 50 , E.C. 50 In some embodiments, the first binding value is 1 nanomolar, 10 nanomolar, 100 nanomolar, 1 micromolar, 10 micromolar, 100 micromolar, or millimolar.
[0212] In some embodiments, the evaluation layer distinguishes between three activity classes, a first activity class (first classification) being a training compound with respect to the target polymer that has an IC of greater than a first binding value. 50 , E.C. 50 or KI, and the second activity class (second classification) is the IC of the training compound with respect to the target polymer that is between the first binding value and the second binding value. 50 , E.C. 50 or KI, and the third activity class (third classification) is an IC of the training compound for the target polymer that is less than the second binding value. 50 , E.C. 50 or KI, where the first bond value is other than the second bond value.
[0213] In some embodiments, the evaluation layer includes a logistic regression cost layer across three activity classes, where a first activity class (first classification) is a training compound with an IC for the target polymer that exceeds a first binding value. 50 , E.C. 50or KI, and the second activity class (second classification) is the IC of the training compound with respect to the target polymer that is between the first binding value and the second binding value. 50 , E.C. 50 or KI, and the third activity class (third classification) is an IC of the training compound for the target polymer that is less than the second binding value. 50 , E.C. 50 or KI, where the first bond value is other than the second bond value.
[0214] In some embodiments, the scorer 30 comprises a fully connected single or multi-layer perceptron. In some embodiments, the scorer comprises a support vector machine, random forest, nearest neighbor. In some embodiments, the scorer 30 assigns a numerical score indicating the strength (or certainty or probability) of classifying the inputs into various output categories. In some examples, the categories are classified according to binders and non-binders, or alternatively, potency levels (e.g., IC<1 molar, <1 millimolar, <100 micromolar, <10 micromolar, <1 micromolar, <100 nanomolar, <10 nanomolar, <1 nanomolar). 50 , E.C. 50 or KI efficacy).
[0215] Use case. Below are sample use cases, provided for illustrative purposes only, illustrating some applications of some embodiments of the present disclosure. Other uses may be contemplated, and the examples provided below are non-limiting and may be subject to variations, omissions, or may include additional elements.
[0216] Hit Discovery. Pharmaceutical companies spend millions of dollars screening compounds to discover novel promising drug leads. Large compound collections are tested to discover the few compounds that have any interaction with the disease target of interest. Unfortunately, wet lab screening is subject to experimental error, and in addition to the cost and time to perform assay experiments, the collection of large screening collections imposes significant challenges through storage constraints, shelf stability, or chemical costs. Even the largest pharmaceutical companies only have hundreds of thousands to millions of compounds, versus tens of millions of commercially available molecules and hundreds of millions of simulable molecules.
[0217] A potentially more efficient alternative to physical experimentation is virtual high-throughput screening. Just as physics simulations can help aerospace engineers evaluate possible wing designs before the models are physically tested, computational molecular screening can focus experimental testing on a small subset of highly promising molecules. This can reduce the cost and time of screening, reduce false negatives, improve success rates, and / or cover a broader swath of the chemical environment.
[0218] In this application, a protein target may be provided as an input to the system. A large set of molecules may also be provided. For each molecule, a binding affinity to the protein target is predicted. The resulting scores may be used to rank the molecules, with the highest scoring molecules being most likely to bind the target protein. Optionally, the ranked molecule list may be analyzed for clusters of similar molecules, or larger clusters may be used as stronger predictors of molecular binding, or molecules may be selected across clusters to ensure diversity in validation experiments.
[0219] Off-target side effects prediction. Many drugs can be found to have side effects. These side effects are often due to interactions with biological pathways other than the pathway responsible for the drug's therapeutic effect. These off-target side effects can be unpleasant or dangerous and may limit the patient population in which the drug is safe to use. Therefore, off-target side effects are an important criterion for evaluating which drug candidates to further develop. Although it is important to characterize the interactions of drugs with many alternative biological targets, such tests can be expensive and time-consuming to develop and perform. Computational prediction can make such processes more efficient.
[0220] In applying embodiments of the present invention, a panel of biological targets associated with important biological responses and / or side effects can be constructed. The system can then be configured to predict binding to each protein in the panel in turn. Strong activity against a particular target (i.e., activity as strong as a compound known to activate off-target proteins) may implicate the molecule in side effects due to off-target effects.
[0221] Toxicity prediction. Toxicity prediction is a special case of off-target side effect prediction that is of particular importance. Approximately half of drug candidates in late-stage clinical trials fail due to unacceptable toxicity. As part of the new drug approval process (and before a drug candidate can be tested in humans), the FDA requires toxicity test data for a set of targets including cytochrome P450 liver enzymes (whose inhibition may lead to toxicity from drug-drug interactions) or the hERG channel (whose binding may lead to QT prolongation leading to ventricular arrhythmias and other adverse cardiac effects).
[0222] In toxicity prediction, the system identifies off-target proteins as major anti-targets (e.g., CYP450, hERG, or 5-HT 2BThe molecules may be configured to constrain the proteins to be targets (receptors). Binding affinities for drug candidates may then be predicted for these proteins. Optionally, the molecules may be analyzed to predict a set of metabolites (subsequent molecules produced by the body during metabolism / degradation of the original molecule), which may also be analyzed for binding to the anti-target. Problematic molecules may be identified and modified to avoid toxicity, or development on the molecular series may be halted to avoid the expenditure of additional resources.
[0223] Optimizing Potency. One of the primary requirements for a drug candidate is strong binding to its disease target. Screening rarely finds a compound that binds strongly enough to be clinically effective. Thus, initial compounds are subjected to a lengthy process of optimization in which medicinal chemists iteratively modify the molecular structure to propose new molecules with increased strength of target binding. Each new molecule is synthesized and tested to determine whether the changes successfully improved binding. Systems can be configured to facilitate this process by replacing physical testing with computational predictions.
[0224] In this application, a disease target and a set of lead molecules can be input into the system. The system can be configured to generate a lead set binding affinity prediction. Optionally, the system can highlight differences between the candidate molecules that may help inform the reasons for predicted differences in binding affinity. A medicinal chemist user can use this information to suggest a new set of molecules with hopefully improved activity against the target. These new alternative molecules can be analyzed in the same manner.
[0225] Optimization of selectivity. As discussed above, molecules tend to bind many proteins with different strengths. For example, the binding pockets of protein kinases (which are well-known chemotherapy targets) are very similar, and most kinase inhibitors affect many different kinases. This means that many biological pathways are simultaneously modified, resulting in a "dirty" pharmaceutical profile and many side effects. Thus, a key challenge in the design of many drugs is not the activity itself, but the specificity: the ability to selectively target one protein (or a subset of proteins) from a set of possibly closely related proteins.
[0226] The system can reduce the time and cost to optimize the selectivity of a candidate drug. In this application, the user inputs two sets of proteins. One set describes the proteins on which the compound should be active, and the other set describes the proteins on which the compound should be inactive. The system can be configured to make molecular predictions for all of the proteins in both sets and establish a profile of interaction strengths. Optionally, these profiles can be analyzed to suggest explanatory patterns in the proteins. The user can use the information generated by the system to consider structural changes to the molecules that improve the relative binding to the different protein sets and design new candidate molecules with better specificity. Optionally, the system can be configured to highlight differences between the candidate molecules that may help inform the reasons for predicted differences in selectivity. The proposed candidates can be analyzed iteratively to further refine the specificity of their respective activity profiles.
[0227] Fitness Functions for Automated Molecular Design: Automated tools to perform the aforementioned optimizations are invaluable. Successful molecules require optimization and balance between potency, selectivity, and toxicity. "Scaffold hopping" (when the activity of the lead compound is retained but the chemical structure is significantly altered) can result in improved pharmacokinetic, pharmacodynamic, toxicity, or intellectual property profiles. Algorithms such as random generation of molecules, growing molecular fragments to fill a given binding site, genetic algorithms to "mutate" and "cross-breed" populations of molecules, and replacing portions of molecules with bioisosteric substitutions exist to iteratively suggest new molecules. Drug candidates generated by each of these methods must be evaluated against the multiple objectives described above (potency, selectivity, toxicity), and can be incorporated into an automated molecular design system, just as the technology can be beneficial for each of the aforementioned manual settings (binding prediction, selectivity, side effects, and toxicity prediction).
[0228] Repurposing of Drugs. All drugs have side effects, and sometimes these side effects are beneficial. The best-known example may be aspirin, which is commonly used as a headache treatment, but is also used for cardiovascular health. Drug repositioning can greatly reduce the cost, time, and risk of drug discovery, since the drug has already been shown to be safe in humans and is optimized for rapid absorption and favorable stability in patients. Unfortunately, drug repositioning is largely accidental. For example, sildenafil (Viagra) was developed as a blood pressure lowering drug and serendipitously observed to be an effective treatment for erectile dysfunction. Computational prediction of off-target effects can be used in the context of drug repurposing to identify compounds that can be used to treat alternative diseases.
[0229] In this application, similar to off-target side effect prediction, the user may assemble a set of possible target proteins, each protein linked to a disease; that is, inhibition of each protein would treat a (possibly different) disease. For example, an inhibitor of cyclooxygenase-2 can alleviate inflammation, while an inhibitor of factor Xa can be used as an anticoagulant. These proteins are annotated with the binding affinity of approved drugs, if present. A set of molecules is then assembled, limiting the set of molecules to those approved or investigated for human use. Finally, for each pair of protein and molecule, the user may use the system to predict the binding affinity. Candidates for repurposing of drugs may be identified if the predicted binding affinity of the molecule is close to the binding affinity of the effective drug for the protein.
[0230] Drug Resistance Prediction. Drug resistance is an inevitable consequence of drug use that puts selective pressure on pathogen populations to rapidly divide and mutate. Drug resistance is seen in diverse pathogens such as viruses (HIV), exogenous microbes (MRSA), and dysregulated host cells (cancer). Over time, a given drug becomes ineffective, whether the drug is an antibiotic or chemotherapy. At that point, intervention can shift to a different, hopefully still more potent drug. In HIV, there is a well-known path of disease progression defined by the mutations the virus accumulates while the patient is being treated.
[0231] There is considerable interest in predicting how pathogens will adapt to medical interventions. One approach is to characterize which mutations will arise in pathogens during treatment. Specifically, a drug's protein target needs to mutate to avoid binding the drug while at the same time continuing to bind its natural substrate.
[0232] In this application, a set of possible mutations of the target protein can be proposed. For each mutation, the resulting protein shape can be predicted. For each of these mutant protein forms, the system can be configured to predict the binding affinity to both the natural substrate and the drug. Mutations that cause the protein to no longer bind the drug but continue to bind to the natural substrate are candidates for conferring drug resistance. These mutated proteins can be used as targets for designing drugs, for example, by using these proteins as inputs into one of these other predictive use cases.
[0233] Personalized medicine. Ineffective drugs should not be administered. In addition to cost and effort, all drugs have side effects. Moral and economic considerations make it essential to give drugs only when the benefits outweigh these harms. It can be important to be able to predict when a drug will be useful. People differ from each other by a small number of mutations. However, small mutations can have immense effects. When these mutations occur at active (orthosteric) or regulatory (allosteric) sites of the disease target, they may prevent the drug from binding and thus inhibit the drug's activity. When the protein structure of a particular person is known (or predicted), the system can be configured to predict whether a drug will be effective, or the system can be configured to predict when a drug will not work.
[0234] In this application, the system may be configured to receive as input the chemical structure of a drug and a particular expressed protein of a particular patient. The system may be configured to predict binding between the drug and the protein, and if the predicted binding affinity of the drug is too weak for a particular patient's protein structure to be clinically effective, the clinician or practitioner may prevent the drug from being unprofitably prescribed to the patient.
[0235] Clinical Trial Design. This application generalizes the personalized medicine use case above to the case of patient populations. When the system can predict whether a drug will be effective for a particular patient phenotype, this information can be used to help design clinical trials. By excluding patients whose particular disease targets cannot be sufficiently affected by the drug, clinical trials can achieve statistical power using fewer patients. Fewer patients directly reduce the cost and complexity of clinical trials.
[0236] In this application, the user may divide a potential patient population into subpopulations characterized by the expression of different proteins (e.g., due to mutations or isoforms). The system may be configured to predict the binding strength of a drug candidate to different protein types. If the predicted binding strength to a particular protein type indicates a required drug concentration below the clinically achievable concentration in the patient (e.g., as based on physical characterization in test tubes, animal models, or healthy volunteers), the drug candidate is predicted to fail for that protein subpopulation. Patients with that protein may then be excluded from the clinical trial.
[0237] Pesticide design. In addition to pharmaceutical uses, the agrochemical industry uses binding predictions in the design of new pesticides. For example, one requirement for an insecticide is to stop a single species of interest without adversely affecting any other species. For environmental safety, one would want to kill weevils without killing bumblebees.
[0238] In this application, the user can input a set of protein structures from the different species under consideration into the system. A subset of proteins can be designated as proteins against which the molecule is active, while the rest can be designated as proteins against which the molecule should be inactive. As in the previous use case, some set of molecules (whether in an existing database or newly generated) will be considered for each target, and the system will return molecules that have the greatest effect on the first group of proteins while avoiding the second group.
[0239] Materials Science. It can be useful to analyze molecular interactions to predict the behavior and properties of new materials. For example, to study solvation, a user can input a repeating crystal structure of a given small molecule and evaluate the binding affinity of another instance of the small molecule on the surface of the crystal. To study polymer strength, a set of polymer strands can be input similar to a protein target structure, and polymer oligomers can be input as small molecules. Thus, the binding affinity between the polymer strands can be predicted by the system.
[0240] Simulation. Simulators often measure the binding affinity of molecules to proteins, since the tendency of a molecule to remain in a region of a protein correlates with the binding affinity of the protein. Accurate descriptions of the features governing binding can be used to identify regions and poses with particularly high or low binding energy. The energy descriptions can be folded into Monte Carlo simulations to account for the molecular motions and occupancy of protein-binding regions. Similarly, stochastic simulators for studying and modeling systems biology can benefit from accurate predictions of how small changes in molecular concentrations affect biological networks. EXAMPLES
[0241] AtomNet® Carbon: Learning physics and geometry gives pause sensitivity to a structure-based virtual high-throughput screening architecture.
[0242] Molecular bioactivity is a property of the ensemble and is determined by the enthalpy and entropy components of receptor-compound complex formation. Structure-based deep learning methods are successful in predicting activity, but can be insensitive to the docked pose, making hit detection less reliable. Furthermore, structure-based deep learning methods often ignore the entropy contribution to the change in free energy. The ensemble approach is successful when the ensemble is sensitive to the pose. In this example, a deep learning multitask architecture with increased sensitivity to the docked pose is described.
[0243] 1. Introduction Vast on-demand chemical libraries such as ENAMINE or Mcule have revolutionized the scale of drug, structure-based virtual high-throughput screening (vHTS) campaigns [1]. To identify "hits" from a library of candidate molecules, structure-based virtual screening methods predict the binding affinity between protein and ligand from their respective docked binding complexes, thereby assuming that the experimentally observed affinity correlates with the protein-ligand interaction. Traditional methods use empirical physics-based approaches that attempt to calculate the binding free energy of complex formation. In contrast, machine learning (ML) and deep learning (DL) approaches are trained on large datasets using explicit (ML) or implicit (DL) features and labels to predict activity. These statistical models generally outperform physics-based approaches in retrospective testing to predict activity.
[0244] Early structure-based DL methods of vHTS, centered on convolutional neural networks (CNNs), represent protein-ligand structures by 3D grids to predict activity [2–5]. Although generally effective [6], the drawback of CNNs is that they are not rotationally invariant and require more parameters than alternative representations. As a result, graph convolutional networks [7], or more generally, message-passing neural networks [8–10], have gained popularity. Recent studies suggest that the performance of structure-based machine learning methods is driven in part by protein chemical features [11, 12, 5]. Rather than responding to specific interactions between ligands and binding sites, models learn characteristic properties of general ligand-protein interactions. This deficiency manifests itself by a decline in predictive performance when the model is confronted with a previously unseen binding site on the same protein, especially when that site partially overlaps with a canonical site. For example, a model may highly rank an ATP-competitive binder of an allosteric site on a kinase. Such limitations severely hinder the discovery of new chemical entities or the ability to target novel sites on proteins.
[0245] Simultaneous training on ligand pose quality and affinity can improve pose sensitivity
[13] . Based on that observation, here we build and present a multi-task architecture for bioactivity prediction that simultaneously evaluates bioactivity and the physics-based Vina
[14] score of the pose. We further condition the bioactivity task on pose quality. Finally, we expose the model to bad poses while disabling the bioactivity labels of bioactive molecules, and thus present bad poses of true binders as negative examples. We demonstrate that our architecture improves pose sensitivity on several rigorous benchmarks.
[0246] 2.1 Neural Network Architecture The system in this example is a graph neural network based architecture with position dependent edges. This is an example of a convolutional neural network 24 of the present disclosure. In this example, we only consider receptor atoms within 7 Å of any ligand atom. We use two graph convolutional layers where (ligand and receptor) atoms are neighbors if they are within 4 Å of each other. We then extract ligand-only features and add two more ligand-only layers. The ligand-only layers are pooled using a sum pooling layer. The pooled features are then used as embeddings for the multitask multi-layer perceptrons (first model 72, second model 74, etc.) at the top of the network. The embeddings generated by the graph neural network are used to predict three outputs in this example: activity, PoseRanker pose quality score, and Vina docking score. This is done in two stages. First, the PoseRanker and Vina score predictions are calculated by passing the embeddings through two independent multi-layer perceptrons. A conditioned embedding is then formed by concatenating the input embeddings with the PoseRanker score predictions and passed to a third multi-layer perceptron to compute activation predictions
[15] . Section 4.3 provided details of the model training parameters.
[0247] 2.2 Data The training dataset consisted of binding affinity measurements collected from publicly available sources such as Chembl or Pubchem, and from commercial databases such as Reaxys or Liceptor. In this example, only quantitative measurements at pKi2(0;11) were considered. If the measured pKi (or IC50) was less than 10 μM, the compound was labeled as active, otherwise the compound was labeled as inactive. Since the number of measured active compounds was greater than the inactive ones, the training dataset was augmented by randomly assigning each of the active compounds as a decoy for another different protein target. In addition, some models used pose negatives (active compounds that posed poorly and were labeled as inactive). For more information, see section 4.2. A set of 12 diverse proteins (D12) was excluded from training, which served as a holdout test set. In addition, all close homologs of the proteins in the D12 set were excluded from the training set (as well as sequences less than 95%). The training set covers over 3800 diverse proteins and considers 4.8M (5.8M) data points without pose negativity in some cases. The holdout set considers about 33000 compounds distributed across 12 proteins. All compounds were docked with the disclosed architecture, CUina
[16] , and the best available poses (as ranked by the PoseRanker model
[10] ) were used for scoring with the DL model.
[0248] 2.3 Numerical experiments To study the pose sensitivity of our model, each of the active target-compound pairs in D12 was scored three times with: i) the top pose, ii) the bad pose, and iii) a physically unrealistic pose obtained from the top pose by random rotation (repeated four times) of the ligand around its center of mass. The good pose was the highest ranked pose by PoseRanker (poses were generated with CUina). The bad pose was the worst ranked pose by PoseRanker. The unrealistic pose was obtained from the good pose by random rotation of the ligand around its center of mass. All poses were scored, then the scores of the bad and unrealistic poses were subtracted from the score of the good pose. The measure of pose sensitivity is the median decrease in activity score between the good pose and the bad / unrealistic pose. Convolutional neural networks can detect features in the perceptual field of the input data. If that field is large and complex enough, the model can detect a collection of atoms that are characteristic of conserved binding sites, e.g., the ATP binding site in a protein kinase. However, limiting the range of the perception field, e.g., by pooling, omits spatial information between detected features. As a result, the model may be biased by detecting chemically irrelevant features provided in the input data, the so-called Picasso problem. To monitor how neighboring binding sites interfere with the inference of the model, we selected a diverse set of known kinase inhibitors (about 300 diverse compounds labeled as active) and compared them with 10 randomly selected compounds from an available screening library (MCULE, as of 2017 / 18 / 10, labeled as inactive). 5The compounds were mixed with 1000 compounds. Each compound was docked into the ATP binding site and into an allosteric site 6-10 Å away from the ATP site. To monitor potential bias of the model, all compounds were also docked into a tentative binding site located on a remote SH2 domain (>50 Å away from the ATP binding site) (Figure 20). The expectation was that a model with good performance would be able to adequately discriminate the kinase inhibitor from background random molecules when docked into the ATP binding site (expect ROC AUC much higher than 0.5). On the other hand, the pause-sensitive model should not be biased by the neighboring ATP site when the compound is docked into the allosteric site (ROC AUC close to 0.5) (Figure 20). To account for any possible bias in the training set, the ROC AUC of the spatially distant binding site located in the SH2 domain was calculated and shown as blue dots in Figure 20.
[0249] 3 Results The results in Figure 21 show that the models studied in this example perform well on the holdout set, with GCN slightly outperforming CNN. However, when both single-task models are used in a virtual screening of allosteric sites in the human ZAP70 protein, both of them will improve on the known ATP site kinase inhibitors. This is because the models do not learn the features of the ligand-receptor interaction, but instead learn independent representations of the ligand and receptor. These learned representations / embeddings are then used for the inference of the models. Because the ATP binding site is in the perceptual field of these two networks, the GCN and CNN models can identify the highly conserved features of the ATP binding site (Figure 20, Figure 22), and make predictions as if the models were asked about the ATP site instead of the less common allosteric site (Figure 22). This result cannot be explained by a biased training set, since screening of binding sites that are spatially distant from the primary site (ATP site) did not result in improvements of kinase inhibitors (SH2 site, Figure 20, Figure 22). This suggests that these two models, CNN and GCN, are protein-chemistry oriented (ligand and receptor representations are used, but one is independent of the other). This is further supported by their insensitivity not only to misalignment (bad pose) of the ligand in the binding site (left panel of Fig. 23), but also to disruption of the ligand-receptor interface (right panel of Fig. 23). This strange behavior has also been observed in previous work on 3D grid-based CNNs [4,13], but no generally applicable solution has been proposed. The main drawback of the PCM model is its innate insensitivity to the pose used for inference. Thus, the solution to the Picasso problem is to ensure that the model is pose sensitive.In this example, the minimum requirements for a model to be considered pose-sensitive are that i) physically unrealistic poses (e.g., overlapping atoms) are penalized compared to poses without physically unrealistic features, ii) poses with ligands outside the binding pocket should be penalized more than poses with ligands in the binding site, and iii) binding sites in the vicinity of the target site should not interfere with the prediction.
[0250] At first, it is not intuitive that a single-task (active) model trained on structural data would not use its structural information on the ligand-receptor interaction. However, this may be true since during training, the main objective is to minimize a specified loss function, and the assumption is that the use of ligand-receptor interactions can give the model an edge in this task. In practice, in silico generated poses are subject to errors and uncertainties, and relying too heavily on them can impair the model's performance. Since the model has no incentive to learn the structural features of the ligand-receptor interaction, the model often ignores them. Thus, training a multitask model, where an additional task requires embedding that is structure-sensitive, should, in theory, alleviate the problem. This is indeed the case, as can be seen for the MT model in Figures 20 and 22. We see that adding docking score regression as another task (model MT-1) already leads to a model that penalizes clearly inappropriate poses (unrealistic poses), reducing the amount of improved kinase inhibitors with top hits for the allosteric site of the hZAP70 protein (Figure 22). Since the bad poses are misaligned but have no atomic collisions, the MT-1 model still cannot distinguish between the good and bad poses used in screening (left panel of Fig. 23). Interestingly, this problem cannot be solved by adding pose quality regression as a third task or by conditioning the activation task on pose quality alone, models MT-2 and MT-3 (Fig. 23). This is because the models are only shown good poses and cannot learn images of what bad poses look like.
[0251] To compensate for the missing information, a data augmentation technique called pose negatives is used. Pose negatives are examples that were originally labeled as positive data points and used with the best available pose. However, the worst available poses (according to any metric, in this case the PoseRanker score) can be selected and presented as negative examples to the model whose labeling has been changed. With this approach, it was observed that the models (MT-4a and MT-4b) were able to penalize both physically unrealistic and bad poses (Figure 23). Moreover, the same models also mitigated the Picasso problem. However, in this case, it was observed that the lack of conditioning of activations on pose quality led to models that were more prone to the Picasso problem (Figure 22).
[0252] 4 Conclusion The multitask architecture results in a model that can predict the biological activity of compounds and also fully exploit the provided structural data for inference. We force the model to learn orthogonal tasks and regularize the final model. The proposed solution is generally applicable to both 3D grid-based and graph-based models (data not shown). This approach opens up the field of deep learning and structure-based drug discovery to novel binding sites and previously undruggable proteins.
[0253] This work was developed in the context of efforts to reduce the costs and development times associated with early-stage drug discovery. Success in this area may, in the long term, improve access to medicines and reduce healthcare costs. It should be recognized that the training dataset described here consists of publicly available data and therefore necessarily reflects biases in the allocation of research funding to various diseases and health conditions. We hope that efforts to improve the generalizability of models across protein binding sites will help alleviate such limitations in the training data.
[0254] 4.1 Conditional Multitasking Architecture In practice, there were no restrictions on the loss functions that could be used for regression tasks (MSE, MAE, Huber, Log-Cosh, etc.) and classification tasks (BCE, hinge loss, squared hinge, locality loss, etc.). The auxiliary tasks were i) shared embedding x em Transform it into the task output si,
number
number
[0255] 4.2 Data Augmentation with Pause Negativity Using CUina docking, 64 poses were generated for each ligand-target pair. The poses were then sorted according to their quality using PoseRanker
[10] and the top 16 poses were selected. The highest ranked poses were used in training and scoring as good poses, whereas the last (16th) pose was used as a pose negative and considered to be inactive (non-binder).
[0256] 4.3 Training Each model was trained for 10 epochs. For every neural network architecture, six models were trained, each using the 5th / 6th data as a training set, with the 1st / 6th removed for crossfold validation. Each data crossfold contains clusters of proteins that share more than 70% sequence similarity. Models were trained using the ADAM optimizer with a learning rate of lr=0:001, and targets were sampled with replacement in proportion to the number of active compounds associated with that target (targets with no measured active compounds were removed from the training set).
[0257] References [1] Irwin and Shoichet,2016,“Docking Screens for Novel Ligands Conferring New Biology:Miniperspective.Journal of Medicinal Chemistry,” 59(9):4103-4120,May 2016.ISSN 167 0022-2623,1520-4804.doi:0.1021 / acs.jmedchem.5b02008.URL https: / / pubs.acs.org / doi / 10.1021 / acs.jmedchem.5b02008.
[0258] [2] Wallach, Dzamba, and Heifets, 2015, “AtomNet:A Deep Convolutional Neural Network for Bioactivity Prediction in Structure-based Drug Discovery,” arXiv:1510.02855 [cs,q-bio,stat].
[0259] [3] Ragoza et al,2017,“Protein-Ligand Scoring with Convolutional Neural Networks,Journal of Chemical Information and Modeling 57(4),pp.942-957.
[0260] [4] Stepniewska-Dziubinska et al.,2018,“Development and evaluation of a deep learning model for protein-ligand binding affinity prediction,” Bioinformatics,34(21),pp.3666-3674.
[0261] [5] Boyles et al.,2019,“Learning from the ligand:using ligand-based features to improve binding affinity prediction,” Bioinformatics,page btz665.
[0262] [6] Hsieh et al.,2019,“Miro1 Marks Parkinson’s Disease Subset and Miro1 Reducer Rescues Neuron Loss in Parkinson’s Models.Cell Metabolism,” 30(6),pp.1131-1140.
[0263] [7] Kipf and Welling,2017,“Semi-Supervised Classification with Graph Convolutional Networks,” arXiv:1609.02907 [cs,stat],February 2017.URL http: / / arxiv.org / abs / 1609.02907.arXiv:1609.02907.
[0264] [8] Feinberg et al.,2018,“PotentialNet for Molecular Property Prediction,” ACS Central Science 4(11),pp.1520-1530.
[0265] [9] Lim et al.,2019,“Predicting Drug-Target Interaction Using a Novel Graph Neural Network with 3D Structure-Embedded Graph Representation,” Journal of Chemical Information and Modeling,59(9),pp.3981-3988.
[0266]
[10] Stafford et al.,2021,“Enriching Ligand Pose Quality for Dynamic Proteins in Virtual High Throughput Screens,” doi:10.33774 / chemrxiv-2021-t6xkj.URL https: / / chemrxiv.org / engage / chemrxiv / article-details / 614b905e39ef6a1c36268003.
[0267]
[11] Siege et al.,2019,“In Need of Bias Control:Evaluating Chemical Data for Machine Learning in Structure-Based Virtual Screening,” Journal of Chemical Information and Modeling 59(3),pp.947-961.
[0268]
[12] Chen et al.,2019,“Hidden bias in the DUD-E dataset leads to misleading performance of deep learning in structure-based virtual screening,” PLOS ONE 14(8):e0220113.
[0269]
[13] Francoeur et al.,2020,“Three-Dimensional Convolutional Neural Networks and a Cross-Docked Data Set for Structure-Based Drug Design,” Journal of Chemical Information and Modeling 60(9),pp.4200-4215.
[0270]
[14] Trott and Olson,2010,“AutoDock Vina:Improving the speed and accuracy of docking with a new scoring function,efficient optimization,and multithreading,” Journal of Computational Chemistry 31(2) pp.455-461.
[0271]
[15] Long et al.,2018,“Conditional Adversarial Domain Adaptation,” arXiv:1705.10667 [cs],December 2018.
[0272]
[16] Morrison et al.,2020,“CUina:An Efficient GPU Implementation of AutoDock Vina,” August 2020.URL https: / / blog.atomwise.com / efficient-gpu-implementation-of-autodock-vina.
[0273] conclusion For purposes of explanation, the foregoing description has been set forth with reference to specific implementations. However, the above illustrative discussion is not intended to be exhaustive or to limit the implementation to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. The implementations have been chosen and described to best explain the principles and their practical applications, thereby enabling those skilled in the art to best utilize the implementations suited to the particular use contemplated and various implementations with various modifications.
Claims
1. A computer system (100) for characterizing an interaction between a test compound and a target polymer (38), said computer system comprising: one or more processors (49); a memory (92 / 90) addressable by said one or more processors, said memory storing at least one program (36) for execution by said one or more processors, said at least one program comprising: (A) acquiring a plurality of atomic coordinates (40) of the target polymer, wherein the plurality of atomic coordinates includes atomic coordinates of at least 400 atoms; (B) obtaining a training dataset (44) comprising a respective electronic description of each training compound (46) in a plurality of training compounds, the plurality of training compounds comprising at least 100 compounds, each respective electronic description comprising: (i) a corresponding positive pose (48) of the corresponding training compound with respect to a plurality of atomic spatial coordinates combined with a corresponding first positive interaction score (50) representing a first binding coefficient or in silico pose quality score of the corresponding training compound to the target polymer; (ii) obtaining a training dataset including corresponding negative poses (60) of the corresponding training compounds with respect to the plurality of atomic spatial coordinates combined with corresponding first negative interaction scores (62) representing second binding coefficients or in silico pose quality scores of the corresponding training compounds to the target polymer; (C) training at least a first model (72), said first model having a first plurality of parameters (73), said first plurality of parameters including more than 400 parameters, said training comprising, for each corresponding training compound in said plurality of training compounds: (i) comparing the corresponding first positive interaction score of the corresponding training compound with respect to the target polymer with an output from the first model upon inputting into the first model a corresponding positive score of the corresponding positive pose of the corresponding training compound with respect to the target polymer, wherein the corresponding positive score is obtained by inputting the corresponding positive pose into a trained neural network (24); (ii) comparing the corresponding first negative interaction score of the corresponding training compound with respect to the target polymer with an output from a first model upon inputting into the first model a corresponding negative score of the corresponding negative pose of the corresponding training compound with respect to the target polymer, wherein the corresponding negative score is obtained by inputting the corresponding negative pose into the trained neural network; and training, whereby the first plurality of parameters are adjusted and an output of at least the first model provides a binding coefficient or an in silico quality score for the interaction between the test compound and the target polymer.
2. 2. The computer system of claim 1, wherein the first model is a first fully connected neural network.
3. the training is a regression task in which the first plurality of parameters are adjusted by backpropagation through an associated loss function; The corresponding first positive interaction score is related to the corresponding first negative interaction score by the formula: B = N × A, During the ceremony, A is the corresponding positive interaction score, B is the corresponding negative interaction score, 2. The computer system of claim 1, wherein N is a real number greater than zero and less than one.
4. 4. The computer system of claim 3, wherein the associated loss function is a mean squared error loss function, a mean absolute error loss function, a Huber loss function, a Log-Cosh loss function, or a quantile loss function.
5. the corresponding first positive interaction score and the corresponding first negative interaction score each represent a binding coefficient; The computer system of claim 3 , wherein the corresponding first positive interaction score is an in vitro measurement of the binding coefficient of the corresponding training compound to the target polymer.
6. The first positive interaction score is the IC of each of the training compounds with respect to the target polymer. 50 , E.C. 50 , Kd, KI, or pKI.
7. each respective electronic description in the training dataset further comprises a corresponding positive activity score (56) for the corresponding positive pose of the corresponding training compound and a corresponding negative activity score (68) for the corresponding negative pose of the corresponding training compound; The training (C) of at least the first model further comprises jointly training a second model (74) with the first model, the second model having a second plurality of parameters (75), and the training includes, for each corresponding training compound in the plurality of training compounds: (iii) comparing the corresponding positive activity scores of the corresponding training compounds with the output from the second model when the corresponding positive scores of the corresponding positive poses of the corresponding training compounds with respect to the target polymer are input into the second model; 10. The computer system of claim 1, further comprising: (iv) comparing the corresponding negative activity score of the corresponding training compound with the output from the second model when the corresponding negative score of the corresponding negative pose of the corresponding training compound with respect to the target polymer is input into the second model, thereby adjusting the second plurality of parameters such that the second model provides an activity of the interaction between the test compound and the target polymer.
8. each respective electronic description in the training dataset further comprises a corresponding positive activity score for the corresponding positive pose of the corresponding training compound and a corresponding negative activity score for the corresponding negative pose of the corresponding training compound; The training of at least the first model (C) further comprises jointly training a second model with the first model, the second model having a second plurality of parameters, and the training includes, for each corresponding training compound in the plurality of training compounds: (iii) comparing the corresponding positive activity score of the corresponding training compound with the output from the second model when the corresponding positive score of the corresponding positive pose of the corresponding training compound with respect to the target polymer and the corresponding first positive interaction score are jointly input into the second model; (iv) comparing the output from the second model when the corresponding negative activity score of the corresponding training compound with the corresponding negative score of the corresponding negative pose of the corresponding training compound with respect to the target polymer and the corresponding first negative interaction score are jointly input into the second model, thereby adjusting the second plurality of parameters, and the second model provides an activity score of the test compound with respect to the target polymer.
9. the training of the first model is a regression task in which the first plurality of parameters are adjusted by backpropagation through a first associated loss function; 10. The computer system of claim 8, wherein the training of the second model is a classification task in which the second plurality of parameters are tuned by backpropagation through a second associated loss function.
10. each respective electronic description in the training dataset further comprises a corresponding second positive interaction score (58) for the corresponding positive pose of the corresponding training compound and a corresponding second negative interaction score (70) for the corresponding negative pose of the corresponding training compound; each respective electronic description in the training dataset further comprises a corresponding positive activity score for the corresponding positive pose of the corresponding training compound and a corresponding negative activity score for the corresponding negative pose of the corresponding training compound; The training (C) of at least the first model further comprises jointly training a second model (74) and a third model (76) with the first model, the second model having a second plurality of parameters (75), the third model having a third plurality of parameters (77), and the training includes, for each corresponding training compound in the plurality of training compounds: (iii) comparing the corresponding second positive interaction scores of the corresponding training compounds with respect to the target polymer with the output from the second model when the corresponding positive scores of the corresponding positive poses of the corresponding training compounds with respect to the target polymer are input into the second model; (iv) comparing the corresponding second negative interaction scores of the corresponding training compounds with respect to the target polymer with the output from the second model when the corresponding negative scores of the corresponding negative poses of the corresponding training compounds with respect to the target polymer are input into the second model, thereby adjusting the second plurality of parameters; (v) comparing the corresponding positive activity score of the corresponding training compound with the output from the third model when jointly inputting into the third model: (a) the corresponding positive score of the corresponding positive pose of the corresponding training compound with respect to the target polymer, (b) the output of the first model when inputting the corresponding positive score of the corresponding positive pose of the corresponding training compound, and (c) the output of the second model when inputting the corresponding positive score of the corresponding positive pose of the corresponding training compound; (vi) comparing the output from the first model when the corresponding negative activity score of the corresponding training compound is jointly input to the third model with (a) the corresponding negative score of the corresponding negative pose of the corresponding training compound with respect to the target polymer, (b) the output of the first model when the corresponding negative score of the corresponding negative pose of the corresponding training compound is input, and (c) the output of the second model when the corresponding negative score of the corresponding negative pose of the corresponding training compound is input, thereby adjusting the third plurality of parameters of the third model, and an output of the third model provides the activity of the test compound with respect to the target polymer.
11. the training of the first model is a first regression task in which the first plurality of parameters are adjusted by backpropagation through a first associated loss function; the training of the second model is a second regression task in which the second plurality of parameters are adjusted by backpropagation through a second associated loss function; 11. The computer system of claim 10, wherein the training of the third model is a classification task in which the third plurality of parameters are tuned by backpropagation through a third associated loss function.
12. the corresponding first positive interaction score and the corresponding first negative interaction score each represent an in silico quality score of the corresponding training compound with the target polymer; the corresponding second positive interaction score and the corresponding second negative interaction score each represent a binding coefficient of the corresponding training compound to the target polymer; 12. The computer system of claim 11, wherein the corresponding positive activity score is a first binary activity score and the corresponding negative activity score is a second binary activity score.
13. the first associated loss function is a mean squared error loss function, a mean absolute error loss function, a Huber loss function, a Log-Cosh loss function, or a quantile loss function; the second associated loss function is a mean squared error loss function, a mean absolute error loss function, a Huber loss function, a Log-Cosh loss function, or a quantile loss function; The computer system of claim 12 , wherein the third associated loss function is a binary cross-entropy loss function, a hinge loss function, or a squared hinge loss function.
14. 1. A method for characterizing an interaction between a test compound and a target polymer (38), said method comprising: In a computer system having a memory, (A) acquiring a plurality of atomic coordinates (40) of the target polymer, wherein the plurality of atomic coordinates includes atomic coordinates of at least 400 atoms; (B) obtaining a training dataset (44) comprising a respective electronic description of each training compound in a plurality of training compounds (46), said plurality of training compounds comprising at least 100 compounds, each respective electronic description comprising: (i) a corresponding positive pose of the corresponding training compound with respect to a plurality of atomic spatial coordinates (48) combined with a corresponding first positive interaction score (50) representing a first binding coefficient or in silico pose quality score of the corresponding training compound to the target polymer; (ii) obtaining a training dataset including corresponding negative poses (60) of the corresponding training compounds with respect to the plurality of atomic spatial coordinates combined with corresponding first negative interaction scores (62) representing second binding coefficients or in silico pose quality scores of the corresponding training compounds to the target polymer; (C) training at least a first model (72), said first model having a first plurality of parameters (73), said first plurality of parameters including more than 400 parameters, said training comprising, for each corresponding training compound in said plurality of training compounds: (i) comparing the corresponding first positive interaction score of the corresponding training compound with respect to the target polymer with an output from the first model upon inputting into the first model a corresponding positive score of the corresponding positive pose of the corresponding training compound with respect to the target polymer, wherein the corresponding positive score is obtained by inputting the corresponding positive pose into a trained neural network (24); (ii) comparing the corresponding first negative interaction score of the corresponding training compound with respect to the target polymer with an output from a second model upon inputting into the first model a corresponding negative score of the corresponding negative pose of the corresponding training compound with respect to the target polymer, wherein the corresponding negative score is obtained by inputting the corresponding negative pose into the trained neural network, thereby adjusting the first plurality of parameters and training wherein an output of at least the first model provides a binding coefficient or an in silico pose quality score for the interaction between the test compound and the target polymer.
15. 1. A non-transitory computer-readable storage medium storing instructions that, when executed by a computer system, cause the computer system to perform a method for characterizing an interaction between a test compound and a target polymer, the method comprising: (A) obtaining a plurality of atomic coordinates of the target polymer, wherein the plurality of atomic coordinates includes atomic coordinates of at least 400 atoms (40); (B) obtaining a training dataset (44) comprising a respective electronic description of each training compound (46) in a plurality of training compounds, the plurality of training compounds comprising at least 100 compounds, each respective electronic description comprising: (i) a corresponding positive pose (48) of the corresponding training compound with respect to a plurality of atomic spatial coordinates combined with a corresponding first positive interaction score (50) representing a first binding coefficient or in silico pose quality score of the corresponding training compound to the target polymer; (ii) obtaining a training dataset including corresponding negative poses (60) of the corresponding training compounds with respect to the plurality of atomic spatial coordinates combined with corresponding first negative interaction scores (62) representing second binding coefficients or in silico pose quality scores of the corresponding training compounds to the target polymer; (C) training at least a first model (72), said first model having a first plurality of parameters (73), said first plurality of parameters including more than 400 parameters, said training comprising, for each corresponding training compound in said plurality of training compounds: (i) comparing the corresponding first positive interaction score of the corresponding training compound with respect to the target polymer with an output from the first model upon inputting into the first model a corresponding positive score of the corresponding positive pose of the corresponding training compound with respect to the target polymer, wherein the corresponding positive score is obtained by inputting the corresponding positive pose into a trained neural network (24); (ii) comparing the corresponding first negative interaction score of the corresponding training compound with respect to the target polymer with an output from the first model upon inputting the corresponding negative score of the corresponding negative pose of the corresponding training compound with respect to the target polymer into the first model, wherein the corresponding negative score is obtained by inputting the corresponding negative pose into a trained neural network (24), thereby adjusting the first plurality of parameters and training at least the output of the first model providing a binding coefficient or an in silico pose quality score for the interaction between the test compound and the target polymer.