Characterizing interactions between compounds and polymers using pose ensembles

By employing a conditional multitasking architecture with attention mechanisms in vHTS machine learning models, the insensitivity to poses in traditional methods is addressed, resulting in enhanced bioactivity prediction and interaction characterization.

JP2025515347APending Publication Date: 2025-05-14ATOMWISE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024563511
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-04-29
Filing Date
2023-03-17
Publication Date
2025-05-14

AI Technical Summary

Technical Problem

Traditional structure-based virtual high-throughput screening (vHTS) machine learning methods are insensitive to different poses of compounds bound to target polymers, leading to inaccurate characterization of interactions.

Method used

A conditional multitasking architecture for vHTS machine learning models that predicts biological activity from multiple poses, incorporating attention mechanisms to exploit hidden correlations in pose distributions.

Benefits of technology

Improves bioactivity prediction by enforcing pose sensitivity, leading to more accurate characterization of interactions between test compounds and target polymers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025515347000001_ABST
    Figure 2025515347000001_ABST
Patent Text Reader

Abstract

A system and method for characterizing an interaction between a compound and a polymer includes obtaining a plurality of atomic coordinate sets, each of which includes a compound bound to the polymer at a corresponding pose in a plurality of poses. Each respective atomic coordinate set, or an encoding thereof, is input sequentially into a neural network to obtain a corresponding initial embedding as an output, thereby obtaining a plurality of initial embeddings, each initial embedding corresponding to an atomic coordinate set in the plurality of atomic coordinate sets. An attention mechanism is applied to the plurality of initial embeddings in a concatenated fashion to obtain an attention embedding. A pooling function is applied to the attention embeddings to derive a pooled embedding. The pooled embeddings are input into a model to obtain an interaction score for the interaction between the compound and the polymer.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 336,841, entitled “CHARACTERIZATION OF INTERACTIONS BETWEEN COMPOUNDS AND POLYMERS USING POSE ENSEMBLES,” filed April 29, 2022, which is incorporated herein by reference.

[0002] The present application is directed to the use of models to characterize the interactions between test compounds and target polymers. [Background technology]

[0003] Essentially, biological systems function through the physical interaction of molecules such as compounds with target polymers. Structure-based virtual high-throughput screening (vHTS) methods have been used to characterize the interaction between candidate (test) compounds and target polymers through machine learning approaches. Such characterization can report, for example, continuous or categorical activity indicators, PKa, or any other suitable metric to characterize the interaction between the candidate compound and the target polymer.

[0004] One shortcoming of vHTS machine learning methods is that they do not fully consider the enthalpy and entropy components of acceptor-ligand complex formation. Structure-based deep learning methods typically predict bioactivity from static docked ligand poses. However, these approaches ignore the entropy contribution to the free energy change. Predicting bioactivity from an ensemble of docked poses can overcome this limitation, but requires the model to recognize and be sensitive to different poses. However, traditional machine learning methods tend to be insensitive to different poses. Figure 19 illustrates this insensitivity, where a machine learning model such as a convolutional neural network erroneously favors a pose that has all the correct components but is fundamentally wrong overall. Figure 18 illustrates a situation where the left pose and the right pose have the same parts, two eyes, two eyebrows, nose, lips, and overall shape of the head. Thus, it can be seen that it is difficult to teach a machine learning model that the left pose is the correct pose. Thus, there is an inherent pose insensitivity in traditional vHTS machine learning methods. This pause insensitivity can lead to erroneous or inaccurate characterization of the interaction between the test compound and the target polymer. For example, this pause insensitivity can lead a vHTS machine learning approach, which provides a categorical activity label for each compound in a screening library, to mislabel a certain percentage of the compounds in the screening library.

[0005] Given the above background, what is needed in the art is a method for imposing pause sensitivity on vHTS machine learning methods such that they are sensitive to bad pauses. Summary of the Invention

[0006] The present disclosure addresses the problems identified in the Background Art by utilizing a vHTS machine learning model that simultaneously predicts bioactivity from multiple poses. The conditional multi-task architecture of the disclosed vHTS machine learning model enforces sensitivity to different ligand poses, and the architecture includes an attention mechanism to exploit hidden correlations in the pose distribution. The disclosed vHTS machine learning model improves bioactivity predictions compared to baseline models that predict only from static ligand poses.

[0007] Thus, one aspect of the present disclosure provides a computer system for characterizing the interaction between a test compound and a target polymer. The computer system comprises one or more processors and a memory addressable by the one or more processors. The memory stores at least one program for execution by the one or more processors. The at least one program includes instructions for obtaining a plurality of atomic coordinate sets. Each respective atomic coordinate set in the plurality of atomic coordinate sets includes a test compound bound to the target polymer at a corresponding pose in the plurality of poses. In some embodiments, each respective atomic coordinate set in the plurality of atomic coordinate sets includes atomic coordinates for at least 30 atoms.

[0008] For each respective atomic coordinate set in the plurality of atomic coordinate sets, the at least one program includes instructions for inputting the respective atomic coordinate set or an encoding of the respective atomic coordinate set into a first neural network to obtain a corresponding initial embedding as an output of the first neural network, thereby obtaining a plurality of initial embeddings. Each initial embedding in the plurality of initial embeddings corresponds to an atomic coordinate set in the plurality of atomic coordinate sets. The first neural network includes more than 400 parameters.

[0009] The at least one program further includes instructions for concatenatively applying an attention mechanism to the multiple initial embeddings, thereby obtaining an attention embedding.

[0010] The at least one program further includes instructions for applying a pooling function to the attention embeddings to derive a pooled embedding.

[0011] The at least one program further comprises instructions for inputting the pooled embeddings into a first model, thereby obtaining a first interaction score for the interaction between the test compound and the target polymer. The first model comprises more than 400 parameters.

[0012] In some embodiments, the first interaction score represents a binding coefficient of the test compound to the target polymer. In some such embodiments, the binding coefficient is the IC 50 , E.C. 50、 Kd, KI, or pKI.

[0013] In some embodiments, the first interaction score represents an in silico quality score of the test compound relative to the target polymer.

[0014] In some embodiments, the first model is a fully connected second neural network.

[0015] In some embodiments, the at least one program further comprises instructions for inputting the pooled embeddings into a second model, thereby obtaining a second interaction score for the interaction between the test compound and the target polymer. In such embodiments, the first model is a first fully-connected neural network, the second model is a second fully-connected neural network, the first interaction score represents an in silico quality score of the test compound for the target polymer, and the second interaction score represents an in silico quality score of the test compound for the target polymer. In some such embodiments, the at least one program further comprises instructions for inputting the first interaction score and the second interaction score into a third model to obtain a third interaction score, the third model is a third fully-connected neural network. In some such embodiments, the third interaction score is a discrete binary activity score having a first value if the test compound is determined to be inactive by the third model and a second value if the test compound is determined to be active by the third model.

[0016] In some embodiments, the target polymer is an assembly of proteins, polypeptides, polynucleic acids, polyribonucleic acids, polysaccharides, or any combination thereof.

[0017] In some embodiments, each set of atomic coordinates in the plurality of atomic coordinate sets comprises three-dimensional coordinates {x, ..., x} for at least a portion of a polymer from a crystal structure of the target polymer resolved at 2.5 Å resolution or better or 3.3 Å resolution or better. N}.

[0018] In some embodiments, each atomic coordinate set in the plurality of atomic coordinate sets comprises a collection of three-dimensional coordinates for at least a portion of the target polymer as determined by nuclear magnetic resonance, neutron diffraction, or cryo-electron microscopy.

[0019] In some embodiments, the first interaction score is a binary score, and a first value of the binary score represents an IC of the test compound for the target polymer that is above a first threshold. 50 , E.C. 50、 The second value of the binary score represents the IC of the test compound against the target polymer below the first threshold, representing Kd, KI, or pKI. 50 , E.C. 50 , Kd, ​​KI, or pKI.

[0020] In some embodiments, the compound satisfies two or more, three or more, or all four of Lipinski's rule of five: (i) five or fewer hydrogen bond donors, (ii) ten or fewer hydrogen bond acceptors, (iii) molecular weight less than 500 daltons, and (iv) LogP less than 5.

[0021] In some embodiments, the test compounds are organic compounds having a molecular weight of less than 500 daltons, less than 1000 daltons, less than 2000 daltons, less than 4000 daltons, less than 6000 daltons, less than 8000 daltons, less than 10000 daltons, or less than 20000 daltons.

[0022] In some embodiments, the test compounds are organic compounds having a molecular weight between 400 and 10,000 daltons.

[0023] In some embodiments, the plurality of atomic coordinate sets consists of 3 to 64 poses. In some embodiments, the plurality of atomic coordinate sets consists of 2 to 64 poses.

[0024] In some embodiments, the first neural network is a convolutional neural network. In some such embodiments, each respective atomic coordinate set in the plurality of atomic coordinate sets comprises atomic coordinates for at least 80 atoms.

[0025] In some embodiments, the first neural network is a graph neural network. In some such embodiments, the graph neural network is characterized by an initial embedding layer and a plurality of interaction layers each contributing to an interaction data structure in the plurality of interaction data structures for each atom in a respective set of atomic sets of atomic coordinates of corresponding poses in the plurality of poses pooled to form a corresponding initial embedding of the corresponding pose.

[0026] In some embodiments, the first neural network is an equivariant neural network or a message passing neural network.

[0027] In some embodiments, the first neural network includes multiple graph convolution blocks, each of which uses multiple radial graphs to account for connectivity in a respective set of atomic coordinates.

[0028] In some embodiments, the first neural network is 1×10 6 Contains parameters.

[0029] In some embodiments, the corresponding initial embedding comprises a data structure having between 128 and 768 values. In some embodiments, the corresponding initial embedding comprises a data structure having more than 100 values. In some embodiments, the corresponding initial embedding comprises a data structure having more than 80 values, more than 100 values, more than 120 values, more than 140 values, or more than 160 values. In some embodiments, the corresponding initial embedding comprises a data structure of between 100 values ​​and 2000 values.

[0030] In some embodiments, the plurality of initial embeddings includes a first plurality of values, and applying the attention mechanism includes (i) inputting the first plurality of values ​​into an attention neural network, thereby obtaining a first plurality of weights, each weight in the first plurality of weights corresponding to a respective value in the first plurality of values, and (ii) weighting each respective value in the first plurality of values ​​by a corresponding weight in the plurality of weights, thereby obtaining the attention embedding. In some such embodiments, a sum of the first plurality of weights is 1, and each weight in the first plurality of weights is a scalar value between 0 and 1.

[0031] In some embodiments, the pooling function collapses the attention embedding into a pooled embedding by applying a statistical function to combine portions of the attention embeddings that represent different poses in the multiple poses to form the pooled embedding. In some such embodiments, the attention embedding includes a corresponding plurality of values ​​for a corresponding plurality of elements for each respective pose in the multiple poses, and the statistical function is a maximum function that takes a maximum over the corresponding elements of each respective pose represented in the attention embedding to form the pooled embedding. In some embodiments, the attention embedding includes a corresponding plurality of values ​​for a corresponding plurality of elements for each respective pose in the multiple poses, and the statistical function is an averaging function that averages the corresponding elements of each respective pose represented in the attention embedding to form the pooled embedding.

[0032] In some embodiments, the first model is a regression task and the first interaction score quantifies the interaction between the test compound and the target polymer.

[0033] In some embodiments, the first model is a classification task and the first interaction score classifies an interaction between the test compound and the target polymer.

[0034] Another aspect of the disclosure provides a method for characterizing an interaction between a test compound and a target polymer. The method includes obtaining a plurality of atomic coordinate sets. Each respective atomic coordinate set in the plurality of atomic coordinate sets includes a test compound bound to the target polymer at a corresponding pose in the plurality of poses. In some embodiments, each respective atomic coordinate set in the plurality of atomic coordinate sets includes atomic coordinates for at least 30 atoms. Further, in the method, for each respective atomic coordinate set in the plurality of atomic coordinate sets, the respective atomic coordinate set, or an encoding of the respective atomic coordinate set, is input to a first neural network to obtain a corresponding initial embedding as an output of the first neural network. In this manner, a plurality of initial embeddings are obtained, and each initial embedding in the plurality of initial embeddings corresponds to an atomic coordinate set in the plurality of atomic coordinate sets. In some such embodiments, the first neural network comprises more than 400 parameters. The method further includes applying an attention mechanism to the plurality of initial embeddings in a concatenated manner, thereby obtaining an attention embedding. The method further includes applying a pooling function to the attention embedding to derive a pooled embedding. The method further comprises inputting the pooled embeddings into a first model, thereby obtaining a first interaction score for the interaction between the test compound and the target polymer. In some embodiments, the first model comprises more than 400 parameters.

[0035] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores instructions that, when executed by a computer system, cause the computer system to perform a method for characterizing an interaction between a test compound and a target polymer. The method includes obtaining a plurality of atomic coordinate sets, each of which includes a test compound bound to the target polymer at a corresponding pose in the plurality of poses. In some embodiments, each of which includes atomic coordinate sets for at least 30 atoms. The method further includes, for each of which, inputting the respective atomic coordinate set or an encoding of the respective atomic coordinate set into a first neural network to obtain a corresponding initial embedding as an output of the first neural network. In this manner, a plurality of initial embeddings are obtained, and each of which corresponds to an atomic coordinate set in the plurality of atomic coordinate sets. In some embodiments, the first neural network includes more than 400 parameters. The method further includes applying an attention mechanism to the plurality of initial embeddings in a concatenated manner, thereby obtaining an attention embedding. The method further comprises applying a pooling function to the attention embedding to derive a pooled embedding. The method further comprises inputting the pooled embedding into a first model, thereby obtaining a first interaction score for the interaction between the test compound and the target polymer. In some embodiments, the first model comprises more than 400 parameters.

[0036] In the drawings, embodiments of the disclosed system and method are shown by way of example, and it is to be expressly understood that the description and drawings are for illustrative purposes only and are intended as an aid to understanding and are not intended as a definition of the limits of the disclosed system and method. [Brief description of the drawings]

[0037] [Figure 1A]1 illustrates a computer system according to some embodiments of the present disclosure. [Figure 1B] 1 illustrates a computer system according to some embodiments of the present disclosure. [Figure 2A] 1 illustrates a method for characterizing an interaction between a test compound and a target polymer according to some embodiments of the present disclosure. [Figure 2B] 1 illustrates a method for characterizing an interaction between a test compound and a target polymer according to some embodiments of the present disclosure. [Figure 2C] 1 illustrates a method for characterizing an interaction between a test compound and a target polymer according to some embodiments of the present disclosure. [Figure 2D] 1 illustrates a method for characterizing an interaction between a test compound and a target polymer according to some embodiments of the present disclosure. [Figure 2E] 1 illustrates a method for characterizing an interaction between a test compound and a target polymer according to some embodiments of the present disclosure. [Figure 2F] 1 illustrates a method for characterizing an interaction between a test compound and a target polymer according to some embodiments of the present disclosure. [Diagram 3] FIG. 1 is a schematic diagram of an exemplary training compound in pose relative to a target polymer, according to some embodiments of the present disclosure. [Figure 4] 1 is a schematic diagram of a geometric representation of input features in the form of a three-dimensional grid of voxels, according to some embodiments of the present disclosure. [Diagram 5] FIG. 2 is a diagram of a compound encoded onto a two-dimensional grid of voxels, according to some embodiments of the present disclosure. [Figure 6] FIG. 2 is a diagram of a compound encoded onto a two-dimensional grid of voxels, according to some embodiments of the present disclosure. [Figure 7] FIG. 7 is a diagram of the visualization of FIG. 6 with voxels numbered, according to some embodiments of the present disclosure. [Figure 8] 1A-1C are schematic diagrams of geometric representations of input features in the form of coordinate positions of atom centers, according to some embodiments of the present disclosure; [Figure 9A] 1 illustrates a system for characterizing the interaction between a test compound and a target polymer, according to one embodiment of the present disclosure. [Figure 9B] 1 shows a system for characterizing an interaction between a test compound and a target polymer according to another embodiment of the present disclosure. [Figure 9C] 1 illustrates a system for characterizing an interaction between a test compound and a target polymer according to another embodiment of the present disclosure, where the first neural network is a graph-based neural network. [Figure 10] According to one embodiment of the present disclosure, there is provided a system for characterizing an interaction between a test compound and a target polymer, the characterization being (i) a binary discrete activity and (ii) a pKi, wherein the system is trained using poses to train the compound. [Figure 11] A system for characterizing an interaction between a test compound and a target polymer according to one embodiment of the present disclosure, wherein the characterization is pKi, the pKi is partially conditioned for activity, and the system is trained using a pose to train the compound. [Figure 12] According to one embodiment of the present disclosure, there is provided a system for characterizing an interaction between a test compound and a target polymer, the characterization being activity, the activity being conditioned, in part, on both pKi and pose quality score, and the system being trained using poses to train the compound. [Figure 13] According to one embodiment of the present disclosure, there is provided a system for characterizing an interaction between a test compound and a target polymer, the characterization being activity, the activity being conditioned, in part, on both pKi and binding mode score, and the system being trained using poses to train the compound. [Figure 14]According to one embodiment of the present disclosure, there is provided a system for characterizing the interaction between a test compound and a target polymer, the characterization being activity and two different compound binding mode scores, and the system is trained using poses to train the compound. [Figure 15] According to one embodiment of the present disclosure, a system for characterizing the interaction between a test compound and a target polymer, the characterization being activity, two different compound binding mode scores, and pKi, and the system is trained using poses to train the compound. [Figure 16A] According to one embodiment of the present disclosure, there is provided a system for characterizing an interaction between a test compound and a target polymer, wherein the characterization is activity, the activity being conditioned, in part, on pKi and binding mode scores, and the system is trained using poses to train the compound. [Figure 16B] According to one embodiment of the present disclosure, there is provided a system for characterizing an interaction between a test compound and a target polymer, the characterization being activity, the activity being conditioned, in part, on pKi and two different binding mode scores, and the system being trained using poses to train the compound. [Figure 17] 1 is a depiction of applying multiple function computation elements (g1, g2, ...) to voxel inputs (x1, x2, ..., x100) and using g() to jointly compose the function computation element outputs, according to some embodiments of the present disclosure. [Figure 18] 1 illustrates the insensitivity faced by machine learning models in characterizing the pose of compounds with respect to target polymers according to the prior art. [Figure 19] Demonstrating the insensitivity of traditional machine learning models to the quality of the compound-polymer pose, as shown in the figure, the best possible pose receives the same score as a bad pose by the machine learning model, and an unrealistic pose receives the same score as the best possible pose by the machine learning model. [Figure 20]1 illustrates an active task conditioned on PoseRanker and Vina scores according to one embodiment of the present disclosure. [Figure 21A] We provide performance statistics for the disclosed architectures (o3-2.8.0 and o4-2.8.0) against other architectures (n8b-long and n8b-maxlong). [Figure 21B] We provide performance statistics for the disclosed architectures (o3-2.8.0 and o4-2.8.0) against other architectures (n8b-long and n8b-maxlong).

[0038] Like reference numbers refer to corresponding parts throughout the several views. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0039] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0040] The present disclosure provides systems and methods for characterizing interactions between compounds and polymers. In the method, a plurality of atomic coordinate sets are obtained. Each of these atomic coordinate sets includes a compound bound to the polymer at a corresponding pose in the plurality of poses. Each respective atomic coordinate set, or an encoding thereof, is input sequentially into a neural network to obtain a corresponding initial embedding as an output. In this manner, a plurality of initial embeddings are calculated. Each initial embedding corresponds to an atomic coordinate set in the plurality of atomic coordinate sets. An attention mechanism is applied to the plurality of initial embeddings in a concatenated fashion to obtain an attention embedding. In some embodiments, the attention mechanism is a neural network trained on test data to emphasize some parts of the plurality of initial embeddings while de-emphasizing some parts of the plurality of initial embeddings. A pooling function is applied to the attention embedding to derive a pooled embedding. Thus, the pooling function collapses all the initial embeddings representing the plurality of poses into a single composite embedding representing all the poses in the plurality of poses. The pooled embedding is input into a model to obtain an interaction score for the interaction between the compound and the polymer.

[0041] It should also be understood that terms such as first, second, etc. may be used herein to describe various elements, but these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first object may be referred to as a second object, and similarly, a second object may be referred to as a first object, without departing from the scope of the present invention. A first object and a second object are both objects, but are not the same object.

[0042] The terms used in this disclosure are only for the purpose of describing particular embodiments and are not intended to limit the invention. When used in the description of the invention and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. The term "and / or" as used herein will also be understood to refer to and include any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms "comprises" and / or "comprising" as used herein specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0043] As used herein, the term "if" may be interpreted to mean "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "when determined" or "when [a described condition or event] is detected" may be interpreted to mean "at the time of determining" or "in response to determining" or "upon detection of [a described condition or event]," or "when [a described condition or event] is detected," depending on the context.

[0044] 1 shows a computer system 100 for characterizing interactions between test compounds and target polymers, which can be used, for example, as a binding affinity prediction system to generate accurate predictions regarding the binding affinity of one or more test compounds with a target polymer.

[0045] With reference to Figure 1, in a typical embodiment, computer system 100 comprises one or more computers. For purposes of illustration in Figure 1, computer system 100 is represented as a single computer that includes all of the disclosed functionality of computer system 100. However, the present disclosure is not so limited. The functionality of computer system 100 may be distributed across any number of networked computers and / or may reside on each of multiple networked computers and / or virtual machines. Those skilled in the art will appreciate that a variety of different computer topologies are possible for computer system 100, and all such topologies are within the scope of the present disclosure.

[0046] With the above in mind and referring to Figure 1, computer system 100 comprises one or more processing units (CPUs) 59, a network or other communication interface 84, a user interface 78 (e.g., including an optional display 82 and an optional keyboard 80 or other form of input device), memory 92 (e.g., random access memory, persistent memory, or a combination thereof), one or more magnetic disk storage and / or persistent devices 90 optionally accessed by one or more controllers 88, one or more communication buses 12 for interconnecting the aforementioned components, and a power supply 79 for providing power to the aforementioned components. To the extent that components of memory 92 are not persistent, data in memory 92 may be seamlessly shared with non-volatile memory 90 or portions of memory 92 that are non-volatile / persistent using known computing techniques such as caching. Memory 92 and / or memory 90 may include mass storage located remotely relative to central processing unit(s) 59. In other words, some data stored in memory 92 and / or memory 90 may actually be hosted on a computer that is external to computer system 100, but can be accessed electronically by computer system 100 via the Internet, an intranet, or other form of network or electronic cable using network interface 84. In some embodiments, computer system 100 utilizes models that execute from memory associated with one or more graphical processing units to improve system speed and performance. In some alternative embodiments, computer system 100 utilizes models that execute from memory 92 rather than memory associated with the graphical processing units.

[0047] The memory 92 of the computer system 100 includes: an optional operating system 34 that includes procedures for handling various basic system services; and A spatial data evaluation module 36 for characterizing interactions between test compounds (or training compounds) and the target polymer; target polymer data 38, including structural data of the target polymer (a plurality of atomic spatial coordinates 40 of the target polymer) and, optionally, active site information 42; a training dataset 44 comprising a plurality of electronic descriptions, each electronic description 46 in the plurality of electronic descriptions corresponding to one training compound in a plurality of training compounds (and / or test compounds), (i) a plurality of poses of the corresponding compounds, each respective pose 48 of the corresponding compounds comprising (a) a corresponding set of atomic spatial coordinates 49 detailing the atomic coordinates of the corresponding compound at the respective pose relative to the spatial coordinates 40 of the target polymer 38, (b) an optional corresponding voxel map 5 detailing the atomic interactions of the corresponding compound at the respective pose relative to the target polymer according to the corresponding set of atomic coordinates; 2, and (c) a plurality of poses of the corresponding compounds represented by a corresponding set of atomic coordinates 49 and / or optional corresponding vectors 54 encoding interactions between the corresponding compounds at each pose relative to the target polymer according to a corresponding voxel map 52, (ii) a (first) interaction score 50 between the corresponding compounds and the target polymer 38, (iii) an optional activity score 56 between the corresponding compounds and the target polymer 38, and (iv) an optional (second) interaction score 58 between the corresponding compounds and the target polymer 38; a first neural network 72 including a plurality of parameters, each respective output of the first neural network providing an initial embedding 74 corresponding to the set of atomic coordinates 49; an attention mechanism 77 that is collectively applied, in a concatenated fashion, to the initial embeddings 74 of each pose 48 of a corresponding compound (a particular training or test compound) to derive an attention embedding 79; a pooling function 81 having a number of parameters 83, the pooling function being applied to the attention embeddings 79 to derive a pooled embedding 85 having a number of embedding elements 87; ● A first model 89 having a plurality of parameters 91 that is applied to the pooled embedding 85 to (i) obtain a first interaction score of an interaction between a corresponding compound (a particular training or test compound) and the target polymer, and / or (ii) condition any other single model and / or group of models; an optional second model 93 including a plurality of parameters 95, the output of which is used to (i) provide a second interaction score for the interaction between the corresponding compound (particular training or test compound) and the target polymer, and / or (ii) condition any other single model and / or group of models; - optionally a third model 97 including a third plurality of parameters 99, the output of the third model being used to (i) provide a third interaction score for the interaction between the corresponding compound (particular training or test compound) and the target polymer, and / or (ii) condition any other single model and / or group of models; ●Optionally, store any number of additional xth models, each such additional xth model including a corresponding plurality of parameters, and the output of the xth model is used, at least in part, to (i) provide a characterization of the interaction between a corresponding compound (a particular training or test compound) and the target polymer, and / or (ii) the xth model is used to condition any other single models and / or groups of models.

[0048] In some implementations, one or more of the above identified data elements or modules of computer system 100 are stored in one or more of the aforementioned memory devices and correspond to sets of instructions for performing the functions described above. The above identified data, modules or programs (e.g., sets of instructions) need not be implemented as separate software programs, procedures, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various implementations. In some implementations, memory 92 and / or 90 (and optionally 52) optionally stores a subset of the above identified modules and data structures. Furthermore, in some embodiments, memory 92 and / or 90 (and optionally 52) stores additional modules and data structures not described above. In some embodiments, first neural network 72 is replaced with another form of model.

[0049] Disclosed herein is a system for characterizing interactions between test compounds and target polymers, and a method for carrying out such characterization is detailed with reference to FIG. 2 and discussed below.

[0050] Block 200. Referring to block 200 of FIG. 2A, a computer system 100 for characterizing interactions between a test compound and a target polymer is provided. As described above in connection with FIG. 1, the computer system 100 comprises one or more processors 59 and a memory 90 / 92 addressable by the one or more processors. The memory stores at least one program for execution by the one or more processors. The at least one program includes instructions as detailed below.

[0051] Blocks 202-218. Referring to block 202 of FIG. 2A, a plurality of atomic coordinate sets are obtained. Each respective atomic coordinate set 49 in the plurality of atomic coordinate sets includes a test compound bound to the target polymer at a corresponding pose 48 in the plurality of poses. In other words, referring to FIG. 1A, each respective atomic coordinate set 49 includes both the atomic coordinates of the test compound and at least a subset of the spatial coordinates of the target polymer considered by the first neural network. For example, in some embodiments, the atomic coordinate set 49 consists of the atomic coordinates of the test compound and the atomic coordinates of a portion of the target polymer that constitutes an active site to which the test compound is docked. In some embodiments, the target polymer includes a plurality of active sites, and the test compound is docked in one of the active sites. FIG. 3 illustrates a pose 48 of a test compound in an active site of a target polymer 38. Each respective atomic coordinate set in the plurality of atomic coordinate sets includes atomic coordinates for at least 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, 250, or 300 atoms. In some embodiments, an atomic coordinate set in the plurality of atomic coordinate sets includes atomic coordinates for at least 400 atoms of the target polymer in addition to the atomic coordinates of the test compound. In some embodiments, an atomic coordinate set in the plurality of atomic coordinate sets includes atomic coordinates for at least 25 atoms, at least 50 atoms, at least 100 atoms, at least 200 atoms, at least 300 atoms, at least 400 atoms, at least 1000 atoms, at least 2000 atoms, or at least 5000 atoms of the target polymer in addition to the atomic coordinates of the test compound. In some embodiments, only the coordinates of the active site of the target polymer 38 where the ligand is expected to bind to the target polymer are present in each respective atomic coordinate set in the multiple atomic coordinate sets, in addition to the coordinates of the corresponding test compound.

[0052] Referring to block 204 of FIG. 2A, in some embodiments, the target polymer is an assembly of proteins, polypeptides, polynucleic acids, polyribonucleic acids, polysaccharides, or any combination thereof.

[0053] In some embodiments, the target polymer 38 is a macromolecule made up of repeating residues. In some embodiments, the target polymer 38 is a natural material. In some embodiments, the target polymer 38 is a synthetic material. In some embodiments, the target polymer 38 is an elastomer, shellac, amber, natural or synthetic rubber, cellulose, bakelite, nylon, polystyrene, polyethylene, polypropylene, polyacrylonitrile, polyethylene glycol, or a polysaccharide.

[0054] In some embodiments, the target polymer 38 is a heteropolymer (copolymer). A copolymer is a polymer derived from two (or more) monomer species, as opposed to a homopolymer, where only one monomer is used. Copolymerization refers to the method used to chemically synthesize a copolymer. Examples of copolymers include, but are not limited to, ABS plastic, SBR, nitrile rubber, styrene acrylonitrile, styrene-isoprene-styrene (SIS), and ethylene vinyl acetate. Because copolymers contain at least two types of building blocks (structural units, or particles as well), copolymers can be classified based on how these units are arranged along the chain. These include alternating copolymers, in which A and B units regularly alternate. See, for example, Jenkins, 1996, “Glossary of Basic Terms in Polymer Science,” Pure Appl. Chem. 68(12):2287-2311, which is incorporated herein by reference in its entirety. Additional examples of copolymers include copolymers that contain repeating sequences (e.g., (ABABBAAAABBB) n) are periodic copolymers with A and B units arranged in a cyclic manner. Additional examples of copolymers are statistical copolymers, in which the arrangement of monomer residues in the copolymer follows statistical laws. See, for example, Painter, 1997, Fundamentals of Polymer Science, CRC Press, 1997, p14, which is incorporated herein by reference in its entirety. Yet another example of a copolymer that can be evaluated using the disclosed systems and methods is a block copolymer, which includes two or more homopolymer subunits linked by covalent bonds. The linking of the homopolymer subunits may require an intermediate non-repeating subunit, known as a junction block. Block copolymers with two or three distinct blocks are called diblock and triblock copolymers, respectively.

[0055] In some embodiments, the target polymer 38 comprises 50 or more, 100 or more, 150 or more, 200 or more, 300 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more atoms.

[0056] In some embodiments, the target polymer 38 is actually a plurality of polymers (e.g., 2 or more, 3 or more, 10 or more, 100 or more, 1000 or more, or 5000 or more polymers), where each polymer in the plurality does not all have the same molecular weight. In some such embodiments, the target polymers 38 in the plurality of polymers share at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% sequence identity and fall into weight ranges with a corresponding chain length distribution. In some embodiments, the target polymer 38 is a branched polymer molecule that includes a backbone with one or more substituent side chains or branches. Types of branched polymers include, but are not limited to, star polymers, comb polymers, brush polymers, dendronized polymers, ladders, and dendrimers. See, for example, Rubinstein et al., 2003, Polymer physics, Oxford; New York: Oxford University Press. p. 6, which is incorporated herein by reference in its entirety.

[0057] In some embodiments, the target polymer is a polypeptide. As used herein, the term "polypeptide" refers to two or more amino acids or residues linked by a peptide bond. The terms "polypeptide" and "protein" are used interchangeably herein and include oligopeptides and peptides. An "amino acid", "residue" or "peptide" refers to any of the 20 standard structural units of proteins, as known in the art, including imino acids such as proline and hydroxyproline. The names of amino acid isomers may include D, L, R, and S. The definition of amino acid includes unnatural amino acids. Thus, selenocysteine, pyrrolysine, lanthionine, 2-aminoisobutyric acid, gamma-aminobutyric acid, dehydroalanine, ornithine, citrulline, and homocysteine, as non-limiting examples, are all considered amino acids. Other variants or analogs of amino acids are known in the art. Thus, a polypeptide may include synthetic peptide analog structures, such as peptoids. See Simon et al., 1992, Proceedings of the National Academy of Sciences USA, 89, 9367, which is incorporated herein by reference in its entirety. See also Chin et al., 2003, Science 301, 964, and Chin et al., 2003, Chemistry & Biology 10, 511, each of which is incorporated herein by reference in its entirety.

[0058] The target polymers 38 evaluated according to some embodiments of the disclosed systems and methods may also have any number of post-translational modifications. Thus, target polymers 38 include those polymers that have been modified by acylation, alkylation, amidation, biotinylation, formylation, gamma-carboxylation, glutamylation, glycosylation, glycylation, hydroxylation, iodination, isoprenylation, lipoylation, cofactor addition (e.g., heme, flavin, metals, etc.), addition of nucleosides and their derivatives, oxidation, reduction, pegylation, phosphatidylinositol addition, phosphopantetheinylation, phosphorylation, pyroglutamate formation, racemization, addition of amino acids by tRNA (e.g., arginylation), sulfation, selenoylation, ISGylation, sumoylation, ubiquitination, chemical modifications (e.g., citrullination and deamidation), and processing by other enzymes (e.g., proteases, phosphatases, and kinases). Other types of post-translational modifications are known in the art and are within the scope of the target polymer 38 of the present disclosure.

[0059] In some embodiments, the target polymer 38 is a surfactant. A surfactant is a compound that reduces the surface tension of a liquid, the interfacial tension between two liquids, or the interfacial tension between a liquid and a solid. Surfactants can function as detergents, wetting agents, emulsifiers, foaming agents, and dispersants. Surfactants are usually organic compounds that are amphiphilic, meaning that they contain both hydrophobic groups (their tails) and hydrophilic groups (their heads). Thus, surfactant molecules contain both water-insoluble (or oil-soluble) and water-soluble components. Surfactant molecules diffuse into water and, when water is mixed with oil, adsorb to the air-water interface or the oil-water interface. The insoluble hydrophobic groups can extend from the bulk water phase into the air or oil phase, while the water-soluble head groups remain in the water phase. This alignment of surfactant molecules at the surface modifies the surface properties of the water at the water / air or water / oil interface.

[0060] Examples of ionic surfactants include ionic surfactants such as anionic, cationic, or zwitterionic (ampotelic) surfactants, in some embodiments, the target entity 58 is a reverse micelle or a liposome.

[0061] In some embodiments, the target polymer 38 is a fullerene. A fullerene is any molecule composed entirely of carbon in the form of a hollow sphere, ellipsoid, or tube. Spherical fullerenes are also called buckyballs and resemble the balls used in soccer. Cylindrical ones are called carbon nanotubes or buckytubes. Fullerenes are similar in structure to graphite, which consists of stacked graphene sheets of interlocking hexagonal rings, but may also contain pentagonal (or sometimes heptagonal) rings.

[0062] 2A , block 206, in some embodiments, each atomic coordinate set 49 in the plurality of atomic coordinate sets comprises three-dimensional coordinates {x1, ..., x2, ... N In some embodiments, the portion of the polymer from the crystal structure of the target polymer consists of atomic coordinates for less than 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, 250, or 300 atoms. In some embodiments, the portion of the polymer from the crystal structure of the target polymer consists of atomic coordinates for less than 400 atoms of the target polymer in addition to the atomic coordinates of the test compound. In some embodiments, the portion of the polymer from the crystal structure of the target polymer consists of atomic coordinates for less than 25 atoms, less than 50 atoms, less than 100 atoms, less than 200 atoms, less than 300 atoms, less than 400 atoms, less than 1000 atoms, less than 2000 atoms, or less than 5000 atoms of the target polymer.

[0063] In some embodiments, each atomic coordinate set 49 in the plurality of atomic coordinate sets is a set of three-dimensional coordinates {x1,...,x2,...} for at least a portion of a polymer from a structure prediction program such as AlphaFold2. N See Jumper et al., 2021, “Highly accurate protein structure prediction with AlphaFold,” Nature 596, pp. 583-589, which is incorporated herein by reference.

[0064] Referring to block 208 of FIG. 2A, in some embodiments, each atomic coordinate set 49 in the plurality of atomic coordinate sets comprises a collection of three-dimensional coordinates of at least a portion of the target polymer determined by nuclear magnetic resonance, neutron diffraction, or cryo-electron microscopy. In some embodiments, this portion of the polymer comprises atomic coordinates of less than 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, 250, or 300 atoms. In some embodiments, this portion of the polymer comprises atomic coordinates of less than 400 atoms of the target polymer in addition to the atomic coordinates of the test compound. In some embodiments, the portion of the polymer comprises atomic coordinates of less than 25 atoms, less than 50 atoms, less than 100 atoms, less than 200 atoms, less than 300 atoms, less than 400 atoms, less than 1000 atoms, less than 2000 atoms, or less than 5000 atoms of the target polymer. In some embodiments, the collection of three-dimensional coordinates includes 10 or more, 20 or more, or 30 or more atomic structures of at least a portion of the target polymer that have a backbone RMSD of 1.0 Å or more, 0.9 Å or more, 0.8 Å or more, 0.7 Å or more, 0.6 Å or more, 0.5 Å or more, 0.4 Å or more, 0.3 Å or more, or 0.2 Å or more.

[0065] In some embodiments, the target polymer 38 includes two different types of polymers, such as a nucleic acid bound to a polypeptide. In some embodiments, the natural target polymer includes two polypeptides bound to each other. In some embodiments, the natural target polymer under study includes one or more metal ions (e.g., a metalloproteinase having one or more zinc atoms). In such cases, the metal ions and / or small organic molecules may be included in the atomic coordinates 40 of the target polymer.

[0066] In some embodiments, target polymer 38 is a polymer having 10 or more, 20 or more, 30 or more, 50 or more, 100 or more, 100-1000, or less than 500 residues present in the target polymer.

[0067] In some embodiments, the atomic coordinates of the target polymer 38 are determined using modeling methods such as ab initio methods, density functional methods, semi-empirical and empirical methods, molecular mechanics, chemical mechanics, or molecular mechanics.

[0068] In some embodiments, each respective set of atomic coordinates 49 is represented by the Cartesian coordinates of the centers of atoms that make up the target polymer 38. In some alternative embodiments, each respective set of atomic coordinates 49 is represented by the electron density of the target polymer, e.g., as measured by X-ray crystallography. For example, in some embodiments, the atomic coordinates 40 are represented by 2F observed -F calculated Includes electron density maps, F observed is the amplitude of the observed structure factor of the target polymer, F calculated is the amplitude of the structure factor calculated from the calculated atomic coordinates of the target polymer 38.

[0069] In various other embodiments, each respective set of atomic coordinates 49 is obtained from any of a variety of sources, including, but not limited to, structural ensembles generated by solution NMR, co-complexes as interpreted from X-ray crystallography, neutron diffraction, cryo-electron microscopy, sampling from computer simulations, homology modeling, rotamer library sampling, or any combination thereof.

[0070] Referring to block 210 of FIG. 2A, in some embodiments, the compound satisfies two or more, three or more, or all four of Lipinski's rule of five: (i) five or less hydrogen bond donors, (ii) ten or less hydrogen bond acceptors, (iii) molecular weight less than 500 Daltons, and (iv) LogP less than 5. See Lipinski, 1997, Adv. Drug Del. Rev. 23, 3, which is incorporated herein by reference in its entirety. In some embodiments, the test compound satisfies one or more criteria in addition to Lipinski's rule of five. For example, in some embodiments, the test compound has five or less aromatic rings, four or less aromatic rings, three or less aromatic rings, or two or less aromatic rings.

[0071] Referring to block 214 of Figure 2A, in some embodiments, the test compound is an organic compound having a molecular weight of less than 500 Daltons, less than 1000 Daltons, less than 2000 Daltons, less than 4000 Daltons, less than 6000 Daltons, less than 8000 Daltons, less than 10000 Daltons, or less than 20000 Daltons. However, some embodiments of the disclosed systems and methods do not have a limit on the size of the test compound. For example, in some embodiments, the test compound is itself a large polymer, such as an antibody.

[0072] Referring to block 216 of FIG. 2A, in some embodiments, the test compound is an organic compound having a molecular weight between 400 Daltons and 10,000 Daltons.

[0073] Referring to block 218 of FIG. 2A, in some embodiments, the plurality of atomic coordinate sets consists of 3-64 poses. In some embodiments, the target polymer 38 is a polymer having active sites, and each of the poses is obtained by docking a test compound into the active sites of the target polymer. In some embodiments, the test compound is docked onto the target polymer 38 multiple times to form a plurality of poses. In some embodiments, the test compound is docked onto the target polymer 38 two, three, four, five or more times, ten or more times, fifty or more times, one hundred or more times, or one thousand or more times. Each such docking represents a different pose of the test compound docked onto the target polymer 38. In some embodiments, the target polymer 38 is a polymer having active sites, and the test compound is docked into the active sites in each of a plurality of different ways, each such way representing a different pose. In some embodiments, the target polymer includes a plurality of active sites, and the test compound is docked into one of the active sites in each of a plurality of different ways, each such way representing a different pose. In some such embodiments, separate studies are also conducted individually on one or more of the other active sites of the target polymer using the systems and methods of the present disclosure.

[0074] In some embodiments, each pose of the test compound is determined by AutoDock Vina. See Trott and Olson, "AutoDock Vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization and multithreading," Journal of Computational Chemistry 31 (2010) 455-461. In some embodiments, one docking program is used to determine some of the poses of the test compound, and another docking program is used to determine other poses of the test compound.In some embodiments, Quick Vina 2(Alhossary et al.,2015,“Fast,accurate,and reliable molecular docking with QuickVina,”Bioinformatics 31:13,pp.2214-2216),VinaLC(Zhang et al.,2013,“Message Passing Interface and Multithreading Hybrid for Parallel Molecular Docking of Large Databases on Petascale High Performance Computing Machines,”J.Comput.Chem.DOI:10.1002 / jcc.23214), Smina(Koes et al,,2013,“Lessons learned in empirical scoring with smina from the CSAR 2011 benchmarking exercise,”Journal of chemical information and modeling 53:8,pp.1893-1904), or CUina(Morrison et al. al.,“Efficient GPU Implementation of AutoDock Vina,”COMP poster No. 3,432,389) were used to determine the pose of the test compounds, each of which is incorporated herein by reference.

[0075] In some embodiments, the multiple atomic coordinate sets are an ensemble from an ensemble docking algorithm such as that disclosed in Stafford et al., 2022, “AtomNet PoseRanker: Enriching Ligand Pose Quality for Dynamic Proteins in Virtual High-Throughput Screens,” Journal of Chemical Information and Modeling 62, pp. 1178-1189, which is incorporated herein by reference. In some such embodiments, the ensemble consists of 3-64, 4-128, 5-32, more than 5, or 8-25 structurally similar poses.

[0076] In some embodiments, each pose of the docked test compound is scored against multiple (e.g., 2-100) different conformations of the target protein. In some embodiments, each pose (e.g., in the ensemble of poses) is scored against a fixed conformation of the target protein.

[0077] In some embodiments, the test compound is docked to the target polymer 38 by either random pose generation techniques or biased pose generation. In some embodiments, the test compound is docked to the target polymer 38 by Markov chain Monte Carlo sampling. In some embodiments, such sampling allows for sufficient flexibility of the test compound in the docking calculations and a scoring function that is the sum of the interaction energy between the test compound and the target polymer 38 and the conformational energy of the test compound. See, e.g., Liu and Wang, 1999, "MCDOCK: A Monte Carlo simulation approach to the molecular docking problem," Journal of Computer-Aided Molecular Design 13, 435-451, which is incorporated herein by reference. In some such embodiments, the pose represented by the multiple atomic coordinate sets is the pose that receives the top score relative to all other poses tested (e.g., the top 256 scores, thus 256 poses; the top 128 scores, thus 128 poses; the top 64 scores, thus 64 poses; the top 32 scores, thus 32 poses, etc.).

[0078] In some embodiments, algorithms such as DOCK (Shoichet, Bodian, and Kuntz, 1992, "Molecular docking using shape descriptors," Journal of Computational Chemistry 13(3), pp. 380-397, and Knegtel et al., 1997 "Molecular docking to ensembles of protein structures," Journal of Molecular Biology 266, pp. 424-440, each of which is incorporated herein by reference) are used to find multiple poses of the test compound relative to the target polymer 38. Such algorithms model the target polymer 38 and the test compound as rigid bodies. The docked conformations are searched using complementary surfaces to find poses 48.

[0079] In some embodiments, AutoDOCK (Morris et al., 2009, “AutoDock4 and AutoDockTools4: Automated Docking with Selective Receptor Flexibility,” J. Comput. Chem. 30(16), pp. 2785-2791; Sotriffer et al., 2000, “Automated docking of ligands to antibodies: methods and applications,” Methods: A Companion to Methods in Enzymology 20, pp. 280-291, and Morris et al., 1998, “Automated Docking Using a Lamarckian Genetic Algorithm and Empirical Binding Free Energy Function,” Journal of Computational Chemistry 20, pp. 2771-2779, each of which is incorporated herein by reference) is used. 19:1639-1662, each of which is incorporated herein by reference) to find multiple poses of the test compound relative to the target polymer 38. AutoDOCK uses a kinetic model of the ligand and supports Monte Carlo, simulated annealing, Lamarckian genetic algorithms, and genetic algorithms. Thus, in some embodiments, multiple different poses of the test compound are obtained by Markov chain Monte Carlo sampling, simulated annealing, Lamarckian genetic algorithms, or genetic algorithms using a docking scoring function.

[0080] In some embodiments, an algorithm such as FlexX (Rarey et al., 1996, "A Fast Flexible Docking Method Using an Incremental Construction Algorithm," Journal of Molecular Biology 261, pp. 470-489, incorporated herein by reference) is used to find multiple poses of the test compound relative to the target polymer. FlexX uses a greedy algorithm to perform a sequential construction of the test compound at the active site of the target polymer 38. Thus, in some embodiments, multiple different poses of the test compound are obtained by the greedy algorithm.

[0081] In some embodiments, an algorithm such as GOLD (Jones et al., 1997, "Development and Validation of a Genetic Algorithm for flexible Docking," Journal Molecular Biology 267, pp. 727-748, incorporated herein by reference) is used to find multiple poses of the test compound relative to the target polymer 38. GOLD stands for Genetic Optimization for Ligand Docking. GOLD builds a genetically optimized hydrogen bond network between the test compound and the target polymer 38.

[0082] In some embodiments, molecular mechanics is performed on the target polymer (or a portion thereof, such as the active site of the target polymer) and the test compound to identify multiple poses. During the molecular mechanics run, the atoms of the target polymer and the test compound are allowed to interact for a fixed period of time, providing a view of the dynamic evolution of the system. The trajectories of the atoms in the target polymer and the test compound are determined by numerically solving Newton's equations of motion for the interacting particle system, and the forces between the particles and their respective potential energies are calculated using interatomic potentials or molecular mechanics force fields. See Alder and Wainwright, 1959, "Studies in Molecular Dynamics. I. General Method," J. Chem. Phys. 31(2):459, and Bibcode, 1959, J. Ch. Ph. 31, 459A, doi:10.1063 / 1.1730376, each of which is incorporated herein by reference. Thus, in this manner, the molecular mechanics run generates trajectories of the target polymer and each test compound over time. The trajectory includes trajectories of atoms in the target polymer and the test compound. In some embodiments, a subset of the multiple different poses is obtained by taking snapshots of the trajectory over a period of time. In some embodiments, the poses are obtained from snapshots of several different trajectories, each trajectory including a different molecular mechanics run of the target polymer interacting with the test compound. In some embodiments, prior to running the molecular mechanics, the test compound is first docked into the active site of the target polymer using a docking technique.

[0083] Blocks 220-238. Referring to block 220 of FIG. 2B, for each respective atomic coordinate set 49 in the plurality of atomic coordinate sets, the respective atomic coordinate set, or the encoding of the respective atomic coordinate set, is input to a first neural network 72 to obtain a corresponding initial embedding 74 as an output of the first neural network, thereby obtaining a plurality of initial embeddings 74-1, ..., 74-N. Each initial embedding 74 in the plurality of initial embeddings 74-1, ..., 74-N corresponds to an atomic coordinate set 49 in the plurality of atomic coordinate sets 49-1-1, ..., 49-1-N. In some embodiments, the first neural network 72 includes more than 400 parameters. In some embodiments, the respective atomic coordinate set, or the encoding of the respective atomic coordinate set, input to the first neural network 72 consists of 20 bits to 20,000 bits of information. In some embodiments, each set of atomic coordinates, or the encoding of each set of atomic coordinates, input to the first neural network 72 includes 20 bits, 40 bits, 60 bits, 80 bits, 100 bits, 200 bits, 300 bits, 400 bits, 500 bits, 600 bits, 700 bits, 800 bits, 900 bits, or 1000 bits of information. In some embodiments, each set of atomic coordinates, or the encoding of each set of atomic coordinates, input to the first neural network 72 includes 2000 bits, 4000 bits, 6000 bits, 8000 bits, or 10,000 bits of information. In some embodiments, the first neural network 72 includes more than 400 parameters, more than 1000 parameters, more than 2000 parameters, more than 5000 parameters, more than 10,000 parameters, more than 100,000 parameters, or more than 1×10 6 In some embodiments, the amount of information in each set of atomic coordinates input to the first neural network combined with the number of parameters of the neural network may require more than 10,000 calculations, more than 100,000 calculations, 1×10 calculations, or more than 1×10 calculations to calculate the initial embedding 74 using the first neural network 74.6 More than 5×10 calculations 6 More than 1×10 calculations 7 This results in a performance of more than 1000 calculations.

[0084] Referring to block 222 of Figure 2B, in some embodiments, the first neural network 72 is a convolutional neural network. Referring to block 224 of Figure 2B, in some such embodiments, each respective atomic coordinate set 49 in the plurality of atomic coordinate sets includes atomic coordinates of at least 80 atoms (e.g., atoms of the target compound in addition to atoms of the target protein considered by the first neural network).

[0085] In some embodiments, each set of atomic coordinates 49 is converted into a corresponding voxel map 52. The corresponding voxel map 52 thus represents the test compound relative to the target polymer 38 at the corresponding pose 48. In some such embodiments, the voxel map 52 is expanded into a corresponding vector 54 and input to a first neural network 72. In some such embodiments, the first neural network 72 is a convolutional neural network that provides as an output a corresponding initial embedding 74 for the corresponding pose 48. In some such embodiments, the corresponding vector 54 referenced above is a one-dimensional vector. In some embodiments, the corresponding vector 54 includes 10 or more elements, 20 or more elements, 100 or more elements, 500 or more elements, 1000 or more elements, or 10,000 or more elements. In some such embodiments, each such element is represented by a different bit in a data structure that is input to the first neural network.

[0086] In some embodiments, the first neural network 72 is any of the convolutional neural networks disclosed in Wallach et al., 2015, “AtomNet: A Deep Convolutional Neural Network for Bioactivity Prediction in Structure-based Drug Discovery,” arXiv:1510.02855v1, or U.S. Pat. Nos. 11,080,570, 10,546,237, 10,482,355, 10,002,312, or 9,373,059, each of which is incorporated herein by reference. Further details regarding using a convolutional neural network as the first neural network 72 to obtain corresponding initial embeddings 74 for corresponding poses 48 of the test compound relative to the target polymer 38 are disclosed below in the section entitled “Using a Convolutional Neural Network as the First Neural Network 72.”

[0087] In some embodiments, the first neural network 72 is an equivariant neural network. Non-limiting examples of equivariant convolutional neural networks include Thomas et al., 2018, “Tensor field networks: Rotation-and translation-equivariant neural networks for 3D point clouds,” arXiv:1802.08219; Anderson et al., 2019, “Cormorant: Covariant Molecular Neural Networks,” Neural Information Processing Systems; Johannes et al., 2020, “Directional Message Passing For Molecular Graphs,” International Conference on Learning Representations; Townshend et al., 2021, “ATOM3D: Tasks On Molecules in Three Dimensions,” International Conference on Learning Representations; Jing et al., 2009, “Learning from Protein Structure with Geometric Vector Perceptrons,” arXiv:2009.01411; and Satorras et al., 2021, “E(n)Equivariant Graph Neural Networks,” arXiv:2102.09844, each of which is incorporated herein by reference in its entirety.

[0088] Referring to block 228 of Figure 2C, in some embodiments, the first neural network 72 is a graph neural network. Referring to block 230 of Figure 2C, in some such embodiments, the graph neural network is characterized by an initial embedding layer and multiple interaction layers that each contribute to an interaction data structure in multiple interaction data structures for each atom in a respective set of atomic sets of atomic coordinates 49 of corresponding poses 48 in the multiple poses that are pooled to form a corresponding initial embedding 74 of the corresponding pose 48.

[0089] 9B illustrates an embodiment of the present disclosure where the first neural network 72 is a graph neural network (GCN). In some embodiments, the GCN takes the three-dimensional coordinates of the protein-ligand poses 48 along with one-hot atomic encoding that simultaneously identifies the element type, the target protein / compound, and the hybridization state. In some embodiments, the connectivity is defined purely by radial functions without the use of chemical bonds. In some embodiments, a radial cutoff at layer l is used.

number

number

number

number

number

number

[0090] In some embodiments, the convolution kernel W l (r)

number

number

number

number

number

number

number

number

number

number

number

number

number

[0091] FIG. 20 illustrates a system for characterizing the interaction between a test compound and a target polymer, where the characterization is based on (i) pose ranking and (ii) Vina score, as well as activity. In this embodiment, the covalently embedded z read =g(f 0…L ), PoseRanker score

number

number

number

number

number

[0092] Additional non-limiting examples of graph convolutional neural networks include Behler Parrinello, 2007, “Generalized Neural-Network Representation of High Dimensional Potential-Energy Surfaces,” Physical Review Letters 98, 146401; Chmiela et al., 2017, “Machine learning of accurate energy-conserving molecular force fields,” Science Advances 3(5):e1603015; Schuett et al., 2017, “SchNet: A continuous-filter convolutional neural network for modeling quantum interactions,” Advances in Neural Information Processing Systems 30, pp. 992-1002; Feinberg et al., 2018, “PotentialNet for Molecular Property Prediction,” ACS Cent. Sci. 4, 11, 1520-1530; and Stafford et al., “AtomNet PoseRanker: Enriching Ligand Pose Quality for Dynamic Proteins in Virtual High Throughput Screens,” https: / / chemrxiv.org / engage / chemrxiv / article-details / 614b905e39ef6a1c36268003, each of which is incorporated herein by reference.

[0093] 2C, block 232, in some embodiments, the first neural network 72 is an equivariant neural network or a message passing neural network. See Bao and Song, 2020, “Equivariant Neural Networks and Equivarification,” arXiv:1906.07172v4, and Gilmer et al., 2020, “Message Passing Neural Networks,” In: Schuett et al. (eds), Machine Learning Meets Quantum Physics, Lecture Notes in Physics 968, Springer, Cham, each of which is incorporated herein by reference.

[0094] 2C, in some embodiments, the first neural network 72 includes multiple graph convolution blocks, each using multiple radial graphs to consider connectivity at a respective set of atomic coordinates. For example, in some such embodiments, the first neural network 72 includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more graph convolution blocks, each using multiple radial graphs to consider connectivity at a respective set of atomic coordinates.

[0095] Referring to block 236 of FIG. 2C, in some embodiments, the first neural network 72 is 6 In some embodiments, the first neural network 72 includes 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 10,000, 50,000, 100,000, or 1×10 parameters. 6 Contains more than one parameter.

[0096] Referring to block 238 of FIG. 2C, in some embodiments, the corresponding initial embedding 74 includes a data structure having between 128 and 768 values.

[0097] Blocks 240-246. Referring to block 240 of FIG. 2C, an attention mechanism 77 is applied to the multiple initial embeddings (74-1-74-P) in a concatenated form, thereby obtaining an attention embedding 79. In some embodiments, the multiple initial embeddings consist of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more than 20 initial embeddings. In some embodiments, the attention mechanism is a mapping of a query (multiple initial embeddings in a concatenated form) and a set of key-value pairs to an output (attention embedding 79) where the query, key, value, and output are all vectors. In some such embodiments, the output (attention embedding 79) is computed as a weighted sum of values, where the weight assigned to each value is computed by a fitness function between the query (multiple initial embeddings in a concatenated form) and the corresponding key.

[0098] Thus, according to block 240, each of the initial embeddings 74 for each of the poses of the compound are concatenated together and applied to the attention mechanism. For example, if there are five poses 48 for a compound resulting in five initial embeddings 74, the five initial embeddings are z cat are linked together to form this z catis applied to an attention mechanism 77 to obtain attention embeddings 79. Exemplary attention mechanisms are described in Chaudhari et al., July 12, 2021 “An Attentive Survey of Attention Models,” arXiv:1904-02874v3, and Vaswani et al., “Attention is All You Need,” 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, California, USA, each of which is incorporated herein by reference. The attention mechanism 77 exploits the inference that some parts of the pose 48 are more important than other parts, and therefore some parts (elements or sets of elements) in the initial embeddings 74 are more important than other parts. For example, in an example where each initial embedding consists of 20 elements, elements 1-4 and 9-15 may contain more information regarding the characterization of the interaction between the compound and the target polymer than elements 5-9 and 16-20. The attention mechanism is trained to find such observations using the training compounds, and then this learned (trained) observation is applied to the initial embeddings 74 of the test compounds to form the attention embeddings. Thus, the attention mechanism allows a model downstream of the attention mechanism (e.g., the first model 89) to recognize certain parts of the input embeddings (e.g., z attn , z pool ), which helps to effectively perform the task at hand (characterizing the interaction between a test compound and a target polymer).

[0099] Referring to block 244 of FIG. 2D , in some embodiments, the plurality of initial embeddings 74 (e.g., in a concatenated form) includes a first plurality of values, and applying the attention mechanism 77 includes (i) inputting the first plurality of values ​​into an attention neural network to obtain a first plurality of weights, where each weight in the first plurality of weights corresponds to a respective value in the first plurality of values, and (ii) weighting each respective value in the first plurality of values ​​by a corresponding weight in the plurality of weights, thereby obtaining an attention embedding. Thus, for example, consider the case where there are five poses 48 for a compound, resulting in five initial embeddings 74, each of the five initial embeddings including 20 values. Thus, the concatenation of the five initial embeddings 74 results in a pooled vector (z shown in FIG. 9A ) having 100 elements, where each element has a value. cat ) results. cat As a result of inputting z cat It returns 100 weights, one for each element of z cat Adjust the corresponding values ​​in the z attn Therefore, z cat If z has 100 elements, then in some embodiments, attn z also has 100 elements. attn The value of element 1 of is (a)z cat (b) the value of element 1 of z cat is multiplied by the weight of element 1 of z attn The value of element 2 of (a)z cat (b) the value of element 2 of z cat with the weight of element 2, and so on.

[0100] 2D, in some such embodiments, the first plurality of weights sums to 1 (or some other constant value), and each weight in the first plurality of weights is a scalar value between 0 and 1 (or some other constant value). cat In the above example where the attention neural network has 100 elements and therefore returns 100 weights, the sum of the 100 weights will be 1 (or some other constant value) according to block 246, which are scalar values ​​(or some other constant value) between 0 and 1. In embodiments with an attention neural network, the attention neural network is jointly trained with at least the first neural network 72 and the first model 89 against known indicators (e.g., pKa, activity, binding score) of the multiple training compounds. In some embodiments, the multiple training compounds are 25 or more training compounds, 100 or more training compounds, 500 or more training compounds, 1000 or more training compounds, 10,000 or more training compounds, 100,000 or more training compounds, or 1×10 6 training compounds.

[0101] Blocks 248-256. Referring to block 248 of FIG. 2D, a pooling function 81 is applied to the attention embeddings 79 to derive a pooled embedding 85. Examples of pooling functions include, but are not limited to, average, sum, or max pooling. Thus, there are five poses 48 for a compound resulting in five initial embeddings 74, each of which contains 20 values, and thus z cat , and z attn Consider the example above where z has 100 elements. attn Apply the pooling function 81 to z pool Consider the case where the pooling function is an averaging function. In this example, z pool has 20 elements, z pool The first element of z attn is the average of elements 1, 21, 41, 61, and 81 of z pool The second element of zattn is the average of elements 2, 22, 42, 62, and 82 of, …, z pool The 20th element of z attn In other words, referring to block 250 of FIG. 2D , in some such embodiments, the pooling function 81 applies a statistical function to the z attn By combining the corresponding parts of the attention embeddings of z to form a pooled embedding, attn The attention embeddings of 79 are pooled embeddings of 85, z pool Fold it into.

[0102] Now, considering again the example above where there are five poses 48 for a compound resulting in five initial embeddings 74, each of the five initial embeddings contains 20 values, thus z cat , and z attn has 100 elements, and z pool To generate z attn The pooling function 81 applied to z is a sum function. In this example, z pool has 20 elements, z pool The first element of z attn is the sum of elements 1, 21, 41, 61, and 81 of z pool The second element of z attn is the sum of elements 2, 22, 42, 62, and 82 of z pool The 20th element of z attn In other words, referring to block 250 of FIG. 2D , in some such embodiments, the pooling function 81 applies a statistical function to the z attn We combine the corresponding parts of the attention embeddings 79 of z (now weighted by the attention mechanism) to form a pooled embedding 85. attn Collapse to pooled embedding 85.

[0103] Referring to block 252 of FIG. 2D, in some embodiments, attention embedding 79, z attn z includes corresponding values ​​for corresponding elements for each respective pose 48 in the plurality of poses, and the statistical pooling function combines the attention embedding 79, z attn , where z is a maximal function that takes the maximum over the corresponding elements of each respective pose 48 represented in . To illustrate, consider again the example above where there are five poses 48 for a compound resulting in five initial embeddings 74. Each of the five initial embeddings contains 20 values, and thus z cat , and z attn has 100 elements, z pool To generate z attn The pooling function 81 applied to z is the max function. In this example, pool has 20 elements, z pool The first element of z attn is the maximum value among elements 1, 21, 41, 61, and 81 of z pool The second element of z attn is the maximum value among elements 2, 22, 42, 62, and 82 of, ..., z pool The 20th element of z attn is the maximum value among elements 20, 40, 60, 80, and 100.

[0104] 2D, block 256, in some embodiments, the attention embedding 79 includes corresponding values ​​for corresponding elements for each respective pose 48 in the plurality of poses, and the statistical pooling function is an averaging function that averages corresponding elements for each respective pose 48 represented in the attention embedding 79 to form the pooled embedding 85. Such an embodiment is similar to the max pooling function described above, except that an averaging function is applied instead of a max function.

[0105] Blocks 258-280. Referring to block 258 of FIG. 2E, the pooled embeddings 85 are input into a first model 89, thereby obtaining a first interaction score for the interaction between the test compound and the target polymer. In some embodiments, the first model 89 includes more than 400 parameters 91. In some embodiments, the first model 89 includes more than 400 parameters, more than 1000 parameters, more than 2000 parameters, more than 5000 parameters, more than 10,000 parameters, more than 100,000 parameters, or more than 1×10 6 In some embodiments, the amount of information in the pooled embeddings input to the first model 89 combined with the number of parameters of the first model 89 may require more than 10,000 calculations, more than 100,000 calculations, more than 1×10 calculations, or more than 1×10 parameters to calculate the first interaction score. 6 More than 5×10 calculations 6 More than 1×10 calculations 7 This results in a performance of more than 1000 calculations.

[0106] In some embodiments, the system 100 or one or more programs hosted, stored, or addressable by the system 100 can characterize the interaction between the test compound and the target polymer 38. In some such embodiments, this characterization is a discrete (e.g., discrete binary) activity score. In other words, the characterization is categorical. For example, referring to block 262 of FIG. 2E, in some embodiments, the first model 89 is a classification task and the first interaction score classifies the interaction between the test compound and the target polymer. In some such embodiments, the characterization (e.g., the first interaction score) is discrete binary and the computer system provides one value, e.g., "1", when the test compound is determined to be active against the target polymer by the in silico methods disclosed herein, and provides another value, e.g., "0", when the test compound is determined to be not active against the target polymer.

[0107] In some embodiments, the characterization (e.g., the first interaction score) is on a discrete scale other than binary. For example, in some embodiments, the characterization provides a first value, e.g., "0", if the test compound is determined by the in silico methods disclosed herein to have activity below a first threshold, a second value, e.g., "1", if the test compound is determined to have activity between the first and second thresholds, and a third value, e.g., "2", if the test compound is determined to have activity above the second threshold. In such embodiments, the first and second thresholds are selected to have values ​​that are predetermined and constant for a particular experiment (e.g., for a particular evaluation of a particular database of test compounds against a particular target polymer) and that prove useful in identifying suitable test compounds from a database of test compounds for activity against the test polymer. For example, in some embodiments, any of the thresholds disclosed herein are designed to identify 0.1 percent or less, 0.5 percent or less, 1 percent or less, 2 percent or less, 5 percent or less, 10 percent or less, 20 percent or less, or 50 percent or less of a database of test compounds as active against a target polymer, and the database of test compounds may be greater than 100 compounds, greater than 1000 compounds, greater than 10,000 compounds, greater than 100,000 compounds, greater than 1×10 6 Compounds, 10x10 6 The compound may include one or more compounds.

[0108] In alternative embodiments, the system 100 or one or more programs hosted, stored, or addressable by the system 100 can characterize the interaction between the test compound and the target polymer 38 as an activity on a continuous scale. That is, the system 100 or one or more programs hosted, stored, or addressable by the system 100 provide a number on a continuous scale indicating the activity of the test compound against the target polymer. The continuous scale activity value is useful, for example, to compare the activity of each test compound in the database of test compounds against the target polymer assigned by the trained spatial data evaluation module 36. As an example, referring to block 260 of FIG. 2E, in some embodiments, the first model 89 is a regression task and the first interaction score quantifies the interaction between the test compound and the target polymer.

[0109] The disclosed systems and methods are not limited to characterizing the interaction between a test compound and a target polymer 38 as a continuous or discrete scale activity. In alternative embodiments, the system 100 or one or more programs hosted, stored, or addressable by the system 100 may characterize the interaction between a test compound and a target polymer as an IC of the test compound on the target polymer. 50 , E.C. 50 , Kd, ​​KI, or pKI on a continuous or discrete (categorical) scale.

[0110] Although a binary discrete scale and a discrete scale having three possible outcomes have been identified, the present disclosure is not limited to these two examples of discrete scales for characterizing the interaction between a test compound and a target polymer 38. Indeed, any discrete scale can be used for characterizing the interaction between a test compound and a target polymer 38, including, as non-limiting examples, discrete scales having 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 different outcomes.

[0111] Referring to block 264 of FIG. 2E, in some embodiments, the first interaction score represents a binding coefficient of the test compound to the target polymer.

[0112] Referring to block 266 of FIG. 2E, in some embodiments, the first interaction score is the IC 50 , E.C. 50 , Kd, ​​KI, or pKI. IC 50 , E.C. 50 Measured binding coefficients such as Kd, KI, and pKI are generally described in Huser ed., 2006, High-Throughput-Screening in Drug Discovery, Methods and Principles in Medicinal Chemistry35 and Chen ed., 2019, A Practical Guide to Assay Development and High-Throughput Screening in Drug Discovery, each of which is incorporated herein by reference.

[0113] Referring to block 270 of FIG. 2F, in some embodiments, the first interaction score is a binary score, and a first value of the binary score is an IC of the test compound for the target polymer that is above a first threshold. 50 , E.C. 50、 The second value of the binary score represents the IC of the test compound against the target polymer below the first threshold, representing Kd, KI, or pKI. 50 , E.C. 50 , Kd, ​​KI, or pKI.

[0114] Referring to block 272 of FIG. 2F, in some embodiments, the first interaction score represents an in silico quality score of the test compound relative to the target polymer.

[0115] Referring to block 274 of FIG. 2F, in some embodiments, the first model 89 is a fully connected second neural network, also known as a multilayer perceptron (MLP). In some embodiments, the MLP is a type of feedforward artificial neural network (ANN) that includes at least three layers: an input layer, a hidden layer, and an output layer of nodes. In such embodiments, except for the input nodes, each node is a neuron that uses a nonlinear activation function. Further disclosure of suitable MLPs that function as the first model 89 in some embodiments of the present disclosure can be found in Vang-mata ed., 2020, Multilayer Perceptrons: Theory and Applications, Nova Science Publishers, Hauppauge, New York.

[0116] 2F, block 276, in some embodiments, the pooled embeddings 85 are also input into a second model 93, thereby obtaining a second interaction score for the interaction between the test compound and the target polymer. In such embodiments, the first model 89 is a first fully-connected neural network and the second model 93 is a second fully-connected neural network, where the first interaction score represents an in silico quality score of the test compound relative to the target polymer and the second interaction score represents an in silico quality score of the test compound relative to the target polymer.

[0117] 10 illustrates a system for characterizing the interaction between a test compound and a target polymer, according to one embodiment of the present disclosure, in which the characterization is (i) binary discrete activity and (ii) pKi, and the system is trained using training compounds with known activity and pKi. In this training, a multitask loss function is calculated and the loss function is calculated for all tasks (in the case of FIG. 10, binary discrete activity and pKi) and N b is minimized over mini-batches of training compounds,

number

number

[0118] In the system of Figure 10, the pKi model and the activity model are independent of each other. In some embodiments, the pKi model is trained as a regression task using a loss function such as mean squared error on the pKi values ​​of the training compounds, while the activity model is trained as a classification task using a loss function such as binary cost entropy on the known binary discrete activity values ​​of the training compounds.

[0119] As shown in Figure 10, in some embodiments, the pooled embeddings 85 are input to both a first model 89 (to provide a characterization of the interaction between the test compound and the target polymer in the form of a calculated pKi value) as well as a second model 93 (to provide a characterization of the interaction between the test compound and the target polymer in the form of the activity of the test compound against the target polymer 38). Thus, in the embodiment illustrated in Figure 10, the characterization of the interaction between the test compound and the target polymer is both a pkI score (e.g., a discrete binary score or a scalar score) and an activity score (e.g., a classification as a "good binder", "bad binder", etc.). While the second model 89 calculates a pKi in the embodiment shown in Figure 10, in other embodiments having the topology of Figure 10, the second model calculates an IC instead of a pKi. 50 , E.C. 50 , Kd, ​​or KI.

[0120] 11 is a system for characterizing an interaction between a test compound and a target polymer according to one embodiment of the present disclosure, where the characterization is pKi, where the pKi is conditioned, in part, on activity, and the system is trained using known pKi and activity of the training compounds. In the system of FIG. 11, the pKi model is conditioned on the activity model. In some embodiments, the pKi model is trained as a regression task using a loss function such as mean squared error on the pKi values ​​of the training compounds, while the activity model is trained as a classification task using a loss function such as binary cost entropy on the activity values ​​of the training compounds.

[0121] As shown in FIG. 11, in some embodiments, the pooled embeddings 85 are input to both the first model 89 (through edge 1102) as well as the second model 93 (through edge 1104). Additionally, the output of the second model 93, which is a calculation of the activity of the test compound against the target polymer, is input to the first model 89 via edge 1106. In some embodiments, this characterization provided by the second model 93 is an activity score for the test compound. In some embodiments, this activity score is a discrete binary score, e.g., where a "1" indicates that the test compound is active against the target polymer and a "0" indicates that the test compound is inactive against the target polymer. In some embodiments, the activity score provided by the second model 93 is a scalar. Thus, the first model 89 receives both the output of the second model 93 and the pooled embeddings 85. The first model 89 uses both of these inputs to determine a characterization of the interaction between the test compound and the target polymer (e.g., in the form of the pKi of the test compound against the target polymer 37). The conditioning of the pKi calculation of the first model 89 on both the pooled embedding 85 and the second model 93 serves to improve the performance of the first model 89 in characterizing the test compound. The second model 89 calculates a pKi in the embodiment shown in FIG. 11, but in other embodiments having the topology of FIG. 11, the second model calculates an IC instead of a pKi. 50 , E.C. 50 , Kd, ​​or KI.

[0122] As shown in FIG. 12, in some embodiments, the pooled (shared) embeddings 85 are input to both the first model 89 (via edge 1202) and the second model 93 (via edge 1204). Additionally, the output of the first model 89, which is a calculation of the pKi of the test compound against the target polymer, is input to the second model 93 via edge 1206. Thus, the second model 93 receives both the output of the first model 89 and the pooled embeddings 85. The second model 93 uses both of these inputs to determine a characterization of the interaction between the test compound and the target polymer (e.g., in the form of an activity score for the test compound). In some embodiments, this activity score is a discrete binary score, e.g., where a "1" indicates that the test compound is active against the target polymer and a "0" indicates that the test compound is inactive against the target polymer. In some embodiments, the activity score provided by the second model 93 is a scalar. The conditioning of the activity scores of the second model 93 on both the pooled embeddings 85 and the output of the first model 93 serves to improve the performance of the second model 83 in characterizing test compounds. The first model 89 calculates a pKi in the embodiment shown in Figure 12, but in other embodiments having the topology of Figure 12, the first model 89 calculates an IC instead of a pKi. 50 , E.C. 50 , Kd, ​​or KI.

[0123] Referring to block 278 of Figure 2F, in some such embodiments, the first interaction score and the second interaction score are input into a third model 97 to obtain a third interaction score. In some such embodiments, the third model 97 is a third fully connected neural network. Referring to block 280 of Figure 2F, in some such embodiments, the third interaction score is a discrete binary activity score having a first value when the test compound is determined to be inactive by the third model 97 and a second value when the test compound is determined to be active by the third model 97. For example, Figure 13 shows a system for characterizing the interaction between a test compound and a target polymer according to one embodiment of the present disclosure, where the characterization is activity (through activity model 97), where activity is conditioned, in part, on both pKi (through pKi model 89) and binding mode score (through PoseRanker model 93), where the pKi model is trained using known pKi values ​​for the training compounds, and the PoseRanker model is trained using binding mode scores for the training compounds. For purposes of training the activity model 97 of Figure 13, in some embodiments, each training compound is labeled as active if its pKi (or IC50) is less than 10 μM, and is labeled as inactive otherwise.For purposes of training a PoseRanker model, in some embodiments, the binding mode scores of training compounds are determined by docking the training compounds with a docking program such as CUina [Morrison et al., 2020, “CUina: An Efficient GPU Implementation of AutoDock Vina. August 2020. URL https: / / blog.atomwise.com / efficient-gpu-implementation-of-autodock-vina.] and then analyzing the binding mode scores of the training compounds with the PoseRanker model [Stafford et al., 2021, “AtomNet PoseRanker: Enriching Ligand Pose Quality for Dynamic Proteins in Virtual High Throughput Screens,” doi:10.33774 / chemrxiv-2021-t6xkj. URL https: / / chemrxiv.org / engage / chemrxiv / article-details / 614b905e39ef6a1c36268003. Thus, in such an embodiment, the binding mode scores used to train the PoseRanker model 93 are the PoseRanker rankings.

[0124] In the system of FIG. 13, the activity model 97 is conditioned against both the pKi model 89 and the PoseRanker model 93. In some embodiments, the pKi model and the PoseRanker model (Stafford et al., 2022, “AtomNet PoseRanker: Enriching Ligand Pose Quality for Dynamic Proteins in Virtual High-Throughput Screens,” Journal of Chemical Information and Modeling 62, pp. 1178-1189, which is incorporated herein by reference) are trained as a regression task using a loss function such as mean squared error, while the activity model is trained as a classification task using a loss function such as binary cost entropy.

[0125] Thus, as shown in Figure 13, in some embodiments, the pooled (shared) embeddings 85 are input to both a first model 89 (through edge 1302) as well as a second model 93 (through edge 1304). Additionally, the output of the first model 89, which is a calculation of the pKi of the test compound against the target polymer, is input to a third model 97 through edge 1306. Additionally, the output of the second model 93, which is a calculation of the pose quality of the test compound against the target polymer represented by the pooled embeddings 85, is also input to the third model 97 through edge 1308. In Figure 13, the second model is referred to as the PoseRanker model. Throughout this disclosure, the terms "PoseNet" and "PoseRanker" are used interchangeably. The PoleRanker (PoseNet) model is described in more detail in Stafford et al., 2022, “AtomNet PoseRanker: Enriching Ligand Pose Quality for Dynamic Proteins in Virtual High-Throughput Screens,” Journal of Chemical Information and Modeling, Volume 62, pp. 1178-1189, which is incorporated herein by reference. Additionally, the pooled embeddings 85 are input to a third model. Thus, the third model 97 receives the output of the first model 89, the output of the second model 93, and the pooled embeddings 85. The third model 97 uses all of these inputs to determine a characterization of the interaction between the test compound and the target polymer (e.g., in the form of an activity score for the test compound). In some embodiments, the activity score is a discrete binary score, e.g., where “1” indicates that the test compound is active against the target polymer and “0” indicates that the test compound is inactive against the target polymer. The first model 89 calculates pK in the embodiment shown in FIG. 13, but in other embodiments having the topology of FIG. 13, the first model 89 calculates IC instead of pK. 50 , E.C. 50 , Kd, ​​or KI.

[0126] 14 is a system for characterizing interactions between a test compound and a target polymer according to one embodiment of the present disclosure, where the characterization is activity and two different compound binding mode scores, and the system is trained using training compounds with known activity scores. In the system of FIG. 14, the activity model is conditioned against a pose quality score model. In some embodiments, the pose quality model is trained as a regression task using a loss function such as mean squared error, while the activity model is trained as a classification task using a loss function such as binary cost entropy.

[0127] 15 illustrates a system for characterizing an interaction between a test compound and a target polymer, according to one embodiment of the present disclosure, where the characterizations are activity, two different compound binding mode scores, and pKi, and the system is trained using combined positive and negative poses for the training compounds. In the system of FIG. 15, the activity model is conditioned against the pose quality score model. In some embodiments, the pose quality model is trained as a regression task using a loss function such as mean squared error, while the activity model is trained as a classification task using a loss function such as binary cost entropy.

[0128] In the embodiment shown in FIG. 16A, a pose of a test compound, for example in the form of a set of atomic coordinates 49 or vectors 54, is introduced into a first neural network 72 to finally result in a pooled embedding 85 according to FIG. 9A. This pooled embedding 85 is input to a first model 89 (via edge 1630), a second model 93 (via edge 1610), and a third model 97 (via edge 1620). Furthermore, the output of the second model 93 of the test compound (which is a calculation of an interaction score, such as a pose quality score) is input to the first model 89 through edge 1640. Furthermore, the output of the third model 97 of the test compound (which is a calculation of an interaction score, such as a pKi) is input to the first model 89 through edge 1650. Thus, the first model receives the outputs of the third model, the second model, and the pooled embedding 84. The first model 89 uses each of these inputs to collectively determine a characterization of the interaction between the test compound and the target polymer. In some embodiments, the characterization is an activity score for the test compound. In some embodiments, the activity score is a discrete binary score, e.g., where "1" indicates that the test compound is active against the target polymer and "0" indicates that the test compound is inactive against the target polymer. In some embodiments, the activity score provided by the first model 89 is a scalar.

[0129] In some embodiments, referring to FIG. 16A, the pooled embeddings 85 are used to predict three outputs: activity (through the first model 89), CUina pose quality score (through the second model 93), and pKi score (through the third model 97). For a description of the fully connected CUina pose quality model, see Morrison et al. “Efficient GPU Implementation of AutoDock Vina,” COMP poster 3432389, which is incorporated herein by reference. This is performed in two stages in the embodiment shown in FIG. 16A. First, the CUina and pKi score predictions are calculated by passing the pooled embeddings 85 through the second model 93 and the third model 97. Second, the conditioned embeddings 1690 are formed by concatenating (i) the pooled embeddings 85, (ii) the resulting second model 93 score predictions from the first stage, and (iii) the third model 97 score predictions from the first stage. This embedding 1690 is then passed to a first model 89, in the form of a multi-layer perceptron, to calculate an activity prediction for the test compound. In some embodiments, rather than simply concatenating (i) the pooled embedding 85, (ii) the resulting second model 93 score prediction from the first stage, and (iii) the third model 97 score prediction from the first stage, the embedding 1690 represents a multiplication of the three components against each other, or some other mathematical combination of these three components. In such embodiments, the product of the multiplication of the three components, or some other mathematical combination of the three components, is input to the third model as the embedding 1690. In some embodiments, rather than concatenating the three components, the embedding 1690 transforms each of the three sources, and this transformation serves as an input to the first model 89. More generally, the embedding 1690 may perform any mathematical function on all or any portion of any of the inputs to the embedding 1690, including, but not limited to, multiplication, concatenation, linear or non-linear transformations, to form a conditioning embedding that is passed on to the first model 72.The third model 97 estimates the pKi in the embodiment illustrated in Figure 16A, however, in other embodiments having the topology of Figure 16A, the third model 97 estimates the IC of the test compound against the target polymer instead of the pKi. 50 , E.C. 50 It is understood that the present invention estimates the Kd, Kd, ​​or K I.

[0130] 16B, it is possible to condition the first model 89 against additional models as well. Thus, in FIG. 16B, the first model 89 is conditioned against the pooled embeddings 85 as well as the output of a second model 93 providing a CUina score for the test compound against the target polymer, a third model 97 providing a pKi score for the test compound against the target polymer, and a fourth model 990 providing a PoseRanker score for the test compound against the target polymer. While the third model 97 estimates pKi in the embodiment illustrated in FIG. 16B, in other embodiments having the topology of FIG. 16B, the third model 97 estimates the IC of the test compound against the target polymer instead of the pKi. 50 , E.C. 50, Kd, ​​or KI. Referring to FIG. 16B, in some such embodiments, the first model, the second model, the third model, and the fourth model 990 are each fully connected neural networks. Such fully connected neural networks are also known as multilayer perceptrons (MLPs). In some embodiments, the MLP is a type of feed-forward artificial neural network (ANN) that includes at least three layers: input, hidden, and output layers of nodes. In such embodiments, except for the input nodes, each node is a neuron that uses a nonlinear activation function. In some embodiments, further disclosure regarding suitable MLPs that function as the first model 72 can be found in Vang-mata ed., 2020, Multilayer Perceptrons: Theory and Applications, Nova Science Publishers, Hauppauge, New York, which is incorporated herein by reference. Referring to FIG. 16B, in some embodiments, the corresponding activation scores provided by the first model 89 are binary activity scores. For example, in some embodiments, an activity score having a value of "1" means that the test compound is "active" in inhibiting the activity or function (e.g., enzymatic activity) of the target polymer, and an activity score of "0" means that the test compound does not inhibit the activity or function of the target polymer.

[0131] Representative test compounds and training compounds. Several different architectures or systems have been described for characterizing the interaction between a test compound and a target polymer. Before each such architecture or system can be used to characterize the interaction between a test compound and a target polymer, it is trained on a training compound. The significant difference between the test compounds and the training compounds is that the training compounds are labeled (e.g., with complementary binding data to the target polymer obtained from a wet lab binding assay or the like) and such labels are used to train the first neural network 72, the attention mechanism 77, the first model 89, the second model 93, and any third and subsequent models, whereas each test compound is unlabeled or no labels are used and the first neural network 72, the attention mechanism 77, the first model 89, the second model 93, and any third and subsequent models of the present disclosure are used to characterize the interaction between each test compound and the target polymer. In other words, the training compounds have already been characterized by labeling (characterization of the interaction between the training compounds and the target polymer), and such characterization is used to train the models of the present disclosure so that they can characterize the interaction between the test compounds and the target polymer. The interaction between the test compounds and the target polymer is typically not characterized prior to application of the first neural network 72 and other models of the present disclosure. In a typical embodiment, the characterization of the interaction between the training compounds and the target polymer available is binding data for the target polymer 38 obtained by a wet lab binding assay.

[0132] Training a predictive model. In some embodiments, a predictive model according to the present disclosure, such as the predictive models collectively depicted in FIG. 9A, is trained to receive compound geometric data input (e.g., compound poses 48) and output a characterization of the interaction between the compound and a target polymer. For example, in some embodiments, each of a plurality of poses for each of a plurality of training compounds (e.g., 50 or more training compounds, 100 or more training compounds, 100,000 or more training compounds) having known binding data to the target polymer are run sequentially through the model depicted in FIG. 9A, and the model provides a single value for each respective training compound.

[0133] In some such embodiments, the system of the present disclosure (e.g., the system illustrated in FIG. 9A) outputs one of two possible activity classes for each training compound for a given target polymer. For example, the single value provided by the system of the present disclosure for each respective training compound is in a first activity class (e.g., binder) if its number is below a predetermined threshold, and in a second activity class (e.g., non-binder) if its number is above a predetermined threshold. The activity class assigned by the system of the present disclosure is compared to the actual activity class as represented by the training compound binding data. In a typical non-limiting embodiment, such training compound binding data is from an independent web lab binding assay. The error in the activity class assignment made by the system of the present disclosure, as verified against the binding data, is then back-propagated through the parameters of each of the models of the system of the present disclosure (e.g., the first neural network 72, the attention mechanism 77, and / or the first model 89, etc.) to train the system. In an exemplary embodiment, the models of the present disclosure are trained by stochastic gradient descent with the AdaDelta adaptive learning method (Zeiler, 2012, “ADADELTA: an adaptive learning rate method,” 'CoRR, vol. abs / 1212.5701, incorporated herein by reference) taking into account the combined data for errors in the activation class assignments made by the model, and the back-propagation algorithm as presented in Rumelhart et al., 1988, “Neurocomputing: Foundations of research,” ch. Learning Representations by Back-propagating Errors, pp. 696-699, Cambridge, MA, USA: MIT Press, incorporated herein by reference.In some such embodiments, the two possible activity classes are each determined by a binding constant greater than a given threshold amount (e.g., IC of the training compound for the target polymer greater than 1 nanomolar, 10 nanomolar, 100 nanomolar, 1 micromolar, 10 micromolar, 100 micromolar, or 1 millimolar). 50 , E.C. 50 , or KI) and a binding constant below a given threshold amount (e.g., IC 50 , E.C. 50 , or KI).

[0134] In some embodiments, the system of the present disclosure outputs one of a plurality of possible activity classes (e.g., three or more activity classes, four or more activity classes, five or more activity classes) for each training compound for a given target polymer. For example, the single value provided by the system of the present disclosure for each respective training compound is in a first activity class if its number falls in a first range, in a second activity class if its number falls in a second range, in a third activity class if its number falls in a third range, etc. The activity class assigned by the system of the present disclosure is compared to the actual activity class as represented by the training compound binding data of other forms of training data. The error of the activity class assignment made by the system of the present disclosure is used to train the system of the present disclosure using the techniques discussed above, such that it is validated against the binding data (or other forms of measured or independently calculated data). In some embodiments, each respective classification in the plurality of classifications is determined by the IC of the training compound for the target polymer. 50 , E.C. 50 , pkA, or KI range.

[0135] In some embodiments, the classification of the plurality of training compounds by the system of the present disclosure is compared against training data (e.g., binding data or other independently measured data for the training compounds) using non-parametric techniques. For example, the system of the present disclosure is used to rank the plurality of training compounds for a given property (e.g., binding to a given target polymer), and this rank order is compared against a rank order provided by training data obtained by wet lab binding assays of the plurality of training compounds. This gives rise to the ability to train the system of the present disclosure against the error of the calculated rank order using the system error correction techniques discussed above. In some embodiments, the error (difference) between the ranking by the training compounds by the system of the present disclosure and the ranking of the training compounds determined by the binding data (or other independently measured data for the training compounds) is calculated using the Wilcoxon Mann Whitney function (Wilcoxon signed rank test) or other non-parametric test, and this error is used to further train the system of the present disclosure (e.g., the first neural network 72, the attention mechanism 77, and / or the first model 89, etc.).

[0136] In some embodiments, model training may include modifying parameters of one or more component models, which may be further constrained using various forms of regularization, such as L1, L2, weight decay, and dropout.

[0137] In one embodiment, any of the models disclosed herein may optionally have their parameters (e.g., weights) tuned (adjusted to potentially minimize the error between the predicted binding affinity and / or classification of the system and the reported binding affinity and / or classification of the training data) if the training data is labeled (e.g., with binding data). To minimize the error function, various methods can be used, such as gradient descent, which may include, but are not limited to, logarithmic loss, sum of squares error, hinge loss, etc. These methods may include second-order methods or approximations, such as momentum, Hessian-free estimation, Nesterov's accelerated gradient, Adagrad, etc. Unlabeled developmental pre-training and labeled discriminative training may be combined.

[0138] A convolutional neural network is used as the first neural network 72. To characterize the interaction between the test compound and the target polymer, in some embodiments, a voxel map 52 is created for each pose 48 of the compound. In some embodiments, the voxel map 52 is created by (i) sampling the compound and the target polymer 38 in the pose 48 on a three-dimensional grid basis, thereby forming a corresponding three-dimensional uniform space-filling honeycomb with a corresponding plurality of space-filling (three-dimensional) polyhedral cells, and (ii) for each respective three-dimensional polyhedral cell in the corresponding plurality of three-dimensional cells, populating voxels (a discrete set of regularly spaced polyhedral cells) in the respective voxel map based on the properties (e.g., chemical properties) of the respective three-dimensional polyhedral cell. Thus, for a particular pose 48 of a particular compound, a corresponding voxel map 52 is created. Examples of space-filling honeycombs include cubic honeycombs with parallelepiped cells, hexagonal prism honeycombs with hexagonal prism cells, rhombic dodecahedrons with rhombic dodecahedral cells, elongated dodecahedrons with elongated dodecahedral cells, and truncated octahedrons with truncated octahedral cells.

[0139] In some embodiments, the space-filling honeycomb is a cubic honeycomb with cubic cells, the dimensions of such voxels determining their respective resolution. For example, a resolution of 1 Å may be selected, meaning that each voxel represents a corresponding cube of geometric data with 1 Å dimensions in such embodiments (e.g., 1 Å×1 Å×1 Å in the respective height, width, and depth of each cell). However, in some embodiments, a finer grid spacing (e.g., 0.1 Å, or even 0.01 Å) or a coarser grid spacing (e.g., 4 Å) is used, which spacing results in an integer number of voxels to cover the input geometric data. In some embodiments, sampling occurs at a resolution of 0.1 Å to 10 Å. By way of example, for a 40 Å input cube, with a resolution of 1 Å, such an arrangement would result in 40×40×40=64,000 input voxels.

[0140] In some embodiments, the atomic features generated in sampling (i) are arranged in a single voxel of the respective voxel map, with each voxel of the plurality of voxels representing at most one atomic feature. In some embodiments, the atomic features consist of an enumeration of atom types. As an example, some embodiments of the disclosed system and method are configured to represent the presence of all atoms in a given voxel of the voxel map 52 as different numbers of its entries, e.g., if carbon is present in the voxel, a value of 6 is assigned to the voxel since carbon has atomic number 6. However, such encoding may imply that atoms with close atomic numbers behave similarly, which may not be particularly useful in some applications. Furthermore, the behavior of elements may be more similar within a group (a column on the periodic table), and thus such encoding imposes additional work for the convolutional neural network 24 to decode.

[0141] In some embodiments, atomic features are encoded in the voxels as binary categorical variables. In such embodiments, atom types are encoded in what is called "one-hot" encoding, where every atom type has a separate channel. Thus, in such embodiments, each voxel has multiple channels, where at least a subset of the multiple channels represent atom types. For example, one channel in each voxel may represent carbon, while another channel in each voxel may represent oxygen. If the three-dimensional grid element corresponding to a given voxel contains a given atom type, then the channel for that atom type in the given voxel is assigned a first value of the binary categorical variable, such as "1," and if the three-dimensional grid element corresponding to the given voxel does not contain that atom type, then the channel for that atom type in the given voxel is assigned a second value of the binary categorical variable, such as "0," in the given voxel.

[0142] There are over 100 elements, but most are not encountered in biology. However, representing even the most common biological elements (e.g., H, C, N, O, F, P, S, Cl, Br, I, Li, Na, Mg, K, Ca, Mn, Fe, Co, Zn) can result in 18 channels per voxel, or 10,483 x 18 = 188,694 inputs to the acceptor field. Thus, in some embodiments, each respective voxel in the voxel map includes multiple channels, each channel within the multiple channels representing a different property that may occur in the three-dimensional space-filling polyhedral cell corresponding to the respective voxel. The number of possible channels for a given voxel is even higher in embodiments where additional features of the atom (e.g., partial charges, presence of ligands for protein targets, electronegativity, or SYBYL atom type) are further presented as independent channels for each voxel, requiring more input channels to distinguish between otherwise equivalent atoms.

[0143] In some embodiments, each voxel has 5 or more input channels. In some embodiments, each voxel has 15 or more input channels. In some embodiments, each voxel has 20 or more input channels, 25 or more input channels, 30 or more input channels, 50 or more input channels, or 100 or more input channels. In some embodiments, each voxel has 5 or more input channels selected from the descriptors included in Table 1 below. For example, in some embodiments, each voxel has 5 or more channels, each encoded as a binary categorical variable, each such channel representing a SYBYL atom type selected from Table 1 below. For example, in some embodiments, each respective voxel in the voxel map includes a channel for a C.3 (sp3 carbon) atom type, meaning that if the grid in space of a given test object-target object (or training object-target object) complex represented by the respective voxel contains sp3 carbon, the channel adopts a first value (e.g., "1"), and otherwise is a second value (e.g., "0"). [Table 1-1] [Table 1-2]

[0144] In some embodiments, each voxel includes 10 or more input channels, 15 or more input channels, or 20 or more input channels selected from the descriptors included in Table 1 above. In some embodiments, each voxel includes a channel for a halogen.

[0145] In some embodiments, a structural protein-ligand interaction fingerprint (SPLIF) score is generated for each compound pose 48. In such embodiments, the SPLIF scores are used as additional inputs to the underlying first neural network 72 or are individually encoded in the voxel map. For a description of SPLIF, see Da and Kireev, 2014, J. Chem. Inf. Model. 54, pp. 2555-2561, "Structural Protein-Ligand Interaction Fingerprints (SPLIF) for Structure-Based Virtual Screening: Method and Benchmark Study", which is incorporated herein by reference in its entirety. SPLIF implicitly encodes all possible interaction types that may occur between interacting fragments of a compound (test compound or training compound) and a target polymer 38 (e.g., π-π, CH-π, etc.). In a first step, the compound (test compound or training compound)-target polymer 38 are examined for intermolecular contacts. The two atoms are within a specified threshold distance between them (e.g., 4.5 Å). For each such intermolecular atom pair, the respective atom of the compound and the target polymer atom are expanded into a circular fragment (e.g., a fragment containing the atom in question and its successive neighbors up to a certain distance). Each type of circular fragment is assigned an identifier. In some embodiments, such identifiers are encoded in individual channels in each voxel. In some embodiments, an extended connectivity fingerprint to the first nearest neighbor (ECFP2) can be used, as defined in Pipeline Pilot software. See Pipeline Pilot, ver. 8.5, Accelrys Software Inc., 2009, which is incorporated herein by reference. ECFP holds information about all atom / bond types and uses one unique integer identifier to represent one substructure (e.g., circular fragment). The SPLIF fingerprint encodes all the circular fragment identifiers it contains.In some embodiments, the SPLIF fingerprints are not encoded into individual voxels, but serve as separate and independent inputs in the neural networks described below.

[0146] In some embodiments, instead of or in addition to SPLIFs, a Structural Interaction Fingerprint (SIFt) is calculated for each pose of a given compound (test or training compound) relative to the target polymer and provided separately as an input to the first neural network 72 or encoded within the voxel map 52. For the calculation of SIFt, see Deng et al., 2003, "Structural Interaction Fingerprint (SIFt): A Novel Method for Analyzing Three-Dimensional Protein-Ligand Binding Interactions," J. Med. Chem. 47(2), pp. 337-344, which is incorporated herein by reference.

[0147] In some embodiments, instead of or in addition to SPLIF and SIFT, an atom pair-based interaction fragment (APIF) is calculated for each pose of a given compound (test compound or training compound) with respect to the target polymer 38 and provided independently as an input to the first neural network or individually encoded in the voxel map. For calculation of APIF, see Perez-Nueno et al., 2009, "APIF: a new interaction fingerprint based on atom pairs and its application to virtual screening," J. Chem. Inf. Model. 49(5), pp. 1245-1260, which is incorporated herein by reference.

[0148] The data representation may be encoded in a manner that allows for expression of various structural relationships associated with the molecule / protein, for example. The geometric representation may be implemented in various ways and topographies according to various embodiments. The geometric representation is used for data visualization and analysis. For example, in one embodiment, the geometry may be represented using voxels laid out on various topographies, such as 2-D, 3-D Cartesian / Euclidean space, 3-D non-Euclidean space, manifolds, etc. For example, FIG. 4 shows a sample three-dimensional grid structure 400 including a series of subcontainers, according to one embodiment. Each subcontainer 402 may correspond to a voxel. A coordinate system may be defined for the grid such that each subcontainer has an identifier. In some embodiments of the disclosed systems and methods, the coordinate system is a Cartesian system in 3-D space, but in other embodiments of the system, the coordinate system may be any other type of coordinate system, such as an oblate spheroid, a cylindrical or spherical coordinate system, a polar coordinate system, other coordinate systems designed for various manifolds and vector spaces, among others. In some embodiments, the voxels may have a particular value associated with each, which may be represented, for example, by applying a label and / or determining the respective positioning, among other things.

[0149] In some embodiments, the first neural network 72 is a convolutional neural network that requires a fixed input size. In some such embodiments of the disclosed systems and methods, the geometric data (e.g., the voxel map 52 and / or the atomic coordinate set 49) is cropped to fit within an appropriate bounding box. For example, a cube of 25-40 Å to a side may be used. In some embodiments where a compound (test compound or training compound) is docked to the active site of the target polymer 38, the center of the active site serves as the center of the cube. In some such embodiments, a square cube of fixed dimensions centered on the active site of the target polymer 38 is used to divide the space into a voxel grid, although the disclosed systems are not so limited. In some embodiments, any of a variety of shapes are used to divide the space into a voxel grid. In some embodiments, a polyhedron, such as a rectangular prism, a polyhedral shape, etc., is used to divide the space. In one embodiment, the grid structure may be configured to resemble an arrangement of voxels. For example, each substructure may be associated with a channel for each atom being analyzed. Also, a coding method may be provided for representing each atom numerically.

[0150] In some embodiments, the voxel map takes into account the factor of time (e.g., along with a molecular dynamics run of a compound docked to a target polymer) and therefore may be four-dimensional (X, Y, Z, and time).

[0151] In some embodiments, other implementations such as pixels, points, polygons, polyhedra, or any other type of shape of multiple dimensions (e.g., 3D, 4D, etc. shapes) may be used instead of voxels.

[0152] In some embodiments, the geometric data is normalized by selecting the origin of the X, Y, and Z coordinates to be the center of mass of the binding site of the target polymer 38, as determined by a cavity flooding algorithm. For representative details of such algorithms, see Ho and Marshall, 1990, "Cavity search: An algorithm for the isolation and display of cavity-like binding regions," Journal of Computer-Aided Molecular Design 4, pp. 337-354, and Hendlich et al., 1997, "Ligsite: automatic and efficient detection of potential small molecule-binding sites in proteins," J. Mol. Graph. Model 15:6, each of which is incorporated herein by reference in its entirety. Alternatively, in some embodiments, the origin of the voxel map is centered on the center of mass of the entire co-complex (of the compounds docked in their respective poses relative to the target polymer). In some embodiments, the origin of the voxel map is centered on the center of mass of the compound (test compound or training compound). In some embodiments, the origin of the voxel map is centered on the center of mass of the target polymer 38. The basis vectors can be optionally selected to be the principal moments of inertia of the entire co-complex, the target polymer only, or the compound (test compound or training compound) only. In some embodiments, the target polymer 38 has an active site, and sampling samples the compound (test compound or training compound) at poses within the active site of the target polymer on a three-dimensional grid base where the center of mass of the active site is taken as the origin, and a corresponding three-dimensional uniform honeycomb for sampling represents a portion of the polymer and compound (test compound or training compound) centered on the center of mass. In some embodiments, the uniform honeycomb is a regular cubic honeycomb, and the target polymer and the portion of the compound (test compound or training compound) are cubes of predetermined fixed dimensions.In such embodiments, the use of a cube of predetermined fixed dimensions ensures that relevant portions of the geometric data are used to ensure that each voxel map is the same size. In some embodiments, the predetermined fixed dimensions of the cube are NÅ×NÅ×NÅ, where N is an integer or real number between 5 and 100, an integer between 8 and 50, or an integer between 15 and 40. In some embodiments, the uniform honeycomb is a prismatic honeycomb, and the target polymer and a portion of the compound (test compound or training compound) are aligned along the predetermined fixed dimensions of the prismatic QÅ×RÅ×SÅ, where Q is a first integer between 5 and 100, R is a second integer between 5 and 100, and S is a third integer or real value between 5 and 100, and at least one number in the set {Q,R,S} is not equal to another value in the set {Q,R,S}.

[0153] In one embodiment, every voxel has one or more input channels, which may have various values ​​associated with them, which in a simple implementation may be on / off, and may be configured to code for a certain type of atom. The atom type may indicate the element of the atom, or the atom type may be further refined to distinguish characteristics of other atoms. The atoms present may then be coded at each voxel. The various types of coding may be utilized using various techniques and / or methodologies. As an exemplary coding method, the atomic number of the atom may be utilized, resulting in one value per voxel from 1 for hydrogen to 118 for ununoctium (or any other element).

[0154] However, as discussed above, other encoding methods may be utilized, such as "one-hot encoding," where each voxel has many parallel input channels, each of which is either on or off and encodes for a certain type of atom. The atom types may indicate the elemental character of the atom, or the atom types may be further refined to distinguish other atomic characteristics. For example, the SYBYL atom type distinguishes a singly bonded carbon from a doubly bonded carbon, a triple bonded carbon, or an aromatic carbon. For a description of the SYBYL atom types, see Clark et al., 1989, "Validation of the General Purpose Tripos Force Field, 1989, J. Comput. Chem. 10, pp. 982-1012, which is incorporated herein by reference.

[0155] In some embodiments, each voxel further comprises one or more channels for distinguishing between atoms that are part of the target polymer 38 or cofactor and atoms that are part of a compound (test compound or training compound). For example, in one embodiment, each voxel further comprises a first channel for the target polymer 38 and a second channel for the compound (test compound or training compound). If an atom in the portion of the space represented by the voxel is from the target polymer 38, the first channel is set to a value such as "1", otherwise it is zero (e.g., because the portion of the space represented by the voxel does not contain any atoms or contains one or more atoms from the compound). Furthermore, if an atom in the portion of the space represented by the voxel is from a compound (test compound or training compound), the second channel is set to a value such as "1", otherwise it is zero (e.g., because the portion of the space represented by the voxel does not contain any atoms or contains one or more atoms from the target polymer 38). Similarly, other channels may additionally (or alternatively) specify further information such as partial charge, polarizability, electronegativity, solvent accessible space, and electron density. For example, in some embodiments, an electron density map of the target polymer is overlaid with the three-dimensional coordinate set, and the creation of a voxel map further samples the electron density map. Examples of suitable electron density maps include, but are not limited to, multiple isomorphous replacement maps, single isomorphous replacement with anomalous signal maps, single wavelength anomalous dispersion maps, multi-wavelength anomalous dispersion maps, and 2Fo-Fc maps (260). See McRee, 1993, Practical Protein Crystallography, Academic Press, which is incorporated herein by reference.

[0156] In some embodiments, voxel encoding according to the disclosed systems and methods may include additional optional encoding refinements. The following two are provided as examples:

[0157] In a first encoding refinement, memory requirements can be reduced by reducing the set of atoms represented by a voxel (e.g., by reducing the number of channels represented by a voxel) based on the fact that most elements occur rarely in biological systems. Atoms can be mapped to share the same channel in a voxel either by combining rare atoms (thus, may have little impact on the performance of the system) or by combining atoms with similar properties (thus, may minimize inaccuracies from combinations). In some embodiments, 2, 3, 4, 5, 6, 7, 8, 9, or 10 different atoms share the same channel in a voxel.

[0158] An improvement to the encoding is to allow a voxel to represent the location of an atom by partially activating neighboring voxels. This results in partial activation of neighboring neurons in the subsequent neural network, moving from one-hot encoding to "several-warm" encoding. For example, 3 When the grid is placed, the van der Waals diameter is 3.5 Å and therefore the volume is 22.4 Å. 3 It may be exemplary to consider a chlorine atom where voxels within the chlorine atom are completely filled and voxels at the edge of the atom are only partially filled. Thus, channels representing chlorine in partially filled voxels are turned on in proportion to the amount of such voxels that fall within the chlorine atom. For example, if 50% of the voxel volume falls within the chlorine atom, then the channels in the voxel representing chlorine are activated 50%. This may result in a "smoothed" and more accurate representation relative to discrete one-hot coding. Thus, in some embodiments, the atomic features generated in the sampling are distributed to a subset of voxels in the voxel map, where this subset of voxels includes two or more voxels, three or more voxels, five or more voxels, ten or more voxels, or twenty-five or more voxels. In some embodiments, the atomic features consist of an enumeration of atom types (e.g., one of the SYBYL atom types).

[0159] Thus, voxelization (rasterization) of the encoded geometric data (docking of test or training compounds onto the target polymer) is based on various rules that are applied to the input data.

[0160] Figures 5 and 6 provide illustrations of two compounds 502 encoded on a two-dimensional grid 500 of voxels, according to some embodiments. Figure 5 provides two compounds superimposed on the two-dimensional grid. Figure 6 provides one-hot encoding, using different shading patterns to encode the presence of oxygen, nitrogen, carbon, and empty space, respectively. As noted above, such encoding may be referred to as "one-hot" encoding. Figure 6 shows the grid 500 of Figure 5 with compound 502 omitted. Figure 7 provides an illustration of the two-dimensional grid of voxels of Figure 6, with the voxels numbered.

[0161] In some embodiments, feature shapes are represented in forms other than voxels. Figure 8 provides illustrations of various representations in which features (e.g., atomic centers) are represented as 0-D points (representation 802), 1-D points (representation 804), 2-D points (representation 806), or 3-D points (representation 808). Initially, the spacing between points may be chosen randomly. However, as the predictive model is trained, the points may move closer together or further apart.

[0162] In some embodiments, the input representation of the first neural network 72 may be in the form of a 1D array of features, including but not limited to three-dimensional coordinates.

[0163] The voxel map is expanded into a corresponding vector. In some embodiments, each voxel map 52 is optionally expanded into a corresponding vector 54. In some embodiments, each such vector is a one-dimensional vector. For example, in some embodiments, 20 Å each side is centered on the active site of the target polymer 38 with the compound (target compound or training compound) docked in the pose, sampled at a three-dimensional fixed grid spacing of 1 Å to form corresponding voxels of the voxel map, which holds basic structural features of the voxels, such as atom types, as described above, as well as, optionally, more complex compound-target polymer descriptors in each channel. In some embodiments, the voxels of this three-dimensional voxel map are expanded into one-dimensional floating point vectors.

[0164] In some embodiments, the vectorized representation of the voxel map is input to the first neural network 72. In some embodiments, the vectorized representation of the voxel map is stored in GPU memory along with the first neural network 72. This provides the advantage of processing the vectorized representation of the voxel map through the first neural network 72 at a faster rate. However, in other embodiments, any or all of the vectorized representation of the voxel map and the first neural network 72 are within the memory 92 of the system 100 or are simply addressable by the system 92 over a network. In some embodiments, any or all of the vectorized representation of the voxel map and the first neural network 72 are within a cloud computing environment.

[0165] In some embodiments, each vector 54 is provided to a graphical processing unit memory, which includes a network architecture including a first neural network 72 in the form of a convolutional neural network comprising an input layer for sequentially receiving the vectors, a number of convolutional layers, and optionally a scorer. The number of convolutional layers includes an initial convolutional layer and a final convolutional layer. In some embodiments, the convolutional neural network is not in the GPU memory, but in the memory 92 of the system 100. In some embodiments, the voxel map 52 is not vectorized before being input to the first neural network 72.

[0166] In some embodiments where the first neural network 72 is a convolutional neural network, a convolutional layer of the multiple convolutional layers in the network includes a set of learnable filters (also referred to as kernels). Each filter has a fixed three-dimensional size that is convolved (stepped at a predetermined step rate) across the depth, height, and width of the input volume of the convolutional layer, and calculates a dot product (or other function) between the filter's entries (weights, or more generally, parameters) and the input, thereby creating a multidimensional activation map for that filter. In some embodiments, the step rate of the filter is 1 element, 2 elements, 3 elements, 4 elements, 5 elements, 6 elements, 7 elements, 8 elements, 9 elements, 10 elements, or more than 10 elements of the input space. Thus, when a filter size is 5 3 In some embodiments, the filter computes the dot product (or other mathematical function) between consecutive cubes of input space with a depth of 5 elements, a width of 5 elements, and a height of 5 elements, for a total of 125 input space values ​​per voxel channel.

[0167] The input space to the initial convolutional layer (e.g., output from the input layer) is formed from either the voxel map 52 or the vectorized representation 54 of the voxel map. In some embodiments, the vectorized representation of the voxel map is a one-dimensional vectorized representation of the voxel map that serves as the input space to the initial convolutional layer. Nonetheless, when the filter convolves its input space, and the input space is a one-dimensional vectorized representation of the voxel map, the filter still obtains elements from the one-dimensional vectorized representation that represent a corresponding continuous cube of fixed space within the target polymer 38-compound complex. In some embodiments, the filter uses bookkeeping techniques to select those elements from within the one-dimensional vectorized representation that form the corresponding continuous cube of fixed space within the target polymer 38-compound complex. Thus, in some instances, this entails employing a non-contiguous subset of elements within the one-dimensional vectorized representation to obtain element values ​​for the corresponding continuous cube of fixed space within the target polymer 38-compound complex.

[0168] In some embodiments, the filter is initialized (e.g., to Gaussian noise) or trained to have 125 corresponding weights (per input channel) for taking a dot product (or some other form of mathematical operation, such as the function disclosed in FIG. 17) of the 125 input spatial values ​​to calculate a first single value (or set of values) of the activation layer corresponding to the filter. In some embodiments, the values ​​calculated by the filter are summed, weighted, and / or biased. To calculate the sum value of the activation layer corresponding to the filter, the filter is then stepped (convolved) in one of the three dimensions of the input volume by a step rate (stride) associated with the filter, where a dot product (or some other form of mathematical operation, such as the mathematical function disclosed in FIG. 17) of the filter weights and the 125 input spatial values ​​(per channel) is taken at a new location within the input volume. This stepping (convolution) is repeated until the filter has sampled the entire input space according to the step rate. In some embodiments, the edges of the input space are zero-padded to control the spatial volume of the output space generated by the convolution layer. In typical embodiments, each of the filters of a convolutional layer thus canvasses the entire three-dimensional input volume, thereby forming a corresponding activation map. The collection of activation maps from the filters of a convolutional layer collectively forms the three-dimensional output volume of one convolutional layer, thereby serving as the three-dimensional (three spatial dimensions) input of a subsequent convolutional layer. Thus, every entry in the output volume can also be interpreted as the output of a single neuron (or set of neurons) that looks at a small region in the input space to the convolutional layer and shares parameters with neurons in the same activation map. Thus, in some embodiments, a convolutional layer of multiple convolutional layers has multiple filters, each filter of the multiple filters having a stride of N with stride Y (in the three spatial dimensions). 3 where N is an integer greater than or equal to 2 (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10) and Y is a positive integer (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10).

[0169] Each of the multiple convolutional layers is associated with a different set of weights, or more generally, a different set of parameters. More specifically, each of the multiple convolutional layers includes multiple filters, each filter including independent parameters (e.g., weights). In some embodiments, the convolutional layers are of dimension 5 or more. 3 For a given convolutional layer, the convolutional layer has 128 filters, i.e., 128×5×5×5 or 16,000 parameters (e.g., weights) per channel in the voxel map. Thus, if there are five channels in the voxel map, the convolutional layer has 16,000×5 parameters (e.g., weights), i.e., 80,000 parameters (e.g., weights). In some embodiments, such parameters (and, optionally, biases) of some or all of the filters in a given convolutional layer may be tied together, e.g., constrained to be the same.

[0170] In response to input of the respective vectors, the input layer supplies a first plurality of values ​​to the initial convolutional layer as a first function of the values ​​in the respective vectors, the first function being, optionally, computed using a graphical processing unit. In some embodiments, computer system 100 has two or more graphical processing units, each such graphical processing unit being used simultaneously to facilitate computation of first neural network 72.

[0171] Each respective convolutional layer other than the final convolutional layer supplies intermediate values ​​to another convolutional layer in the plurality of convolutional layers as a respective second function of (i) a different set of parameters (e.g., weights) associated with the respective convolutional layer and (ii) the input values ​​received by the respective convolutional layer. In some embodiments, the respective second functions are calculated using a graphical processing unit. For example, in some embodiments, each respective filter of each convolutional layer canvases an input volume (in three spatial dimensions) to the convolutional layer according to the characteristic three-dimensional stride of the convolutional layer, and at each respective filter position, takes a dot product (or some other mathematical function) of the filter parameters (e.g., weights) of the respective filter and the values ​​of the input volume (a continuous cube that is a subset of the total input space) at the respective filter position, thereby generating a calculated point (or set of points) on the activation layer corresponding to the respective filter position. The activation layers of the filters of each convolutional layer collectively represent the intermediate values ​​of the respective convolutional layers.

[0172] In some embodiments, the convolutional neural network has one or more activation layers. In some embodiments, the activation layers are layers of neurons that apply a non-saturating activation function f(x)=max(0,x). This increases the non-linearity of the decision function and the overall network without affecting the receptive field of the convolutional layers. In other embodiments, the activation layers use other functions to increase the non-linearity, such as the saturated hyperbolic tangent function f(x)=tanh, f(x)=│tanh(x)│, and the sigmoid function f(x)=(1+e -x ) -1In some embodiments for the neural network, non-limiting examples of other activation functions included in other activation layers may include, but are not limited to, logistic (or sigmoid), softmax, Gaussian, Boltzmann weighted averaging, absolute value, linear, rectified linear, bounded rectified linear, soft rectified linear, parameterized rectified linear, mean, maximum, minimum, any vector norm LP (for p=1, 2, 3, ..., ∞), sign, square, square root, biquadratic, inverse quadratic, inverse biquadratic, multiharmonic spline, and thin plate spline.

[0173] The network learns filters in the convolutional layer that activate when it sees some particular type of feature at some spatial location in the input. In some embodiments, the initial parameters (e.g., weights) of each filter in the convolutional layer are obtained by training the convolutional neural network on a compound training library. Thus, the operation of the convolutional neural network may result in more complex features than those historically used to make binding affinity predictions. For example, a filter in a given convolutional layer of the network that serves as a hydrogen bond detector may be able to not only recognize that hydrogen bond donors and acceptors are at a given distance and angle, but also recognize that the biochemical environment around the donor and acceptor strengthens or weakens the bond. Additionally, the filters in the network may be trained to effectively distinguish binders from non-binders in the underlying data.

[0174] As mentioned above, in some embodiments, the first neural network 72 is configured to deploy three-dimensional convolutional layers. The input region to the lowest level convolutional layer may be a cube (or other continuous region) of voxel channels from the receptive field. Higher convolutional layers evaluate the outputs from lower convolutional layers while still making their outputs a function of bounded regions of voxels that are close to each other (in 3-D Euclidean distance). In one embodiment, the first neural network 72 is configured to apply regularization techniques to reduce the tendency of the model to overfit the training data.

[0175] Zero or more network layers in the above convolutional neural network may consist of pooling layers. Similar to convolutional layers, pooling layers are a set of function calculations that apply the same function across different spatially localized patches of the input. In the case of pooling layers, the output is given by a pooling operator, e.g., some vector norm LP for p=1, 2, 3, ..., ∞, across multiple voxels. In some embodiments, the pooling is done per channel, while in other embodiments, the pooling is done across channels. In some embodiments, the pooling divides the input space into a set of three-dimensional boxes and outputs some other mathematical pooling operation, such as max or average pooling, for each such sub-region. The pooling operation provides a form of translation invariance. The function of the pooling layer is to gradually reduce the spatial size of the representation to reduce the amount of parameters and calculations in the network, and thus also to control overfitting. In some embodiments, a pooling layer is inserted between successive convolutional layers in the above convolutional neural network. Such a pooling layer operates independently on each depth slice of the input and varies the size spatially. In addition to, or instead of, max pooling, the pooling unit can also perform other functions such as average pooling or even L2-norm pooling.

[0176] Zero or more of the layers in the convolutional neural network may consist of normalization layers, such as local response normalization or local contrast normalization, which may be applied across channels at the same location or to a particular channel across multiple locations. These normalization layers may promote diversity in the response of multiple function calculations to the same input.

[0177] In some embodiments, there is no scorer: rather, the convolutional neural network outputs an initial embedding 74 for the input pose 48 rather than a score.

[0178] Use case. Below are sample use cases, provided for illustrative purposes only, illustrating some applications of some embodiments of the present disclosure. Other uses may be contemplated, and the examples provided below are non-limiting and may be subject to variations, omissions, or may include additional elements.

[0179] Hit Discovery. Pharmaceutical companies spend millions of dollars screening compounds to discover novel promising drug leads. Large compound collections are tested to discover the few compounds that have any interaction with the disease target of interest. Unfortunately, wet lab screening is subject to experimental error, and in addition to the cost and time to perform assay experiments, the collection of large screening collections imposes significant challenges through storage constraints, shelf stability, or chemical costs. Even the largest pharmaceutical companies have between only a few hundred thousand and a few million compounds, versus tens of millions of commercially available molecules and hundreds of millions, billions, or even trillions of simulatable molecules. Examples of commercially available molecular databases include MCULE (Kiss et al., 2012, “Http: / / Mcule.Com: A Public Web Service for Drug Discovery,” J. Cheminformatics 4(1), p. 17.) and ENAMINE (Irwin et al., 2016, “Docking Screens for Novel Ligands Conferring New Biology,” J. Med. Chem. 59(9), pp. 4103-4120).

[0180] A potentially more efficient alternative to physical experimentation is virtual high-throughput screening. Just as physics simulations can help aerospace engineers evaluate possible wing designs before the models are physically tested, computational molecular screening can focus experimental testing on a small subset of highly promising molecules. This can reduce the cost and time of screening, reduce false negatives, improve success rates, and / or cover a broader swath of the chemical environment.

[0181] In this application, a protein target may be provided as an input to the system. A large set of compounds may also be provided. For each compound, a binding affinity to the protein target is predicted. The resulting scores may be used to rank the compounds, with the best scoring compounds being most likely to bind to the target protein. Optionally, the ranked compound list may be analyzed for clusters of similar compounds, and larger clusters may be used as stronger predictors of compound binding, or compounds may be selected across clusters to ensure diversity in validation experiments.

[0182] Off-target side effects prediction. Many drugs can be found to have side effects. These side effects are often due to interactions with biological pathways other than the pathway responsible for the drug's therapeutic effect. These off-target side effects can be unpleasant or dangerous and may limit the patient population in which the drug is safe to use. Off-target side effects are therefore an important criterion for evaluating which drug candidates to further develop. It is important to characterize the interactions of drugs with many alternative biological targets, but such tests can be expensive and time-consuming to develop and perform. Computational prediction can make this process more efficient.

[0183] In applying one embodiment of the present disclosure, a panel of protein targets associated with significant biological responses and / or side effects can be constructed. The system can then be configured to predict binding to each protein target in the panel in turn. Strong activity against a particular target (i.e., activity as strong as a compound known to activate off-target proteins) can implicate the molecule in side effects due to off-target effects.

[0184] Toxicity prediction. Toxicity prediction is a special case of off-target side effect prediction that is of particular importance. Approximately half of drug candidates in late-stage clinical trials fail due to unacceptable toxicity. As part of the new drug approval process (and before a drug candidate can be tested in humans), the FDA requires toxicity test data for a set of targets including cytochrome P450 liver enzymes (whose inhibition may lead to toxicity from drug-drug interactions) or the hERG channel (whose binding may lead to QT prolongation leading to ventricular arrhythmias and other adverse cardiac effects).

[0185] In toxicity prediction, the system identifies off-target proteins as major anti-targets (e.g., CYP450, hERG, or 5-HT 2B Acceptors) of the original compound. Binding affinities for drug candidates can then be predicted for these proteins. Optionally, the compounds can be analyzed to predict a set of metabolites (subsequent molecules produced by the body during metabolism / degradation of the original compound), which can also be analyzed for binding to the anti-target. Problematic compounds can be identified and modified to avoid toxicity, or development on a molecule series can be halted to avoid further waste of resources.

[0186] Optimizing Potency. One of the primary requirements for a drug candidate is strong binding to its disease target. Screening rarely finds a compound that binds strongly enough to be clinically effective. Thus, an initial compound seeds a lengthy optimization process, in which medicinal chemists iteratively modify the compound's molecular structure to propose new compounds with increased strength of target binding. Each new compound is synthesized and tested to determine whether the changes successfully improved binding. Systems can be configured to facilitate this process by replacing physical testing with computational predictions.

[0187] In this application, a disease protein target and a set of lead compounds can be input into the system. The system can be configured to generate a set of lead binding affinity predictions. Optionally, the system can highlight differences between the candidate compounds that may help inform reasons for predicted differences in binding affinity. A medicinal chemist user can use this information to suggest new sets of compounds, preferably with improved activity against the target. These new alternative compounds can be analyzed in the same manner.

[0188] Optimizing selectivity. As mentioned above, compounds tend to bind to a host of proteins with varying strengths. For example, the binding pockets of protein kinases (which are common chemotherapy and inflammation targets) are very similar, and most kinase inhibitors affect many different kinases. This means that many different biological pathways are modified simultaneously, resulting in a "dirty" pharmaceutical profile and many side effects. Thus, a key challenge in the design of many drugs is not activity per se, but specificity: the ability to selectively target one protein (or a subset of proteins) from a set of possibly closely related proteins.

[0189] The system can reduce the time and cost to optimize the selectivity of a candidate drug. In the present application, a user can input two sets of target proteins. One set describes the target proteins on which the compound should be active, while the other set describes the target proteins on which the compound should be inactive. The system can be configured to make predictions for the compound against all of the proteins in both sets and establish a profile of interaction strengths. Optionally, these profiles can be analyzed to suggest explanatory patterns in the target proteins. The user can use the information generated by the system to consider structural modifications to the compound that improve the relative binding to the different target protein sets and design new candidate compounds with better specificity. Optionally, the system can be configured to highlight differences between the candidate compounds that may help inform the reasons for the differences in predicted selectivity. The proposed candidates can be analyzed iteratively to further refine the specificity of their respective activity profiles.

[0190] Fitness Functions for Automated Molecular Design. Automated tools to perform the aforementioned optimizations are invaluable. Successful drug compounds require optimization and balance between potency, selectivity, and toxicity. "Scaffold hopping" (when the activity of the lead compound is retained but the chemical structure is significantly altered) can result in improved pharmacokinetic, pharmacodynamic, toxicity, or intellectual property profiles. Algorithms exist that iteratively suggest new compounds, such as random generation of compounds, growing molecular fragments to fill a given binding site, genetic algorithms to "mutate" and "crossbreed" a population of compounds, and exchanging fragments of compounds with bioisosteric substitutions. Compounds generated by each of these methods can be evaluated against the multiple objectives mentioned above (potency, selectivity, toxicity) and can be incorporated into an automated compound design system in the same way that the technology can inform each of the aforementioned manual settings (binding prediction, selectivity, side effect and toxicity predictions).

[0191] Repurposing of Drugs. Drugs typically have side effects, and sometimes these side effects are beneficial. For example, aspirin, commonly used as a headache remedy, is also taken for cardiovascular health. Drug repositioning can significantly reduce the cost, time, and risk of drug discovery because the drug has already been shown to be safe in humans and is optimized for rapid absorption and favorable stability in patients. Unfortunately, drug repositioning is largely accidental. For example, sildenafil (Viagra) was developed as a blood pressure lowering drug and serendipitously observed to be an effective treatment for erectile dysfunction. Computational prediction of off-target effects can be used in the context of drug repurposing to identify compounds that can be used to treat alternative diseases.

[0192] In this disclosure, similar to off-target side effect prediction, the user can assemble a set of possible target proteins, each of which is linked to a disease. That is, inhibition of each target protein treats a (possibly different) disease; for example, an inhibitor of cyclooxygenase-2 may provide relief from inflammation, while an inhibitor of factor Xa may be used as an anticoagulant. These target proteins are annotated with the binding affinity of approved drugs, if present. A set of compounds is then constructed, and this set is limited to compounds that have been approved or investigated for use in humans. Finally, for each pair of target proteins and compounds, the user can use the system to predict the binding affinity. Candidates for repurposing the drug may be identified if the predicted binding affinity of the molecule is close to the binding affinity of the effective drug for the protein.

[0193] Drug Resistance Prediction. Drug resistance is an inevitable consequence of drug use that puts selective pressure on pathogen populations to rapidly divide and mutate. Drug resistance is seen in diverse pathogens such as viruses (HIV), exogenous microbes (MRSA), and dysregulated host cells (cancer). Over time, a given drug becomes ineffective, whether the drug is an antibiotic or chemotherapy. At that point, intervention can shift to a different, hopefully still more potent drug. In HIV, there is a well-known path of disease progression defined by the mutations the virus accumulates while the patient is being treated.

[0194] There is considerable interest in predicting how pathogens will adapt to medical interventions. One approach is to characterize which mutations will arise in pathogens during treatment. Specifically, a drug's protein target needs to mutate to avoid binding the drug while at the same time continuing to bind its natural substrate.

[0195] In this application, a set of possible mutations of the target protein can be proposed. For each mutation, the resulting protein shape can be predicted. For each of these mutant protein forms, the system can be configured to predict the binding affinity to both the natural substrate and the drug. Mutations that cause the protein to no longer bind the drug but continue to bind to the natural substrate are candidates for conferring drug resistance. These mutated proteins can be used as targets for designing drugs, for example, by using these proteins as inputs into one of these other predictive use cases.

[0196] Personalized medicine. Ineffective drugs should not be administered. In addition to cost and effort, all drugs have side effects. Moral and economic considerations make it essential to give drugs only when the benefits outweigh these harms. It can be important to be able to predict when a drug will be useful. People differ from each other by a small number of mutations. However, small mutations can have immense effects. When these mutations occur at active (orthosteric) or regulatory (allosteric) sites of the disease target, they may prevent the drug from binding and thus inhibit the drug's activity. When the protein structure of a particular person is known (or predicted), the system can be configured to predict whether a drug will be effective, or the system can be configured to predict when a drug will not work.

[0197] In this application, the system may be configured to receive as input the chemical structure of a drug and a particular expressed protein of a particular patient. The system may be configured to predict binding between the drug and the protein, and if the predicted binding affinity of the drug is too weak for a particular patient's protein structure to be clinically effective, the clinician or practitioner may prevent the drug from being unprofitably prescribed to the patient.

[0198] Clinical Trial Design. This application generalizes the personalized medicine use case above to the case of patient populations. When the system can predict whether a drug will be effective for a particular patient phenotype, this information can be used to help design clinical trials. By excluding patients whose particular disease targets cannot be sufficiently affected by the drug, clinical trials can achieve statistical power using fewer patients. Fewer patients directly reduce the cost and complexity of clinical trials.

[0199] In this application, the user may divide a potential patient population into subpopulations characterized by the expression of different proteins (e.g., due to mutations or isoforms). The system may be configured to predict the binding strength of a drug candidate to different protein types. If the predicted binding strength to a particular protein type indicates a required drug concentration below the clinically achievable concentration in the patient (e.g., as based on physical characterization in test tubes, animal models, or healthy volunteers), the drug candidate is predicted to fail for that protein subpopulation. Patients with that protein may then be excluded from the clinical trial.

[0200] Pesticide design. In addition to pharmaceutical uses, the agrochemical industry uses binding predictions in the design of new insecticides. For example, one consideration for pesticides is that they stop a single species of interest without adversely affecting any other species. For environmental safety, one would want to kill weevils without killing bumblebees.

[0201] In this application, the user can input a set of target protein structures from different species under consideration into the system. A subset of the target proteins can be identified as target proteins against which the compound should be active, while the rest are identified as target proteins against which the compound should be inactive. As in the previous use case, several sets of compounds (whether in an existing database or newly generated) are considered for each target protein, and the system identifies compounds with maximum efficacy against the first group of target proteins while avoiding the second group of target proteins.

[0202] Materials Science. It can be useful to analyze molecular interactions to predict the behavior and properties of new materials. For example, to study solvation, a user can input a repeating crystal structure of a given small molecule and evaluate the binding affinity of another instance of the small molecule on the surface of the crystal. To study polymer strength, a set of polymer strands can be input similar to a protein target structure, and polymer oligomers can be input as small molecules. Thus, the binding affinity between the polymer strands can be predicted by the system.

[0203] Simulation. Simulators often measure the binding affinity of compounds to proteins, since the tendency of a compound to remain in a region of a target protein correlates with its binding affinity there. Accurate descriptions of the features governing binding can be used to identify regions and poses with particularly high or low binding energy. Energy descriptions can be incorporated into molecular dynamics simulations to describe molecular motions and occupancy of protein binding regions. Similarly, stochastic simulators for studying and modeling systems biology can benefit from accurate predictions of how small changes in compound concentration affect biological networks. EXAMPLES

[0204] FIG. 21A provides the area under the curve (AUC) receiver operating characteristic (ROC) as a function of training iterations for two architectures of the present disclosure, o3-2.8.0 (2106) and o4-2.8.0 (2108), illustrated in FIG. 9B, versus two other deep learning neural networks, architectures n8b-long (2102) and n8b-max-long (2104). The n8b-long and n8b-max-long architectures are of the type disclosed in U.S. Provisional Patent Application No. 63 / 251,142, entitled "Characterization of Interactions Between Compounds and Polymers Using Negative Pose Data and Model Conditioning," filed October 1, 2021, which is incorporated herein by reference, and which requires a pose for each training compound, including a "positive pose" and a "negative pose." Architectures n8b-long and n8b_max-long do not apply the attention mechanism 77 to multiple initial embeddings 74, each representing a different pose 48 of the compound, to obtain attention embeddings 79, whereas architectures o3-2.8.0 and o4-2.8.0 do. For the calculations summarized in FIG. 21A, 3,097 different target proteins were utilized and a library of 5,562,818 compound-protein pairs was available for model training. All 3,097 different target proteins were used for this training. The 3,097 different target proteins were considered to be non-allosteric. As described above with reference to Figure 10, each of the four models (n8b-long, n8b-max-ling, o3-2.8.0, and o4-2.8.0) was trained using a respective mini-batch of compound-protein pairs, where each mini-batch of compound-protein pairs was selected from among the available compound-protein pairs using a graphical processing unit (GPU), and the number of compound-protein pairs in a mini-batch was limited by the GPU memory size. For the training referenced in Figure 21A, each mini-batch consisted of poses for 64 different compound-protein pairs.Each compound used in training had a known binary activity label for at least some of the 3,097 target proteins, either "active" if the pKa of the compound for the target protein was less than 10 μM, or "inactive" if the pKa of the compound for the target protein was greater than 10 μM. Thus, each compound used in the training summarized in FIG. 21A consisted of an activity label for at least some of the 3097 target proteins. During the training summarized in FIG. 21A, if the activity of a particular training compound for a particular target protein was not known, the corresponding compound-protein pair was not included in the training. For each respective compound-protein pair of each mini-batch, the error exhibited by the respective model in calculating the activity label for each training compound for the paired target protein pair of each compound-protein pair was used to refine the parameters of each model as part of the interaction. Thus, each iteration represented training resulting from all compound-protein pairs of each mini-batch. As shown in Figure 21A, the ROC AUC performance of each model in predicting the activity of the training compound in the compound-protein pair is obtained from 100000 to 5,000,000 such iterations. Figure 21A shows that the ROC AUC statistics of the models of the present disclosure, represented by curves 2106 and 2108, consistently improve as the number of iterations increases, while the ROC AUC statistics of models n8b-long and n8b-max-long, represented by curves 2102 and 2104, show overtraining such that their ROC AUC values ​​decrease after about 500000 iterations.

[0205] FIG. 21B provides the AUC ROC statistics for two architectures of the present disclosure, o3-2.8.0 (2106) and o4-2.8.0 (2108), compared to two other deep learning neural networks, architectures-long (2102) and n8b-max-long (2104), on the allosteric benchmark AA103. FIG. 21B shows that o3-2.8.0 and o4-2.8.0 have improved AUC ROC statistics compared to architectures-long (2102) and n8b-max-long (2104) on the allosteric benchmark AA103. The calculations summarized in FIG. 21B utilized 103 different target proteins and had a library of 552,011 compound-protein pairs available for model training. All 103 different target proteins were used for this training. The 103 different target proteins were considered to be allosteric. As described above with reference to FIG. 21A, each of the four models (n8b-long, n8b-max-long, o3-2.8.0, and o4-2.8.0) was trained using a respective mini-batch of training compound-protein pairs, where each mini-batch of compound-protein pairs was selected from among the available training compound-protein pairs using a graphical processing unit (GPU), and the number of compound-protein pairs in a mini-batch was limited by the GPU memory size. For the training referred to in FIG. 21B, each mini-batch consisted of all poses for 64 different compound-protein pairs. Each training compound used in model training had a known binary activity label for at least some of the 103 target proteins, either "active" if the pKa of the training compound for the target protein was less than 10 μM, or "inactive" if the pKa of the training compound for the target protein was greater than 10 μM. Thus, each training compound used in the training summarized in FIG. 21B consisted of an activity label for at least some of the 103 target proteins. During training, summarized in FIG. 21B, if the activity of a particular training compound against a particular target protein was not known, the corresponding compound-protein pair was not included in the training.For each respective compound-protein pair of each mini-batch, the error in calculating the activity label for each training compound against the paired target protein pair of each compound-protein pair was used to refine the model parameters as part of the interaction. Thus, each iteration represented the training resulting from all compound-protein pairs of each mini-batch. As shown in FIG. 21B, the ROC AUC performance is obtained from 100000 to 5,000,000 such iterations. FIG. 21B shows that the ROC AUC statistics of the models of the present disclosure, represented by curves 2106 and 2108, consistently improve as the number of iterations increases, while the ROC AUC statistics of models n8b-long and n8b-max-long, represented by curves 2102 and 2104, exhibit overtraining as the number of training iterations increases.

[0206] conclusion For purposes of explanation, the foregoing description has been set forth with reference to specific implementations. However, the exemplary discussion above is not intended to be exhaustive or to limit the implementation to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. The implementations have been chosen and described to best explain the principles and their practical applications, thereby enabling those skilled in the art to best utilize the implementations suited to the particular use contemplated, and various implementations with various modifications.

Claims

1. 1. A computer system for characterizing an interaction between a test compound and a target polymer, the computer system comprising: one or more processors; a memory addressable by said one or more processors, said memory storing at least one program for execution by said one or more processors, said at least one program comprising: (A) obtaining a plurality of atomic coordinate sets, each respective atomic coordinate set in the plurality of atomic coordinate sets comprising the test compound bound to the target polymer at a corresponding pose in a plurality of poses, each respective atomic coordinate set in the plurality of atomic coordinate sets comprising atomic coordinates of at least five atoms; (B) for each respective atomic coordinate set in the plurality of atomic coordinate sets, inputting the respective atomic coordinate set or an encoding of the respective atomic coordinate set into a first neural network to obtain a corresponding initial embedding as an output of the first neural network, thereby obtaining a plurality of initial embeddings, each initial embedding in the plurality of initial embeddings corresponding to an atomic coordinate set in the plurality of atomic coordinate sets, the first neural network including more than 400 parameters, and performing at least 10,000 calculations to calculate each initial embedding in the plurality of initial embeddings; (C) applying an attention mechanism to the multiple initial embeddings in a concatenated manner, thereby obtaining an attention embedding; (D) applying a pooling function to the attention embeddings to derive pooled embeddings; (E) inputting the pooled embeddings into a first model, thereby obtaining a first interaction score for an interaction between the test compound and the target polymer.

2. The computer system of claim 1 , wherein the first interaction score represents a binding coefficient of the test compound to the target polymer.

3. The binding coefficient is the IC 50 , E.C. 50 , Kd, ​​KI, or pKI.

4. The computer system of claim 1 , wherein the first interaction score represents an in silico quality score of the test compound relative to the target polymer.

5. The computer system according to any one of claims 1 to 4, wherein the first model is a fully connected second neural network.

6. the at least one program further comprises instructions for inputting the pooled embeddings into a second model, thereby obtaining a second interaction score for an interaction between the test compound and the target polymer; the first model is a first fully connected neural network; the second model is a second fully connected neural network; the first interaction score represents an in silico quality score of the test compound relative to the target polymer; The computer system of claim 1 , wherein the second interaction score represents an in silico quality score of the test compound relative to the target polymer.

7. 7. The computer system of claim 6, wherein the at least one program further comprises instructions for inputting the first interaction scores and the second interaction scores into a third model to obtain a third interaction score, the third model being a third fully-connected neural network.

8. 8. The computer system of claim 7, wherein the third interaction score is a discrete binary activity score having a first value if the test compound is determined to be inactive by the third model and a second value if the test compound is determined to be active by the third model.

9. The computer system of any one of claims 1 to 8, wherein the target polymer is an assembly of a protein, a polypeptide, a polynucleic acid, a polyribonucleic acid, a polysaccharide, or any combination thereof.

10. Each set of atomic coordinates in the plurality of atomic coordinate sets is a set of three-dimensional coordinates {x,y,z,z} for at least a portion of the target polymer from a crystal structure of the target polymer resolved at 2.5 Å resolution or better or 3.3 Å resolution or better. 1 , …, x N 10. The computer system of claim 1, further comprising:

11. 10. The computer system of claim 1, wherein each atomic coordinate set in the plurality of atomic coordinate sets comprises a collection of three-dimensional coordinates for at least a portion of the target polymer determined by nuclear magnetic resonance, neutron diffraction, or cryo-electron microscopy.

12. the first interaction score is a binary score; A first value of the binary score represents an IC of the test compound against the target polymer that is greater than a first threshold. 50 , E.C. 50 , Kd, ​​KI, or pKI; A second value of the binary score is an IC of the test compound against the target polymer that is below the first threshold. 50 , E.C. 50 , Kd, ​​KI, or pKI.

13. The computer system of any one of claims 1 to 12, wherein the test compound satisfies two or more, three or more, or all four of Lipinski's rule of five: (i) no more than five hydrogen bond donors, (ii) no more than ten hydrogen bond acceptors, (iii) molecular weight less than 500 Daltons, and (iv) Log P less than 5.

14. 14. The computer system of any one of claims 1 to 13, wherein the test compound is an organic compound having a molecular weight of less than 500 Daltons, less than 1000 Daltons, less than 2000 Daltons, less than 4000 Daltons, less than 6000 Daltons, less than 8000 Daltons, less than 10000 Daltons, or less than 20000 Daltons.

15. The computer system according to any one of claims 1 to 13, wherein the test compound is an organic compound having a molecular weight between 400 and 10,000 daltons.

16. The computer system according to any one of claims 1 to 15, wherein the plurality of atomic coordinate sets consists of 2 to 64 poses.

17. The computer system of any one of claims 1 to 16, wherein the first neural network is a convolutional neural network.

18. 20. The computer system of claim 17, wherein each respective set of atomic coordinates in the plurality of atomic coordinate sets includes atomic coordinates for at least 80 atoms.

19. The computer system of any one of claims 1 to 16, wherein the first neural network is a graph neural network.

20. 20. The computer system of claim 19, wherein the graph neural network is characterized by an initial embedding layer and a plurality of interaction layers each contributing to an interaction data structure in a plurality of interaction data structures for each atom in a respective set of atomic sets of atomic coordinates of the corresponding pose in the plurality of poses pooled to form the corresponding initial embedding of the corresponding pose.

21. The computer system of any one of claims 1 to 16, wherein the first neural network is an equivariant neural network or a message passing neural network.

22. 17. The computer system of claim 1, wherein the first neural network includes a plurality of graph convolution blocks, each of which uses a plurality of radial graphs to account for connectivity in the respective sets of atomic coordinates.

23. The first neural network is 1×10 6 A computer system according to any one of claims 1 to 22, comprising parameters.

24. The computer system of any preceding claim, wherein the corresponding initial embedding comprises a data structure containing 100 or more values.

25. The plurality of initial embeddings includes a first plurality of values, and the applying of the attention mechanism includes: (i) inputting the first plurality of values ​​into an attention neural network to obtain a first plurality of weights, each weight in the first plurality of weights corresponding to a respective value in the first plurality of values; (ii) weighting each respective value in the first plurality of values ​​by the corresponding weight in the plurality of weights, thereby obtaining the attention embedding.

26. 26. The computer system of claim 25, wherein a sum of the first plurality of weights is one and each weight in the first plurality of weights is a scalar value between zero and one.

27. 27. The computer system of claim 1, wherein the pooling function collapses the attention embedding into the pooled embedding by applying a statistical function to combine portions of the attention embedding representing different poses of the plurality of poses to form the pooled embedding.

28. 28. The computer system of claim 27, wherein the attention embedding includes a corresponding plurality of values ​​for a corresponding plurality of elements for each respective pose in the plurality of poses, and the statistical function is a maximum function that takes a maximum value over corresponding elements of each respective pose represented in the attention embedding to form the pooled embedding.

29. 28. The computer system of claim 27, wherein the attention embedding includes a corresponding plurality of values ​​for a corresponding plurality of elements for each respective pose in the plurality of poses, and the statistical function is an averaging function that averages the corresponding elements for each respective pose represented in the attention embedding to form the pooled embedding.

30. 2. The computer system of claim 1, wherein the first model is a regression task and the first interaction score quantifies the interaction between the test compound and the target polymer.

31. 2. The computer system of claim 1, wherein the first model is a classification task and the first interaction score classifies the interaction between the test compound and the target polymer.

32. 32. The computer system of claim 1, wherein each respective atomic coordinate set in the plurality of atomic coordinate sets includes atomic coordinates for at least 15 atoms, at least 20 atoms, at least 25 atoms, or at least 30 atoms.

33. 33. The computer system of claim 1, wherein the first neural network performs at least 100,000 calculations to compute each initial embedding in the plurality of initial embeddings.

34. The first neural network uses at least 1×10 6 A computer system according to any one of claims 1 to 32, for performing calculations.

35. The computer system of any one of claims 1 to 34, wherein the first model performs at least 10,000 calculations to calculate the first interaction score.

36. The computer system of any one of claims 1 to 34, wherein the first model performs at least 100,000 calculations to calculate the first interaction score.

37. The first model is used to calculate the first interaction score. 6 A computer system according to any one of claims 1 to 34, for performing calculations.

38. 38. The computer system of claim 1, wherein the first model includes more than 400 parameters and the first model performs more than 1000 calculations to calculate the first interaction score.

39. 38. The computer system of claim 1, wherein the first model includes more than 400 parameters and the first model performs more than 10,000 calculations to calculate the first interaction score.

40. 1. A method for characterizing an interaction between a test compound and a target polymer, the method comprising: (A) obtaining a plurality of atomic coordinate sets, each respective atomic coordinate set in the plurality of atomic coordinate sets comprising the test compound bound to the target polymer at a corresponding pose in a plurality of poses, each respective atomic coordinate set in the plurality of atomic coordinate sets comprising atomic coordinates of at least five atoms; (B) for each respective atomic coordinate set in the plurality of atomic coordinate sets, inputting the respective atomic coordinate set or an encoding of the respective atomic coordinate set into a first neural network to obtain a corresponding initial embedding as an output of the first neural network, thereby obtaining a plurality of initial embeddings, each initial embedding in the plurality of initial embeddings corresponding to an atomic coordinate set in the plurality of atomic coordinate sets, the first neural network including more than 400 parameters, and performing at least 10,000 calculations to calculate each initial embedding in the plurality of initial embeddings; (C) applying an attention mechanism to the initial embeddings in a concatenated manner, thereby obtaining an attention embedding; and (D) applying a pooling function to the attention embeddings to derive pooled embeddings; and (E) inputting the pooled embeddings into a first model, thereby obtaining a first interaction score for an interaction between the test compound and the target polymer.

41. 1. A non-transitory computer readable storage medium storing instructions that, when executed by a computer system, cause the computer system to perform a method for characterizing an interaction between a test compound and a target polymer, the method comprising: (A) obtaining a plurality of atomic coordinate sets, each respective atomic coordinate set in the plurality of atomic coordinate sets comprising the test compound bound to the target polymer at a corresponding pose in a plurality of poses, each respective atomic coordinate set in the plurality of atomic coordinate sets comprising atomic coordinates of at least five atoms; (B) for each respective atomic coordinate set in the plurality of atomic coordinate sets, inputting the respective atomic coordinate set or an encoding of the respective atomic coordinate set into a first neural network to obtain a corresponding initial embedding as an output of the first neural network, thereby obtaining a plurality of initial embeddings, each initial embedding in the plurality of initial embeddings corresponding to an atomic coordinate set in the plurality of atomic coordinate sets, the first neural network including more than 400 parameters, and performing at least 10,000 calculations to calculate each initial embedding in the plurality of initial embeddings; (C) applying an attention mechanism to the initial embeddings in a concatenated manner, thereby obtaining an attention embedding; and (D) applying a pooling function to the attention embeddings to derive pooled embeddings; and (E) inputting the pooled embeddings into a first model, thereby obtaining a first interaction score for an interaction between the test compound and the target polymer.