Machine learning enabled analytics of intermolecular interfaces
Machine learning analysis of intermolecular interface features addresses the challenge of optimizing protein sequences by determining binding affinity and specificity, even in low data scenarios, enhancing the development of therapeutic proteins.
Patent Information
- Application Number
- PCT/US2025/026258
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-25
- Filing Date
- 2025-04-24
- Publication Date
- 2025-10-30
AI Technical Summary
Conventional computational methods struggle to accurately assess and optimize protein sequences for therapeutic applications due to the vast and sparsely populated solution space of protein sequences, especially in low data regimes with scarce empirical data, making it difficult to generate protein sequences capable of binding to targets with sufficient affinity and specificity.
A system and method that utilizes machine learning to analyze intermolecular interface features, such as geometric and energetic features, independent of the monomer sequence, to determine properties like binding affinity and specificity, by generating a structural ensemble of possible conformations and applying feature and property computation models.
Enables accurate assessment and optimization of protein sequences by leveraging interpretable interface features, even in low data regimes, improving the likelihood of developing viable therapeutic proteins.
Smart Images

Figure US2025026258_30102025_PF_FP_ABST
Abstract
Description
MACHINE LEARNING ENABLED ANALYTICS OF INTERMOLECULAR INTERFACESCROSS REFERENCE TO RELATED APPLICATION
[0001] This application claims priority to U.S. Provisional Application No. 63 / 638,915, filed on April 25, 2024 and entitled “COMPUTATIONAL PROTEIN-TO- PROTEIN ANALYTICS,” the disclosure of which is incorporated by reference in its entirety.TECHNICAL FIELD
[0002] The subject matter described herein relates generally to protein design and more specifically to machine learning enabled techniques for analyzing intermolecular interfaces.BACKGROUND
[0003] Molecular binding is a process in which two or more molecules interact to form a molecular structure, such as a protein-ligand complex. For example, molecular binding can include intermolecular interactions (e.g., electrostatics, van der Waals forces, and / or the like) between two molecules. The portion of one molecule that participates in the chemical bonding with the other molecule can include one or more binding sites. For instance, in the case of a protein-ligand complex formed by the binding interaction between a protein molecule and a ligand, each binding site in the protein molecule may include one or more amino acid residues that participate in the binding interaction with the ligand. Where the ligand is another protein molecule, such as a viral or tumor antigen, the binding sites in the ligand may also include one or more of the amino acid residues that participate in the binding interaction.SUMMARY
[0004] Systems, methods, and articles of manufacture, including computer program products, are provided for computational intermolecular interface analytics. In one aspect, thereis provided a system including at least one data processor and at least one memory. The at least one memory may store instructions, which when executed by the at least one data processor, result in operations comprising: generating a molecule design comprising a sequence of monomers; determining one or more interface features of an intermolecular interface between the molecule design and a binding target, where the one or more interface features are determined by at least identifying one or more contacts at the intermolecular interface, where each contact of the one or more contacts comprises a pair of monomers at the intermolecular interface having a threshold interaction energy, and determining, based at least on the one or more contacts, the one or more interface features; and determining, based at least on the one or more interface features, a property of the molecule design.
[0005] In another aspect, there is provided a computer-implemented method for computational intermolecular interface analytics. The method may include: generating a molecule design comprising a sequence of monomers; determining one or more interface features of an intermolecular interface between the molecule design and a binding target, where the one or more interface features are determined by at least identifying one or more contacts at the intermolecular interface, where each contact of the one or more contacts comprises a pair of monomers at the intermolecular interface having a threshold interaction energy, and determining, based at least on the one or more contacts, the one or more interface features; and determining, based at least on the one or more interface features, a property of the molecule design.
[0006] In another aspect, there is provided a computer program product including a non-transitory computer-readable medium storing instructions. The instructions may cause operations when executed by at least one data processor. The operations may include: generating a molecule design comprising a sequence of monomers; determining one or moreinterface features of an intermolecular interface between the molecule design and a binding target, where the one or more interface features are determined by at least identifying one or more contacts at the intermolecular interface, where each contact of the one or more contacts comprises a pair of monomers at the intermolecular interface having a threshold interaction energy, and determining, based at least on the one or more contacts, the one or more interface features; and determining, based at least on the one or more interface features, a property of the molecule design.
[0007] In some variations, one or more features disclosed herein including the following features can optionally be included in any feasible combination.
[0008] In some variations, the one or more contacts are identified by at least determining an interaction energy for each pair of monomers at the intermolecular interface between the molecule design and the binding target, and identifying, based at least on the interaction energy between each pair of monomers at the intermolecular interface between the molecule design and the binding target, the one or more contacts.
[0009] In some variations, wherein the determining the one or more interface features further includes identifying, based at least on an interaction energy of each contact of the one or more contacts, one or more hotspot contacts, wherein a contact is identified as a hotspot contact where an interaction energy between a pair of interacting monomers comprising the contact satisfies one or more thresholds, and determining the one or more interface features to include one or more of a total quantity, an average density, a distribution over specific intermolecular interface regions, an aggregate energy value, and a summary statistic of energy distribution of the one or more hotspot contacts.
[0010] In some variations, the one or more interface features include one or more of a total quantity, an average density, a distribution over specific intermolecular interface regions, a minimum energy, an aggregate energy value and a summary statistic of energy distribution of the one or more contacts.[OH] In some variations, the determining the one or more interface features further includes determining a location of each contact of the one or more contacts, determining, based at least on the location of each contact of the one or more contacts, a functional role of each contact of the one or more contacts, and determining the one or more interface features further corresponds to the functional role of each contact of the one or more contacts.
[0012] In some variations, the location of each contact of the one or more contacts include a location of each monomer of the pair of interacting monomers comprising the contact.[0131 In some variations, the functional role of a contact is structure-determining where both monomers of the pair of interacting monomers comprising the contact are located in a complementarity determining region (CDR) of the molecule design.
[0014] In some variations, the functional role of a contact is structure-determining where one monomer of the pair of interacting monomers comprising the contact is located in a complementarity determining region (CDR) of the molecule design and another monomer of the pair of interacting monomers comprising the contact is located in a framework region (FW) of the molecule design.
[0015] In some variations, the functional role of a contact is affinity-determining or specificity-determining where one monomer of the pair of interacting monomers comprising the contact is located in the molecule design and another monomer of the pair of interacting monomers comprising the contact is located in the binding target.
[0016] In some variations, the determining the one or more interface features further includes identifying, based at least on the one or more contacts, one or more interaction motifs present at the intermolecular interface between the molecule design and the binding target, and determining the one or more interface features to include one or more of a total quantity, an average density, a distribution over specific intermolecular interface regions, an aggregate energy value, and a summary statistic of energy distribution of the one or more interaction motifs.
[0017] In some variations, each interaction motif of the one or more interaction motifs comprises a pattern of one or more contacts observed in a threshold quantity of molecule designs.
[0018] In some variations, the one or more interaction motifs include a favorable interaction motif observed in a threshold quantity of molecular complexes exhibiting one or more desirable properties.
[0019] In some variations, the one or more interaction motifs include an unfavorable interaction motif observed in molecule complexes failing to exhibit one or more desirable properties or exhibiting one or more undesirable properties.
[0020] In some variations, the one or more interaction motifs include a false positive interaction motif present in a threshold quantity of computationally generated molecular complexes but absent from a threshold quantity of molecular complexes observed in nature.
[0021] In some variations, the one or more interaction motifs include one or more preferred pairings of monomers.
[0022] In some variations, the determining the one or more interface features further include determining a source of interaction energy between the pair of monomers comprisingeach contact of the one or more contacts in the intermolecular interface between the molecule design and the binding target, and determining the one or more interface features to include the source of interaction energy for each contact of the one or more contacts.
[0023] In some variations, the source of interaction energy include sidechain-specific atomic interactions or backbone-mediated atomic interactions.
[0024] In some variations, the determining the one or more interface features further include determining a type of interaction energy between the pair of monomers comprising each contact of the one or more contacts, determining, based at least on the type of interaction energy, a functional role of each contact of the one or more contacts, and determining the one or more interface features to include the functional role of each contact of the one or more contacts.
[0025] In some variations, the determining the one or more interface features further include determining one or more specificity determining interatomic forces as giving rise to an interaction energy between the pair of monomers comprising a contact, and determining, based at least on the one or more specificity determining interatomic forces, a functional role of the contact as specificity-determining.
[0026] In some variations, the one or more specificity determining interatomic forces include hydrogen-bonding, electrostatics, and nonpolar van der Waals forces.
[0027] In some variations, the determining the one or more interface features further include determining one or more nonspecific interatomic forces as giving rise to an interaction energy between the pair of monomers comprising a contact, and determining, based at least on the one or more nonspecific interatomic forces, a functional role of the contact as neutral, affinity-determining, and / or structure-determining.
[0028] In some variations, the one or more nonspecific interatomic forces include solvation interactions.
[0029] In some variations, a structural ensemble including a plurality of three- dimensional structures of a molecular complex comprising the molecule design bound to the binding target is generated. The one or more interface features of the intermolecular interface between the molecule design and the binding target are determined based at least on the structural ensemble.
[0030] In some variations, the one or more interface features are determined as a distribution of values across the plurality of three-dimensional structures comprising the structural ensemble.
[0031] In some variations, the generating the structural ensemble includes iteratively adjusting a three-dimensional structure of the molecular complex to generate a plurality of possible three-dimensional structures of the molecular complex, and identifying, for inclusion in the structural ensemble, one or more possible three-dimensional structures of the plurality of possible three-dimensional structures whose energy satisfies one or more criteria.
[0032] In some variations, a feature computation model is applied to determine the one or more interface features of the intermolecular interface between the molecule design and the binding target. The feature computation model outputs a feature vector including one or more numerical values corresponding to the one or more interface features. A property computation model is applied to determine, based at least on the feature vector, the property of the molecule design.
[0033] In some variations, the property of the molecule design includes a binding affinity and / or a binding specificity between the molecule design and the binding target.
[0034] In some variations, the molecule design comprises a biopolymer and the binding target comprises a small molecule or another biopolymer.
[0035] In some variations, the molecule design comprises a protein molecule and the binding target is another protein molecule.
[0036] In some variations, the molecule design comprises an antibody and the binding target comprises an antigen.
[0037] Implementations of the current subject matter can include, but are not limited to, methods consistent with the descriptions provided herein as well as articles that comprise a tangibly embodied machine-readable medium operable to cause one or more machines (e.g., computers, etc.) to result in operations implementing one or more of the described features. Similarly, computer systems are also described that may include one or more processors and one or more memories coupled to the one or more processors. A memory, which can include a non- transitory computer-readable or machine-readable storage medium, may include, encode, store, or the like one or more programs that cause one or more processors to perform one or more of the operations described herein. Computer implemented methods consistent with one or more implementations of the current subject matter can be implemented by one or more data processors residing in a single computing system or multiple computing systems. Such multiple computing systems can be connected and can exchange data and / or commands or other instructions or the like via one or more connections, including, for example, to a connection over a network (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, or the like), via a direct connection between one or more of the multiple computing systems, etc.
[0038] The details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims. While certain features of the currently disclosed subject matter are described for illustrative purposes in relation to the binding interactions between a polymer and a monomer, it should be readily understood that such features are not intended to be limiting. The claims that follow this disclosure are intended to define the scope of the protected subject matter.DESCRIPTION OF THE DRAWINGS
[0039] The accompanying drawings, which are incorporated in and constitute a part of this specification, show certain aspects of the subject matter disclosed herein and, together with the description, help explain some of the principles associated with the subject matter disclosed herein. The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawings will be provided by the Office upon request and payment of the necessary fee. In the drawings,
[0040] FIG. 1 depicts a system diagram illustrating an example of a molecule design system, in accordance with some example embodiments;
[0041] FIG. 2 depicts a flowchart illustrating an example of a process for computational molecular analytics using interface features, in accordance with some example embodiments;
[0042] FIG. 3 A depicts a flowchart illustrating an example of a process for computational intermolecular interface analytics, in accordance with some example embodiments;
[0043] FIG. 3B depicts a flowchart illustrating another example of a process for computational intermolecular interface analytics, in accordance with some example embodiments;
[0044] FIG. 4A depicts a screenshot illustrating an example selection of interface features of an intermolecular interface between a molecule design and a binding target, in accordance with some example embodiments;
[0045] FIG. 4B depicts graphs illustrating the distributions in the values of different interface features across multiple antibody-antigen interfaces and the density of each value, in accordance with some example embodiments;
[0046] FIG. 5A depicts graphs illustrating the outcome of filter-based approaches to identifying binders as applied to known binders, in accordance with some example embodiments;
[0047] FIG. 5B depicts a graph illustrating the outcome of filter-based approaches to identifying binders as applied to known binders, in accordance with some example embodiments;
[0048] FIG. 5C depicts a graphs illustrating the distributions of a property for binders that are antibody-antigen and general protein-to-protein interfaces as well as graphs illustrating the distribution of the same property for two different sets of designed antibodies that are variations of two different seed molecules in accordance with some example embodiments;
[0049] FIG. 5D depicts a graph illustrating the distributions of a conformity score computed using a selection of interface features having varying significance assigned by a property computation model, in accordance with some example embodiments;
[0050] FIG. 5E depicts the varying significance of an example selection of interface features in differentiating between binders and nonbinders as determined by a property computation model, in accordance with some example embodiments;
[0051] FIG. 6A depicts a decision tree illustrating an example of a classification hierarchy resulting from a property computation model being implemented with a random forest classifier architecture, in accordance with some example embodiments;
[0052] FIG. 6B depicts a graph illustrating a receiver operating characteristic (ROC) curve of an example of a property computation model implemented with a random forest classifier architecture, in accordance with some example embodiments;
[0053] FIG. 6C depicts a confusion matrix of an example of a property computation model implemented with a random forest classifier architecture, in accordance with some example embodiments;
[0054] FIG. 7A depicts a graph illustrating the weights from a linear layer of a property computation model implemented with a logistic regression classifier architecture, in accordance with some example embodiments;
[0055] FIG. 7B depicts a confusion matrix of a property computation model implemented with a logistic regression classifier architecture, in accordance with some example embodiments;
[0056] FIG. 8A depicts a graph illustrating the distribution of minimum, average, and median contact interaction energy, in accordance with some example embodiments;
[0057] FIG. 8B depicts a graph illustrating the average quantity of hotspot contacts across the complementarity determining regions (CDR) of various antibodies, in accordance with some example embodiments;
[0058] FIG. 9A depicts a graph illustrating the relationship between proportion of contacts and the interaction energy threshold used to identify contacts, in accordance with some example embodiments;
[0059] FIG. 9B depicts a graph illustrating the relationship between proportion of contacts and the interaction energy threshold used to identify contacts, in accordance with some example embodiments;
[0060] FIG. 10A depicts graphs illustrating a comparison of the distribution of interface features assigned more positive weights and the distribution of interface features assigned more negative weights by a linear regressor property computation model, in accordance with some example embodiments;
[0061] FIG. 10B depicts graphs illustrating the shift that is present in the distributions of different interface features for binders and nonbinders, in accordance with some example embodiments;
[0062] FIG. 11 depicts schematic diagrams illustrating examples of contacts for which the interaction energy of the constituent pair of interacting monomers is differentiated between total interaction energy and sidechain-specific interaction energy, in accordance with some example embodiments;
[0063] FIG. 12A depicts a heatmap illustrating the amino acid residue pairing preferences present at the intermolecular interface between antibodies and antigens, in accordance with some example embodiments;
[0064] FIG. 12B depicts a graph illustrating the amino acid residue pairing preferences for intramolecular interactions and intermolecular interactions between antibodies and antigens, in accordance with some example embodiments;
[0065] FIG. 12C depicts schematic diagrams illustrating examples of interaction motifs present at the intermolecular interface between antibodies and antigens, in accordance with some example embodiments;
[0066] FIG. 12D depicts a schematic diagram illustrating a comparison of interaction motifs identified within an intermolecular interface between a molecule design and a binding target and similar interaction motifs mined from a dataset of known structures, in accordance with some example embodiments;
[0067] FIG. 13 A depicts a schematic diagram illustrating additional examples of structure-determining contacts in a complementarity determining region (CDR) of another antibody, in accordance with some example embodiments;
[0068] FIG. 13B depicts a heatmap illustrating the prevalence of sidechain-specific hydrogen-bonds between contacts based on the respective locations of the interacting amino acid residues, in accordance with some example embodiments;
[0069] FIG. 13C depicts schematic diagrams illustrating the different functional roles of hydrogen bonds, in accordance with some example embodiments;
[0070] FIG. 13D depicts a visualization of examples of critical contacts and the corresponding interaction energies in a seed complex strucutre, in accordance with some example embodiments;
[0071] FIG. 14A depicts a comparison of the conventional distance-based metrics and the contact energy-based interface feature for quantifying an intermolecular interface between a molecule design and a binding target, in accordance with some example embodiments;
[0072] FIG. 14B depicts a heatmap illustrating the quantity of sidechain-specific hydrogen-bond driven contacts across different regions of an antibody, in accordance with some example embodiments;
[0073] FIG. 14C depicts another heatmap illustrating the quantity of backbone- mediated hydrogen-bond driven contacts across different regions of an antibody, in accordance with some example embodiments;
[0074] FIG. 14D depicts a network graph illustrating an example of a hydrogen-bond network formed by contacts with interacting amino acid residues in one or more complimentarity determining regions (CDRs) of an antibody, in accordance with some example embodiments;
[0075] FIG. 15 depicts a graph illustrating the frequency with which each canonical amino acid residues is part of a contact and part of a hotspot contact, in accordance with some example embodiments; and
[0076] FIG. 16 depicts a graph illustrating the cosine similarity of the feature vectors of possible non-redundant non-self-pairs of antibody-antigen interfaces across different antibodies, in accordance with some example embodiments;
[0077] FIG. 17 depicts a block diagram illustrating an example of a computing system, in accordance with some example embodiments.
[0078] When practical, similar reference numbers are used to refer to same or similar items in the drawings.DETAILED DESCRIPTION
[0079] Biopolymers are large molecules (or macromolecules) naturally produced by living organisms. One example of a biopolymer is a protein molecule, which may include one or more polypeptide chains formed by amino acid residues linked by peptide bonds. Biopolymers, such as protein molecules, may modulate a multitude of biological processes including, forexample, enzymatic reactions, molecular transport, biological pathway regulation and execution, cell growth and proliferation, nutrient uptake, morphology, motility, intercellular communication, and / or the like. In some cases, a biopolymer may modulate a biological process by interacting with another molecule, such as another biopolymer. For example, the interaction between antibodies and antigens (e.g., viral antigens, tumor antigens, and / or the like) is a critical part of a body’s immune response in which exposure to antigens triggers the production of antibodies to bind to and neutralize the antigens. In the case of therapeutic proteins, such as antibodies, chimeric antigen receptors (CARs), enzymes, hormones, and cytokines, the specificity of the interaction between a therapeutic protein and its therapeutic binding target may be leveraged for the diagnosis and treatment of a wide range of diseases with fewer side effects. Accordingly, one primary objective of large molecule drug discovery (LMDD) is to engineer therapeutic proteins capable of binding to therapeutic binding targets with sufficient affinity and specificity. Deciphering the dynamics of intermolecular interactions, which influence the binding affinity and specificity between a therapeutic protein and its binding target, is therefore an essential endeavor in the context of large molecule drug discovery (LMDD).
[0080] A typical drug discovery workflow includes two high-level phases: lead identification followed by lead optimization (LO). The objective of lead identification (LI) is to identify lead molecules, which in the context of large molecule drug discovery (LMDD) would be protein molecules exhibiting at least some level of fitness as a protein therapeutic. For example, in some cases, a lead molecule may be an antibody identified through an animal immunization campaign as having one or more desirable properties, such as binding affinity towards an antigen (e.g., a viral antigen, a tumor antigen, and / or the like). However, despite having the one or more desirable properties, the lead molecule in its original form is unlikely tobecome a viable protein therapeutic due to the presence of one or more undesirable properties, such as insufficient human-ness, poor expression, immunogenicity, in vivo instability, and / or the like. To increase the likelihood that further drug development efforts result in a viable protein therapeutic, the lead molecule may therefore undergo lead optimization (LO), which includes modifying the underlying sequence of amino acid residues, for example, by inserting, deleting, and / or changing the type of one or more constituent amino acid residues, such that the resulting molecules exhibit better properties than the lead molecule.
[0081] To accelerate drug development and reduce reliance on expensive wet lab resources, a suite of computation tools have been developed for tasks such as in silica protein design and virtual screening. For example, a protein design computation model may be trained to generate protein molecules de novo or, alternatively, modify a lead molecule to optimize one or more properties of the lead molecule. Nevertheless, conventional computational tools, including those leveraging artificial intelligence and machine learning, are thwarted by the scale and complexity of designing and optimizing protein molecules capable of being developed into safe and effective therapeutics. For instance, the solution space of every possible protein sequences is vast but sparsely populated with protein sequence with the requisite desirable properties. Of the 20wpossible protein sequences for protein sequences with an N quantity of amino acid residues, a disproportionately small proportion of those 20wpossible protein sequences exhibit the requisite desirable properties. An even smaller proportion of the 20wpossible protein sequences have empirical data corroborating the presence (or absence) of desirable properties such as binding affinity, binding specificity, biological activity, developability, and / or the like.
[0082] To the extent sufficient empirical data is available to train a protein design computation model to achieve adequate performance when modifying the sequence of a lead molecule to optimize one or more properties thereof, the available empirical data is inadequate to train the protein design computation model for the de novo generation of protein molecules. In the latter case, no lead molecules (or known sequences of amino acid residues) with desirable properties (e.g., from animal immunization campaigns) may exist. Thus, there is desire for a protein design computation model that can generate novel protein sequences from scratch based on the target molecule alone. Given the current low data regime in which empirical data on protein sequences known to bind to the target molecule is scarce, the task of generating protein sequences capable of binding to target molecule with sufficient affinity and specificity de novo is infeasible with current state-of-the-art methodologies. Notably, the dearth of empirical data renders the properties of de novo generated protein sequences difficult to accurately assess in silico, which forestalls efforts to virtually screen such protein sequences for those with a greater likelihood of developing into viable protein therapeutics. Intermolecular interfaces are typically defined based on the two “sides” of interacting molecules (e.g. chains “A” and “C” in a biopolymer complex such as a protein complex). In some cases, an intermolecular interface may be be defined based on particular expectations and / or algorithmically, such as knowing than an antibody is the “binder” and the antigen is the “target”. In some cases, interface residues may be identified by distance. For example, in some cases, interface residues may be identified as the alpha-carbon residues on side 1 within 1 nanometer of at least one alpha-carbon residue on side 2 (or vice versa). Alternatively, as another example, interface residues may be identified as residues for which at least one heavy atom on side 1 is within 4.5-5.0 A of at least one heavy atom on side 2 (or vice versa). Another alternative approach for identifying interface residuesmay be based on the change in solvent-accessible surface area (SASA) upon complexation. In such a method, the total solvent-accessible surface area (SASA) may be computed for the complex, then the two sides of the complex are separated by a large distance (e.g., in an in silica simulation) and the total solvent-accessible surface area (SASA) is recomputed. The difference in these two values may indicate how much protein surface area is buried upon complexation, which can be one indicator of binding affinity and specificity (e.g., larger interfaces can form more interactions and these more interactions all requiring proper satisfaction helps enforce specificity). It is possible to assign how much solvent-accessible surface area (SASA) changes per-residue before one or more thresholds are set to determine how much change in solvent- accessible surface area (SASA) is required to qualify a residue as being an interface residue (e.g., any residue with a >5 A2SASA difference upon complexation is in the interface).[0831 Various embodiments of the present disclosure overcome the limitations of existing drug design protocols, including computational methodologies leveraging artificial intelligence (Al) and machine learning (ML), by evaluating one or more properties of a biopolymer, such as a protein molecule, based on one or more interface features of the intermolecular interface between the biopolymer and a binding target (e.g., another biopolymer). As used herein, the term “intermolecular interface” may refer to the respective portions of two molecules (e.g., two biopolymers) engaged in a binding interaction between the two molecules. In some cases, the intermolecular interface may be between two protein molecules. Where the intermolecular interface is between an antigen and an antibody, for example, the intermolecular interface may include the epitope of the antigen and the paratope (or complementarity determining regions (CDRs)) of the antibody. In some cases, the one or more interface features may be used to evaluate one or more properties of a biopolymer generated computationally by amolecule design computation model. For instance, in some cases, the molecule design computation model may be applied to generate the monomer sequence of a biopolymer (e.g., sequence of amino acid residues in instances where the biopolymer is a protein molecule) de novo or, alternatively, by modifying the monomer sequence of another biopolymer (e.g., a lead molecule identified through an animal immunization campaign). In the context of computational drug design, the molecule design computation model may be applied to generate the protein sequence of a potential therapeutic protein (e.g., an antibody, a chimeric antigen receptor (CAR), an enzyme, a hormone, a cytokine, and / or the like). In some cases, the viability of the protein sequence as a therapeutic protein, and therefore its suitability for further development as such, may be contingent on the protein sequence exhibiting certain properties. Where the protein sequences is an antibody, for example, the protein sequence should exhibit sufficient binding affinity (or strength of the binding interaction) towards an antigen (e.g., a viral antigen, a tumor antigen, and / or the like).
[0084] In some example embodiments, the properties of a biopolymer generated by the molecule design computation model, such as the binding affinity of a protein molecule towards a binding target (e.g., another biopolymer), may be assessed based on the one or more interface features of the intermolecular interface between the biopolymer and the binding target. For example, in instances where the biopolymer is an antibody, the binding affinity of the antibody towards an antigen (e.g., viral antigen, tumor antigen, and / or the like) may be determined based on one or more interface features of the intermolecular interface between the antibody and the antigen. As described in more details below, in some cases, these interface features may include biophysical features such as energetic features and geometric features, which are independent of the monomer sequence forming the biopolymer (e.g., the sequence ofamino acid residue forming the protein molecule). That the properties of the biopolymer (e.g., protein molecule) may be determined in silica independent of the underlying monomer sequence (e.g., protein sequence) may be especially advantageous in a low data regime in which the available empirical data is insufficient to support reliable monomer sequence-based virtual screening of the biopolymer. That is, in instances where the properties of the biopolymer cannot be accurately determined due to its monomer sequence being outside of the distribution of the empirical data available to train a conventional property computation model focused on the monomer sequence alone, the properties of the biopolymer can still be accurately determined based on interface features. For instance, even when the monomer sequence of the biopolymer is out-of-distribution (OOD), the interface features may still be within the distribution of known intermolecular interfaces, such as protein-to-protein interfaces (PPIs), for which more empirical data may be available to improve the performance of a property computation model operating on interface features to determine the properties of the biopolymer.
[0085] In some example embodiments, a feature computation model may be applied to determine one or more interface features of the intermolecular interface between a biopolymer (e.g., a protein molecule) and a binding target (e.g., a small molecule or another biopolymer). In some cases, the feature computation model may compute interface features, such as geometric features and energetic features, that are non-trivial and impossible to extract from static molecular structures. Moreover, the one or more interface features may be interpretable, with well-defined distributions (of values) that differ between biopolymers (e g., protein molecules) exhibiting certain properties (e.g., binding affinity towards a particular antigen) and biopolymers that fail to exhibit the properties. As such, the correlations between the one or more interface features and the properties of the biopolymer (e g., binding affinity of the protein molecule)computed therefrom may be observable and explicable. Accordingly, as described in more details below, the one or more interface features may enable the structural bioinformatic analysis of biopolymers along novel dimensions inaccessible to conventional distance or angle-based descriptors of the intermolecular interface computed from atomic coordinates alone.
[0086] In some example embodiments, the feature computation model may determine one or more interface features of the intermolecular interface between a biopolymer (e.g., a protein molecule) and a binding target (e.g., a small molecule or another biopolymer) by performing conformational sampling and energy minimization. For example, in some cases, the feature computation model may generate, for a biopolymer, a structural ensemble (or conformational ensemble) including multiple possible three-dimensional structures (or conformations) that can be adopted by the biopolymer when bound to the binding target. In some cases, the feature computation model may determine the possible three-dimensional structures (or conformations) by performing Monte Carlo sampling, which may include iteratively adjusting the three-dimensional structure (or conformation) of the biopolymer to identify those three-dimensional structures (or conformations) whose energy satisfies one or more criteria. For instance, in some cases, a possible three-dimensional structure (or conformation) may be added to the structural ensemble (or conformational ensemble) if the energy of the three-dimensional structure (or conformation) satisfies one or more thresholds. Alternatively and / or additionally, a possible three-dimensional structure (or conformation) may be added to the structural ensemble (or conformational ensemble) if that three-dimensional structure (or conformation) is one of an N quantity of the three-dimensional structures (or conformations) sampled with the lowest energy.
[0087] In some example embodiments, one or more interface features may be determined based on the possible three-dimensional structures (or conformations) of the biopolymer bound to the binding target within the structural ensemble (or conformational ensemble). It should be appreciated that the structural ensemble (or conformational ensemble) may reflect the flexible nature of biopolymers (e.g., protein molecules) in solution, which typically lack a fixed tertiary structure in the absence of an interaction partner such as a binding target (e.g., a small molecule or another biopolymer). Thus, instead of extracting features from a single static three-dimensional structure (or conformation), which yields merely an average across a structural ensemble of low temperature three-dimensional structures (or conformations), various implementations of the present disclosure extract features from multiple independent three-dimensional structures (or conformations) to better capture the natural structural variability of biopolymers in solution. As such, features extracted from individual static three-dimensional structures are an unreliable approximation whose accuracy can fluctuate significantly across different categories of three-dimensional structures (or conformations). Such inaccuracies are reduced by the feature computation model computing one or more interface features as a distribution across multiple possible three-dimensional structures (or conformations) within the structural ensemble (or conformational ensemble).
[0088] In some example embodiments, a property computation model may be trained to determine, based at least on the one or more interface features of the intermolecular interface between a biopolymer (e.g., a protein molecule) and a binding target (e.g., a small molecule or another biopolymer), the binding affinity between the biopolymer and the binding target. In some cases, the feature computation model may perform conformational sampling and energy minimization in order determine the one or more interface features of the intermolecularinterface between the biopolymer and the binding target. For example, in some cases, the feature computation model may output a feature vector containing one or more numerical values corresponding to the one or more interface features. In some cases, the property computation model may operate on the feature vector to perform one or more predictive tasks. For instance, in some cases, the property computation model may perform a classification task and generate an output identifying the biopolymer as a binder (or nonbinder) of the binding target. Where the biopolymer is an antibody and the binding target is an antigen (e.g., a viral antigen, a tumor antigen, and / or the like), for example, the output of the property computation model may indicate whether the antibody is a binder (or nonbinder) of the antigen. Alternatively and / or additionally, the property computation model may perform a regression task and generate an output indicating the binding affinity (or strength of the binding interaction) between the biopolymer (e.g., antibody) and the binding target (e.g., antigen). It should be appreciated that various embodiments of the property computation model disclosed herein determines the properties of the biopolymer based on its interface properties rather than the monomer sequence (e.g., protein sequence) of the biopolymer. This independence from the monomer sequence of the biopolymer may be especially advantageous in a low data regime where little to no empirical data is available to train the property computation model to reliably determine the properties of the biopolymer based on its monomer sequence.
[0089] In some example embodiments, the one or more interface features of the intermolecular interface between a biopolymer (e.g., a protein molecule) and a binding target (e.g., a small molecule or another biopolymer) may include one or more geometric features. Examples of the one or more geometric features may include metrics derived from and / or quantifying interface shape complementarity, hydrogen bond satisfaction, interface packingquality, entropic penalty of sidechain burial upon complexion, and / or the like. In some cases, the one or more geometric features may quantify the quality of the intermolecular interface. In this context, a higher quality intermolecular interface may support a stronger and more stable binding interaction between the biopolymer and its binding target. Conversely, a lower quality intermolecular interface may engender a weaker and less stable binding interaction between the biopolymer and its binding target. Accordingly, it should be appreciated that the one or more geometric features may correlate with the binding affinity (or the strength of the binding interaction) between the biopolymer and its binding target. Moreover, as described in more details below, the property computation model may determine, based at least on the one or more geometric features of the intermolecular interface, the binding affinity between the biopolymer and its binding target.[0901 In some example embodiments, the one or more interface features of the intermolecular interface between a biopolymer (e.g., a protein molecule) and a binding target (e g., a small molecule or another biopolymer) may include one or more energetic features. In some cases, the one or more energetic features may be derived from and / or quantify the interatomic forces driving the binding interaction between the biopolymer (e.g., generated by the molecule design computation model) and its binding target (e.g., another biopolymer). For example, in some cases, the one or more energetic features may be capture a variety of interatomic forces, such as electrostatic forces, hydrogen bonding, van der Waals forces, hydrophobic effects, and / or the like. In some cases, the one or more geometric features may correlate with the binding affinity (or the strength of the binding interaction) between the biopolymer and its binding target. Moreover, in some cases, the one or more geometric features may be quantified on a continuous scale that explicitly captures atomic forces from atomiccoordinates, thus allowing the contact between two monomers (e.g., amino acid residues in the biopolymer and the binding target) to be defined based on physical properties rather than strict thresholds subject to experimental accuracy and uncertainty. Accordingly, as described in more details below, the property computation model may determine, based at least on the one or more energetic features of the intermolecular interface, the binding affinity between the biopolymer and its binding target.
[0091] FIG. 1 depicts a system diagram illustrating an example of a molecule design system 100, in accordance with some example embodiments. Referring to FIG. 1, the molecule design system 100 may include a featurization engine 110, a molecule design engine 120, a molecule analysis engine 130, and a client device 140 with a user interface 145. In the example of the molecule design system 100 shown in FIG. 1, the featurization engine 110, the molecule design engine 120, the molecule analysis engine 130, and the client device 140 may be communicatively coupled via a network 150. In some cases, the client device 140 may be a processor-based device including, for example, a workstation, a desktop computer, a laptop computer, a high-performance computer, a smartphone, a tablet computer, a wearable apparatus, and / or the like. In some cases, the network 150 may be a wired network and / or a wireless network including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, and / or the like.
[0092] In some example embodiments, the featurization engine 1 10 may include a feature computation model 115 that determines one or more interface features of an intermolecular interface between two molecules, such as a biopolymer (e.g., a protein molecule) and a binding target (e.g., a small molecule or another biopolymer). For example, in some cases,the feature computation model 115 may be applied to determine one or more interface features of an intermolecular interface between a molecule design 112 and a binding target 114. In some cases, the molecule design 112 may be generated by a molecule design computation model 125 at the molecule design engine 120. For instance, in some cases, the molecule design computation model 125 may be applied to generate the molecule design 112 based on the binding target 114 alone in instances the molecule design computation model 125 is applied to generate the molecule design 112 de novo. Alternatively, the molecule design computation model 125 may generate the molecule design 112 by at least modifying a lead molecule 122.
[0093] In some cases, the molecule design computation model 125 may generate the molecule design 112 by at least determining the protein sequence (or primary structure) forming the molecule design 112, which may include identifying the type of each of monomer (e.g., amino acid residues for a protein molecule) in the monomer sequence of the molecule design 112. In some cases, the molecule design computation model 125 may jointly determine the identity (or type) of each monomer and the three-dimensional structure (or conformation) adopted by the corresponding monomer sequence (e.g., sequence of amino acid residues for a protein molecule), which may include the atomic coordinates (e.g., Cartesian coordinates (x, y, z)) of the atoms (e.g., heavy atoms) forming each constituent monomer. Alternatively, in some cases, the molecule design computation model 125 may first determine the monomer sequence (e.g., the identify (or type) of each monomer) of the molecule design 112 before determining the three-dimensional structure (or conformation) adopted by the monomer sequence (e.g., the atomic coordinates (e.g., Cartesian coordinates (x,y, z)) of the atoms (e.g., heavy atoms) forming each constituent monomer).
[0094] In the example shown in FIG. 1, the feature computation model 1 15 may operate on the atomic coordinates (e.g., Cartesian coordinates (x, y, z)) of one or more atoms (e.g., heavy atoms) forming the molecule design 112 and the binding target 114. For example, in some cases, the featurization engine 110 may receive, from the molecule design engine 120, the atomic coordinates (e.g., Cartesian coordinates (x,y, z)) of the one or more atoms (e.g., heavy atoms) forming the molecule design 112 and the binding target 114. In some cases, in addition to the atomic coordinates (e.g., Cartesian coordinates (x, y, z)) of the one or more atoms (e.g., heavy atoms) forming the molecule design 112 and the binding target 114, the featurization engine 110 may also receive information identifying the identity (or type) of each atom (e.g., heavy atom) and / or the monomers (e.g., amino acid residues) formed therefrom. In some cases, the feature computation model 115 may determine, based at least on the atomic coordinates (e.g., Cartesian coordinates (x, y, z)) of one or more atoms (e.g., heavy atoms) forming the molecule design 112 and the binding target 114, one or more interface features 116 of the intermolecular interface between the molecule design 112 and the binding target 114. For instance, in some cases, the feature computation model 115 may generate a vector containing one or more values corresponding to the one or more interface features 116 of the intermolecular interface between the molecule design 112 and the binding target 114. As described in more details below, the feature computation model 115 may perform conformational sampling and energy minimization in order to determine the one or more interface features 116. Moreover, in some cases, the one or more interface features 116 may include one or more geometric features and energetic features which, unlike conventional angle- and distance-based descriptors of intermolecular interfaces, are non-trivial and impossible to extract from static molecular structures.
[0095] In some example embodiments, the one or more interface features 116 may include interface features, such as solvent-accessible surface area (SASA), determined based on the coordinates and atomic composition of the molecule design 112. Alternatively and / or additionally, in some cases, the one or more interface features 116 may interface features corresponding to or determined based on contacts present in the intermolecular interface between the molecule design 112 and the binding target 114. As used herein, the term “contact” may refer to a pair of contacting or interacting monomers (e.g., amino acid residues). In some cases, a contact (or monomer pair) may include one monomer from the molecule design 112 and another monomer from the binding target 114. As described in more details below, a contact may be associated with a functional role. For instance, in some cases, a contact between the molecule design 112 and the binding target 114 may contribute to the binding affinity (or the strength of the binding interaction) between the molecule design 112 and the binding target 114 as well as the binding specificity of the molecule design 112 towards the binding target 114. Alternatively, it is also possible for a contact (or monomer pair) to include two monomers from the molecule design 112. In some cases, a contact within the molecule design 112 may determine and stabilize at least a portion of the three-dimensional structure adopted by the molecule design 112. For instance, in instances where the molecule design 112 is an antibody, a contact that includes two amino acid residues from a complementarity determining region (CDR) of the molecule design 112, such as those interacting through hydrogen bonding (e.g., to form a hydrogen-bond network), may determine and / or stabilize the three-dimensional structure (or conformation) adopted by the complementarity determining region (CDR) loops of the antibody. Similarly, a contact that includes an amino acid residue from the complementarity determining region (CDR) and another amino acid residue from the framework region may also determineand / or stabilize the three-dimensional structure (or conformation) adopted by the complementarity determining region (CDR) loops of the antibody.
[0096] Conventional approaches for identifying contacts (or interacting monomers) rely on atomic distance or angle-based metrics (e.g., distance or angle between the atoms in each monomer), which disregard key biophysical attributes such as interaction strength and the nature of the interaction (e.g., affinity versus specificity). The interaction between spatially proximate monomers (e.g., amino acid residues), for example, do not necessarily interact to an extent necessary to determine the three-dimensional structure (or conformation) of the molecule design 112 or its binding affinity or specificity towards the binding target 114. As such, identifying contacts (or interacting monomers) based on spatial proximity alone lacks structural context, rendering the results insufficiently reliable for downstream analytics. Contrastingly, various embodiments of the feature computation model 115 described herein may classify a pair of monomers as contacts based on the interaction energy (e.g., magnitude of the interaction energy) between the two constituent monomers. In some cases, the monomer pair may be classified as contacts based on a continuous scale of interaction energy that explicitly captures atomic forces from the spatial coordinates (e.g., Cartesian coordinates (x, y, z)) of the atoms (e.g., heavy atoms) in each constituent monomer. For example, energy scaling plots demonstrate that the small subset of monomers (e.g., amino acid residues) at intermolecular interfaces that meet the “contact” energy threshold may represent significant proportion (e.g., 95-97%) of the total energy across the intermolecular interface. Accordingly, instead of the monomer pair being classified as contacts based on conventional metrics (e.g., atomic distance-based thresholds), which are subject to experimental accuracy and uncertainty, the feature computation model 115 may define contacts based on biophysical properties such as interaction energy (e.g., magnitudeof interaction energy). In doing so, the feature computation model 115 may differentiate between “hotspot” contacts, which are pairs of interacting monomers contributing significantly more to binding affinity (e.g., towards the binding target 114) compared to typical (or nonhotspot) contacts, which are weaker but more numerous and are important for binding specificity.
[0097] In some cases, a “hotspot” contact may be a contact with a total binding energy satisfying one or more thresholds (e.g., <-6). It should be appreciated that the definition of a contact and a hotspot contact may differ slightly. For example, a contact may be any residue that forms an interaction of at least -0.6 in total binding energy to a single other residue. Thus, a residue that forms interactions with 3 target residues with respective binding energies of -0.2, - 0.3, -0.1 is not a “contact” despite the total binding energy being -0.6 because those individual interactions are of “nonspecific” type interactions that contribute to overall binding but not through particular interactions to particular residues. Contrastingly, a “hotspot” contact may refer to any contact with total binding energy of < -6.0 to all residues on the binding target. This “hotspot” residue may have some interactions with target residues that are individually smaller than -0.6, though typically specific target “contacts” comprise the majority of the total “hotspot” energy, as the large magnitude of “hotspot” binding energies is commonly achieved through their simultaneous interaction with several target residues rather than a single particularly powerful pair interaction. In other words, a hotspot contact may be any contact that is >N-fold stronger than a non-hotspot contact. This definition differentiates a hotspot contact as having an order lOx stronger impact than a typical, non-hotspot contact.
[0098] In some example embodiments, the feature computation model 115 may determine, for each contact, one or more functional roles. For example, in some cases, a contactmay be identified as neutral or, alternatively, as contributing to the three-dimensional structure (or conformation), binding affinity, and / or binding specificity of the molecule design 112. As noted, in some cases, the functional role of a contact may be determined by the respective locations of the constituent monomers (e.g., amino acid residues). In some cases, the functional role of the contact may be further determined based on the types of interaction energies (or energy terms) present between the constituent monomers (e.g., amino acid residues). For instance, in some cases, the feature computation model 115 may identify contacts based on specific interaction energies (e.g., magnitude of specific interaction energies), which arise from interatomic forces critical for binding specificity such as hydrogen-bonding, electrostatics, and nonpolar van der Waals forces. In some cases, a single monomer pair may be classified as contacts (or noncontacts) based on individual specific interaction energies or combinations of two or more such specific interaction energies. Moreover, monomer pairs classified as contacts (or noncontacts) based on specific interaction energies may be differentiated from those classified as contacts (or noncontacts) based on nonspecific interaction energies (e.g., magnitude of nonspecific interaction energies) such as solvation interactions. In some cases, the interaction energy for identifying contacts within the intermolecular interface between the molecule design 112 and the binding target 114 may exclude contributions from those arising from certain types of interatomic forces. Where downstream analytics include determining the binding affinity of the molecule design 112, for example, the interaction energy may exclude those arising from nonspecific interaction energies, such as the aforementioned solvation interactions.
[0099] In some example embodiments, the feature computation model 115 may generate, based at least on the contacts in the intermolecular interface between the molecule design 112 and the binding target 114, the one or more interface features 116 of theintermolecular interface as a whole. For example, in some cases, the one or more interface features 116 may include one or more of a total quantity, an average density, a distribution over specific intermolecular interface regions (e.g., individual complementarity determining region (CDR) loops), an aggregate energy value (e.g., minimum, median, maximum, mode, range, and / or the like), and a summary statistic of energy distribution (e.g., skew, kurtosis, and / or the like) of the contacts identified within the intermolecular interface between the molecule design 112 and the binding target 114. In some cases, the feature computation model 115 may differentiate contacts based on the source of the interaction energies in the molecule design 112. For instance, where the molecule design 112 is a biopolymer, such as a protein molecule, the interaction energy of a contact (e g., the pair of interacting amino acid residues) at the intermolecular interface may include its total interaction energy as well as sidechain-specific interaction energy and backbone-mediated interaction energy. In this context, sidechain-specific interaction energy may arise from the chemical properties of the monomer sidechains (e.g., amino acid sidechains (or R-groups)) that protrude from the backbone of the monomer. The sidechains present in the molecule design 112 may engage in interactions, such as hydrogen bonds, hydrophobic interactions, and ionic interactions, that determine the three-dimensional structure (or conformation) as well as the function of the molecule design 112. Meanwhile, backbone-mediated interaction energies arise from the chemical properties of the backbone of the monomer. Those interactions, which may include hydrogen bonds between the atoms in the backbone, may determine the secondary structure (e.g., alpha helices, beta sheets, and / or the like) of the molecule design 112, in instances where the molecule design 112 is a biopolymer (e.g., a protein molecule).
[0100] In some example embodiments, the location of the contact may also determine its functional role. For example, in instances where the molecule design is a protein molecule, such as an antibody, a contact in which one amino acid residue is in the molecule design and the other amino acid residue is in the binding target may contribute to the affinity and / or specificity of the molecule design. A contact with two interacting amino acid residues (e.g., forming a hydrogen-bond network) in the complementarity determining regions (CDRs) of the molecule design may determine and / or stabilize the three-dimensional structure (or conformation) of the complementarity determining region (CDR) loops of the molecule design. A contact with one interacting amino acid residue in the complementarity determining region (CDR) and another amino acid residue in the framework region may also determine and / or stabilize the three- dimensional structure (or conformation) adopted by the complementarity determining region (CDR) loops of the molecule design.
[0101] In some example embodiments, the feature computation model 115 may identify, based at least on the contacts in the intermolecular interface between the molecule design 112 and the binding target 114, the presence (or absence) of one or more interaction motifs. In this context, an “interaction motif’ may refer to a recurring pattern of monomers (e.g., amino acid residues) observed in the intermolecular interfaces of molecular complexes in which, for example, a biopolymer is bound to a binding target. In some cases, an interaction motif may include one or more contacts, each of which being a pair of interacting monomers, that are commonly observed in the intermolecular interfaces of molecular complexes. In some cases, the feature computation model 115 may identify, based at least on the contacts in the intermolecular interface between the molecule design 112 and the binding target 114, the presence (or absence) of one or more favorable interaction motifs associated with a higher quality intermolecularinterface for supporting a stronger and more stable binding interaction between the molecule design 112 and the binding target 114. For example, in some cases, a favorable interaction motif may be one that is present in more than a threshold quantity of intermolecular interfaces between molecules exhibiting one or more desirable properties (e.g., sufficient binding affinity towards the binding target 114). Alternatively and / or additionally, the feature computation model 115 may identify, based at least on the contacts in the intermolecular interface between the molecule design 112 and the binding target 114, the presence (or absence) of one or more unfavorable interaction motifs associated with a lower quality intermolecular interface that engenders a weaker and less stable binding interaction between the molecule design 112 and the binding target 114. For instance, in some cases, an unfavorable interaction motif may be one that is found in more than a threshold quantity of intermolecular interfaces between molecules failing to exhibit the one or more desirable properties or, alternatively, exhibiting one or more undesirable properties (e.g., insufficient binding affinity towards the binding target 114). In some cases, an unfavorable interaction motif may be a false positive interaction motif, which may be common in computationally generated molecular complexes but rare in complexes observed in nature.
[0102] Referring again to FIG. 1, in some cases, the molecule analysis engine 130 may include a property computation model 135 that determines, based at least on one or more interface features of the intermolecular interface between a biopolymer (e.g., a protein molecule) and a binding target (e.g., a small molecule or another biopolymer), one or more properties of the biopolymer. In the example shown in FIG. 1, the property computation model 135 may determine, based at least on the one or more interface features 116 of the intermolecular interface between the molecule design 112 and the binding target 114, a property prediction 136 including one or more properties of the molecule design 112. For example, in instances where themolecule design 112 corresponds to an antibody (or at least a portion thereof) and the binding target 114 is an antigen (e.g., a viral antigen, a tumor antigen, and / or the like), the property prediction 136 may include the binding affinity of the molecule design 112 towards the binding target 114. In instances where the binding target 114 is the same molecule design 112 (e.g., the same antibody), the property prediction 136 may indicate a propensity of the molecule design 112 for self-association in a solution containing a population of the molecule design 112. In some cases, whether the molecule design 112 advances to a subsequent stage of the drug development process may be determined based on the property prediction 136, such as the molecule design 112 exhibiting sufficiently high binding affinity or low propensity for selfassociation. Alternatively and / or additionally, the property prediction 136 of the molecule design 112 may be leveraged as an annotation (or label) for the molecule design 112 such that the molecule design 112 may serve to augment available training data. In instances where biopolymers are being generated de novo and their corresponding monomer sequences (e.g., protein sequences) are out of the distribution of existing training data, the inclusion of the molecule design 112 as an additional training sample may improve the predictive performance of the computation model trained with the training data.
[0103] FIG. 2 depicts a flowchart illustrating an example of a process 200 for computational molecular analytics using interface features, in accordance with some example embodiments. Referring to FIGS. 1-2, in some example embodiments, the process 200 may be performed by the molecule design system 100 including, for example, the featurization engine 110, the molecule design engine 120, and the molecule analysis system 130. For example, in some cases, the featurization engine 110 may apply the feature computation model 115 to determine the one or more interface features 116 of the interm olecular interface between themolecule design 112 and the binding target 114. In some cases, the molecule design 112 may be generated by the molecule design engine 120 applying the molecule design computation model 125. In some cases, the molecule design computation model 125 may generate the molecule design 112 de novo such that the monomer sequence (e.g., protein sequence) of the molecule design 112 and, in some cases, the corresponding three-dimensional structure (or conformation), are generated based on the binding target 114. Alternatively, the molecule design computation model 125 may generate the molecule design 112 by modifying the lead molecule 122, for example, to improve one or more properties of the lead molecule 122. In some cases, the one or more interface features 116 of the interm olecular interface between the molecule design 112 and the binding target 114 may include one or more geometric features and energetic features. In some cases, the one or more interface features 116 may capture biophysical attributes of the intermolecular interface that are inaccessible to conventional atomic distance- and angle-based metrics. For instance, in some cases, the one or more interface features 116 may include one or more contacts (or pairs of interacting monomers) identified based on interaction energies arising from a variety of interatomic forces as well as the functional role of each contact (e.g., neutral, structure-determining, affinity-determining, specificity-determining, and / or the like).
[0104] At 202, a molecule design including a sequence of monomers is generated. In some example embodiments, a molecule design computation model may be applied to generate the molecule design. For example, in instances where the molecule design is a biopolymer, such as a protein molecule, the molecule design computation model may be trained to generate the monomer sequence (e.g., protein sequence) of the molecule design. In some cases, the molecule design computation model may also generate the three-dimensional structure (or conformation) that is adopted by the monomer sequence (e.g., protein sequence) of the molecule design, whichmay include the atomic coordinates (e.g., Cartesian coordinates (x,y, z)) of the atoms (e.g., heavy atoms) forming each constituent monomer. In some cases, the molecule design computation model may first generate the monomer sequence (e.g., protein sequence) of the molecule design before determining the three-dimensional structure (or conformation) adopted by the monomer sequence of the molecule design. Alternatively, the molecule design computation model may be trained to jointly determine (or co-design) the monomer sequence (e.g., protein sequence) of the molecule design and the corresponding three-dimensional structure (or conformation).
[0105] In some cases, the molecule design computation model may generate the molecule design, including the corresponding monomer sequence and three-dimensional structure (or conformation), de novo, meaning that the molecule design may be generated based on a binding target alone. Alternatively, the molecule design computation model may generate the molecule design by modifying a lead molecule. Where the molecule design is an antibody, for example, the molecule design computation model may modify a lead molecule identified through an animal immunization campaign as having one or more desirable properties, such as binding affinity toward a particular antigen (e g., viral antigen, tumor antigen, and / or the like). In this latter case, the molecule design computation model may be trained to apply modifications, for example, to the monomer sequence (e.g., protein sequence) of the lead molecule, that further improves the one or more desirable properties and / or eliminates one or more undesirable properties.
[0106] At 204, one or more interface features of an intermolecular interface between the molecule design and a binding target is determined. In some example embodiments, a feature computation model may be applied to determine the one or more interface features of theintermolecular interface between the molecule design and the binding target. For example, in some cases, the feature computation model may perform conformational sampling and energy minimization to generate a structural ensemble (or conformational ensemble) including multiple possible three-dimensional structures (or conformations) that can be adopted by the molecule design bound to the binding target. In instances where the molecule design is a biopolymer (e.g., a protein molecule, the structural ensemble (or conformational ensemble) may reflect the flexible nature of biopolymers (e.g., protein molecules) in solution, which typically lack a stable three- dimensional structure (or conformation) in the absence of an interaction partner such as a binding target (e.g., a small molecule or another biopolymer). In some cases, the feature computation model may determine the possible three-dimensional structures (or conformations) by performing Monte Carlo sampling, which may include iteratively adjusting the three- dimensional structure (or conformation) of the molecule design to identify those three- dimensional structures (or conformations) with the lowest energy. In some cases, one or more interface features may be determined based on the possible three-dimensional structures (or conformations) of the molecule design bound to the binding target within the structural ensemble (or conformational ensemble). For instances, in some cases, the one or more interface features may be determined as a distribution across multiple possible three-dimensional structures (or conformations) within the structural ensemble (or conformational ensemble) to reduce the inaccuracies associated with extracting features from individual static three-dimensional structures, which can fluctuate significantly across different categories of three-dimensional structures (or conformations).
[0107] In some example embodiments, the feature computation model may determine, based on the structural ensemble (or conformation ensemble) of the molecule design,the one or more interface features of the intermolecular interface between the molecule design and the binding target. In some cases, the one or more interface features may include geometric features as well as energetic features. In some cases, the one or more interface features may include one or more contacts (or pairs of interacting monomers), which may be further differentiated by their respective functional roles. For example, in some cases, a contact (or a pair of interacting monomers) may be identified based on interaction energies (or magnitude of interaction energies) arising from a variety of interatomic forces, such as electrostatic forces, hydrogen bonding, van der Waals forces, hydrophobic effects, and / or the like. In some cases, a pair of monomers may be identified as a contact (or a pair of interacting monomers) if the interaction energies between the two monomers satisfies one or more thresholds. In some cases, the contact may be associated with one or more functional roles, such as neutral, structuredetermining, affinity-determining, specificity-determining, and / or the like. Moreover, in some cases, the contact may be classified based on the source of the interaction energy between constituent monomers including, for example, total interface energy, sidechain-specific interaction energy, backbone-mediated interaction energy, and / or the like.
[0108] In some example embodiments, the feature computation model 115 may output a feature vector containing one or more values corresponding to the one or more interface features of the intermolecular interface between the molecule design and the binding target. In some cases, the feature vector may include values corresponding to one or more geometric features, which may include metrics quantifying or derived from interface shape complementarity, hydrogen bond satisfaction, interface packing quality, entropic penalty of sidechain burial upon complexion, and / or the like. In some cases, the feature vector may also include values corresponding to one or more energetic features, which may be derived fromand / or quantify various interatomic forces such as electrostatic forces, hydrogen bonding, van der Waals forces, hydrophobic effects, and / or the like. In some cases, the feature vector may also include one or more features derived from the contacts identified within the intermolecular interface. Examples of such features may include one or more of a total quantity, an average density, a distribution over specific intermolecular interface regions (e.g., individual complementarity determining region (CDR) loops), an aggregate energy value (e.g., minimum, median, maximum, mode, range, and / or the like), and a summary statistic of energy distribution (e.g., skew, kurtosis, and / or the like) of the contacts identified within the intermolecular interface between the molecule design and the binding target. In some cases, the feature vector may further include a quantity of interaction motifs, which may be identified based on the contacts present in the intermolecular interface between the molecule design and the binding target. In some cases, the quantity of interaction motifs may include favorable interaction motifs typically found in the intermolecular interface between molecules exhibiting sufficient binding affinity, unfavorable interaction motifs typically found in the intermolecular interface between molecules without sufficient binding affinity, and / or false positive interaction motifs common in computationally generated molecular complexes but rarely observed in nature.
[0109] At 206, a property of the molecule design is determined based at least on the one or more interface features. In some example embodiments, the property computation model may be trained to determine, based at least on the one or more interface features of the intermolecular interface between the molecule design and the binding target, one or more properties of the molecule design. For example, in instances where the molecule design is a biopolymer, such as a protein molecule (e.g., an antibody), the property computation model may determine, based at least on the one or more interface features, the binding affinity of thebiopolymer towards the binding target. In instances where the molecule design and the binding target are the same protein molecule (e.g., antibody), the property computation model may determine, based at least on the one or more interface features, the propensity of the molecule design for self-association in a solution containing a population of the molecule design. In some cases, whether the molecule design advances to a subsequent stage of the drug development process may be determined based on one or more properties of the molecule design determined by the property computation model, such as the molecule design exhibiting sufficiently high binding affinity or low propensity for self-association. Alternatively and / or additionally, the one or more properties of the molecule design determined by the property computation model may be used as an annotation (or label) for the molecule design such that the molecule design may serve to augment the available training data. This training data augmentation may be advantageous in low data regimes where empirical data on the properties of biopolymers, particularly therapeutic proteins such as antibodies, is scarce. Low data regimes may be especially detrimental for de novo drug design in which the monomer sequence of a de novo generated biopolymer is likely to be out of the distribution of existing training data. In such scenarios, the inclusion of the molecule design as an additional training sample may improve the predictive performance of the computation model trained with the training data.
[0110] FIG. 3A depicts a flowchart illustrating an example of a process 300 for computational intermolecular interface analytics, in accordance with some example embodiments. Referring to FIGS. 1-2 and 3 A, in some example embodiments, the process 300 may be performed by the molecule design system 100 including, for example, the featurization engine 110. In some cases, the process 300 may implement operation 204 of the process 200 shown in FIG. 2 in which the featurization engine 110 applies the feature computation model 115to determine the one or more interface features 116 of the intermolecular interface between the molecule design 112 and the binding target 114. In some cases, the one or more interface features 116 of the interm olecular interface between the molecule design 112 and the binding target 114 may include one or more geometric features and energetic features. In some cases, the one or more interface features 116 may capture biophysical attributes of the intermolecular interface that are inaccessible to conventional atomic distance- and angle-based metrics. For example, in some cases, the one or more interface features 116 may include one or more contacts (or pairs of interacting monomers) identified based on interaction energies arising from a variety of interatomic forces as well as the functional role of each contact (e.g., neutral, structuredetermining, affinity-determining, specificity-determining, and / or the like).[OHl] At 302, a structural ensemble including a plurality of possible three- dimensional structures of a molecule design bound to a binding target is generated. In some cases, operation 302 may also include generating a possible three-dimensional structures of a molecule design bound to a binding target. In some example embodiments, a feature computation model 115 may perform conformational sampling and energy minimization in order to generate the structural ensemble (or conformational ensemble), which includes multiple possible three-dimensional structures (or conformations) of the molecule design bound to the binding target. Biopolymers (e.g., protein molecules) in solution tend to be flexible and will often lack a stable three-dimensional structure (or conformation) in the absence of an interaction partner, such as the binding target, which may be another biopolymer (e.g., protein molecule). The flexible nature of biopolymers means that it is possible for the same monomer sequence (e.g., protein sequence) to fold into multiple different three-dimensional structures (or conformations). Accordingly, in some cases, the feature computation model may samplemultiple different three-dimensional structures (or conformations) of the molecular complex formed by the molecule design binding to the binding target and identify those molecular complexes whose energy satisfies one or more thresholds as the most plausible three- dimensional structures (or conformations) of the molecule design bound to the binding target. For example, in some cases, the feature computation model may perform Monte Carlo sampling, which may include generating multiple different molecular complexes and determining the energy of each molecular complex before selecting those with the lowest energy (e.g., energy minimization). In doing so, the feature computation model may generate the structural ensemble (or conformational ensemble) to include one or more of the molecular complexes whose energy satisfies the one or more thresholds (e.g., lowest energy). As described in more details below, one or more interface features of the molecular interface between the molecule design and the binding target may be determined as a distribution across the different three-dimensional structures (or conformations) in the structural ensemble (or conformational ensemble). Doing so may yield more accurate and reliable results than extracting interface features from a single, static three-dimensional structure (or conformation).
[0112] At 304, one or more interface features of an intermolecular interface between the molecule design and the binding target are determined based at least on the structural ensemble. In some example embodiments, the feature computation model may determine, based at least on the three-dimensional structures (or conformations) of the molecule design bound to the binding target in the structural ensemble (or conformational ensemble), one or more interface features of the intermolecular interface between the molecule design and the binding target. For example, in some cases, the feature computation model may determine the one or more features as a distribution across the different three-dimensional structures (or conformations) in thestructural ensemble (or conformational ensemble), meaning that the same interface features may be extracted from multiple different three-dimensional structures (or conformations) of the molecule design bound to the binding target. In some cases, the one or more interface features may include one or more geometric features including, for example, one or more metrics quantifying or derived from interface shape complementarity, hydrogen bond satisfaction, interface packing quality, entropic penalty of sidechain burial upon complexion, and / or the like. In some cases, the one or more interface features may also include one or more energetic features, which may be derived from and / or quantify a variety of interatomic forces such as electrostatic forces, hydrogen bonding, van der Waals forces, hydrophobic effects, and / or the like. The aforementioned geometric features and energetic features may provide biophysicsbased descriptors of the intermolecular interface between the molecule design and the binding target such that the intermolecular interface may be analyzed along novel dimensions inaccessible to conventional distance- and angle-based metrics.
[0113] In some example embodiments, the one or more interface features may include contacts, or pairs of interacting monomers (e.g., amino acid residues), in the intermolecular interface between the molecule design and the binding target. In some cases, a contact, or a pair of interacting monomers, may be identified based at least on the interaction energy (or magnitude of the interaction energy) between the constituent monomers. In some cases, the one or more interface features may include a total quantity, an average density, a distribution over specific intermolecular interface regions (e.g., individual complementarity determining region (CDR) loops), an aggregate energy value (e.g., minimum, median, maximum, mode, range, and / or the like), and a summary statistic of energy distribution (e.g., skew, kurtosis,and / or the like) of the contacts identified within the intermolecular interface between the molecule design and the binding target.
[0114] In some cases, a contact, or pair of interacting monomers (e.g., amino acid residues), may be classified based on the location of its interaction energy and / or constituent monomers. For example, where the molecule design is a biopolymer, such as a protein molecule, a contact (e.g., the pair of interacting amino acid residues) at the intermolecular interface may be classified based on the source of its interaction energy (e.g., total, sidechainspecific, backbone-mediated, and / or the like). In some cases, the contact may be further classified based on its functional role, which may be determined based at least in part on the type of atomic forces giving rise to the interaction energy between the constituent monomers. For instance, in some cases, the contact may be associated with one or more functional roles, such as neutral, structure-determining, affinity-determining, specificity-determining, and / or the like. In some cases, the contact may be identified as specificity determining if the interaction energy between the constituent monomers arise from interatomic forces critical for binding specificity such as hydrogen-bonding, electrostatics, and nonpolar van der Waals forces. Contrastingly, where the interaction energy between the constituent monomers arises from nonspecific interaction energies, such as solvation interactions, the contact may be identified as neutral, affinity-determining, or structural-determining.
[0115] In some example embodiments, the one or more interface features may include one or more interaction motifs, which the feature computation model may identify based on the contacts (or interacting monomers) present in the intermolecular interface between the molecule design and the binding target. As noted, an interaction motif may refer to a recurring pattern of monomers (e.g., amino acid residues) observed in the intermolecular interfaces ofmolecular complexes (e.g., biopolymers bound to binding targets). In some cases, the feature computation model may further classify each interaction motif present in the intermolecular interface as a favorable interaction motif indicative of a higher quality intermolecular interface, an unfavorable interaction motif indicative of a lower quantity molecular interface, and / or a false positive interaction motif found in computationally generated molecular complexes but rarely observed in nature. In some cases, the one or more interface features may include a total quantity, an average density, a distribution over specific intermolecular interface regions (e.g., individual complementarity determining region (CDR) loops), an aggregate energy value (e.g., minimum, median, maximum, mode, range, and / or the like), and a summary statistic of energy distribution (e.g., skew, kurtosis, and / or the like) of the interaction motifs (or different types of interaction motifs) identified within the intermolecular interface between the molecule design and the binding target.
[0116] FIG. 3B depicts a flowchart illustrating an example of a process 350 for computational intermolecular interface analytics, in accordance with some example embodiments. Referring to FIGS. 1-2 and 3A-B, in some example embodiments, the process 350 may be performed by the molecule design system 100 including, for example, the featurization engine 110. In some cases, the process 350 may implement operation 304 of the process 300 shown in FIG. 3A in which the featurization engine 110 applies the feature computation model 115 to determine the one or more interface features 116 of the intermolecular interface between the molecule design 1 12 and the binding target 114. In some cases, the one or more interface features 116 of the intermolecular interface between the molecule design 112 and the binding target 114 may include one or more geometric features and energetic features. For example, in some cases, the one or more interface features 116 may include contacts, or pairs ofinteracting monomers, identified based on interaction energies arising from a variety of interatomic forces as well as the functional role of each contact (e.g., neutral, structuredetermining, affinity-determining, specificity-determining, and / or the like). Furthermore, in some cases, the one or more interface feature 116 may include interaction motifs, including favorable interaction motifs, unfavorable interaction motifs, and false positive interaction motifs, identified based on the contacts present in the intermolecular interface between the molecule design 112 and the binding target 114.
[0117] At 352, an interaction energy is determined for each pair of monomers in an intermolecular interface between a molecule design and a binding target. In some example embodiments, a feature computation model may be applied to determine one or more interface features of the intermolecular interface between the molecule design and the binding target. In some cases, the one or more interface features may include one or more geometric features and energetic features which, unlike conventional angle- and distance-based descriptors of intermolecular interfaces, are non-trivial and impossible to extract from static molecular structures. For example, in some cases, the feature computation model may determine the one or more interface features as a distribution across a structural ensemble (or conformational ensemble) including multiple different possible three-dimensional structures (or conformations) adopted by the molecule design bound to the binding target. Accordingly, for an individual three-dimensional structure (or conformation) from the structural ensemble (or conformational ensemble), the feature computation model may determine the interaction energy, such as the magnitude of the interaction energy, between each pair of monomers (e.g., amino acid residues) at the intermolecular interface between the molecule design and the binding target.
[0118] The interaction energy between a pair of monomers (e g., amino acid residues) may be driven by a variety of interatomic forces including, for example, electrostatic forces, hydrogen bonding, van der Waals forces, hydrophobic effects, and / or the like. In some cases, the interaction energy between a pair of monomers (e.g., amino acid residues) may capture biophysical attributes absent from conventional distance- and angle-based descriptors characterizing the intermolecular interface. Moreover, in some cases, the interaction energy between a pair of monomers (e.g., amino acid residues) may exclude contribution from those arising from certain types of interatomic forces, such as those associated with functional roles irrelevant to downstream analytics. For example, where the downstream analytics include determining the binding affinity of the molecule design, the interaction energy used to identify contacts at the intermolecular interface between the molecule design and the binding target may exclude contributions from nonspecific interaction energies, such as solvation interactions, that may not contribute to a sufficient extent to the binding affinity of the molecule design.
[0119] At 354, one or more contacts are identified based at least on the interaction energy between each pair of monomers in the intermolecular interface between the molecule design and the binding target. In some example embodiments, the feature computation model may identify a pair of monomers (e.g., amino acid residues) in the intermolecular interface as contacts, or a pair of interacting monomers, based on the interaction energy (e.g., magnitude of the interaction energy) between the pair of monomers. For example, in some cases, a pair of monomers in the intermolecular interface may be identified as contacts (or a pair of interacting monomers) if the interaction energy (e.g., magnitude of the interaction energy) between the pair of monomers satisfies one or more thresholds. In some cases, doing so may classify the pair of monomers as contacts based on a continuous scale of interaction energy that explicitly capturesatomic forces from the spatial coordinates (e.g., Cartesian coordinates (x,y, z)) of the atoms(e.g., heavy atoms) in each constituent monomer. In some cases, the interaction energy -based identification of contacts may further enable the identification of “hotspot” contacts, which are pairs of interacting monomers that contribute significantly more to the binding affinity between the molecule design and the binding target than other, non-hotspot contacts. The latter may be more numerous (than hotspot contacts) along the intermolecular interface but contribute less towards binding affinity and more towards binding specificity than the hotspot contacts. This differentiation between hotspot and non-hotspot contacts is consistent with the observation from energy scaling plots demonstrate that a small subset of monomers (e.g., amino acid residues) at intermolecular interfaces that meets the “contact” energy threshold may represent significant proportion (e.g., 95-97%) of the total energy across the intermolecular interface. In some cases, a contact at the intermolecular interface between the molecule design and the binding target may be classified based on the underlying source of the interaction energy between the constituent monomers. In instances where the molecule design is a biopolymer, for example, the source of the interaction energy be the overall interface, sidechain-specific, or backbone-mediated. As described in more details below, the contacts identified at the intermolecular interface between the molecule design and the binding target may be further classified based on their locations, the underlying interatomic forces, and functional roles.
[0120] At 356, a location of each contact of the one or more contacts is determined. In some example embodiments, the location of a contact may include the location of one or more hotspot contacts, which contribute more to the binding affinity between the molecule design and the binding target than non-hotspot contacts. In some cases, the location of the contact may include the location of the constituent monomers. For instances, in some cases, the pair ofinteracting monomers forming the contact may both be located in the molecule design or, alternatively, in the molecule design and the binding target. In some cases, one or both of the interacting monomers forming the contact may be located in specific regions of the molecule design. In instances where the molecule design is an antibody, for example, both interacting monomers in the contact may be located within a complementarity determining region (CDR) of the molecule design. Alternatively, one monomer may be located within a complementarity determining region (CDR) while the other monomer is located within a framework region of the molecule design. As described in more details below, the location of the monomers forming the contact may determine its functional role, for example, as neutral, specificity-determining, affinity-determining, structure-determining, and / or the like.
[0121] At 358, a functional role of each contact of the one or more contacts is determined. In some example embodiments, the feature computation model may determine, for each contact of the one or more contacts, one or more corresponding functional roles, such as neutral, structure-determining, affinity-determining, specificity-determining, and / or the like. In some cases, the feature computation model may determine the functional role of a contact based on the location of the pair of interacting monomers forming the contact. For example, where the molecule design is an antibody, a contact in which one monomer is in the molecule design and the other monomer is in the binding target may be affinity-determining and / or specificity determining. Contrastingly, a contact in which both monomers are in the molecule design, the contact may be structure-determining. For instance, a contact in which both monomers are in a complementarity determining region (CDR) of the molecule design, such as those interacting through hydrogen bonding (e.g., to form a hydrogen-bond network), may determine and / or stabilize the three-dimensional structure (or conformation) adopted by the complementaritydetermining region (CDR) loops of the molecule design. Similarly, a contact that includes a monomer from a complementarity determining region (CDR) of the molecule design and another monomer from the framework region of the molecule design may also determine and / or stabilize the three-dimensional structure (or conformation) adopted by the complementarity determining region (CDR) loops of the molecule design. As described in more details below, the feature computation model may determine one or more interface features of the intermolecular interface between the molecule design and the binding target based on the contacts present in the intermolecular interface.
[0122] At 360, one or more interaction motifs present in the intermolecular interface are identified based at least on the one or more contacts. In some example embodiments, the feature computation model may identify, based at least on the one or more contacts present in the intermolecular interface, one or more interaction motifs, which are recurring pattern of monomers (e.g., amino acid residues) observed in the intermolecular interfaces of molecular complexes. In some cases, an interaction motif may include two or more interacting monomers (e.g., amino acid residues) at the intermolecular interface. Moreover, an interaction motif may be classified as a favorable interaction motif associated with a stronger or more stable binding interaction between the molecule design and the binding target, an unfavorable interaction motif associated with a weaker or less table binding interaction between the molecule design and the binding target, or a false positive interaction motif common in computationally generated molecular complexes but rarely observed in nature. In some cases, the feature computation model may determine the quantity and distribution of favorable interaction motifs, unfavorable interaction motifs, and / or false positive interaction motifs present at the intermolecular interface between the molecule design and the binding target. As described in more details below, the oneor more interface features of the intermolecular interface between the molecule design and the binding target may be determined based at least on the interaction motifs (or types thereof) present at the intermolecular interface.
[0123] At 362, one or more interface features of the intermolecular interface between the molecule design and the binding target are determined based at least on the location of the one or more contacts, the functional role of each contact of the one or more contacts, and / or the one or more interaction motifs. In some example embodiments, the feature computation model may generate a feature vector including values corresponding to the one or more interface features of the intermolecular interface between the molecule design and the binding target. In some cases, the one or more interface features may include one or more of a total quantity, an average density, a distribution over specific intermolecular interface regions (e.g., individual complementarity determining region (CDR) loops), an aggregate energy value (e.g., minimum, median, maximum, mode, range, and / or the like), and a summary statistic of energy distribution (e.g., skew, kurtosis, and / or the like) of the contacts identified within the intermolecular interface between the molecule design and the binding target. In some cases, the one or more interface features may also include the quantity, relative proportions, and / or distribution of the contacts with different functional roles (e.g., neutral, structure-determining, affinity-determining, specificity-determining, and / or the like) identified within the intermolecular interface. In some cases, the one or more interface features may include the quantity, relative proportions, and / or distribution of the interaction motifs (e.g., favorable interaction motifs, unfavorable interaction motifs, false positive interaction motifs, and / or the like) identified within the intermolecular interface. As noted, the aforementioned interface features may enable the molecule design,including the intermolecular interface between the molecule design and binding target, to be analyzed along novel dimensions inaccessible to conventional distance- and angle-based metrics.
[0124] FIG. 4A depicts a screenshot 400 illustrating an example selection of interface features of an intermolecular interface between a molecule design and a binding target, in accordance with some example embodiments. Referring to FIG. 4A, the screenshot 400 may include a description of the molecule design and the binding target, including the respective monomer sequences, which in this example are the sequences of amino acid residues forming a portion of an antibody (e.g., heavy chain variable region and light chain variable region) and an antigen. In some cases, a feature computation model, such as the feature computation model 115 shown in FIG. 1, may be applied to determine one or more interface features based on a structural ensemble (or conformational ensemble) of the possible three-dimensional structures (or conformations) of the molecule design (e.g., antibody) bound to the binding target (e.g., antigen). For example, in some cases, the feature computation model may perform, based on the spatial coordinates (e.g., Cartesian coordinates (x, y, z)) of the atoms (e.g., heavy atoms) forming the molecule design and the binding target, conformational sampling and energy minimization in order to generate the structural ensemble (or conformational ensemble) before determining the one or more interface features as a distribution across the different three-dimensional structures (or conformations) within the structural ensemble (or conformational ensemble). The screenshot 400 in FIG. 4A shows that the output of the feature computation model may include a numerical value corresponding to each interface feature. In some cases, the output of the feature computation model may be a feature vector populated by these numerical values. FIG. 4B shows the distribution in the values of each interface feature across known protein molecule complexes.
[0125] In some example embodiments, a property computation model may be applied to determine, based at least on the one or more interface features of the intermolecular interface between the molecule design and the binding target, one or more properties of the molecule design. For example, in some cases, the property computation model may be a machine learning model (e.g., a classifier model, a regressor model, and / or the like) that has been trained to determine the one or more properties of the molecule design based on the numerical values of the one or more interface features included in the feature vector. In instances where the molecule design is an antibody and the binding target is an antigen, the property computation model may be trained to perform a classification task and classify the molecule design as a binder (or nonbinder) or, alternatively, perform a regression task and determine the binding affinity (or strength of the binding interaction) between the molecule design and the binding target.
[0126] It should be appreciated that the property computation model, whether implemented as a classifier model or a regressor model, may outperform conventional methodologies for determining the properties of the molecule design. [Let’s include a few sentences here to set up the relationship b / w FIGs. 5A & 5B, and how they pave the way for FIG. 5C and on. I’d like for us to capture Rob’s comment (@
[0048] ) on how ML trained on feature vectors can help account for the various ways to achieve binding and perhaps identify binders more readily. I realize the spec refers to the feature vectors and training - just connect it more directly to improvements over traditional filtering methods here. Thank you!]
[0127] To illustrate, FIG. 5A depicts the outcome of filter-based approaches to differentiating between binders and nonbinders. For example, the plots in FIG. 5A show the distributions for individually filtering molecule designs to median or better for specific onedimensional features. For example, in the top left of FIG 4B, the dashed line indicating themedian of dG cross can be observed to be around -50, which is now the upper limit of dG cross in FIG 5A. Referring now to FIG. 5B, some features may be hypothesized as differentiating between binders and nonbinders. When applied individually, FIG. 5B shows that approximately 50% of the molecule designs are selected in each case (“marginal” curve of 5B, where the second feature is <50% because it’s an integer value that remains either above or below 50%). However, when every filter is applied in combination, the resulting number of molecule designs that pass exponentially drops to 0 (“cumulative” in FIG. 5B). Thus, as shown in FIG. 5B, filtering based on the values of the interface properties individually (“marginally”) removes only approximately 50% of binders, but applying in combination (“cumulatively”) removes nearly all antibodies known to be true binders of the binding target. As such, conventional filtering-based approaches are ineffective, as almost all known true binders are excluded.
[0128] FIG. 5C shows why filter-based methodologies fail to differentiate between binders and nonbinders. Graph 510 in FIG. 5C shows that one of the properties (relative hydrophobicity) is only partially overlapping for antibody -antigen interfaces versus all other protein-protein interfaces, but these are all still binders. Scientifically, this has more to do with establishing that antibodies have their own particular properties (e.g. resulting from their loop- mediated interfaces which are rarer in general protein-to-protein interfaces). In graphs 520 and 530, the same property (relative hydrophobicity) is shown for two different sets of designed antibodies that are variations of 2 different seed molecules. These graphs show that this property by itself is not actually very good at distinguishing binders from nonbinders. Furthermore, graphs 520 and 530 show that two different antibodies that are both binders may exhibit different distributions for the identical property. Thus, filtering molecule designs based on this property(relative hydrophobicity) in isolation is likely to remove a significant number of binders.
[0129] FIG. 5D depicts a graph 540 illustrating the distributions of a conformity score calculated by a property computation model, such as the property computation model 135, for a selection of interface features, in accordance with some example embodiments. As shown in FIG. 5D, the distributions of the conformity score may enable a more accurate differentiation between binders and nonbinders. In some cases, various embodiments of the property computation model disclosed herein may be trained to assign varying significance to each interface feature such that the resulting conformity score enables an accurate differentiation between binders and nonbinders. For example, in some cases, the training of the property computation model may include adjusting one or more parameters of the property computation model, which include the weights that are applied to each interface feature when computing the conformity score. In some cases, the one or more parameters may be adjusted such that the conformity scores output by the property computation model enables a differentiation between binders and nonbinders. For instance, in the example shown in FIG. 5D, the property computation model may be trained to output a lower conformity score for nonbinders and a higher conformity score for binders. The relative importance of each interface feature in the example selection of interface features, which may correspond to the weights learned by the property computation model through training, is depicted in FIG. 5E.
[0130] FIG. 6A depicts a decision tree illustrating an example of a classification hierarchy resulting from the aforementioned property computation model, such as the property computation model 135 shown in FIG. 1, being implemented with a random forest classifier architecture. The performance of the property computation model in differentiating between binders and nonbinders when implemented with the random forest classifier architecture is illustrated in the receiver operating characteristic (ROC) curve shown in the graph 610 shown inFIG. 6B and the confusion matrix 620 shown in FIG. 6C. The graph 610 in FIG. 6B plots the rate of false positive outputs to the rate of the true positive outputs from the property computation model. According to the graph 610 in FIG. 6B, the property computation model, when implemented with a random forest classifier architecture, is able to achieve an area under the curve of 0.72, as well as an accuracy of 0.65, a precision of 0.64, a recall of 0.61, and a Fl score of 0.62. Meanwhile, the confusion matrix 620 depicts a comparison of the counts of true positives (or true binders), true negatives (or true nonbinders), false positives (or false binders), and false negatives (or false nonbinders) in the output of the property computation model. As shown in FIG. 6C, the property computation model, when implemented with a random forest classifier architecture, achieved 187 false negatives (or false nonbinders) and 163 false positives (or false binders) compared to 289 true positives (or true binders) and 353 true negatives (or true nonbinders).
[0131] In some example embodiments, the aforementioned property computation model may also be implemented with a logistic regression classifier architecture. FIG. 7A depicts a graph 700 illustrating the weights from a linear layer of the property computation model implemented with a logistic regression classifier architecture. As shown in the graph 700, the property computation model may be trained to apply different weights to different interface features of the intermolecular interface between a molecule design and a binding target. As noted in the graph 700, the interface features, even those that appears to be bimodal (or having two distinct peak values in their distributions), are in fact unimodal (or having a single peak in their distributions), and may thus contribute negligibly to the classification of binders and nonbinders (e.g. the distributions of the feature dSASA_polar in graphs 520 and 530, which has almost 0.0 feature weight in graph 700). The performance of the property computation modelimplemented with a logistic regression classifier architecture is illustrated by the confusion matrix 750 shown in FIG. 7B. As shown in FIG. 7B, when implemented with a logistic regression classifier architecture, the property computation model achieved 348 false negatives (or false nonbinders) and 328 false positive (or false binders) relative to 653 true negatives (or true nonbinders) and 656 true positives (or true binders). The property computation model also achieved an accuracy of 0.67, precision of 0.65, recall of 0.65, Fl score of 0.66, and an area under the curve (AUC) of 0.66.
[0132] In some example embodiments, the one or more interface features of the intermolecular interface between a molecule design and a binding target may include “hotspot” contacts, which are pairs of interacting monomers contributing more to the binding affinity of the molecule design towards the binding target compared to non-hotspot contacts, which are weaker but more numerous and more important for binding specificity. The interaction energy between hotspot contacts may contribute to a significant portion of the total energy across the intermolecular interface. FIG. 8A depicts a graph 800 illustrating the distribution of minimum contact energy (e.g. hotspots), average contact energy, and median contact energy in the intermolecular interface between the complementarity determining region (CDR) of an antibody and an antigen. Hotspot contacts may concentrate in certain portions of the three-dimensional structure of the molecule design. This phenomenon is illustrated in graph 850 of FIG. 8B, which shows the average number of hotspot contacts across a sampling of a few thousand different antibodies (e g., contacts with an interaction energy < —6) on the heavy chain first complementarity determining region (Hl), heavy chain second complementarity determining region (H2), heavy chain third complementarity determining region (H3), light chain first complementarity determining region (LI), light chain second complementarity determiningregion (L2), and light chain third complementarity determining region (L3). As shown in FIG. 8B, the largest number of hotspot contacts are found in the heavy chain third complementarity determining region (H3) of an antibody.
[0133] In some example embodiments, a feature computation model (e.g., the feature computation model 115 shown in FIG. 1) may identify, based on the interaction energy between each pair of monomers (e.g., amino acid residues) in the intermolecular interface between a molecule design and a binding target, one or more contacts (or pairs of interacting molecules) at the intermolecular interface. In some cases, a pair of monomers (e.g., amino acid residues) in the intermolecular face between the molecule design and the binding target may be identified as a contact (or pair of interacting molecules) if the interaction energy between the two monomers satisfies one or more thresholds. In some cases, a higher interaction energy threshold may identify a smaller quantity of the monomer pairs at the intermolecular interface as contacts (or pairs of interacting monomers) while a lower interaction energy threshold may identify a larger quantity of the monomer pairs at the intermolecular interface as contacts (or pairs of interacting monomers). FIG. 9A depicts a graph 900 illustrating the individual traces for 1000 randomly sampled interfaces from over 8300 molecular complexes while FIG. 9B depicts a graph 950 illustrating the average and standard deviation for the 8300 traces. FIGS. 9A-B illustrate the relationship between the interaction energy threshold and the proportion of monomer pairs identified as contacts.
[0134] In some example embodiments, a property computation model (e.g., the property computation model 135 shown in FIG. 1) may determine one or more properties of a molecule design (e.g., a biopolymer such as a protein molecule) based on one or more interface features of the intermolecular interface between the molecule design and a binding target (e.g., asmall molecule or another biopolymer). In some cases, the training of the property computation model may include adjusting one or more parameters of the property computation model, such as the weights that are applied in one or more layers of the property computation model, to place greater emphasis on those interface features that increases the predictive performance of the property computation model. For example, in some cases, an interface feature that enables a more accurate prediction (e.g., classification, regression, and / or the like) of the binding affinity between the molecule design and the binding target may be associated with a higher weight than those that fail to enable an accurate binding affinity prediction. To further illustrate, FIG. 10A depicts graphs illustrating a comparison of the distribution of interface features assigned more positive weights and the distribution of interface features assigned more negative weights by a linear regressor property computation model. FIG. 10B depicts graphs illustrating the shift that is present in the distributions of different interface features for binders and nonbinders. As shown in FIGS. 10A-B, an interface feature whose distribution of values for binders overlaps less with the distribution of values for nonbinders may be assigned a higher weight than an interface feature whose distribution of values for binders overlaps more with the distribution of values for nonbinders. This may be due at least in part to an interface feature with a significant overlap in the distribution of binders and the distribution of nonbinders being less effective at enabling a differentiation between binders and nonbinders.
[0135] In some example embodiments, a feature computation model (e.g., the feature computation model 115 shown in FIG. 1) may identify one or more interface property, including the source of the interaction energy between the pair of interacting monomers forming a contact at the interm olecular interface between a molecule design and a binding target. For example, the interaction energy of the contact may be classified as total interaction energy, sidechain-specificinteraction energy, and backbone-mediated interaction energy. FIG. 11 depicts examples in which the interaction energy of a contact is split between total interaction energy and sidechainspecific interaction energy.
[0136] In some example embodiments, the one or more interface properties determined by the feature computation model (e.g., the feature computation model 115 shown in FIG. 1) may include the presence (or absence) of one or more interaction motifs identified based on the contacts (or pairs of interacting monomers) present at the intermolecular interface between a molecule design and a binding target. In some cases, the one or more interaction motifs may include one or more pairs of monomers (e.g., preferred pairs of amino acid residues) that are commonly observed in the intermolecular interfaces of molecular complexes. FIG. 12A depicts a heatmap 1200 illustrating the amino acid residue pairing preferences that are present at the intermolecular interface between antibodies and antigens. As shown in FIG. 12A, some amino acid residue pairs may occur more frequently at the intermolecular interface of antigens and antibodies than others. The graph 1250 in FIG. 12B further differentiates between amino acid residue pairing preferences for intramolecular interactions (e.g., within the antibody) and intermolecular interactions (e.g., between the antibody and the antigen). FIG. 12C depicts additional examples of interaction motifs, in this case one that includes aspartic acid (Asp) in the antibody and asparagine (N) and serine (S) in the antigen. As shown in FIG. 12C, the same interaction motif may be present at the intermolecular interface of two antigen-antibody complexes, despite differences in modalities and the positioning of the aspartic acid (Asp). Accordingly, in some cases, an interaction motif that includes two or more pairs of monomers, such as two or more pairs of interacting amino acid residues, may be classified as an interaction motif if the interaction motif occurs with sufficient frequency in known intermolecularinterfaces. Alternatively, an interaction motif may be assigned a more granular classification corresponding to the frequency with which the interaction motif is observed in known intermolecular interfaces. FIG. 12D depicts a schematic diagram illustrating a comparison of the interaction motifs identified by the feature computation model and similar interaction motifs mined from a dataset of known structures. In some cases, one or more properties of a molecule design, such as its binding affinity towards a binding target, may be determined (e.g., by the property computation model 135 shown in FIG. 1) based on a comparison of the interaction motifs identified within the intermolecular interface between the molecule design and the binding target (along with the respective contributions sidechain-specific interactions and backbone-mediated interactions) and those interaction motifs mined from the dataset of known structures. In some cases, for example, the one or more properties of the molecule design may be determined based on the presence (or absence) of favorable motifs, unfavorable motifs, and / or false positive motifs mined from the dataset of known structures.
[0137] In some example embodiments, the feature computation model (e.g., the feature computation model 115 shown in FIG. 1) may classify each contact in accordance with its functional role including, for example, structural-determining, affinity-determining, specificity-determining, and / or the like. In some cases, the locations of the interacting monomers (e.g., amino acid residues) in a contact and the interatomic forces driving the interaction may determine the functional role of the contact. For example, where the interaction energy between the interacting monomers arises from nonspecific interaction energies, such as solvation interactions, the contact may be identified as neutral, affinity-determining, or structural-determining. In some cases, where both interacting monomers are in a complementarity determining region (CDR) of the molecule design and interacts throughhydrogen-bonds (e.g., to form a hydrogen-bond network), the contact may determine and / or stabilize the three-dimensional structure (or conformation) adopted by the complementarity determining region (CDR) loops of the molecule design. In cases where one interacting monomer is in a complementarity determining region (CDR) of the molecule design and the other interacting monomer is in the framework region of the molecule design, the contact may also determine and / or stabilize the three-dimensional structure (or conformation) adopted by the complementarity determining region (CDR) loops of the molecule design.
[0138] FIG. 13 A depicts another example of an intermolecular interface with contacts within the complementarity determining region (CDR). The example in FIG. 13 A includes 22 sidechain-mediated hydrogen-bonds across 18 out of the 77 (or -23%) of the amino acid residues in the complementarity determining region. FIG. 13B depicts a heatmap 1300 illustrating the prevalence of sidechain-specific hydrogen-bonds between contacts based on the respective locations of the interacting amino acid residues (e.g., heavy chain CDR1 (Hl), heavy chain CDR2 (H2), heavy chain CDR3 (H3), heavy chain CDR4 (H4), light chain CDR1 (LI), light chain CDR2 (L2), light chain CDR3 (L3), light chain CDR4 (L4), and framework region (FW)). Meanwhile, the heatmap 1350 shown in FIG. 13B illustrates the prevalence backbone-mediated hydrogen-bonds between contacts based on the respective locations of the interacting amino acid residues (e.g., heavy chain CDR1 (Hl), heavy chain CDR2 (H2), heavy chain CDR3 (H3), heavy chain CDR4 (H4), light chain CDR1 (LI), light chain CDR2 (L2), light chain CDR3 (L3), light chain CDR4 (L4), and framework region (FW)).
[0139] It should be appreciated that contacts driven by hydrogen-bonds between one amino acid residue in the complementarity determining region (CDR) of an antibody and another amino acid residue in the framework region (FW) of the antibody may be relevant to thehumanization of the antibody in instances where the antibody originates from a non-human source (e.g., an animal immunization campaign). For example, FIG. 13C depicts the three- dimensional structures (or conformations) of antibodies before and after affinity maturation to illustrate the different structural roles of hydrogen-bonds. FIG. 13C shows that, in some cases, sequence-conserving hydrogen-bonds and compensatory hydrogen-bonds may be likely to preserve the structure of the complementarity determining region (CDR). Contrastingly, new hydrogen-bonds may mediate alternative structures. Accordingly, in some example embodiments, the modification of a non-human antibody, for example, to humanize the antibody for therapeutic uses, may be guided by the hydrogen-bond driven contacts present in the intermolecular interface. It should be appreciated that hydrogen bonds are one example of structural stabilizing contacts and other types of contacts may also serve to stabilize the three- dimensional structure (or conformation) of a molecule design, such as a protein molecule. For instance, hydrophobic interactions between a complementarity determining region (CDR) sidechain and framework region sidechain may also serve to stabilize a particular conformation if the former is not exclusively affinity-determining.
[0140] The schematic diagram depicted in FIG. 13D provides a visualization of so- called critical contacts and the corresponding interaction energies in a seed complex structure, which may include a lead molecule (e.g., a non-human antibody identified through an immunization campaign) bound to a binding target (e.g., an antigen). In this context, a “critical” contact may be one identified, for example, based on the underlying interatomic forces and / or locations of the constituent monomers, as requiring conservation (or change) during the modification of the lead molecule.
[0141] FIG. 14A depicts a comparison of the conventional distance-based metrics and the contact energy-based interface feature for quantifying an intermolecular interface between a molecule design and a binding target, in accordance with some example embodiments. Referring to FIG. 14A, the graph 1400 is a distogram depicting the distance between every pair of amino acid residues in the intermolecular interface of an antigen-antibody complex formed by the binding interaction between an antibody and an antigen. As shown in the graph 1400, conventional distance-based metrics fail to provide any context for why the antibody adopts a particular three-dimensional structure (or conformation). Contrastingly, the graph 1425 illustrates the distribution of pairwise interaction energies present in the intermolecular interface of the antigen-antibody complex. As noted, in some cases, one or more properties of the antibody, such as its binding affinity towards the antigen, may be determined based on the location and the functional role of contacts identified based on interaction energy as well as, in some cases, the source of the interaction energy (e.g., sidechain-specific, backbone-mediated, and / or the like).
[0142] In some cases, the distribution of interaction energies shown in the graph 1425, including certain “hotspot” contacts that contributes significantly more to the binding interaction between the antigen and the antibody, may capture the biophysics of the three- dimensional structure (or conformation) adopted by the antibody. For example, contacts in which both of the interacting monomer (e.g., amino acid residues) are in the antibody, such as the complementarity determining region (CDR) and / or the framework region (FR) of the antibody, may determine and / or stabilize the three-dimensional structure (or conformation) of the complementarity determining region (CDR) of the antibody. This phenomenon is illustrated in FIGS. 14B-14C. In FIG. 14B, the heatmap 1450 depicts the average quantity of sidechain-specific hydrogen-bond driven contacts across different regions of an antibody, in accordance with some example embodiments. These sidechain-specific hydrogen-bond driven contacts may include interacting amino acid residues located across different regions of an antibody, including, heavy chain CDR1 (Hl), heavy chain CDR2 (H2), heavy chain CDR3 (H3), heavy chain CDR4 (H4), light chain CDR1 (LI), light chain CDR2 (L2), light chain CDR3 (L3), light chain CDR4 (L4), and framework region (FW). The heatmap 1450 shows that contacts in which both interacting amino acid residues are in the heavy chain CDR3 (H3) may contribute most to stabilizing the structure of the complementarity determining region (CDR) of the antibody followed by those contacts in which both interacting amino acid residues are in the heavy chain CDR4 (H4) and the framework region (FW) of the antibody. The magnitude of interaction energies present in sidechain-specific hydrogen-bond driven contacts may be sequencedependent, meaning that changing the type (or identity) of one or more animo acid residues in the antibody may change the magnitudes of the interaction energies and, consequently, the three- dimensional structure (or conformation) adopted by the complementarity determining region (CDR) of the antibody.
[0143] FIG. 14C depicts a heatmap 1475 illustrating the average quantity of backbone-mediated hydrogen-bond driven contacts across different regions of an antibody, in accordance with some example embodiments. The aforementioned backbone-mediated hydrogen-bond driven contacts may include interacting amino acid residues located across different regions of an antibody, including, heavy chain CDR1 (Hl), heavy chain CDR2 (H2), heavy chain CDR3 (H3), heavy chain CDR4 (H4), light chain CDR1 (LI), light chain CDR2 (L2), light chain CDR3 (L3), light chain CDR4 (L4), and framework region (FW). As shown in FIGS. 14B-C, backbone-mediated interaction energies may be fewer than sidechain-specificinteraction energies for contacts across various complementarity determining regions (CDR) and framework regions (FW) of the antibody. Nevertheless, backbone-mediated hydrogen-bond driven contacts in which one interacting amino acid residue is in the light chain CDR1 (LI) and the other interacting amino acid residue is in the light chain CDR4 (L4) may exhibit largest backbone-mediated interaction energy, and may therefore contribute more to determining and / or stabilizing the three-dimensional structure (or conformation) of the complementarity determining region (CDR) of the antibody.
[0144] FIG. 14D depicts a network graph 1480 illustrating an example of a hydrogenbond network formed by contacts in which the interacting amino acid residues are in different complementarity determining regions of the antibody. It should be appreciated that these intercomplementarity determining region (CDR) hydrogen bonds may form a hydrogen-bond network that stabilizes and / or determines the three-dimensional structure (or conformation) of the complementarity determining regions (CDRs) of the antibody.
[0145] In some example embodiments, the one or more interface features determined by the feature computation model and, in some cases, the one or more properties of a molecule design (e.g., biopolymer) determined by the property computation model therefrom, may serve to augment available training data. For example, in instances where molecular complexes including a molecule design (e.g., a biopolymer such as a protein molecule) and a binding target (e.g., a small molecule or another biopolymer) are generated computationally, such as by one or more computation models, these molecular complexes may be annotated (or labeled) based on one or more interface features. As such, computationally generated molecular complexes and the annotations (or labels) determined based on one or more interface features (e.g., contacts at the intermolecular interface identified on the basis of interaction energies) may be used astraining samples, which may be especially advantageous in a low data regime where molecular complexes with empirical data are scarce. In some cases, the one or more interface features, which capture the biophysics governing the intermolecular interface, may be interpretable and, as the graphs in FIG. 4B shows, possess well-defined distributions (of values). Some interface features, such as the contacts and hotspot contacts present in intermolecular interfaces, may themselves give rise to annotations (or labels) for training and / or evaluation of individual molecules (e.g., biopolymers such as protein molecules) and molecular complexes (e.g., antigenantibody complexes). For instance, some monomers (e.g., amino acid residues) or pairings of monomers may contribute more to structural stability, binding affinity, and / or binding specificity than other monomers or pairings of monomers. This phenomenon is illustrated in the graph 1500 shown in FIG. 15, which shows that some amino acid residues (e.g., of the twenty canonical amino acid residues encoded by the standard genetic code) occur more frequently in contacts (or pairs of interacting amino acid residues) and hotspot contacts (or contacts exhibiting a threshold interaction energy between the constituent amino acid residues) than others. In some cases, each type of amino acid residue may be annotated (or labeled) based on how frequently the amino acid residue appears in a contact, a hotspot contact, and / or the like.
[0146] In some example embodiments, the feature computation model may generate, for the intermolecular interface between a molecule design (e.g., a biopolymer such as a protein molecule) and a binding target (e.g., a small molecule or another biopolymer), a feature vector populated by values corresponding to the one or more interface features of the intermolecular interface. For example, in some cases, the feature vector may include one or more values corresponding to a total quantity, an average density, a distribution over specific intermolecular interface regions (e.g., individual complementarity determining region (CDR) loops), anaggregate energy value (e.g., minimum, median, maximum, mode, range, and / or the like), and a summary statistic of energy distribution (e.g., skew, kurtosis, and / or the like) of the contacts, hotspot contacts, and interaction motifs identified within the intermolecular interface between the molecule design and the binding target. In some cases, the feature vectors generated by the feature computation model may provide a more complete description of the space of possible intermolecular interfaces. FIG. 16 depicts a graph 1600 illustrating the distribution of the pairwise cosine similarity between the feature vectors associated with the intermolecular interface of antigen-antibody complexes. In instances where a computation model is trained to generate molecular complexes, such as antigen-antibody complexes, the loss function of the computation model may be configured to promote intermolecular interfaces that are geometrically similar to those with certain desirable properties (e.g., binding affinity) across the entirety of the viable interface feature-combination space. As noted, various examples of the interface features described herein are sequence-agnostic fashion. Accordingly, the ability to train and / or evaluate in a sequence-agnostic manner may be especially advantageous in a low data regime in which empirical data for molecular complexes is scarce. In a low data regime and in other scenarios in which a computation model is required to perform out-of-distribution inference, it may be more advantageous to train the computation model to learn biophysical fundamentals, as captured by various examples of interface features described herein.
[0147] FIG. 17 depicts a block diagram illustrating an example of a computing system 1700, in accordance with some example embodiments. In some example embodiments, the computing system 1700 can be used to implement the analysis controller 110 and / or any components therein. As shown in FIG. 17, the computing system 1700 can include a processor 1710, a memory 1720, a storage device 1730, and input / output devices 1740. The processor1710, the memory 1720, the storage device 1730, and the input / output devices 1740 can be interconnected via a system bus 1750. The processor 1710 is capable of processing instructions for execution within the computing system. Such executed instructions can implement one or more components of, for example, analysis controller 110. In some example embodiments, the processor 1710 can be a single-threaded processor. Alternately, the processor 1710 can be a multi -threaded processor. The processor 1710 is capable of processing instructions stored in the memory 1720 and / or on the storage device 1730 to display graphical information for a user interface provided via the input / output device 1740.
[0148] The memory 1720 is a computer readable medium such as volatile or nonvolatile that stores information within the computing system 1700. The memory 1720 can store data structures representing configuration object databases, for example. The storage device 1730 is capable of providing persistent storage for the computing system 1700. The storage device 1730 can be a floppy disk device, a hard disk device, an optical disk device, a tape device, a solid-state drive, and / or other suitable persistent storage means. The input / output device 1740 provides input / output operations for the computing system 1700. In some example embodiments, the input / output device 1740 includes a keyboard and / or pointing device. In various implementations, the input / output device 1740 includes a display unit for displaying graphical user interfaces.DEFINITIONS
[0149] Unless otherwise defined, all terms of art, notations and other scientific terms or terminology used herein are intended to have the meanings commonly understood by those of skill in the art to which this disclosure pertains. In some cases, terms with commonly understood meanings are defined herein for clarity and / or for ready reference, and the inclusion of such definitions herein should not necessarily be construed to represent a substantial difference overwhat is generally understood in the art. Many of the techniques and procedures described or referenced herein are well understood and commonly employed using conventional methodology by those skilled in the art.
[0150] As used herein, the conformation of a protein molecule (or a protein conformation) is a three-dimensional arrangement of residues in a sequence of amino acid residues forming the protein molecule. A residue is an organic molecule that includes an alpha carbon linked to an amino group, a carboxyl group, a hydrogen atom, and a variable component called a side chain. As used herein, a protein-ligand complex is a structure formed as a result of intermolecular interactions (e.g., electrostatics, van der Waals forces, and / or the like) between a protein molecule (having a particular conformation) and a ligand. In some cases, the protein molecule and the compound in the coupled protein-structure-compound can be bound to one another (e.g., distance between the center of mass of the compound and the center of mass of the nitrogen atoms in the perturbation set of the protein molecule is smaller than a threshold distance value). In some cases, the protein molecule and the compound in the coupled protein- structurecompound are dissociated or not bound to one another (e.g., distance between the center of mass of the compound and the center of mass of the nitrogen atoms in the perturbation set of the protein molecule is greater than the threshold distance value). As used herein, structural information associated with a particular conformation of a protein molecule includes the spatial arrangement and / or orientation of atoms in each constituent amino acid residue forming the protein molecule (e g., relative spatial locations and orientation of atoms in the protein molecule).
[0151] The singular form “a”, “an”, and “the” include plural references unless the context clearly dictates otherwise. For example, the term “a cell” includes one or more cells,including mixtures thereof. “A and / or B” is used herein to include all of the following alternatives: “A”, “B”, “A or B”, and “A and B”.
[0152] It is understood that aspects and embodiments of the disclosure described herein include “comprising”, “consisting”, and “consisting essentially of’ aspects and embodiments.
[0153] As used herein, “comprising” is synonymous with “including”, “containing”, or “characterized by”, and is inclusive or open-ended and does not exclude additional, unrecited elements or method steps. Any recitation herein of the term “comprising”, particularly in a description of components of a composition or in a description of steps of a method, is understood to encompass those compositions and methods consisting essentially of and consisting of the recited components or steps. As used herein, “consisting of’ excludes any elements, steps, or ingredients not specified in the claimed composition or method. As used herein, “consisting essentially of” does not exclude materials or steps that do not materially affect the basic and novel characteristics of the claimed composition or method.
[0154] Where a range of values is provided, it is understood by one having ordinary skill in the art that all ranges disclosed herein encompass any and all possible sub-ranges and combinations of sub-ranges thereof. Any listed range can be easily recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, etc. As will also be understood by one skilled in the art all language such as “up to”, “at least”, “greater than”, “less than”, and the like include the number recited and refer to ranges which can be subsequently broken down into sub-ranges as discussed above. As will be understood by one skilled in the art,a range includes each individual member. Thus, for example, a group having 1-3 articles refers to groups having 1, 2, or 3 articles. Similarly, a group having 1-5 articles refers to groups having 1, 2, 3, 4, or 5 articles, and so forth.
[0155] Certain ranges are presented herein with numerical values being preceded by the term “about.” The term “about” is used herein to provide literal support for the exact number that it precedes, as well as a number that is near to or approximately the number that the term precedes. In determining whether a number is near to or approximately a specifically recited number, the near or approximating unrecited number may be a number which, in the context in which it is presented, provides the substantial equivalent of the specifically recited number. If the degree of approximation is not otherwise clear from the context, “about” means either within plus or minus 10% of the provided value, or rounded to the nearest significant figure, in all cases inclusive of the provided value.
[0156] Headings, e.g., (a), (b), (i) etc., are presented merely for ease of reading the specification and claims. The use of headings in the specification or claims does not require the steps or elements be performed in alphabetical or numerical order or the order in which they are presented.
[0157] It is appreciated that certain features of the disclosure, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the disclosure, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub-combination. All combinations of the embodiments pertaining to the disclosure are specifically embraced by the present disclosure and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all sub-combinations of thevarious embodiments and elements thereof are also specifically embraced by the present disclosure and are disclosed herein just as if each and every such sub-combination was individually and explicitly disclosed herein.
[0158] Non-transitory computer program products (i.e., physically embodied computer program products) are also described that store instructions, which when executed by one or more data processors of one or more computing systems, causes at least one data processor to perform operations herein. Similarly, computer systems are also described that may include one or more data processors and memory coupled to the one or more data processors. The memory may temporarily or permanently store instructions that cause at least one processor to perform one or more of the operations described herein. In addition, methods can be implemented by one or more data processors either within a single computing system or distributed among two or more computing systems. Such computing systems can be connected and can exchange data and / or commands or other instructions or the like via one or more connections, including a connection over a network (e g. the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, or the like), via a direct connection between one or more of the multiple computing systems, etc.
[0159] The subject matter described herein may be embodied in systems, apparatus, methods, and / or articles depending on the desired configuration. For example, apparatuses and / or processes described herein can be implemented using one or more of the following: a processor executing program code, an application-specific integrated circuit (ASIC), a digital signal processor (DSP), an embedded processor, a field programmable gate array (FPGA), and / or combinations thereof. These various implementations may include implementation in one or more computer programs that are executable and / or interpretable on a programmable systemincluding at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device. These computer programs (also known as programs, software, software applications, applications, components, program code, or code) include machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the term “machine-readable medium” refers to any computer program product, computer-readable medium, computer-readable storage medium, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine- readable medium that receives machine instructions. Similarly, systems are also described herein that may include a processor and a memory coupled to the processor. The memory may include one or more programs that cause the processor to perform one or more of the operations described herein.
[0160] Although a few variations have been described in detail above, other modifications or additions are possible. In particular, further features and / or variations may be provided in addition to those set forth herein. Moreover, the implementations described above may be directed to various combinations and subcombinations of the disclosed features and / or combinations and subcombinations of several further features disclosed above. In addition, the logic flow depicted in the accompanying figures and / or described herein does not require the particular order shown, or sequential order, to achieve desirable results. Other implementations may be within the scope of the following claims. Furthermore, the specific values provided in the foregoing are merely examples and may vary in some implementations.
[0161] Although various aspects of the present disclosure are set out in the claims, other aspects of the present disclosure comprise other combinations of features from the described implementations with the features of the claims, and not solely the combinations explicitly set out in the claims.
Claims
CLAIMSWhat is claimed is:
1. A computer-implemented method, comprising: generating a molecule design comprising a sequence of monomers; determining one or more interface features of an intermolecular interface between the molecule design and a binding target, where the one or more interface features are determined by at least identifying one or more contacts at the intermolecular interface, where each contact of the one or more contacts comprises a pair of monomers at the intermolecular interface having a threshold interaction energy, and determining, based at least on the one or more contacts, the one or more interface features; and determining, based at least on the one or more interface features, a property of the molecule design.
2. The method of claim 1, wherein the one or more contacts are identified by at least determining an interaction energy for each pair of monomers at the intermolecular interface between the molecule design and the binding target, and identifying, based at least on the interaction energy between each pair of monomers at the intermolecular interface between the molecule design and the binding target, the one or more contacts.
3. The method of claim 2, wherein the determining the one or more interface features further includes identifying, based at least on an interaction energy of each contact of the one or morecontacts, one or more hotspot contacts, wherein a contact is identified as a hotspot contact where an interaction energy between a pair of interacting monomers comprising the contact satisfies one or more thresholds, and determining the one or more interface features to include one or more of a total quantity, an average density, a distribution over specific intermolecular interface regions, an aggregate energy value, and a summary statistic of energy distribution of the one or more hotspot contacts.
4. The method of any of claims 1 to 3, wherein the one or more interface features include one or more of a total quantity, an average density, a distribution over specific intermolecular interface regions, a minimum energy, an aggregate energy value and a summary statistic of energy distribution of the one or more contacts.
5. The method of any of claims 1 to 4, wherein the determining the one or more interface features further includes determining a location of each contact of the one or more contacts, determining, based at least on the location of each contact of the one or more contacts, a functional role of each contact of the one or more contacts, and determining the one or more interface features further corresponds to the functional role of each contact of the one or more contacts.
6. The method of claim 5, wherein the location of each contact of the one or more contacts include a location of each monomer of the pair of interacting monomers comprising the contact.
7. The method of claim 6, wherein the functional role of a contact is structuredetermining where both monomers of the pair of interacting monomers comprising the contact are located in a complementarity determining region (CDR) of the molecule design.
8. The method of any of claims 6 to 7, wherein the functional role of a contact is structure-determining where one monomer of the pair of interacting monomers comprising the contact is located in a complementarity determining region (CDR) of the molecule design and another monomer of the pair of interacting monomers comprising the contact is located in a framework region (FW) of the molecule design.
9. The method of any of claims 6 to 8, wherein the functional role of a contact is affinity-determining or specificity-determining where one monomer of the pair of interacting monomers comprising the contact is located in the molecule design and another monomer of the pair of interacting monomers comprising the contact is located in the binding target.
10. The method of any of claims 1 to 9, wherein the determining the one or more interface features further includes identifying, based at least on the one or more contacts, one or more interaction motifs present at the intermolecular interface between the molecule design and the binding target, and determining the one or more interface features to include one or more of a total quantity, an average density, a distribution over specific intermolecular interface regions, an aggregate energy value, and a summary statistic of energy distribution of the one or more interaction motifs.
11. The method of claim 10, wherein each interaction motif of the one or more interaction motifs comprises a pattern of one or more contacts observed in a threshold quantity of molecule designs.
12. The method of any of claims 10 to 11, wherein the one or more interaction motifs include a favorable interaction motif observed in a threshold quantity of molecular complexes exhibiting one or more desirable properties.
13. The method of any of claims 10 to 12, wherein the one or more interaction motifs include an unfavorable interaction motif observed in molecule complexes failing to exhibit one or more desirable properties or exhibiting one or more undesirable properties.
14. The method of any of claims 10 to 13, wherein the one or more interaction motifs include a false positive interaction motif present in a threshold quantity of computationally generated molecular complexes but absent from a threshold quantity of molecular complexes observed in nature.
15. The method of any of claims 10 to 14, wherein the one or more interaction motifs include one or more preferred pairings of monomers.
16. The method of any of claims 1 to 15, wherein the determining the one or more interface features further include determining a source of interaction energy between the pair of monomers comprising each contact of the one or more contacts in the intermolecular interface between the molecule design and the binding target, and determining the one or more interface features to include the source of interaction energy for each contact of the one or more contacts.
17. The method of claim 16, wherein the source of interaction energy include sidechain-specific atomic interactions or backbone-mediated atomic interactions.
18. The method of any of claims 1 to 17, wherein the determining the one or more interface features further include determining a type of interaction energy between the pair of monomers comprising each contact of the one or more contacts, determining, based at least on the type of interaction energy, a functional role of eachcontact of the one or more contacts, and determining the one or more interface features to include the functional role of each contact of the one or more contacts.
19. The method of claim 18, wherein the determining the one or more interface features further include determining one or more specificity determining interatomic forces as giving rise to an interaction energy between the pair of monomers comprising a contact, and determining, based at least on the one or more specificity determining interatomic forces, a functional role of the contact as specificity-determining.
20. The method of claim 19, wherein the one or more specificity determining interatomic forces include hydrogen-bonding, electrostatics, and nonpolar van der Waals forces.
21. The method of any of claims 18 to 20, wherein the determining the one or more interface features further include determining one or more nonspecific interatomic forces as giving rise to an interaction energy between the pair of monomers comprising a contact, and determining, based at least on the one or more nonspecific interatomic forces, a functional role of the contact as neutral, affinity-determining, and / or structure-determining.
22. The method of claim 21, wherein the one or more nonspecific interatomic forces include solvation interactions.
23. The method of any of claims 1 to 22, further comprising: generating a structural ensemble including a plurality of three-dimensional structures of a molecular complex comprising the molecule design bound to the binding target; and determining, based at least on the structural ensemble, the one or more interface featuresof the intermolecular interface between the molecule design and the binding target.
24. The method of claim 23, wherein the one or more interface features are determined as a distribution of values across the plurality of three-dimensional structures comprising the structural ensemble.
25. The method of any of claims 23 to 24, wherein the generating the structural ensemble includes iteratively adjusting a three-dimensional structure of the molecular complex to generate a plurality of possible three-dimensional structures of the molecular complex, and identifying, for inclusion in the structural ensemble, one or more possible three- dimensional structures of the plurality of possible three-dimensional structures whose energy satisfies one or more criteria.
26. The method of any of claims 1 to 25, further comprising: applying a feature computation model to determine the one or more interface features of the intermolecular interface between the molecule design and the binding target, wherein the feature computation model outputs a feature vector including one or more numerical values corresponding to the one or more interface features; and applying a property computation model to determine, based at least on the feature vector, the property of the molecule design.
27. The method of any of claims 1 to 26, wherein the property of the molecule design includes a binding affinity and / or a binding specificity between the molecule design and the binding target.
28. The method of any of claims 1 to 27, wherein the molecule design comprises a biopolymer and the binding target comprises a small molecule or another biopolymer.
29. The method of any of claims 1 to 28, wherein the molecule design comprises a protein molecule and the binding target is another protein molecule.
30. The method of any of claims 1 to 29, wherein the molecule design comprises an antibody and the binding target comprises an antigen.
31. A system, comprising: at least one data processor; and at least one memory storing instructions, which when executed by the at least one data processor, result in operations comprising the method of any of claims 1 to 30.
32. A non-transitory computer readable medium storing instructions, which when executed by at least one data processor, result in operations comprising the method of any of claims 1 to 30.
Citation Information
Patent Citations
Polypeptides Capable of Forming Homo-Oligomers with Modular Hydrogen Bond Network-Mediated Specificity and Their Design
US20190112345A1