Systems and methods for predicting protein-protein interaction using machine learning techniques

WO2026015381A8PCT designated stage Publication Date: 2026-02-19TRIANA BIOMEDICINES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/036421
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-09
Filing Date
2025-07-03
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Predicting protein-protein interactions (PPIs) based solely on structure remains a challenge in biology, as existing methods struggle to accurately identify descriptors that describe the physics of interaction between proteins and predict binding poses.

Method used

A machine learning approach utilizing regression and classification models to predict binding poses of proteins, employing a systematic filtering of docking poses through hotspot prediction analysis and clustering, with descriptors characterizing interaction interface features to identify native and near-native binding configurations.

Benefits of technology

The method effectively predicts biologically relevant protein-protein complexes by distinguishing between native, near-native, and other binding poses, enhancing the accuracy of protein interaction prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025036421_19022026_PF_FP_ABST
    Figure US2025036421_19022026_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods are disclosed for building a regression-based machine learning algorithm to predict native and near-native binding confirmations between two or more proteins. They include receiving training data, the training data representing a plurality of protein to protein docking poses based on one or more hotspots and a set of descriptors characterizing interaction interface features of one or more of the plurality of protein to protein docking poses. The systems and methods further specify a dedicated training data set based on a variety of single domain interactions. The training data set comprises an information space required to create dedicated regression and classification models in order to identify near native binding configurations. A specific and diverse set of interaction interface descriptors is calculated to provide the foundation for both, the regression and classification models.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.136883-01220 SYSTEMS AND METHODS FOR PREDICTING PROTEIN-PROTEIN INTERACTION USING MACHINE LEARNING TECHNIQUES RELATED APPLICATIONS

[0001] This application claims the benefit of priority to US Provisional Application Serial No. 63 / 669,129, filed July 9, 2024, the entire contents of which are incorporated herein by reference. BACKGROUND

[0001] Understanding protein-protein interactions (PPIs) is critical to understanding the biological processes in our everyday lives. In order to understand the interactions between proteins, and in particular the ability to predict the interaction between proteins, an in-depth understanding of protein docking, protein folding, and protein stability is essential. The ability to predict the interaction between different proteins hinges on correctly identifying descriptors that describe the physics of interaction between the proteins, descriptors that describe amino acid sequences that exist at an interface between the proteins, and descriptors that describe the interaction matrices responsible for the biding of proteins. To this end identifying the descriptors that significantly contribute to the interaction of proteins, and developing models that predict the interaction of proteins and that determine whether the proteins will actually bind based on the descriptors is particularly important to our understanding of how proteins interact with one another. 1 ME1\53706606.v1Attorney Docket No.136883-01220 SUMMARY

[0002] The instant application discloses, among other things, a method directed to training a regression-based machine learning algorithm to predict native and near-native binding confirmations between two or more proteins. The method can include a step of receiving training data, where the training data represents a plurality of protein to protein docking poses based on one or more hot spots. The method can further include a step of receiving a set of descriptors, the set of descriptors characterizing interaction interface features of one or more of the plurality of protein to protein docking poses. The method can further include a step of training, using the set of descriptors and the training data, the regression-based machine learning algorithm to generate a trained machine learning regression model, the trained machine learning regression model is able to identify one or more binding poses of two or more proteins.

[0003] The instant application also discloses, a method directed to determining native and non- native binding confirmation of two or more proteins. The method can include a step of receiving a set of descriptors, the set of descriptors characterizing interaction interface features of one or more of the plurality of protein to protein docking poses. The method can further include a step of receiving data representing a plurality of protein to protein docking poses based on one or more hot spots. The method can further include performing one or more of the steps of the trained regression-based machine learning algorithm in order to identify one or more binding poses of two proteins. 2 ME1\53706606.v1Attorney Docket No.136883-01220

[0004] The instant application also discloses, a method of training a classification-based machine learning algorithm to classify native and near native binding confirmations between two or more proteins. The method can include a step of receiving training data, the training data representing a plurality of protein to protein docking poses based on one or more hot spots. The method can also include receiving a set of descriptors, the set of descriptors characterizing interaction interface features of one or more of the plurality of protein to protein docking poses. The method can also include a step of training, using the set of descriptors and the training data, a classification-based machine learning algorithm to generate a trained machine learning classification model, the trained machine learning classification model is able to evaluate a possible binding pose of two proteins and indicate if the possible binding pose is native, near native or a combination thereof.

[0005] The instant application also discloses, a method for classifying native and near native binding confirmations between two or more proteins. The method can include a step of receiving data representing a plurality of protein to protein docking poses based on one or more hot spots. The method can further include a step of receiving a set of descriptors, the set of descriptors characterizing interaction interface features of one or more of the plurality of protein to protein docking poses. The method can further include executing the trained machine learning classification model in order to evaluate a possible binding configuration of two proteins and indicate if the possible binding configuration is native, near native or a combination thereof.

[0006] The instant application also discloses, a computer implemented method for predicting binding poses between two or more proteins. The method can include a step of receiving a set of descriptors characterizing interaction interface features of a plurality of protein to protein docking poses. The method can further include a step of receiving data representative of the two or more 3 ME1\53706606.v1Attorney Docket No.136883-01220 proteins. The method can further include a step of generating as a first output a predicted protein- protein interaction for a first protein and a second protein using a trained machine learning regression model, the trained machine learning regression model taking as an input the set of descriptors and the data representative of the two or more proteins. The method can further include a step of generating as a second output a classification of a protein-protein interaction for the first protein and the second protein using a trained machine learning classification model, the trained machine learning classification model taking as an input the set of descriptors and the data representative of the two or more proteins. The method can further include a step of generating a predicted binding pose between the first protein and the second protein based on the first output generated by the trained machine learning regression model and the second output generated by the trained machine learning classification.

[0007] The instant application also discloses, a system comprising a memory holding computer readable instructions, and a processing device for executing the computer readable instructions, thereby causing the processing device to perform one or more operations of receiving a set of descriptors characterizing interaction interface features of a plurality of protein to protein docking poses. The one or more operations can further include receiving data representative of the two or more proteins. The one or more operations can further include generating as a first output a predicted protein-protein interaction for a first protein and a second protein using a trained machine learning regression model, the trained machine learning regression model taking as an input the set of descriptors and the data representative of the two or more proteins. The one or more operations can further include generating as a second output a classification of a protein-protein interaction for the first protein and the second protein using a trained machine learning 4 ME1\53706606.v1Attorney Docket No.136883-01220 classification model, the trained machine learning classification model taking as an input the set of descriptors and the data representative of the two or more proteins. The one or more operations can further include generating a predicted binding pose between the first protein and the second protein based on the first output generated by the trained machine learning regression model and the second output generated by the trained machine learning classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The foregoing features of the present disclosure will be apparent from the following Detailed description of the present disclosure, taken in connection with the accompanying drawings, in which:

[0009] FIG.1 is a diagram illustrating an example embodiment of a protein to protein interaction binding pose prediction machine learning system in accordance with the present disclosure;

[0010] FIG. 2A is a workflow diagram illustrating overall processing steps carried out by the disclosed systems and methods for predicting a binding pose in accordance with the present disclosure;

[0011] FIG.2B is a flowchart illustrating example processing steps carried out by the disclosed systems and methods for predicting a protein to protein interaction binding pose in accordance with the present disclosure;

[0012] FIG.3 is a chart of different unconstrained protein docking engines and their performance based on the number of docked complexes and number of poses in accordance with the present disclosure; 5 ME1\53706606.v1Attorney Docket No.136883-01220

[0013] FIG. 4 illustrates a process flow for using a regression model to predict a protein-protein interaction and using a classification model to predict true negative, true positive, false negative, and false positive protein-protein interactions in accordance with the present disclosure;

[0014] FIG.5 is a table of Protein Interaction Domains (PIDs) suitable for use with embodiments taught herein for predicting protein to protein interactions;

[0015] FIG. 6 illustrates a diagram illustrating different binding partners for the HECT domain suitable for use in some embodiments in accordance with the present disclosure;

[0016] FIG. 7 is a chart depicting a tossing rate of a test set of protein-protein docked poses corresponding to predicted protein-protein docking cutoff values in accordance with the present disclosure;

[0017] FIG. 8 is a chart depicting a protein-protein docking distribution of a training data set corresponding to protein-protein docking accuracy values in accordance with the present disclosure;

[0018] FIG. 9 is a chart depicting a protein-protein docking distribution of a training data set of binding and non-binding proteins corresponding to protein-protein docking accuracy values in accordance with the present disclosure;

[0019] FIG. 10 is a portion of a first histogram illustrating a comparison between near native protein-protein docking scores generated by conventional docking engine and near native protein- protein docking scores as taught herein; 6 ME1\53706606.v1Attorney Docket No.136883-01220

[0020] FIG. 11 is a portion of a second histogram illustrating a comparison between near native protein-protein docking scores generated by conventional docking engine and near native protein- protein docking scores as taught herein;

[0021] FIG. 12 is an example diagram illustrating computer hardware and network components on which the system can be implemented; and

[0022] FIG.13 is an example block diagram of an example computing device that can be used to perform one or more steps of the methods provided herein. DETAILED DESCRIPTION

[0023] The present disclosure relates to machine learning systems and machine learning methods for predicting protein-protein interaction. Example systems and methods are described in detail below in connection with FIGS. 1-13. In particular, the systems and methods disclosed herein specify a dedicated training data set based on a variety of single domain interactions. The training data set can comprise an information space required to create dedicated regression and classification models in order to identify near native binding configurations. A specific and diverse set of interaction interface descriptors can be calculated to provide the foundation for both, a regression and classification models.

[0024] PPIs are physical contacts of high specificity established between two or more protein molecules as a result of binding events steered by interactions that include, but not limited to electrostatic forces, hydrogen bonding and the hydrophobic effect. Predicting interactions between proteins and other biomolecules solely based on structure remains a challenge in biology. The 7 ME1\53706606.v1Attorney Docket No.136883-01220 machine learning systems and machine learning methods of the present disclosure utilize a regression model and a classification model to predict binding poses of two or more proteins as well as indicate if the one or more predicted binding poses is or represents a native binding pose, a near native binding pose or a combination thereof. The classification model and regression model utilize descriptors characterizing interaction interface features of each docking pose identified in a filtered corpus of docking poses. The filtering of the docking poses is performed using a hotspot prediction analysis. The hotspot prediction analysis is performed on docking poses from an unconstrained docking analysis of two or more proteins. The generated docking poses of the two or more proteins can be clustered together according to their relative orientation relative to a base protein or a receptor protein. Within each cluster, the poses are ranked according to their predicted binding affinity obtained from the regression model. This systematic approach can advantageously predict the binding poses of two proteins, proteins A and B, facilitating the identification of potential protein-protein complexes with biological significance.

[0025] As used herein, “protein–protein interactions” (PPIs) are specific, physical, and intentional interactions between the interfaces of two or more proteins as the result of biomolecular events / forces. The interaction interface should be non-generic, i.e., evolved for a purpose distinct from generic functions such a protein production, degradation, aggregate formation, and the like. In one aspect, the biomolecular events / forces include one or more non-covalent interactions such as, e.g., hydrogen bonding, electrostatic interactions, hydrophobic interactions, etc.

[0026] As used herein, a “hot spot” refers to a specific region on a protein surface, the specific region being more likely than not to result in a useful protein to protein interaction. More 8 ME1\53706606.v1Attorney Docket No.136883-01220 specifically, a “hot spot” refers to a collection of residues that makes a significant contribution to the binding free energy of a protein.

[0027] As used herein, a “surface patch” is defined by a collection of mesh elements (e.g., a polygon mesh element having vertices, edges, faces, etc.) pooled as a result of application of district criterion (e.g., a collection of surface points with similar and / or predefined geometric / chemical properties). An example of a “surface patch” is discussed below in relation to FIGS.3A and 3B.

[0028] As used herein, an “interaction patch” is defined as a collection of surface points of a coherent region of a particular type (e.g., positive charge, negative charge, hydrophobicity) involved in protein-protein interactions.

[0029] As used herein, a “biomolecule” refers to a molecule which is produced by a living organism and includes, but is not limited to carbohydrates, proteins, nucleic acids (DNA and RNA), lipids and polysaccharides.

[0030] As used herein, a “molecular surface” refers to a surface which an exterior probe-sphere touches as it is rolled over the spherical atoms of that molecule.

[0031] As used herein, a “neural network” refers to an artificial neural network having an input layer, one or more hidden layers, and an output layer. Each of the input layer, hidden layers, and the output layer may include a plurality of nodes, or artificial neurons, of the artificial neural network. And each of the plurality of nodes may be connected to one, more, or all of the other plurality of nodes with an associated weight and threshold. If the output of any individual node is 9 ME1\53706606.v1Attorney Docket No.136883-01220 above the specified threshold value, that node is activated, sending data to the next layer of the network. Otherwise, no data is passed along to the next layer of the network.

[0032] As used herein, “near-native” refers to an energetically favored confirmation of a protein, or complex of proteins (e.g., as in the case of protein-protein interactions forming a complex of two or more proteins), which is predicted based on the methods and systems taught herein. In some aspects, “near-native” is the predicted energetically favored confirmation of a protein, or complex of proteins which is most likely to occur in nature or that exists in nature.

[0033] As used herein, “native” refers to the energetically favored confirmation of a protein, or complex of proteins (e.g., as in the case of protein-protein interactions forming a complex of two or more proteins), that exists in nature. In certain aspect, the native confirmation is predicted based on the methods and systems taught herein.

[0034] As used herein, “glue” or “molecular glue” refers to synthetic small molecules which enhance the ubiquitination of a target protein by inducing and / or stabilizing the interaction between the target protein and an DE3 ligase have been reported and are referred to as “molecular glues”. See e.g., Guoqiang Dong et al., J. Med. Chem.2021, 64, 15, 10606–10620. Molecular glues are distinct from PROTACS in that PROTACS are bifunctional molecules, i.e., the molecules have one moiety that selectively binds a target protein and another moiety that recruits an E3 ubiquitin ligase, each moiety being separated from each other on the molecule by a linker. Molecular glues, on the other hand, are small chemical entities made up of a single molecule. While both PROTACS and molecular glues are effective, molecular glues are inherently more advantageous than PROTACS because they resemble traditional small molecule 10 ME1\53706606.v1Attorney Docket No.136883-01220 drug design principles. Being a smaller molecule, a molecular glue intrinsically poses qualities such as better solubility and membrane permeability, based on basic druggability principles. Molecular glues encourage, induce, or stabilize the interaction, or amplify the affinity, between proteins of interest.

[0035] Those skilled the art will appreciate that a protein structure is built as a plurality of chains of amino acids that are folded into a unique 3D shape. The plurality of chains of amino acids can be divided into side chains and a main chain (also referred to as a protein backbone). The amino acids are small organic molecules that consist of an alpha (central) carbon atom linked to an amino group, a carboxyl group, a hydrogen atom, and a variable component. An alpha carbon atom linked to a variable component forms a side chain. Within a protein, multiple amino acids are linked together by peptide bonds, thereby forming a long chain. Once linked in the protein, an individual amino acid is called a residue, the linked series of carbon, nitrogen, and oxygen atoms are known as the main chain or protein backbone, and the linked series of carbon atoms and variable components are known collectively as side chains. Multiple side chains have a great variety of chemical structures and properties. It is the combined effect of all of the amino acid side chains in a protein that ultimately determines its three-dimensional structure, its chemical reactivity and propensity to engage a protein-biomolecule interaction.

[0036] Turning to the drawings, FIG. 1 is a diagram illustrating an example embodiment of a system 100 of the present disclosure for predicting protein to protein interactions. The system 100 can be embodied as a central processing unit 102 (processor) in communication with a storage 104. Storage 104 can be a file system in some embodiments. Storage 104 can be memory in some embodiments. In some embodiments, storage 104 can comprise one or more hard disks. Further 11 ME1\53706606.v1Attorney Docket No.136883-01220 still, in some other embodiments, storage 104 can be a database. In some embodiments, the storage 104 is also referred to as storage 104. The processor 102 can include, but is not limited to, a computer system, a server, a personal computer, a cloud computing device, a smart phone, or any other suitable device programmed to carry out the processes disclosed herein. Still further, the system 100 can be embodied as a customized hardware component such as a field-programmable gate array (“FPGA”), an application-specific integrated circuit (“ASIC”), embedded system, or other customized hardware components without departing from the spirit or scope of the present disclosure. It should be understood that FIG. 1 is just one potential configuration, and that the system 100 of the present disclosure can be implemented using a number of different configurations.

[0037] As taught herein, when predicting protein-protein interactions the whole protein is modeled as a 3D model. The storage 104 can include various types of data including, but not limited to, protein to protein pairing structures, training data that includes domain-domain interaction configurations between proteins, for example, domain-domain interaction configurations between proteins facilitated by single chains, three-dimensional (3D) models, each 3D model representing a whole structure of a protein, preprocessed geometric property data and / or chemical property data associated with one or more 3D models, descriptor sets characterizing the interaction interface features of docking poses, some descriptor sets are made up of descriptors to facilitate computational speed, for example, 0.1 seconds per complex, some descriptor sets are made up of descriptors to facilitate graph networks and physics based calculations with a computational speed of about one to two seconds per complex. As will be described below in more detail, the one or more descriptors in the descriptor sets can include general descriptors describing the protein 12 ME1\53706606.v1Attorney Docket No.136883-01220 interfaces, descriptors describing the physics between atoms or molecules at the protein interfaces, pairwise sequence matrices, sequence profiles. In some embodiments the one or more descriptors can also include docking scores as well.

[0038] The storage 104 can also include system code 106. The system code 106 can include a structure preparation engine 132, an unconstrained docking engine 134, a pose filtering engine 136, a descriptor calculation engine 138, a classifier model 140, a regression model 142, and a clustering engine 144. In some embodiments the regression model 142 can be a trained neural network, for example a multilayer perceptron (MLP) neural network.

[0039] In some embodiments, the process of converting the 3D surface patches into 2D feature descriptors (also referred to as a dimensionality reduction) can be performed by several neural network algorithms including, but not limited to, multidimensional scaling (MDS) algorithms, singular value decomposition (SVD) algorithms, squeeze-and-excitation (SE) network algorithms, principal component analysis (PCA) algorithms. The dimensionality reduction projects pattern of proximities among a set of features (e.g., geometric properties and / or chemical properties) by providing feature values and distances between the feature values.

[0040] The storage 104 can also include one or more outputs (i.e., results) from various components, engines and models of the system 100 (e.g., outputs from the structure preparation engine 132, the unconstrained docking engine 134, the pose filtering engine 136, the descriptor calculation engine 138, the classifier model 140, the regression model 142, the clustering model 144, and / or other suitable components of the system 100). 13 ME1\53706606.v1Attorney Docket No.136883-01220

[0041] The system 100 includes the system code 106 (non-transitory, computer-readable instructions including executable code) stored on a computer-readable medium, for example storage 104 in FIG. 13, and executable by the hardware processor 102 or one or more computer systems. In some embodiments, the storage 104 is implemented as the storage 104. The system code 106 can include various custom-written software modules that carry out the steps / processes described herein, and can include, but is not limited to, the structure preparation engine 132, the docking engine 134, the filtering engine 136, the descriptor engine 138, the classifier model 140, the regression model 142, and the clustering model 144. Each component of the system 100 is described with respect to FIGS.2-13.

[0042] The system code 106 can be programmed using any suitable programming language including, but not limited to, C, C++, C#, Java, Python, or any other suitable language. Additionally, the system code 106 can be distributed across multiple computer systems in communication with each other over a communications network, and / or stored and executed on a cloud computing platform and remotely accessed by a computer system in communication with the cloud platform. The system code 106 can communicate with the storage 104, which can be stored on the same computer system as the system code 106, or on one or more other computer systems in communication with the system code 106.

[0043] FIG. 2A is a computational workflow diagram 200 illustrating overall processing steps carried out by the system 100 for predicting a binding pose of a protein to protein interaction as taught herein. The workflow diagram 200 for predicting a binding pose of a protein to protein interaction of two proteins to form protein-protein complexes involves a systematic series of steps, seamlessly integrating structural preparation, molecular docking, hotspot prediction, machine- 14 ME1\53706606.v1Attorney Docket No.136883-01220 learning descriptor calculation, and the application of classifier models, regression models and clustering models as taught herein. The following description provides a description of each step in the workflow.

[0044] At block 202 of the workflow 200, a protein of interest (POI) or proteins of interest (POI), for example Protein A, or Protein B or Proteins A and B, are identified and used as input for the structure preparation engine 132.

[0045] At structure preparation block 204, a process of preparing three-dimensional structures of two or more proteins, for example Protein A and Protein B, that may interface and bind together is started based on the POI. Preparation can include tasks such as removing water molecules from the three-dimensional structures, adding missing atoms or residues to the three-dimensional structures, and optimizing the three-dimensional structures for accurate representation in subsequent computational analyses. The process in block 204 may be implemented by structure preparation engine 132.

[0046] At unconstrained protein docking block 206 of the workflow 200, the docking engine 134 can receive as input the prepared three-dimension structures from the structure preparation process 204 as well as receive input pairing structures from a non-transitory storage medium, for example, a storage 104 The contents of the storage 104 can include a set of target proteins, for example, E3 ligases, whose structures have been prepared to avoid redundant operations prior to protein docking. The structure of a protein of interest (POI) is also prepared in order to ensure that the structures of the target proteins for docking have similar structural quality. 15 ME1\53706606.v1Attorney Docket No.136883-01220

[0047] The protein of interest (POI) as well as the docking engine 134 can be selected via a user interface in order to model an unconstrained docking of the two proteins. In some embodiments, known docking engines such as PatchDock, or PIPER, or FTDock can be executed to perform the unconstrained docking of Protein A and Protein B. The selected docking engine 134 generates a series of docking poses based on the various spatial orientations and conformations between the two proteins. The resolution for changing the orientation against each other is often in the order of 1 Angstrom ensuring a maximum coverage of the interaction space of the two proteins. In some embodiments, between 100 – 150K poses can be generated for a given protein-protein pair. The process in block 204 may be implemented by the structure preparation engine 132. In some embodiments, more than one docking engine is used, for example, two, three, four or more.

[0048] At pose filtering block 210 of the workflow 200, the docking poses from the unconstrained protein docking block 206 can be filtered based on a hotspot prediction analysis. Hotspots refer to regions or areas on the two proteins that are likely to have or form a stable interaction. Hotspots are important for forming stable interactions. The procedure of predicting hotspots is described in U. S. Patent Application No.18 / 376729 and also described below. The pose filtering performed at the pose filtering block 210 is executed by the hotspot prediction engine 130.

[0049] In some embodiments, instead of using arbitrary radii for determining the surface patches, the hot spot prediction model 130 can use interaction patches. An interaction patch can be defined as a collection of surface points of a coherent region of a particular type (e.g., positive charge, negative charge, hydrophobicity) involved in protein-protein interactions. For example, in some embodiments, a hydrophobic patch can be calculated by projecting a hydrophobic potential of each atom onto a protein surface. 16 ME1\53706606.v1Attorney Docket No.136883-01220

[0050] In some embodiments, the hot spot prediction model 130 can utilize an energy decomposition method (e.g., an eigenvalue decomposition method) to decompose an interaction energy matrix associated with a protein (e.g., an interaction matrix that includes residue information accounting for van der Waals energy, electrostatic energy, hydrogen bonds energy, hydrophobic interaction, or some combination thereof) into eigenvalues to identify residues within the protein which contribute significantly to the stability of the protein and / or having strong couplings to interact with a biomolecule. For example, the components of the eigenvector associated with the lowest eigenvalue indicate which residues are likely to be responsible for the stability and for the rapid folding of the protein.

[0051] The hot spot prediction model 130 can then determine for at least one of the plurality of different surface patches at least one of a geometric property or a chemical property. In some embodiments, the hot spot prediction model 130 determines for each of the plurality of different surface patches at least one of a geometric property or a chemical property. For example, for each point of a surface patch (e.g., a vertex of each mesh element in the surface patch), a geometric property calculator can calculate geometric properties, and a chemical property calculator can compute chemical properties. Examples of a geometric property can include a shape index (describes a shape around each point on the surface patch with respect to the local curvature), and a distance-dependent curvature (describes a relationship between a distance to a center of a surface patch and surface normals of each point and the center point).

[0052] Examples of a chemical property can include properties associated with electrostatics, hydrophobicity, and hydrogen bonds. 17 ME1\53706606.v1Attorney Docket No.136883-01220

[0053] The hot spot prediction model 130 processes at least one of a collection of geometric properties collected from one or more of the plurality of different surface patches or a collection of the chemical properties collected from the one or more of the plurality of different surface patches or both are used to predict one or more hot spots on a surface of the protein that are highly likely involved in an interaction between the protein and a biomolecule. The hot spot prediction engine 130 can input the 3D surface patches having geometric properties and / or chemical properties to the descriptor engine 138. The descriptor engine 138 converts the 3D surface patches into feature descriptors (e.g., a number, a vector, a matrix or a string).

[0054] In some embodiments, the hot spot prediction model 130 can compare an output with a hot sport threshold to determine whether or not an input surface patch is highly likely to be a hot spot. A hot spot threshold refers to a value or a value range indicating that an input surface patch is highly likely to be a hot spot. For example, if the hot spot prediction model 130 determines that an output of the neural network satisfies the hot spot threshold, the hot spot prediction model 130 determines that an input surface patch is highly likely to be a hot spot. If the hot spot prediction model 130 determines that an output of the neural network dissatisfies the hot spot threshold, the hot spot prediction model 130 determines that an input surface patch is not likely to be a hot spot.

[0055] In some embodiments, the hot spot prediction model 130 can predict one or more hot spots of a particular type (e.g., positive charge, negative charge, or hydrophobicity). For example, the hot spot prediction model 130 can place the predicted hot spots into a classification indicative of a particular type. 18 ME1\53706606.v1Attorney Docket No.136883-01220

[0056] The hotspot prediction engine 130 is able to retain poses that exhibit favorable interactions at predicted hotspots, and discard the poses that present unfavorable or unlikely interactions at the predicted hotspots. The process at block 210 can increase the likelihood of identifying biologically relevant binding configurations. The hotspot prediction engine 130 can determine a set of residues with the highest propensity prediction scores on a side of a designated receptor molecule for selection of possible binding partners within a maximum of 10 Angstrom distance to one of these hotspot residues. That is, the hotspot prediction engine 130 can identify the designated receptor molecule to be used to bind a protein to another protein based on the set of residues with the highest propensity prediction scores. The hotspot prediction engine 130 can consider proteins that are within 10 angstroms of the hot sport residues. All other binding configurations that do not match the filtering criteria are discarded.

[0057] At descriptor calculation block 212, of the workflow 200, descriptors characterizing interaction interface features of each docking pose output from the pose filtering 210 can be determined or computed by descriptor calculation engine 138. These descriptors capture structural and energetic properties, serving as input features for machine-learning models such as the regression model 142 and the classifier model 140. Examples of descriptors may include interface area, electrostatic potentials, and van der Waals interactions. In addition to these descriptors there can be additional descriptors including the following: interface_asaGeneral Surface area (in A2) between the two interacting r t in inME1\53706606.v1Attorney Docket No.136883-01220 median_iface_atom_dist General Median distance of the interacting atom pairs within the complex nr_res_pairs General Number of interacting residues engaged in protein / protein interaction nr_iface_residues General Number of unique residues in the interaction interface median_iface_res_distGeneral Median distance of the interactin residue airs of hotspot_score dfire_score ev_score mj3H_score mj2H_score mj2_scoremj1_score Docking Score pisa_scoreDocking Score PISA Protein Docking Score sd_score Docking Score An electrostatics and Van der Waals based scoring function as described in the SwarmDock publication, but using AMBER94 force-field charges and parameters. tobi_score Docking Score TOBI Protein Docking Score rrce55_res_score Docking Score Residue contact potentials at 5.5 Angstrom distance rrce65_res_score Docking Score Residue contact potentials at 6.5 Angstrom distance 20 ME1\53706606.v1Attorney Docket No.136883-01220 rrce75_res_score Docking Score Residue contact potentials at 7.5 Angstrom distance rrce85 res scoreDocking Score Residue contact potentials at 8.5 Angstrom ts ts sME1\53706606.v1Attorney Docket No.136883-01220 quasi_pot_solv2 Pair Matrices Quasichemical energy of transfer of amino acids from water to the protein environment22 ME1\53706606.v1Attorney Docket No.136883-01220 polarity_barley Sequence Profiles polarity granthamSequence

[0058] The protein interaction descriptors may be classified into distinct categories that showcase diversity in the setup of the interface characterization. These may include, but are not limited to, physics-based assessments such as the size of the binding surface area of the two binding partners (two binding proteins), electrostatics energy of the binding complex, hydrophobicity scores and salt bridge or hydrogen bond energetics. Other interface characterizations can include pairwise residue propensity scales used in modeling or simulating protein folding or in protein stability and mutability assessments that delineate certain preferences for interacting residue pairs that can be informative for a particular binding complex. In addition, residue propensities for many important aspects of protein integrity, such as stability, mutability, hydrophobicity, surface transfer energy for example, can be assessed for each binding residue within the binding complex. In some embodiments, the descriptors are grouped into sets. For example, a first set of descriptors may include descriptors selected for computational speed (0.1 seconds per complex). As another example, a second set of descriptors may include descriptors selected for use in graph networks and sophisticated physics based calculations (computation speed of about 1-2 seconds per complex). As another example, a third set of descriptors may include sequence based descriptors. 23 ME1\53706606.v1Attorney Docket No.136883-01220

[0059] An example of a group of descriptors selected for computational speed, for example, the first set, can include generic descriptors, physics-based descriptors, docking scores, pairwise sequence matrices and sequence profiles. These descriptors can be referred to as fast descriptors. An example of a fast descriptor can be a generic descriptor, which can include interface ASA, nr of interface atom / res pairs, median atom / res interface distances, or hotspot overlap (ML-based interface prediction). Another example of a generic descriptor is a physics-based descriptor which can include electrostatics, hydrophobicity, dipole moment, inertia moment, clashing, HBond, Salt- Bridge normalized against interface surface area. Yet another example of a general descriptor is a docking score. The docking score can include fast scores, from other dockers (pisa, MJ3H, MJ2H, MJ2, MJ1, tobi, sd, dfire). Another example of a general descriptor can include pairwise sequence matrices, in which only interface residues are considered. And another example of a general descriptor can include sequence profiles, which can include aggregation, disorder propensity, and mutability.

[0060] The calculated descriptors can be used by the classifier model 140, and the regression model 142 to determine whether the binding poses from the binding pose prediction engine 146 represent a near-native binding mode or not. The classifier model 140 and the regression model 142 are trained using the calculated machine-learning descriptors as described above. The training data set can be used as a learning base, and can be made up of 1068 protein complexes composed of two interacting chains.

[0061] In some embodiments, the training data set can include one or more of the following features. For example, the training data set can include non-redundant features, no internal skew, a complete or nearly complete coverage of observations, or an equal distributions of the observed 24 ME1\53706606.v1Attorney Docket No.136883-01220 events. In some embodiments, the training data set can include features that enable the machine- learning models to accurately predict weak-binding protein to protein interactions. In particular, the training data set can include data associated with binding events facilitated by single chains as observed in domain to domain interactions. Those skilled in the art that some of the data is available from publicly available databases, for example, the 3DID database at https: / / 3did.irbbarcelona.org. The 3DID database is a public resource for 3-dimensional protein to protein binding cataloging of various aspects of protein to protein interactions (e.g. interaction surface, binding topology, interacting residues etc.). In addition, the various binding modes are also grouped by protein family associations. The training data set can include a minimum set of protein structures while covering a broad range of distinct binding modes within the population of single chain / single chain interaction types. Consequently, in some embodiments, antibody / antigen e m the training data set, by definition single cha

[0062] T n Data Bank (PDB) tha odels can be trained o code and the two intera els disclosed herein areME1\53706606.v1Attorney Docket No.136883-01220 1F3V-B-A 1F80-B-D 1FFV-E-D 1FM0-E-D 1FQV-G-H 1G73-B-C 1GGP-B-A 1GK9-B-A 1GLA-G-F 1GO3-E-F 1HE1 A 1I4EAB 1IARBA 1I FAI 1IRAYX26 ME1\53706606.v1Attorney Docket No.136883-01220 2RVB-B-A 2UUY-A-B 2UYZ-A-B 2UZI-R-H 2V0X-B-A 2VRW-B-A 2VXQ-H-A 2W1T-A-B 2W2N-A-P 2W84-A-B 2WY AB 2WY7A 2WY A 2X 4A 2X AB27 ME1\53706606.v1Attorney Docket No.136883-01220 4E5Z-A-B 4EAI-A-B 4EDW-H-V 4EIG-A-B 4EIZ-B-C 4FQI-H-B 4G27-R-B 4G28-R-B 4G6U-A-B 4GI3-A-C 4H PAB 4HEPA 4HHHB 4HI AB 4I B28 ME1\53706606.v1Attorney Docket No.136883-01220 5VAY-D-H 5VEB-A-X 5VPG-A-D 5VWY-A-B 5VXJ-I-J 5WHE-A-B 5WPL-J-K 5WSV-A-B 5X1F-N-P 5X2N-A-H 5Z A Z BD Z HA Z HBD Z LA29 ME1\53706606.v1Attorney Docket No.136883-01220 1M9E-A-D 1M9X-B-C 1MAH-A-F 1MCT-A-I 1MVF-B-D 1NLN-A-B 1NTM-A-I 1NVI-E-D 1NW9-B-A 1OP9-B-A 1PN DB 1PNBBA 1PPFEI 1PP BI 1P ZAB30 ME1\53706606.v1Attorney Docket No.136883-01220 3BNW-A-B 3BP6-B-A 3BS5-A-B 3BTP-A-B 3BWV-B-A 3CBX-A-B 3CFI-E-F 3CI6-B-A 3CIT-A-B 3CJD-A-B 3DB BA DB BA DD AB DHIAE DTNAB31 ME1\53706606.v1Attorney Docket No.136883-01220 4LVO-A-C 4LZX-A-B 4M0W-A-B 4M3K-A-B 4M7E-A-C 4MQS-A-B 4MRT-A-C 4N1C-C-B 4N78-B-F 4NBB-A-D 4PK A 4PL AB 4P AB 4 IAD 4 D2AE32 ME1\53706606.v1Attorney Docket No.136883-01220 6FPO-M-T 6FV0-A-F 6G5Z-B-D 6GFZ-D-E 6GMN-A-B 6GWN-A-C 6H02-A-B 6H46-A-B 6H6Y-A-E 6H72-B-D 6IR2AB I THA IVZAH IYB D 4PAB

[0063] Si ible binding configura or both is to separate mber is often yielded in l 142 use the descriptor pose filtering engine 13

[0064] At more bindingposes of two or more proteins based on the descriptor set or sets. At block 216, the regression 33 ME1\53706606.v1Attorney Docket No.136883-01220 model 142 can predict the quality of with respect to the observable and offers more insights into the predicted binding mode. For example, classifier model 140 can determine that a pose between two proteins represents a native pose or a near-native pose or a combination thereof and thus their binding pose constitutes a native or near-native or combination thereof binding mode. The regression model 142 can determine that the quality of the determination of the binding pose, by determining a coefficient of determination (R2). When the coefficient of determination exceeds a predefined threshold, the combination of descriptors used by the regression model to predict the binding pose indicates that the binding pose can adequately be predicted by the combination of descriptors and therefore the regression model provides a certain level of quality in predicting the binding pose of the two proteins.

[0065] The classifier model 140 and the regression model 142 can be trained using the calculated machine-learning descriptors as described above as well as trained as follows. A training data set can be curated to be used as the learning base or corpus, and can comprise 1000 protein complexes composed of two interacting chains. The learning base can be used to train both the regression model 142 and the classifier model 140. Training of the machine-learning models (the regression model 142 and the classifier model 140) involves utilizing a labeled dataset of known protein- protein complexes to enable the models to generalize and predict the binding characteristics of new poses. The dataset used to produce the regression model 142 is made up of 1.1 million docking poses obtained from the benchmark of 1000 input structures. As observable, the DockQ protein docking evaluation score is used ranging from 0.0 (invalid model) to 1.0 (correct binding). Often, docking poses with a DockQ value of larger than 0.23 and 0.5 are considered acceptable, medium 34 ME1\53706606.v1Attorney Docket No.136883-01220 models within the range of 0.5 and 0.75 and very good models identified by DockQ values of 0.75 and higher. For the classification model 140, a DockQ cutoff value of 0.5 is employed.

[0066] The docking engine(s) 134 can produce a large number of possible binding configurations for protein-protein interactions. As a result the machine-learning models (the regression model 142 and the classifier model 140) are trained to separate plausible binding pose configurations from the native, near-native or combination thereof binding pose configurations for which small number are often yielded in response to the docking engine executing a docking run. The machine- learning models are trained to determine as many native, near-native or a combination thereof binding pose configurations. The machine-learning models (the regression model 142 and the classifier model 140) however do recognize and appropriately flag predicted, native, near-native or combination thereof binding pose configurations for the appropriate protein-protein pairings.

[0067] The regressor model 142 and the classifier model 140 are built on histogram-based gradient boosting engines. These fast estimators first bin input samples into integer-valued bins (typically 256 bins) which can reduce the number of splitting points to consider and allows the models to leverage integer-based data structures (histograms) instead of relying on sorted continuous values when building tree histogram.

[0068] At clustering block 218, the clustering model 144 can cluster and rank the output poses of the classifier model 140 and the regression model 142. Clustering is a procedure to pool similar binding configurations and use representatives of each cluster to further evaluate and rank. The docking poses that have structural similarity are clustered together based on the output of the regression model 142. Within each cluster, the poses can be ranked according to their predicted 35 ME1\53706606.v1Attorney Docket No.136883-01220 binding affinity obtained at the classifier model 140. This can allow for the identification of predicted native, near-native or a combination thereof binding poses that are both structurally diverse, and energetically favorable to interact. The clustering and ranking results are output at workflow block 220.

[0069] FIG. 2B is a flowchart illustrating example processing steps carried out by the disclosed machine learning systems and methods for predicting native, near-native or a combination there of protein to protein interactions in accordance with the present disclosure. At step 302, data associated with two or more three-dimensional (3D) models of a protein structure may be received. Each of the two or more 3D models corresponds to a different protein. At step 304, the received data is processed using one or more docking engines, to cause one or more unconstrained docking poses to be generated between the two or more different proteins corresponding to the 3D models. The resolution for changing the orientation against each other is usually in the order of 1 Angstrom ensuring a maximum coverage of the interaction space of the two entities. Often, 100 – 150K poses are generated for a given protein-protein pair. At step 306, the one or more docking poses can be filtered based on hot spot prediction between the two or more different proteins where the hot spots correspond to regions on the protein surfaces suitable for forming stable interactions between the two or more different proteins. At step 308, a set of protein interaction descriptors can be determined between each of the two or more different proteins. In some embodiments, the descriptors can include the presence of a small molecule including a molecular glue via a small molecule. At step 310, the classifier model 140 is able to evaluate a possible binding pose and predict a binding pose quality for each of the filtered poses with respect to the observable and offers more insights into the predicted binding mode, for example, if the predicted binding mode 36 ME1\53706606.v1Attorney Docket No.136883-01220 corresponds to native, near-native or combination thereof. At step 312, the regression model 140 using the descriptors sets predicts and identifies one or more binding poses of two or more proteins. At block 314, the clustering model 144 clusters the generated docking poses based on structural similarity. Within each cluster, the poses are ranked according to their predicted binding affinity obtained from the regression model 140.

[0070] FIG. 3 is a table 400 of different protein docking engines and their performance based on the number of docked complexes and number of poses in accordance with the present disclosure. The machine learning systems and methods disclosed herein in are based on the use of a docking engine able to produce results with high dockQ scores. Table 400 includes a protein docker 302 column which includes the name of the different docker engines that are available to use to simulate, or model, the protein-protein interface between two or more proteins. The table 400 also includes a number of configurations 304 column which includes the number of docked complexes that were simulated, or modeled, using any of the respective docking engines. The table 400 also includes a number of poses 306 column associated with the binding of two or more proteins along one or more interface between the two proteins.

[0071] The table 400 also includes a DockQ column 308 that includes the number of poses in which the accuracy of the docking along one or more interfaces between two or more proteins exceeds a docking score of 0.75. The DockQ score can be a score to evaluate the accuracy of the docking interface between the two or more proteins. The values of the DockQ score can range between 0.0 and 1.0 with a score of 1.0 representing a perfect location and a score of 0.0 representing a location that does not support docking. In other words, a DockQ score of 1.0 indicates that the accuracy between the two or more proteins along a protein-protein interface is 37 ME1\53706606.v1Attorney Docket No.136883-01220 completely correct. In other words, the location at which the two or more proteins are bound to one another is 100% accurate.

[0072] In some instances, the methods disclosed herein to simulate the interaction between proteins can generate clashes which occur when two non-bonded atoms are impossibly close to each other. This can happen when the van der Waals radii of the atoms overlap; that is, when two atoms are occupying the same space. If the methods disclosed herein simulate the interaction between proteins with clashes then more lower energy structures (e.g., more near native structures) can be generated than there would be if there were no clashes because there will be more options for the protein to interact. That is the model can force an interaction between proteins that otherwise might not occur. As a result the DockQ score in the DockQ > .75310 column could include the number of dockings that would result in clashes between the two or more proteins at a location, along an interface between the two or more proteins, in which the accuracy of the docking of the two or more proteins at the location along the interface exceeds 0.75. Consequently, the simulation or modeling of the specific location of the protein-protein interface may be slightly inaccurate. As a result, the table 400 includes a column for the two or more proteins that do dock and include clashes, with a DockQ score that exceeds .75.

[0073] The DockQ score can be based at least in part on a root mean square deviation between residues of chains on each of the two or more proteins at the location at which the two or more proteins interface. This metric can be referred to as the interface root mean square deviation (iRMSD). The DockQ score can further be based at least in part on a Ligand Root Mean Square (LRMSD) deviation. The LRMSD is based at least in part on the root mean square deviation of the shorter of a protein chain (ligand) of the model after superposition of the longer chain of the 38 ME1\53706606.v1Attorney Docket No.136883-01220 protein (receptor). The LRMSD is a measure of the average distance between the atoms of superimposed proteins. In some embodiments, the LRMSD is a measure of the difference between a crystal conformation of the ligand conformation and a docking prediction. The DockQ score can further be based at least in part on fraction of conserved native contacts (fNAT). The fNAT is the number of native (correct) residue–residue contacts between two or more proteins in the docked complex divided by the number of contacts in the original complex. The fNAT reflects the overlap between any interfaces shared between an original complex and a docked complex. A fNAT value of 1.0 indicates a complete overlap between the interfaces shared between the original and the docked complexes, and a value of 0.0 indicates that there are no overlapping interfaces between the original and docked complexes. The DockQ score can be based on at least in part on the iRMSD, LRMSD, fNAT, and a first distance (d1) and a second distance (d2) between. More specifically the DockQ score can be expressed as follows:!"#$%(&')*, +-. / !, 0-. / !, 12, 134 = (&')* 5 -. / !67)89:(+.- / !, 1245;-. / !67)89:(0.- / !, 1344<>.

[0074] In some embodiments, the docking engine selected by the user can be PIPER, PatchDock, or FTDock.

[0075] FIG. 4 illustrates alternative visual way of representing the workflow of FIG.s 2A and 2B in order to predict a protein-protein interaction and using the regression model 142 and theclassification model 140 in accordance with the present disclosure. FIG. 4 provides additional detail on a training or test set 501 which can be used by one or more of the models described herein to generate one or more scoring functions called protein interaction descriptors as described above. As discussed above, fast descriptors can include generic or general descriptors, sequence profiles 39 ME1\53706606.v1Attorney Docket No.136883-01220 descriptors, pairwise sequence descriptors, physics-based descriptors, and docker or docking scores. The protein interaction descriptors 503 can be used to train the regression model 142 and the classifier model 140.

[0076] The training or test set 501 can include data representing a combination of good and bad protein to protein binding interactions. Good protein-protein interactions in which the binding between the proteins is particularly strong, and bad protein-protein interactions in which the binding between the proteins is not particularly strong. The differences in binding mode can be categorized based at least in part on interface residues between the proteins. For example, the training data may be based on the interface residues associated with the Homologous to the E6- AP Carboxyl Terminus (HECT) domain. The categorization of the binding modes can be based on topology and the number of cases.

[0077] The regression model 142 predicts the binding poses for different protein pairs. For instance, as shown in the scatter plot associated with regression model 142, the horizontal axis can represent the different cutoff values of protein docking evaluation scores and the vertical axis can represent the number of near native poses associated with a given cutoff value. The relationship between the cutoff values and the protein docking evaluation scores is a continuous one in which for any value in the range of 0-1 for a given cutoff value there is a corresponding number of predicted native, near native or combination thereof poses that the regression model 142 can predict with some level of accuracy.

[0078] The classification model 140 can generate a binary prediction about whether the obtained binding configuration constitutes a native, near-native or combination thereof binding based on 40 ME1\53706606.v1Attorney Docket No.136883-01220 the protein interaction descriptors. Using the cutoff values of protein docking evaluation scores as a variable to determine what constitutes a native, near-native or combination thereof binding or not, a matrix of values can be created in which the actual near, near-native or combination thereof bindings can be compared to those predicted by the classification model 140 in a truth table 507, in order to evaluate the performance of the classification model 140 to accurately predict a native, near-native or combination thereof binding.

[0079] The truth table 507 includes four entries. The classification model 140 assigns a value of “0” to the native or near-native or combination thereof binding between at least one pair of proteins which means that the classification model 140 has predicted that there is no native or near-native or combination thereof binding between the at least one pair of proteins. The classification model 140 classifies a binding between the at least one pair of proteins as native or near-native or combination thereof based at least in part on a cutoff value of a protein docking evaluation score. As noted above, and reiterated here, the protein docking evaluation score is an observable value that is used to determine when there is an actual binding between the at least one pair of proteins. The classification model 140 is trained using training data in training / test set 501, and the training data is used to generate a distribution of the protein docking evaluation score. The distribution is a frequency associated with bins of the protein docking evaluation scores. That is, several bins of the protein docking evaluation scores can be created where the values in each of the bins is less than 1 but greater than 0. The distribution of the protein docking evaluation score can be arranged with the bins having the highest frequency of non-binding between the at least one pair of proteins to the bins having the lowest frequency of non-binding between the least one pair of proteins. This is illustrated in distribution 900 of FIG.8. 41 ME1\53706606.v1Attorney Docket No.136883-01220

[0080] The classification model 140 is trained using the training data to assign a value of “1” to the native or near-native or combination thereof binding between the least one pair of proteins which means that the classification model 140 has predicted that there is a native or near-native or combination thereof binding between the at least one pair of proteins. The classification model 140 is trained to predict that there is a native or near-native or combination thereof binding between the at least one pair of proteins based on the protein docking evaluation score being greater than 0.5. As noted above and reiterated here the protein docking evaluation score ranges from 0.0 (invalid model) to 1.0 (correct binding). Docking poses with a protein docking evaluation score between than 0.23 and 0.5 are considered acceptable, and scores between of 0.5 and 0.75 are considered medium, and scores exceeding 0.75 are considered to be very good in correctly identifying that there is a native or near-native or combination thereof biding between the at least one protein pair. The classification model 140 classifies a binding between the at least one protein pair as native or near-native or combination thereof if the value produced by the classification model 140 exceeds a protein docking evaluation score of 0.5.

[0081] The classification model 140 is trained using the training data to assign a value of “0” to the native or near-native or combination thereof binding between the least one pair of proteins which means that the classification model 140 has predicted that there is not a native or near-native or combination thereof binding between the at least one pair of proteins. The classification model 140 classifies a binding between the at least one protein pair as being non native or near-native or combination thereof if the value produced by the classification model 140 is less than a protein docking evaluation score of 0.5. 42 ME1\53706606.v1Attorney Docket No.136883-01220

[0082] Truth table 507 includes four entries each of which corresponds to a protein docking evaluation score and a predicted value for the native or near-native or combination thereof binding between the at least one pair of proteins. A true negative value (TN) corresponds to a situation in which the classification model 140 correctly predicts that there is not a native or near-native or combination thereof binding between the at least one pair of proteins when there in fact is not an actual native or near-native or combination thereof binding between the at least one pair of proteins as determined by the protein docking evaluation score. The TN arises when the classification model 140 assigns a value of 0 to a given binding between the at least one protein pair thereby indicating that the binding is not a native or near-native or combination thereof binding and the protein docking evaluation score is equal to 0 as well.

[0083] FIG.5 illustrates a list of different binding partners for one or more proteins in accordance with the present disclosure. FIG.5 is an overview of some of the more dominant protein families along with the respective number of structures within the families that are used to compile the 3D interaction database (3DID). The training data used disclosed herein includes these families excluding Immunoglobulins.

[0084] The biding partners can include, but are not limited to a V-set domain with 425 partners, Ras domain with 105 partners, a ubiquitin domain with 103 partners, Pkinase domain with 98 partners, WD40 domain with 90 partners, SBC_pac_1 domain with 81 partners, C1-set domain with 80 partners, an ANK_2 domain with 67 partners, a fn3 domain with 63 partners, and a RNA_pol_Rpb1_1 domain with 56 partners. The listed domains are available in the database of three-dimensional interacting domains (3did) which is publicly available at https: / / 3did.irbbarcelona.org / . 43 ME1\53706606.v1Attorney Docket No.136883-01220

[0085] The various domains are high-resolution three-dimensional structural templates representing three-dimensional interacting domains. In particular the various domains represent interactions between two globular domains as well as domain-peptide interactions. The structural templates are a representative collection of protein interactions that can be used to discern realistic binding interactions from non-binding interaction between proteins. The training data from the training or test set 501 can be generated using 2500 protein domain to domain binding configurations. The 2500 protein domain to domain binding configurations can be used to predict hot spots. And 1500 of the protein domain to domain binding configurations can be used to as a set of training data for a machine learning based pose ranking. The set of training data can include interactions mediated between two single chains of two distinct domains. In some embodiments, the set of training data does not include antibody / antigen interactions, peptide / protein interactions and multichain / single chain domain interactions.

[0086] FIG.6 illustrates a line diagram 700 illustrating the relationship between the HECT domain and different binding partners in accordance with the present disclosure. In particular the line diagram 700 shows the relationship between HECT and ubiquitin, HECT and ww, HECT and DUF913, HECT and UQ_con, as well as the interaction between ubiquitin and UQ_con.

[0087] FIG.7 is a chart depicting a histogram 800 of a tossing rate of a test set of protein-protein docked poses corresponding to predicted protein-protein docking cutoff values in accordance with the present disclosure. The tossing rate of the test set of protein-protein docked poses in a test set using PIPER. The histogram 800 includes a rate 801 for each of the DockQ Bins 803. The DockQ Bins 803 include bins of DockQ values. The PIPER test set included 124,000 docked poses, representing over 10% of the entire training data set. Additionally, an independent evaluation set 44 ME1\53706606.v1Attorney Docket No.136883-01220 of 564 complexes including 42 from E3 ligases in PDB and 71 from known molecular glue systems, were also included in the PIPER test. Each complex of the evaluation set is comprised of an average of 60000 docked poses. The training process for the regression method includes a dataset compilation of 953 complexes, with 1,002,620 docked poses used for training and 123,919 used for testing.

[0088] For each of the dockQ Bins 803 there is a percentage, frequency, or rate 801 associated with the occurrence of the dockQ scores in that bin. And for each bin there can be any combination of bindings that are native, near-native or a combination thereof 805, non near-native 807, and / or bindings that were discarded or tossed 809 by the regression model 142 as not being adequate to accurately predict the protein-to-protein interaction between the at least two-proteins.

[0089] The regression model 142 can include an Extended Tree Regressor, a MLP Regressor (Multiple Layer Perceptron), and a histogram based Gradient Boosting Regressor (HistGBR). In some embodiments, when a pDockQ cutoff value of 0.4 is applied to the test set more than 97% of all near-native docking configurations are retained, while tossing more than 90% of non-near native docking configurations.

[0090] FIG.8 is a chart depicting a protein-protein docking distribution 900 of a training data set corresponding to protein-protein docking accuracy values in accordance with the present disclosure. The protein-protein docking distribution 900 can include a frequency 901 at which a docQ values in docQBins 903 occur in the training data set. As discussed above, the classification model 140 based on a 2-class prediction system using a dockQ cutoff value of 0.5 is used to predict a protein-protein docking. In order to ensure that the predictions are accurate, training data sets 45 ME1\53706606.v1Attorney Docket No.136883-01220 and testing sets are created to evaluate the classification model's generalization on unseen data. Additionally, steps are taken to address class imbalances through weight adjustments during training, preventing bias towards the majority class. The training and test sets can be the same as for the regression model. HistGBC (Gradient Boosting Classifier with histogram-based learning) can be used as the machine learning algorithm. In some embodiments, relevant features contributing to the prediction task, involving scaling, normalization, or encoding categorical variables can be identified as the classification model is being executed. Further, in some embodiments class weight balancing can be employed to enhance the classification model's ability to discern patterns in a minority class. A minority class can be a subset of the training data that has a dockQ score that is greater than 0.5, and where the size of the subset of the training data with the dockQ score exceeding 0.5 is less than half the size of the entire training data.

[0091] FIG.9 is a chart depicting a protein-protein docking distribution 1000 of a training data set of binding and non-binding proteins corresponding to protein-protein docking accuracy values in accordance with the present disclosure. The protein-protein docking distribution 1000 can be divided into two segments. The first segment can correspond to all dockQ values that are less than 0.5 and the segment can correspond to all dockQ values that are greater than 0.5.

[0092] FIG. 10 is a histogram 1100 illustrating a comparison between protein-protein docking scores generated by a docking model and protein-protein interface near-native prediction values in accordance with the present disclosure. Approximately 1000 complexes were used in the training data set for PIPER and approximately 500 complexes that were not included in the training data set were used to assess model predictions. The classifier model 140 (URANK-classifier) correctly classifies near native poses with a 99% hit rate, and retains 97% of all near-native poses 46 ME1\53706606.v1Attorney Docket No.136883-01220 with a classifier model 140 median accuracy 95%. The regression model 142 (URANK-regressor) correctly predicts near-native poses with a 98% hit rate, and retains 92% of near-native poses for a dockQ cutoff of 0.4 with a root mean square error (RMSE) or 0.12. Whereas the top 150 docking scores for poses IFACE, DFIRE, TOBI, and PISA failed to retain near-native poses with a hit rate of less than 2%.

[0093] FIG. 11 is a histogram illustrating a comparison between protein-protein docking scores generated by a docking engine and protein-protein interface near-native prediction values in accordance with the present disclosure. The histogram 1200 shows the docking scores poses IFACE, DFIRE, TOBI, and PISA in addition to the hit rates for the classifier model 140 (URANK- classifier) and regression model 142 (URANK-regressor). The best performing docking score pose is DFIRE and generates a hit rate of 43%, whereas the classification model 140 correctly classifies a binding between two proteins with a hit rate of 100%, and retains 95% of all near-native bindings. The median model accuracy of the classification model 140 is 95%. The regression model 142 correctly predicts the protein-to-protein interface with a hit rate of 99%, retains 81% of near-native bindings, with dockQ cutoff value of 0.4 and a RMSE=0.12.

[0094] FIG. 12 is an example diagram illustrating computer hardware and network components on which the system 1900 can be implemented. The system 1900 can include a plurality of computational servers 1902a-1902n having at least one processor (e.g., one or more graphics processing units (GPUs), microprocessors, central processing units (CPUs), tensor processing units (TPUs), application-specific integrated circuits (ASICs), etc.) and memory for executing the computer instructions and methods described above (which can be embodied as system code 106). The system 1900 can also include a plurality of data storage servers 104a-104n for storing data. 47 ME1\53706606.v1Attorney Docket No.136883-01220 The computational servers 1902a-1902n, the data storage servers 1904a-1904n, and the user device 1910 can communicate over a communication network 1908. The system 1900 needs not be implemented on multiple devices, and indeed, the system 1900 can be implemented on a single device (e.g., a personal computer, server, mobile computer, smart phone, etc.) without departing from the spirit or scope of the present disclosure.

[0095] FIG.13 is an example block diagram of an example computing device 102 that can be used to perform one or more steps of the methods and execute one or models provided by example embodiments. The computing device 102 includes one or more non-transitory computer-readable media for storing one or more computer-executable instructions or software for implementing example embodiments. The non-transitory computer-readable media can include, but are not limited to, one or more types of hardware memory, non-transitory tangible media (for example, one or more magnetic storage disks, one or more optical disks, one or more USB flash drives), and the like. For example, memory 1106 included in the computing device 102 can store computer- readable and computer-executable instructions or software for implementing example embodiments. The computing device 102 also includes processor 2002 and associated core 2004, and optionally, one or more additional processor(s) 2002’ and associated core(s) 2004’ (for example, in the case of computer systems having multiple processors / cores), for executing computer-readable and computer-executable instructions or software stored in the memory 2006, in the storage 104 and other programs or software for controlling system hardware. Processor 2002 and processor(s) 2002’ can each be a single core processor or multiple core (2004 and 2004’) processor. The computing device 102 also includes a graphics processing unit (GPU) 2005. In some embodiments, the computing system 102 includes multiple GPUs. 48 ME1\53706606.v1Attorney Docket No.136883-01220

[0096] Virtualization can be employed in the computing device 102 so that infrastructure and resources in the computing device can be shared dynamically. A virtual machine 2014 can be provided to handle a process running on multiple processors so that the process appears to be using only one computing resource rather than multiple computing resources. Multiple virtual machines can also be used with one processor.

[0097] Memory 2006 can include a computer system memory or random access memory, such as DRAM, SRAM, EDO RAM, and the like. Memory 2006 can include other types of memory as well, or combinations thereof. A user can interact with the computing device 102 through a visual display device 2018, such as a touch screen display or computer monitor, which can display one or more user interfaces 2019. The visual display device 2018 can also display other aspects, transducers and / or information or data associated with example embodiments. The computing device 102 can include other I / O devices for receiving input from a user, for example, a keyboard or any suitable multi-point touch interface 2008, a pointing device 2010 (e.g., a pen, stylus, mouse, or trackpad). The keyboard 2008 and the pointing device 2010 can be coupled to the visual display device 2018. The computing device 102 can include other suitable conventional I / O peripherals.

[0098] The computing device 102 can also include one or more storage 104, such as a hard-drive, CD-ROM, cloud storage or other computer readable media, for storing data and computer-readable instructions, applications, and / or software that implements example operations / steps of the system (e.g., the systems 100 and 1000) as described herein, or portions thereof, which can be executed to generate user interface 2019 on display 2018. Example storage 104 can also store one or more databases for storing any suitable information required to implement example embodiments. The databases can be updated by a user or automatically at any suitable time to add, delete or update 49 ME1\53706606.v1Attorney Docket No.136883-01220 one or more items in the databases. Example storage 104 can store one or more databases 2026 for storing provisioned data, and other data / information used to implement example embodiments of the systems and methods described herein.

[0099] The system code 106 as taught herein may be embodied as an executable program and stored in the storage 104 and the memory 2006. The executable program can be executed by the processor to perform the in-situ inspection as taught herein.

[0100] The computing device 102 can include a network interface 2012 configured to interface via one or more network devices 2022 with one or more networks, for example, Local Area Network (LAN), Wide Area Network (WAN) or the Internet through a variety of connections including, but not limited to, standard telephone lines, LAN or WAN links (for example, 802.11, T1, T3, 56kb, X.25), broadband connections (for example, ISDN, Frame Relay, ATM), wireless connections, controller area network (CAN), or some combination of any or all of the above. The network interface 2012 can include a built-in network adapter, network interface card, PCMCIA network card, card bus network adapter, wireless network adapter, USB network adapter, modem or any other device suitable for interfacing the computing device 102 to any type of network capable of communication and performing the operations described herein. Moreover, the computing device 102 can be any computer system, such as a workstation, desktop computer, server, laptop, handheld computer, tablet computer (e.g., the iPad®tablet computer), mobile computing or communication device (e.g., the iPhone®communication device), or other form of computing or telecommunications device that is capable of communication and that has sufficient processor power and memory capacity to perform the operations described herein. 50 ME1\53706606.v1Attorney Docket No.136883-01220

[0101] The computing device 102 can run any operating system 2016, such as any of the versions of the Microsoft® Windows® operating systems, the different releases of the Unix and Linux operating systems, any version of the MacOS® for Macintosh computers, any embedded operating system, any real-time operating system, any open source operating system, any proprietary operating system, any operating systems for mobile computing devices, or any other operating system capable of running on the computing device and performing the operations described herein. In some embodiments, the operating system 2016 can be run in native mode or emulated mode. In some embodiments, the operating system 2016 can be run on one or more cloud machine instances.

[0102] It should be understood that the operations and processes described above and illustrated in the figures can be carried out or performed in any suitable order as desired in various implementations. Additionally, in certain implementations, at least a portion of the operations can be carried out in parallel. Furthermore, in certain implementations, less than or more than the operations described can be performed.

[0103] In describing example embodiments, specific terminology is used for the sake of clarity. For purposes of description, each specific term is intended to at least include all technical and functional equivalents that operate in a similar manner to accomplish a similar purpose. Additionally, in some instances where a particular example embodiment includes multiple system elements, device components or method steps, those elements, components or steps may be replaced with a single element, component or step. Likewise, a single element, component or step may be replaced with multiple elements, components or steps that serve the same purpose. Moreover, while example embodiments have been shown and described with references to 51 ME1\53706606.v1Attorney Docket No.136883-01220 particular embodiments thereof, those of ordinary skill in the art will understand that various substitutions and alterations in form and detail may be made therein without departing from the scope of the present disclosure. Further still, other embodiments, functions and advantages are also within the scope of the present disclosure. 52 ME1\53706606.v1

Claims

Attorney Docket No.136883-01220 CLAIMS WHAT IS CLAIMED IS:

1. A method for training a regression-based machine learning algorithm to predict native and near-native binding confirmations between two or more proteins, the method comprising: receiving training data, the training data representing a plurality of protein to protein docking poses based on one or more hotspots; receiving a set of descriptors, the set of descriptors characterizing interaction interface features of one or more of the plurality of protein to protein docking poses; training, using the set of descriptors and the training data, the regression-based machine learning algorithm to generate a trained machine learning regression model, the trained machine learning regression model is able to identify one or more binding poses of two or more proteins.

2. The method of claim 1, wherein the trained machine learning regression model is able to evaluate the one or more binding poses and predict a docking pose quality metric.

3. The method of any one of claims 1-2, wherein descriptors of the set of descriptors are selected based on structural and energetic properties of a plurality of proteins.

4. The method of any one of claims 1-3, wherein descriptors of the set of descriptors are selected based on graph networks and physics-based calculations. 5 The method of any one of claims 2-4, wherein the set of descriptors further comprises general descriptors, physics-based descriptors, docker scores, pairwise sequences, and sequence profiles.

6. The method of any one of claims 1-5, wherein the regression-based machine learning algorithm is an extended tree regressor algorithm.

7. The method of any one of claims 1-6, wherein the regression-based machine learning algorithm is a multiple layer perceptron algorithm. 53 ME1\53706606.v1Attorney Docket No.136883-01220 8. The method of any one of claims 1-7, wherein the regression-based machine learning algorithm is a histogram-based gradient boosting regressor algorithm.

9. The method of any one of claims 1-8, wherein the training data comprises protein-protein interaction data for a plurality of proteins, each of the plurality of proteins having a plurality of docking poses.

10. The method of any one of claims 1, further comprising predicting the hotspots using a neural network trained on at least one of geometric property data or chemical property data projected onto an interaction surface of a protein to predict the predicted hotspots on a surface of the protein.

11. The method of any one of claims 1-10, wherein the trained machine learning regression model further receives a molecular glue as an input.

12. The method of claim 11, wherein the molecular glue enhances an affinity between two proteins.

13. The method of claim 1, wherein the one or more hotspots are predicted hotspots.

14. The method of any one of claims 1-13, wherein the one or more binding poses of two or more proteins represents a protein to protein complex that exists in native or near-native form.

15. A method for determining native and non-native binding confirmation of two or more proteins, the method comprising: receiving a set of descriptors, the set of descriptors characterizing interaction interface features of one or more of the plurality of protein to protein docking poses; receiving data representing a plurality of protein to protein docking poses based on one or more hotspots; and executing the trained machine learning regression model according to any of claims 1-14 to identify one or more binding poses of two proteins. 54 ME1\53706606.v1Attorney Docket No.136883-01220 16. A method for training a classification-based machine learning algorithm to classify native and near native binding confirmations between two or more proteins, the method comprising: receiving training data, the training data representing a plurality of protein to protein docking poses based on one or more hotspots; receiving a set of descriptors, the set of descriptors characterizing interaction interface features of one or more of the plurality of protein to protein docking poses; training, using the set of descriptors and the training data, a classification-based machine learning algorithm to generate a trained machine learning classification model, the trained machine learning classification model is able to evaluate a possible binding pose of two proteins and indicate if the possible binding pose is native, near native or a combination thereof.

17. The method of claim 16, wherein the classification-based machine learning algorithm is a gradient boosting classifier with histogram-based learning.

18. The method of any of claims 16-17, further comprising performing class weight balancing to enable the trained machine learning classification model to discern a pattern in a minority class.

19. The method of any of claims 16-18, wherein the set of descriptors comprises general descriptors, physics-based descriptors, docker scores, pairwise sequences, and sequence profiles.

20. The method of any of claims 16-19, wherein the training data comprises protein-protein interaction data for a plurality of protein complexes, each of the plurality of protein complexes having a plurality of docking poses.

21. The method of any of claims 16-20, further comprising predicting the one or more hotspots using a neural network trained on at least one of geometric property data or chemical property data projected onto an interaction surface of a protein to predict the predicted hotspots on a surface of the protein.

22. The method of any of claims 16-21, wherein the trained machine learning classification model further receives a molecular glue as an input. 55 ME1\53706606.v1Attorney Docket No.136883-01220 23. The method of claim 22, wherein the molecular glue enhances an affinity between two proteins.

24. A method for classifying native and near native binding confirmations between two or more proteins, the method comprising: receiving data representing a plurality of protein to protein docking poses based on one or more hotspots; receiving a set of descriptors, the set of descriptors characterizing interaction interface features of one or more of the plurality of protein to protein docking poses; and executing the trained machine learning classification model according to any of claims 15-23 to evaluate a possible binding configuration of two proteins and indicate if the possible binding configuration is native, near native or a combination thereof.

25. A computer implemented method for predicting binding poses between two or more proteins, the method comprising: receiving a set of descriptors characterizing interaction interface features of a plurality of protein to protein docking poses; receiving data representative of the two or more proteins; generating as a first output a predicted protein-protein interaction for a first protein and a second protein using a trained machine learning regression model, the trained machine learning regression model taking as an input the set of descriptors and the data representative of the two or more proteins; generating as a second output a classification of a protein-protein interaction for the first protein and the second protein using a trained machine learning classification model, the trained machine learning classification model taking as an input the set of descriptors and the data representative of the two or more proteins; and 56 ME1\53706606.v1Attorney Docket No.136883-01220 generating a predicted binding pose between the first protein and the second protein based on the first output generated by the trained machine learning regression model and the second output generated by the trained machine learning classification.

26. The method of claim 25, further comprising: preparing three-dimensional structures of the first protein and the second protein; and generating a filtered set of docking poses based on identified hotspots for the first protein and the second protein by discarding any docking poses that do not satisfy a filtering criteria, wherein generating the set of descriptors for each of the first protein and the second protein is based on the filtered set of docking poses.

27. The method of any of claims 25-26, wherein the first output is a predicted docking pose for the first protein and the second protein.

28. The method of any of claims 25-27, further comprising: clustering a plurality of predicted docking poses into a plurality of clusters based on structural similarity; and ranking the predicted docking poses within each of the plurality of clusters according to a predicted binding affinity for each of the predicted docking poses obtained from the trained machine learning regression model.

29. The method of any one of claims 25-28, wherein the predicted protein-protein interaction represents a protein to protein complex that exists in native or near-native form.

30. A system comprising: a memory holding computer readable instructions; and a processing device for executing the computer readable instructions, the computer readable instructions controlling the processing device to perform operations comprising: 57 ME1\53706606.v1Attorney Docket No.136883-01220 receiving a set of descriptors characterizing interaction interface features of a plurality of protein to protein docking poses; receiving data representative of the two or more proteins; generating as a first output a predicted protein-protein interaction for a first protein and a second protein using a trained machine learning regression model, the trained machine learning regression model taking as an input the set of descriptors and the data representative of the two or more proteins; generating as a second output a classification of a protein-protein interaction for the first protein and the second protein using a trained machine learning classification model, the trained machine learning classification model taking as an input the set of descriptors and the data representative of the two or more proteins; and generating a predicted binding pose between the first protein and the second protein based on the first output generated by the trained machine learning regression model and the second output generated by the trained machine learning classification.

31. The system of claim 30, wherein the predicted binding pose between the first protein and the second protein represents a protein to protein complex that exists in native or near-native form. 58 ME1\53706606.v1