System and method for predicting protein-protein interactions

The neural network-based system for predicting PPIs improves accuracy by using refined geometric and chemical property data to identify hotspots in protein structures, addressing the challenges of conventional methods in predicting weak interactions.

JP2025536243APending Publication Date: 2025-11-05TRIANA BIOMEDICINES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025520000
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-07
Filing Date
2023-10-04
Publication Date
2025-11-05

AI Technical Summary

Technical Problem

Conventional methods for predicting protein-protein interactions (PPIs) face challenges due to the complexity of protein structures and the need for specific interactions, often leading to inaccurate predictions, especially in weak binding scenarios, and lack of clear definitions for geometric and chemical properties.

Method used

The system employs a neural network trained on improved geometric and chemical property data, using 3D models of proteins to predict hotspots for interactions, with refined surface patch definitions and flexible feature normalization, incorporating weak binding scenarios and removing residue constraints.

Benefits of technology

This approach enhances the accuracy of PPI predictions by identifying critical interaction sites, improving the model's ability to handle weak interactions and reducing errors in geometric and chemical property calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536243000001_ABST
    Figure 2025536243000001_ABST
Patent Text Reader

Abstract

Provided herein are systems and methods for predicting protein-protein interactions to predict one or more hot spots on the surface of a protein. One example method can include receiving a 3D model representing the entire structure of the protein. The method can include determining distinct surface patches associated with the 3D model. The method can include determining at least one of geometric properties or chemical properties for at least one of the distinct surface patches. The method can further include assigning each node of the surface patch input feature to include a chemical property. The method can include using a neural network to process at least one of a collection of geometric properties collected from one or more of the distinct surface patches or a collection of chemical properties collected from one or more of the distinct surface patches to predict one or more hot spots on the surface of the protein.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 414,233, filed October 7, 2022, which is incorporated herein by reference in its entirety. [Background technology]

[0002] Protein-protein interactions (PPIs) underlie most biological processes and play a central role in the normal function of proteins in all living organisms. Predicting these interactions is crucial for understanding most biological processes (e.g., DNA replication and transcription, protein synthesis and secretion, signal transduction, and metabolism) and for developing new drugs. Because proteins are large molecules with complex three-dimensional structures, PPIs are highly specific in that they require numerous favorable interactions (e.g., proper hydrogen bonding, electrostatic interactions, and hydrophobicity) between each partner. Therefore, predicting PPIs using computational methods can be challenging and resource-intensive. Summary of the Invention

[0003] The present disclosure relates to improved systems and methods for predicting protein-protein interactions by predicting one or more hotspots as defined below.

[0004] The systems and methods taught herein address some of the technical problems of conventional protein-protein interaction prediction by using computational methods in conjunction with neural network training. Conventional systems and methods for predicting PPIs have several drawbacks. For example, conventional neural networks learn from general protein benchmark datasets when implementing machine learning-based algorithms, but do not specifically recognize small molecule binding sites or weak protein-protein bonds. Furthermore, due to at least simplistic and unclear definitions of the parameters used to calculate geometric and chemical properties, it is often difficult to introduce fundamental changes to the input features and retrain the neural network. For example, identifying the correct binding pose of a protein complex and systematically distinguishing it from the extremely large pool of plausible binding configurations is widely accepted as a highly complex task. Conventional algorithms and methods based on physical principles are computationally feasible using scoring functions with various levels of abstraction, which often leads to incorrect predictions of binding poses. To address the problems of conventional systems and methods, embodiments of the present disclosure improve the accuracy of calculating the geometric and chemical properties of entire three-dimensional (3D) structures representing entire proteins. Improvements in 3D structures representing whole proteins are achieved by providing clear definitions of parameters used to calculate geometric and chemical properties, more accurate parameters for calculating geometric and chemical properties (e.g., atomic partial charges and radii, atomic slogP propensity values, etc.), and more accurate molecular surface representations. Additional improvements for predicting PPIs as taught herein include modifying chemical features that inform protein binding and allowing flexibility in normalizing input features, allowing for the assignment of user-defined feature weights used to optimize the neural network. As discussed in more detail below, training datasets for training networks to predict PPIs are improved. For example, training datasets as taught herein treat whole proteins as 3D models.Furthermore, the training dataset as taught herein allows the model to take into account weak binding scenarios of PPIs. Still further, the training dataset as taught herein departs from the conventional practice of relying on (and, in some embodiments, removing) residue constraints when predicting weak binding scenarios of PPIs. The ability of the models taught herein to minimize or remove residue constraints provides the ability to add small molecule ligand-related data to the training set.

[0005] In one embodiment, the present disclosure provides an example method for predicting one or more hot spots on the surface of a protein that are likely to be involved in PPIs. The method includes receiving a 3D model representing the entire structure of the protein. The method includes determining a plurality of different surface patches associated with the 3D model. A surface patch as used herein is defined below. The method includes determining at least one of a geometric property or a chemical property for each of the plurality of different surface patches. The method further includes processing, using a neural network, at least one of a collection of geometric properties collected from one or more of the plurality of different surface patches or a collection of chemical properties collected from one or more of the plurality of different surface patches to predict one or more hot spots on the surface of the protein that are likely to be involved in interactions between the protein and a biomolecule. A biomolecule as used herein is defined below.

[0006] In another embodiment, the present disclosure provides an exemplary method for training a neural network for predicting protein-biomolecule interactions. The method includes training the neural network using a training set having a corpus of 3D models. Each 3D model represents the entire structure of an identified protein and has a plurality of distinct surface patches. Each of the plurality of distinct surface patches includes at least one of geometric or chemical properties associated with the identified protein. The method further includes deploying the trained neural network to predict one or more hotspots on the surface of the protein that are likely to be involved in interactions between the protein and a biomolecule (e.g., protein, RNA, DNA, or the like).

[0007] The foregoing features of the present disclosure will become apparent from the following detailed description of the disclosure, taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a diagram illustrating one exemplary embodiment of the system of the present disclosure. [Figure 2A] FIG. 2A is a flow chart illustrating the overall processing steps performed by the system of the present disclosure. [Figure 2B] FIG. 2B is a flowchart illustrating exemplary processing steps performed by the system of the present disclosure for predicting hot spots. [Figure 3A] FIG. 3A illustrates an example surface patch in a molecular surface representation of a protein. [Figure 3B] FIG. 3B is an enlarged view of one of the surface patches of FIG. 3A. [Figure 4A] FIG. 4A illustrates an exemplary molecular surface representation showing the positive charge of the protein of FIG. 3A. [Figure 4B] FIG. 4B illustrates an exemplary molecular surface representation showing the negative charge of the protein of FIG. 3A. [Figure 5]FIG. 5 illustrates an exemplary molecular surface representation depicting the hydrophobic properties of the protein of FIG. 3A. [Figure 6A] FIG. 6A illustrates an exemplary molecular surface representation showing the hydrogen bond acceptor region of the protein of FIG. 3A. [Figure 6B] FIG. 6B illustrates an exemplary molecular surface representation showing the hydrogen bond donor region of the protein of FIG. 3A. [Figure 7] FIG. 7 illustrates the hydrogen bond geometries used in the hydrogen bond energy potential given in equation (1). [Figure 8] FIG. 8 illustrates an exemplary molecular surface representation of a protein interface surface. [Figure 9] FIG. 9 is an exemplary flowchart illustrating the neural network training process performed by the system of the present disclosure. [Figure 10] 10A-10B are Table 1 showing an exemplary training set of protein-protein interaction pairs. [Figure 11] 11A-11C are Table 2 showing an exemplary training set of protein-DNA / RNA interaction pairs. [Figure 12] FIG. 12 is an example diagram illustrating computer hardware and network components on which the system may be implemented. [Figure 13] FIG. 13 is an illustrative block diagram of an example computing device that may be used to perform one or more steps of the methods provided herein. DETAILED DESCRIPTION OF THE INVENTION

[0009] The present disclosure relates to systems and methods for predicting protein-protein interactions. Exemplary systems and methods are described in detail below with reference to Figures 1-13.

[0010] PPI is a highly specific physical contact established between two or more protein molecules as a result of binding events driven by interactions including, but not limited to, electrostatic forces, hydrogen bonding, and hydrophobic effects. Predicting interactions between proteins and other biomolecules based solely on structure remains a challenge in biology. The disclosed system and method utilizes a neural network trained on geometric and / or chemical property data projected onto the protein's interaction surface to predict one or more hot spots on the surface of a protein. The system and method offer several advantages over conventional methods, including, but not limited to, optimized surface projections using improved geometric and / or chemical property data in electrostatics, hydrophobicity, and hydrogen bonding. Disclosed herein is an improved training set that improves the accuracy of models representing protein structures. In some embodiments, the improved protein model is capable of predicting weak protein interaction scenarios (e.g., 10-fold correlations indicating weak binding and / or low affinity). -4 A dissociation constant K greater than M D ) and have improved surface patch definitions instead of arbitrary radii, as further described below with respect to Figures 2A and 2B, or the like. In some embodiments, the improved protein models have weak or strong interactions between the protein and the biomolecule.

[0011] As used herein, a "protein-protein interaction" (PPI) is a specific, physical, and intentional interaction between the interfaces of two or more proteins as a result of a biomolecular event / biomolecular force. The interaction interface should be non-generic, i.e., evolved for a purpose other than general functions such as protein production, degradation, aggregate formation, and the like. In one embodiment, the biomolecular event / biomolecular force comprises one or more covalent or non-covalent interactions, such as hydrogen bonding, electrostatic interactions, hydrophobic interactions, etc.

[0012] As used herein, "hot spots" refer to specific regions on the surface of a protein that are more likely not to result in useful protein-protein interactions. More specifically, "hot spots" refer to collections of residues that make a significant contribution to the binding free energy of the protein.

[0013] As used herein, a "surface patch" is defined by a collection of mesh elements (e.g., polygonal mesh elements having vertices, edges, faces, etc.) pooled as a result of applying area criteria (e.g., a collection of surface points having similar and / or predetermined geometric / chemical properties). An example of a "surface patch" is discussed below in connection with Figures 3A and 3B.

[0014] As used herein, an "interaction patch" is defined as a collection of surface points of coherent regions of a particular type (eg, positive charge, negative charge, hydrophobicity) that are involved in protein-protein interactions.

[0015] As used herein, "biomolecule" refers to molecules produced by living organisms, including, but not limited to, carbohydrates, proteins, nucleic acids (DNA and RNA), lipids, and polysaccharides.

[0016] As used herein, "molecular surface" refers to the surface that an external probe sphere contacts as it rolls over the spherical atoms of the molecule.

[0017] As used herein, "neural network" refers to an artificial neural network having an input layer, one or more hidden layers, and an output layer. Each node (or artificial neuron) connects to another node and has an associated weight and threshold. If the output of any individual node is above a specified threshold, that node is activated and sends data to the next layer of the network; otherwise, no data is passed on to the next layer of the network. A convolutional neural network (CNN) is a class of neural networks in which hidden layers contain convolutional layers that convolve the input and pass the result to the next layer; pooling layers that reduce the size of the data by combining the outputs of neuron clusters in one layer into a single neuron in the next layer; and fully connected layers that connect all neurons in one layer to all neurons in another layer.

[0018] Turning to the drawings, FIG. 1 is a diagram illustrating one example embodiment of a system 100 of the present disclosure. System 100 can be embodied as a central processing unit 102 (processor) in communication with a database 104. Processor 102 can include, but is not limited to, a computer system, a server, a personal computer, a cloud computing device, a smartphone, or any other suitable device programmed to execute the processes disclosed herein. Still further, system 100 can be embodied as a customized hardware component, such as a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), an embedded system, or other customized hardware component, without departing from the spirit or scope of the present disclosure. Of course, FIG. 1 is merely one potential configuration, and system 100 of the present disclosure can be implemented using many different configurations.

[0019] As taught herein, when predicting protein-protein interactions, the entire protein is modeled as a 3D model. Database 104 includes various types of data, including, but not limited to, three-dimensional (3D) models (each 3D model representing the entire structure of a protein), preprocessed geometric and / or chemical property data associated with one or more 3D models, trained neural networks, and one or more outputs from various components of system 100 (e.g., output from surface projection engine 110, geometric property calculator 112, chemical property calculator 114, neural network training engine 120, training set generator 122, training module 124, neural network module 126, hotspot prediction engine 130, and / or other suitable components of system 100).

[0020] Protein structures are constructed as chains of amino acids that fold into unique 3D shapes. These chains can be divided into side chains and the main chain (also called the protein backbone). Amino acids are small organic molecules consisting of an alpha (central) carbon atom attached to an amino group, a carboxyl group, a hydrogen atom, and a variable moiety. The alpha carbon atom attached to the variable moiety forms the side chain. Within proteins, multiple amino acids are linked together by peptide bonds to form long chains. When combined in a protein, individual amino acids are called residues, and the connected series of carbon, nitrogen, and oxygen atoms is known as the main chain or protein backbone, and the connected series of carbon atoms and variable moieties is collectively known as the side chain. The side chains have a wide variety of chemical structures and properties. It is the combined effect of all the amino acid side chains in a protein that ultimately determines the protein's three-dimensional structure, chemical reactivity, and tendency to engage in protein-biomolecular interactions.

[0021] System 100 includes system code 106 (non-transitory computer-readable instructions) stored on a computer-readable medium (e.g., storage 1124 of FIG. 13 ) and executable by hardware processor 102 or one or more computer systems. System code 106 can include various custom-use software modules that perform the steps / processes described herein, and can include, but are not limited to, a surface projection engine 110, a geometric property calculator 112, a chemical property calculator 114, a neural network training engine 120, a training set generator 122, a training module 124, a neural network module 126, and a hotspot prediction engine 130. Each component of system 100 is described with respect to FIGS. 2-13 .

[0022] The system code 106 can be programmed using any suitable programming language, including, but not limited to, C, C++, C#, Java, Python, or any other suitable language. Furthermore, the system code 106 can be distributed across multiple computer systems that communicate with each other through a communications network and / or can be stored and executed on a cloud computing platform and remotely accessed by computer systems that communicate with the cloud platform. The system code 106 can communicate with a database 104, which can be stored on the same computer system as the system code 106 or on one or more other computer systems that communicate with the system code 106.

[0023] 2A is a flowchart illustrating the overall processing steps 200 performed by the system 100 of the present disclosure. In step 202, the system 100 receives a 3D model representing the entire structure of a protein (e.g., the protein structure described above with respect to FIG. 1). For example, the system 100 can retrieve the 3D model from the database 104.

[0024] In step 204, the system 100 determines a plurality of different surface patches associated with the 3D model. For example, the surface projection engine 110 of the system 100 can calculate a molecular surface (e.g., a solvent-excluded surface, a solvent-accessible surface, a discretized molecular surface, etc.) from the 3D model and generate a molecular surface representation for visualizing the molecular surface. A molecular surface is defined above. The generated molecular surface representation (e.g., a polygonal mesh having a plurality of mesh elements) can include a plurality of different surface patches. For example, the surface patches can be the result of collecting surface points based on a predefined geodesic radius (e.g., 5 angstroms (Å), 9 Å, 12 Å, or other suitable geodesic radius greater than 12 Å). The geodesic radius is the distance from the center of a geodesic circle on the surface to a point on the geodesic circle. The system 100 can determine the geodesic radius based on a particular application or a particular interaction type. For example, in some applications (e.g., PPI search, pocket classification), system 100 may select a geodesic radius of 12 Å to cover the surface area of ​​numerous PPIs. In some embodiments, system 100 may select a geodesic radius of 5 Å or 9 Å to generate small surface patches. In some embodiments, instead of selecting a surface patch with a predefined geodesic radius, system 100 determines an interaction patch as a collection of coherent regions of distinct biophysical types (e.g., positive and negative charges, hydrophobicity) involved in protein-protein interactions. System 100 may also consider only vertices of surface patches that have the same biophysical type within a predefined distance radius (e.g., 5 Å or 9 Å) in the neural network.

[0025] In some embodiments, the system 100 can input information from the same surface patch into a neural network for processing. Examples of molecular surface representations and surface patches are shown in Figures 3A and 3B.

[0026] In step 206, the system 100 determines at least one of a geometric property or a chemical property for at least one of the plurality of different surface patches. In some embodiments, the system 100 determines at least one of a geometric property or a chemical property for each of the plurality of different surface patches. For example, for each point of the surface patch (e.g., a vertex of each mesh element within the surface patch), the geometric property calculator 112 of the system 100 can calculate the geometric property, and the chemical property calculator 114 of the system 100 can calculate the chemical property. Examples of geometric properties can include a shape index (which describes the shape around each point on the surface patch in terms of local curvature) and a distance-dependent curvature (which describes the relationship between the distance to the center of the surface patch and the surface normal of each point and the center point).

[0027] Examples of chemical properties include electrostatic, hydrophobic, and properties associated with hydrogen bonding.

[0028] Electrostatically associated properties can include atomic partial charges and radii assigned by a protein force field (e.g., the Amber99 force field or other suitable AMBER (Assisted Model Building and Energy Refinement) force field) rather than the small molecule force field used in conventional systems and methods. Use of a protein force field by the methods and systems taught herein can improve the accuracy of electrostatically associated properties. System 100 can set a pH value and eliminate collisions prior to electrostatic calculations, which can reduce errors caused by the coarse surface grids used in conventional systems. An example of a molecular representation with electrostatic properties is described with reference to Figures 4A and 4B.

[0029] Properties associated with hydrophobicity can include atom slogP (increasing confidence in atom or hybrid partitioning for n-octanol / water) propensity values ​​that are atom-based and independent of natural amino acid context, allowing for reliable predictions using existing small molecules instead of using Kyte-Doolittle residue propensities for hydrophobicity used in traditional methods, which are simple and rely on reduced context. An example of a molecular representation with hydrophobic properties is described with reference to FIG. 5.

[0030] As taught herein, properties associated with hydrogen bonding can include negative values ​​representing donors, positive values ​​representing acceptors, hydrogen bond geometries defined by established force field definitions, and hydrogen bond energies defined by established force field definitions as described below. In comparison, conventional methods that are prone to errors caused by subtle variations in surface atoms often use unclear definitions for hydrogen bond geometries and energy scales. Examples of molecular surface representations with hydrogen bond properties are described with reference to Figures 6A and 6B.

[0031] In some embodiments, the hydrogen bond energy calculation is based on the following equation (1):

number

[0032] where sp 3 Donor and sp 3 For acceptor pairs,

number

number

number

number

[0033] In some embodiments, the system 100 also refines the interface input definition. The system 100 considers surface vertices of atoms that contact other chains of the protein (e.g., other side chains or the main chain). As a result, computational efficiency is improved, since at least atoms that do not contact other chains are not calculated. An exemplary molecular surface representation 800 representing atoms that contact other chains of the protein is illustrated with respect to FIG. 8.

[0034] In step 208, the system 100 uses a neural network 126 to process at least one of a collection of geometric properties collected from one or more of the plurality of different surface patches, or a collection of chemical properties collected from one or more of the plurality of different surface patches, which are used to predict one or more hotspots on the surface of the protein that are likely to be involved in interactions between the protein and a biomolecule. The hotspot prediction engine 130 of the system 100 can input the 3D surface patches having geometric and / or chemical properties into a neural network (e.g., a convolutional neural network, geometric deep learning, or other similar algorithm) through an input layer, a hidden layer (e.g., a convolutional layer followed by a series of fully connected layers), and an output layer. The neural network converts the 3D surface patches into feature descriptors (e.g., numbers, vectors, matrices, or strings) and further processes the feature descriptors to predict one or more hotspots.

[0035] In some embodiments, the hotspot prediction engine 130 can compare the output of the neural network to a hotspot threshold to determine whether the input surface patch is likely to be a hotspot. The hotspot threshold refers to a value or range of values ​​that indicates that the input surface patch is likely to be a hotspot. For example, if the hotspot prediction engine 130 determines that the output of the neural network satisfies the hotspot threshold, the hotspot prediction engine 130 determines that the input surface patch is likely to be a hotspot. If the hotspot prediction engine 130 determines that the output of the neural network does not satisfy the hotspot threshold, the hotspot prediction engine 130 determines that the input surface patch is unlikely to be a hotspot.

[0036] In some embodiments, the neural network can predict one or more hot spots of a particular type (e.g., positive charge, negative charge, or hydrophobicity). For example, the neural network can place the predicted hot spots into a classification indicating a particular type.

[0037] In some embodiments, the process of converting 3D surface patches into 2D feature descriptors (also called dimensionality reduction) can be performed by several neural network algorithms, such as, but not limited to, multidimensional scaling (MDS) algorithm, singular value decomposition (SVD) algorithm, compressed excitation (SE) network algorithm, principal component analysis (PCA) algorithm, etc. Dimensionality reduction projects patterns of proximity between a set of features (e.g., geometric and / or chemical properties) by providing feature values ​​and distances between feature values.

[0038] In some embodiments, surface patches with geometric and chemical properties (also called input features to the neural network) can be normalized to be in the range of -1 to 1 to reduce errors and allow for the assignment of user-defined feature weights used during neural network optimization.

[0039] In some embodiments, instead of using an arbitrary radius to determine surface patches, the system 100 can use interaction patches. An interaction patch can be defined as a collection of surface points of coherent regions of a particular type (e.g., positive charge, negative charge, hydrophobicity) involved in protein-protein interactions. For example, in some embodiments, hydrophobic patches can be calculated by projecting the hydrophobic potential of each atom onto the protein surface. Positive patches can be calculated by projecting the positive hydrophilic potential of each atom onto the protein surface. Negative patches can be calculated by projecting the negative hydrophilic potential of each atom onto the protein surface. An example of applying this concept in the context of identifying and predicting protein aggregation hotspots is published by Sankar et al., “AggScore: Prediction of aggregation-prone regions in proteins based on the distribution of surface patches,” Proteins (2018), 86:1147-1156. In some embodiments, the feature space of the neural network can be fed with information about members of the same interaction patch with an interaction radius of 5.0 Å instead of patches based on an arbitrary radius, which can also refine the neural network training as described with respect to FIG. 9.

[0040] In some embodiments, system 100 utilizes an energy decomposition method (e.g., an eigenvalue decomposition method) to decompose an interaction energy matrix associated with a protein (e.g., an interaction matrix containing residue information describing van der Waals energy, electrostatic energy, hydrogen bond energy, hydrophobic interactions, or some combination thereof) into eigenvalues ​​to identify residues within the protein that significantly contribute to protein stability and / or have strong connections for interacting with biomolecules. For example, the components of the eigenvector associated with the lowest eigenvalues ​​indicate which residues are likely responsible for protein stability and rapid folding. An example of this concept is discussed in Tiana et al., "Understanding the determinants of stability and folding of small globular proteins from their energetics," Protein Science (2004), 13:113-124 (where driver residues for protein stabilization and folding are identified), and is further demonstrated in the prediction of antibody / antigen interactions in Peri et al., "Surface energetics and protein-protein interactions: analysis and mechanistic implications," Scientific Reports (2016), 6:240-35.

[0041] In some embodiments, the system 100 can adjust the density of the molecular surface representation (e.g., the surface grid density) to increase the resolution of the surface grid. An example is described with respect to Figures 3A and 3B.

[0042] In some embodiments, system 100 can retain hydrogen atoms throughout the entire process, including the model creation process, training process, deployment process, and / or various applications, while conventional methods remove hydrogen atoms within interfaces (e.g., surface patches involved in PPIs). Retaining hydrogen atoms ensures consistent treatment of the protein system throughout the entire process, which results in greater accuracy of electrostatics and allows implicit assessment of the feasibility of collisions and binding configurations.

[0043] 2B is a flow chart illustrating processing steps 210 performed by the system of the present disclosure to predict hotspots. In step 212, system 100 receives a user-input structure of a protein, e.g., a 3D model as taught herein. An example is described with respect to the 3D model of FIG. 1 and step 202 of FIG. 2A.

[0044] In step 214, system 100 performs structure preparation, including chain assignment, as needed, to calculate the molecular surface of the protein. An example is described with reference to step 204 of FIG. 2A. In some embodiments, the input protein structure is the result of querying one or more public databases, such as the Protein Data Bank. However, protein structures available from public databases often contain incomplete information necessary for surface property calculations. Preparation of a protein structure for use as input as taught herein can be performed using a software application (e.g., Schrodinger Protein Wizard) that can complete protein structure information, for example, by adding hydrogen and potentially missing side chain atoms, assigning partial charges, and adjusting the charge based on the respective system pH (default is pH 7.1). In some embodiments where the input protein structure is composed of three or more input chains, system 100 can perform chain assignment, in which system 100 determines which of the input chains are in contact with each other and which combinations of these interacting chains will be used in subsequent surface calculations.

[0045] In step 216, the system 100 performs a surface calculation to determine a number of different surface patches, an example of which is described with respect to step 204 of Figure 2A.

[0046] In step 218, the system 100 performs property calculations to calculate geometric and chemical properties 220 of the protein, including shape index, electrostatics, hydrophobicity, and hydrogen bonding, an example of which is described with respect to step 206 of Figure 2A.

[0047] In step 222, system 100 normalizes the calculated property values, for example, to the range

[0001] , [-1 0], [-1 1], or other suitable normalization range. Examples are described with respect to FIGS. 4-6.

[0048] In step 224, system 100 assigns normalized values ​​to one or more of the plurality of different surface patches (e.g., each surface patch or a portion of a surface patch). Examples are described with respect to step 206 of Figure 2A. In some embodiments, system 100 can assign non-normalized values ​​to one or more of the plurality of different surface patches.

[0049] In step 226, the system 100 performs geodesic reduction to convert the surface patches into input vectors for the neural network model 126. Geodesic reduction is a method of dimensionality reduction, which can be done by projecting surface features (hydrogen bond donors / acceptors, electrostatic charge propensity, hydrophobic propensity, curvature) into a two-dimensional form that is better suited as input to neuron vectors for the neural network model 126.

[0050] In step 228 , the system 100 feeds the input vector into the neural network model 126 .

[0051] In step 232, system 100 predicts one or more hot spots on the surface of the protein, an example of which is described with respect to step 208 of Figure 2A.

[0052] Figure 3A illustrates exemplary surface patches 330A and 330B (also referred to as patches) in a molecular surface representation 340. Figure 3B is a close-up of surface patch 330A of Figure 3A. A portion of surface patch 330A overlaps a portion of surface patch 330B. Surface patches 330A and 330B have geodesic radii 332A and 332B (shown in Figure 3A), respectively. It should be understood that the shapes of surface patches 330A and 330B are shown for illustrative purposes, and that the shapes of surface patches can vary based on the protein.

[0053] Figure 4A illustrates an example molecular surface representation 400A representing positive charge values ​​on the protein of Figure 3 A. Each vertex of the mesh surface of molecular surface representation 400A is assigned a normalized charge value between 0 (indicating no charge) and 1 (the maximum normalized positive charge value indicating a strong positive charge).

[0054] Figure 4B illustrates an example molecular surface representation 400B representing negative charge values ​​on the protein of Figure 3A. Each vertex of the mesh surface of molecular surface representation 400A is assigned a normalized charge value between 0 (indicating no charge) and 1 (the maximum normalized negative charge value indicating a strong negative charge). Those skilled in the art will understand that other scales can be used to represent normalized charge values. For example (not shown in Figures 4A and 4B), charge values ​​can be normalized between -1 (the maximum normalized negative charge value indicating a strong negative charge) and 1 (the maximum normalized positive charge value indicating a strong positive charge).

[0055] Figure 5 illustrates an example molecular surface representation 500 representing a hydrophobic region on the protein of Figure 3A. Each vertex of the mesh surface of molecular surface representation 500 is assigned a hydrophobicity scalar value based on the atomic slogP propensity values. The hydrophobicity scalar value is normalized to be between 0 (indicating no hydrophilicity) and 1 (the maximum normalized hydrophobicity scalar value, indicating strong hydrophobicity). In some embodiments (not shown in Figure 5), the hydrophobicity scalar value can be normalized to be between -1 (the maximum normalized hydrophobicity scalar value, indicating strong hydrophilicity) and 1 (the maximum normalized hydrophobicity scalar value, indicating strong hydrophobicity).

[0056] FIG. 6A illustrates an exemplary molecular surface representation 600A representing hydrogen bond acceptor regions on the protein of FIG. 3A. All possible donors and acceptors with surface exposure are considered for hydrogen bond calculations. In some embodiments, less than all possible donors and acceptors with surface exposure can be considered for hydrogen bond calculations. Hydrogen bond geometries and energies are calculated based on Equation (1) and FIG. 7. A normalized value between 0 (indicating no acceptor) and 1 (the maximum normalized hydrogen bond energy value for an acceptor indicating an acceptor with a strong hydrogen bond) was assigned to each vertex of the mesh surface of molecular surface representation 600A.

[0057] FIG. 6B illustrates an exemplary molecular surface representation 600B representing the hydrogen bond donor region of the protein of FIG. 3A. All possible donors and acceptors with surface exposure are considered for hydrogen bond calculations. In some embodiments, less than all possible donors and acceptors with surface exposure can be considered for hydrogen bond calculations. Hydrogen bond geometries and energies are calculated based on Equation (1) and FIG. 7. A normalized value between 0 (indicating no donor) and 1 (the maximum normalized hydrogen bond energy value for a donor indicating a donor with a strong hydrogen bond) was assigned to each vertex of the mesh surface of molecular surface representation 600B. In some embodiments (not shown in FIGS. 6A and 6B), negative values ​​represent donors and positive values ​​represent acceptors. For example, values ​​can be normalized between −1 (the maximum normalized hydrogen bond energy value for a donor indicating a donor with a strong hydrogen bond) and 1 (the maximum normalized hydrogen bond energy value for an acceptor indicating a donor with a strong hydrogen bond).

[0058] Figure 7 illustrates the hydrogen bond geometry 700 used in the hydrogen bond energy potential given by equation (1), as described above. θ is the donor (N: nitrogen atom) 704-hydrogen 706-acceptor (O: oxygen atom) 708 angle, φ is the hydrogen 706-acceptor 708-base (C: carbon atom) 710 angle, d is the donor 704-acceptor 708 distance, and r is the hydrogen 706-acceptor 708 distance. Γ (not shown in Figure 7) is the angle between the normal to the plane defined by the bonds from the donor 704 and acceptor.

[0059] 8 illustrates an example molecular surface representation 800, which represents a protein interface surface on a protein. Each vertex of the mesh surface of molecular surface representation 800 is assigned a value that indicates the corresponding atom that contacts other chains of the protein. Only the surface vertices of atoms that contact other chains are considered for interface prediction, which is a prediction of surface patches involved in PPIs.

[0060] FIG. 9 is an example flowchart illustrating a neural network training process 900 performed by the system 100 of the present disclosure.

[0061] In step 902, the system 100 trains a neural network using a training set having a corpus of 3D models. Each 3D model represents the entire structure of an identified protein in the context of PPI (e.g., the protein structure described above with reference to FIG. 1). Each 3D model has a plurality of distinct surface patches. Each of the plurality of distinct surface patches includes at least one of geometric or chemical properties associated with the identified protein. For example, the neural network training engine 120 of the system 100 can perform the training step 900. The training set generator 122 can obtain the entire protein structure from an external source (e.g., a public source / database). In some embodiments, the entire protein structure can be based on a protein benchmark (e.g., Protein Benchmark V5) used in protein-protein docking (e.g., Vreven, T. et al., Journal of Molecular Biology, 2015, vol. 427, pp. 3031-41).

[0062] The training set generator 122 generates various 3D models of entire protein structures, surface patches with known / calculated geometric and / or chemical properties for the particular 3D models, calculated geometric and / or chemical properties for the particular surface patches, molecular surfaces with known / calculated geometric and / or chemical properties for the particular 3D models, known / labeled / identified interacting protein pairs with binder proteins and target proteins for various protein interaction scenarios (e.g., weak PPI, strong PPI, other suitable PPI), various scenarios of protein-biomolecular interactions, including, but not limited to, the protein-protein interaction pairs listed in Table 1 shown in Figures 10A-10D. Training sets can be created that include, but are not limited to, other known / labeled / identified interacting protein-biomolecule pairs for the purpose of the present invention, other protein-biomolecule interaction pairs (e.g., protein-deoxyribonucleic acid (DNA) / ribonucleic acid (RNA) interaction pairs) listed in Table 2 shown in Figures 11A-11C, known / labeled / identified interacting patch pairs having binder patches and target patches, known / labeled / identified non-interacting sets having target protein / biomolecule / patch and random protein / biomolecule / patch, known / labeled / identified interaction types associated with the above data, and other suitable application-specific training data. In some embodiments, known / labeled / identified interacting protein-biomolecule pairs, known / labeled / identified interacting patch pairs, and / or known / labeled / identified non-interacting sets may be found from public sources / databases, such as the RCSB protein database. In some embodiments, ligands, DNA, metals, and / or crystal shrinkage may be removed from the training set.

[0063] The training module 124 can provide a training set into the neural network to be trained. For example, the training module 124 can provide interacting proteins (e.g., two single-chain proteins with no ligand, no DNA, no metal, and / or no crystal contacts), other interacting protein-biomolecule pairs, non-interacting proteins, and / or other non-interacting protein-biomolecule groups into the neural network. The training module 124 can adjust weights and other parameters in the neural network during the training process to reduce discrepancies between the neural network's output and expected outputs. The trained neural network can be stored in the database 104 or the neural network module 126.

[0064] In step 904, the system 100 deploys the trained neural network to predict one or more hot spots on the surface of a protein. For example, the neural network training engine 120 can select a group of the training set as a validation set and apply the trained neural network to the validation set to evaluate the trained neural network. In another example, the system 100 can deploy the trained neural network to predict hot spots on the surface of a protein (e.g., an unidentified protein, a protein input by a user, an unknown protein, or a random protein).

[0065] 12 is an example diagram illustrating computer hardware and network components upon which system 1000 may be implemented. System 1000 may include multiple computation servers 1002a-1002n having at least one processor (e.g., one or more graphics processing units (GPUs), microprocessors, central processing units (CPUs), tensor processing units (TPUs), application-specific integrated circuits (ASICs), etc.) and memory for executing the computer instructions and methods described above (which may be embodied as system code 106). System 1000 may also include multiple data storage servers 1004a-1004n for storing data. The computation servers 1002a-1002n, the data storage servers 1004a-1004n, and the user device 1010 may communicate over a communications network 1008. System 1000, of course, need not be implemented on multiple devices, and in fact system 1000 can be implemented on a single device (e.g., a personal computer, a server, a mobile computer, a smartphone, etc.) without departing from the spirit or scope of the present disclosure.

[0066] 13 is an illustrative block diagram of an example computing device 102 that can be used to implement one or more steps of a method provided by an illustrative embodiment. The computing device 102 includes one or more non-transitory computer-readable media for storing one or more computer-executable instructions or software for implementing the illustrative embodiment. The non-transitory computer-readable media may include, but are not limited to, one or more types of hardware memory, non-transitory tangible media (e.g., one or more magnetic storage disks, one or more optical disks, one or more USB flash drives), and the like. For example, memory 1106 included within the computing device 102 may store computer-readable, computer-executable instructions or software for implementing the illustrative embodiment. The computing device 102 also includes a processor 1102 and associated cores 1104, and optionally one or more additional processors 1102′ and associated cores 1104′ (e.g., in the case of a computer system with multiple processors / cores), for executing the computer-readable, computer-executable instructions or software stored in the memory 1106, as well as other programs for controlling system hardware. Processor 1102 and processor 1102' can each be a single core processor or a multi-core (1104 and 1104') processor. Computing device 102 also includes a graphics processing unit (GPU) 1105. In some embodiments, computing system 102 includes multiple GPUs.

[0067] Virtualization can be employed in computing device 102 to enable dynamic sharing of infrastructure and resources within the computing device. Virtual machines 1114 may be provided to handle processes running on multiple processors so that the processes appear to be using only one computing resource rather than multiple computing resources. Multiple virtual machines can also be used on a single processor.

[0068] The memory 1106 may include computer system memory or random access memory, such as DRAM, SRAM, EDO RAM, and the like. The memory 1106 may also include other types of memory, or combinations thereof. A user may interact with the computing device 102 through a visual display device 1118, such as a touchscreen display or computer monitor, which may display one or more user interfaces 1119. The visual display device 1118 may also display other aspects, transducers, and / or information or data associated with the illustrative embodiments. The computing device 102 may include other input / output devices, such as a keyboard or any suitable multi-point touch interface 1108, a pointing device 1110 (e.g., a pen, stylus, mouse, or trackpad) for receiving input from a user. The keyboard 1108 and pointing device 1110 may be coupled to the visual display device 1118. The computing device 102 may include other suitable conventional input / output peripherals.

[0069] The computing device 102 may also include one or more storage devices 1124, such as a hard drive, CD-ROM, or other computer-readable medium, for storing data and computer-readable instructions, applications, and / or software that implement example operations / steps of the systems (e.g., systems 100 and 1000) or portions thereof as described herein, which may be executed to generate a user interface 1119 on a display 1118. The example storage device 1124 may also store one or more databases for storing any suitable information needed to implement the example embodiments. The databases may be updated by a user or automatically at any suitable time to add, delete, or update one or more items in the databases. The example storage device 1124 may store one or more databases 1126 for storing provided data and other data / information used to implement example embodiments of the systems and methods described herein.

[0070] System code 106 as taught herein may be embodied as an executable program and stored in storage 1124 and memory 1106. The executable program can be executed by a processor to perform in-situ inspection as taught herein.

[0071] Computing device 102 may include a network interface 1112 configured to interface with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), or the Internet) via one or more network devices 1122 through various connections, including, but not limited to, standard telephone lines, LAN or WAN links (e.g., 802.11, T1, T3, 56 kb, X.25), broadband connections (e.g., ISDN, Frame Relay, ATM), wireless connections, controller area networks (CAN), or any combination of any or all of the above. Network interface 1112 may include a built-in network adapter, a network interface card, a PCMCIA network card, a card bus network adapter, a wireless network adapter, a USB network adapter, a modem, or any other device suitable for interfacing computing device 102 to any type of network having the communications capabilities and the ability to perform the operations described herein. Furthermore, the computing device 102 may be any computer system, such as a workstation, desktop computer, server, laptop, handheld computer, tablet computer (e.g., an iPad® tablet computer), mobile computing or communication device (e.g., an iPhone® communication device), or other form of computing or telecommunications device having communications capabilities and sufficient processor power and memory capacity to perform the operations described herein.

[0072] Computing device 102 may run any operating system 1116, such as any version of the Microsoft® Windows® operating system, various releases of Unix and Linux operating systems, any version of MacOS® for Macintosh computers, any embedded operating system, any real-time operating system, any open source operating system, any proprietary operating system, any operating system for mobile computing devices, or any other operating system that runs on a computing device and is capable of performing the operations described herein. In some embodiments, operating system 1116 may run in native mode or in an emulated mode. In some embodiments, operating system 1116 may run on one or more cloud machine instances.

[0073] It will be understood that the operations and processes described above and illustrated in the figures can be performed or carried out in any suitable order, as desired, in various implementations. Additionally, in certain implementations, at least a portion of the operations can be performed in parallel. Moreover, in certain implementations, fewer or more operations than those described can be performed.

[0074] In describing the exemplary embodiments, specific terms are used for the sake of clarity. For descriptive purposes, each specific term is intended to include, at a minimum, all technical and functional equivalents that operate in a similar manner to accomplish a similar purpose. Additionally, in some instances where a particular exemplary embodiment includes multiple system elements, device components, or method steps, those multiple elements, components, or steps may be replaced with a single element, component, or step. Similarly, a single element, component, or step may be replaced with multiple elements, components, or steps that serve the same purpose. Furthermore, while exemplary embodiments have been shown and described with reference to specific embodiments thereof, those skilled in the art will recognize that various substitutions and modifications in form and detail may be made therein without departing from the scope of the present disclosure. Furthermore, other embodiments, features, and advantages are also within the scope of the present disclosure.

Claims

1. 1. A computer-implemented method for predicting one or more hot spots on the surface of a protein, comprising: receiving a three-dimensional (3D) model representing the overall structure of the protein; determining a plurality of distinct surface patches associated with the 3D model; determining at least one of a geometric property or a chemical property for at least one of the plurality of different surface patches; assigning a plurality of input features to each node of a first surface patch of the plurality of surface patches, the plurality of input features including one or more chemical features; and processing at least one of a collection of geometric properties collected from one or more of the plurality of different surface patches or a collection of chemical properties collected from one or more of the plurality of different surface patches using a neural network to predict one or more hot spots on the surface of the protein.

2. The computer-implemented method of claim 1 , wherein the geometric properties include a shape index and a distance-dependent curvature.

3. 2. The computer-implemented method of claim 1, wherein the chemical properties include atomic partial charges assigned by a protein force field, atomic radii assigned by the protein force field, atomic slogP propensity values ​​independent of natural amino acid context, negative values ​​representing donors, positive values ​​representing acceptors, hydrogen bond geometries defined by an established force field definition, and hydrogen bond energies defined by the established force field definition.

4. The computer-implemented method of claim 1 , wherein each of the plurality of different surface patches comprises a collection of surface points having similar chemical properties.

5. 1. A computer-implemented method for training a neural network to predict one or more hot spots on the surface of a protein, comprising: training a neural network using a training set having a corpus of three-dimensional (3D) models, each 3D model representing the entire structure of an identified protein and having a plurality of distinct surface patches, each of the plurality of distinct surface patches comprising at least one of geometric or chemical properties associated with the identified protein; and deploying the trained neural network to predict the one or more hotspots on the surface of the protein.

6. 1. A system for predicting one or more hot spots on the surface of a protein, comprising: a memory storing one or more instructions; a processor configured or programmed to execute the one or more instructions stored in the memory, the processor being configured to execute the instructions stored in the memory to perform the steps of any one of claims 1 to 4.

7. A non-transitory computer-readable medium storing instructions for predicting one or more hotspots on the surface of a protein, which, when executed by a processor, causes the processor to execute the instructions stored in a memory to perform the steps described in any one of claims 1 to 4.

8. 1. A system for training a neural network to predict one or more hot spots on the surface of a protein, comprising: a memory storing one or more instructions; a processor configured or programmed to execute the one or more instructions stored in the memory, the processor configured to execute the instructions stored in the memory to perform the steps of claim 5.

9. A non-transitory computer-readable medium storing instructions for training a neural network to predict one or more hot spots on the surface of a protein, which, when executed by a processor, causes the processor to execute the instructions stored in memory to perform the steps described in claim 5.