Graph neural networks for predicting micro pka
A deep graph neural network is employed to predict micro pKa values by processing molecular graph representations and incorporating functional group data, addressing the inefficiencies of traditional methods and achieving rapid and accurate predictions with reduced computational resources.
Patent Information
- Application Number
- PCT/EP2024/083113
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2024-11-21
- Publication Date
- 2025-06-26
AI Technical Summary
Existing methods for predicting micro pKa values are computationally intensive and inefficient, particularly when dealing with complex molecules, as they require solving complex equations and have high memory usage.
The use of a deep graph neural network that processes graph representations of molecules in different states to predict micro pKa values, incorporating functional group data for improved accuracy, and leveraging machine learning models to bypass the need for complex equation solving.
This approach enables fast and accurate prediction of micro pKa values with significantly lower computational resources and memory usage compared to traditional methods, facilitating high-throughput screening in drug discovery and materials science.
Smart Images

Figure EP2024083113_26062025_PF_FP_ABST
Abstract
Description
GRAPH NEURAL NETWORKS FOR PREDICTING MICRO PKABACKGROUND
[0001] This specification relates to predicting properties of a target molecule using a neural network.
[0002] Machine learning models receive an input and generate an output, e.g., a predicted output, based on the received input. Some machine learning models are parametric models and generate the output based on the received input and on values of the parameters of the model. Some machine learning models are deep models that employ multiple layers of models to generate an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non-linear transformation to a received input to generate an output.SUMMARY
[0003] This specification generally describes a system implemented as computer programs on one or more computers in one or more locations that can predict the micro pKa for each of one or more target atoms in a molecule.
[0004] Throughout this specification, a “molecule” can refer to a collection of atoms which are bonded together through chemical bonds. A molecule can be represented in any of a variety of possible ways, e.g., as a sequence of characters, e.g., a Simplified Molecular Input Line Entry System (SMILES) string.
[0005] An “embedding” refers to an ordered collection of numerical values, e.g., a vector, matrix, or other tensor of numerical values.
[0006] An “intermediate output” of a machine learning model refers to an output generated by one or more hidden layers of the machine learning model, e.g., one or more hidden layers of a neural network, or, more generally, a set of values that (i) are generated by the machine learning model as part of processing a given input to generate a predicted output for the given input but that (ii) are not the predicted output for the given input.
[0007] A “ring” in a molecule can refer to an arrangement of some or all of the atoms in the molecule in a cyclic configuration, where a series of atoms are connected by chemical bonds (e.g., single, double or hydrogen bonds) in a closed loop. An atom can be a member of more than one ring in some molecules.
[0008] A pKa may alternatively be referred to as an acid-base dissociation constant.
[0009] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages.
[0010] The system described in this application leverages a graph neural network to accurately generate respective micro pKa predictions for target atoms of target molecules. By more accurately predicting micro pKas, the system can enable improved performance on a variety of downstream tasks that benefit from accurate pKa predictions.
[0011] More specifically, the system described in this specification can effectively integrate a deep graph neural network, i.e., one that has significantly more graph neural network layers (also referred to as message passing blocks or message passing steps) than other graph neural networks that have been used for pKa prediction. For example, the system can use a graph neural network that has over ten graph neural network layers, e.g., between six and fifty layers, and that generates more accurate micro pKa predictions than systems that use graph neural networks with three or four graph neural network layers.
[0012] For example, when predicting the micro pKa of a given protonation site, the system can process respective graph representations of two different states of the target molecule, e.g., a protonated state in which a hydrogen atom (proton) has been gained by the target molecule at the protonation site and an unprotonated state of the target molecule in which a hydrogen atom (proton) has not been gained by the target molecule, using the graph neural network and make the prediction from the outputs for the two different states. This can allow the graph neural network to effectively incorporate information about the two possible states of the molecule when making a given micro pKa prediction. The unprotonated state of the target molecule can be electrically neutral, negatively charged, or even positively charged. For example, the target molecule may in some cases be an anion, e.g., such that protonation corresponds to the target molecule becoming electrically neural. Alternatively, the target molecule may itself have already been protonated one or more times, e.g., such that it is positively charged prior to (further) protonation.
[0013] Optionally, in order to further improve the accuracy of the micro pKa predictions, the system can incorporate features from a model that operates on functional group data for the target atom. That is, the system can generate the micro pKa predictions from both features that are the output of the graph neural network and features that are the output of another machine learning model that processes an encoding of the target atom that characterizes whether the target atom is included in various functional groups, e.g., amine, amide, carboxyl, nitrile, alcohol, phosphate, thiols and so on. By including the functional group features in the generation of the pKa prediction, the system incorporates information that iscomplementary to the information that is encoded in the graph network features, thereby improving the accuracy of the micro pKa predictions. As one example, the system can, through the functional group features, incorporate information about whether an atom is a Lewis base or has a partial negative charge and is therefore a more favorable site for protonation, or a Lewis acid or has a partial positive charge and is therefore a less favorable site for protonation.
[0014] The system described in this specification can leverage a machine learning model, in particular, a graph neural network, to generate predicted micro pKa values with significantly higher computational efficiency (e.g., in terms of memory and computing power) than approaches based on density functional theory (DFT) calculations, or ab initio calculations, or molecular dynamics (MD) simulations, or quantum mechanics / molecular mechanics (QM / MM) methods, e.g., for the reasons set forth next.
[0015] First, approaches based on DFT and ab initio calculations require significant computational resources due to their need to solve complex equations describing electron interactions at each atomic site. For instance, calculating the pKa of a single molecule with DFT might take hours or even days on a high-performance computing cluster. This high computational demand scales poorly with the size and complexity of molecules, especially for systems with multiple ionizable groups or extensive solvation effects. In contrast, the graph neural network described in this specification is trained on a dataset of molecules with known pKa values and can generalize learned relationships to make new predictions in seconds, as it bypasses the need to solve complex equations directly.
[0016] Additionally, the graph neural network-based approach for predicting micro pKa may be much less memory-intensive than MD simulations or QM / MM hybrid approaches. Memory usage in MD simulations grows with the number of atoms, time steps, and solvent molecules in the system, while the graph neural network require memory only to store the model parameters and input data. Once the graph neural network is trained, it can make predictions on new data with a small memory footprint. This is especially beneficial for large datasets, where processing power and memory usage are major constraints.
[0017] Another efficiency benefit of the graph neural network-based approach for predicting micro pKa is scalability. With alternative approaches, predicting the micro pKa for a new set of molecules involves starting the computation process from scratch. However, the graph neural network, once trained, can be deployed to predict micro pKa values for many molecules at once, allowing for high-throughput screening. This capability is particularly valuable in drug discovery and materials science, where there is a need to evaluate thousandsor even millions of compounds. Updating an graph neural network with new data is also straightforward, making it adaptable as more experimental pKa values become available.
[0018] In summary, the efficiency of the graph neural network approach for micro pKa prediction lies in its ability to approximate complex calculations without needing to reperform the underlying physics each time. The ability to predict new micro pKa values quickly and with minimal computational resources makes the graph neural network approach described in this specification a practical, scalable approach.
[0019] According to a first aspect, there is provided a method performed by one or more computers, the method comprising: obtaining data specifying a target molecule and a target atom of the target molecule; generating a first graph representation of the target molecule that represents a first state of the target molecule as a graph of the target molecule, wherein the graph comprises a plurality of nodes and a plurality of edges, wherein each node represents a respective atom in the target molecule and each edge connects a respective pair of nodes in the graph and represents a bond between a pair of atoms that are represented by the respective pair of nodes; processing the first graph representation using a graph neural network to generate a first graph network output; generating a second graph representation of the target molecule that represents a second state of the target molecule as a graph of the target molecule; processing the second graph representation using the graph neural network to generate a second graph network output; and generating, from at least the first graph network output and the second graph network output, a predicted micro pKa of the target atom of the target molecule. The method may therefore be for determining one or more micro pKas of a target molecule. The method can be performed as part of a method for determining pharmacokinetic or pharmacodynamic properties of a target molecule, e.g. a drug candidate.
[0020] In some implementations, the second graph representation also represents the target molecule as a graph that includes a plurality of nodes and edges, where each node represents a respective atom of the target molecule and each edge connects two nodes in the graph and represents a bond between the two nodes in the graph that are connected by the edge. That is, the second graph representation and the first graph representation can have the same graph structure, although the number and / or configuration of the nodes and edges of the graphs may differ. The first (or second) graph representations can comprise a respective node embedding of each node in the first (or second) graph and a respective edge embedding of each edge in the first (or second) graph.
[0021] In some implementations, the method further comprises: obtaining functional group data that identifies, for each of a plurality of functional groups, whether the target atom isincluded in the function group; generating latent features of the target atom from the functional group data, wherein generating, from at least the first graph network output and the second graph network output, a predicted micro pKa of the target atom of the target molecule comprises: generating, from at least the first graph network output, the second graph network output, and the latent features, the predicted micro pKa of the target atom of the target molecule.
[0022] In some implementations, generating latent features from the functional group data comprises: generating, from the functional group data, an encoding of the target atom; and processing the encoding of the target atom using a machine learning model, wherein the latent features are an intermediate output of the machine learning model.
[0023] In some implementations, the machine learning model has been trained to predict the micro pKa of the target atom from the encoding.
[0024] In some implementations, the first state of the target molecule is a protonated state of the target molecule.
[0025] In some implementations, generating, from at least the first graph network output and the second graph network output, the predicted micro pKa of the target atom of the target molecule comprises: generating combined features from at least the first graph network output and the second graph network output; and processing the combined features using a prediction neural network head to generate the predicted micro pKa of the target atom of the target molecule.
[0026] In some implementations, the second state of the target molecule is an unprotonated state of the target molecule.
[0027] In some implementations, the first graph representation comprises a respective node embedding of each node in the first graph and a respective edge embedding of each edge in the first graph.
[0028] In some implementations, the respective node embedding for each node in the first graph identifies whether or not the atom represented by the node is the target atom.
[0029] In some implementations, the respective node embedding for each node in the first graph identifies an atom type of the atom represented by the node.
[0030] In some implementations, the respective edge embedding for each edge in the first graph identifies a bond type of the bond represented by the edge.
[0031] In some implementations, the respective edge embedding for each edge in the first graph identifies whether the bond represented by the edge is a conjugated bond.
[0032] In some implementations, the respective edge embedding for each edge in the first graph identifies whether the bond represented by the edge is a rotatable bond.
[0033] In another aspect, there is provided a method of obtaining a drug, the method comprising: for each of a plurality of candidate drug molecules, using the method of the first aspect to predict a respective one or more micro pKas for the candidate drug molecule; and using the predicted micro pKas to select at least one of the candidate drug molecules as the drug.
[0034] In some implementations, the method comprises, for each of the candidate drug molecules, using the one or more predicted micro pKas to determine one or more pharmacokinetic or pharmacodynamic properties (e.g., biological half-lives) for the candidate drug molecule; and using the predicted pharmacokinetic or pharmacodynamic properties of the candidate drug molecules to select the one or more candidate drug molecules.
[0035] In another aspect, there is provided a method of obtaining a ligand, wherein the ligand is a drug or a ligand of an industrial enzyme. The method comprises: using the method of the first aspect to predict one or more micro pKas for protonation of a target molecule at a corresponding one or more target atoms; using the one or micro pKas to determine a protonation state defining whether the one or more target atoms of the target molecule are protonated under biological conditions (e.g., a particular pH or range of pHs, such as from 6.5-7.5); for each of one or more candidate ligands: (a) determining a predicted structure of a complex comprising the candidate ligand and the target molecule in the protonation state; and (b) evaluating an interaction of the candidate ligand with the target protein molecule dependent on the predicted structure. The method further comprises selecting one or more of the candidate ligands as the ligand dependent on a result of the evaluating. The target molecule may comprise a receptor or enzyme, and wherein the ligand is an agonist or antagonist of the receptor or enzyme. The ligand may comprise an antibody or aptamer and the target molecule comprises an antibody or aptamer target, in particular a virus or cancer cell protein, and the antibody or aptamer can then bind to the antibody or aptamer target to provide a therapeutic effect.
[0036] The method may further comprise synthesizing the drug or ligand, and optionally, testing for biological activity of the ligand or drug in vitro or in vivo. For example the ligand may be tested for ADME (absorption, distribution, metabolism, excretion, e.g., by a living organism or cell culture or tissue model) and / or toxicological properties, to screen out unsuitable ligands, e.g., in vivo or in vitro. The testing may include, e.g., bringing the candidate drug molecule (which may be a small molecule (e.g., having a molecular mass ofless than 1000 daltons), a polypeptide or polynucleotide ligand, and so on) into contact with the target molecule and measuring a change in expression or activity of the target molecule. One or more ligands or drugs can then be selected based on the results of the testing.
[0037] According to another aspect, there are provided one or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the methods described herein.
[0038] According to another aspect there is provided a system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the methods described herein.
[0039] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0040] FIG. 1 shows an example prediction system.
[0041] FIG. 2 is a flow diagram of an example process for predicting the micro pKa of a target atom in a target molecule.
[0042] FIG. 3 is a block diagram that shows an example of the operation of the prediction system.
[0043] FIG. 4 is a flow diagram of an example process for training the graph neural network.
[0044] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION
[0045] FIG. 1 shows an example prediction system 100. The prediction system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations in which the systems, components, and techniques described below are implemented.
[0046] The system 100 is configured to process data defining a target molecule 102 and data defining a target atom 104 of the target molecule 102 to generate a prediction of the micro acid-base dissociation constant (pKa) 114 of the target atom 104 of the target molecule 102.
[0047] That is, rather than predicting the macro pKa of the entire target molecule 102, the system 100 can predict the micro pKa of any given atom of the target molecule 104.
[0048] Generally, the macro pKa of a molecule measures the dissociation ability of the whole molecule.
[0049] The micro pKa, on the other hand, is a micro acid-base dissociation constant that measures the loss or gain of a proton at a specific protonation site, i.e., from a specific target atom, within the molecule. In this specification, loss or gain of a proton is also referred to as loss or gain of a hydrogen atom, but such references should be understood as including a concomitant decrease (for protonation) or increase (for deprotonation) in the total charge of the molecule.
[0050] The target molecule 102 can be any appropriate molecule, e.g., an organic molecule, an inorganic molecule, a small molecule (e.g., a molecule with a molecular mass of less than 1000 daltons), a synthetic molecule (i.e., a molecule that is not known to occur in nature), and so forth. In some examples, the target molecule can be a biomolecule, such as a deoxyribonucleic acid (DNA), a ribonucleic acid (RNA), a protein (polypeptide), carbohydrate, sugar, and so on.
[0051] The system 100 can receive the data defining the target molecule 102 and the target atom 104 from any appropriate source, e.g., from a user or from another system, by way of an appropriate interface, e.g., an application programming interface (API) or a user interface (e.g., a graphical user interface).
[0052] While this specification describes that the system 100 receives data specifying a single target atom 104 within the target molecule 102, the system 100 can generate respective micro pKa predictions for many different target atoms within the given target molecule 102. For example, in some implementations, when receiving data specifying the target molecule 102, the system can, either in parallel or in sequence, generate respective micro pKa predictions for each “heavy” atom within the target molecule 102. Generally, a “heavy” atom is any atom within the target molecule 102 that is not a hydrogen atom.
[0053] As another example, in some other implementations, the system receives the data specifying the target molecule 102 and a specific subset of target atoms 104 that are of interest, and can then generate, either in parallel or in sequence, respective micro pKa predictions for each atom within the specific subset.
[0054] After generating the micro pKa predictions 114 for each target atom 104 of the target molecule 102, the system 100 can, e.g., store the predictions 114 in a memory, or transmit data defining the predictions 114 over a data communication network, or provide datadefining the predictions 114 directly to a downstream system that performs downstream processing based on the predictions 114.
[0055] In some implementations, a user can input data defining a target molecule 102 into a user interface made available on a user device by the system 100. The user device can be, e.g., a personal computer or a tablet or a smartphone. The user interface by which the user inputs the target molecule 102 can be presented to the user, on the user device, as part of an application running on the user device. The target molecule input into the user interface of the user device can be transmitted to the system 100, e.g., by transmission over an appropriate data communications network, e.g., the internet.
[0056] The system 100 can process the target molecule to generate a set of pKa predictions for one or more target atoms for the target molecule, and then output data identifying the set of pKa predictions for the target molecule. For instance, the system 100 can transmit the data identifying the set of pKa predictions for the target molecule back to the user device, e.g., over a data communications network, and a user can access the set of pKa predictions by way of the user device, e.g., by viewing data characterizing the set of pKa predictions using an application running on the user device.
[0057] Micro pKa predictions generated by the system 100 can be used in any of a variety of possible downstream applications. A few example use cases for pKa predictions generated by the system 100 are described next. The pKa predictions generated by the system 100 can be used for drug discovery. Drug discovery can involve identifying specific molecules within the body that are involved in a human or animal disease process. These molecules are often proteins, such as enzymes, receptors, or signaling proteins, that play a key role in the disease's development or progression. A ligand, often a small molecule (e.g., with a molecular weight equal to or less than 900 daltons), peptide, or antibody, can be selected to bind specifically to an identified target protein. When a drug that includes the ligand is administered to a patient, the ligand can bind to the target protein with high affinity and in doing so contribute to achieving a therapeutic effect in the patient. For instance, if the target molecule is an enzyme involved in a disease process, the ligand can inhibit its activity, thus disrupting the disease pathway. More generally, the interaction between the ligand and the target molecule can activate, inhibit, or alter the function of the target molcule to achieve a therapeutic effect. For example, the ligand can be an agonist or antagonist of a receptor of the target molecule (e.g., protein).
[0058] As one example, the micro pKa predictions generated by the system 100 can be used to screen candidate drug molecules for suitability as a drug, e.g., by using the micro pKapredictions to determine one or more pharmacokinetic or pharmacodynamic properties, such as such as ADMET (absorption, distribution, metabolism, excretion and toxicity properties, e.g., under in vitro or in vivo conditions), for each of the candidate drug molecules. Candidate drug molecules that have suitable pharmacokinetic or pharmacodynamic properties can then by physically synthesized for in vitro or in vivo testing.
[0059] Another use for the pKa predictions generated by the system 100 in drug discovery is to use is in determining a protonation state of a target molecule (e.g., a protein). For example, one or micro pKas for a target molecule can be used to determine a protonation state defining whether the one or more target atoms of the target molecule are protonated under biological conditions, e.g., in vitro or in vivo, or intracellular conditions. The 3D structure of the target molecule in the protonation state can then be determined and the interaction of the 3D structure with candidate ligands assessed. For example, a predicted structure of a complex comprising a candidate ligand and the target molecule in the protonation state can be determined for each of a plurality of candidate ligands. The interaction of each candidate ligand with the target protein molecule can then be assessed dependent on the predicted structure. The structures of the target molecule and / or complexes can, for example, be predicted using a structure prediction machine learning model, which can be implemented in any of a variety of possible ways. For instance, the structure prediction machine learning model can be based on the AlphaFold2 model, as described in Jumper, John, et al., “Highly accurate protein structure prediction with AlphaFold.” Nature 596.7873 (2021):583-589. As another example, the structure prediction machine learning model 120 can be based on the AlphaFold3 model, as described in Abramson, Josh, et al., “Accurate structure prediction of biomolecular interactions with AlphaFold 3.” Nature (2024): 1-3. As another example, the structure prediction machine learning model 120 can be based on the RoseTTAFold model, as described in Baek, Minkyung, et al., “Accurate prediction of protein structures and interactions using a three-track neural network.” Science 373.6557 (2021):871-876.
[0060] In some implementations a candidate (e.g. polypeptide or polynucleotide) drug molecule or ligand may include: an isolated antibody or aptamer, a fragment of an isolated antibody or aptamer, a single variable domain antibody, a bi- or multi-specific antibody, a multivalent antibody, a dual variable domain antibody, an immuno-conjugate, a fibronectin molecule, an adnectin, an DARPin, an avimer, an affibody, an anticalin, an affilin, a protein epitope mimetic or combinations thereof. A candidate (polypeptide) ligand may include an antibody with a mutated or chemically modified amino acid Fc region, e.g., which prevents or decreases ADCC (antibody-dependent cellular cytotoxicity) activity and / or increases half-lifewhen compared with a wild type Fc region. Candidate (polypeptide or polynucleotide) drug molecules or ligands may include antibodies with different CDRs (Complementarity- Determining Regions).
[0061] As another example, a downstream system can use the micro pKa predictions generated by the system 100 to perform molecular dynamics simulations, e.g., to find the most likely protonation site in a given solvent at a given pH.
[0062] As another example, the downstream system can use the micro pKa predictions to evaluate predictions generated by other models. For example, a downstream system can use the micro pKa predictions to analyze structural binding modes, e.g., to determine, given a predicted three-dimensional structure produced by a structure prediction model for a given molecule, whether certain sites within the binding modes are likely able to be protonated or not.
[0063] As another example, the downstream system can use the micro pKa predictions to guide the generation of a molecule by a generative model. For example, the downstream system can use the system to predict the micro pKas of various candidate molecules proposed by the generative model and then filter out candidates that do not have acceptable micro pKas. An example of a protein design generative model in Zambaldi et al., “De novo design of high-affinity protein binders with AlphaProteo,” arXiv: 2409.08022.
[0064] Generally, the system 100 can generate the micro pKa prediction 114 for the target atom 104 of the target molecule 102 using a graph neural network 110.
[0065] In particular, the system can generate a graph representation 116 of the target molecule 102. The graph representation 116 represents the target molecule 102 as a graph of nodes connected by edges, where each node represents one of the atoms in the target molecule 102 and each edge represents a bond between two atoms within the target molecule 102.
[0066] In some implementations, all of the atoms in the molecule are represented as respective nodes in the graph. In some other implementations, only the “heavy” atoms in the molecule are represented as nodes in the graph, e.g., so that the hydrogen atoms are not explicitly represented as nodes in the graph.
[0067] More specifically, the graph representation 116 includes a respective node embedding for each of the nodes in the graph and a respective edge embedding for each of the edges in the graph.
[0068] Node and edge embeddings will be described in more detail below with reference to FIG. 2.
[0069] The system 100 then processes the graph representation 116 using the graph neural network 110 to generate a graph network output 118 and generates the micro pKa prediction 114 from at least the graph network output 118.
[0070] Using the graph neural network 110 to generate the micro pKa prediction 114 will be described in more detail below with reference to FIGS. 2 and 3.
[0071] FIG. 2 is a flow diagram of an example process 200 for generating a prediction of the macro pKa of a target atom of a target molecule. For convenience, the process 200 will be described as being performed by a system of one or more computers located in one or more locations. For example, a prediction system, e.g., the prediction system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 200.
[0072] The system receives data specifying a target molecule and a target atom of the target molecule for which the micro pKa is to be predicted (step 202). That is, the target atom is designated as a protonation site on the target molecule for the purposes of predicting the micro pKa.
[0073] The system generates a first graph representation of a first state of the target molecule (step 204).
[0074] As described above, the first graph representation represents the target molecule as a graph that includes a plurality of nodes and edges, where each node represents a respective atom of the target molecule and each edge connects two nodes in the graph and represents a bond between the two nodes in the graph that are connected by the edge.
[0075] More specifically, the graph representation includes respective node features for each node in the graph and edge features for each edge in the graph.
[0076] Generally, the node embedding for a given node can be a combination, e.g., concatenation, of multiple different feature embeddings that each represent a feature of the corresponding atom.
[0077] In particular, the features of the atom include a feature that identifies whether the atom is the target atom or not, i.e., whether the atom is the protonation site for the computation of the micro pKa. For example, the feature embedding for this feature can be a binary feature that is equal to “1” if the atom is the target atom and equal to “0” if the atom is not the target atom (or vice versa).
[0078] The features can also include any of a variety of additional features.
[0079] For example, the additional features can also include an atom type feature that identifies the type of the atom, e.g., from a set of possible types of atoms that can be presentin molecules processed by the system. For example, the feature embedding for this feature can be a “one-hot” encoding that assigns a “1” to the type of the atom and a “0” to all other atom types in the set of possible atom types.
[0080] Thus, the node embedding for any given node in the graph identifies whether the node represents the target atom and also includes other information characterizing the atom represented by the node.
[0081] Generally, the edge embedding for a given edge can be a combination, e.g., concatenation, of multiple different feature embeddings that each represent a feature of the corresponding bond.
[0082] The bond features can generally include any of a variety of features that characterize the bond.
[0083] For example, the bond features can include a bond type feature that identifies the type of the bond, e.g., from a set of possible types of bonds that can be present between two atoms in molecules processed by the system. For example, the feature embedding for this feature can be a “one-hot” encoding that assigns a “1” to the type of the bond and a “0” to all other atom types in the set of possible atom types.
[0084] As yet another example, the bond features can include a feature that identifies whether the bond is a rotatable bond. For example, the feature embedding for this feature can be a binary feature that is equal to “1” if the bond is a rotatable bond and equal to “0” if the bond is not a rotatable bond (or vice versa).
[0085] As another example, the bond features can include a feature that identifies whether the bond is conjugated. For example, the feature embedding for this feature can be a binary feature that is equal to “1” if the bond is a conjugated bond (e.g., a bond that is part of conjugated system in the target molecule, e.g., a bond in a sequence of alternating double and single bonds) and equal to “0” if the bond is not a conjugated bond (or vice versa).
[0086] As another example, the bond features can include a feature that identifies the stereo type of the bond, e.g., the stereochemistry of the bond, more specifically the spatial arrangement of atoms or groups around the bond, e.g., E / Z isomerism, chirality (R / S) configuration, axial / equatorial, and so forth.
[0087] As another example, the bond features can include a feature that identifies whether the bond is rotatable or not, e.g., because it forms part of a conjugated system or a cyclic system, or where there is steric hindrance preventing rotation of the bond.
[0088] As yet another example, the bond features can include a feature that identifies whether the bond is in a ring, e.g., a homocycle, a heterocycle, an aromatic ring, a furanose ring, a macrocycle and so on.
[0089] In some implementations, the graph representation is a “two-dimensional” graph representation that does not consider the three-dimensional structure of the molecule.
[0090] In some other implementations, however, the graph representation is a “three- dimensional” graph representation that does consider the three-dimensional (“3D”) structure of the molecule.
[0091] In particular, in these implementations, the system obtains data specifying a respective 3D position of each of the atoms that are represented by nodes in the graph. For example, the system can receive the 3D positions of the atoms from an external system or can generate the 3D positions of the atoms by processing data specifying the target molecule using an appropriate structure prediction machine learning model.
[0092] The system can then include, as one of the features for each of the bonds, a feature that specifies the 3D distance between (i) the 3D position of one of the atoms connected by the bond and (ii) the 3D position of the other atom connected by the bond. For example, the feature embedding of this feature can be a regressed distance value or can be a “one-hot” encoding that assigns the distance value to one of multiple distance ranges, i.e., assigns a “1” to the range to which the distance value is assigned and a “0” to all other ranges.
[0093] As indicated above, the feature representation characterizes a “first” state of the target molecule.
[0094] For example, the first state can be a protonated state in which a hydrogen atom has been gained by the target molecule at the protonation site, i.e., a hydrogen atom has been bonded to the target atom (and the charge of the target molecule decreased by one electron).
[0095] When the graph includes a respective node for each of the atoms in the molecule, the system can generate a representation of the protonated state by adding a node to the graph representing the gained a hydrogen atom and an edge to the graph representing the bond between the target atom and the gained hydrogen atom. The system can then add corresponding node and edge features to the graph representation.
[0096] When the graph includes nodes only for each of the heavy atoms in the molecule, the addition of another hydrogen atom does not require modifying the nodes in the graph. In these cases, the system can reflect the gain of the hydrogen atom by modifying the node features, edge features, or both to account for the additional hydrogen atom. For example, the addition of the hydrogen atom can affect whether some of the bonds within the moleculeare conjugated. The system can then update the feature embeddings for the corresponding edges to indicate the change in conjugation of the corresponding bonds.
[0097] As another example, the first state can be an unprotonated state of the target molecule, i.e., a state in which the target molecule has not gained a hydrogen atom at the target atom. In this example, the system does not modify the edge features or the node features in the graph representation.
[0098] The system processes the graph representation of the first state of the target molecule using a graph neural network to generate a first graph network output (step 206). Generally, the first graph network output is a feature vector that characterizes features of the graph neural network that are relevant to predicting micro pKa.
[0099] In particular, the graph neural network processes the graph representation to generate a respective first updated node embedding for each node in the graph and, optionally, a respective first updated edge embedding for each edge in the graph. For example, the graph neural network can optionally first generate an initial updated node embedding for each node in the graph by processing the node embeddings using a node encoder neural network and generate an initial updated edge embedding for each edge in the graph by processing the edge embeddings using an edge encoder neural network. The graph neural network can process the initial updated node and edge embeddings using a sequence of graph neural network layers, each of which performs a variant of “message passing” to update at least the node embeddings that are received as input by the graph neural network layer. Examples of graph neural networks can be found in Battaglia et al., “Relational inductive biases, deep learning, and graph networks”, arxiv: 1806.01261.
[0100] The system can then generate the first graph network output from at least the first updated embedding for the node that represents the target atom. For example, the system can generate the first graph network output by applying pooling, e.g., global average pooling, on the embeddings for all of the nodes in the graph. As another example, the system can use the first updated embedding for the nod that represents the target atom as the first graph network output.
[0101] The graph neural network can generally have any appropriate graph neural network architecture with any number of graph neural network layers. For example, the graph neural network can be a message passing neural network (MPNN), a graph convolutional neural network (GCNN), or a graph simulator network (GSN). In some implementations, however, the graph neural network is a “deep” graph neural network with more than ten graph neural network layers, e.g., with sixteen, fifty, or one hundred graph neural network layers.
[0102] In some implementations, the graph neural network architecture is optimized to improve performance of the graph neural network on the prediction task. For example, the graph neural network can include tailored activation functions (TATs). Graph neural networks with TATs are described in more detail in Zaidi, et al, PRE-TRAINING VIA DENOISING FOR MOLECULAR PROPERTY PREDICTION, arXiv:2206.00133.
[0103] Optionally, the system also generates a graph representation of a second state of the target molecule (step 208) and processes this graph representation using the graph neural network to generate a second graph network output (step 210). For example, if the first state is the unprotonated state of the target molecule, the second state can be the protonated state of the target molecule. If the first state is the protonated state of the target molecule, the second state can be the unprotonated state of the target molecule. That is the second state of the target molecule may be deprotonated relative to the first state of the target molecule (or vice versa). In some examples, protonation or deprotonation of the target molecule can be accompanied by rearrangement of the atoms or chemical bonds in the target molecule. In some examples, the first and second states of the target molecule correspond to different tautomeric states of the target molecule (i.e., the proton transfer may be intramolecular). As described above, the second graph representation represents the target molecule as a graph that includes a plurality of nodes and edges, where each node represents a respective atom of the target molecule and each edge connects two nodes in the graph and represents a bond between the two nodes in the graph that are connected by the edge. Thus, where the second state corresponds to a deprotonated state of the target molecule, there may be one fewer nodes and one fewer edges, corresponding to a loss of a hydrogen atom.
[0104] Further optionally, the system generates latent features of the target atom (step 212). Generally, the system generates the latent features based on which functional groups the target atom is included in.
[0105] One example technique for generating the latent features is described in more detail below with reference to FIG. 3.
[0106] The system then generates the micro pKa for the target atom from at least the first graph network output (step 214).
[0107] In particular, when the system does not generate the second graph network output or the latent features, the system can generate the micro pKa from the first graph network output.
[0108] When the system generates the second graph network output but not the latent features, the system can generate the micro pKa from the first and second graph network outputs.
[0109] When the system generates the latent features but not the second graph network output, the system can generate the micro pKa from the first graph network output and the latent features.
[0110] When the system generates both the latent features and the second graph network output, the system can generate the micro pKa from the first and second graph network outputs and the latent features.[OHl] An example technique for generating the micro pKa is described in more detail below.
[0112] FIG. 3 is a block diagram that shows an example 300 of the operation of the prediction system.
[0113] In the example 300, given a target atom 104 of a target molecule 102, the system 100 generates a first graph network output 310, a second graph network output 320, and latent features 340, and then generates the predicted micro pKa 114 from the first graph network output 310, the second graph network output 320, and the latent features 340.
[0114] In particular, the system 100 processes a graph representation 302 of the protonated state of the target molecule 102 using the graph neural network 110 to generate the first graph network output 310 and a graph representation 304 of the unprotonated state of the target molecule 102 to generate the second graph network output 320.
[0115] To generate the latent features 340, the system 100 maintains data identifying a predetermined set of functional groups. A “functional group” is a specific grouping of atoms within a molecule. For example, the function groups can be a set of a variety of acidic and basic functional groups that appear in molecules processed by the system. Examples of such functional groups include sulphonic acid, sulphide, phenol, and so on.
[0116] The system 100 generates an encoding vector 306 that identifies, for each functional group in the predetermined set, whether the target atom 104 is included in the functional group. For example the encoding vector 306 can be a binary encoding vector (also referred to as a “bit” vector) that includes a respective entry for each functional group in the set, with the entry for a given functional group being equal to 1 if the target atom 104 is included in the functional group and equal to 0 if the target atom 104 is not included in the functional group.
[0117] The system 100 then processes the encoding vector 306 using a machine learning model 330 to generate the latent features 340. In particular, the latent features 340 can be an intermediate output of the machine learning model. For example, the machine learningmodel can have been trained to predict the micro pKa of the target atom from the encoding vector 306. As a particular example, the model 330 can be a linear model that applies one or more linear layers as part of predicting the micro pKa. The latent features 340 can then be the outputs of one of the linear layers. As another particular example, the model 330 can be a MLP. The latent features 340 can then be the outputs of one the intermediate layers of the MLP.
[0118] The system then combines the first graph network output 320, the second graph network output 320, and the latent features 340 to generate combined features 350 and processes the combined feature 350 using a prediction neural network head 360 to generate the predicted pKa 114.
[0119] A neural network “head,” as used in this specification, is a collection of one or more neural network layers. For example, the prediction neural network head 360 can be a multilayer perceptron (MLP) that processes the combined features 350 to generate the predicted pKa 114.
[0120] The system can combine the first graph network output 320, the second graph network output 320, and the latent features 340 in any of a variety of ways.
[0121] For example, the system can generate the combined features 350 by concatenating the first graph network output 320, the second graph network output 320, and the latent features 340.
[0122] As another example, the system can first average the first graph network output 320 and the second graph network output 320 to generate an averaged output and then concatenate the averaged output and the latent features 340.
[0123] When the system 100 does not generate the latent features 340, the system canc use the combination of the first graph network output 320 and the second graph network output 320 as the combined features 350.
[0124] When the system 100 does not generate the second graph network output 320, the system can use the combination of the first graph network output 320 and the latent features 340 as the combined features 350.
[0125] When the system 100 does not generate the second graph network output 320 or the latent features 340, the system can use the first graph network output 320 as the combined features 350.
[0126] Prior to using the graph neural network and the prediction neural network head for predicting micro pKas, the system 100 or another training system trains the graph neuralnetwork and the prediction neural network head on training data to determine trained values of the network parameters of the neural networks.
[0127] One example technique for training the graph neural network and the prediction neural network head is described below.
[0128] FIG. 4 is a flow diagram of an example process 400 for training the graph neural network. For convenience, the process 400 will be described as being performed by a system of one or more computers located in one or more locations. For example, a prediction system, e.g., the prediction system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 400.
[0129] The system receives synthetic training examples (step 402). Each training example incudes data identifying a target molecule and, for each of one or more target atoms within the target molecule, a respective target micro pKa.
[0130] The training examples are referred to as “synthetic” because micro pKas of individual atoms of a given molecule are generally unable to be experimentally measured. That is, only one or more macro pKas of the given molecule are able to be determined experimentally. That is, certain molecules may have multiple macro pKas.
[0131] To account for this, the system generates (or obtains from another system) “synthetic” training examples in which the target micro pKas are computed, e.g., using density functional theory (DFT) or another computational approach. For example, a target micro pKa can be determined from a free energy change (e.g., Gibb’s free energy change) and associated equilibrium constant for protonation of the corresponding target atom of the target molecule. The free energy change can be determined using path integral molecular dynamics techniques, for example.
[0132] The system trains the graph neural network (and the prediction head) on the synthetic training examples (step 404). For example, the system can perform this training to minimize a loss function, e.g., a mean-squared error, that measures differences between predicted micro pKas generated using the graph neural network and corresponding target micro pKas in the training examples.
[0133] The system obtains experimental training examples (step 406) and trains the neural network on the experimental training examples (step 408). Generally, each experimental training example can include data specifying a training molecule and an experimentally measured macro pKa for the training molecule. Experimental training examples are available from publicly available databases, such as ChEMBL(https: / / www.ebi. ac.uk / chembl / ), which is a manually curated database of bioactive molecules with drug-like properties.
[0134] In some implementations, the system can leverage the fact that monoprotic molecules have micro and macro pKas that are the same. Thus, the experimental training examples can be monoprotic molecules and the system can use experimentally measured macro pKas for the molecules as the target micro pKas for this training. Experimentally measured macro pKas can be obtained by many quantitative chemical analysis techniques, such as titration or volumetric analysis.
[0135] In some other implementations, for at least some of the experiment training examples, the system can regress a macro pKa for a given training molecule, e.g., by enumerating all protonation states and predicting all corresponding micro pKas. The system can then compute a loss between the regressed macro pKa and an experimentally measured macro pKa for the given training molecule. That is, the system can predict respective micro pKas for each of a plurality of target atoms in the training molecule and combine the predicted micro pKas to obtain a predicted macro pKa that is compared with e.g., experimentally measured macro pKas, when training the neural network. For example, the macro pKa can be determined using equilibrium constants corresponding to the predicted micro pKas to calculate an overall equilibrium constant (and hence macro pKa) for protonation / deprotonation of the target molecule.
[0136] Optionally, when performing step 404, when performing step 408, or both, the system can incorporate one or more auxiliary losses.
[0137] One example of an auxiliary loss is a loss that requires the system to predict, from the first network output, the second network output, or both, a total charge of the target molecule when the molecule is in the corresponding state. This auxiliary loss can be applied when performing step 404, step 408, or both.
[0138] Another example of an auxiliary loss is a loss that requires the system to predict, from the first network output, the second network output, or the combined features, a free energy of the molecule when the molecule is in the corresponding state. This auxiliary loss can be applied when performing step 404, i.e., when the free energy can readily be calculated as part of the micro pKa calculation.
[0139] Another example of an auxiliary loss is a denoising loss that requires the system to “denoise” a noisy version of the target molecule. An example of such a denoising loss is described in more detail in Godwin, et al, SIMPLE GNN REGULARISATION FOR 3D MOLECULAR PROPERTY PREDICTION & BEYOND, arXiv:2106.07971.
[0140] Optionally, instead of or in addition to using the denoising loss as an auxiliary loss, the system can pre-train the graph neural network on a denoising objective prior toperforming step 404. This type of pre-training is described in more detail in Zaidi, et al, PRE-TRAINING VIA DENOISING FOR MOLECULAR PROPERTY PREDICTION, arXiv:2206.00133.
[0141] Training a neural network (e.g., the graph neural network or the prediction neural network head) to optimize an objective function, e.g., to minimize a loss function, can include determining gradients of the objective function with respect to the parameters of the neural network, and then using the gradients to update the parameter values of the neural network. The system can determine the gradients of the objective function with respect to the parameters of the neural network, e.g., using backpropagation. The system can use the gradients to update the parameter values of the neural network using the update rule of an appropriate gradient descent optimization algorithm, e.g., RMSprop or Adam.
[0142] This specification uses the term “configured” in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions.
[0143] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine- readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially-generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.
[0144] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0145] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub-programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.
[0146] In this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.
[0147] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.
[0148] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit.Generally, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.
[0149] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0150] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.
[0151] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and compute-intensive parts of machine learning training or production, i.e., inference, workloads.
[0152] Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework, or a Jax framework.
[0153] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.
[0154] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.
[0155] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0156] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations beperformed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0157] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
Claims
CLAIMS1. A method performed by one or more computers, the method comprising: obtaining data specifying a target molecule and a target atom of the target molecule; generating a first graph representation of the target molecule that represents a first state of the target molecule as a graph of the target molecule, wherein the graph comprises a plurality of nodes and a plurality of edges, wherein each node represents a respective atom in the target molecule and each edge connects a respective pair of nodes in the graph and represents a bond between a pair of atoms that are represented by the respective pair of nodes; processing the first graph representation using a graph neural network to generate a first graph network output; generating a second graph representation of the target molecule that represents a second state of the target molecule as a graph of the target molecule; processing the second graph representation using the graph neural network to generate a second graph network output; and generating, from at least the first graph network output and the second graph network output, a predicted micro pKa of the target atom of the target molecule.
2. The method of claim 1, further comprising: obtaining functional group data that identifies, for each of a plurality of functional groups, whether the target atom is included in the function group; generating latent features of the target atom from the functional group data, wherein generating, from at least the first graph network output and the second graph network output, a predicted micro pKa of the target atom of the target molecule comprises: generating, from at least the first graph network output, the second graph network output, and the latent features, the predicted micro pKa of the target atom of the target molecule.
3. The method of claim 2, wherein generating latent features from the functional group data comprises: generating, from the functional group data, an encoding of the target atom; and processing the encoding of the target atom using a machine learning model, whereinthe latent features are an intermediate output of the machine learning model.
4. The method of claim 3, wherein the machine learning model has been trained to predict the micro pKa of the target atom from the encoding.
5. The method of any preceding claim, wherein the first state of the target molecule is a protonated state of the target molecule.
6. The method of any preceding claim, wherein generating, from at least the first graph network output and the second graph network output, the predicted micro pKa of the target atom of the target molecule comprises: generating combined features from at least the first graph network output and the second graph network output; and processing the combined features using a prediction neural network head to generate the predicted micro pKa of the target atom of the target molecule.
7. The method of any preceding claim, wherein the second state of the target molecule is an unprotonated state of the target molecule.
8. The method of any preceding claim, wherein the first graph representation comprises a respective node embedding of each node in the first graph and a respective edge embedding of each edge in the first graph.
9. The method of claim 8, wherein the respective node embedding for each node in the first graph identifies whether or not the atom represented by the node is the target atom.
10. The method of claim 8 or 9, wherein the respective node embedding for each node in the first graph identifies an atom type of the atom represented by the node.
11. The method of any one of claims 8-10, wherein the respective edge embedding for each edge in the first graph identifies a bond type of the bond represented by the edge.
12. The method of any one of claims 8-11, wherein the respective edge embedding for each edge in the first graph identifies whether the bond represented by the edge is aconjugated bond.
13. The method of any one of claims 8-12, wherein the respective edge embedding for each edge in the first graph identifies whether the bond represented by the edge is a rotatable bond.
14. A method of obtaining a drug, the method comprising: for each of a plurality of candidate drug molecules, using the method of any one of the preceding claims to predict a respective one or more micro pKas for the candidate drug molecule; and using the predicted micro pKas to select at least one of the candidate drug molecules as the drug.
15. The method of claim 14, wherein using the predicted micro pKas to select at least one of the candidate drug molecules as the drug comprises: for each of the candidate drug molecules, using the one or more predicted micro pKas to determine one or more pharmacokinetic or pharmacodynamic properties for the candidate drug molecule; and using the predicted pharmacokinetic or pharmacodynamic properties of the candidate drug molecules to select the one or more candidate drug molecules.
16. A method of obtaining a ligand, wherein the ligand is a drug or a ligand of an industrial enzyme, the method comprising: using the method of any one of claims 1-13 to predict one or more micro pKas for protonation of a target molecule at a corresponding one or more target atoms; using the one or micro pKas to determine a protonation state defining whether the one or more target atoms of the target molecule are protonated under biological conditions; for each of one or more candidate ligands:(a) determining a predicted structure of a complex comprising the candidate ligand and the target molecule in the protonation state; and(b) evaluating an interaction of the candidate ligand with the target protein molecule dependent on the predicted structure; and selecting one or more of the candidate ligands as the ligand dependent on a result of the evaluating.
17. The method of claim 16, wherein the target molecule comprises a receptor or enzyme, and wherein the ligand is an agonist or antagonist of the receptor or enzyme, or wherein the ligand comprises an antibody or aptamer and the target molecule comprises an antibody or aptamer target, in particular a virus or cancer cell protein, and wherein the antibody or aptamer binds to the antibody or aptamer target to provide a therapeutic effect.
18. The method of any one of claims 14-17, further comprising synthesizing the drug or ligand.
19. A method of claim 18, further comprising testing for biological activity of the ligand or drug in vitro or in vivo.
20. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective method of any one of claims 1-17.
21. A system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of any one of claims 1-17.
Citation Information
Cited By
Molecular property prediction using molecule protonation states
WO2026131690A1