Synthetic data generation using matched molecular pairs
The system generates synthetic training examples using matched molecular pairs to address data scarcity in machine learning models, enhancing predictive performance for molecular property predictions and reactivity classification.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- ISOMORPHIC LABS LTD
- Filing Date
- 2025-10-28
- Publication Date
- 2026-05-07
AI Technical Summary
Training machine learning models for molecular property predictions and reactivity classification tasks is challenging due to the scarcity of experimental training examples, leading to insufficient data and potential overfitting or poor predictive performance.
A system that generates synthetic training examples using matched molecular pairs to define transformation functions, allowing for the generation of large numbers of synthetic data without requiring generative machine learning models, thereby overcoming data scarcity and computational inefficiencies.
The system efficiently produces synthetic training examples that effectively train machine learning models for molecular property predictions and reactivity classification, reducing computational resources and improving predictive performance.
Smart Images

Figure EP2025081180_07052026_PF_FP_ABST
Abstract
Description
Isomorphic Labs Limited F&R Ref.: 53672-0030W01 PCT ApplicationSYNTHETIC DATA GENERATION USING MATCHED MOLECULAR PAIRSBACKGROUND
[0001] This specification relates to processing data using machine learning models.
[0002] Machine learning models receive an input and generate an output, e.g., a predicted output, based on the received input. Some machine learning models are parametric models and generate the output based on the received input and on values of the parameters of the model.
[0003] Some machine learning models are deep models that employ multiple layers of models to generate an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non-linear transformation to a received input to generate an output.SUMMARY
[0004] This specification describes a system implemented as computer programs on one or more computers in one or more locations that can generate a dataset of synthetic training examples including a chemical structure and one or more molecular properties of the chemical structure, using matched molecular pairs. In this specification, a synthetic training example refers to a training example that is generated computationally rather than experimentally, e.g., an example that is predicted using computational techniques.
[0005] In this specification, a matched molecular pair refers to a pair of molecules that differ by a transformation, e.g., a structural modification. For example, the structural modification can involve a change to an atom, functional group, or substituent, e.g., an atom or group of atoms attached to the core structure of a molecule. As another example, the structural modification can involve a change to one or more atoms of the core structure of a molecule. Matched molecular pairs can be useful for determining the impact of a structural change to one or more molecular properties. As an example, the one or more molecular properties can be ADMET properties, e.g., absorption, distribution, metabolism, excretion, and toxicity characteristics of a chemical structure within the context of drug development.
[0006] In particular, the system can define a transformation function by identifying one or more matched molecular pairs that differ by a particular transformation, e.g., structural modification, and can determine the impact of the transformation function by aggregating the experimentally observed changes in the molecular property values between the respective molecules in each of the identified one or more matched molecular pairs as the predicted change in the molecular property values that results from the transformation.
[0007] The system can then identify candidate molecules that are candidates for the application of the particular transformation function based on the inclusion of a target chemical structure associated with the transformation function, and can apply the transformation specified by the transformation function to the candidate molecules. In particular, the system can modify each candidate molecule according to the transformation function to generate a synthetic training input molecule and can apply the determined impact in the molecular property for the transformation function to yield a predicted molecular property value for the synthetic training example.
[0008] More specifically, the system can identify a number of transformation functions using experimentally observed matched molecular pairs and can generate synthetic training examples using the transformation functions. The system can assemble a synthetic training dataset for training a machine learning model using the synthetic training examples, e.g., by a machine learning technique. As an example, the system can use the synthetic training examples to train a machine learning model to generate predicted molecular property values for an input molecule.
[0009] As another example, the system can identify transformations between reactants in a collection of reactions and generate synthetic training reactions using the transformations. In particular, the system can identify pairs of reactions in which one or more reactants are modified in the same way, e.g., in which at least one reactant in the first reaction and one reactant in the second reaction are matched molecular pairs, and can define a transformation function based on the transformation to the reactants. In this case, the system can train a machine learning model to classify whether an input set of molecules is reactive, e.g., whether an input set of molecules can react to form a reaction product, e.g. comprising one or more product molecules).
[0010] According to a first aspect there is provided a method for obtaining data that identifies a collection of molecules and, for each molecule in the collection of molecules, a molecular property value for the molecule, processing the data identifying the collection of molecules to identify a plurality of matched molecular pairs, processing the plurality of matched molecular pairs to generate data defining a set of transformation functions, wherein each transformation function is defined by at least: (i) a first molecular fragment, (ii) a second molecular fragment, (iii) an inclusion criterion defining a class of molecules, and (iv) a predicted change in the molecular property value resulting from replacing the first molecular fragment by the second molecular fragment in a molecule that satisfies the inclusion criterion, generating a plurality of synthetic training examples for training a machine learning model using the set oftransformation functions, wherein the machine learning model is configured to process data characterizing an input molecule to generate a predicted molecular property value for the input molecule, and training the machine learning model on the plurality of synthetic training examples by a machine learning training technique.
[0011] In some implementations, each matched molecular pair includes a pair of molecules comprising: (i) a first molecule, and (ii) a second molecule that differs from the first molecule in that one molecular fragment in the first molecule is replaced by another molecular fragment in the second molecule.
[0012] In some implementations, for each transformation function, a candidate molecule satisfies the inclusion criterion for the transformation function if: (i) the candidate molecule includes the first molecular fragment associated with the transformation function, and (ii) a chemical structure of a portion of the candidate molecule that includes the first molecular fragment and all atoms in the candidate molecule that are separated from the first molecular fragment by at most a threshold number of bonds in the candidate molecule matches a target chemical structure associated with the transformation function.
[0013] In some implementations, the threshold number of bonds is at least two bonds.
[0014] In some implementations, for each transformation function, a candidate molecule satisfies the inclusion criterion for the transformation function if: (i) the candidate molecule includes the first molecular fragment associated with the transformation function, and (ii) a molecular fingerprint of the candidate molecule excluding the first molecular fragment matches a target molecular fingerprint associated with the transformation function.
[0015] In some implementations, processing the plurality of matched molecular pairs to generate data defining the set of transformation functions comprises, for each transformation function, identifying a subset of the plurality of matched molecular pairs for which: (i) the first molecule in each matched molecular pair includes a same first fragment, (ii) the second molecule in each matched molecular pair differs from the corresponding first molecule in that the same first fragment is replaced by a same second fragment, (iii) the first molecule in each matched molecular pair satisfies a same inclusion criterion, and (iv) the subset of the plurality of matched molecular pairs includes at least two matched molecular pairs, and generating the transformation function based on the subset of the plurality of matched molecular pairs.
[0016] In some implementations, processing the plurality of matched molecular pairs to generate data defining the set of transformation functions comprises, for each transformation function, determining, for each matched molecular pair in the subset of the plurality of matched molecular pairs that are used to generate the transformation function, a delta between themolecular property values of: (i) the first molecule, and (ii) the second molecule, in the matched molecular pair, and determining the predicted change in the molecular property value for the transformation function based on a measure of central tendency of the deltas.
[0017] In some implementations, the measure of central tendency is a mean.
[0018] In some implementations, processing the plurality of matched molecular pairs to generate data defining the set of transformation functions further comprises, for each transformation function determining that a measure of dispersion of the deltas does not exceed a maximum threshold.
[0019] In some implementations, the measure of dispersion is a standard deviation.
[0020] In some implementations, generating the plurality of synthetic training examples for training the machine learning model using the set of transformation functions comprises, for each of a plurality of original molecules in the collection of molecules, determining, for each transformation function, whether the original molecule satisfies the inclusion criterion of the transformation function, and in response to determining that the original molecule satisfies the inclusion criterion of a transformation function: generating a new molecule by replacing the first molecule fragment with the second molecular fragment in the original molecule, generating a molecular property value for the new molecule as a sum of: (i) the molecular property value for the original molecule, and (ii) the predicted change in the molecular property value that is specified by the transformation function, and generating a synthetic training example that includes: (i) a training input to the machine learning model that characterizes the new molecule, and (ii) a target output of the machine learning model that specifies the molecular property value of the new molecule.
[0021] In some implementations, each synthetic training example includes: (i) a training input to the machine learning model that characterizes a molecule, and (ii) a target output of the machine learning model that specifies a target molecular property value of the molecule.
[0022] In some implementations, training the machine learning model using the plurality of synthetic training examples by the machine learning training technique comprises, for each synthetic training example: training the machine learning model to reduce a discrepancy between: (i) a predicted molecular property value generated by processing the training input of the synthetic training example using the machine learning model, and (ii) the target molecular property value specified by the synthetic training example.
[0023] In some implementations, the machine learning model comprises a neural network.
[0024] According to another aspect there is provided a method for predicting a molecular property value, comprising obtaining data characterizing an input molecule, and processing thedata characterizing the input molecule using a machine learning model that has been trained by the method of any of the aforementioned implementations, to generate a predicted molecular property value for the input molecule.
[0025] According to another aspect there is provided a method for obtaining chemical reaction data that identifies a collection of chemical reactions, processing the chemical reaction data to generate data defining a set of transformation functions, wherein each transformation function is defined by at least: (i) a first molecular fragment, and (ii) a second molecular fragment, and wherein generating each transformation function comprises determining that, for at least a threshold number of pairs of chemical reactions in the collection of chemical reactions: (a) at least one reactant in a first chemical reaction of the pair of chemical reactions includes the first molecular fragment, and (b) a second set of reactants of a second chemical reaction of the pair of chemical reactions differs from a first set of reactants of the first chemical reaction of the pair of chemical reactions in that each instance of the first molecular fragment in the first set of reactants is replaced by the second molecular fragment in the second set of reactants, generating a plurality of synthetic training examples for training a machine learning model using the set of transformation functions, wherein the machine learning model is configured to process data characterizing an input set of molecules to classify whether the input set of molecules is reactive, and training the machine learning model using the plurality of synthetic training examples by a machine learning training technique.
[0026] In some implementations, each chemical reaction in the collection of chemical reactions further comprises a label defining a reaction type of the reaction, and wherein processing the chemical reaction data to generate data defining the set of transformation functions further comprises, for each reaction type, selecting a subset of the chemical reaction data corresponding with chemical reactions that have the label defining the reaction type, determining one or more transformation functions associated with the reaction type using the subset of the chemical reaction data.
[0027] In some implementations, the threshold number of pairs of chemical reactions is at least three.
[0028] In some implementations, each synthetic training example comprises: (i) a training input to the machine learning model that characterizes a set of molecules, and (ii) a target output of the machine learning model that identifies the set of molecules as being reactive.
[0029] In some implementations, generating the plurality of synthetic training examples for training the machine learning model using the set of transformation functions comprises, for each of a plurality of original chemical reactions in the collection of chemical reactions,determining, for each transformation function, whether the original chemical reaction has at least one reactant that includes the first molecular fragment of the transformation function, and, in response to determining that the original chemical reaction has at least one reactant that includes the first molecular fragment of the transformation function, generating a new chemical reaction by replacing each instance of the first molecular fragment in a set of reactants and in a set of products of the original chemical reaction with the second molecular fragment of the transformation function, and generating a synthetic training example that includes: (i) a training input to the machine learning model that characterizes at least the set of reactants of the new chemical reaction, and (ii) a target output of the machine learning model that identifies the set of reactants as being reactive.
[0030] In some implementations, the training input to the machine learning model characterizes the set of reactants and the set of products of the new chemical reaction.
[0031] In some implementations, the method further comprises identifying the reaction type associated with each transformation function, determining whether the original chemical reaction has the label defining the reaction type, and in response to determining that the original chemical reaction has the label defining the reaction type, generating the new chemical reaction.
[0032] In some implementations, training the machine learning model using the plurality of synthetic training examples by the machine learning training technique comprises, for each synthetic training example, training the machine learning model to reduce a discrepancy between: (i) a classification generated by the machine learning model by processing the training input of the training example, and (ii) the target output of the machine learning model that identifies the set of molecules characterized by the training input as being reactive.
[0033] In some implementations, the machine learning model comprises a neural network.
[0034] According to another aspect there is provided a method of determining whether an input set of molecules is reactive, comprising obtaining data characterizing an input set of molecules, and processing the data characterizing the input set of molecules using a machine learning model that has been trained by the method of any of the aforementioned methods, to generate a classification output that classifies whether the input set of molecules is reactive.
[0035] In another aspect, there is provided a system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform the method of any one of the example implementation methods described.
[0036] In another aspect, there is provided a computer storage medium encoded with a computer program, the program comprising instructions that are operable, when executed by data processing apparatus, to cause the data processing apparatus to perform the method of any one of the example implementation methods described.
[0037] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages.
[0038] Biological and chemical applications of machine learning that rely on molecular data often suffer from data scarcity since gathering real-world biological and chemical data requires expensive and time-consuming experiments. Additionally, experiments usually yield relatively small datasets, presenting a data bottleneck problem.
[0039] Training a machine learning model to perform a molecular property prediction task requires large numbers of training examples for use in training the machine learning model. In particular, effectively predicting molecule properties requires the machine learning model to implicitly encode, in its parameter values, complex biochemical principles and relationships that connect chemical structures to associated molecule properties. Training a machine learning model to learn the required biochemical principles and relationships directly from data (by a machine learning training technique) requires the training examples used for training the machine learning model to characterize molecules with large numbers of diverse chemical structures and associated molecule properties.
[0040] Training a machine learning model to perform a molecular property prediction task using an insufficient number of training examples may cause the training to fail to converge, or may cause the machine learning model to overfit the training examples (and thus fail to effectively generalize to generating predictions for previously unseen molecules), or may cause the trained machine learning model to exhibit poor predictive performance.
[0041] Thus, a technical problem that arises in the context of training machine learning models to perform molecular property predictions tasks is that large numbers of training examples are required for training but that experimental training examples are scarce and difficult to obtain. Without a sufficient number of training examples, training a machine learning model to perform a molecular property prediction task may be technically infeasible.
[0042] The system described in this specification provides a technical solution to the technical problem of training a machine learning model to perform a molecular property prediction task in the absence of a large number of experimental training examples. In particular, the system can generate synthetic training examples by leveraging matched molecular pairs as a template for identifying how a given transformation function would impact the molecule, e.g., matchedmolecular pairs can be used to determine relative changes in molecular property values since they are useful indicators for structure-activity relationship (SAR) analysis. More specifically, the system can identify groups of matched molecular pairs to define transformation functions with associated predicted changes in molecular property values that are determinative of the actual impact of the transformation function. The system can therefore use the transformation functions to generate synthetic data for candidate molecules of the same class, e.g., that include the same target chemical structure, associated with the transformation function.
[0043] The system implements a fast, computationally efficient approach for generating synthetic training examples for training a machine learning model to perform a molecular property prediction task. More specifically, the operations performed by the system to identify matched molecular pairs, aggregate groups of matched molecular pairs to define transformation functions, and then apply those transformation functions to generate synthetic training examples do not require auxiliary prediction models, or machine learning, or complex statistical analyses. Rather, the system implements fast, deterministic, and stable numerical operations based on comparing and matching large numbers of complementary chemical structures in order to generate synthetic training examples.
[0044] An alternative approach to generating synthetic training examples may involve training a generative machine learning model to generate chemical structures of new molecules for inclusion in synthetic training examples. However, such an approach is highly computationally intensive as it requires training a standalone generative machine learning model, which may be parametrized by millions or billions of parameters in order to capture and internally represent the underlying distribution of chemical structures and associated chemical properties in a vast space of possible chemical structures. In contrast, the system described in this specification can identify and leverage transformation functions to facilitate the generation of synthetic training examples, which does not require any generative machine learning model (or any other auxiliary predictive model), and therefore the system can generate large numbers of synthetic training examples while consuming relatively fewer computational resources, e.g., memory and computing power. Further the system avoids technical issues involved with training and using generative machine learning models for generating synthetic training examples such as numerical issues, instability, overfitting, and so forth.
[0045] Likewise, a similar technical problem can arise in the context of training machine learning models to perform reactivity classification tasks for a set of input molecules. This also requires a large number of training examples, where experimental training examples are scarce and difficult to obtain. In fact, the space of reactions is even more complex than the space ofallowable chemical structures for a molecule, e.g., since there is often more than one molecule involved as a reactant in a potential reaction, thereby rendering the generation of synthetic reaction examples computationally infeasible, especially without large generative machine learning models.
[0046] The system of this specification can provide for the generation of synthetic reaction data by identifying pairs of reactions in which one or more reactants are modified in the same way, and can define a transformation function based on the modification to the reactants. The system can then apply the transformation function to candidate reactions that include the reactant of the first reaction in the transformation function to generate a synthetic reaction. In this case, since the transformation functions were determined from known reactions, the system can identify transformations that will not impact the success of the reaction, and therefore can generate synthetic reactions in which the reactions are likely to be successful.
[0047] Similarly, in this case, the operations performed by the system to identify and apply transformation functions to generate synthetic training examples do not require auxiliary prediction models, or machine learning, or complex statistical analyses. Rather, the system implements fast, deterministic, stable numerical operations based on comparing and matching large numbers of complementary chemical structures across reactants in a collection of reactions in order to generate synthetic training examples.
[0048] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0049] FIG. 1A is a system diagram of an example matched molecular pair (MMP) data augmentation system being used to generate synthetic training examples for a collection of molecules.
[0050] FIG. IB is a system diagram of the example matched molecular pair (MMP) data augmentation system of FIG. 1 A being used to generate synthetic training examples for a collection of chemical reactions.
[0051] FIG. 2 provides an example of how an MMP data augmentation system can calculate a predicted change in a molecular property value for a transformation function.
[0052] FIG. 3 illustrates how an MMP data augmentation system can determine whether to apply a set of transformation functions to a set of molecules.
[0053] FIG. 4 is a flow diagram of an example process for generating synthetic training examples using MMPs and training a machine learning model using the synthetic training examples to generate a predicted property value for an input molecule.
[0054] FIG. 5 is a flow diagram of an example process for generating synthetic training examples using MMPs based on a collection of chemical reactions and training a machine learning model using the synthetic training examples to determine whether an input collection of molecules is reactive.
[0055] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION
[0056] FIG. 1A shows an example matched molecular pair (MMP) data augmentation and training system 100. The system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations in which the systems, components, and techniques described below are implemented.
[0057] The system 100 can be used to generate a dataset including computationally-generated molecules and predicted associated molecular property values as synthetic training examples. In particular, the system 100 can define respective transformation functions by identifying the transformation for a group of matched molecular pairs and can apply each transformation function on candidate molecules that satisfy an inclusion criterion defining a class of molecules for the transformation function to generate molecules with predicted molecular property values based on the transformation function.
[0058] The system 100 can obtain a collection of molecules 110 from a molecule database 105. The molecular database 105 can maintain one or more collections of molecules, e.g., libraries of molecules. In particular, the collections of molecules can include a number of data objects or data files each specifying data that characterizes a respective molecule, e.g., chemical structure data and associated molecular property data for a molecule.
[0059] For example, the molecule database 105 can include a collection of molecules 110 that are each specified by a Simplified Molecular Input Line Entry System (SMILES) string or by an International Chemical Identifier (InChi) string, e.g., which both provide a representation of the chemical structure of a molecule in a one-dimensional string. As another example, the molecule database 105 can include a collection of structured data files, e.g., respective protein databank files, a MolBlock file, or chemical markup files, etc.
[0060] The molecule database 105 can include ADMET properties, e.g., absorption, distribution, metabolism, excretion, and toxicity characteristics of a molecule within thecontext of drug development, as the associated molecular property data. For example, the molecular property values can include an intrinsic or in vivo clearance, e.g., how quickly the molecule can be metabolized theoretically or in an actual biological system, a solubility, a permeability, and so forth. As another example, the molecular property values can include a measure of duration of a drug in a biological system, e.g., a CYP enzyme inhibition or a P- glycoprotein efflux, or indicators of safe use, e.g., a no-observed-adverse-effect dose or a measure of potassium channel inhibition, e.g., hERG inhibition. As yet another example, the molecular property values can include a bioavailability, a volume of distribution, e.g., a measure of how widely a molecule distributes in tissue, and a measure of plasma protein binding. As a further example, the molecular property values can include a partition coefficient (LogP), e.g., a measure of the lipophilicity of the molecule, e.g., an indicator of a molecule’s ability to be absorbed in a lipid-containing biological system, and a distribution coefficient (LogD), e.g., a measure of the lipophilicity of the molecule at different acidities (pH). As yet another example, the molecular property values can include a value defining a more specific structure activity relationship, e.g. a measure of blood-brain barrier penetration, or characterizing a substrate or inhibitor of an enzyme. In general, the molecular property values can include any quantifiable physical, chemical, or biological, e.g. physicochemical or pharmacological, property of a molecule.
[0061] For example, the molecule database 105 can include molecule data from a publicly- accessible data source, e.g. ChEMBL, PubChem (NCBI, the National Center for Biotechnology Information), or DrugBank. As another example, the molecule database 105 can include molecule data from a private data source, e.g., an entity. In some cases, experiments may be performed to generate data for molecule database 105.
[0062] The system 100 can process the collection of molecules 110 using a matched molecular pair (MMP) identification engine 130. In particular, the system 100 can use the MMP identification engine 130 to identify subsets of the collection of molecules 110 as matched molecular pairs 135, e.g., pairs of molecules with the same chemical environment (e.g. scaffold, core, or framework) that differ by the same structural modification (e.g. substituent or bioisosteric), as will be described in more detail below.
[0063] In this specification, a transformation function is defined in part, or completely, by a transformation, e.g., a structural modification, between a first and second molecular fragment. For example, a transformation can involve a change to one or more adjacent atoms, functional group, or substituent, e.g., an atom or group of atoms attached to the core structure of a molecule. As another example, the structural modification can involve a change to one or moreatoms of the core structure of a molecule. As an example, a transformation can involve a modification of a first fragment methyl group to a second fragment ethyl group, a modification of a first fragment isobutyl group to a second fragment methoxy group, or a modification of an oxygen to a nitrogen atom.
[0064] A transformation function can also be defined by an inclusion criterion specifying which molecules are eligible for the transformation, e.g., based on the chemical environment of molecules used to define the transformation function. In particular, the MMP identification engine 130 can compare the structural data for molecules in the collection of molecules 110 to identify specific structural modifications between molecules with the same chemical environment as a transformation function, e.g., a same portion of the molecule that is separated from the molecular fragment modified in the transformation function by a threshold number of bonds. The system can then associate the transformation function with the chemical environment, e.g., by including the chemical environment in the transformation function as a target chemical structure.
[0065] More specifically, the engine 130 can identify molecules of the same class, e.g., with the same chemical environment 132. In particular, the chemical environment 132 can include the atoms present across a threshold number of bonds, e.g., two, four, or six bonds. For example, in FIG. 1A the chemical environment 132 shows a central atom (labelled “1”) and layers of atoms (“2”, “3”, “4”) that move away from the atom bond-by-bond (and thus in this example the chemical environment comprises a chemical environment defined by a number of bonds). In implementations, the system 100 can define the chemical environment 132 as part of the inclusion criterion that must be satisfied by a molecule in order to apply a particular transformation function, e.g., as a target chemical structure, e.g., as will be described in more detail below.
[0066] For example, the engine 130 can iterate through sets of molecules of the same class, e.g., where each set includes the same chemical environment 132 as specified by a certain threshold number of bonds, and can use each set of molecules with the chemical environment 132 to identify different transformation functions using the transformation(s), e.g., the structural modifications between a first fragment and a second fragment of the molecules of the same class, e.g., with the same core structure or a same chemical environment as defined by a number of atoms in the core structure. In particular, the engine 130 can identify matched molecular pairs by pairing up molecules with the same target chemical structure in which a same first fragment is modified to a same second fragment.
[0067] More specifically, at each iteration, the engine 130 can select identified molecules with the same chemical environment 132. The system 100 can then generate pairs of molecules defined by different transformation functions based on a modification from a first fragment to a second fragment, e.g., a transformation as defined by a given fragment pair 134, and can organize the pairs of molecules defined by the same transformation function, e.g., the same chemical environment 132 and fragment pair 134, as a determined subset of the matched molecular pairs 135.
[0068] In some cases, the engine 130 can identify one or more chemical environments and fragments in each of the molecules by fragmentizing the collection of molecules 110 into specific portions that can be used to identify each chemical environment 132 and fragment 134. For example, the engine 130 can fragmentize the collection of molecules 110 using Hussain- Rea fragmentation, as is described in Hussain J, Rea C. Computationally efficient algorithm to identify matched molecular pairs (MMPs) in large datasets. J Chem Inf Model. 2010 Mar 22;50(3):339-48. doi: 10.1021 / ci900450m. PMID: 20121045.
[0069] A transformation function is also defined by a predicted molecular property value change as a result of the transformation between the fragment pair 142. More specifically, the system 100 can use the determined matched molecular pairs 135 to define a predicted change 142, e.g., AProperty 142, in one or more molecular property values that can be attributed to the application of each transformation function.
[0070] In particular, the system 100 can process the associated molecular property data for the matched molecular pairs 135 for each transformation function to determine changes to molecular property values that are associated with the transformation function, e.g., using the molecular property value calculation engine 140. More specifically, for each molecular property value associated with the first molecule in each matched molecular pair for a particular transformation function, the engine 140 can analyze the impact of the transformation function on the molecular property using the molecular property value of the corresponding second molecule in the matched molecular pair.
[0071] For example, the engine 140 can identify the matched molecular pairs for a first transformation function and can determine a delta, e.g., the change or difference, in a particular molecular property value between the first molecule and the second molecule in each molecular matched pair of the transformation function. The engine 140 can then aggregate the deltas across the matched molecular pairs for each respective property value for the first transformation function to determine the predicted changes in molecular property values associated with the application of the transformation function.
[0072] In the particular example depicted, the engine 140 can determine the predicted change in the respective molecular property value 142 for the transformation function, e.g., using a measure of central tendency. As an example, the engine 140 can compute a mean of the deltas or a median of the deltas as the AProperty 142.
[0073] More specifically, the system 100 can determine the predicted change in the respective molecular property value 142, e.g., for each of the molecular property values, for each transformation function. The system 100 can then maintain the respective transformation function with the corresponding predicted changes in molecular property values as a template, e.g., for generating synthetic training examples, e.g., by applying the transformation function and the predicted change to relevant candidate molecules.
[0074] In some cases, the molecular property value engine 140 can determine a measure of dispersion for the deltas, e.g., a standard deviation value, skew, interquartile range, or mean absolute deviation. For example, the engine 140 can compare the measure of dispersion to a maximum threshold dispersion value to ensure that the measure of dispersion of the deltas does not exceed the threshold.
[0075] In particular, in the case that the measure of dispersion exceeds the maximum threshold value, the engine 140 can determine that the transformation function is not reliable for use in generating synthetic training examples, e.g., since the predicted change in the molecular property value is not precise, determining the predicted molecular property value for a synthetic training example using the predicted change is not necessarily reliable. In the case that the system 100 deems that a transformation function is unreliable, the system 100 can either discard the transformation function, e.g., instead of maintaining it in a database as described below, or refrain from using the transformation function for generating synthetic training examples.
[0076] The molecular property value engine 140 can, in general, comprise any Quantitative Structure-Activity Relationship (QSAR) model, or it can comprise a machine learning model. There are many examples; merely to illustrate a couple are SwissADME (“SwissADME: a free web tool to evaluate pharmacokinetics, drug-likeness and medicinal chemistry friendliness of small molecules”, Sci. Rep. (2017) 7:42717); and pkCSM (Pires et al., “pkCSM: Predicting Small-Molecule Pharmacokinetic and Toxicity Properties Using Graph-Based Signatures”, Journal of Medicinal Chemistry, 58(9), 4066-4072).
[0077] An example for generating the deltas, the predicted change in the molecular property value, and the measure of dispersion will be described in more detail with respect to FIG. 2.
[0078] The system 100 can maintain the transformation functions, e.g., each including the transformation between the first fragment and second fragment, the target chemical structure,e.g., the chemical environment used to define the class of molecule for the transformation function, and the corresponding predicted molecular property changes associated with the transformation function, in a transformation database 150. As an example, the transformation database 150 can include separate table for different types of transformation functions, e.g., a functional group transformation function table, a single atom transformation function table, a substituent transformation function table, etc.
[0079] For example, the tables can include rows for each transformation function, e.g., each row can include columns specifying the chemical structure modification defined by the transformation function, the target chemical structure associated with the transformation function, e.g., the chemical environment used to define the class of molecules to which the transformation function applies, and columns for each predicted change in molecular property value, e.g., a particular row-column index can include the predicted molecular property value change 142 for the transformation function.
[0080] In some cases, the tables can also include one or more statistics determined by the molecular property value engine 140. For example, the tables can also include columns for the number of matched molecular pairs that were used to define the transformation function, a list of the deltas, the measure of dispersion, etc.
[0081] The system 100 can use the maintained transformation functions with associated predicted molecular property value change(s) to generate synthetic training examples. In particular, the system 100 can generate a number of computationally-generated molecules using the transformation application engine 160. The engine 160 can apply each transformation function to relevant molecules that satisfy an inclusion criterion defining a class of molecules, e.g., by including the target chemical structure that was used to identify the transformation function, and can generate one or more corresponding predicted molecular property values based on the transformation function using the corresponding AProperty 142 for the transformation function.
[0082] For example, the engine 160 can iteratively apply each of the identified transformation functions, e.g., by iterating through the rows of each table in the database 160. In some cases, the system 100 can determine whether to apply a particular transformation function, e.g., by comparing the number of identified matched molecular pairs used to define the transformation function to a threshold number of pairs. In the case that the number of identified matched molecular pairs does not satisfy, e.g., exceed, the threshold number of pairs, the engine 160 can determine to not apply the transformation function to generate synthetic training examples.
[0083] As an example, the transformation engine 160 can process a set of original molecules 120 to determine the relevant molecules, e.g., the candidate molecules 165, that are applicable for each respective transformation function. In some cases, the system 100 can select the set of original molecules 120 from the collection of molecules 110, e.g., as the molecules that were not identified in any of the matched molecular pairs 135.
[0084] The engine 160 can identify candidate molecules 165 that satisfy an inclusion criterion for applying each transformation function, by identifying candidate molecules 165 of the same class of molecule that was used to define the transformation function. More specifically, the system can identify candidate molecules 165 that include the target chemical structure associated with the transformation function maintained in the transformation database 160. In particular, the engine 160 can identify candidate molecules 165 that include the first molecular fragment associated with the transformation function 170, e.g., a functional group, atom, or substituent, and the target chemical structure defining the class of molecule necessary for applying the transformation function 170.
[0085] As an example, the engine 160 can first identify molecules that include the first molecular fragment of the transformation function 170 and then determine whether any of the molecules including the first molecular fragment include the target chemical structure of the inclusion criterion. More specifically, the engine 160 can determine that all atoms in the candidate molecule 165 within the chemical environment defined by the threshold number of bonds that were used to define the transformation function 170 match the target chemical structure.
[0086] In some cases, the engine 160 can evaluate a molecular fingerprint of the first molecular fragment to determine whether the portion of the candidate molecule that excludes the first molecular fragment, e.g., the core structure, matches a target molecular fingerprint associated with the transformation function 170. In this context, a molecular fingerprint refers to an encoded digital representation of a molecular structure, e.g., a binary fixed-length vector where each entry represents the presence or absence of a structural chemical element. An example for determining whether to apply different transformation functions to a set of molecules will be described in more detail with respect to FIG. 3.
[0087] The engine 160 can then apply the transformation function 170 to generate a synthetic molecule including the modified fragment 174 and can determine the corresponding predicted molecular property value(s) for the candidate molecules using the predicted change in the respective molecular property value(s) 172. In particular, the engine 150 can generate new respective molecular property value(s) for a synthetic molecule by adding the molecularproperty value for the original un-transformed molecule and the predicted change in the molecular property value 172 specified for the transformation function 170.
[0088] The system 100 can maintain the generated synthetic training examples in a synthetic training example database 180. For example, the synthetic training examples can include the computationally-generated molecules with corresponding predicted molecular property value(s) or both. The data in the synthetic training example database 180 can be used to train a machine learning model 190 to predict molecular property values for a given input molecule, e.g., the system 100 can augment the data from the collection of molecules 110 using the computationally-generated molecules.
[0089] The machine learning model 190 can be any appropriate type of machine learning model, e.g., a neural network, a random forest, or a support vector machine, and can have any appropriate machine learning architecture that enables the machine learning model to perform its described functions, e.g., processing data characterizing a molecule to generate a predicted value of a molecular property of the molecule. For instance, in an implementation where the machine learning model is implemented as a neural network, then the neural network can have any appropriate number of neural network layers (e.g., 1 layer, 5 layers, or 10 layers) of any appropriate type (e.g., fully-connected layers, attention layers, convolutional layers, etc.) connected in any appropriate configuration (e.g., as a linear sequence of layers, or as a directed graph of layers). Merely as some examples, a sequence processing neural network, such as a Transformer neural network or a recurrent neural network, can be used to process sequence format data such as SMILES or InChi data, or a graph neural network can be used to process a MolBlock file.
[0090] As an example, the system 100 can train the machine learning model 190 using the synthetic training examples in the database 180 to reduce a discrepancy between: (i) a predicted molecular property value generated by processing the computationally-generated molecule of the synthetic training example using the machine learning model 190, and (ii) the target molecular property value specified by the synthetic training example. For example, the system 100 can train the machine learning model 190 on the set of training examples by a machine learning training technique to optimize an objective function that measures the discrepancy between the target and predicted output in any appropriate way, e.g., using a cross-entropy loss or a mean squared error loss.
[0091] The system can train the machine learning model 190 at each of a number of training iterations by any appropriate training technique until a training termination criterion is met. In the case that the model 190 is a neural network, the model 190 can be trained by calculatingand backpropagating gradients of an objective function to update parameter values of the model, e.g., using the update rule of any appropriate gradient descent optimization algorithm, e.g., RMSprop or Adam.
[0092] FIG. IB shows the example matched molecular pair (MMP) data augmentation and training system 100 of FIG. 1. While described above with respect to generating synthetic training examples for a collection of chemical reactions, as another example, the system 100 can be used to generate synthetic training examples for a collection of chemical reactions.
[0093] In this case, the system 100 can receive a collection of reactions 112, e.g., where each reaction includes a set of reactant molecules and a set of product molecules. Similarly to the collection of molecules 110, each of the molecules in the set of reactants and the set of products can be specified by a Simplified Molecular Input Line Entry System (SMILES) string, by an International Chemical Identifier (InChi) string, by structured data files, e.g., respective protein databank files, MolBlock files, or chemical markup files, etc.
[0094] In some cases, each reaction in the collection of reactions can be identified by reaction type, e.g., e.g., an Amide coupling reaction, a Suzuki reaction, a Heck reaction, etc. In particular, each reaction can include a label that indicates the type of reaction.
[0095] The system 100 can identify transformations between pairs of reactions, e.g., using the matched molecular pair (MMP) identification engine 130 and the reaction identification engine 144. In this case, the MMP identification engine 130 can identify matched molecular pairs 135 amongst reactants and products in the collection of reactions 112, e.g., as is described with respect to FIG. 1. In particular, the system 100 can identify the modification between a first fragment and a second fragment in a reactant in a first and second reaction, respectively, as a transformation and can ensure that the transformation is consistent across the set of reactants and the set of products in the first and second reactions. For example, the system 100 can identify a matched molecular pair between A to F for a first reaction including reactant A-B- C-D reacting with J-K-L to yield product A-B-C-D-J-K-L and a second reaction including F- B-C-D reacting with J-K-L to yield product F-B-C-D-J-K-L, e.g., where A was transformed to F in both the reactants and the product (optionally under the same or equivalent reaction conditions, such as pressure or temperature). In this case, molecular matched pairs 135 are identified using the reactants, but the approach can be applied to the products as well, e.g., as will be described in more detail below. This approach can also be used where a catalyst is present, e.g. considering the catalyst as part of the reaction conditions, or one or more different catalysts as parts of the reaction conditions.
[0096] The reaction identification engine 144 can then identify one or more pairs of reactions in the collection of reactions 112 using the molecular matched pairs 135. More specifically, the reaction identification engine 144 can identify multiple pairs of reactions in the collection of reactions 112, based on a transformation between a first reaction and a second reaction where the first reaction includes at least one reactant with a first molecular fragment and the second reaction includes at least one reactant with a second molecular fragment.
[0097] In particular, the reactant that includes the first molecular fragment and the reactant that includes the second molecular fragment can be considered a matched molecular pair, e.g., based on the defined transformation between fragments and the inclusion of the target chemical environment. For example, the fragments can be functional groups, one or more adjacent atoms, or substituents, and the engine 144 can pair up reactants from the first and second reactions with the same target chemical structure in which any instances of the same first fragment are modified to a same second fragment. The engine 144 can identify one or more matched molecular pairs between pairs of reactions, e.g., using MMPs 135 between respective sets of reactants for the pair of reactions, and can define the transformation between the one or more first fragments of the reactants and the products of the first reaction and the one or more corresponding second fragments of the reactants and the products of the second fragments as a reaction transformation function.
[0098] As an example, the engine 144 can identify a transformation function that includes a single modification between the first and second reactions, e.g., between an ethyl group and a methyl group for a reactant or a product with the same core structure between the first and the second reaction. As another example, the engine 144 can identify a transformation function that includes multiple modifications between the first and second reactions, e.g., an isobutyl group to a hydroxyl group for a first reactant or a first product with the same core structure and a methoxy group to an ethyl group for a second reactant or a second product with the same core structure between the first and second reactions.
[0099] In the case that the chemical reaction data is labelled by reaction type, the engine 144 can evaluate groups of reactions per reaction type to identify the transformation functions. More specifically, the engine 144 can identify reactants across pairs of reactions of the same reaction type that are matched molecular pairs and can define a transformation function using the matched molecular pairs for the reaction type, e.g., if there is a threshold number of pairs of chemical reactions of the chemical reaction type that include the transformation between the first and second fragment of the matched molecular pair. In this case, the transformation function can include the label specifying the reaction type for the transformation function.
[0100] In the particular example depicted, the system 100 can maintain the transformation functions between reactions, in a reaction transformation database 155. Since the reaction identification engine 144 identifies the reactions with matched molecular pair reactants and products from a collection of reactions 112, the identified reactions that correspond with the transformation function 146 are known reactive reactions.
[0101] In this case, the system 100 can distinguish between reactive reactions and non-reactive reactions with a reactivity indicator, e.g., a binary indicator indicative of whether or not the reaction is reactive. For example, the reactive reactions corresponding with the identified transformation functions can be associated with a positive reactivity indicator, e.g., a value of 1, and non-reactive reactions can be associated with a negative reactivity indicator, e.g., a value of 0 or -1. In some implementations, a reaction can be considered as reactive, e.g., if it results in the reaction product with at least some defined yield, under unspecified conditions or, optionally, under particular reaction conditions.
[0102] In some cases, the reaction identification engine 144 can determine a change in reactivity based on the transformation function. For example, the engine 144 can compare a measure of reactivity between the pairs of reactions, e.g., a reaction rate, an activation energy, or a bond dissociation energy. In this case, the engine 144 can determine a reactivity value that specifies whether there was an aggregated increase or decrease in a measure of reactivity for the pairs of reactions that include a matched molecular pair from a particular transformation function as the reactivity for the transformation function. In this case, the system 100 can maintain the change in reactivity with the transformation function in the reaction transformation database 155.
[0103] The system 100 can then apply the transformation functions maintained in the reaction transformation database 155 using the transformation engine 162 to generate synthetic training reactions for each transformation function. In particular, the transformation engine 162 can process a set of original reactions 120 to determine the corresponding relevant original reactions, e.g., the candidate reaction(s) 167, that are applicable for each respective transformation function.
[0104] In particular, the engine 162 can identify the reactions from the set of original reactions 120 with candidate reactant(s) and product(s) 167 that match the reactant(s) and product(s) that are modified in the transformation function 175, e.g., the transformation function 175 can specify a modification to a single fragment or multiple fragments. In the example given above, the engine 162 can identify all reactions with A and modify A to F, e.g., the engine 162 can identify and modify a third reaction including A-S-T-D reacting with J-K-L to yield productA-S-T-D-J-K-L to a new fourth reaction including F-S-T-D reacting with J-K-L to yield product F-S-T-D-J-K-L.
[0105] In the case that the transformation function 175 is defined for a particular reaction type, the transformation engine 162 can additionally determine that the candidate reach on(s) and product(s) 167 are of the same reaction type as the transformation function 175. In particular, the engine 162 can determine whether each candidate reaction has the same label as the transformation function 175.
[0106] For example, the engine 162 can iteratively apply each of the identified transformation functions, e.g., by iterating through the rows of each table in the database 155. In some cases, the system 100 can determine whether to apply a particular transformation function, e.g., by comparing the number of reactions that include the transformation function to a threshold number of reactions. In the case that the number of reactions for the transformation function does not satisfy, e.g., exceed, the threshold number of reactions, the engine 162 can determine to not apply the transformation function.
[0107] More specifically, the system 100 can apply the transformation function 175 to each of the candidate reactant(s) and product(s) 169 in the identified candidate reach on(s) 167 to generate modified reactant(s) or product(s) 178, e.g., reactant(s) or product(s) with the second fragment specified by the transformation function 175. In particular, the system 100 can then generate a new chemical reaction by replacing each instance of the candidate reactant(s) and product(s) 169 in the candidate reach on(s) 167 with the modified reactant(s) or product(s) 178 of the transformation function 175. The system 100 can then assign a positive reactivity indicator 176 to each of the new chemical reactions, e.g., since the transformation function was determined between known reactions, reactions generated using the transformation function 175 are considered to be reactive as well.
[0108] The system 100 can maintain the generated synthetic reaction examples in a synthetic reaction database 182. The data in the synthetic reaction database 182 can be used to train a machine learning model 192 to predict a reactivity indicator classifying whether an input set of reactants, and, in some cases, an input set of products is reactive.
[0109] Since all of the training examples in the database 182 are reactive, e.g., all the synthetic reactions have a positive reactivity indicator 176, the system 100 can obtain a set of non- reactive reactions 185 with corresponding negative reactivity indicators to ensure the training is not imbalanced and results in a model 192 that is able to predict the reactivity indicator. In particular, the system 100 can train the machine learning model 192 on a training datasetincluding a subset of the synthetic training reactions in the database 182 and a subset of known non-reactive reactions 185.
[0110] The machine learning model 192 can be any appropriate type of machine learning model, e.g., a neural network, a random forest, or a support vector machine, and can have any appropriate machine learning architecture that enables the machine learning model to perform its described functions, e.g., processing data characterizing a set of reactants, or the set of reactants and the corresponding set of products, to generate a predicted reactivity indicator that classifies whether or not a reaction will occur. For instance, in an implementation where the machine learning model is implemented as a neural network, then the neural network can have any appropriate number of neural network layers (e.g., 1 layer, 5 layers, or 10 layers) of any appropriate type (e.g., fully-connected layers, attention layers, convolutional layers, etc.) connected in any appropriate configuration (e.g., as a linear sequence of layers, or as a directed graph of layers).[OHl] As an example, the system 100 can train the machine learning model 192 to process a set of reactants and a corresponding set of products to generate a predicted reactivity indicator. In this case, the system 100 can train the machine learning model 192 on the training dataset to reduce a discrepancy between: (i) a predicted reactivity indicator generated by processing the set of reactants and the set of products of the synthetic training example using the machine learning model, and (ii) the target reactivity indicator specified by the synthetic training example.
[0112] As another example, the system 100 can train the machine learning model 192 to process the set of reactants to generate a predicted reactivity indicator. In this case, the system 100 can train the machine learning model 192 on the training dataset to reduce a discrepancy between: (i) a predicted reactivity indicator generated by processing the set of reactants of the synthetic training example using the machine learning model, and (ii) the target reactivity indicator specified by the synthetic training example.
[0113] For example, the system 100 can train the machine learning model 192 on the training dataset including the synthetic reaction training examples by a machine learning training technique to optimize an objective function that measures the discrepancy between the target and predicted output in any appropriate way, e.g., using a cross-entropy loss or a mean squared error loss.
[0114] The system can train the machine learning model 192 at each of a number of training iterations by any appropriate training technique until a training termination criterion is met. In the case that the model 192 is a neural network, the model 192 can be trained by calculatingand backpropagating gradients of an obj ective function to update parameter values of the model 192 , e.g., using the update rule of any appropriate gradient descent optimization algorithm, e.g., RMSprop or Adam.
[0115] FIG. 2 provides an example of calculating the predicted change in molecular property value for a transformation function to a molecule. For example, the molecular property value engine 140 of FIG. 1A can calculate the predicted change for the molecular property value as depicted in panel 200.
[0116] For example, the system 100 can identify the matched molecular pairs 210 for the transformation function 220 between each first molecule 212 and second molecule 214 pair, e.g., the pairs 230-240, 232-242, 234-244, and 236-246. In this case, the transformation function 220 defines a modification between a first fragment, e.g., fragment A 222, and a second fragment, e.g., fragment B 224, for four molecule pairs of the same class, e.g., that include the same target chemical structure. In particular, the chemical environment used to define the transformation function 220 was a core structure including a nitrogen and the start of a carbon ring.
[0117] In the particular example depicted, the molecular property value is the log solubility, e.g., a measure of the molecule’s solubility in water or another solvent. The system can determine the deltas 226, e.g., the change in the log solubility between the first molecule 212 and the second molecule 214 in each matched molecular pair 210. More specifically, in order to quantify the impact of the transformation function 240, the system can calculate the difference between the log solubility between molecule 240 and molecule 230, molecule 242 and molecule 232, molecule 244 and molecule 234, and molecule 246 and molecule 236 as the deltas 226.
[0118] The system can then aggregate the deltas 226 to determine the predicted change in the molecular property value, e.g., the predicted change in log solubility 250 corresponding with the transformation function 220. In this case, the system can compute the mean of the deltas 226 as the predicted change in the log solubility 250.
[0119] As described with respect to FIG. 1, the system can maintain the predicted change in log solubility 250 with the transformation function 220, e.g., in the transformation database 150 of FIG. 1, and can apply the transformation function 220 on candidate molecules that satisfy an inclusion criterion defining the class of molecules for the transformation function 220 based on the first fragment, e.g., fragment A 222, and the target chemical structure for the transformation function 220. The system can then use the predicted change in log solubility 250 to determine the predicted log solubility for each synthetic molecule the system generatesby applying the transformation function 220 to a candidate molecule, e.g., by adding the predicted log solubility 250 to the log solubility for the candidate molecule.
[0120] FIG. 3 illustrates how an example MMP data augmentation system can determine whether to apply different transformation functions to a set of molecules.
[0121] In the particular example depicted, the system maintains the transformation functions 310, e.g., transformation A 312, transformation B 314, and transformation C 316, e.g., in the transformation database 150 of FIG. 1. More specifically, each transformation function 312, 314, and 316 is maintained with the target chemical structure 320, e.g., the highlighted chemical environment defined by a threshold number of bonds, and the first and second fragment defining the transformation function.
[0122] As an example, transformation A 312 involves modifying fragment A 330 to fragment B 342 for candidate molecules with the target chemical structure 322, e.g., as defined by a one bond threshold chemical environment. As another example, transformation B 314 involves modifying fragment A 330 to fragment C 344 for candidate molecules with the target chemical structure 324, e.g., as defined by a two bond threshold chemical environment. As yet another example, transformation C 316 involves modifying fragment A 330 to fragment D 346 for candidate molecules with the target chemical structure 326, e.g., as defined by a three bond threshold chemical environment.
[0123] For example, the transformation engine 160 of FIG. 1 can determine whether to apply a transformation function as depicted in panel 350. In particular, the system can identify one or more candidate molecules, e.g., the candidate molecules 362, 364, 366, and 368 and can evaluate whether or not each is eligible for transformation A 312, B 314, C 316, or any combination of A 312, B 314, and C 316 based on an inclusion criterion determined by the inclusion of the first fragment 330 and the respective target chemical structure 320 for each transformation.
[0124] In particular, the system can determine whether each molecule is a candidate for a transformation by identifying whether the molecule includes the first fragment 330 and one or more of the respective target chemical structures, e.g., the target chemical structures 322, 324, and 326, associated with the transformation function. More specifically, the system can determine that molecule 362 is a candidate molecule for transformation A 312, B 314, and C 316 based on the presence of the first fragment 330 and the target chemical structures 322, 324, and 326.
[0125] Similarly, the system can determine that molecules 364 and 366 are candidate molecules for transformation A 312 based on the presence of the first fragment 330 and thetarget chemical structure 322. As another example, the system can determine that the molecule 368 is a candidate molecule for transformation A 312 and transformation B 314 based on the presence of the first fragment 330 and the target chemical structure 322 and 324.
[0126] After determining whether a molecule is a candidate for a particular transformation function, the system can apply the relevant transformation function(s), e.g., by modifying the first fragment A 330 to either fragment B 342, C 344, or D 346. For example, the system can generate three computationally-generated molecules using the candidate molecule 362, e.g., by replacing the fragment A 330 with fragment B 342, fragment C 344, and D 346 respectively, two computationally-generated molecules using the candidate molecule 368, e.g., by replacing the fragment A 330 with fragment B 342 and fragment C 344, and one synthetic molecule for each of the molecules 364 and 366, e.g., by replacing the fragment A 330 with fragment B 342.
[0127] FIG. 4 is a flow diagram of an example process for generating synthetic training examples using MMPs and training a machine learning model using the synthetic training examples. For convenience, the process 400 will be described as being performed by a system of one or more computers located in one or more locations. For example, a matched molecular pair data augmentation and training system, e.g., the system 100 of FIG. 1A, appropriately programmed in accordance with this specification, can perform the process 400 to generate a predicted property value for an input molecule.
[0128] The system can obtain data that identifies a collection of molecules, and a corresponding molecular property value for the molecule (step 410), e.g. for each of the molecules, data that characterizes the respective molecule. As an example, the data for each of the molecules in the collection of molecules can include a Simplified Molecular Input Line Entry System (SMILES) string, e.g., which provides a representation of the chemical structure of the molecule in a one-dimensional string, a molecular fingerprint of the molecule or a structured data file, e.g., a protein databank file, a MolBlock file, chemical markup file, etc.
[0129] The system can then process the data identifying the collection of molecules to identify a number of matched molecular pairs (MMPs) (step 420). In this case, each matched molecular pair includes a (i) first molecule and (ii) a second molecule that differs from the first molecule in that one molecular fragment in (of) the first molecule is replaced by another molecular fragment in (of) the second molecule.
[0130] The system can process the number of matched molecular pairs to generate data defining a set of transformation functions (step 430). In particular, each transformation function can be defined by at least a (i) first molecular fragment, a (ii) second molecular fragment, an (iii) inclusion criterion defining a class of molecules, and (iv) a predicted change in themolecular property value resulting from replacing the first molecular fragment by the second molecular fragment in a molecule that satisfies the inclusion criterion. More specifically, generating the data defining the set of transformation functions can involve identifying a subset of the number of matched molecular pairs and generating the transformation function based on the subset of the number of matched molecular pairs. For example, the system can identify the subset of matched molecular pairs for which (i) the first molecule in each matched molecular pair includes a same first fragment, (ii) the second molecule in each matched molecular pair differs from the corresponding first molecule in that the same first fragment is replaced by a same second fragment, (iii) the first molecule in each matched molecular pair satisfies a same inclusion criterion, and (iv) the subset of the number of matched molecular pairs includes at least two matched molecular pairs.
[0131] In particular, the system can determine a delta between the molecular property values of (i) the first molecule and (ii) the second molecule for each matched molecular pair in the subset of the number of matched molecular pairs that are used to generate the transformation function, and can determine the predicted change in the molecular property value for the transformation function based on a measure of central tendency, e.g., a mean, of the deltas. As another example, the system can determine that a measure of dispersion of the deltas does not exceed a maximum threshold, e.g., a standard deviation.
[0132] Moreover, a candidate molecule can satisfy the inclusion criterion for the transformation function if (i) the candidate molecule includes the first molecular fragment associated with the transformation function, and (ii) a chemical structure of a portion of the candidate molecule that includes the first molecular fragment and all atoms in the candidate molecule that are separated from the first molecular fragment by at most a threshold number of bonds in the candidate molecule matches a target chemical structure associated with the transformation function. As an example, the threshold number of bonds can be at least two bonds. Furthermore, a candidate molecule can satisfy the inclusion criterion for the transformation function if (i) the candidate molecule includes the first molecular fragment associated with the transformation, and (ii) a molecular fingerprint of the candidate molecule excluding the first molecular fragment matches a target molecular fingerprint associated with the transformation function.
[0133] The system can generate a number of synthetic training examples for training a machine learning model to predict the molecular property for an input molecule using the set of transformation functions (step 440). Each synthetic training example can include (i) a traininginput to the machine learning model that characterizes a molecule and (ii) a target output of the machine learning model that specifies a target molecular property value of the molecule.
[0134] For example, the system can apply the set of transformation functions on each of a number of original molecules in the collection of molecules. In particular, the system can determine, for each transformation function, whether the original molecule satisfies the inclusion criterion of the transformation function, and in response to determining that the original molecule satisfies the inclusion criterion of the transformation function, can generate a synthetic training example using the original molecule and the transformation function. More specifically, in the case that the original molecule satisfies the inclusion criterion, the system can generate a new molecule by replacing the first molecule fragment with the second molecular fragment in the original molecule. The system can then generate a molecular property value for the new molecule as a sum of (i) the molecular property value for the original molecule, and (ii) the predicted change in the molecular property value that is specified by the transformation function, and can generate a training example including a training input that characterizes the new molecule and the molecular property value for the new molecule as the target output of the machine learning model.
[0135] The system can then train the machine learning model using the number of synthetic training examples to generate a predicted molecular property value for an input molecule by a machine learning training technique (step 450). For example, the system 100 can augment the data from the collection of molecules 110 with the computationally-generated molecules. In particular, the system can process the training input using the machine learning model to generate a predicted molecular property value. In this case, the system can train the machine learning model to reduce a discrepancy between (i) the predicted molecular property value and the (ii) the target molecular property value specified by the synthetic training example.
[0136] For example, the machine learning model can be a neural network, and the system can update values of a set of parameters of the neural network using an objective function based on the discrepancy. In particular, the objective function can measure the discrepancy in any appropriate way, e.g., using a cross-entropy loss or a mean squared error loss. For example, the system can train the machine learning model by calculating and backpropagating gradients of the objective function using the update rule of any appropriate gradient descent optimization algorithm, e.g., RMSprop or Adam.
[0137] There is also described a computer-implemented method of predicting a molecular property value that involves obtaining data characterizing an input molecule, and processingthe data characterizing the input molecule using a machine learning model that has been trained as described above, to generate a predicted molecular property value for the input molecule.
[0138] As previously described, the molecular property value can be a predicted value for any quantifiable chemical, physical, or biological characteristic of a molecule, e.g. a physicochemical or pharmacological property of a molecule; or the molecular property value can be a weighted combination of values of such characteristics, e.g. a weighted combination of ADMET values.
[0139] Such a system may be used in many ways. For example, a set of input molecules may be filtered based on one or more of their predicted properties to obtain a subset of the molecules (containing less than all of the molecules) as candidate molecules, e.g. candidate drug molecules (e.g. a molecule that interacts with a drug target to produce a therapeutic effect). Some or all of these may afterwards be further investigated e.g. screened, in silico and / or after synthesizing a selected molecule, in vitro or in vivo. Similarly, the predicted properties can be used to select molecules for inclusion in a library for further screening. Or the predicted properties of the set of input molecules can be used to prioritize further investigation of a molecule, in silico, and / or after synthesizing the molecule, in vitro or in vivo. In general, a molecule may be selected, e.g. for synthesis, based on its predicted properties and / or excluded, e.g. for synthesis, based on its predicted properties, e.g. based on poor predicted pharmacokinetics.
[0140] Optionally a molecule can be selected based on values multiple predicted properties, e.g. for high permeability and low toxicity. As another example the candidate molecules can be candidate pesticide or herbicide molecules, in which case they can be selected for predicted toxicity to an animal or plant, but can also be selected for lack of toxicity to some other animals or plants.
[0141] FIG. 5 is a flow diagram of an example process for generating synthetic training examples using MMPs based on a collection of chemical reactions and training a machine learning model using the synthetic training examples. For convenience, the process 500 will be described as being performed by a system of one or more computers located in one or more locations. For example, a matched molecular pair data augmentation and training system, e.g., the system 100 of FIG. IB, appropriately programmed in accordance with this specification, can perform the process 500 to train a machine learning model to determine whether an input collection of molecules is reactive.
[0142] The system can obtain chemical reaction data that identifies a collection of chemical reactions (step 510). In particular, the system can receive data, e.g. structured data, that definesthe set of reactants and set of products of each chemical reaction. For example, such structured data can be comprise an input-output key -value dictionary, or be based on a relative positioning of the reactant data before a delimiter that signifies the start of the product data. In some cases, the chemical reaction data can be labeled by reaction type, e.g., each chemical reaction in the collection of chemical reactions can include a label defining a reaction type of the reaction, e.g., an Amide coupling reaction, a Suzuki reaction, a Heck reaction, etc. As an example, the data for each of the molecules in a chemical reaction can include a Simplified Molecular Input Line Entry System (SMILES) string, e.g., which provides a representation of the chemical structure of the molecule in a one-dimensional string, or a structured data file, e.g., a protein databank file, a MolBlock file, chemical markup file, etc.
[0143] The system can process the chemical reaction data to generate data defining a set of transformation functions (step 520). In particular, each transformation function can be defined by at least a first and second molecular fragment, e.g., between the first and second molecular fragment in a set of reactants and a set of products. More specifically, generating each transformation function can involve determining that (a) at least one reactant in a first chemical reaction of the pair of chemical reactions includes the first molecular fragment, and (b) a second set of reactants of a second chemical reaction of the pair of chemical reactions differs from a first set of reactants of the first chemical reaction of the pair of chemical reactions in that each instance of the first molecular fragment in the first set of reactants is replaced by the second molecular fragment in the second set of reactants, e.g., for at least a threshold number of pairs of chemical reactions in the collection of chemical reactions. As an example, the threshold number of pairs can be at least three pairs of chemical reactions.
[0144] In the case that the chemical reaction data is labelled by reaction type, the system can select a subset of the chemical reaction data corresponding with the chemical reactions that have the label defining the reaction type. The system can then determine one or more transformation functions associated with the reaction type using the subset of the chemical reaction data. More specifically, the system can identify reactants across pairs of reactions of the same reaction type that are matched molecular pairs and can define a transformation function using the matched molecular pairs for the reaction type, e.g., if there is a threshold number of pairs of chemical reactions of the chemical reaction type that include the transformation between the first and second molecule of the matched molecular pair.
[0145] The system can generate a number of synthetic training examples for training a machine learning model to predict whether data characterizing an input set of molecules is reactive using the set of transformations (step 530). In particular, each synthetic training example can include(i) a training input to the machine learning model that characterizes a set of reactants, and in some cases, a corresponding set of products, and (ii) a target output of the machine learning model that identifies the set of reactants and the set of products as being reactive.
[0146] For example, the system can apply the set of transformation functions on each of a number of original chemical reactions in the collection of chemical reactions. In particular, the system can determine, for each transformation function, whether the original chemical reaction has at least one reactant that includes the first molecular fragment of the transformation function, and in response to determining that the original chemical reaction has at least one reactant that includes the first molecular fragment of the transformation function, can generate a synthetic training example.
[0147] In the case that the chemical reaction data is labelled by reaction type, the system can additionally determine whether the original chemical reaction has the label defining the reaction type before applying the transformation function. In particular, the system can identify the reaction type associated with each transformation function, can determine whether the original chemical reaction has the label defining the reaction type, and, in response to determining that the original chemical reaction has the label defining the reaction type, can generate the new chemical reaction.
[0148] More specifically, in the case that the original molecule satisfies the inclusion criterion, the system can generate a new chemical reaction by replacing each instance of the first molecular fragment in a set of reactants and the set of products of the original chemical reaction with the second molecular fragment of the transformation function. The system can then generate a synthetic training example that includes (i) data that characterizes a set of reactants, and in some cases, the corresponding set of products of the new chemical reaction as the training input, and (ii) a target output that identifies the set of reactants as being reactive, or, in the case that the training input includes the set of products, the set of reactants and the set of products as being reactive.
[0149] In particular, in some cases, each training input can include both a set of reactants and a corresponding set of products, e.g., the system can configure the machine learning model to process a set of reactants and a set of products to generate a predicted reactivity indicator. In other cases, each training input can include the set of reactants, e.g., without the corresponding set of products. In this case, the system can configure the machine learning model to process the set of reactants to generate a predicted reactivity indicator.
[0150] The system can then train the machine learning model using the number of synthetic training examples to classify whether an input set of molecules is reactive by a machinelearning training technique (step 540). In particular, each input set of molecules can include a set of reactants, or a set of reactants and a corresponding set of products. More specifically, the system can assemble a training dataset using the new chemical reactions of the synthetic training examples and non-reactive chemical reactions, e.g., since all of the new chemical reactions are reactive.
[0151] The system can process the set of molecules characterized by the training input using the machine learning model to generate a reactivity indicator indicative of the reaction defined by the training input as being classified as reactive or not reactive. In particular, the system can train the machine learning model to reduce a discrepancy between (i) the generated reactivity indicator using the training input and the (ii) the target reactivity indicator specified by the training example, e.g., either a positive reactivity indicator of a synthetic training example or a negative reactivity indicator of a non-reactive chemical reaction.
[0152] For example, the machine learning model can be a neural network, and the system can update values of a set of parameters of the neural network using an objective function based on the discrepancy. In particular, the objective function can measure the discrepancy in any appropriate way, e.g., using a cross-entropy loss or a mean squared error loss. For example, the system can train the machine learning model by calculating and backpropagating gradients of an objective function using the update rule of any appropriate gradient descent optimization algorithm, e.g., RMSprop or Adam.
[0153] There is also described a computer-implemented method of determining whether the input set of molecules is reactive that involves obtaining data characterizing an input set of molecules, and processing the data characterizing the input set of molecules using a machine learning model that has been trained as described above, to generate a classification output that classifies whether or not the input set of molecules is reactive to form a reaction product, e.g. one or more product molecules. Optionally, a physical reaction can then be performed to confirm that input set of molecules is reactive as predicted, i.e. to physically synthesize one or more products of the reaction.
[0154] Such a system may be used in many ways. For example, as well as predicting reaction success it can also be used to screen reactants, e.g. in a set of candidate reactants, e.g. to find reactants for synthesizing a drug or other molecule, or for planning a retrosynthesis route for a product molecule. As another example the prediction of reaction success can be used to prioritize possible combinations of reactants for experimental verification that they can produce a desired reaction product.
[0155] This specification uses the term “configured” in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions.
[0156] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine- readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially-generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.
[0157] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0158] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as astand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub-programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.
[0159] In this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.
[0160] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.
[0161] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.
[0162] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way ofexample semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0163] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.
[0164] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and computeintensive parts of machine learning training or production, i.e., inference, workloads.
[0165] Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework, or a Jax framework.
[0166] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.
[0167] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. Therelationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.
[0168] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0169] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0170] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
Claims
CLAIMS1. A method performed by one or more computers, the method comprising: obtaining data that identifies a collection of molecules and, for each molecule in the collection of molecules, a molecular property value for the molecule; processing the data identifying the collection of molecules to identify a plurality of matched molecular pairs; processing the plurality of matched molecular pairs to generate data defining a set of transformation functions, wherein each transformation function is defined by at least:(i) a first molecular fragment;(ii) a second molecular fragment;(iii) an inclusion criterion defining a class of molecules; and(iv) a predicted change in the molecular property value resulting from replacing the first molecular fragment by the second molecular fragment in a molecule that satisfies the inclusion criterion; generating a plurality of synthetic training examples for training a machine learning model using the set of transformation functions, wherein the machine learning model is configured to process data characterizing an input molecule to generate a predicted molecular property value for the input molecule; and training the machine learning model using the plurality of synthetic training examples by a machine learning training technique.
2. The method of claim 1, wherein each matched molecular pair includes a pair of molecules comprising: (i) a first molecule, and (ii) a second molecule that differs from the first molecule in that one molecular fragment in the first molecule is replaced by another molecular fragment in the second molecule.
3. The method of any preceding claim, wherein for each transformation function, a candidate molecule satisfies the inclusion criterion for the transformation function if:(i) the candidate molecule includes the first molecular fragment associated with the transformation function; and(ii) a chemical structure of a portion of the candidate molecule that includes the first molecular fragment and all atoms in the candidate molecule that are separated from the first molecular fragment by at most a threshold number of bonds in the candidate molecule36matches a target chemical structure associated with the transformation function.
4. The method of claim 3, wherein the threshold number of bonds is at least two bonds.
5. The method of any one of claims 1-2, wherein for each transformation function, a candidate molecule satisfies the inclusion criterion for the transformation function if:(i) the candidate molecule includes the first molecular fragment associated with the transformation function; and(ii) a molecular fingerprint of the candidate molecule excluding the first molecular fragment matches a target molecular fingerprint associated with the transformation function.
6. The method of any preceding claim, wherein processing the plurality of matched molecular pairs to generate data defining the set of transformation functions comprises, for each transformation function: identifying a subset of the plurality of matched molecular pairs for which:(i) the first molecule in each matched molecular pair includes a same first fragment;(ii) the second molecule in each matched molecular pair differs from the corresponding first molecule in that the same first fragment is replaced by a same second fragment;(iii) the first molecule in each matched molecular pair satisfies a same inclusion criterion; and(iv) the subset of the plurality of matched molecular pairs includes at least two matched molecular pairs; and generating the transformation function based on the subset of the plurality of matched molecular pairs.
7. The method of claim 6, wherein processing the plurality of matched molecular pairs to generate data defining the set of transformation functions comprises, for each transformation function: determining, for each matched molecular pair in the subset of the plurality of matched molecular pairs that are used to generate the transformation function, a delta between the molecular property values of: (i) the first molecule, and (ii) the second molecule, in the matched molecular pair; and37determining the predicted change in the molecular property value for the transformation function based on a measure of central tendency of the deltas.
8. The method of claim 7, wherein the measure of central tendency is a mean.
9. The method of any one of claims 7-8, wherein processing the plurality of matched molecular pairs to generate data defining the set of transformation functions further comprises, for each transformation function: determining that a measure of dispersion of the deltas does not exceed a maximum threshold.
10. The method of claim 9, wherein the measure of dispersion is a standard deviation.
11. The method of any preceding claim, wherein generating the plurality of synthetic training examples for training the machine learning model using the set of transformation functions comprises, for each of a plurality of original molecules in the collection of molecules: determining, for each transformation function, whether the original molecule satisfies the inclusion criterion of the transformation function; and in response to determining that the original molecule satisfies the inclusion criterion of a transformation function: generating a new molecule by replacing the first molecule fragment with the second molecular fragment in the original molecule; generating a molecular property value for the new molecule as a sum of: (i) the molecular property value for the original molecule, and (ii) the predicted change in the molecular property value that is specified by the transformation function; and generating a synthetic training example that includes: (i) a training input to the machine learning model that characterizes the new molecule, and (ii) a target output of the machine learning model that specifies the molecular property value of the new molecule.
12. The method of any preceding claim, wherein each synthetic training example includes: (i) a training input to the machine learning model that characterizes a molecule, and (ii) a target output of the machine learning model that specifies a target molecular propertyvalue of the molecule.
13. The method of claim 12, wherein training the machine learning model using the plurality of synthetic training examples by the machine learning training technique comprises, for each synthetic training example: training the machine learning model to reduce a discrepancy between: (i) a predicted molecular property value generated by processing the training input of the synthetic training example using the machine learning model, and (ii) the target molecular property value specified by the synthetic training example.
14. The method of any preceding claim, wherein the machine learning model comprises a neural network.
15. A computer-implemented method of predicting a molecular property value, comprising: obtaining data characterizing an input molecule; and processing the data characterizing the input molecule using a machine learning model that has been trained by the method of any of claims 1-14, to generate a predicted molecular property value for the input molecule.
16. A method performed by one or more computers, the method comprising: obtaining chemical reaction data that identifies a collection of chemical reactions; processing the chemical reaction data to generate data defining a set of transformation functions, wherein each transformation function is defined by at least: (i) a first molecular fragment, and (ii) a second molecular fragment, and wherein generating each transformation function comprises: determining that, for at least a threshold number of pairs of chemical reactions in the collection of chemical reactions:(a) at least one reactant in a first chemical reaction of the pair of chemical reactions includes the first molecular fragment; and(b) a second set of reactants of a second chemical reaction of the pair of chemical reactions differs from a first set of reactants of the first chemical reaction of the pair of chemical reactions in that each instance of the first molecular fragment in the first set of reactants is replaced by the second molecular fragment in the second set of reactants;generating a plurality of synthetic training examples for training a machine learning model using the set of transformation functions, wherein the machine learning model is configured to process data characterizing an input set of molecules to classify whether the input set of molecules is reactive; and training the machine learning model using the plurality of synthetic training examples by a machine learning training technique.
17. The method of claim 16, wherein each chemical reaction in the collection of chemical reactions further comprises a label defining a reaction type of the reaction, and wherein processing the chemical reaction data to generate data defining the set of transformation functions further comprises, for each reaction type: selecting a subset of the chemical reaction data corresponding with chemical reactions that have the label defining the reaction type; determining one or more transformation functions associated with the reaction type using the subset of the chemical reaction data.
18. The method of any one of claims 16-17, wherein the threshold number of pairs of chemical reactions is at least three.
19. The method of any one of claims 16-18, wherein each synthetic training example comprises: (i) a training input to the machine learning model that characterizes a set of molecules, and (ii) a target output of the machine learning model that identifies the set of molecules as being reactive.
20. The method of any one of claims 16-19, wherein generating the plurality of synthetic training examples for training the machine learning model using the set of transformation functions comprises, for each of a plurality of original chemical reactions in the collection of chemical reactions: determining, for each transformation function, whether the original chemical reaction has at least one reactant that includes the first molecular fragment of the transformation function; and in response to determining that the original chemical reaction has at least one reactant that includes the first molecular fragment of the transformation function: generating a new chemical reaction by replacing each instance of the firstmolecular fragment in a set of reactants and in a set of products of the original chemical reaction with the second molecular fragment of the transformation function; and generating a synthetic training example that includes: (i) a training input to the machine learning model that characterizes at least the set of reactants of the new chemical reaction, and (ii) a target output of the machine learning model that identifies the set of reactants as being reactive.
21. The method of claim 20, wherein the training input to the machine learning model characterizes the set of reactants and the set of products of the new chemical reaction.
22. The method of any one of claims 20-21, when dependent on claim 17, further comprising: identifying the reaction type associated with each transformation function; determining whether the original chemical reaction has the label defining the reaction type; and in response to determining that the original chemical reaction has the label defining the reaction type, generating the new chemical reaction.
23. The method of any one of claims 16-22, wherein training the machine learning model using the plurality of synthetic training examples by the machine learning training technique comprises, for each synthetic training example: training the machine learning model to reduce a discrepancy between: (i) a classification generated by the machine learning model by processing the training input of the training example, and (ii) the target output of the machine learning model that identifies the set of molecules characterized by the training input as being reactive.
24. The method of any one of claims 16-23, wherein the machine learning model comprises a neural network.
25. A computer-implemented method of determining whether an input set of molecules is reactive, comprising: obtaining data characterizing an input set of molecules; and41processing the data characterizing the input set of molecules using a machine learning model that has been trained by the method of any of claims 16-24, to generate a classification output that classifies whether the input set of molecules is reactive.
26. A system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of any one of claims 1-25.
27. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective method of any one of claims 1-25.42