Generating training data using a generative machine learning model
A generative machine learning model generates training examples with biological context from textual data, addressing inefficiencies in existing data sets and enhancing the performance of molecular property prediction models.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- DEEPMIND TECH LTD
- Filing Date
- 2024-10-22
- Publication Date
- 2026-04-30
AI Technical Summary
Existing data sets for molecular property prediction tasks, such as protein-protein interactions, are inefficient and time-consuming to curate and lack sufficient biological context, limiting the performance of prediction machine learning models.
A system using a generative machine learning model processes textual data to automatically generate training examples with biological context, enabling large-scale and efficient production of training data for molecular property prediction tasks.
The system enhances the performance of prediction machine learning models by providing a larger number and variation of training examples with biological context, improving their ability to generalize to unseen inputs.
Smart Images

Figure EP2024079855_30042026_PF_FP_ABST
Abstract
Description
[0001] GENERATING TRAINING DATA USING A GENERATIVE MACHINE LEARNING MODEL
[0002] BACKGROUND
[0003] This specification relates to methods for training a prediction machine learning model to perform a molecular property prediction task.
[0004] A protein includes a sequence of amino acids. An amino acid is an organic compound which includes an amino functional group and a carboxyl functional group, as well as a side-chain (i.e., group of atoms) that is specific to the amino acid.
[0005] Predictions can be made using machine learning models. Machine learning models receive an input and generate an output, e.g., a predicted output, based on the received input. Some machine learning models are parametric models and generate the output based on the received input and on values of the parameters of the model. Some machine learning models are deep models that employ multiple layers of models to generate an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non-linear transformation to a received input to generate an output.
[0006] SUMMARY
[0007] This specification generally describes a system implemented as computer programs on one or more computers in one or more locations that trains a prediction machine learning model to perform a molecular property prediction task.
[0008] The system described in this specification can generate training examples for training a prediction machine learning model to perform a molecular property prediction task. The system can include a generative machine learning model that has been trained to perform a language modeling task.
[0009] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages.
[0010] Prediction machine learning models can be trained to perform molecular property prediction tasks, such as predicting whether a protein-protein interaction occurs for a given molecular system that includes a set of proteins. For example, given two amino acid sequences, the prediction machine learning model can predict a likelihood of a protein-protein interaction occurring. (Throughout this specification, a “protein-protein interaction” can refer to a physical contact and specific binding between two or more protein molecules, e.g., through non-covalent bonds, which can influence various biological functions, including signal transduction, molecular recognition, and structural assembly). Training a prediction machine learning model to perform a molecular property prediction task requires a large amount of training data. In addition, a prediction machine learning model can often perform a molecular property prediction task with higher accuracy when given more information, e.g., biological context, about the molecular system.
[0011] Manually curating a data set of molecular properties, such as a database of information related to protein-protein interactions, can be inefficient and time-consuming. Thus many existing data sets of molecular properties include protein-protein interaction results for a limited number of protein sets and a limited amount of biological context for each protein set. Although existing data sets may report a protein-protein interaction result for each protein set in the existing data set, the existing data sets do not include other biological context for each protein set that may make the reported protein-protein interaction possible. For example, a protein-protein interaction can be contingent on the existence of other molecules or ions, the size or subtype of proteins, or modifications to the proteins.
[0012] The system described in this specification can efficiently generate training examples from text sequences. For example, each text sequence can describe one or more molecular systems. The system can generate training examples using a generative machine learning model by processing a text sequence using the generative machine learning model to generate a model output that identifies, for each of one or more molecular systems described in the text sequence, one or more molecular properties of the molecular system that are extracted from the text sequence. The system can generate a respective training example for one or more of the molecular systems based on the one or more molecular properties. For example, the system can include biological context for a protein set in the training input of the training example. Thus the system can generate a large number of training examples for a molecular property prediction task that include biological context in the training inputs.
[0013] The system can train a prediction machine learning model on the generated training examples, resulting in better performance at inference compared to a prediction machine learning model trained on existing data sets. For example, training the prediction machine learning model on a larger number and greater variation of training examples, and training examples with a larger amount of biological context, allows the prediction machine learning model to generalize better to previously unseen inputs at inference. The prediction machine learning model can then be used in a range of applications, such as computational drug or ligand design, diagnostic antibody or aptamer marker screening, identification of the presence of a protein or nucleic acid mis-folding disease, or in determining structures of molecular systems.
[0014] The system described in this specification thus provides a technical solution to a technical problem, in particular, by enabling large scale, automated generation of training examples for training a prediction machine learning model to perform a molecular property prediction task. The system can perform large scale, automated generation of training examples by, for example, processing an underlying corpus of free form textual data using a generative machine learning model trained to perform a language modeling task.
[0015] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.
[0016] BRIEF DESCRIPTION OF THE DRAWINGS FIG. 1 A is a block diagram of an example training system.
[0017] FIG. IB is a block diagram of another example training system.
[0018] FIG. 2 is a flow diagram of an example process for training a prediction machine learning model.
[0019] FIG. 3 is a flow diagram of an example process for generating a model output using a generative machine learning model.
[0020] FIG. 4 is a flow diagram of an example process for generating a training example.
[0021] Like reference numbers and designations in the various drawings indicate like elements.
[0022] DETAILED DESCRIPTION FIG. 1A shows an example training system 100. The training system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations in which the systems, components, and techniques described below are implemented. The training system 100 can generate training examples such as the training examples 130a-n. Each training example 130a-n includes molecular properties of a molecular system. For example, each training example 130a-n can include a corresponding training input 132a-n that includes a first proper subset of molecular properties of the molecular system, and a corresponding target output 134a-n that includes a second proper subset of molecular properties of the molecular system. A proper subset of the molecular properties of a molecular system includes one or more but not all of the molecular properties of the molecular properties of the molecular system.
[0023] Each of the training examples 130a-n can be associated with a different molecular property prediction task. For example, the training example 130a can be for a particular molecular property prediction task. The training input 132a and target output 134a can include an appropriate first subset and second subset of molecular properties of the molecular system, respectively, for the particular molecular property prediction task. The training example 130b can be for a different molecular property prediction task. The training input 132b can include a subset of molecular properties that is appropriate for the different molecular property prediction task, where the subset of the training input 132b is different from the first subset of the training input 132a. The target output 134b can include a subset of molecular properties that is appropriate for the different molecular property prediction task, where the subset of the training output 134b is different from the second subset of the target output 134a. The training input 132b and target output 134b can include an appropriate first subset and second subset of molecular properties of the molecular system, respectively, for a different molecular property prediction task, that include different first subsets and second subsets, respectively, from the training input 132a and target output 134a.
[0024] Once the training system 100 has generated training examples for multiple molecular systems, the training system 100 can train a prediction machine learning model 150 on the training examples. The prediction machine learning model 150 can be configured to perform one or more molecular property prediction tasks, e.g., processing a model input that characterizes a molecular system in accordance with current values of parameters of the prediction machine learning model 150 to generate an output that includes data characterizing predicted molecular properties of the molecular system.
[0025] Some example molecular property prediction tasks include predicting whether molecules included in a given molecular system interact to form a molecule complex, predicting a binding affinity of molecules included in a given molecular system, predicting a location of one or more amino acid residues included in proteins in a given molecular system that are located at one or more interfaces in the given molecular system, or predicting a likelihood of a protein-protein interaction occurring for one or more given locations of amino acid residues. Example training inputs and target outputs for different molecular property prediction tasks are described with reference to FIG. 2.
[0026] The prediction machine learning model 150 can be implemented in any of a variety of possible ways. For instance, the prediction machine learning model 150 can be based on the AlphaFold2 model, as described in Jumper, John, et al. "Highly accurate protein structure prediction with AlphaFold." Nature 596.7873 (2021): 583-589. As another example, the prediction machine learning model 150 can be based on the AlphaFold3 model, as described in Abramson, Josh, et al. "Accurate structure prediction of biomolecular interactions with AlphaFold 3." Nature (2024): 1-3.
[0027] More generally, the prediction machine learning model 150 can be implemented as any appropriate type of machine learning model that can perform a molecular property prediction task, e.g., by processing data characterizing a molecular system to generate a prediction for one or more molecular properties of the molecular system. For instance, the prediction machine learning model can be implemented as a neural network model, or a random forest model, or a support vector machine model, and so forth. In implementations where the prediction machine learning model is implemented as a neural network, the prediction machine learning model can include any appropriate types of neural network layers (e.g., fully connected layers, message passing layers, convolutional layers, attention layers, recurrent layers, pooling layers, and so forth), in any appropriate number (e.g., 5 layers, or 10 layers, or 50 layers), and connected in any appropriate configuration (e.g., as a directed graph of layers).
[0028] As part of generating the training examples, the system 100 obtains a text sequence 102. The text sequence 102 includes natural language text that describes at least one molecular system. For example, the text sequence 102 can include the text of a scientific journal article or patent that describes at least one molecular system.
[0029] In some examples, the system 100 can obtain the text sequence 102 from a larger set of text sequences. The larger set of text sequences can include text sequences that each include natural language text. For each text sequence in the larger set of text sequences, the system 100 can process the text sequence to determine whether the text sequence describes at least one molecular system. For example, the system can provide the text sequence and an instruction to determine whether the text sequence describes at least one molecular system as input to a generative machine learning model to generate an output specifying whether the text sequence describes at least one molecular system. For each text sequence in the larger set of text sequences, in response to determining that the text sequence describes at least one molecular system, the system 100 can include the text sequence in a subset of text sequences. The system 100 can obtain the text sequence 102 from the subset of text sequences.
[0030] The system 100 generates one or more training examples 130a-n from the text sequence 102. To generate the training examples 130a-n, the system 100 can use a training data generation system 110 as described below with reference to FIG. IB.
[0031] The system 100 can generate one or more training examples for each of multiple text sequences that each describe at least one molecular system. As a particular example, the multiple text sequences can include a large number, e.g., at least 10,000, or at least 100,000, or at least 1 million text sequences. Training the prediction machine learning model 150 on training examples generated by the system 100 results in better performance at inference compared to a prediction machine learning model trained on a limited amount of training data. For example, training the prediction machine learning model on a larger number and greater variation of training examples allows the prediction machine learning model to generalize better to previously unseen inputs at inference.
[0032] The training data generation system can generate any appropriate number of training examples by processing the text sequences 102, e.g., at least 10,000, or at least 100,000, or at least 1 million training examples. Generating the training examples involves processing the text sequences 102 using a generative machine learning model, as will be described in more detail throughout this specification. The generative machine learning model, having been previously trained to perform a language modeling task, can operate rapidly at inference. For example, the generative machine learning model can perform a small number, e.g., less than five, or less than three, of forward passes through a text sequence to generate training examples corresponding to the molecular systems described in the text sequence. The generative machine learning model can thus generate large numbers of training examples, e.g., at least 10,000 training examples, in a short period of time, e.g., less than 15 hours, less than 10 hours, less than eight hours, or less than five hours. FIG. IB shows the example training system 100 described above with reference to FIG.
[0033] 1A. In particular, in the example of FIG. IB, the system 100 includes an input system 160, a generative machine learning model 170, and a molecular property processing system 180.
[0034] The system 100 obtains the text sequence 102 as described above with reference to FIG.
[0035] 1A.
[0036] To generate the training examples 130a-n, the system 100 processes the text sequence 102 using the training data generation system 110.
[0037] The training data generation system 110 can include an input system 160. The input system 160 can generate an input 162 for the generative machine learning model 170. For example, the input system 160 can receive the text sequence 102 and generate an input 162 that includes the text sequence 102.
[0038] In some examples, the input system 160 can also include a respective molecular system identifier 158 for each of one or more target molecular systems, and an instruction to extract molecular properties of the one or more target molecular systems from the text sequence 102, in the input 162. The respective molecular system identifier 158 can include, for example, data identifying one or more of: a protein name of each of one or more proteins in the respective molecular system, a domain name of each of one or more proteins in the respective molecular system, a residue span of each of one or more proteins in the respective molecular system, or an amino acid sequence of each of one or more proteins in the respective molecular system. In the example of FIG. IB, the molecular system identifier 158 can include a protein name for each of a set of proteins. A set of proteins includes two or more protein molecules. As a particular example, a set of proteins can include a pair of proteins.
[0039] As a particular example, the respective molecular system identifier 158 can include data identifying a set of proteins in the respective molecular system. For example, the respective molecular system identifier 158 can include a respective amino acid sequence of each of the proteins.
[0040] In some examples, the respective molecular system identifier 158 can include multiple names for each molecule in the respective molecular system. For example, a particular protein can have many names. The respective molecular system identifier 158 can include multiple names for the particular protein. In some examples, the input system 160 can obtain one or more of the multiple names for the particular protein from a database. For example, the input system 160 can obtain a protein name for the particular protein, e.g., from a user or using a generative machine learning model as described below. The input system 160 can query a database using the protein name to obtain one or more additional names for the particular protein.
[0041] In some examples, the respective molecular system identifier 158 can include data identifying a compound name of each of any compounds (e.g., organic compound, organometallic compound, or inorganic compound) in the respective molecular system.
[0042] In some examples, the input system 160 can obtain at least one of the one or more respective molecular system identifiers from a user, e.g., through a user interface of a user device. In some examples, the input system 160 can obtain at least one of the one or more respective molecular system identifiers by processing the text sequence 102 using a generative machine learning model. For example, the input system 160 can provide the text sequence 102 and an instruction to determine a molecular system identifier for each of the target molecular systems described in the text sequence 102 to generate the one or more respective molecular system identifiers.
[0043] In some examples, the input system 160 can also include an instruction to extract one or more types of molecular properties from the text sequence 102 in the input 162. Example inputs to the generative machine learning model 170 are described with reference to FIG. 2 and 3.
[0044] The training data generation system 110 processes the input 162 using the generative machine learning model 170 to generate a model output 172. The model output 172 identifies, for each of one or more molecular systems described in the text sequence 102, one or more molecular properties 174 of the molecular system that are extracted from the text sequence 102. Processing the input 162 using the generative machine learning model 170 is described below with reference to FIG. 2.
[0045] In some examples, the training data generation system 110 can process multiple inputs for the text sequence 102. An example of processing a first input and a second input to generate the model output 172 is described below with reference to FIG. 3.
[0046] The generative machine learning model 170 is trained to perform a language modeling task. For example, the generative machine learning model 170 can include a language model neural network. The language model neural network can have any of a variety of Transformerbased neural network architectures, e.g., encoder-only Transformer architectures, encoder- decoder Transformer architectures, decoder-only Transformer architectures, other attentionbased architectures, and so on.
[0047] Examples of such architectures include those described in Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683, 2019; TomB Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020; Aakanksha Chowdhery, et al. PaLM: Scaling Language Modeling with Pathways, arXiv preprint arXiv: 2204.02311; Rohan Anil, et al. Palm 2 technical report. arXiv preprint arXiv: 2305.10403, 2023, and Gemini Team, et al., Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
[0048] The language model neural network can be configured to generate output sequences made up of tokens from a vocabulary. In some examples, the vocabulary of tokens can include any of a variety of tokens that represent text symbols or other symbols. Lor example, the vocabulary of tokens can include one or more of characters, sub-words, words, punctuation marks, numbers, or other symbols that appear in a corpus of natural language text and / or computer code.
[0049] Additionally, or alternatively, the vocabulary of tokens can include tokens that can represent data other than text, such as images, videos, or audio. Lor example, the vocabulary of tokens can include image tokens that represent a discrete set of image patch embeddings of an image that can be generated by an image encoder neural network based on processing the image patches of the image. As another example, the vocabulary of tokens can include audio tokens that represent code vectors in a codebook of a quantizer, e.g., a residual vector quantizer.
[0050] Prior to using the language model neural network to generate model outputs, the language model neural network is pre-trained e.g., by the system 100 or by one or more other systems.
[0051] In particular, the system 100 or the other system(s) pre-trains the language model neural network on a language modeling task, e.g., a task that requires predicting, given a current sequence of text tokens, the next token that follows the current sequence in the training data. Equivalently, the language modeling task can require, for each given unlabeled text sequence in a training data set, predicting a text sequence that followed the given unlabeled text sequence in a corresponding document. As a particular example, the language model neural network can be pre-trained on a maximum-likelihood objective on a large dataset of text, e.g., text that is publicly available from the Internet or another text corpus.
[0052] After training, the system 100 can be configured to use the language model neural network to perform any appropriate machine learning task.
[0053] As an example, the language model neural network can generate text sequences, i.e., each output sequence generated by the language model neural networks is a sequence of text tokens from a vocabulary of text tokens that includes, e.g., one or more of characters, sub- words, words, punctuation marks, numbers, or other symbols that appear in natural language text.
[0054] In some cases, the language model neural networks can receive a context input, also referred to as a prompt, and generate an output sequence that is a response to the context input.
[0055] For example, the context input can be an input sequence of text and the output sequence is another sequence of text, e.g., a completion of the input sequence of text, a paraphrase of the input sequence of text, a response to a question posed in the input sequence, or a sequence of text that is about a topic specified by the input sequence of text. As another example, the context input can be an input other than text, e.g., an image, and the output sequence can be text that describes the input.
[0056] As another example, the context input represents data to be compressed, e.g., image data, text data, audio data, or any other type of data; and the output sequence is a compressed version of the data. The tokens included in the output sequence can include any representation of compressed data, e.g., symbols or embeddings to be decoded by a respective neural network.
[0057] In some examples, the generative machine learning model 170 can have been further trained on a fine-tuning dataset that includes multiple fine-tuning examples. Each fine-tuning example can include a fine-tuning input that includes a training text sequence and a ground-truth model output identifying, for each of one or more training molecular systems described in the training text sequence, one or more molecular properties of the training molecular system.
[0058] In some examples, the one or more molecular properties of the ground-truth model output can include one or more amino acid subsequences of protein molecules of a training molecular system. For example, the ground-truth model output can include data characterizing the start and end of one or more amino acid subsequences. The system 100 can obtain the data characterizing the start and end of one or more amino acid subsequences from a text sequence that describes the training molecular system. The training molecular system can be one of multiple training molecular systems for which the start and end of one or more amino acid subsequences that are involved in a physical experiment are described in a text sequence.
[0059] The training data generation system 110 processes the model output 172 using the molecular property processing system 180 to generate the training examples 130a-n. For one or more of the molecular systems of the model output 172, the molecular property processing system 180 can generate a respective training example based on the one or more molecular properties of the molecular system. In some examples, the molecular property processing system 180 can generate multiple training examples for each of the one or more molecular systems, e.g., training examples for different molecular property prediction tasks. In the example of FIG. IB, the molecular property processing system 180 can generate the training examples 130a-n based on the molecular properties 174 for a particular molecular system of the model output 172. Generating a training example is described in further detail below with reference to FIG. 4.
[0060] FIG. 2 is a flow diagram of an example process 200 for training a prediction machine learning model. For convenience, the process 200 will be described as being performed by a system of one or more computers located in one or more locations. For example, a training system, e.g., the training system 100 depicted in FIGS. 1A-1B, appropriately programmed in accordance with this specification, can perform the process 200.
[0061] The system obtains multiple text sequences (step 202). Each text sequence can describe one or more molecular systems. The system can obtain the text sequences, e.g., by accessing one or more databases of patents, scientific publications, and so forth. Each text sequence can include any appropriate number of tokens (e.g., characters, word pieces, or words), e.g., 100, 1000, or 10,000 tokens. Generally different text sequences can include different numbers of tokens.
[0062] The system generates multiple training examples (step 204). The system generates multiple training examples using a generative machine learning model by performing steps 206-208 for each text sequence.
[0063] For each text sequence, the system processes an input including the text sequence using the generative machine learning model to generate a model output (step 206). The model output identifies, for each of one or more molecular systems described in the text sequence, one or more molecular properties of the molecular system that are extracted from the text sequence. In some examples, the input can also include an instruction to extract one or more types of molecular properties from the text sequence. For example, the input can include an instruction to extract one or more types of molecular properties from the text sequence for each molecular system described in the text sequence.
[0064] In some examples, the input can include a respective molecular system identifier for each of one or more target molecular systems. The input can include an instruction to extract one or more types of molecular properties of the one or more target molecular systems from the text sequence.
[0065] In some examples, the input can also include an instruction to format the molecular properties into a structured data record. In some examples, the input can also include one or more extraction examples. Each extraction example can include an example text sequence and an example structured data record for an example molecular system described by the example text sequence.
[0066] As a particular example, a structured data record can include key-value pairs. For example, the key-value pair for a molecular property can include a label for the type of the molecular property and the molecular property.
[0067] In some examples, the input can include one or more key-value pair examples. Each keyvalue pair example can include a label for a type of molecular property and an example molecular property.
[0068] In some examples, the instruction to format the molecular properties into a structured data record can include a template of the structured data record and an instruction to fill in the template. For example, the template can include a list of keys with empty values to be filled in. For example, the value for a key can be filled in with a molecular property of the type identified by the label of the key, or a “null” value. As a particular example, the instruction can include:
[0069] You are a helpful research assistant. Your task is to read a PDF and fill in the entries to the following blank JSON:
[0070] acc:,
[0071] start.:,
[0072] end:,
[0073] reference fid:,
[0074] reference source :, reference Jitml:,
[0075] term_name: "protein binding",
[0076] ecjiame:,
[0077] construct alterations:
[0078] [{
[0079] term_name:,
[0080] start:,
[0081] end:,
[0082] position:
[0083] }, >...............
[0084] In some implementations, processing an input to generate a model output can include processing multiple inputs using the generative machine learning model to generate the model output, as described in further detail below with reference to FIG. 3.
[0085] Some example types of molecular properties include a respective identifier for each molecule (e.g., protein, small molecule, ribonucleic acid (RNA) molecule, or deoxyribonucleic acid (DNA) molecule) or ion involved in a physical experiment conducted for a molecular system, a number of molecules or ions involved in the physical experiment, or an experimental method of the physical experiment.
[0086] As a particular example, the input can include an instruction to extract molecular properties for molecular systems that include two or more protein molecules that interact to form a protein complex. The instruction can identify multiple types of molecular properties to be extracted from the text sequence.
[0087] For example, the molecular properties can include a respective identifier (e.g., name, database identifier) for each protein molecule included in the protein complex. The respective identifier can include a name or a database identifier for the protein molecule.
[0088] As another example, the molecular properties can include a number of additional molecules involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex. As another example, the molecular properties can include a respective identifier (e.g., name, database identifier) for each of any additional molecules involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex. Additional molecules can include, for example, protein molecules, small molecules, RNA, or DNA molecules. As another example, the molecular properties can include a number of ions involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex. The molecular properties can also include a respective identifier (e.g., name) for any ions involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex.
[0089] As another example, the molecular properties can include locations and types of any modified amino acid residues of the protein molecules in the protein complex. For example, the location can include the residue identifier for the modified amino acid residue. The type of modified amino acid residue can include, for example, phosphorylation, acetylation, or methylation.
[0090] As another example, the molecular properties can include one or more amino acid subsequences of the protein molecules of the molecular system. For example, the one or more amino acid subsequences of the protein molecules can be amino acid subsequences involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex.
[0091] As another example, the molecular properties can include stoichiometry (e.g., numbers or proportions) of one or more amino acid subsequences of the protein molecules of the molecular system. For example, the stoichiometry can indicate a number of copies of each amino acid subsequence involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex.
[0092] As another example, the molecular properties can include a respective identifier for an isoform (e.g., protein variant) of each protein molecule of the molecular system. For example the respective identifier can identify a member of a set of proteins that originate from a single gene or gene family.
[0093] As another example, the molecular properties can include a location of one or more amino acid residues included in the protein molecules in the molecular system that are located at one or more interfaces in the molecular system. For example, the location of the one or more amino acid residues can include the residue identifier for the start and end of the interface.
[0094] As another example, the molecular properties can include a binding affinity of the protein molecules of the molecular system. As another example, the molecular properties can include a strength of the molecular system.
[0095] As another example, the molecular properties can include an experimental method of a physical experiment conducted to determine that the protein molecules interact to form the protein complex.
[0096] In some examples, the model output includes a “null” value for one or more of the molecular properties. For example, if the text sequence does not describe a particular type of molecular property, the model output can include a “null” value for the molecular property.
[0097] For each text sequence, the system generates, for one or more of the molecular systems, a respective training example for training a prediction machine learning model (step 208). For example, the system generates the respective training example based on the one or more molecular properties of the molecular system, as described below in further detail with reference to FIG. 4.
[0098] In some implementations, the system generates a respective training example for each of the molecular systems. In some implementations, the system determines to generate a training example for a molecular system based on a confidence measure of the generative machine learning model. In some examples, the generative machine learning model can hallucinate, e.g., generate a model output that identifies molecular properties not described in the text sequence. Training the prediction machine learning model on training examples generated from such model outputs can result in compromised training of the prediction machine learning model and poor performance at inference. If the confidence measure of the generative machine learning model in the model output indicates a low amount of confidence, the system can determine that the model output is a hallucination. Thus the system can use the confidence measure to identify possible hallucinations and refrain from generating training examples from possible hallucinations, and training the prediction machine learning model on training examples generated from possible hallucinations.
[0099] For example, for each of the molecular systems, the system can determine a confidence measure of the generative machine learning model in the model output that identifies the molecular properties of the molecular system. In particular, the system can determine the confidence measure in the portion of the model output that identifies the molecular properties of the molecular system. The confidence measure can be determined in a number of different ways. In some examples, the system can determine the confidence measure based on the logits (e.g., scores or inputs to a soft-max layer) output by the generative machine learning model for the portion of the model output that identifies the molecular properties of the molecular system. For example, the generative machine learning model can generate an output sequence of tokens. At each position in the output sequence of tokens, the generative model processes (i) the input to the generative machine learning model, and (ii) tokens at any preceding output positions in the output sequence of tokens, to generate a score distribution over a set of possible tokens. The generative machine learning model then selects a token to occupy the position based on the score distribution over the set of possible tokens, e.g., by selecting the token with the highest score, or by processing the score distribution using a soft-max layer to generate a probability distribution, and then sampling, e.g., using top-k sampling, nucleus sampling or another sampling technique, a token using the probability distribution. In some examples, the system can generate the confidence measure, e.g., as a function (e.g., a product, sum, or measure of central tendency, such as a mean value) of the scores for the tokens defining the molecular properties of the molecular system.
[0100] As another example, the system can determine the confidence measure based on the entropy of the probability distribution over the set of possible tokens at each position. For example, a higher entropy can indicate a lower measure of confidence, and a lower confidence measure.
[0101] In some examples, the system can determine the confidence measure by providing the portion of the model output that identifies the molecular properties of the molecular system and the text sequence as input to the generative machine learning model. For example, the system can also provide an instruction to generate a confidence measure for the portion of the model output given the text sequence.
[0102] The system can determine to generate a training example corresponding to the molecular system only if the confidence measure of the generative machine learning model in the model output that identifies the molecular properties of the molecular system satisfies a threshold. For example, in response to determining that the confidence measure satisfies the threshold, the system generates the training example for the molecular system as described with reference to FIG. 4. In some examples, the threshold is a predetermined threshold. In some examples, the system can tune the threshold. In some examples, the system can determine to further train, e.g., fine-tune, the generative machine learning model if the system determines that the confidence measure has not satisfied the threshold for a threshold number or proportion of molecular systems.
[0103] In response to determining that the confidence measure does not satisfy the threshold, the system does not generate the training example for the molecular system. For example, for one or more text sequences of the multiple text sequences, the system can determine that a training example should not be generated for a molecular system described in the text sequence because the confidence measure of the generative machine learning model in the model output that identifies the molecular properties of the molecular system does not satisfy the threshold. Thus, by using the confidence measure to identify possible hallucinations generated by the generative machine learning model, the system can refrain from generating training examples from such hallucinations, and refrain from training the prediction machine learning model on training examples generated from such hallucinations, which can result in compromised training and poor performance for the prediction machine learning model.
[0104] The system trains the prediction machine learning model to perform a molecular property prediction task (step 210). For example, the system can train the prediction machine learning model on the multiple training examples generated in step 208 by a machine learning training technique.
[0105] For example, the system can apply a gradient descent with backpropagation training technique that uses, e.g., a stochastic gradient descent, RMSprop, or Adam optimizer, or another known or learned optimizer, to optimize an objective function that is appropriate for the molecular property prediction task of the training examples. The exact forms of the objective function may vary across different tasks, but typically, the objective function measures a quality of the target output, e.g., that measures a difference between the training output of the prediction machine learning model generated based on the training input of a training example and the known, target output (or another target output that is derived from the known, target output) of the training example. A cross-entropy loss function, e.g., in the case of classification tasks, and a mean squared error (MSE) loss function, e.g., in the case of regression tasks, are examples of suitable objective functions that can be used by the system during the training. FIG. 3 is a flow diagram of an example process 300 for generating a model output using a generative machine learning model. For convenience, the process 300 will be described as being performed by a system of one or more computers located in one or more locations. For example, a training system, e.g., the training system 100 of FIGS. 1A-1B, appropriately programmed in accordance with this specification, can perform the process 300.
[0106] The system performs steps 302 and 304 to generate a model output that identifies, for each of one or more molecular systems described in the text sequence, one or more molecular properties of the molecular system, in a structured data record. As an example, the structured data record can include a key-value pair for each of one or more molecular properties for each of the one or more molecular systems. For example, for each molecular system, each key-value pair can include a label for the type of molecular property, and the molecular property of the molecular system. As a particular example, the structured data record can have a JavaScript Object Notation (JSON) structure.
[0107] The system processes a first input using the generative machine learning model to generate a first model output (step 302). The first input can include (i) the text sequence, and (ii) an instruction to generate a free form (e.g., unstructured text) summary of content in the text sequence describing molecular systems. The first model output includes a free form summary of content from the text sequence describing molecular systems.
[0108] The system processes a second input using the generative machine learning model to generate a second model output (step 304). The second input includes (i) the first model output that includes the free form summary of the content from the text sequence describing molecular systems, and (ii) an instruction to format the free form summary into a structured data record. The second model output includes a structured data record.
[0109] In some examples, the second input to the generative machine learning model can include one or more key-value pair examples. Each key-value pair example can include a label for a type of molecular property and an example molecular property. In some examples, the second input can also include an instruction to generate the structured data record according to the keyvalue pair examples.
[0110] In some examples, the instruction to format the free form summary into a structured data record can include a template of the structured data record and an instruction to fill in the template. For example, the template can include a formatted list of keys. FIG. 4 is a flow diagram of an example process 400 for generating a training example. For convenience, the process 400 will be described as being performed by a system of one or more computers located in one or more locations. For example, a training system, e.g., the training system 100 of FIGS. 1A-1B, appropriately programmed in accordance with this specification, can perform the process 400.
[0111] The system includes a first proper subset of the molecular properties of the molecular system in a training input of the training example (step 402). In some examples, the system includes data derived from one or more molecular properties of the first proper subset.
[0112] In some examples, the first proper subset includes data identifying each molecule (e.g., protein molecules, small molecules) included in the molecular system. For example, if the molecular system includes one or more proteins, the first proper subset can include a respective amino acid sequence of each of the one or more proteins. If the molecular system includes one or more nucleic acids, the first proper subset can include a respective nucleic acid sequence of each of the one or more nucleic acids. If the molecular system includes one or more small molecules, the first proper subset can define a respective chemical structure of each of the one or more small molecules, e.g., by respective Simplified Molecular Input Line Entry System (SMILES) strings. In these examples, the system can derive the data identifying each molecule from the respective identifiers for the molecules included in the model output.
[0113] In some examples, the first proper subset can include other molecular properties or data derived from other molecular properties, such as a number of additional molecules involved in a physical experiment conducted to determine that protein molecules interact to form a protein complex, a number of ions involved in a physical experiment conducted to determine that protein molecules interact to form the protein complex; a respective identifier for any ions involved in a physical experiment conducted to determine that protein molecules interact to form the protein complex; locations and types of any modified amino acid residues of protein molecules in the protein complex; one or more amino acid subsequences of protein molecules of the molecular system; stoichiometry of one or more amino acid subsequences of protein molecules of the molecular system; a respective identifier for an isoform of each protein molecule of the molecular system; a location of one or more amino acid residues included in protein molecules in the molecular system that are located at one or more interfaces in the molecular system; a binding affinity of protein molecules of the molecular system; a strength of the molecular system; or an experimental method of a physical experiment conducted to determine that protein molecules interact to form the protein complex.
[0114] The system includes a second proper subset of the molecular properties of the molecular system in a target output of the training example (step 404).
[0115] In some examples, the second proper subset includes data characterizing one or more of: whether molecules included in the molecular system interact to form a molecule complex; a binding affinity of molecules included in the molecular system; or one or more amino acid residues included in proteins in the molecular system that are located at one or more interfaces in the molecular system.
[0116] Example training inputs and target outputs for different molecular property prediction tasks are described below. The first proper subset and second proper subset of molecular properties can include different molecular properties for different molecular property prediction tasks.
[0117] As a particular example, the molecular property prediction task can include predicting whether molecules included in a given molecular system interact to form a molecule complex. For example, the molecular property prediction task can include predicting whether a proteinprotein interaction occurs for a given set of protein molecules to form a protein complex. The training input can include data identifying the set of protein molecules.
[0118] In some examples, the training input can also include any one or more of a number of additional molecules involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex, a respective identifier for each of any additional molecules involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex, a number of ions involved in a physical experiment conducted to determine that protein molecules interact to form the protein complex, a respective identifier for any ions involved in a physical experiment conducted to determine that protein molecules interact to form the protein complex, locations and types of any modified amino acid residues of protein molecules in the protein complex, one or more amino acid subsequences of protein molecules of the molecular system, stoichiometry of one or more amino acid subsequences of protein molecules of the molecular system, a respective identifier for an isoform of each protein molecule of the molecular system, a location of one or more amino acid residues included in protein molecules in the molecular system that are located at one or more interfaces in the molecular system, or an experimental method of a physical experiment conducted to determine that protein molecules interact to form the protein complex.
[0119] The target output can include whether molecules identified in the training input interact to form a molecule complex. For example, the target output can include a binary indicator that indicates whether a protein-protein interaction occurs for the set of protein molecules identified in the training input. In some examples, the system can derive the binary indicator from the binding affinity of the protein molecules.
[0120] As another example, the molecular property prediction task can include predicting a binding affinity of molecules included in a given molecular system. For example, the molecular property prediction task can include predicting a binding affinity between protein molecules. The training input can include data identifying the set of protein molecules.
[0121] In some examples, the training input can also include any one or more of a number of additional molecules involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex, a respective identifier for each of any additional molecules involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex, a number of ions involved in a physical experiment conducted to determine that protein molecules interact to form the protein complex, a respective identifier for any ions involved in a physical experiment conducted to determine that protein molecules interact to form the protein complex, locations and types of any modified amino acid residues of protein molecules in the protein complex, one or more amino acid subsequences of protein molecules of the molecular system, stoichiometry of one or more amino acid subsequences of protein molecules of the molecular system, a respective identifier for an isoform of each protein molecule of the molecular system, a location of one or more amino acid residues included in protein molecules in the molecular system that are located at one or more interfaces in the molecular system, or an experimental method of a physical experiment conducted to determine that protein molecules interact to form the protein complex.
[0122] The target output can include data characterizing a binding affinity of molecules included in the molecular system identified by the training input. For example, the system can include the binding affinity of the protein molecules in the target output.
[0123] As another example, the molecular property prediction task can include predicting a location of one or more amino acid residues included in proteins in a given molecular system that are located at one or more interfaces in the molecular system. That is, the molecular property prediction task can include predicting which interface residues are involved in a protein-protein interaction in the given molecular system. The training input can include at least data identifying each protein molecule.
[0124] In some examples, the training input can include any one or more of a number of additional molecules involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex, a respective identifier for each of any additional molecules involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex, a number of ions involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex, a respective identifier for any ions involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex, locations and types of any modified amino acid residues of the protein molecules in the protein complex, one or more amino acid subsequences of the protein molecules of the molecular system, stoichiometry of one or more amino acid subsequences of the protein molecules of the molecular system, a respective identifier for an isoform of each protein molecule of the molecular system, a binding affinity of the protein molecules of the molecular system, a strength of the molecular system, or an experimental method of a physical experiment conducted to determine that the protein molecules interact to form the protein complex.
[0125] The target output can include data characterizing one or more amino acid residues included in proteins in the molecular system that are located at one or more interfaces in the molecular system identified by the training input. For example, the system can include the location of one or more amino acid residues included in the protein molecules in the molecular system that are located at one or more interfaces in the molecular system in the target output.
[0126] As another example, the molecular property prediction task can include predicting a likelihood of a protein-protein interaction occurring for one or more given locations of amino acid residues. That is, the molecular property prediction task can include predicting a likelihood of a protein-protein interaction occurring given interface residues of protein molecules. The training input can include at least data identifying locations of amino acid residues.
[0127] In some examples, the training input can also include any one or more of a number of additional molecules involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex, a respective identifier for each of any additional molecules involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex, a number of ions involved in a physical experiment conducted to determine that protein molecules interact to form the protein complex, a respective identifier for any ions involved in a physical experiment conducted to determine that protein molecules interact to form the protein complex, locations and types of any modified amino acid residues of protein molecules in the protein complex, one or more amino acid subsequences of protein molecules of the molecular system, stoichiometry of one or more amino acid subsequences of protein molecules of the molecular system, a respective identifier for an isoform of each protein molecule of the molecular system, a location of one or more amino acid residues included in protein molecules in the molecular system that are located at one or more interfaces in the molecular system, or an experimental method of a physical experiment conducted to determine that protein molecules interact to form the protein complex.
[0128] The target output can include a likelihood of a protein-protein interaction occurring for the one or more given locations of amino acid residues of the training input. For example, the system can derive the likelihood from the binding affinity of the protein molecules.
[0129] As used in this specification, a“molecular system” includes at least one organic or organometallic compound, polymer, protein, ligand, or nucleic acid. In some examples, a molecular system can also include an inorganic compound (e.g., crystalline inorganic compound, semiconductor, magnetic or superconductor compound, or bioinorganic compound) or alloy. For example, the molecular property prediction task can comprise predicting one or more material properties of an inorganic compound, semiconductor or alloy. The systems and methods described in this specification therefore have applications in materials science and metallurgy, for example.
[0130] Predictions of the prediction machine learning model can be validated experimentally, or by the use of computational techniques. For example, predicted protein-protein interactions may be evaluated using standard molecular (e.g. protein-ligand) docking or molecular dynamics software. For example, in some implementations, the biological activity of the molecular systems may be tested in vitro and / or in vivo. For example the candidate molecules may be tested for ADME (absorption, distribution, metabolism, excretion) and / or toxicological properties, to screen out unsuitable ligands. The testing may include bringing components of the molecular system into contact and measuring a change in expression or activity, e.g., bringing a candidate small molecule, polypeptide or polynucleotide ligand into contact with a target molecule (e.g. protein) and measuring a change in expression or activity of the target molecule.
[0131] This specification uses the term “configured” in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions.
[0132] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially-generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.
[0133] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub-programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.
[0134] In this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.
[0135] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.
[0136] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.
[0137] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0138] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.
[0139] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and computeintensive parts of machine learning training or production, i.e., inference, workloads.
[0140] Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework, or a Jax framework.
[0141] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.
[0142] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.
[0143] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0144] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0145] Features of the invention are also described in the following numbered clauses. In some implementations, the features described in these numbered clauses may be combined with features described above.
[0146] 1. A method performed by one or more computers, the method comprising:
[0147] obtaining a plurality of text sequences;
[0148] generating a plurality of training examples using a generative machine learning model trained to perform a language modeling task, comprising, for each text sequence:
[0149] processing an input comprising the text sequence using the generative machine learning model to generate a model output that identifies, for each of one or more molecular systems described in the text sequence, one or more molecular properties of the molecular system that are extracted from the text sequence; and
[0150] generating, for one or more of the molecular systems, a respective training example for training a prediction machine learning model based on the one or more molecular properties of the molecular system; and
[0151] training the prediction machine learning model to perform a molecular property prediction task by training the prediction machine learning model on the plurality of training examples by a machine learning training technique.
[0152] 2. The method of clause 1, wherein obtaining the plurality of text sequences comprises obtaining the plurality of text sequences from a larger set of text sequences, comprising:
[0153] for each text sequence in the larger set of text sequences:
[0154] processing the text sequence to determine whether the text sequence describes at least one molecular system; and
[0155] in response to determining that the text sequence describes at least one molecular system, including the text sequence in the plurality of text sequences to be processed using the generative machine learning model.
[0156] 3. The method of clause 1 or 2, wherein the input to the generative machine learning model comprises: (i) a respective molecular system identifier for each of one or more target molecular systems, and (ii) an instruction to extract molecular properties of the one or more target molecular systems from the text sequence.
[0157] 4. The method of clause 3, wherein each molecular system identifier comprises data identifying one or more of: a compound name, a protein name, a domain name, a residue span, or an amino acid sequence.
[0158] 5. The method of any of clauses 1-4, wherein the input to the generative machine learning model comprises an instruction to extract one or more types of molecular properties from the text sequence.
[0159] 6. The method of clause 5, wherein the input to the generative machine learning model further comprises an instruction to format the molecular properties into a structured data record.
[0160] 7. The method of clause 5 or 6, wherein the input to the generative machine learning model further comprises one or more extraction examples, each extraction example comprising an example text sequence and an example structured data record for an example molecular system described by the example text sequence.
[0161] 8. The method of any one of clauses 1-7, wherein the model output identifies the one or more molecular properties in a structured data record, and wherein processing an input comprising the text sequence using the generative machine learning model to generate a model output comprises:
[0162] processing a first input that comprises: (i) the text sequence, and (ii) an instruction to generate a free form summary of content in the text sequence describing molecular systems, using the generative machine learning model to generate a first model output that includes a free form summary of content from the text sequence describing molecular systems; and processing a second input that comprises: (i) the first model output that includes the free form summary of the content from the text sequence describing molecular systems, and (ii) an instruction to format the free form summary into a structured data record, using the generative machine learning model to generate a second model output that includes a structured data record. 9. The method of clause 8, wherein the second input to the generative machine learning model comprises one or more key-value pair examples, each key-value pair example comprising a label for a type of molecular property and an example molecular property.
[0163] 10. The method of any one of clauses 1-9, wherein the generative machine learning model has been further trained on a fine-tuning dataset comprising a plurality of fine-tuning examples, each comprising a fine-tuning input comprising a training text sequence and a ground-truth model output identifying, for each of one or more training molecular systems described in the training text sequence, one or more molecular properties of the training molecular system.
[0164] 11. The method of any one of clauses 1-10, wherein the input to the generative machine learning model comprises an instruction to extract molecular properties for molecular systems that include two or more protein molecules that interact to form a protein complex;
[0165] wherein the instruction identifies a plurality of types of molecular properties to be extracted from the text sequence, including one or more of:
[0166] a respective identifier for each protein molecule included in the protein complex; a number of additional molecules involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex;
[0167] a respective identifier for each of any additional molecules involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex;
[0168] a number of ions involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex;
[0169] a respective identifier for any ions involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex;
[0170] locations and types of any modified amino acid residues of the protein molecules in the protein complex;
[0171] one or more amino acid subsequences of the protein molecules of the molecular system;
[0172] stoichiometry of one or more amino acid subsequences of the protein molecules of the molecular system;
[0173] a respective identifier for an isoform of each protein molecule of the molecular system;
[0174] a location of one or more amino acid residues included in the protein molecules in the molecular system that are located at one or more interfaces in the molecular system;
[0175] a binding affinity of the protein molecules of the molecular system; a strength of the molecular system; or
[0176] an experimental method of a physical experiment conducted to determine that the protein molecules interact to form the protein complex.
[0177] 12. The method of any one of clauses 1-11, wherein generating, for one or more of the molecular systems, a respective training example for training a prediction machine learning model based on the one or more molecular properties of the molecular system comprises, for each of the one or more molecular systems:
[0178] including a first proper subset of the molecular properties of the molecular system in a training input of the training example; and
[0179] including a second proper subset of the molecular properties of the molecular system in a target output of the training example.
[0180] 13. The method of clause 12, wherein the first proper subset of the molecular properties that are included in the training input of the training example comprise data identifying each molecule included in the molecular system.
[0181] 14. The method of clause 12 or 13, wherein the second proper subset of the molecular properties that are included in the target output of the training example comprise data characterizing one or more of:
[0182] whether molecules included in the molecular system interact to form a molecule complex;
[0183] a binding affinity of molecules included in the molecular system; or
[0184] one or more amino acid residues included in proteins in the molecular system that are located at one or more interfaces in the molecular system. 15. The method of any one of clauses 1-14, wherein the molecular property prediction task comprises one or more of: predicting whether molecules included in a given molecular system interact to form a molecule complex, predicting a binding affinity of molecules included in a given molecular system, predicting a location of one or more amino acid residues included in proteins in a given molecular system that are located at one or more interfaces in the given molecular system, and predicting a likelihood of a protein-protein interaction occurring for one or more given locations of amino acid residues.
[0185] 16. The method of any one of clauses 1-15, wherein generating, for one or more of the molecular systems, a respective training example for training the prediction machine learning model based on the one or more molecular properties of the molecular system comprises, for each of the one or more molecular systems:
[0186] determining a confidence measure of the generative machine learning model in the model output that identifies the molecular properties of the molecular system; and
[0187] determining to generate a training example corresponding to the molecular system only if the confidence measure of the generative machine learning model in the model output that identifies the molecular properties of the molecular system satisfies a threshold.
[0188] 17. The method of clause 16, further comprising, for one or more text sequences of the plurality of text sequences, determining that a training example should not be generated for a molecular system described in the text sequence because the confidence measure of the generative machine learning model in the model output that identifies the molecular properties of the molecular system does not satisfy the threshold.
[0189] 18. The method of any one of clauses 1-17, wherein the plurality of text sequences comprise at least 10,000 text sequences.
[0190] 19. The method of any one of clauses 1-18, further comprising:
[0191] for each of a plurality of candidate molecular systems, using the trained prediction machine learning model to perform the molecular property prediction task by processing a respective input characterizing the candidate molecular system to determine a prediction of one or more molecular properties of the candidate molecular system; and
[0192] selecting at least one molecular system of the plurality of candidate molecular systems based on the predictions of the one or more molecular properties.
[0193] 20. The method of clause 19, further comprising generating an instruction to synthesize the selected at least one molecular system.
[0194] 21. The method of clause 19 or 20, further comprising synthesizing the selected at least one molecular system.
[0195] 22. A method of obtaining a drug or a ligand for an industrial enzyme, the method comprising the method of any one of clauses 1-18 and:
[0196] for each of one or more candidate ligands, using the trained prediction machine learning model to perform the molecular property prediction task by processing a respective input characterizing a molecular system comprising the candidate ligand and a target molecule to determine a corresponding prediction of one or more molecular properties of the molecular system; and
[0197] selecting, at least one of the candidate ligands as the drug or ligand based on the predictions of the one or more molecular properties for the one or more candidate ligands.
[0198] 23. The method of clause 22, wherein the target molecule comprises a receptor or enzyme, and wherein the ligand is an agonist or antagonist of the receptor or enzyme.
[0199] 24. The method of clause 22 or 23, wherein the candidate ligand and the target molecule are each proteins and the molecular property prediction task comprises determining whether a protein-protein interaction occurs to form a protein complex.
[0200] 25. The method of any one of clause 22-24, wherein the target molecule comprises a receptor or enzyme, and wherein the ligand is an agonist or antagonist of the receptor or enzyme; or wherein the ligand comprises an antibody or aptamer and the target molecule comprises an antibody or aptamer target, in particular a virus or cancer cell protein, and wherein the antibody or aptamer binds to the antibody or aptamer target to provide a therapeutic effect.
[0201] 26. A method of obtaining a diagnostic antibody or aptamer marker of a disease, the method comprising the method of any one of clauses 1-18 and:
[0202] selecting a target molecule;
[0203] for each of one or more candidate antibodies or aptamers, using the trained prediction machine learning model to perform the molecular property prediction task by processing a respective input characterizing a molecular system comprising the candidate antibody or aptamer and a target molecule to determine a corresponding prediction of one or more molecular properties of the molecular system; and
[0204] selecting one of the one or more of the candidate antibodies or aptamers as the diagnostic antibody or aptamer marker based on the on the predictions of the one or more molecular properties for the one or more candidate antibodies or aptamers.
[0205] 27. The method of clause 26, further comprising synthesizing the ligand or diagnostic antibody or aptamer marker;
[0206] 28. The method of clause 27, further comprising testing biological activity of the ligand or diagnostic antibody or aptamer marker in vitro and / or in vivo.
[0207] 29. A method of identifying the presence of a protein or nucleic acid mis-folding disease, the method comprising the method of any one of clauses 1-18, wherein the molecular property prediction task is to determine a predicted structure of a protein or nucleic acid, the method further comprising:
[0208] using the trained prediction machine learning model to perform the molecular property prediction task by processing a respective input characterizing a molecular system comprising a nucleic acid or protein to determine a corresponding predicted structure for the nucleic acid or protein;
[0209] obtaining a structure of a version of the protein or nucleic acid obtained from a human or animal body; comparing the predicted structure of the protein or nucleic acid with the structure of a version of the protein or nucleic acid obtained from a human or animal body; and identifying the presence of a protein or nucleic mis-folding disease dependent upon a result of the comparison.
[0210] 30. A method of determining the structure of a molecular system, the method comprising the method of any one of clauses 1-18, wherein the molecular property prediction task is to determine a predicted structure of a molecular system, the method further comprising:
[0211] applying an experimental technique to a physical sample comprising the molecular system to measure experiment signals dependent on a structure of the molecular system; using the trained prediction machine learning model to perform the molecular property prediction task by processing a respective input characterizing the molecular system to determine a corresponding predicted structure for the molecular system;
[0212] using the experiment signals and the predicted structure of the molecular system to determine the structure of the molecular system.
[0213] 31. The method of clause 30, wherein the experimental technique comprises one or more of: x-ray crystallography, nuclear magnetic resonance, and electron microscopy.
[0214] 32. A system comprising:
[0215] one or more computers; and
[0216] one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the method of any one of clauses 1-27.
[0217] 33. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform the method of any one of clauses 1-27. 34. A system comprising:
[0218] one or more computers; and
[0219] one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:
[0220] obtaining a plurality of text sequences;
[0221] generating a plurality of training examples using a generative machine learning model trained to perform a language modeling task, comprising, for each text sequence:
[0222] processing an input comprising the text sequence using the generative machine learning model to generate a model output that identifies, for each of one or more molecular systems described in the text sequence, one or more molecular properties of the molecular system that are extracted from the text sequence; and
[0223] generating, for one or more of the molecular systems, a respective training example for training a prediction machine learning model based on the one or more molecular properties of the molecular system; and
[0224] training the prediction machine learning model to perform a molecular property prediction task by training the prediction machine learning model on the plurality of training examples by a machine learning training technique.
[0225] 35. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:
[0226] obtaining a plurality of text sequences;
[0227] generating a plurality of training examples using a generative machine learning model trained to perform a language modeling task, comprising, for each text sequence:
[0228] processing an input comprising the text sequence using the generative machine learning model to generate a model output that identifies, for each of one or more molecular systems described in the text sequence, one or more molecular properties of the molecular system that are extracted from the text sequence; and
[0229] generating, for one or more of the molecular systems, a respective training example for training a prediction machine learning model based on the one or more molecular properties of the molecular system; and
[0230] training the prediction machine learning model to perform a molecular property prediction task by training the prediction machine learning model on the plurality of training examples by a machine learning training technique.
[0231] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
Claims
CLAIMS1. A method performed by one or more computers, the method comprising:obtaining a plurality of text sequences;generating a plurality of training examples using a generative machine learning model trained to perform a language modeling task, comprising, for each text sequence:processing an input comprising the text sequence using the generative machine learning model to generate a model output that identifies, for each of one or more molecular systems described in the text sequence, one or more molecular properties of the molecular system that are extracted from the text sequence; andgenerating, for one or more of the molecular systems, a respective training example for training a prediction machine learning model based on the one or more molecular properties of the molecular system; andtraining the prediction machine learning model to perform a molecular property prediction task by training the prediction machine learning model on the plurality of training examples by a machine learning training technique.
2. The method of claim 1, wherein obtaining the plurality of text sequences comprises obtaining the plurality of text sequences from a larger set of text sequences, comprising:for each text sequence in the larger set of text sequences:processing the text sequence to determine whether the text sequence describes at least one molecular system; andin response to determining that the text sequence describes at least one molecular system, including the text sequence in the plurality of text sequences to be processed using the generative machine learning model.
3. The method of any preceding claim, wherein the input to the generative machine learning model comprises: (i) a respective molecular system identifier for each of one or more target molecular systems, and (ii) an instruction to extract molecular properties of the one or more target molecular systems from the text sequence.
4. The method of claim 3, wherein each molecular system identifier comprises data identifying one or more of: a compound name, a protein name, a domain name, a residue span, or an amino acid sequence.
5. The method of any preceding claim, wherein the input to the generative machine learning model comprises an instruction to extract one or more types of molecular properties from the text sequence.
6. The method of claim 5, wherein the input to the generative machine learning model further comprises an instruction to format the molecular properties into a structured data record.
7. The method of claim 6, wherein the input to the generative machine learning model further comprises one or more extraction examples, each extraction example comprising an example text sequence and an example structured data record for an example molecular system described by the example text sequence.
8. The method of any preceding claim, wherein the model output identifies the one or more molecular properties in a structured data record, and wherein processing an input comprising the text sequence using the generative machine learning model to generate a model output comprises:processing a first input that comprises: (i) the text sequence, and (ii) an instruction to generate a free form summary of content in the text sequence describing molecular systems, using the generative machine learning model to generate a first model output that includes a free form summary of content from the text sequence describing molecular systems; and processing a second input that comprises: (i) the first model output that includes the free form summary of the content from the text sequence describing molecular systems, and (ii) an instruction to format the free form summary into a structured data record, using the generative machine learning model to generate a second model output that includes a structured data record.
9. The method of claim 8, wherein the second input to the generative machine learning model comprises one or more key-value pair examples, each key-value pair example comprising a label for a type of molecular property and an example molecular property.
10. The method of any preceding claim, wherein the generative machine learning model has been further trained on a fine-tuning dataset comprising a plurality of fine-tuning examples, each comprising a fine-tuning input comprising a training text sequence and a ground-truth model output identifying, for each of one or more training molecular systems described in the training text sequence, one or more molecular properties of the training molecular system.
11. The method of any preceding claim, wherein the input to the generative machine learning model comprises an instruction to extract molecular properties for molecular systems that include two or more protein molecules that interact to form a protein complex;wherein the instruction identifies a plurality of types of molecular properties to be extracted from the text sequence, including one or more of:a respective identifier for each protein molecule included in the protein complex; a number of additional molecules involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex;a respective identifier for each of any additional molecules involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex;a number of ions involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex;a respective identifier for any ions involved in a physical experiment conducted to determine that the protein molecules interact to form the protein complex;locations and types of any modified amino acid residues of the protein molecules in the protein complex;one or more amino acid subsequences of the protein molecules of the molecular system;stoichiometry of one or more amino acid subsequences of the protein molecules of the molecular system;a respective identifier for an isoform of each protein molecule of the molecular system;a location of one or more amino acid residues included in the protein molecules in the molecular system that are located at one or more interfaces in the molecular system;a binding affinity of the protein molecules of the molecular system; a strength of the molecular system; oran experimental method of a physical experiment conducted to determine that the protein molecules interact to form the protein complex.
12. The method of any preceding claim, wherein generating, for one or more of the molecular systems, a respective training example for training a prediction machine learning model based on the one or more molecular properties of the molecular system comprises, for each of the one or more molecular systems:including a first proper subset of the molecular properties of the molecular system in a training input of the training example; andincluding a second proper subset of the molecular properties of the molecular system in a target output of the training example.
13. The method of claim 12, wherein the first proper subset of the molecular properties that are included in the training input of the training example comprise data identifying each molecule included in the molecular system.
14. The method of any one of claims 12-13, wherein the second proper subset of the molecular properties that are included in the target output of the training example comprise data characterizing one or more of:whether molecules included in the molecular system interact to form a molecule complex;a binding affinity of molecules included in the molecular system; orone or more amino acid residues included in proteins in the molecular system that are located at one or more interfaces in the molecular system.
15. The method of any preceding claim, wherein the molecular property prediction task comprises predicting whether molecules included in a given molecular system interact to form a molecule complex, predicting a binding affinity of molecules included in a given molecular system, predicting a location of one or more amino acid residues included in proteins in a given molecular system that are located at one or more interfaces in the given molecular system , or predicting a likelihood of a protein-protein interaction occurring for one or more given locations of amino acid residues.
16. The method of any preceding claim, wherein generating, for one or more of the molecular systems, a respective training example for training the prediction machine learning model based on the one or more molecular properties of the molecular system comprises, for each of the one or more molecular systems:determining a confidence measure of the generative machine learning model in the model output that identifies the molecular properties of the molecular system; anddetermining to generate a training example corresponding to the molecular system only if the confidence measure of the generative machine learning model in the model output that identifies the molecular properties of the molecular system satisfies a threshold.
17. The method of claim 16, further comprising, for one or more text sequences of the plurality of text sequences, determining that a training example should not be generated for a molecular system described in the text sequence because the confidence measure of the generative machine learning model in the model output that identifies the molecular properties of the molecular system does not satisfy the threshold.
18. The method of any preceding claim, wherein the plurality of text sequences comprise at least 10,000 text sequences.
19. A system comprising:one or more computers; andone or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one ormore computers, cause the one or more computers to perform operations of the respective method of any one of claims 1-18.
20. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective method of any one of claims 1-18.
Citation Information
Patent Citations
Machine learning systems for automated pharmaceutical molecule identification
US20220101972A1
Machine learning based methods of analysing drug-like molecules
US20220383992A1