In context learning for molecular inverse design

A generative machine learning model iteratively updates molecule exemplars with predicted properties to efficiently generate novel molecules with desired characteristics, addressing inefficiencies in molecular design by minimizing resource consumption and enhancing computational efficiency.

WO2025245338A1PCT designated stage Publication Date: 2025-11-27SANOFI SA(FR)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/030559
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-23
Filing Date
2025-05-22
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

Existing molecular design processes are resource-intensive, time-consuming, and inefficient in generating novel molecules with specific properties tailored to address scientific or industrial challenges, particularly in drug development, due to the high cost and scarcity of high-quality experimental data.

Method used

A system utilizing a generative machine learning model processes a set of molecule exemplars and design parameters to generate novel molecules, iteratively updating the exemplars with computationally predicted properties, reducing the need for extensive training and experimental data, and optimizing computational resources.

Benefits of technology

The system efficiently generates diverse, valid molecules with desired properties, such as drug-like characteristics, by leveraging in-context learning and active selection of output molecules for exemplars, thereby reducing computational and laboratory resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025030559_27112025_PF_FP_ABST
    Figure US2025030559_27112025_PF_FP_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for molecular inverse design. In one aspect, a method comprises: generating a model input to a generative machine learning model, wherein the model input includes data defining: a set of molecule exemplars, wherein each molecule exemplar corresponds to a respective example molecule and includes: (i) a chemical structure of the example molecule, and (ii) a respective molecule property value of the example molecule for each molecule property in a set of molecule properties; and a set of molecule design parameters that specify desired characteristics of output molecules to be generated by the generative machine learning model; and processing the model input using the generative machine learning model to generate a model output that includes data defining a respective chemical structure of each output molecule in a set of output molecules.
Need to check novelty before this filing date? Find Prior Art

Description

IN CONTEXT LEARNING FOR MOLECULAR INVERSE DESIGNCROSS-REFERENCE TO RELATED APPLICATION|0001] This application claims priority to EP Application No. 24315243.6, filed on May 23, 2024, the disclosure of which is incorporated herein by reference in its entirety.BACKGROUND

[0002] This specification relates to processing data using machine learning models.

[0003] Machine learning models receive an input and generate an output, e.g., a predicted output, based on the received input. Some machine learning models are parametric models and generate the output based on the received input and on values of the parameters of the model.

[0004] Some machine learning models are deep models that employ multiple layers of models to generate an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non-linear transformation to a received input to generate an output.

[0005] This specification also relates to molecular inverse design. Molecular inverse design is a computational approach used to predict molecular structures that meet specified criteria, e.g., specific properties tailored to address particular scientific or industrial challenges.SUMMARY[0006| This specification describes a system implemented as computer programs on one or more computers in one or more locations that can generate a set of output molecules using a generative machine learning model. In particular, the generative machine learning model can process a model input that includes a set of molecule exemplars, each exemplar including the chemical structure of an example molecule and a set of corresponding molecule properties for the molecule, and a set of design parameters that specify one or more characteristics or properties of the molecules to be generated, e.g., as part of molecular inverse design.

[0007] The system can leverage the capabilities of the generative machine learning model to understand and approximate the real distribution of chemical space, which contains an almost infinite variety of potential molecular configurations. More specifically, the generative machine learning model can process the set of molecule exemplars to leam adistribution in chemical space and can generate new compounds that fit the learned chemical distribution, e.g.. to generate molecules that adhere to the underlying rules governing the input data.|0008] In particular, the system can generate a set of molecule exemplars and a set of molecule design parameters specifying desired values or ranges of values for the molecule properties in the set of molecule properties as input to the generative machine learning model. As an example, the set of design parameters can specify optimal drug-like properties suitable for preclinical and clinical evaluation. In some cases, one or more of the output molecules can be modified using a textual query, e.g., that identifies a specific structure change for the output molecule.|(IO09] According to a first aspect there is provided a method for generating a model input to a generative machine learning model, wherein the model input includes data defining: a set of molecule exemplars, wherein each molecule exemplar corresponds to a respective example molecule and includes: (i) a chemical structure of the example molecule, and (ii) a respective molecule property value of the example molecule for each molecule property in a set of molecule properties, and a set of molecule design parameters that specify desired characteristics of output molecules to be generated by the generative machine learning model, processing the model input using the generative machine learning model and in accordance with values of a set of generative machine learning model parameters to generate a model output that includes data defining a respective chemical structure of each output molecule in a set of output molecules, and outputting data identifying the set of output molecules.[001 | In some implementations, the model input to the generative machine learning model further comprises a chemical structure of an input molecule, and each output molecule has a chemical structure that is a modified version of the chemical structure of the input molecule.[0011| In some implementations, for each of one or more molecule properties in the set of molecule properties, the set of molecule design parameters specify a target range of values of the molecule property.

[0012] In some implementations, for each of a plurality of molecule exemplars in the set of molecule exemplars, the molecule property7values included in the molecule exemplar are experimental molecule property values determined by physical experiments on instances of the corresponding example molecule.[00131 In some implementations, the method further comprises updating the model input to the generative machine learning model, comprising generating a respective new molecule exemplar corresponding to each of one or more output molecules in the set of output molecules generated by the generative machine learning model, and adding each new molecule exemplar to the set of molecule exemplars included in the model input to the generative machine learning model, and processing the updated model input using the generative machine learning model and in accordance with values of the set of generative machine learning model parameters to generate an updated model output that includes data defining a respective chemical structure of each new output molecule in a set of new output molecules.|0014] In some implementations, generating a respective new molecule exemplar corresponding to each of one or more output molecules in the set of output molecules generated by the generative machine learning model comprises, for each of one or more output molecules in the set of output molecules, determining a respective predicted molecule property value of the output molecule for each molecule property in the set of molecule properties, and including the predicted molecule property values of the output molecule in the new molecule exemplar corresponding to the output molecule.

[0015] In some implementations, determining the respective predicted molecule property value of the output molecule for each molecule property in the set of molecule properties comprises, for each of one or more molecule properties processing data characterizing the output molecule using a property' prediction machine learning model to generate the predicted molecule property value of the output molecule as an output of the property prediction machine learning model.

[0016] In some implementations, determining the respective predicted molecule property’ value of the output molecule for each molecule property- in the set of molecule properties comprises, for each of one or more molecule properties performing a computational simulation of the output molecule to determine the predicted molecule property value of the output molecule.[0017| In some implementations, for each output molecule in the set of output molecules, the model output of the generative machine learning model defines a respective predicted molecule property value of the output molecule for each of one or more molecule properties in the set of molecule properties, and wherein determining the respective predicted molecule property value of the output molecule for each molecule property in the set of molecule properties comprises, for each of one or more molecule properties,extracting the predicted molecule property value of the output molecule from the model output of the generative machine learning model.(0018| In some implementations, the updated model input to the generative machine learning model comprises a plurality of molecule exemplars that include experimental molecule property7values of example molecules determined by performing phy sical experiments, and a plurality of molecule exemplars that include predicted molecule property values of output molecules determined using computational prediction techniques.

[0019] In some implementations, updating the model input to the generative machine learning model comprises selecting a proper subset of the set of output molecules according to one or more selection criteria based on molecule property values of the output molecules, and wherein generating the respective new molecule exemplar corresponding to each of one or more output molecules in the set of output molecules comprises generating a respective new molecule exemplar for only output molecules in the selected proper subset of the set of output molecules.

[0020] In some implementations, selecting a proper subset of the set of output molecules according to one or more selection criteria based on molecule property values of the output molecules comprises determining, for each output molecule, a respective binding affinity of the output molecule for a target protein, and selecting the proper subset of the set of output molecules based at least in part on their respective binding affinities for the target protein.

[0021] In some implementations, selecting the proper subset of the set of output molecules based at least in part on their respective binding affinities for the target protein comprises determining, for each output molecule, that the output molecule should be selected for inclusion in the proper subset of the set of output molecules only if the binding affinity of the output molecule for the target protein satisfies a threshold.

[0022] In some implementations, outputting data identifying the set of molecules comprises, for one or more of the output molecules, providing data representing the output molecule to a user, receiving, from the user, a textual query comprising a request to modify7the output molecule, processing a model input that comprises: (i) data representing the chemical structure of the output molecule, and (ii) the textual query from the user, using a molecule modification machine learning model and in accordance with values of a set of molecule modification machine learning model parameters, to generate a model output that defines a chemical structure of a modified molecule, wherein themodified molecule is a version of the output molecule that has been modified in accordance with the request included in the textual query from the user, and outputting data identifying the modified molecule.|0023] In some implementations, the molecule modification machine learning model is a generative machine learning model that has been trained to perform a language modeling task.100241 In some implementations, the request to modify the output molecule comprises one or more of a request to modify a scaffold structure of the output molecule, or a request to modify a fragment structure of the output molecule, or a request to modify physiochemical properties of the output molecule, or a request to modify biochemical properties of the output molecule, or a request to add a functional group to the output molecule, or a request to remove a functional group from the output molecule, or a request to replace a functional group in the output molecule, or a request to change a steric hindrance of the output molecule, or a request to change a polarity of the output molecule, or a request to change a hydrophobicity of the output molecule.[0025| In some implementations, the model input to the generative machine learning model is represented as a sequence of input tokens, where each input token in the sequence of tokens is selected from a set of possible tokens.

[0026] In some implementations, for each molecule exemplar, the chemical structure of the corresponding example molecule is represented in the sequence of input tokens by a respective Simplified Molecular Input Line Entry System (SMILES) string.

[0027] In some implementations, the model output generated by the generative machine learning model is represented as a sequence of output tokens where each output token in the sequence of output tokens is selected from a set of possible tokens.

[0028] In some implementations, for each output molecule, the chemical structure of the output molecule is represented in the sequence of output tokens by a respective Simplified Molecular Input Line Entry Sy stem (SMILES) string.

[0029] In some implementations, processing the model input using the generative machine learning model and in accordance with values of the set of generative machine learning model parameters to generate the model output that includes data defining the respective chemical structure of each output molecule in the set of output molecules comprises sequentially generating the sequence of output tokens, one token at a time, starting from a first token in the sequence of output tokens.[00301 In some implementations, for each of one or more positions in the sequence of output tokens, generating the output token at the position in the sequence of output tokens comprises processing data comprising: (i) a sequence of input tokens representing the model input, and (ii) any output token that precedes the position in the sequence of output tokens, using the generative machine learning model to generate a score distribution over the set of possible tokens, and selecting the output token at the position in the sequence of output tokens using the score distribution over the set of possible tokens.[0031 [ In some implementations, selecting the output token at the position in the sequence of output tokens using the score distribution over the set of possible tokens comprises selecting a token having a highest score under the score distribution over the set of possible tokens as the output token at the position in the sequence of output tokens.[0032} In some implementations, selecting the output token at the position in the sequence of output tokens using the score distribution over the set of possible tokens comprises randomly sampling a token from the set of possible tokens, in accordance with the score distribution over the set of possible tokens, to select the output token at the position in the sequence of output tokens.[0033| In some implementations, wherein the generative machine learning model is implemented as a neural network.

[0034] In some implementations, the neural network comprises a plurality of attention layers.

[0035] In some implementations, the generative machine learning model has been trained to perform a language modeling task.[0036J In some implementations, training the generative machine learning model to perform the language modeling task comprises obtaining a set of training examples, wherein each training example comprises at least a sequence of output tokens, and training the generative machine learning model on the set of training examples, comprising, for each training example, training the generative machine learning model to generate, for each position in the sequence of output tokens, the output token occupying the position by processing at least any output tokens that precede the position in the sequence of output tokens.

[0037] In some implementations, the set of molecule properties comprises one or more of: binding affinity for a target protein, molecular weight, topological polar surface area, lipophilicity, synthetic accessibility, solubility, permeability, absorption, distribution, metabolism, excretion, toxicity, or stability.[0038| In some implementations, outputting data identifying the set of output molecules comprises storing data identifying the set of output molecules in a memory, or providing a representation of the set of output molecules for presentation on a display, or transmitting data identifying the set of output molecules over a data communications network.

[0039] In some implementations, further comprising selecting one or more output molecules from the set of output molecules for physical synthesis.100401 In some implementations, selecting one or more output molecules from the set of output molecules for physical synthesis comprises determining a respective binding affinity of each output molecule for a target protein, and selecting one or more output molecules from the set of output molecules for physical synthesis based at least in part on the binding affinities of the output molecules for the target protein.

[0041] In some implementations, further comprising physically synthesizing each of the one or more output molecules that are selected for physical synthesis.

[0042] In some implementations, further comprising, for each of the one or more output molecules that are selected for physical synthesis, performing physical experiments on one or more physically synthesized instances of the output molecule to evaluate one or more molecule properties of the output molecule.

[0043] In another aspect, a system comprising one or more computers and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the methods described herein.[00441 In another aspect, one or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective methods described herein.

[0045] In another aspect, a molecule having a chemical structure of an output molecule that is defined by a model output generated by a generative machine learning model according to the respective methods described herein.[0046[ Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages.

[0047] The system described in this specification can generate data defining valid, novel, and diverse sets of output molecules using a generative machine learning model that processes a model input that includes a set of molecule design parameters specifying desired characteristics of the output molecules. The system can thus be used to generatenovel molecules that possess specific properties tailored to address particular scientific or industrial challenges. For instance, the system can generate data identifying molecules that are predicted to bind to a protein that is a therapeutic target for treating a disease and that has desirable, drug-like properties, e.g., absorption, distribution, metabolism, excretion, and toxicity properties.[0048} The system can provide the generative machine learning model with a model input that includes a set of molecule design parameters (as noted above) and also a set of molecule exemplars. Each molecule exemplar corresponds to a respective example molecule and includes data defining: (i) a chemical structure of the example molecule, and (ii) a respective molecule property value of the example molecule for each molecule property in a set of molecule properties. The generative machine learning model can leverage the set of molecule exemplars included in the model input to implicitly correlate patterns in chemical structures with corresponding molecule properties, and then use these inferred correlations to generate output molecules that satisfy the molecule design parameters.[0049} The generative machine learning model can leverage foundational training, e.g., to perform a language modeling task involving next token prediction, to generate output molecules satisfying the molecule design parameters without necessarily being explicitly trained (or fine-tuned) to perform the task of molecular inverse design. The system can thus use the generative machine learning model to perform the task of molecular inverse design while consuming fewer computational resources (e.g., memory and computing power) compared to alternative approaches that, instead of including the molecule exemplars in the model input, instead require training the generative machine learning model on the molecule exemplars using a machine learning training technique. The generative machine learning model can include large numbers of model parameters (e.g., millions or billions of model parameters), and by reducing or removing the need for training the generative machine learning model on the molecule exemplars, the system can enable dramatically increased computational efficiency and reduction in computational resource consumption.[0050} High quality experimental data characterizing molecule properties is expensive and hard to obtain, e.g., as generating such data requires physically synthesizing molecules and performing w et lab experiments. However, the performance of the generative machine learning model, e.g., in generating molecules that satisfy the set of design parameters, depends critically on the amount and quality of the molecule exemplarsprovided in the model input to the generative machine learning model. To address this issue, the system can iteratively augment the set of molecule exemplars included in the model input to the machine learning model to include new molecule exemplars for molecules previously generated by the generative machine learning model. More specifically, at each of multiple iterations, the system can process a current model input using the generative machine learning model to generate a set of output molecules, determine predicted properties of the output molecules, generate new molecule exemplars that include the output molecules and their predicted properties, update the model input to include the new molecule exemplars, and then provide the updated model input for processing at the next iteration. The system can thus iteratively expand the set of molecule exemplars to include not only "‘experimental” molecule exemplars (without experimentally determined molecule properties) but also ‘‘synthetic” molecule exemplars (corresponding to output molecules generated by the generative machine learning model and with computationally predicted molecule properties).[0051 | At each iteration, the system may select only a proper subset (i.e.. fewer than all) of the output molecules generated at the iteration for use in generating new molecule exemplars. For instance, the system may determine a respective predicted binding affinity of each output molecule generated at the current iteration for a protein target, and then select only a proper subset of the output molecules having the highest predicted binding affinity for the protein target for use in the new set of molecule exemplars, e.g., for a future iteration. By adaptively selecting only certain output molecules for use in generating new molecule exemplars to be added to the model input, the system can facilitate reduced consumption of computational resources (e.g., memory and computing power), as compared to a system that indiscriminately uses every output molecule for generating new molecule exemplars. More specifically, processing a larger model input by the generative machine learning model requires performing more computational operations and therefore consumes more computational resources, and thus by limiting the size of the model input (by adaptively selecting only certain output molecules for inclusion in the model input), the system enables reduced consumption of computational resources.[0052 Additionally, the system can save reduce consumption of time and laboratory resources, e.g., assays and personnel-hours, compared to a standard molecular forward design process. In particular, the standard forward design process of iteratively modifying a known molecular structure, synthesizing the modified molecule, and evaluating thesynthesized molecule in an attempt to refine one or more molecular properties in order to fulfill specific criteria, e.g., affinity for biological targets, favorable pharmacokinetic and pharmacodynamic, non-toxicological profiles, etc., can be prohibitively complex, uncertain, resource-intensive, and time-consuming. Furthermore, the system can provide a tool for iterative modification of one or more of the generated output molecules, thereby removing the need to synthesize many versions of an identified molecule structure to evaluate small changes to the chemical compound.[00531 Additionally, the system can reduce the computational resources necessary for processing data representing an input molecule to generate an output molecule by using data that represents the chemical structure of a molecule in two-dimensions, e.g., relative to using data that represents the chemical structure of a molecule in three-dimensions. In particular, the system can process a model input that includes respective Simplified Molecular-Input Line Entry System (SMILES) strings as the structure of each of the molecules in the set of molecule exemplars using a generative machine learning model to generate a set of output molecules including respective SMILES strings as the structure of each of the generated output molecules. More specifically, by relying on processing and generating two-dimensional data representing the chemical structure of the input and output molecule, the generative machine learning model can require less data throughput, e.g., the rate at which data is transferred or processed within a system or network over a certain period of time, and can generate outputs more efficiently, e.g.. with respect to processing and generating three-dimensional data representing the chemical structure.

[0054] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0055] FIG. 1 is a system diagram of an example molecular inverse design system that includes a generative machine learning model.

[0056] FIG. 2 is a flow' diagram of an example process for generating a set of output molecules using a set of molecule exemplars and a set of molecule design parameters.|0057] FIG. 3 illustrates an example model input for the generative machine learning model.

[0058] FIG. 4 demonstrates example results of the generative machine learning model.

[0059] FIG. 5 is a flow diagram of an example process for updating the set of molecule exemplars.

[0060] FIG. 6 demonstrates the improvement in biological activity by updating the set of molecule exemplars, e.g., using the process described in FIG. 5.

[0061] FIG. 7 is a flow diagram for modifying a molecule based on a textual query.

[0062] FIG. 8 illustrates an example of successively modifying a molecule, e.g., using the process 700 of FIG. 7.

[0063] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION

[0064] FIG. 1 shows an example molecular inverse design system 100. The molecular inverse design system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations in which the systems, components, and techniques described below are implemented.

[0065] The molecular inverse design system 100 includes a generative machine learning model 130 configured to process a model input 105 to generate data defining a set of output molecules 140. In particular, the generative machine learning model 130 can process a set of molecule exemplars 110 and a set of molecule design parameters 115 that specify the desired characteristics of the output molecules to be generated by the machine learning model 130, as will be described in more detail below.

[0066] The system 100 can generate the model input 105 for the generative machine learning model 130. In particular, the system 100 can generate a set of molecule exemplars 110. each exemplar including the chemical structure 112 of an example molecule and a set of molecule properties 114 for the example molecule. In this context, each molecule exemplar in the set 110 is an example that provides contextualizing information for a particular molecule. More specifically, each molecule exemplar 110 is a representative example that demonstrates typical structural characteristics or qualities correlated with the set of molecular properties.

[0067] For example, the chemical structure 112 can be a textual representation of the chemical structure of the molecule, e.g., a SMILES (Simplified Molecular-Input Line- Entry System) string, InChi (International Chemical Identifier), an MDL (Molecular Design Limited) file. etc.[0068| As an example, the set of molecule properties 114 can include one or more of a binding affinity for a target protein, molecular weight, partition coefficient, or topological polar surface area (TPSA). As a further example, the set of molecule properties 114 can include one or more of lipophilicity, synthetic accessibility (SA), solubility, or permeability7. In some cases, the system can generate the synthesizability7accessibility score, e.g., as described in Ertl, P. and Schuffenhauer, A. “Estimation of synthetic accessibility score of drugdike molecules based on molecular complexity and fragment contributions.” (J Cheminform 1, 8 (2009) https: / / doi.org / 10.1186 / 1758-2946-l-8).[0069| As yet a further example, the set of molecule properties 114 can include one or more of ADMET (absorption, distribution, metabolism, excretion, toxicity), biological activity (e.g., binding affinity) for a particular target protein (e.g., C50, Ki, or PPAR5 with EC50 data), or stability values. In some cases, the system can generate the biological activity data using a Categorical Boosted regression model. As another example, the set of molecule properties 114 can include a potency measure, e.g., a pCHEMBL measure.[0070} In some cases, one or more experimental molecule property values included in the set of molecule properties 114 can be determined by performing physical experiments on the corresponding example molecule. In other cases, one or more predicted molecule property7values can be predicted, e.g., by performing a computational prediction technique. For example, the predicted molecule property values can be predicted using a property prediction machine learning model 160, or by performing a computational simulation, e g., a quantum mechanical simulation (e g., using ab initio methods or density functional theory ), or a molecular dynamics simulation (e.g., involving using force fields to model molecular interactions and predict the dynamic evolution of molecules over time).

[0071] In particular, the set of molecule exemplars 110 can include N exemplars. As an example, the model input 105 can be a many-shot model input. In this case, N can be a large value, e.g., 200, 500, 2000, etc. As another example, the model input 105 can be a few-shot model input. In this case, N can be 1, 5, or 20 exemplars. In some cases, the ability of the generative machine learning 130 model to generate valid, novel, and diverse molecules with target properties increases with the number of exemplars included in the set of molecule exemplars. In some cases, N can be restricted based on the type of generative machine learning model 130 implemented.

[0072] For example, the system 100 can determine the exemplars included in the set of molecule exemplars 105 using a set of biological activity, e.g., binding affinity, data, e.g..from an obtained biological activity, e.g., binding affinity, dataset, against a set of protein targets. In particular, the system 100 can generate clusters of molecules from the obtained dataset using a similarity measure, e.g., a structural similarity measure based on the chemical structures of each molecule, and can rank the generated clusters from highest to lowed based on the included biological activity e.g., binding affinity7. The system 100 can then select one or more clusters from each dataset for inclusion in the set of molecule exemplars.[0073 J In some cases, the system 100 can additionally perform a property profiling of each cluster, e.g., by evaluating one or more of physiochemical properties, substructure properties, a synthesizability score, etc. to estimate synthetic accessibility or ligand efficiency. In this case, the system 100 can use the synthetic accessibility or ligand efficiency to select a subset of molecules from the dataset that contain chemically attractive and easy to synthesize molecules for inclusion in the set of molecule exemplars 110.[0074| The system 100 can also obtain, e.g., generate, receive, or both, a set of molecule design parameters 115 that specify one or more desired characteristics of the set of output molecules 140 to be generated by the generative machine learning model 130. For example, the system 100 can generate one or more of the molecule design parameters, receive one or more of the molecule design parameters, or a combination of both. In particular, the set of molecule design parameters 115 can specify desired ranges or values for one or more molecule properties in the set of molecule properties 114. For example, the set of molecule design parameters 115 can include a specification that the molecular activity, e.g., binding affinity, for a target protein is greater than a certain threshold, specifying the synthesizability accessibility (SA) score is less than a certain threshold, and / or specifying a target range for the molecular weight of the set of output molecules 140.

[0075] In some cases, the model input 105 can additionally include an input molecule 120, e.g., a lead molecule to serve as a starting point for molecule generation. As an example, the system 100 can select a lead molecule for a particular biological target, e.g., that corresponds with a particular disease based on activity (binding) against a protein target associated with the disease. In the case that the system 100 performs a clustering analysis on an obtained dataset to determine the set of molecule exemplars 110, the system 100 can select the input molecule 120 based on the ranking of the clusters using the biological activity data.[00761 The system 100 can process the model input 105 using the generative machine learning model 130 to generate the set of output molecules 140.

[0077] The generative machine learning model can be any appropriate model that can process a model input (e.g., that includes a set of molecule exemplars and a set of molecule design parameters) to generate one or more samples from a distribution over a space of possible model outputs. The space of possible model outputs can be a space of possible output molecules (and, optionally, predicted properties associated with the output molecules). The space of possible model outputs can be parametrized, e.g., as a space of sequences of tokens, where the tokens can include, e.g., alphanumerical characters.

[0078] The generative machine learning model 130 can have any appropriate machine learning architecture, e.g., a neural network architecture, that can be configured to process a model input 105 that includes the set of molecule exemplars 110, the set of molecule design parameters 115, and, in some cases, the input molecule 120, to generate the set of output molecules 140. In particular, the generative machine learning model 130 can have any appropriate number of neural network layers (e.g., 10 layers, 50 layers, or 100 layers) of any appropriate type (e.g.. fully-connected layers, attention layers, convolutional layers, etc.) connected in any appropriate configuration (e.g., as a linear sequence of layers, or as a directed graph of layers).

[0079] As a particular example, the generative machine learning model 130 can be a language processing model, e.g.. a foundation model such as a transformer large language model (LLM) that has been trained to perform a language modeling task. In this case, the system 100 can formulate the model input 105 as part of an in-context learning process. More specifically, in-context learning can allow a large language model to learn from a specific task at the time of inference, e.g., by leveraging a set of examples in place of finetuning the foundation model. This technique involves presenting a series of inputoutput examples, known as shots, that illustrate or demonstrate typical characteristics or qualities of a particular category, concept, or idea.

[0080] In particular, the model input 105 can be formulated as an input for a large language model (i.e., a generative machine learning model that has been trained to perform a language modeling task). An example prompt for a generative machine learning model implemented as a large language model will be described in more detail with respect to FIG. 2. In this context, each molecule exemplar in the set of molecule exemplars 110 is a representative example, e.g.. a representative example that illustrates typical characteristics of a molecule. The system 100 can supply many examples to thegenerative machine learning model 130 in the model input 105 to provide the context for generating the set of output molecules 140.[0081 | As a further example, the generative machine learning model 130 can be a generative adversarial network or a diffusion network. As another example, the generative machine learning model 130 can have a recurrent neural network architecture that is configured to sequentially process the contents of the model input 105 and trained to perform next element prediction, e.g., to define a likelihood score distribution over a set of next elements. More specifically, the generative machine learning model 130 can be a recurrent neural network (RNN), long short-term memory' (LSTM), or gated-recurrent unit (GRU). As another example, the generative machine learning model 130 can be an encoder-decoder or decoder-only transformer configured to perform parallel processing of the contents of the multimodal input using a multi-headed attention mechanism.

[0082] In some cases, each output molecule in the set of output molecules 140 can have a chemical structure that is a modified version of an input molecule, e.g., the input molecule 120. In particular, in the case that the model input 105 includes an input molecule 120 as a lead molecule, the system 100 can generate the set of output molecules 140 by modifying the input molecule 120 in a number of ways.|0083[ In the particular example depicted, the system 100 can use active learning to improve the outputs generated by the generative machine learning model 130. More specifically, the system 100 can update the model input 105, e.g.. by updating the set of molecule exemplars 110 using one or more of the set of output molecules 140. In particular, the system 100 can select a subset 142, e.g., a proper subset, of the set of output molecules 140 to be included in the updated set of molecule exemplars 110 and can add each new molecule exemplar to the set of molecule exemplars 110. e.g., for future processing using the generative machine learning model 130.

[0084] For example, the system 100 can determine the corresponding set of molecule properties for each of the set of output molecules 140 and can select the subset 142 based on one or more of the determined molecule properties. In particular, the system 100 can process the output set of molecules 140 using a property’ prediction machine learning model 150 to determine the set of molecule properties 114. As another example, the system 100 can determine one or more of the molecule properties 114 by performing physical experiments, or running a computational simulation, etc.

[0085] The property prediction machine learning model 150 can be implemented as any appropriate type of machine learning model (e.g., including one or more of: a neuralnetwork, or a random forest, or a decision tree, or a support vector machine, or a linear regression model) that can be configured to process data representing a molecule to generate one or more predicted molecule properties of the molecule. In particular, when implemented as a neural network model, the property prediction machine learning model 150 can include any appropriate number of neural network layers (e.g., 1 layer, 5 layers, or 10 layers) of any appropriate types (e g., fully -connected layers, attention layers, convolutional layers, etc.) connected in any appropriate configuration (e.g., as a linear sequence of layers, or as a directed graph of layers).

[0086] The system 100 can train the property prediction machine learning model 150 on a set of training examples, where each training example corresponds to a respective training molecule and includes: (i) a training input that includes data characterizing the chemical structure of training molecule, and (ii) a target output that defines one or more target molecule properties of the training molecule, e.g., that corresponds with the set of molecule properties 114. The training input can represent the training molecule in any appropriate way, e.g., as a linear sequence of tokens (e.g., a SMILES string) or as a graph of nodes and edges, where each node represents a respective atom and each edge represents a bond between a pair of atoms or that a pair of atoms are separated by less than a threshold distance. The target molecule properties can be determined, e.g., by physical experiments on physically synthesized instances of the training molecule, or by computational simulations (e.g., quantum mechanical simulations or molecular dynamics simulations), or in any other appropriate way.

[0087] The system 100 can train the property prediction machine learning model 150 on the set of training examples by a machine learning training technique to optimize an objective function. The objective function can measure, for each training example, a discrepancy between: (i) the target molecule properties specified by the training example, and (ii) predicted molecule properties generated by the property prediction machine learning model 150 by processing the training input of the training example. The objective function can measure a discrepancy between target and predicted molecule properties in any appropriate way, e.g., using a cross-entropy loss or a mean squared error loss. The machine learning training technique can be any technique appropriate for training the property prediction machine learning model 150. For instance, for a property prediction machine learning model 150 implemented as a neural network, the machine learning training technique can be a stochastic gradient descent training technique.

[0088] The system 100 can select the selected subset 142 of output molecules based on one or more selection criteria using the determined molecule property values of the set of output molecules 140. In particular, the system 100 can select the top M% of molecules with the highest biological activity, e.g., binding affinity, for a target protein from the set of output molecules 140 and can remove the bottom M% of molecules with the lowest biological activity, e.g., binding affinity, for the target protein from the most recently used set of molecule exemplars 110. For example, M can be 5% 10%, or 25%, or any other appropriate percentage.

[0089] For instance, the system 100 can determine a respective binding affinity with respect to a target protein for each of the molecules in the set of output molecules 140. In particular, the system 100 can run a computational simulation to dock the set of output molecules 110 with the target protein. In this case, the system 100 can select a proper subset of the set of output molecules 140 as the selected subset 142 based at least in part on the respective binding affinities for the target protein, e.g., based on the binding affinity for each of the molecules in the selected subset 142 satisfying a threshold value.

[0090] More specifically, the system 100 can update the molecule exemplars 110 at each of a number of iterations, e.g., wherein each iteration is a processing of the model input 105 using the generative machine learning model 130, as part of an active learning process to improve the generated set of output molecules 140. For example, the system 100 can incorporate a selected subset 142 of the set of output molecules generated at each iteration of processing. As another example, the system 100 can incorporate the selected subset 142 every Y iterations, where Y is a positive integer value greater than one.

[0091] As yet another example, the system 100 can incorporate the selected subset 142 when an aggregated, e.g., mean, median, weighted average, value of one or more of the set of molecule properties 114 in the molecule exemplars 110 satisfies a criteria. In particular, the system 100 can incorporate the selected subset 142 when an aggregated value of the biological activity, e.g., binding affinity for a target protein, of the set of molecule properties 114 is less than a set threshold value.[0092| The system 100 can output the data identifying the set of output molecules 140, e.g., to a user, e.g., on the graphical user interface (GUI) of a user-device. In some cases, in response to outputting the data, the system 100 can receive a textual query 146 from the user that specifies a modification to a selected molecule 144 from the set of output molecules 140. As an example, the textual query’ 146 can specify a modification to the chemical structure of the selected molecule 144. As another example, the textual query146 can specify a modification to a physiochemical or biophysical property of the selected molecule 144.

[0093] In this case, the system 100 can process data representing the chemical structure of the selected molecule 144, e.g., a SMILES string representing the selected molecule 144, and the textual query 146 to modify7the selected molecule 144 according to the textual query 146. In the particular example depicted, the system 100 can process a model input that includes the data representing the chemical structure of the selected molecule 144 and the textual uery 146 using a molecule modification machine learning model 160 to generate a modified molecule output 175.[0094| The molecule modification machine learning model 160 can be any appropriate machine learning model, e.g., a neural network model, that can be configured to process the data representing the chemical structure of the selected molecule 144 and the textual query 146. In particular, when implemented as a neural network model, the molecule modification machine learning model 160 can include any appropriate number of neural network layers (e.g., 1 layer, 5 layers, or 10 layers) of any appropriate type (e.g., fully- connected layers, attention layers, convolutional layers, etc.) connected in any appropriate configuration (e.g., as a linear sequence of layers, or as a directed graph of layers).

[0095] As an example, the system 100 can process the selected molecule 144 and a textual query 146 with a corresponding instruction to add, remove, or replace a functional group of the molecule with a different functional group and generate a modified molecule output 175 with the added, removed, or replaced functional group, respectively. As another example, the system 100 can process the selected molecule 144 and a textual query7146 with a corresponding instruction to change the steric hindrance, the polarity, or the hydrophobicity of the selected molecule 144 and can generate a modified molecule output 175 with the modified steric hindrance, polarity, or hydrophobicity, respectively.

[0096] FIG. 2 is a flow diagram of an example process for generating a set of output molecules using a set of molecule exemplars. For convenience, the process 200 will be described as being performed by a system of one or more computers located in one or more locations. For example, a molecular inverse design system, e.g., the molecular inverse design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 200.|0097] The system can generate a set of molecule exemplars including the chemical structure and respective property values for a set of molecule properties (step 210). and the system can obtain, e.g., generate, receive, both, a set of molecule design parametersspecifying desired characteristics of output molecules to be generated (step 220). In some cases, the system can receive data defining the set of molecule exemplars and the set of molecule design parameters from a user, e.g., by way of a user interface (e.g., a graphical user interface (GUI)) or an application programming interface (API) made available by the system.[0098} For example, the chemical structure of the corresponding example molecule in each molecule exemplar can be represented by a SMILES string. As an example, the set of molecule properties can include one or more of binding affinity for a target protein, molecular weight, topological polar surface area values. As a further example, the set of molecule properties can include one or more of lipophilicity, synthetic accessibility, solubility, and permeability values. As yet a further example, the set of molecule properties can include one or more of absorption, distribution, metabolism, excretion, toxicity7, or stability7values.[0099} In particular, the set of molecule design parameters can specify a target range of values for each molecule property in the set of molecule properties. For example, one or more of the molecule property values included in the molecule exemplars can be experimental molecule property values determined by7physical experiments on instances of the corresponding example molecule. As another example, one or more of the molecule property values included in the molecule exemplars can be computational molecule property values determined by computational prediction techniques.[0100} The system can process the model input including the set of molecule exemplars and the set of molecule design parameters using a generative machine learning model to generate a model output including a set of output molecules (step 230). For example, the chemical structure of the corresponding example molecule in each molecule exemplar can be represented by a SMILES string. In some cases, the model input to the generative machine learning model can include one or more input (lead) molecules, e.g., molecules that exhibit promise for treatment of a particular disease and serve as the starting point for further research and development. In this case, each output molecule in the set of output molecules can have a chemical structure that is a modified version of the chemical structure of an input (lead) molecule.[01011 In some cases, the generative machine learning model can be implemented as an autoregressive language processing network that is configured to sequentially generate each token in a sequence of tokens defining the model output and that is trained to perform next token prediction. In particular, the system can process a sequence of inputtokens selected from a set of possible tokens to generate a sequence of output tokens, e.g., selected from the set of possible tokens. For example, the generative machine learning model can be implemented as a neural network that includes a number of attention layers.|0102] The set of possible tokens can include, e.g., alphabetical characters, numerical characters, special characters, or any other characters that, when combined in a sequence, can be used to represent the chemical structure of a molecule (e.g., as a SMILES string) and molecule properties of the molecule (e.g., binding affinity of the molecule for a target protein, molecular weight, topological polar surface area, and so forth).[0.103] In this case, the system can sequentially generate the sequence of output tokens, one token at a time, starting from a first token in the sequence of output tokens. More specifically, for each position in the sequence of output tokens, the system can process the sequence of input tokens in the model input and any output token that precedes the position in the sequence of output tokens using the generative machine learning model to generate a score distribution over the set of possible tokens. The system can then select the token to occupy the position in the sequence of output tokens using the score distribution of the set of possible tokens. As an example, the system can select the output token having a highest score under the score distribution over the set of possible tokens as the output token at each position in the sequence of output tokens. As another example, the system can randomly sample a token from the set of possible tokens, e.g.. in accordance with the score distribution over the set of possible tokens, to select the output token at the position in the sequence of output tokens.

[0104] For example, the generative machine learning model can be an autoregressive token generation model that has been trained to perform a language modeling task, e.g., a large language model. In this case, training the generative machine learning model can involve obtaining a set of training examples that include at least a sequence of output tokens and training the generative machine learning model to generate, for each position in the sequence of output tokens, the output token occupying the position by processing at least any output tokens that precede the position in the sequence of output tokens.[0105| The system can output data identifying the set of output molecules (step 240). For example, the system can output data identifying the set of output molecules to a user, e.g., by way of a user interface (e.g., a graphical user interface (GUI)) or an application programming interface (API) made available by the system. As another example, the system can provide the set of output molecules for storage in a memory or for transmission over a data communications network.[0.106| In some cases, the output data identifying the set of output molecules can be used for one or more downstream tasks. For example, the system can determine a respective binding affinity of each output molecule for a target protein (e.g., using a property prediction machine learning model, or a computational simulation, or both), and can select one or more output molecules from the set of output molecules for phy sical synthesis based at least in part on the binding affinities of the output molecules for the target protein. The selected output molecules can be physically synthesized and physical experiments can be performed on the physically synthesized instances of each output molecule to evaluate one or more molecule properties of the output molecule, e.g., binding affinity for a target protein, absorption, distribution, metabolism, excretion, toxicity, and so forth.[0107} As another example, the system can identify a subset of the set of output molecules and update the set of molecule exemplars for a future iteration of the process 200. An example process for updating the set of molecule exemplars will be covered in more detail in FIG. 5. After generating an updated set of molecule exemplars, e.g., by generating a respective new molecule exemplar corresponding to some or all of the output molecules generated by the generative machine learning model and adding the new molecule exemplars to the model input to the generative machine learning model, the system can return to step 230.[0108| As yet another example, a user can view the outputted data and select a molecule from the set of output molecules for further modification, e.g., with a molecule modification machine learning model, as will be described in further detail with respect to FIG. 7.[0109[ FIG. 3 illustrates an example model input for the generative machine learning model. In particular, the input 300 is formatted in the style of a prompt, e.g., a directive instruction from a user, e.g., a question, statement, code snippet, or example. In this case, the input 300 includes a general task 310 specify ing that the generative machine learning model generate valid SMILES strings for novel and structurally diverse molecules based on a set of molecule exemplars 320, a set of design parameters 320, a lead molecule 340. and an output format 350.[0110| In particular, each molecule exemplar in the set of molecule exemplars 320 includes a SMILES string that represents the chemical structure of the molecule, and a corresponding set of molecule properties. In this case, the set of molecule properties includes a biological activity, a molecular weight, a synthesizability accessibility score,and a topological surface area. In the particular example depicted, the set of molecule exemplars 320 is formatted as a list of molecule exemplars. As another example, the set of molecule exemplars 320 can be formatted as an array, comma-separated value file, etc.10111] In this case the set of design parameters 335 are represented as part of the prompt.In particular, the set of design parameters 335 are included as candidate requirements 330 of the general task 210. In other cases, the set of design parameters 335 can be included as an array with a configured order, e.g., to specify the ranges based on the relative positioning of the range in the array, or as a dictionary of key -value pairs, e.g., where the molecule property is a key and the value is a specified range for the particular value.[0112 | In the particular example depicted, the task can be to modify the set of molecule exemplars 320 to generate molecules with the specified set of design parameters 335. For example, the set of design parameters 335 can include a list specifying the preferred ranges or values for high activity, e.g., binding affinity, molecular weight, synthesizability, topological polar surface area, and lipophilicity. In some cases, the ordering of the list can specify a relative importance, e.g., user preference, for the set of design parameters 335, e.g., the design parameters specified near the start of the list can be more important specifications than the design parameters specified near the end of the list. In other cases, the ordering of the list does not reflect any preferences for the design parameters 335.

[0113] In this case, the prompt 300 also includes a lead (input) molecule 340. In some cases, the model input can include an input molecule as a seed molecule, e.g., a starting point for modification of the set of molecule exemplars 320. In particular, the generative machine learning model can use the molecule exemplars to understand the underlying distribution of molecule space and can modify the lead molecule 340 in a number of ways, e.g., as represented by the distribution of allowable molecule configurations. In this case, each output of the set of output molecules is a modified version of the chemical structure of the lead molecule 340.

[0114] As an example, the prompt 300 can also include a specified output format 350, e.g., to provide a desired format for the output of the generative machine learning model. In the particular example depicted, the specified output format 350 is a JSON object that includes a list of the output molecules with keys for the “SMILES” string and “predicted activity keys”, e.g., as depicted in the output 360 which contains a list of dictionaries for each output molecule in the set of output molecules, with corresponding SMILES data representing the representing the chemical structure of the output molecule and thepredicted activity value of the output molecule, as specified by the keys. As an example, the system can use a property prediction machine learning model, e.g., the propertyprediction machine learning model 150 of FIG. 1, to generate the corresponding predicted activity values for each of the output molecules.[01151 FIG. 4 demonstrates example results of improving a lead molecule for three lead molecules. For example, a model input fonnatted similar to the input 300 of FIG. 3 can be processed by the generative machine learning model for the three lead molecules 410, 440, and 470, respectively to generate a set of output molecules with improved characteristics, e.g., based on improvements made to one or more of the molecule properties in the set of molecule properties, as indicated by the set of design parameters.| 116] In particular, panels 400. 430, and 460 depict respective example output molecules generated by modifying the lead molecules 410, 440, and 470, respectively, e.g., using separate model inputs. The generated molecules 420, 450, and 480 represent improvements to their respective lead molecules, e.g., as represented by the bolded molecule property values associated with the modified compounds. In this case, not all of the property value changes can be considered improvements, e.g., some of the values are not bolded. More specifically, each output molecule in the generated set of output molecules can satisfy any or all of the molecule design criteria specified by the set of molecule design parameters. In particular, the system can leverage molecule propertyvalues not specified as part of the molecule design parameters as available degrees of freedom, e.g., for exploiting potential variability- to satisfy- any or all of the molecule design criteria.[0117| As an example, panel 400 demonstrates how the generative machine learning model is able to modify the lead molecule 410 with a corresponding set of molecule properties 412 to generate the output molecule 420 with improved molecule properties 422, e.g., that better satisfy the set of molecule design criteria specified in the model input to the generative machine learning model, e.g., the activity- increased to 8.2, the molecular weight decreased to 397.5, and the topological surface area decreased to 108.0. As another example, panel 430 demonstrates how the generative machine learning model is able to modify the lead compound 440 with a corresponding set of molecule properties 442 to generate the output molecule 450 with improved molecule properties 452, e.g., the activity increased to 8.4, the molecular weight decreased to 538.5, and the topological polar surface area decreased to 113.8. As yet another example, panel 460 demonstrates how the generative machine learning model is able to modify the lead compound 470with a corresponding set of molecule properties 472 to generate the output molecule 480 with improved molecule properties 482, e.g., the activity increased to 8.8 and the molecular weight decreased to 412.4.|0118] FIG. 5 is a flow diagram of an example process for updating the set of molecule exemplars included in the model input to the generative machine learning model. For convenience, the process 500 will be described as being performed by a system of one or more computers located in one or more locations. For example, a molecular inverse design system, e.g., the molecular inverse design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 500.

[0011] The system can determine a predicted set of molecule property values for a set of generated output molecules (step 510). In particular, the system can process a current set of molecule exemplars and a set of molecule design parameters using the generative machine learning model to generate a set of output molecules as described in process 200 in FIG. 2. The system can then determine a respective predicted molecule property value for each molecule property in the set of molecule properties for each molecule in the set of output molecules.

[0120] For example, the system can process data characterizing the output molecule using a property prediction machine learning model to generate one or more predicted molecule property values of the output molecule as an output of the property prediction machine learning model. As another example, the system can perform a computational simulation of the output molecule to determine one or more predicted molecule property values. As yet another example, the system can extract the predicted molecule property7values of the output molecule from the model output of the generative machine learning model. As a further example, the system can determine one or more experimental molecule property values by performing physical experiments.

[0121] The system can select one or more output molecules generated by the generative machine learning model as new molecule exemplars (step 520). In particular, the system can select a proper subset of the set of output molecules according to one or more selection criteria based on the determined molecule property values of the set of output molecules. The determined values of the set of output molecules can be included in the new molecule exemplar, e.g., if selected, the system can incorporate the determined set of molecule property values for the example molecule in the new molecule exemplar. More specifically, the system can generate a respective new molecule exemplar for only output molecules in the selected subset of the set of output molecules.[0.1221 For example, the system can evaluate whether one or more of the determined molecule property values satisfy respective thresholds. As another example, the system can determine a respective binding affinity of the output molecule for a target protein, and can select the proper subset of the set of output molecules based at least in part on their respective binding affinities for the target protein. In some cases, the system can determine that an output molecule should be selected for inclusion in the proper subset of the set of output molecules if the binding affinity of the output molecule for the target protein satisfies a threshold.[0.123] The system can then update the model input to include the new molecule exemplars (step 530), and the system can process the updated model input using the generative machine learning model to generate an updated model output (step 540). In particular, the system can process the model input using the process 200 described in FIG. 2 to generate an updated model output. More specifically, the system can process the updated model input including the new molecule exemplars corresponding to output molecules generated by the generative machine learning model using the generative machine learning model to generate additional output molecules.[0124| For example, the system can repeat the process 500, e.g., by selecting one or more of the output molecules from the set of output molecules as new molecule exemplars. In this case, the steps 510-530 can be repeated at each of a number of iterations, e g., as many times as desired by a user, to update the set of molecule exemplars. In particular, the system can continue iteratively generating new output molecules and updating the model input to the generative machine learning model until a termination criterion is satisfied. The termination criterion can be, e.g., that a predefined number of iterations have been perfonned, or that a predefined number of output molecules satisfying the set of molecule design parameters have been generated. As an example, the system can iteratively include generated molecules with high predicted biological activity as new molecule exemplars, e.g., in order to improve the output of the generative machine learning model.[0125| FIG. 6 demonstrates the improvement in one molecular property, the biological activity, by updating the set of molecule exemplars, e.g., using the process described in FIG. 5. In this case, the biological activity is the binding affinity for a target protein. In particular, FIG. 6 demonstrates the distribution of molecular activities, e.g., binding affinity, against a particular protein target, e.g., the MMP8 protein target.[0.126] Graph 600 is a box-and-whisker plot that demonstrates the biological activity for lead molecules as compared to the subset of datasets that include an additional 25 molecule exemplars selected from the previous iteration’s generated set of output molecules, e.g., the number of molecule exemplars grew by 25 shots of synthetic data in each iteration. More specifically, in each iteration, in addition to 500 shots of experimental data, e.g., as the baseline in iteration 0, the system incorporated 25 of the generated output molecules with the highest predicted activities from the previous iteration’s set of output molecules.[0.127] In particular, graph 600 represents how including the 25 additional generated output molecules with the highest biological activity in the set of molecule exemplars for the next iteration increases the biological activity of the set of output molecules generated. More specifically, graph 600 demonstrates a trend of improvement in biological activity based on the increasing median of the boxes with increasing iterations of updating the set of molecule exemplars using the 25 additional generate output molecules. For example, there is a significant shift in the distribution of biological activities of generated molecules toward the high activity region within a few iterations of updating the set of molecule exemplars with synthetic high-activity molecules, thereby representing how active learning can enhance the capability of the generative machine learning model to generate highly active molecules, e.g., as part of molecular inverse design.

[0128] FIG. 7 is a flow diagram for modifying a molecule based on a textual query. For convenience, the process 700 will be described as being performed by a system of one or more computers located in one or more locations. For example, a molecular inverse design system, e.g., the molecular inverse design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 700.10129] The system can provide data representing the output molecule to a user (step 710). As an example, the system can provide data representing one or more of the output molecules generated by the generative machine learning model to a user, e.g., on the graphical user interface (GUI) of a user-device. In particular, the user-device can render the data representing the output molecule as a graphical representation on the GUI using a rendering engine.

[0130] The system can receive a textual query, e.g., a prompt, from the user including a request to modify the output molecule (step 720). For example, the system can receive a textual query defining a request to modify a scaffold structure of the output molecule or to modify' a fragment structure of the output molecule. As another example, the systemcan receive a textual query defining a request to modify physiochemical properties of the output molecule or to modify biochemical properties of the output molecule. In some cases, the system can receive a textual query defining a request to add a functional group to the output molecule, to remove a functional group from the output molecule, or to replace a functional group in the output molecule. In other cases, the system can receive a request to change a steric hindrance of the output molecule, a request to change a polarity of the output molecule, or a request to change a hydrophobicity of the output molecule.[01311 The system can then process data representing the output molecule chemical structure and the textual query7using a molecule modification machine learning model to generate a modified output (step 730). In particular, the system can process the model input using a generative machine learning model that has been trained to perform a language modeling task. In this case, the generative machine learning model can be a large language model configured to process and understand the semantic meaning of the textual query, e.g., prompt, to modify the output molecule in order to generate a new output molecule with the specified modification provided by the prompt.[0132| In some cases, the molecule modification machine learning model can be the same machine learning model as the generative machine learning model that generates the set of output molecules. In particular, in some implementations, the molecule modification machine learning model can be implemented as an autoregressive generative machine learning model that has been trained to perform a next token prediction task, as described above with reference to the generative machine learning model of FIG. 1 .[0133| The system can output data identifying the modified molecule (step 740), e.g., to the user, e.g., on the graphical user interface (GUI) of a user-device. For example, the userdevice can render the data representing the modified molecule as a graphical representation on the GUI using a rendering engine. In some cases, the process 700 can be repeated, e.g., a user can specify successive textual queries defining a request for a modification of a previously generated output molecule. An example for successively updating an output molecule using successive textual queries is depicted in FIG. 8.[0134| FIG. 8 illustrates how a user can receive the output data identifying a molecule or modified molecule to identify and specify successive changes, e.g., via an additional textual query7. For example, the system can provide data representing an output molecule, e.g., the molecule 805, from the set of output molecules generated by the generative machine learning model. In particular, the system can provide the data representing the molecule 805 to a user-device, e.g., for rendering on a graphical user interface.[0.135| While the generative machine learning model can generate novel candidate molecules with multiple improved target molecule properties, in some cases, one or more small modifications can further improve one or more of the generated set of output molecules. As an example, the system can output a SMILES string representing the chemical structure of the output molecule. In such cases, direct modification of the SMILES string, especially for large complex molecules, can be challenging. The system can provide an interactive design tool for seamless modification of generated molecules using textual queries to facilitate the direct modification of the molecule in a user-friendly way.|O I36| In particular, a user, e.g., the user 810 can provide a textual query specifying a desired modification to the output molecule 805. In this case, the system can receive a first textual query 815 defining a request to replace the hydroxyl group connected to NH with a methyl group. The system can process the data defining the molecule structure 805 and the first textual query' 815 to generate the first modified output molecule 820, e.g., using a molecule modification machine learning model.[0137| The system can then provide the first modified output molecule 820 to the user 810. In the particular example depicted, after receiving the molecule 820, the user 810 can specify a second textual query' 825 defining a request to deprotect the aniline. The system can process the data defining the molecule structure 820 and the second textual query 825 to generate the second modified output molecule 830.

[0138] Likewise, the system can provide the second modified output molecule 830 to the user 810. After receiving the molecule 830, the user 810 can specify7a third textual query' 835 defining a request to replace the sulfonamide group by a disulfide 835. The system can then process the data defining the molecule 830 and the third textual query 835 to generate the third modified output molecule 840.[01391 Thus, the system can enable a user to perform successive modifications to a generated output molecule, e.g., the molecules 805, 820, and 830, in order to improve the output molecule generated by the generative machine learning model. In particular, the system can provide a molecule modification machine learning model to facilitate a user’s ability to make small, targeted changes to the generated molecules. In particular, in the case of the generative machine learning model providing a SMILES string as output, the system can provide a convenient way to modify the SMILES string, e.g., as opposed to directly modifying the SMILES string which requires extensive understanding ofstructural chemistry and SMILES notations. In particular, modifying SMILES strings can be prohibitive for larger molecules with longer SMILES strings.(0140] In some cases, the system can use one or more of the textual queries defining the request for modification as domain expert feedback. In this context, domain expert feedback refers to evaluations provided by users with specialized chemistry domain knowledge that can be used to refine the generative machine learning model's understanding and generation of output molecules in the domain of chemistry. For example, the system can further include the feedback from domain experts, e g., the textual queries pertaining to a generated output molecule in the set of molecule exemplars, to improve the molecule exemplars for subsequent generation iterations.(0141] This specification uses the term “configured" in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions.(0142] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially-generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.

[0143] The term “data processing apparatus" refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, includingby way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.[0i44| A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any fonn of programming language, including compiled or interpreted languages, or declarative or procedural languages: and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub-programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.

[0145] In this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.

[0146] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.

[0147] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of centralprocessing unit. Generally, a central processing unit will receive instructions and data from a read-only memory or a random access memory' or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory^ devices for storing instructions and data. The central processing unit and the memory' can be supplemented by, or incorporated in, special purpose logic circuitry . Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g.. a universal serial bus (USB) flash drive, to name just a few.

[0148] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory , media and memory' devices, including by way of example semiconductor memory devices, e.g.. EPROM. EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.|<H 49] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g.. visual feedback, auditory’ feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user's device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.

[0150] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing commonand compute-intensive parts of machine learning training or production, i.e., inference, workloads.(01511 Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework, or a Jax framework.

[0152] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.[01531 The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.[0154| While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.[0.155| Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.1 156] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.[01571 What is claimed is:

Claims

Attorney Docket No. 46567-1587WO1CLAIMS1. A method performed by one or more computers, the method comprising: generating a model input to a generative machine learning model, wherein the model input includes data defining: a set of molecule exemplars, wherein each molecule exemplar corresponds to a respective example molecule and includes: (i) a chemical structure of the example molecule, and (ii) a respective molecule property value of the example molecule for each molecule property in a set of molecule properties; and a set of molecule design parameters that specify desired characteristics of output molecules to be generated by the generative machine learning model; processing the model input using the generative machine learning model and in accordance with values of a set of generative machine learning model parameters to generate a model output that includes data defining a respective chemical structure of each output molecule in a set of output molecules; and outputting data identify ing the set of output molecules.

2. The method of claim 1, wherein the model input to the generative machine learning model further comprises a chemical structure of an input molecule; and wherein each output molecule has a chemical structure that is a modified version of the chemical structure of the input molecule.

3. The method of any preceding claim, wherein for each of one or more molecule properties in the set of molecule properties, the set of molecule design parameters specify a target range of values of the molecule property7.

4. The method of any preceding claim, wherein for each of a plurality of molecule exemplars in the set of molecule exemplars, the molecule property values included in the molecule exemplar are experimental molecule property values determined by physical experiments on instances of the corresponding example molecule.

5. The method of any preceding claim, further comprising: updating the model input to the generative machine learning model, comprising: generating a respective new molecule exemplar corresponding to each of one orAttorney Docket No. 46567-1587WO1 more output molecules in the set of output molecules generated by the generative machine learning model; and adding each new molecule exemplar to the set of molecule exemplars included in the model input to the generative machine learning model; and processing the updated model input using the generative machine learning model and in accordance with values of the set of generative machine learning model parameters to generate an updated model output that includes data defining a respective chemical structure of each new output molecule in a set of new output molecules.

6. The method of claim 5, wherein generating a respective new molecule exemplar corresponding to each of one or more output molecules in the set of output molecules generated by the generative machine learning model comprises, for each of one or more output molecules in the set of output molecules: determining a respective predicted molecule property value of the output molecule for each molecule property’ in the set of molecule properties; and including the predicted molecule property values of the output molecule in the new molecule exemplar corresponding to the output molecule.

7. The method of claim 6, wherein determining the respective predicted molecule property value of the output molecule for each molecule property in the set of molecule properties comprises, for each of one or more molecule properties: processing data characterizing the output molecule using a property’ prediction machine learning model to generate the predicted molecule property value of the output molecule as an output of the property prediction machine learning model.

8. The method of any one of claims 6-7, wherein determining the respective predicted molecule property value of the output molecule for each molecule property' in the set of molecule properties comprises, for each of one or more molecule properties: performing a computational simulation of the output molecule to determine the predicted molecule property value of the output molecule.

9. The method of any one of claims 6-8, wherein for each output molecule in the set of output molecules, the model output of the generative machine learning model defines a respective predicted molecule property value of the output molecule for each of one or moreAttorney Docket No. 46567-1587WO1 molecule properties in the set of molecule properties; and wherein determining the respective predicted molecule property value of the output molecule for each molecule property in the set of molecule properties comprises, for each of one or more molecule properties: extracting the predicted molecule property value of the output molecule from the model output of the generative machine learning model.

10. The method of any one of claims 5-9, wherein the updated model input to the generative machine learning model comprises: a plurality of molecule exemplars that include experimental molecule property values of example molecules determined by performing physical experiments; and a plurality of molecule exemplars that include predicted molecule property values of output molecules determined using computational prediction techniques.

11. The method of any one of claims 5-10. wherein updating the model input to the generative machine learning model comprises: selecting a proper subset of the set of output molecules according to one or more selection criteria based on molecule property7values of the output molecules; and wherein generating the respective new molecule exemplar corresponding to each of one or more output molecules in the set of output molecules comprises: generating a respective new molecule exemplar for only output molecules in the selected proper subset of the set of output molecules.

12. The method of claim 11 , wherein selecting a proper subset of the set of output molecules according to one or more selection criteria based on molecule property values of the output molecules comprises: determining, for each output molecule, a respective binding affinity of the output molecule for a target protein; and selecting the proper subset of the set of output molecules based at least in part on their respective binding affinities for the target protein.

13. The method of claim 12, wherein selecting the proper subset of the set of output molecules based at least in part on their respective binding affinities for the target protein comprises:Attorney Docket No. 46567-1587WO1 determining, for each output molecule, that the output molecule should be selected for inclusion in the proper subset of the set of output molecules only if the binding affinity of the output molecule for the target protein satisfies a threshold.

14. The method of any preceding claim, wherein outputting data identifying the set of molecules comprises, for one or more of the output molecules: providing data representing the output molecule to a user; receiving, from the user, a textual query comprising a request to modify the output molecule; processing a model input that comprises: (i) data representing the chemical structure of the output molecule, and (ii) the textual query from the user, using a molecule modification machine learning model and in accordance with values of a set of molecule modification machine learning model parameters, to generate a model output that defines a chemical structure of a modified molecule, wherein the modified molecule is a version of the output molecule that has been modified in accordance with the request included in the textual query from the user; and outputting data identify ing the modified molecule.

15. The method of claim 14, wherein the molecule modification machine learning model is a generative machine learning model that has been trained to perform a language modeling task.

16. The method of any one of claims 14-15. wherein the request to modify the output molecule comprises one or more of: a request to modify a scaffold structure of the output molecule, or a request to modify a fragment structure of the output molecule, or a request to modify physiochemical properties of the output molecule, or a request to modify' biochemical properties of the output molecule, or a request to add a functional group to the output molecule, or a request to remove a functional group from the output molecule, or a request to replace a functional group in the output molecule, or a request to change a steric hindrance of the output molecule, or a request to change a polarity of the output molecule, or a request to change a hydrophobicity of the output molecule.

17. The method of any preceding claim, wherein the model input to the generative machine learning model is represented as a sequence of input tokens, where each input token in theAttorney Docket No. 46567-1587WO1 sequence of tokens is selected from a set of possible tokens.

18. The method of claim 17, wherein for each molecule exemplar, the chemical structure of the corresponding example molecule is represented in the sequence of input tokens by a respective Simplified Molecular Input Line Entry System (SMILES) string.

19. The method of any preceding claim, wherein the model output generated by the generative machine learning model is represented as a sequence of output tokens where each output token in the sequence of output tokens is selected from a set of possible tokens.

20. The method of claim 19. wherein for each output molecule, the chemical structure of the output molecule is represented in the sequence of output tokens by a respective Simplified Molecular Input Line Entry System (SMILES) string.

21. The method of any one of claims 19-20, wherein processing the model input using the generative machine learning model and in accordance with values of the set of generative machine learning model parameters to generate the model output that includes data defining the respective chemical structure of each output molecule in the set of output molecules comprises: sequentially generating the sequence of output tokens, one token at a time, starting from a first token in the sequence of output tokens.

22. The method of claim 21, wherein for each of one or more positions in the sequence of output tokens, generating the output token at the position in the sequence of output tokens comprises: processing data comprising: (i) a sequence of input tokens representing the model input, and (ii) any output token that precedes the position in the sequence of output tokens, using the generative machine learning model to generate a score distribution over the set of possible tokens; and selecting the output token at the position in the sequence of output tokens using the score distribution over the set of possible tokens.

23. The method of claim 22, wherein selecting the output token at the position in the sequence of output tokens using the score distribution over the set of possible tokens comprises:Attorney Docket No. 46567-1587WO1 selecting a token having a highest score under the score distribution over the set of possible tokens as the output token at the position in the sequence of output tokens.

24. The method of claim 22, wherein selecting the output token at the position in the sequence of output tokens using the score distribution over the set of possible tokens comprises : randomly sampling a token from the set of possible tokens, in accordance with the score distribution over the set of possible tokens, to select the output token at the position in the sequence of output tokens.

25. The method of any preceding claim, wherein the generative machine learning model is implemented as a neural network.

26. The method of claim 25, wherein the neural network comprises a plurality of attention layers.

27. The method of any preceding claim, wherein the generative machine learning model has been trained to perform a language modeling task.

28. The method of claim 27, wherein training the generative machine learning model to perform the language modeling task comprises: obtaining a set of training examples, wherein each training example comprises at least a sequence of output tokens; and training the generative machine learning model on the set of training examples, comprising, for each training example: training the generative machine learning model to generate, for each position in the sequence of output tokens, the output token occupying the position by processing at least any output tokens that precede the position in the sequence of output tokens.

29. The method of any preceding claim, wherein the set of molecule properties comprises one or more of: binding affinity for a target protein, molecular weight, topological polar surface area, lipophilicity, synthetic accessibility, solubility, permeability, absorption, distribution, metabolism, excretion, toxicity, or stability.Attorney Docket No. 46567-1587WO130. The method of any preceding claim, wherein outputting data identifying the set of output molecules comprises: storing data identifying the set of output molecules in a memory, or providing a representation of the set of output molecules for presentation on a display, or transmitting data identifying the set of output molecules over a data communications network.

31. The method of any preceding claim, further comprising selecting one or more output molecules from the set of output molecules for physical synthesis.

32. The method of claim 31, wherein selecting one or more output molecules from the set of output molecules for physical synthesis comprises: determining a respective binding affinity of each output molecule for a target protein; and selecting one or more output molecules from the set of output molecules for physical synthesis based at least in part on the binding affinities of the output molecules for the target protein.

33. The method of any one of claims 31-32, further comprising physically synthesizing each of the one or more output molecules that are selected for physical synthesis.

34. The method of any one of claims 31-33, further comprising, for each of the one or more output molecules that are selected for physical synthesis: performing physical experiments on one or more physically synthesized instances of the output molecule to evaluate one or more molecule properties of the output molecule.

35. A system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of any one of claims 1-32.

36. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective method of any one of claims 1-32.