Targeted molecule generation with latent reinforcement learning

Reinforcement learning in the latent space of a VAE addresses the challenges of generating molecules with specified properties and structural similarity, enhancing the efficiency and validity of molecular synthesis.

WO2025193897A1PCT designated stage Publication Date: 2025-09-18CELLARITY INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/019689
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-14
Filing Date
2025-03-13
Publication Date
2025-09-18

AI Technical Summary

Technical Problem

Existing methods for generating molecules with pre-specified molecular properties and structural similarity face challenges due to the large, discrete, and unstructured molecular search space, leading to low validity rates and high costs in molecular synthesis and testing.

Method used

Employing reinforcement learning in the latent space of a generative model, such as a variational autoencoder (VAE), to navigate and identify molecules with desired properties by modifying the VAE training algorithm and using a policy gradient algorithm to optimize the latent space for targeted molecule generation.

Benefits of technology

The method effectively generates molecules with desired properties and structural similarity, improving the validity rate and reducing synthesis costs by navigating the latent space efficiently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025019689_18092025_PF_FP_ABST
    Figure US2025019689_18092025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for policy model training identify compounds that each have a value for a physical property in a specified range. A latent space encodes a combinatorial synthesis library as vectors. A target for the property is obtained. Responsive to inputting the target, a first cutoff, and a second cutoff into the policy model, an action vector is obtained. The vector is sampled to identify a plurality of latent space points. Each point is decoded to obtain a plurality of candidate compounds. A value for the property is determined for each of the candidates using their molecular structures. The parameters are updated through application of a loss function to each of the candidate compounds. The function includes a reward term that is determined by application of the target and first and second cutoffs to the value for the property for each candidate compound.
Need to check novelty before this filing date? Find Prior Art

Description

TARGETED MOLECULE GENERATION WITH LATENT REINFORCEMENT LEARNINGTECHNICAL FIELD

[0001] This specification relates generally to using reinforcement learning in the latent space of a generative model to identify compounds that have target properties.CROSS-REFERENCE TO RELATED APPLICATION

[0002] This application claims priority to and benefit of U.S. Provisional Patent Application No. 63 / 565,160, filed on March 14, 2024, the content of which is hereby incorporated by reference in its entirety.BACKGROUND

[0003] Identification of novel compounds that have desirable pharmaceutical properties using an in silico combinatorial synthesis library can be constructed as an optimization problem that searches for compounds that satisfy quantitative requirements. However, optimization in molecular space is difficult, because the search space is large, discrete, and unstructured. Moreover, making and testing new compounds is costly and time consuming, and the number of potential candidates is overwhelming. Only about 108substances have ever been synthesized (see, Kim et al., 2016, “Substance and Compound databases,” Nucleic Acids Res. 44, D1202-D1213), whereas the range of potential drug-like molecules is estimated to be between 1023and IO60. See, Polishchuk et al, 2013, “Estimation of the size of drug-like chemical space based on GDB-17 data,” J. Comput. Aided Mol. Des. 27, 675-679.

[0004] Virtual screening can be used to speed up this optimization problem. Virtual libraries containing thousands to hundreds of millions of candidates can be assayed with first-principles simulations or statistical predictions based on learned proxy models, and only the most promising leads are selected and tested experimentally. As illustrated in Figure 17, Gomez-Bombarelli et al., 2018, “Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules,” ACS Cent. Sc. 4, pp. 268-276, make use of probabilistic generative model, in the form of a variational autoencoder (VAE) with recurrent neural network encoding, to represent virtual libraries in a multidimensional latent space that is a continuous representation of the virtual library. VAEs add stochasticity to the encoder, which combined with a penalty term, encourages all areas of the latent space to correspond to a valid decoding. The VAE has two parts, an encoder that projects molecules into a continuous latent space and a decoder that reconstructs molecules from the latent space. Thisforces the decoder to learn how to decode a wider variety of latent points and find more robust representations. VAEs force the latent space to be a multi-dimensional continuous representation of the virtual library. With such a continuous representation, Gomez-Bombarelli, Id., was able to use gradient based continuous optimization to guide the search for compounds with desirable properties.

[0005] A general challenge in these early efforts is the low validity rate in variational Auto-encoders. Thus, improvement is needed in such efforts to generate molecules with pre-specified molecular properties and the generation of analogs of a given molecule that aim to improve a property of interest while maintaining structural similarity with the original molecule.SUMMARY

[0006] The present disclosure addresses the prior art need for generating molecules with prespecified molecular properties and the generation of analogs of a given molecule that aim to improve a property of interest while optionally maintaining structural similarity with the original molecule. An embodiment of the present disclosure is based on reinforcement learning (RL) that operates in the latent space of a generative machine learning model such as a variation autoencoder (VAE). The present disclosure improves upon early VAE training efforts by modifying the original VAE training algorithm and thus improving their validity rate.

[0007] The disclosed methods operate in the latent space of the generative machine learning model. With this approach, chemical operations are reduced into vector modifications in a continuous multidimensional space. Moreover, the disclosed methods pair the generative model, pre-trained on a dataset of molecules, with reinforcement learning. The RL agent operates in the latent space of the generative model and is trained using a policy gradient algorithm. The objective is to learn an optimal policy for navigating the latent space to identify molecules that exhibit an improved profile with respect to a property of interest.

[0008] One aspect of the present disclosure provides a method of training a policy model to identify a subset of a latent space comprising a plurality of compounds that each have a value for at least a first predetermined physical property, in one or more predetermined physical properties, that is in a first specified range, is provided.

[0009] In some embodiments, the first predetermined physical property is compound molecular weight.

[0010] In some embodiments, the first predetermined physical property is compound distribution coefficient (logD), compound octanol-water partition coefficient (logP), penalized logP, compound acid dissociation constant (pKA), compound acid dissociation constant (kA), or compound topological polar surface area (TPSA). In some embodiments, the first predetermined physical property is drug likeness (QED score). In some embodiments, the first predetermined physical property is permeability.

[0011] In some embodiments, the combinatorial synthesis library comprises 1,000 compounds, 5,000 compounds, 10,000 compounds, 50,000 compounds, 100,000 compounds, 1 million compounds 10 million compounds, 100 million compounds, or a billion compounds.

[0012] A specification of a target and a second cutoff for the first predetermined physical property is obtained.

[0013] Responsive to inputting the target, the second cutoff into the policy model, an action vector of fixed dimension d is obtained as output from the policy model by application of the initial state vector to the policy model in accordance with a first plurality of parameters of the policy model.

[0014] In some embodiments, the policy model is a multi-layer perceptron. In some embodiments, the policy model is a deep neural network. In some embodiments, the first plurality of parameters comprises 10,000 or more parameters, 100,000 or more parameters, or 1 x 106or more parameters.

[0015] In some embodiments, the task is a constraint property optimization problem in which candidate compounds having the desired physical properties are sought subject to the constraint that they share structural similarity with a test compound. In some embodiments the test compound is an organic compound having a molecular weight of less than 2000 Daltons. In some embodiments, the test compound is an organic compound that satisfies each of the Lipinski rule of five criteria. In some embodiments, the test compound is an organic compound that satisfies at least three criteria of the Lipinski rule of five criteria.

[0016] In some embodiments, for the constraint optimization problem an initial state vector of fixed dimension d of the test compound is obtained as output from an encoder model, the encoder model comprising a third plurality of parameters, by inputting a string representation of the test compound into the encoder model. In some embodiments, the third plurality of parameters comprises 10,000 or more parameters, 100,000 or more parameters, or 1 x 106or more parameters.

[0017] In some embodiments, the encoder model is a first recurrent neural network and the decoder model is a second recurrent neural network.

[0018] In some embodiments, the encoder model is a first neural network and the decoder model is a second neural network.

[0019] In some embodiments, the string representation is in a SMARTS, a simplified molecular- input line-entry system (SMILES), DeepSMILES, or SELFIES format.

[0020] The action vector is sampled on a probabilistic basis to identify a plurality of points in the latent space. Each point in the plurality of points is of fixed dimension d.

[0021] In some embodiments, the sampling samples the action vector as a multivariate normal distribution with a mean equal to the action vector and a standard deviation equal to a constant value.

[0022] Each point in the plurality of points is inputted into a decoder model thereby obtaining a corresponding structure of each candidate compound in a plurality of candidate compounds in accordance with a second plurality of parameters associated with the decoder model, each candidate compound in the plurality of candidate compounds corresponding to a point in the plurality of points.

[0023] In some embodiments, the plurality of candidate compounds comprises five or more candidate chemical compounds.

[0024] In some embodiments, each candidate compound in the plurality of candidate compounds is an organic compound having a molecular weight of less than 2000 Daltons. In some embodiments, each candidate compound in the plurality of candidate compounds is an organic compound that satisfies each of the Lipinski rule of five criteria. In some embodiments, each candidate compound in the plurality of candidate compounds is an organic compound that satisfies at least three criteria of the Lipinski rule of five criteria.

[0025] In some embodiments, the second plurality of parameters comprises 10,000 or more parameters, 100,000 or more parameters, or 1 x 106or more parameters.

[0026] A value for the first predetermined physical property is determined for each candidate compound in the plurality of candidate compounds using the corresponding structure of the candidate compound.

[0027] An update to the first plurality of parameters is acquired, through application of a loss function to each candidate compound in the plurality of candidate compounds. The loss functionincludes a reward term, a magnitude of which is determined, at least in part, by an application of the target, a first cutoff and the second cutoff of the first predetermined physical property to the value for the first predetermined physical property for a candidate compound in the plurality of candidate compounds.

[0028] In some embodiments, the application of the target, the first cutoff and second cutoff of the first predetermined physical property to determine a magnitude of the reward term comprises: (a) using a maximum reward for the first predetermined physical property in the loss function when the value for the metric of the first predetermined physical property for the candidate compound is both (1) greater than the target value minus the first cutoff value for the first predetermined physical property and (2) less than the target value plus the first cutoff value for the first predetermined physical property, (b) using a minimum reward in the loss function for the first predetermined physical property when the value for the metric of the first predetermined physical property for the candidate chemical compound is less than the target value minus the second cutoff value for the first predetermined physical property or greater than the target value plus the second cutoff value for the first predetermined physical property, and (c) using a value in the loss function for the first predetermined physical property that is a linear function of the value for the metric of the first predetermined physical property for the candidate chemical compound that produces a reward value between the minimum reward and the maximum reward when the value for the metric of the first predetermined physical property for the candidate chemical compound fails to qualify for the maximum reward or the minimum reward.

[0029] In some embodiments, the one or more predetermined physical properties is a plurality of physical properties, the plurality of compounds each have a value for a second predetermined physical property in the plurality of physical properties that is in a second specified range, the obtaining further comprises obtaining a specification of a target for the second predetermined physical property, the determining a value further determines a value for the second predetermined physical property for each candidate compound in the plurality of candidate compounds using the corresponding structure of the candidate compound, and the magnitude of the reward function is further determined by an application of the target, the first cutoff and the second cutoff of the second predetermined physical property to the value for the second predetermined physical property for the candidate compound.

[0030] In some embodiments, the magnitude of the reward function is an average or minimum of (i) the application of the target, the first cutoff and the second cutoff of the first predetermined physical property to the value for the first predetermined physical property for a candidate compound in theplurality of candidate compounds, and (ii) the application of the target, the first cutoff and the second cutoff of the second predetermined physical property to the value for the second predetermined physical property for the candidate compound.

[0031] In some embodiments, the magnitude of the reward function is a minimum of (i) the application of the target, the first cutoff and the second cutoff of the first predetermined physical property to the value for the first predetermined physical property for a candidate compound in the plurality of candidate compounds, and (ii) the application of the target, the first cutoff and the second cutoff of the second predetermined physical property to the value for the second predetermined physical property for the candidate compound.

[0032] The reward can be either calculated exactly (such as in the case of molecular weight) or estimated using a predictive machine learning model, such as in the case of permeability.

[0033] The policy model is updated with the update to the first plurality of parameters thereby training the policy model to sample the latent space to identify the subset of the latent space comprising the plurality of compounds that have the value for the first predetermined physical property that is in the first specified range.

[0034] Another aspect of the present disclosure provides a computer system for training a policy model to identify a subset of a latent space comprising a plurality of compounds that each have a value for at least a first predetermined physical property, in one or more predetermined physical properties, that is in a first specified range. The computer system comprises at least one processor and a memory storing at least one program for execution by the at least one processor. The at least one program comprises instructions for performing any of the methods and embodiments disclosed herein, and / or any combinations thereof as will be apparent to one skilled in the art. In some embodiments, the at least one program is configured for execution by a computer.

[0035] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium having stored thereon program code instructions that, when executed by a processor, cause the processor to perform a method of training a policy model to identify a subset of a latent space comprising a plurality of compounds that each have a value for at least a first predetermined physical property, in one or more predetermined physical properties, that is in a first specified range. In some embodiments, the program code instructions comprise instructions for performing any of the methods and embodiments disclosed herein, and / or any combinations thereof. In some embodiments, the program code instructions are configured for execution by a computer or a computer system.INCORPORATION BY REFERENCE

[0036] All publications, patents, and patent applications herein are incorporated by reference in their entireties. In the event of a conflict between a term herein and a term in an incorporated reference, the term herein controls.BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The implementations disclosed herein are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings. Like reference numerals refer to corresponding parts throughout the several views of the drawings.

[0038] Figures 1 A and IB illustrate a block diagram illustrating an example of a computing system for targeted molecule generation intended to assist drug design in accordance with some embodiments of the present disclosure.

[0039] Figures 2A, 2B, 2C, 2D, 2E and 2F illustrate an example of a method for training a policy model to identify a subset of a latent space comprising a plurality of compounds that each have a value for at least a first predetermined physical property, in one or more predetermined physical properties, that is in a first specified range, in accordance with some embodiments of the present disclosure in which optional elements are indicated by dashed boxes.

[0040] Figure 3 illustrates the use of reinforcement learning to guide a search of a chemical space, encoded as a continuous latent space, for chemical compounds to identify compounds that have targeted properties, in accordance with some embodiments of the present disclosure

[0041] Figures 4A and 4B illustrate an architecture of a latent reinforcement learning framework in which a policy network (policy model) produces an action vector from a state vector representing a test compound, and where the framework samples the action vector in the latent space to acquire a plurality of latent vectors, decodes the latent vectors into candidate compounds, assesses the candidate compounds for a desired physical property, and updates the policy network in accordance a reward term that rewards for candidate compounds having the desired physical property, in accordance with some embodiments of the present disclosure.

[0042] Figure 5 illustrates a reward function for property sampling in accordance with some embodiments of the present disclosure.

[0043] Figure 6 illustrates the two reward functions for constraint property optimization in accordance with some embodiments of the present disclosure.

[0044] Figure 7 illustrates how the distribution of the molecular weight of candidate compounds has been shifted when using the reinforcement framework of Figure 4 to enforce a desired molecular weight in accordance with some embodiments of the present disclosure.

[0045] Figures 8A and 8B illustrate how the distribution of the molecular weight and logP of candidate compounds has been simultaneously shifted when using the reinforcement framework of Figure 4 to simultaneously enforce a desired molecular weight and logP in accordance with some embodiments of the present disclosure.

[0046] Figures 9A and 9B illustrate optimizing compounds for drug likeness (QED), permeability, and penalized LogP, with reinforcement learning to impose constraint property optimization in accordance with some embodiments of the present disclosure.

[0047] Figures 10A, 10B, 10C, and 10D illustrate improvement of permeability, QED and pLogP after optimizing compounds with reinforcement learning to impose constraint property optimization in accordance with some embodiments of the present disclosure.

[0048] Figure 11 illustrates the benchmarking of the disclosed reinforcement learning in accordance with one embodiment of the present disclosure against two existing molecule optimization tools, GCPN and MolDQN.

[0049] Figure 12 illustrates a policy model, and the inputs and outputs of the policy model in various embodiments in accordance with the present disclosure.

[0050] Figure 13 illustrate example compounds discovered that have various predetermined physical properties in various embodiments in accordance with the present disclosure.

[0051] Figure 14A and 14B illustrate benchmarking the reinforcement model against a common set of 800 molecules for plogP and QED in accordance with an embodiment of the present disclosure.

[0052] Figures 15A and 15B illustrate how the pretrained reinforcement learning model is able to improve pLogP on unseen compounds in accordance with an embodiment of the present disclosure.

[0053] Figure 16 illustrates how, to mitigate posterior collapse, in some embodiments of the present disclosure, a cyclical annealing schedule for a beta term for a Beta-variational autoencoder was used.

[0054] Figure 17 illustrates a variational autoencoder in accordance with the prior art.

[0055] Figure 18 illustrates the predicted permeability of unseen compounds versus the experimental permeability of a permeability model in accordance with the present disclosure.DETAILED DESCRIPTION

[0056] A novel framework for targeted molecule generation intended to assist drug design is provided. The objective is twofold: first, the generation of molecules with pre-specified molecular properties, and second, the generation of analogs of a given molecule aiming at improving a property of interest while maintaining structural similarity with the original molecule. One embodiment of the present disclosure is based on reinforcement learning (RL) that operates in the latent space of a variational autoencoder (VAE). With this approach, chemical operations are reduced into vector modifications in a continuous multi-dimensional space. The method pairs a VAE model, pre-trained on a large dataset of molecules, with RL. The RL agent operates in the latent space of the VAE and is trained using a policy gradient algorithm. The objective is to learn an optimal policy for navigating the latent space to identify molecules that exhibit an improved profile with respect to a property of interest. The capability of the disclosed method to generate molecules with desired properties, as well as analogs of molecules with improved profile is demonstrated. For instance, the disclosed method is used to generate molecules within a pre-specified range of a single molecular property of interest as well as two molecular properties simultaneously, such as molecular weight and LogP. The disclosed method has also been applied to three optimization tasks, that is, optimizing a set of molecules for permeability, drug likeness (QED) and LogP, respectively. The results show that the disclosed method can improve the property of interest while maintaining molecular similarity and is able to generalize beyond the training set. The results demonstrate the applicability of the disclosed method in assisting the optimization of candidate molecules in the drug design process.

[0057] Figure 4 illustrates the reinforcement learning framework. The RL framework has three major components: the policy network (policy model), the reward function, and the generative model. The policy network is the part of the framework that is being trained to perform targeted sampling. The policy network operates in the latent space of the generative model (e.g., a variational autoencoder). The generative model established the connection between the latent space and the chemical space, and its decoder allows the translation of the latent samples (points) to chemical molecules. The reward function provides feedback to the policy network on whether the samples molecules have the desirable properties or not.

[0058] The overall framework is shared among the two tasks, property targeting and constraint property optimization, however the interpretation of certain components is different as shown in Table 1. In property targeting the policy network is being trained to identify where to sample in the latent space or otherwise it identifies latent vectors that correspond to chemical molecules with the desirable properties. In constraint optimization, the policy network is being trained to identify perturbations of the starting latent vector that will result in improvements in the property of interest.

[0059] Table 1: Interpretation of the RL framework components for the two tasks, property targeting and constraint property optimization.

[0060] Reference will now be made in detail to embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to one of ordinary skill in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0061] Definitions

[0062] As used herein, the term “about” or “approximately” can mean within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which can depend in part on how the value is measured or determined, e.g., the limitations of the measurement system. For example, “about” can mean within 1 or more than 1 standard deviation, per the practice in the art. “About” can mean a range of ±20%, ±10%, ±5%, or ±1% of a given value. The term “about” or “approximately” can mean within an order of magnitude, within 5-fold, or within 2-fold, of a value. Where particular values are described in the application and claims, unless otherwise stated the term“about” meaning within an acceptable error range for the particular value should be assumed. The term “about” can have the meaning as commonly understood by one of ordinary skill in the art. The term “about” can refer to ±10%. The term “about” can refer to ±5%.

[0063] As used interchangeably herein, the term “classifier” or “model” refers to a machine learning model.

[0064] In some embodiments, a model is an unsupervised learning model. One example of an unsupervised learning model is cluster analysis. In some embodiments, a model includes supervised machine learning. Nonlimiting examples of supervised learning models include, but are not limited to, logistic regression models, neural networks, support vector machines, Naive Bayes models, nearest neighbors models, random forest models, decision trees, boosted trees, multinomial logistic regression models, linear models, linear regression models, Gradient Boosting models, mixture models, hidden Markov models, Gaussian NB models, linear discriminant analysis models, or any combinations thereof. In some embodiments, a model is a multinomial classifier algorithm. In some embodiments, a model is a 2-stage stochastic gradient descent (SGD) model. In some embodiments, a model is a deep neural network (e.g, a deep-and-wide sample-level model).

[0065] Neural networks. In some embodiments, the model is a neural network (e.g., a convolutional neural network and / or a residual neural network). Neural networks, also known as artificial neural networks (ANNs), include convolutional and / or residual neural networks (deep learning models). In some embodiments, neural networks are machine learning models that are trained to map an input dataset to an output dataset, where the neural network includes an interconnected group of nodes organized into multiple layers of nodes. For example, in some embodiments, the neural network architecture includes at least an input layer, one or more hidden layers, and an output layer. In some embodiments, the neural network includes any total number of layers, and any number of hidden layers, where the hidden layers function as trainable feature extractors that allow mapping of a set of input data to an output value or set of output values. In some embodiments, a deep learning model is a neural network including a plurality of hidden layers, e.g., two or more hidden layers. In some instances, each layer of the neural network includes a number of nodes (or “neurons”). In some embodiments, a node receives input that comes either directly from the input data or the output of nodes in previous layers, and performs a specific operation, e.g., a summation operation. In some embodiments, a connection from an input to a node is associated with a parameter (e.g., a weight and / or weighting factor). In some embodiments, the node sums up the products of all pairs of inputs,xi, and their associated parameters. In some embodiments, the weighted sum is offset with a bias, b. In some embodiments, the output of a node or neuron is gated using a threshold or activation function, f, which, in some instances, is a linear or non-linear function. In some embodiments, the activation function is, for example, a rectified linear unit (ReLU) activation function, a Leaky ReLU activation function, or other function such as a saturating hyperbolic tangent, identity, binary step, logistic, arcTan, softsign, parametric rectified linear unit, exponential linear unit, softPlus, bent identity, softExponential, Sinusoid, Sine, Gaussian, or sigmoid function, or any combination thereof.

[0066] In some implementations, the weighting factors, bias values, and threshold values, or other computational parameters of the neural network, are “taught” or “learned” in a training phase using one or more sets of training data. For example, in some implementations, the parameters are trained using the input data from a training dataset and a gradient descent or backward propagation method so that the output value(s) that the ANN computes are consistent with the examples included in the training dataset. In some embodiments, the parameters are obtained from a back propagation neural network training process.

[0067] Any of a variety of neural networks are suitable for use in accordance with the present disclosure. Examples include, but are not limited to, feedforward neural networks, radial basis function networks, recurrent neural networks, residual neural networks, convolutional neural networks, residual convolutional neural networks, and the like, or any combination thereof. In some embodiments, the machine learning makes use of a pre-trained and / or transfer-learned ANN or deep learning architecture. In some implementations, convolutional and / or residual neural networks are used, in accordance with the present disclosure.

[0068] For instance, a deep neural network model includes an input layer, a plurality of individually parameterized (e.g., weighted) convolutional layers, and an output scorer. The parameters (e.g., weights) of each of the convolutional layers as well as the input layer contribute to the plurality of parameters (e.g., weights) associated with the deep neural network model. In some embodiments, at least 50 parameters, at least 100 parameters, at least 1000 parameters, at least 2000 parameters or at least 5000 parameters are associated with the deep neural network model. As such, deep neural network models require a computer to be used because they cannot be mentally solved. In other words, given an input to the model, the model output needs to be determined using a computer rather than mentally in such embodiments. See, for example, Krizhevsky et al., 2012, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems 2,Pereira, Burges, Bottou, Weinberger, eds., pp. 1097-1105, Curran Associates, Inc.; Zeiler, 2012 “ADADELTA: an adaptive learning rate method,” CoRR, vol. abs / 1212.5701; and Rumelhart et al., 1988, “Neurocomputing: Foundations of research,” ch. Learning Representations by Back-propagating Errors, pp. 696-699, Cambridge, MA, USA: MIT Press, each of which is hereby incorporated by reference.[0069J Neural networks, including convolutional neural networks, suitable for use as models are disclosed in, for example, Vincent et al., 2010, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion,” J Mach Learn Res 11, pp. 3371- 3408; Larochelle etal., 2009, “Exploring strategies for training deep neural networks,” J Mach Learn Res 10, pp. 1-40; and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of which is hereby incorporated by reference. Additional example neural networks suitable for use as models are disclosed in Duda et al., 2001, Pattern Classification, Second Edition, John Wiley & Sons, Inc., New York; and Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, each of which is hereby incorporated by reference in its entirety. Additional example neural networks suitable for use as models are also described in Draghici, 2003, Data Analysis Tools for DNA Microarrays, Chapman & Hall / CRC; and Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, each of which is hereby incorporated by reference in its entirety.

[0070] As used herein, the term “parameter” refers to any coefficient or, similarly, any value of an internal or external element (e.g., a weight and / or a hyperparameter) in an algorithm, model, regressor, and / or classifier that can affect (e.g., modify, tailor, and / or adjust) one or more inputs, outputs, and / or functions in the algorithm, model, regressor and / or classifier. For example, in some embodiments, a parameter refers to any coefficient, weight, and / or hyperparameter that can be used to control, modify, tailor, and / or adjust the behavior, learning, and / or performance of an algorithm, model, regressor, and / or classifier. In some instances, a parameter is used to increase or decrease the influence of an input (e.g., a feature) to an algorithm, model, regressor, and / or classifier. As a nonlimiting example, in some embodiments, a parameter is used to increase or decrease the influence of a node e.g., of a neural network), where the node includes one or more activation functions. Assignment of parameters to specific inputs, outputs, and / or functions is not limited to any one paradigm for a given algorithm, model, regressor, and / or classifier but can be used in any suitable algorithm, model, regressor, and / or classifier architecture for a desired performance. In some embodiments, a parameter has a fixed value. In some embodiments, a value of a parameter is manually and / or automatically adjustable. In someembodiments, a value of a parameter is modified by a validation and / or training process for an algorithm, model, regressor, and / or classifier (e.g.. by error minimization and / or backpropagation methods). In some embodiments, an algorithm, model, regressor, and / or classifier of the present disclosure includes a plurality of parameters. In some embodiments, the plurality of parameters is n parameters, where: n > 2; n > 5; n > 10; n > 25; n > 40; n > 50; n > 75; n > 100; n > 125; n > 150; n > 200; n > 225; n > 250; n > 350; n > 500; n > 600; n > 750; n > 1,000; n > 2,000; n > 4,000; n > 5,000; n > 7,500; n > 10,000; n > 20,000; n > 40,000; n > 75,000; n > 100,000; n > 200,000; n > 500,000, n > 1 x 106, n > 5 x 106, or n > 1 x 107. As such, some embodiments of the algorithms, models, regressors, and / or classifiers of the present disclosure cannot be mentally performed. In some embodiments n is between 10,000 and 1 x 107, between 100,000 and 5 x 106, or between 500,000 and 1 x 106. In some embodiments, the algorithms, models, regressors, and / or classifier of the present disclosure operate in a k-dimensional space, where k is a positive integer of 5 or greater (e.g., 5, 6, 7, 8, 9, 10, etc.). As such, some embodiments of the algorithms, models, regressors, and / or classifiers of the present disclosure cannot be mentally performed.

[0071] The terminology used herein is for the purpose of describing particular cases only and is not intended to be limiting. As used herein, the singular forms “a,” “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, to the extent that the terms “including,” “includes,” “having,” “has,” “with,” or variants thereof are used in either the detailed description and / or the claims, such terms are intended to be inclusive in a manner similar to the term “comprising.”

[0072] Several aspects are described below with reference to example applications for illustration. It should be understood that numerous specific details, relationships, and methods are set forth to provide a full understanding of the features described herein. One having ordinary skill in the relevant art, however, will readily recognize that the features described herein can be practiced without one or more of the specific details or with other methods. The features described herein are not limited by the illustrated ordering of acts or events, as some acts can occur in different orders and / or concurrently with other acts or events. Furthermore, not all illustrated acts or events are required to implement a methodology in accordance with the features described herein.

[0073] Exemplary System Embodiments

[0074] Details of an exemplary system are now described in conjunction with Figure 1. Figure 1 is a block diagram illustrating a system 100 in accordance with some implementations. The device 100 insome implementations includes at least one or more processing units CPU(s) 102 (also referred to as processors), one or more network interfaces 104, a display 106 having a user interface 108, an input device 110, a memory, and one or more communication buses 114 for interconnecting these components. The one or more communication buses 114 optionally include circuitry (sometimes called a chipset) that interconnects and controls communications between system components. The memory may be a non-persistent memory 111, a persistent memory 112, or any combination thereof. The non-persistent memory 111 typically includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, ROM, EEPROM, flash memory, whereas the persistent memory typically includes CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Regardless of its specific implementation, the non-persistent memory 111 and / or the persistent memory 112 comprises at least one non-transitory computer-readable storage medium, and it stores thereon computer-executable executable instructions which can be in the form of programs, modules, and data structures.

[0075] In some embodiments, as shown in Figure 1, the non-persistent memory 111 stores the following:• optional instructions, programs, data or information associated with optional operating system 116, which includes procedures for handling various basic system services and for performing hardware-dependent tasks;• optional instructions, programs, data or information associated with optional network communication module (or instructions) 118 for connecting the system 100 with other devices and / or to a communication network;• programs, data or information associated with a first predetermined physical property 120-1 including a target value 122-1, a first cutoff value 124-1, and a second cutoff value 126-1 for the first predetermined physical property 120-1;• programs, data or information associated with a second predetermined physical property 120- 2;• programs, data or information associated with a policy model 128 for obtaining an action vector 137, the policy model trained by a policy 130 comprising a training loss function 132 thatincludes a reward term 134, the policy model including a first plurality of parameters 136-1, 136-2, . . ., 136-N;• programs, data or information associated with a test compound string representation 138 and its corresponding initial state vector 140;• programs, data or information associated with sampling of the action vector 142 to identify a plurality of points (latent vectors) 144-1, 144-2, ..., 144-X;• programs, data or information associated with an encoder model 148 comprising a plurality of parameters 150-1, 1 0-2, ..., 150-M for encoding the test compound string representation 138 into the corresponding initial state vector 140 of the test compound;• a latent space 152 defined by the encoder and the plurality of parameters 150 of the encoder, where the encoder can encode each compound in a combinatorial synthesis library as an undercomplete vector of fixed dimension d in which each entry in the undercomplete vector is a distribution within the latent space; and• programs, data or information associated with a decoder model 154 comprising a plurality of parameters 156-1, 156-2, ..., 156-Q for obtaining a corresponding structure of each candidate compound in a plurality of candidate compounds in accordance with parameters 156-1, 156-2, ..., 156- Q, each candidate compound in the plurality of candidate compounds corresponding to a point 144 from the latent space 152 from the sampling of the action vector.

[0076] In various implementations, one or more of the above-identified elements are stored in one or more of the previously mentioned memory devices, and correspond to a set of instructions for performing various methods described herein. The above-identified modules, data, or programs (e.g., sets of instructions) may not be implemented as separate software programs, procedures, datasets, or modules, and thus various subsets of these modules and data may be combined or otherwise rearranged in various implementations. In some implementations, the non-persistent memory 111 optionally stores a subset of the modules and data structures identified above. Furthermore, in some embodiments, the memory stores additional modules and data structures not described above. In some embodiments, one or more of the above-identified elements are stored in a computer system, other than that of the system 100, that is addressable by the system 100 so that the system 100 may retrieve all or a portion of such data.

[0077] Although Figure 1 depicts a “system 100,” the figure is intended more as functional description of the various features which may be present in computer systems than as a structural schematic of the implementations described herein. In practice, and as recognized by those of ordinary skill in the art, items shown separately could be combined and some items can be separate. Moreover, although Figure 1 depicts certain data and modules in the memory (which can be non-persistent memory 111 or persistent memory 112), it should be appreciated that these data and modules, or portion(s) thereof, may be stored in more than one memory.

[0078] Now that a computer system for training a policy model 128 to identify a subset of a latent space comprising a plurality of compounds that each have a value for at least a first predetermined physical property has been described in conjunction with Figures 1A and IB, various methods for training the policy model to identify the plurality of compounds are described below in conjunction with Figures 2A through 2F. Figure 2 is focused on property targeting. However, the method of Figure 2 is also applicable to the task to constraint property optimization by using different reward functions, as fully explained below.

[0079] Referring to block 200, a method of training a policy model 128 is provided that identifies a subset of a latent space comprising a plurality of compounds that each have a value for at least a first predetermined physical property, in one or more predetermined physical properties, that is in a first specified range.

[0080] In some embodiments the latent space 152 is created using a generative model that comprises an encoder model 148 and a decoder model 154. The generative model therefore facilitates the connection between chemical space and the latent space. The encoder model 148 translates each chemical compound into a latent vector (z) in a multi-dimensional continuous space (latent space 152), while the decoder model 154 retrieves the molecular structure of the chemical compound from a given latent vector (z').

[0081] In some embodiments, the generative model, comprising the encoder model 148 and the decoder model 154 is a variational autoencoder (VAE) (see, for example, Kingma and Max, 2019, Foundations and Trends in Machine Learning, 12(4), ISSN 1935-8237), a conditional variational autoencoder (CVAE), a stacked denoising deep autoencoder (see, for example, Zheng, 2019, Stacked Denoising Auto-Encoder for Short-Term Load Forecasting: Deep Learning with Stacked Denoising Auto-Encoder Algorithm, Lambert Academic Publishing), a deep recurrent autoencoder, a convolutional autoencoder, or a transformer network. In some embodiments, the autoencoder includesa generative adversarial network (GAN), for example a variational autoencoder generative adversarial network (VAE-GAN) architecture. In some embodiments, the autoencoder is a contractive autoencoder. In some embodiments, the autoencoder is a long short-term memory (LSTM) recurrent neural network. In some embodiments, the LSTM recurrent neural network comprises a stacked bidirectional encoder and a stacked decoder.[0082J In some preferred embodiments, the autoencoder (comprising the encoder model 148 and the decoder model 154) is a VAE, and the latent space 152 is constrained to be a plurality of predefined distributions, such that each value in the latent space is drawn from one of the predefined distributions. In some such embodiments, each such predefined distribution is a Gaussian distribution. Thus, in some such embodiments, for a given chemical compound, the latent space 152 is populated by a plurality of vectors generated by the encoder upon inputting a string representation of the chemical compound into the encoder model 148, such that plotting the values of each such vector generates a Gaussian distribution corresponding to the chemical compound. In fact, in some embodiments, the latent space 152 is trained on by a generative model such as an autoencoder on thousands, hundreds of thousands, millions of chemical compounds, or more, each represented by its own Gaussian distribution generated by the autoencoder. As such, in some embodiments the generative model, comprising the encoder model 148 and the decoder model 154 is a variational autoencoder (VAE), see Kingma and Welling, 2013, “Auto-Encoding Variational Bayes,” arXiv preprint arXiv:1312.6114 which is hereby incorporated by reference, that gives a probabilistic representation of a compound in the latent space 152. Variational autoencoders are further described in Kristiadi, 2016, “Variational Autoencoder: Intuition and Implementation,” and Doersch, 2016, “Tutorial on variational autoencoders.” arXiv preprint arXiv: 1606.05908, each of which is incorporated by reference.

[0083] In some embodiments, both the encoder model 148 and the decoder model 154 are implemented as gated recurrent units (GRU). In some embodiments the GRU are trained with word dropout and simulated annealing. In some embodiments the GRU are trained and have the architecture summarized in Table 2:

[0084] Table 2: GRU training hyper-parameters and architecture statistics

[0085] The VAE is trained by a loss function. In some embodiments the loss function comprises two terms. The first term of the loss function in such embodiments is the reconstruction loss that measures the ability of the generative model to reconstruct the structure of a molecular compound from a latent representation of the molecular compound. The second term of the loss function in such embodiments is the Kullback-Leibler divergence (KL divergence) between the learnt and the true distribution of the space that is assumed to be a Gaussian distribution. See Kingma and Welling, 2013, “Auto-Encoding Variational Bayes,” arXiv preprint arXiv: 1312.6114, which is hereby incorporated by reference.

[0086] In some embodiments the generative model comprising the encoder model 148, latent space 152, and decoder model 152 (e.g., the VAE), is pre-trained on a large and diverse dataset of chemical compounds. Once trained, the parameters of the generative model are not further tuned when training the policy model 128 that imposes reinforcement learning. In some embodiments this training dataset of chemical compounds is a diverse dataset of 10 million chemical structures from the PubChem database (see, Kim et al., 2021, “PubChem in 2021 : new data content and improved web interfaces,” Nucleic Acids Research 8(49)(D1), pp. D1388-D1395, which is hereby incorporated by reference, encompassing a wide range of chemical compounds. In some embodiments the training dataset is standardized to non-isomeric SMILES strings using the RDKit package (RDKit: Open-source cheminformatics, https: / / www.rdkit.org, which is hereby incorporated by reference).

[0087] Because the policy model 128 of the present disclosure operates in the latent space 152, the quality of the latent space is important for the success of the disclosed methods. In particular, the continuity of the latent space is of importance as the reinforcement learning model is moving in the latent space. The continuity of the space can be reflected in the validity rate of the policy model 128, that is the percentage of latent vectors that are successfully decoded into valid molecular structures. VAEs though suffer from a phenomenon known as posterior collapse where the KL term in the lossfunction starts vanishing and the decoder model 154 generates molecules independently of the latent variables. See, for example, Yan et al., 2023, “Molecule Sequence Generation with Rebalanced Variational Autoencoder Loss,” Journal of Computational Biology 30(1), pp. 82-94, which is hereby incorporated by reference. When posterior collapse occurs, the VAE loses its reconstruction ability in the expense of a high validity rate.[0088J To mitigate posterior collapse, in some embodiments of the present disclosure the Beta-VAE generative model was used instead of the standard VAE generative model. Beta-VAE is described in Higgins et al., 2016, “Beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework,” in International Conference on Learning Representations, which is hereby incorporated by reference. The beta-VAE has a modified loss function with respect to the original VAE model. A multiplier (beta term) is introduced in the part of the loss function that corresponds to the KL divergence. The beta term contributes to achieving a trade-off between the reconstruction loss and the latent space representation that in turn affects the validity rate.

[0089] Moreover, as illustrated in Figure 16, to mitigate posterior collapse, in some embodiments of the present disclosure, a cyclical annealing schedule for the beta term for the Beta-VAE was used. See, for example, Yan et al., 2023, “Molecule Sequence Generation with Rebalanced Variational Autoencoder Loss,” Journal of Computational Biology 30(1), pp. 82-94. This is an improvement over the Beta-VAE where the beta term alternates between phases of gradual increase and phases of gradual decrease. The gradual periodic increase of the beta term prevents the vanishing of the KL divergence term.

[0090] Further still, to further mitigate posterior collapse, in some embodiments of the present disclosure, a maximum beta term for the Beta-VAE was used. See, for example, Bowman et al., 2015, “Generating Sentences from a Continuous Space,” in Conference on Computational Natural Language Learning, which is hereby incorporated by reference. In this variation of the cyclical annealing beta- VAE, the beta term is limited to a maximum value of 0.6 to further assist finding a trade-off between the reconstruction loss and the validity rate.

[0091] In some embodiments a word dropout mechanism was used as an additional training modification where randomly selected tokens were masked with a random probability during training time. This serves as a form of regularization during the VAE training and helps with preventing the decoder model 154 from overfitting as well as introducing diversity and novelty in the generated sequences.

[0092] Referring to block 202, in some embodiments, the first predetermined physical property is compound molecular weight.

[0093] Referring to block 204, in some embodiments, the first predetermined physical property is a compound distribution coefficient (logD), compound octanol-water partition coefficient (logP), penalized logP, compound acid dissociation constant (pKA), compound acid dissociation constant (kA), or compound topological polar surface area (TPSA).

[0094] LogD is a distribution constant of the lipophilicity of a compound. In some embodiments LogD is calculated from the chemical structure of a compound (cLogD). Generally, cLogD refers to the calculated distribution constant of the lipophilicity of a molecule. This is typically determined for a compound in an aqueous phase, which can be adjusted to a specific pH using a buffer. This parameter can serve as an estimate of a compound's overall lipophilicity, a value that influences the behavior of the compound in a range of biological processes relevant to a drug discovery, such as solubility, permeability through biological membranes, hepatic clearance, lack of selectivity and / or non-specific toxicity. In some instances, lipophilicity can be used as an estimate of drug likeness. See, Cambridge MedChem Consulting, 2019, “Lipophilicity,” available online at cambridgemedchemconsulting.com / resources / physiochem / logD.html, which is hereby incorporated herein by reference in its entirety.

[0095] The logP is a measure of compound hydrophilicity. The partition coefficient, abbreviated P, is defined as a particular ratio of the concentrations of a solute between the two solvents (a biphase of liquid phases), specifically for un-ionized solutes, and the logarithm of the ratio is thus log P. As used here, following convention, the numerator of this ratio is concentration of the compound in octanol and the denominator is concentration of the compound in water.

[0096] Penalized logP is a logP score that also accounts for ring size and synthetic accessibility. See, for example, Ertl, 2009, “Estimation of synthetic accessibility score of drug-like molecules,” J. Cheminform, which is hereby incorporated by reference.

[0097] In some embodiments the pKA is estimated using the structure of a molecule. See, for example, Seybold and Shields, 2015, “Computational estimation of pKa values,” WIREs Comput Mol Sci 5, pp. 290-297, which is hereby incorporated by reference.

[0098] In some embodiments permeability is obtained from a permeability prediction model trained on a dataset of 1550 datapoints curated from CHEMBL assays. Assay data was filtered for compoundswith assay standard type Papp for the MDCK cell line, apical to basolateral and basolateral to apical assay readouts were averaged and the readouts were normalized to the logarithm of 10A-6 cm / sec. Compounds with multiple readouts with standard deviation > 1 of the logarithm of 10A-6 cm / sec were discarded and compound SMILES were standardized to canonical smiles. The RDKIT 2D descriptors and a XGBOOST model were used with a random train / test split for the model validation. Figure 17 shows the predicted permeability of unseen compounds vs the experimental permeability. The orange shade shows the region of 3 -fold error.

[0099] Compound topological polar surface area (TPSA) makes use of functional group contributions based on a large database of structures, is a measure of the polar surface area of a compound that avoids the need to calculate compound three-dimensional structure or to decide which is the relevant biological conformation or conformations. See, for example, Prasanna and Doerksen et al., 2009, “Topological Polar Surface Area: A Useful Descriptor in 2D-QSAR,” Current Medicinal Chemistry 16(1), pp. 21-41, which is hereby incorporated by reference.

[0100] Referring to block 206, in some embodiments, the first predetermined physical property is drug likeness (QED score). QED scores are described in Bickerton et al., 2012, “Quantifying the chemical beauty of drugs,” Nature Chemistry 4(2), pp. 90-98, which is hereby incorporated by reference.

[0101] Referring to block 208, in some embodiments, the first predetermined physical property is permeability.

[0102] Referring to block 210, in some embodiments, the combinatorial synthesis library comprises 1,000 compounds, 5,000 compounds, 10,000 compounds, 50,000 compounds, 100,000 compounds, 1 million compounds, 10 million compounds, 100 million compounds, a billion compounds, or more.

[0103] Referring to block 212, a specification of a target for the first predetermined physical property is obtained. Figure 5 illustrates a reward function that is useful for property targeting. In property targeting, the goal is to generate compounds (molecules) for which a property of interest (e.g., predetermined physical property), or a combination of properties, takes values within a prespecified range.

[0104] In accordance with Table 1, for property targeting tasks, responsive to inputting specification of the target and a second cutoff into the policy model 128 (policy network), an action vector 137 of fixed dimension d is obtained as output from the policy model by application of the specification of thetarget, a first cutoff, and the second cutoff to the policy model in accordance with a first plurality of parameters 136 of the policy model. In some embodiments, the fixed dimension d is between 32 and 1024. In some embodiments, the fixed dimension dis between 48 and 640. In some embodiments, the fixed dimension d is 32, 64, 128, or 256.

[0105] In accordance with Table 1 and block 214, for constraint property optimization, responsive to inputting an initial state vector of fixed dimension d 140 that is a latent representation of a test compound into the policy model 128 (policy network), an action vector 137 of fixed dimension d is obtained as output from the policy model by application of the initial state vector to the policy model in accordance with a first plurality of parameters 136 of the policy model.

[0106] Referring to block 216, in some embodiments, the policy model is a multi-layer perceptron (MLP). Disclosure on suitable MLPs that can serve as the policy model 120 in some embodiments of the present disclosure is found in Vang-mata ed., 2020, Multilayer Perceptrons: Theory and Applications, Nova Science Publishers, Hauppauge, New York, which is hereby incorporated by reference. In some embodiments, the policy model is a 5-layer neural network with dropout layer as illustrated in Figure 12.

[0107] In some embodiments, the policy network (TT) (policy model) is a multi-layer perceptron (MLP) that is trained via a policy gradient algorithm. The policy network takes as input a state s and proposes an action a (action vector 137). In the case of property targeting, the state corresponds to the target values of the properties for which targeted molecule generation is being performed. In the constraint property optimization task, the state corresponds to the latent vector of the original molecule. The action vector a 137 is a distribution vector with the same dimensionality as the latent space 152 from which a latent vector z is being sampled. In property targeting, the action vector points to where in the latent space to sample to obtain the desirable molecular properties specified by the state vector. In constraint optimization, the action vector 137 is being interpreted as the perturbation that after being applied to a latent space vector 144 will result in an improvement of the property of interest. It can be seen as the movement in the latent space that takes us from an initial compound to a compound with improved properties.

[0108] Referring to block 218, in some embodiments, the policy model is a deep neural network (DNN). In some embodiments, a neural learning model (DNN) is a neural network comprising a plurality of hidden layers, e g., two or more hidden layers. Each layer of the neural network comprises a number of nodes (or “neurons”). A node receives input that comes either directly from the input dataor the output of nodes in previous layers, and performs a specific operation, e.g., a summation operation. In some embodiments, a connection from an input to a node is associated with a parameter (e.g., a weight and / or weighting factor). In some embodiments, the node sums up the products of all pairs of inputs, xi, and their associated parameters. In some embodiments, the weighted sum is offset with a bias, b. In some embodiments, the output of a node or neuron is gated using a threshold or activation function, f, which may be a linear or non-linear function. The activation function may be, for example, a rectified linear unit (ReLU) activation function, a Leaky ReLU activation function, or other function such as a saturating hyperbolic tangent, identity, binary step, logistic, arcTan, softsign, parametric rectified linear unit, exponential linear unit, softPlus, bent identity, softExponential, Sinusoid, Sine, Gaussian, or sigmoid function, or any combination thereof.

[0109] The weighting factors, bias values, and threshold values, or other computational parameters of the neural network, may be “taught” or “learned” in a training phase using one or more sets of training data. For example, the parameters may be trained using the input data from a training data set and a gradient descent or backward propagation method so that the output value(s) that the model computes is consistent with the examples included in the training data set. In some embodiments the parameters are refined using a back propagation neural network training process.

[0110] Referring to block 220, in some embodiments, the first plurality of parameters comprises 10,000 or more parameters, 100,000 or more parameters, or 1 x 106or more parameters. In some embodiments, the first plurality of parameters comprises 1000 model parameters. In some embodiments, the first plurality of parameters consists of between 10 and 10 million parameters. In some embodiments, the first plurality of parameters comprises at least 10, at least 50, at least 100, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 20,000, at least 50,000, at least 100,000, at least 200,000, at least 500,000, at least 1 million, at least 2 million, at least 3 million, at least 4 million or at least 5 million parameters. In some embodiments, the first plurality of parameters consists of no more than 8 million, no more than 5 million, no more than 4 million, no more than 1 million, no more than 500,000, no more than 100,000, no more than 50,000, no more than 10,000, no more than 5000, no more than 1000, or no more than 500 parameters. In some embodiments, the first plurality of parameters consists of from 10 to 5000, from 500 to 10,000, from 10,000 to 500,000, from 20,000 to 1 million, or from 1 million to 5 million parameters. In some embodiments, the first plurality of parameters falls within another range starting no lower than 10 parameters and ending no higher than 8 million parameters.

[0111] Referring to block 222, in some embodiments, the test compound is an organic compound having a molecular weight of less than 2000 Daltons (Da). In some embodiments, the test compound has a molecular weight of at least 10 Da, at least 20 Da, at least 50 Da, at least 100 Da, at least 200 Da, at least 500 Da, at least 1 kDa, at least 2 kDa, at least 3 kDa, at least 5 kDa, at least 10 kDa, at least 20 kDa, at least 30 kDa, at least 50 kDa, at least 100 kDa, or at least 500 kDa. In some embodiments, the test compound has a molecular weight of no more than 1000 kDa, no more than 500 kDa, no more than 100 kDa, no more than 50 kDa, no more than 10 kDa, no more than 5 kDa, no more than 2 kDa, no more than 1 kDa, no more than 500 Da, no more than 300 Da, no more than 100 Da, or no more than 50 Da. In some embodiments, the test compound has a molecular weight of from 10 Da to 900 Da, from 50 Da to 1000 Da, from 100 Da to 2000 Da, from 1 kDa to 10 kDa, from 5 kDa to 500 kDa, or from 100 kDa to 1000 kDa. In some embodiments, the test compound has a molecular weight that falls within another range starting no lower than 10 Daltons and ending no higher than 1000 kDa.

[0112] In some embodiments, a test compound is a small molecule. For instance, in some embodiments, a test chemical compound is an organic compound having a molecular weight of less than approximately 1000 Daltons (e.g., less than 900 Daltons).

[0113] Referring to block 224, in some embodiments, the test compound is an organic compound that satisfies each of the Lipinski rule of five criteria. Referring to block 226, in some embodiments, the test compound is an organic compound that satisfies at least three criteria of the Lipinski rule of five criteria. In some embodiments, the test compound satisfies any two or more rules, any three or more rules, or all four rules of the Lipinski's rule of Five. The Lipinski’s rule of five (e.g., RO5) criteria are a set of guidelines used to evaluate drug likeness, such as to determine whether a respective compound with a respective pharmacological or biological activity has corresponding chemical or physical properties suitable for administration in humans. Lipinski’s rule of five includes the following criteria for determining the drug likeness of a compound: (i) not more than five hydrogen bond donors, (ii) not more than ten hydrogen bond acceptors, (iii) a molecular weight under 500 Daltons, and (iv) a LogP under 5. See, Lipinski, 1997, Adv. Drug Del. Rev. 23, 3, which is hereby incorporated herein by reference in its entirety. In some embodiments, the test compound satisfies one or more criteria in addition to Lipinski's Rule of Five. For example, in some embodiments, the test compound has five or fewer aromatic rings, four or fewer aromatic rings, three or fewer aromatic rings, or two or fewer aromatic rings. In some embodiments, the test compound is an organic compound that satisfies at least two, three or four criteria of the Lipinski rule of five criteria. In some embodiments,the test compound is an organic compound that satisfies zero, one, two, three, or all four criteria of the Lipinski rule of five criteria.

[0114] Referring to block 228, in some embodiments, the initial state vector of fixed dimension d of the test compound is obtained as output from an encoder model 148, comprising a third plurality of parameters, by inputting a string representation of the test compound into the encoder model.

[0115] Referring to block 230, in some embodiments, the third plurality of parameters comprises 10,000 or more parameters, 100,000 or more parameters, or 1 x 106or more parameters. In some embodiments, the third plurality of parameters comprises 10,000 or more parameters, 100,000 or more parameters, or 1 x 106or more parameters. In some embodiments, the third plurality of parameters comprises 1000 model parameters. In some embodiments, the third plurality of parameters consists of between 10 and 10 million parameters. In some embodiments, the third plurality of parameters comprises at least 10, at least 50, at least 100, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 20,000, at least 50,000, at least 100,000, at least 200,000, at least 500,000, at least 1 million, at least 2 million, at least 3 million, at least 4 million or at least 5 million parameters. In some embodiments, the third plurality of parameters consists of no more than 8 million, no more than 5 million, no more than 4 million, no more than 1 million, no more than 500,000, no more than 100,000, no more than 50,000, no more than 10,000, no more than 5000, no more than 1000, or no more than 500 parameters. In some embodiments, the third plurality of parameters consists of from 10 to 5000, from 500 to 10,000, from 10,000 to 500,000, from 20,000 to 1 million, or from 1 million to 5 million parameters. In some embodiments, the third plurality of parameters falls within another range starting no lower than 10 parameters and ending no higher than 8 million parameters.

[0116] Referring to block 232, in some embodiments, the encoder model is a first recurrent neural network and the decoder model is a second recurrent neural network. A recurrent neural network (RNN) is any network whose neurons send feedback signals to each other. See, Grossberg, 1023, “Recurrent neural networks,” Scholarpedia 8(2), p. 1888, which is hereby incorporated by reference.

[0117] Referring to block 234, in some embodiments, the encoder model is a first neural network and the decoder model is a second neural network. Neural networks are described in the definitions section above.

[0118] Referring to block 236, in some embodiments, the string representation is in a SMARTS (SMARTS - A Language for Describing Molecular Patterns,” 2022 on the Internet at daylight.com / dayhtml / doc / theory / theory.smarts.html (accessed Dec. 2020), DeepSMILES (O’Boyleand Dalke, 2018, “DeepSMILES: an adaptation of SMILES for use in machine-learning of chemical structures,” Preprint at ChemRxiv. https: / / doi.org / 10.26434 / chemrxiv.7097960.vl.), self-referencing embedded string (SELFIES) (Krenn et al., 2022, “SELFIES and the future of molecular string representations,” Patterns 3(10), pp. 1-27), or simplified molecular-input line-entry system (SMILES) format. Molecular fingerprinting using SMILES strings is described, for example, in Honda et al., 2019, “SMILES Transformer: Pre-trained Molecular Fingerprint for Low Data Drug Discovery,” arXiv: 1911.04738, which is hereby incorporated herein by reference in its entirety.

[0119] Referring to block 238, the action vector is sampled on a probabilistic basis to identify a plurality of points (latent vectors) 144 in the latent space 150. Given that the latent space is probabilistic, the action vector is also modeled as a distribution. Each point 144 in the plurality of points (latent vectors) is of fixed dimension d.

[0120] Referring to block 240, in some embodiments, the sampling samples the action vector as a multivariate normal distribution with a mean equal to the action vector and a standard deviation equal to a constant value. That is, the action vector is sampled from a multivariate Normal distribution with a mean equal to the action a and a standard deviation a equal to a constant value. a = 7r(s) d ~ N(a, cr) z' = z + dz where a is the action vector 137, 7r(s) is the output of policy model 128 n upon inputting the test compound string representation 138 into the policy model 128, tZ is a point 144 that is acquired by sampling the action vector 137 a from normal distribution N with a mean equal to the action vector 137 and a standard deviation o equal to a constant value.

[0121] To ensure that the perturbation will not result to a completely different molecular structure and instead it will be a structural modification to the input state molecule, the action vector is regularized in some embodiment. Essentially, the regularization of the action vector is used in such embodiments to control the magnitude of the perturbation or otherwise the move in the latent space.

[0122] It should be noted that in the case of targeted sampling, during training, the state vector is not constant. The values are being sampled at random from a uniform distribution such that the model will be trained to generate molecules for any given state vector instead of generating molecules for a prespecified property value or combination of property values. In constrained optimization, the policynetwork is trained to shift a property of interest towards a specific direction that is either increase or decrease.

[0123] In some embodiments, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more points (latent vectors) 144 are sampled from the action vector. In some embodiments, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150 or more points (latent vectors) 144 are sampled from the action vector. In some embodiments, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500 or more points (latent vectors) 144 are sampled from the action vector.

[0124] Referring to block 242, each point (latent vector) 144 in the plurality of points is inputted into a decoder model 154 thereby obtaining a corresponding structure of each candidate compound 146 in a plurality of candidate compounds in accordance with a second plurality of parameters associated with the decoder model, each candidate compound 146 in the plurality of candidate compounds corresponding to a point 144 in the plurality of points. This decoding is illustrated on the right hand side of Figure 4.

[0125] Referring to block 244, in some embodiments, the plurality of candidate compounds comprises 5 or more chemical compounds. In some embodiments, the plurality of candidate compounds comprises 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150 chemical compounds, where each such chemical compound corresponds to a point (latent vector) 144 in the plurality of points. In some embodiments, the plurality of candidate compounds comprises 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150 or more chemical compounds, where each such chemical compound corresponds to a point (latent vector) 144 in the plurality of points. In some embodiments, the plurality of candidate compounds comprises 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500 or more chemical compounds, where each such chemical compound corresponds to a point (latent vector) 144 in the plurality of points.

[0126] Referring to block 246, in some embodiments, each candidate compound in the plurality of candidate compounds is an organic compound having a molecular weight of less than 2000 Daltons. In some embodiments, each candidate compound in the plurality of candidate compounds is an organic compound having a molecular weight of less than 2000 Daltons (Da). In some embodiments, each candidate compound in the plurality of candidate compounds is an organic compound having a molecular weight of at least 10 Da, at least 20 Da, at least 50 Da, at least 100 Da, at least 200 Da, at least 500 Da, at least 1 kDa, at least 2 kDa, at least 3 kDa, at least 5 kDa, at least 10 kDa, at least 20 kDa, at least 30 kDa, at least 50 kDa, at least 100 kDa, or at least 500 kDa. In some embodiments,each candidate compound in the plurality of candidate compounds is an organic compound having a molecular weight of no more than 1000 kDa, no more than 500 kDa, no more than 100 kDa, no more than 50 kDa, no more than 10 kDa, no more than 5 kDa, no more than 2 kDa, no more than 1 kDa, no more than 500 Da, no more than 300 Da, no more than 100 Da, or no more than 50 Da. In some embodiments, each candidate compound in the plurality of candidate compounds has a molecular weight of from 10 Da to 900 Da, from 50 Da to 1000 Da, from 100 Da to 2000 Da, from 1 kDa to 10 kDa, from 5 kDa to 500 kDa, or from 100 kDa to 1000 kDa. In some embodiments, each candidate compound in the plurality of candidate compounds has a molecular weight that falls within another range starting no lower than 10 Daltons and ending no higher than 1000 kDa.

[0127] In some embodiments, each candidate compound 146 in the plurality of candidate compounds is a small molecule. For instance, in some embodiments, each candidate compound in the plurality of candidate compounds is an organic compound having a molecular weight of less than approximately 1000 Daltons (e.g., less than 900 Daltons).

[0128] Referring to block 248, in some embodiments, each candidate compound in the plurality of candidate compounds is an organic compound that satisfies each of the Lipinski rule of five criteria. Referring to block 250, in some embodiments, each candidate compound in the plurality of candidate compounds is an organic compound that satisfies at least three criteria of the Lipinski rule of five criteria. In some embodiments, each candidate compound in the plurality of candidate compounds satisfies any two or more rules, any three or more rules, or all four rules of the Lipinski's rule of Five. In some embodiments, each candidate compound in the plurality of candidate compounds satisfies one or more criteria in addition to Lipinski's Rule of Five. For example, in some embodiments, each candidate compound in the plurality of candidate compounds has five or fewer aromatic rings, four or fewer aromatic rings, three or fewer aromatic rings, or two or fewer aromatic rings. In some embodiments, each candidate compound in the plurality of candidate compounds is an organic compound that satisfies at least two, three or four criteria of the Lipinski rule of five criteria. In some embodiments, each candidate compound in the plurality of candidate compounds is an organic compound that satisfies zero, one, two, three, or all four criteria of the Lipinski rule of five criteria.

[0129] Referring to block 252, in some embodiments, the second plurality of parameters comprises 10,000 or more parameters, 100,000 or more parameters, or 1 x 106or more parameters. In some embodiments, the second plurality of parameters comprises 1000 parameters. In some embodiments, the second plurality of parameters consists of between 10 and 10 million parameters. In someembodiments, the second plurality of parameters comprises at least 10, at least 50, at least 100, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 20,000, at least 50,000, at least 100,000, at least 200,000, at least 500,000, at least 1 million, at least 2 million, at least 3 million, at least 4 million or at least 5 million parameters. In some embodiments, the second plurality of parameters consists of no more than 8 million, no more than 5 million, no more than 4 million, no more than 1 million, no more than 500,000, no more than 100,000, no more than 50,000, no more than 10,000, no more than 5000, no more than 1000, or no more than 500 parameters. In some embodiments, the second plurality of parameters consists of from 10 to 5000, from 500 to 10,000, from 10,000 to 500,000, from 20,000 to 1 million, or from 1 million to 5 million parameters. In some embodiments, the second plurality of parameters falls within another range starting no lower than 10 parameters and ending no higher than 8 million parameters.

[0130] Referring to block 254, a value is determined for the first predetermined physical property for each candidate compound in the plurality of candidate compounds using the corresponding structure of the candidate compound. Methods for calculating the value for a first predetermined physical property of a compound using the corresponding structure of the compound have been described with specific physical properties in blocks 202 through 208 above.

[0131] Referring to block 256, an update to the first plurality of parameters 136 (of the policy model 128) is acquired through application of a training loss function 132 to each candidate compound in the plurality of candidate compounds. The training loss function 132 includes a reward term 134, a magnitude of which is determined, at least in part, by an application of the target, the first cutoff and the second cutoff of the first predetermined physical property to the value for the first predetermined physical property for a candidate compound 146 in the plurality of candidate compounds. In so doing, the reward term 134 evaluates whether an action vector 137, proposed by the policy model 128, is indeed improving the property of interest.

[0132] The purpose of the reward term is to provide feedback to the policy network during training on whether the actions proposed by the network are successful for the task at hand or not. The reward function has a different shape for the two examined tasks, property targeting and constrained optimization, as it is explained in the following.

[0133] Through the reward term 134, the value for the first predetermined physical property for the test compound 138 is compared to the value for the first predetermined physical property of the candidate compounds 146 to determine whether the structural modifications from the proposed actionvector 137 are contributing positively or negatively to improving the first predetermined physical property. In some embodiments, the reward for an action vector 137 is positive if the action vector 137 results in structural modifications that improve the first predetermined physical property and are negative otherwise.

[0134] Figure 5 illustrates a reward function that is useful for property targeting. In property targeting, the goal is to generate compounds (molecules) for which a property of interest (e.g., predetermined physical property), or a combination of properties, takes values within a prespecified range. For instance, consider the case where the predetermined physical property is molecular weight. In this instance, the target value for molecular weight is obtained and this target value represented in Figure 5 at x=0 on the x-axis. The first cutoff is represented ci in Figure 5 while the second cutoff is represented by C2 in Figure 5. The Y-axis shows the value of a reward as a function of the value of the predetermined physical property. Thus, a target value p is specified for the property of interest and two cutoff values ci and C2 such that candidate compounds 146 with property values within [p-ci, p+ci] are deemed as successful and take full reward while candidate compounds 146 with values greater than p+c2 or less than p-C2 are not successful and take the minimum reward. Candidate compounds 146 with property values within [ci, C2] or -[ci, C2] are assigned a reward that drops linearly from the maximum to the minimum value as shown in Figure 5. If more than one property is targeted simultaneously, then the reward for each property is calculated and the individual rewards are aggregated by averaging them in some embodiments.

[0135] In some embodiment the ci property cutoff is dynamic during training to mitigate a sparse reward that will lead to unstable training. More specifically, there is higher tolerance in deviations from the target value when it is challenging for the policy network to find successful action vectors such as at the early stages of training. In some embodiments, the cutoff C2 remains fixed throughout training and equal to 1.

[0136] In some embodiments the property cutoff ci is calculated as follows: if the target value for the property of interest is ptand the property values of the generated molecules are pgthen we get the deviation Ap = ||pt— pg\\ and then we compute ci as the 0.3 quantile of Ap.

[0137] Figure 6 illustrates a reward function for constraint optimization. For this task the goal is to generate candidate compounds 146 with improved values for a predetermined physical property while maintaining molecular similarity with the test compound 138 compound. Without loss of generality, in the following it is assumed that the task is to increase the values of the predetermined physicalproperty. The reward function in accordance with Figure 6 is comprised of two components: a property reward function evaluating the property improvement and a similarity reward function assessing whether molecular similarity with the test compound 138 is being maintained. For a pair of compounds (min, mout) with property values (pin, pout) corresponding to the input and output of the reinforcement learning model (policy model 128), the property reward is defined as f 0, if Pout Pin < 0Rproperty = ] Pout ~ Pin' if 0 < pout ~ Pin < Max reward t Pout - Pin' if0< Pout - Pin < Max reward for a predefined Max reward value. With that reward shape, the agent is being rewarded for identifying actions that result in greater property improvements while improvements that exceed a prespecified threshold are not further rewarded. The similarity reward corresponds to a binary function which takes the value 1 if the similarity constraint is being satisfied, that is if the Tanimoto coefficient between the starting molecule and the new molecule is higher than 8, and 0 otherwise. The overall reward is obtained by multiplying the two rewards.

[0138] Referring to block 258, and with reference to Figure 5, in some embodiments, the application of the target, the first cutoff and second cutoff of the first predetermined physical property to the value for the metric of the first predetermined physical property for the candidate compound to determine a magnitude of the reward term comprises using a maximum reward for the first predetermined physical property in the loss function when the value for the metric of the first predetermined physical property for the candidate compound is both (1) greater than the target value minus the first cutoff value for the first predetermined physical property and (2) less than the target value plus the first cutoff value for the first predetermined physical property. This corresponds to the region denoted by box 502 in Figure 5.

[0139] Further, a minimum reward is used for the loss function for the first predetermined physical property when the value for the metric of the first predetermined physical property for the candidate chemical compound is less than the target value minus the second cutoff value for the first predetermined physical property or greater than the target value plus the second cutoff value for the first predetermined physical property. This corresponds to either region denoted by a box 504 in Figure 5.

[0140] Further a value is used in the loss function for the first predetermined physical property that is a linear function of the value for the metric of the first predetermined physical property for the candidate chemical compound when the value for the metric of the first predetermined physicalproperty for the candidate chemical compound fails to qualify for the maximum reward or the minimum reward as detailed above. That is, the linear function is used when the first predetermined physical property does not fall into boxes 502 or 504 in Figure 5. In such instance, the linear function produces a reward value between the minimum reward and the maximum reward as illustrated in Figure 5.

[0141] Referring to block 260, in some embodiments, the one or more predetermined physical properties is a plurality of physical properties. For instance, in some embodiments, the one or more predetermined physical properties is 2, 3, 4, 5, 6, 7, 8, 9, or 10 physical properties. In such embodiments, the plurality of compounds each have a value for a second predetermined physical property in the plurality of physical properties that is in a second specified range. Moreover, the obtaining block 212 further comprises obtaining a specification of a target for the second predetermined physical property. Further still, the determining a value block 254 further determines a value for the second predetermined physical property for each candidate compound 146 in the plurality of candidate compounds using the corresponding structure of the candidate compound, and the magnitude of the reward term is further determined by an application of the target, the first cutoff and the second cutoff of the second predetermined physical property to the value for the second predetermined physical property for the candidate compound. For instance, referring to block 262, in some embodiments, the magnitude of the reward function is an average or a minimum of (i) the application of the target, the first cutoff and the second cutoff of the first predetermined physical property to the value for the first predetermined physical property for a candidate compound in the plurality of candidate compounds, and (ii) the application of the target, the first cutoff and the second cutoff of the second predetermined physical property to the value for the second predetermined physical property for the candidate compound. In other words, in some embodiments the reward term is an average or minimum of (i) the reward value calculated for the first predetermined property in accordance with Figure 5 and (ii) the reward value calculated for the second predetermined property in accordance with Figure 5. In instances where additional properties are simultaneously used, the reward term is an average of the additional reward terms for these additional predetermined property as well (or a minimum of all the reward terms).

[0142] Referring to block 264, the policy model 129 is updated with the update to the first plurality of parameters 136 thereby training the policy model to sample the latent space 152 to identify the subset of the latent space comprising the plurality of compounds that have the value for the first predetermined physical property that is in the first specified range.

[0143] In accordance with Figure 5, in some embodiments, the training loss function, including the reward term 134, is used to train the policy model 128 using the REINFORCE algorithm with the loss function 132 being equal to the negative value of the expected reward.where 9 denotes the parameters 136 of the policy model n (128) and r(s, a) denotes the reward of taking action (a) (active vector 137) given the state (s) (test compound string representation 138. In such embodiments, the objective is to maximize the expectation of the rewards achieved by the policy model 128 with parameters 9. See, Williams, 1992, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning 8(3), pp. 229-256, which is hereby incorporated by reference. The intention of this objective is to increase the probability of actions that result in high rewards and decrease the probability of non rewarding actions. The parameter 9 (parameters 136) is updated via gradient ascent of the objective, owhere rj is the learning rate, Pg(s, a) is the probability of the policy model 128 with parameters 9 (parameters 136) taking action (a) (action vector 137) given state (s) (test compound string representation 138).

[0144] In some embodiments of the constraint property optimization task (Figure 6), the loss function 132 includes an additional term which is a regularization term on the action vector 137.L — — Ue+ A ■ ||a|| 2

[0145] This additional term is introduced to control the magnitude of the step that the policy 130 is taking away from the test compound 138. Essentially, the regularization is modeling the constraint in the optimization task as the regularization on the action vector 137 means that the latent vectors (points 144) of the generated candidate compounds 146 are in close proximity to those of the original molecules (e.g., test compound 138). It is noted that the molecular similarity constraint in the constraint optimization task is being achieved by both the similarity reward function, and the regularization term in the loss function.

[0146] In some embodiments, the choice of X and the regularization function are specified as follows. The intention behind introducing this term is to ensure that the perturbations suggested by the policy network correspond to small structural changes in the input molecule. The choice of X depends on the similarity constraint 8. We recall here that we are seeking to generate molecules with improved properties and Tanimoto similarity with the starting molecule greater than 8. For higher 8 values, we impose stronger regularization by selecting higher X values. More specifically we set the X value as follows: t <5, if 8 > 0 (0.25, if 6 = 0 which means that even when there is no specific requirement for the Tanimoto similarity 8, regularization is applied such that the perturbations applied to the starting latent vector are not resulting very far from the starting latent vector. Regarding the type of regularization, L2 regularization is being used in the example section below but either Li or L2 regularization can be used.

[0147] While the methods of Figure 2 have been described with respect to a single test compound 138, in practice multiple test compounds can be used to train the policy model 128. For each such test compound, the methods disclosed in Figure 2 are performed. In some embodiments, 10 or more test compounds, 100 or more test compounds, 1000 or more test compounds 10,000 or more test compounds, 100,000 or more test compounds or 1 x 106or more test compounds are used to train the policy model to identify candidate compounds 146 that have one or more desired physical properties, with or without the additional constraint of being structurally similar to the test compounds.

[0148] Examples

[0149] Example 1 - Single property targeting.

[0150] The goal of example 1 was to generate molecules with values within a range of a target value for a property of interest. In this example, a desirable value for the molecular weight (MW) was set and the latent reinforcement learning model 128 (Figures 1A, IB, and 4) was used to generate candidate compounds 146 with molecular weights around the target value. Figure 7 shows the distribution of the molecular weight of compounds that have been generated with and without using reinforcement learning. The blue curve corresponds to the distribution of the molecular weight of compounds that have been sampled from the latent space 152 without performing any optimization. The orange and green curves correspond to the distributions of the molecular weight of molecules thathave been generated by the latent reinforcement learning model (Figures 1A, IB, and 4) after setting as target value for the molecular weight at 250 Da and 365 Da, respectively. As can be seen in Figure 7, the distribution of the molecular weight has been shifted when generating molecules using the latent reinforcement learning model with respect to the original distribution.

[0151] Example 2 - Multi-property targeting.

[0152] In example 2, the reinforcement learning model 128 (Figures 1 A, IB, and 4) was trained to generate molecules with a specific combination of molecular weight and LogP. Figures 8A and 8B illustrate the distributions of the two properties for candidate compounds 146 that have been sampled without optimization (blue curves) and molecules that have been generated by the reinforcement learning model 128 (Figures 1A, IB, and 4) (orange curves) for four different combinations of molecular weight and LogP values.

[0153] It should be noted that targeting of two or more two properties simultaneously is especially challenging. A shift in the distribution of one molecular property will affect the distributions of other properties as well. If the combination of target values in multi-property targeting is underrepresented in the set of molecules that has been used for pretraining the variational autoencoder, then training the latent reinforcement learning method will be challenging. In that case, there will not be enough positive examples to reward the policy network and update its weights.

[0154] Example 3 - Constraint property optimization

[0155] In example 3, for each input test compound, the generated SMILES (candidate compounds 146) at the final iteration step for the last 20 epochs was saved. This means that for each test compound at most 20 candidate compounds are generated, considering that a SMILES string may not be a valid compounds, or a candidate compound may be generated more than once. Figures 9A and 9B report the success rate, that is the percentage of test compounds for which there is at least one candidate compound 146 among the generated candidate compounds with an improved property value and satisfying the similarity constraint. For the succeeded cases, the average improvement and the average Tanimoto similarity between the candidate compound with the highest improvement and the starting test compound is reported. The diversity of the succeeded candidate compounds, that is the number of distinct compounds that have succeeded per input test compound is additionally evaluated, with success being defined as a modification that results in a property improvement while the similarity constraint is satisfied. The validity rate, that is the percentage of generated SMILES(candidate compounds) that correspond to valid chemical compounds among all generated SMILES isalso reported. Finally, an assessment is made as to whether the structural modifications suggested by the policy model 128 are distorting the molecular structures. More specifically, the average relative change in the QED score as well as the average relative change in the synthetic accessibility score is reported. See, Ertl and Schuffenhauer, 2009, “Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions,” Journal of Cheminformatics 1(8), which is hereby incorporated by reference. Figure 13 illustrate example candidate compounds for QED, penalized LogP, and permeability, and the corresponding test compound that were found in example 3.

[0156] Higher QED scores correspond to compounds with more drug-like properties while for the synthetic accessibility score, a lower score is associated with better synthesizability. To assess the difficulty of the optimization problem for each property, the following was performed: for each input molecule (target molecule) 20 molecules (candidate molecules) were sampled by randomly perturbing (adding gaussian noise) the latent representation of each starting molecule and the success rate and average improvement was recorded. The purpose of this was to understand how well the latent reinforcement learning model (policy model 128) performed when compared to random sampling of molecules in the neighborhood of the starting molecule. The results are provided in Figures 10A, 10B, 10C, and 10D.

[0157] Example 4 - Comparison with existing methodologies for molecule optimization.

[0158] A common benchmark has been used in the literature for evaluating tools for molecule optimization. This benchmark evaluates the tools on the task of optimizing a common set of 800 molecules for plogP. The starting molecules have been derived from the ZINC database (see, Irwin et al., 2020, “ZINC20 — A Free Ultralarge-Scale Chemical Database for Ligand Discovery,” J. Chem. Inf. Model 60(12), 6065-6073, which is hereby incorporated by reference) and they have been selected such that they have a low plogP value. The evaluation metrics that have been used include success rate, and average improvement and average similarity among the succeeded cases.

[0159] In this example, as illustrated in Figured 14A and 14B, this benchmark was performed to compare the disclosed latent reinforcement learning method against two existing molecule optimization tools, GCPN (You etal., 2018, “Graph Convolutional Policy Network for Goal-Directed Molecular Graph Generation,” in Advances in Neural Information Processing Systems, 2018) and MolDQN (Zhou et al., 2019, “Optimization of Molecules via Deep Reinforcement Learning,” Scientific Reports 9(1)). Some additional metrics have been introduced in order to assess whether thestructural modifications suggested by the models may be distorting the molecules. In particular, the change in the QED score as well as the change in the synthetic accessibility score was examined to assess whether the molecules drug-like properties or synthesizability may get compromised. The results of this benchmark (top 20 prediction, Tanimoto similarity greater than 0.4) are detailed in Figure 11. As is seen in Figure 11, the disclosed latent reinforcement learning model achieved similar pLogP improvement without compromising drug-likeness (QED) and synthesizability (SA).

[0160] Example 5 - Figures 15A and 15B illustrate how the pretrained reinforcement learning model is able to improve pLogP on unseen compounds.

[0161] CONCLUSION

[0162] Plural instances may be provided for components, operations or structures described herein as a single instance. Finally, boundaries between various components, operations, and data stores are somewhat arbitrary, and particular operations are illustrated in the context of specific illustrative configurations. Other allocations of functionality are envisioned and may fall within the scope of the implementation(s). In general, structures and functionality presented as separate components in the example configurations may be implemented as a combined structure or component. Similarly, structures and functionality presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the implementation(s).

[0163] It will also be understood that, although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first subject could be termed a second subject, and, similarly, a second subject could be termed a first subject, without departing from the scope of the present disclosure. The first subject and the second subject are both subjects, but they are not the same subject.

[0164] The terminology used in the present disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used in the description of the invention and the appended claims, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps,operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0165] As used herein, the term “if’ may be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” depending on the context. Similarly, the phrase “if it is determined” or “if [a stated condition or event] is detected” may be construed to mean “upon determining” or “in response to determining” or “upon detecting (the stated condition or event )” or “in response to detecting (the stated condition or event),” depending on the context.

[0166] The foregoing description included example systems, methods, techniques, instruction sequences, and computing machine program products that embody illustrative implementations. For purposes of explanation, numerous specific details were set forth in order to provide an understanding of various implementations of the inventive subject matter. It will be evident, however, to those skilled in the art that implementations of the inventive subject matter may be practiced without these specific details. In general, well-known instruction instances, protocols, structures and techniques have not been shown in detail.

[0167] The foregoing description, for purposes of explanation, has been described with reference to specific implementations. However, the illustrative discussions above are not intended to be exhaustive or to limit the implementations to the precise forms disclosed. Many modifications and variations are possible in view of the above teachings. The implementations were chosen and described in order to best explain the principles and their practical applications, to thereby enable others skilled in the art to best utilize the implementations and various implementations with various modifications as are suited to the particular use contemplated.

Claims

What is claimed is:

1. A method of training a policy model to identify a subset of a latent space comprising a plurality of compounds that each have a value for at least a first predetermined physical property, in one or more predetermined physical properties, that is in a first specified range, wherein the latent space encodes each compound in a combinatorial synthesis library as an undercomplete vector of fixed dimension d in which each entry in the undercomplete vector is a distribution, the method comprising: at a computer system comprising at least one processor and a memory storing at least one program for execution by the at least one processor, the at least one program comprising instructions for:A) obtaining a specification of a target and a second cutoff for the first predetermined physical property;B) responsive to inputting the specification of the target and a second cutoff for the first predetermined physical property into the policy model, obtaining an action vector of fixed dimension d as output from the policy model by application of the specification of the target, a first cutoff, and the second cutoff for the first predetermined physical property to the policy model in accordance with a first plurality of parameters of the policy model;C) sampling the action vector on a probabilistic basis to identify a plurality of points in the latent space, wherein each point in the plurality of points is of fixed dimension tZ;D) inputting each point in the plurality of points into a decoder model thereby obtaining a corresponding structure of each candidate compound in a plurality of candidate compounds in accordance with a second plurality of parameters associated with the decoder model, each candidate compound in the plurality of candidate compounds corresponding to a point in the plurality of points;E) determining a value for the first predetermined physical property for each candidate compound in the plurality of candidate compounds using the corresponding structure of the candidate compound;F) acquiring an update to the first plurality of parameters, through application of a loss function to each candidate compound in the plurality of candidate compounds, wherein the loss function includes a reward term, a magnitude of which is determined, at least in part, by an application of the target, the first cutoff and the second cutoff of the first predetermined physical property to the value for the first predetermined physical property for a candidate compound in the plurality of candidate compounds; andG) updating the policy model with the update to the first plurality of parameters thereby training the policy model to sample the latent space to identify the subset of the latent space comprising the plurality of compounds that have the value for the first predetermined physical property that is in the first specified range.

2. The method of claim 1, wherein the first predetermined physical property is compound molecular weight.

3. The method of claim 1, wherein the first predetermined physical property is compound distribution coefficient (logD).

4. The method of claim 1, wherein the first predetermined physical property is a compound octanolwater partition coefficient (logP).

5. The method of claim 1, wherein the first predetermined physical property is a penalized logP.

6. The method of claim 1, wherein the first predetermined physical property is a compound acid dissociation constant (pKA), a compound acid dissociation constant (kA), or a compound topological polar surface area (TPSA).

7. The method of claim 1, wherein the first predetermined physical property is a drug likeness (QED score).

8. The method of claim 1, wherein the first predetermined physical property is a permeability.

9. The method of any one of claims 1-8, wherein the combinatorial synthesis library comprises 5,000 compounds.

10. The method of any one of claims 1-8, wherein the combinatorial synthesis library comprises 50,000 compounds.

11. The method of any one of claims 1-8, wherein the combinatorial synthesis library comprises 1 million compounds.

12. The method of any one of claims 1-11, wherein the policy model is a multi-layer perceptron.

13. The method of any one of claims 1-11, wherein the policy model is a deep neural network.

14. The method of any one of claims 1-13, wherein the first plurality of parameters comprises 10,000 or more parameters.

15. The method of any one of claims 1-13, wherein the first plurality of parameters comprises 100,000 or more parameters.

16. The method of any one of claims 1-13, wherein the first plurality of parameters comprises 1 x 106or more parameters.

17. The method of any one of claims 1-16, wherein the sampling C) samples the action vector as a multivariate normal distribution with a mean equal to the action vector and a standard deviation equal to a constant value.

18. The method of any one of claims 1-17, wherein the plurality of candidate compounds comprises 5 or more chemical compounds.

19. The method of any one of claims 1-18, wherein each candidate compound in the plurality of candidate compounds is an organic compound having a molecular weight of less than 2000 Daltons.

20. The method of any one of claims 1-18, wherein each candidate compound in the plurality of candidate compounds is an organic compound that satisfies each of the Lipinski rule of five criteria.

21. The method of any one of claims 1-18, wherein each candidate compound in the plurality of candidate compounds is an organic compound that satisfies at least three criteria of the Lipinski rule of five criteria.

22. The method of any one of claims 1-21, wherein the application of the target, the first cutoff and second cutoff of the first predetermined physical property to determine the magnitude of the reward term comprises:(a) using a maximum reward for the first predetermined physical property in the loss function when the value for the metric of the first predetermined physical property for the candidate compound is both (1) greater than the target value minus the first cutoff value for the first predetermined physical property and (2) less than the target value plus the first cutoff value for the first predetermined physical property,(b) using a minimum reward in the loss function for the first predetermined physical property when the value for the metric of the first predetermined physical property for the candidate chemical compound is less than the target value minus the second cutoff value for the first predetermined physical property or greater than the target value plus the second cutoff value for the first predetermined physical property, and(c) using a value in the loss function for the first predetermined physical property that is a linear function of the value for the metric of the first predetermined physical property for the candidate chemical compound that produces a reward value between the minimum reward and the maximum reward when the value for the metric of the first predetermined physical property for the candidate chemical compound fails to qualify for the maximum reward or the minimum reward.

23. The method of any one of claims 1-22, wherein the one or more predetermined physical properties is a plurality of physical properties, the plurality of candidate compounds each have a value for a second predetermined physical property in the plurality of physical properties that is in a second specified range, the obtaining A) further comprises obtaining a specification of a target for the second predetermined physical property, the determining E) further determines a value for the second predetermined physical property for each candidate compound in the plurality of candidate compounds using the corresponding structure of the candidate compound, and the magnitude of the reward function is further determined by an application of the target, the first cutoff and the second cutoff of the second predetermined physical property to the value for the second predetermined physical property for the candidate compound.

24. The method of claim 23, wherein the magnitude of the reward function is an average or minimum of the application of the target, the first cutoff and the second cutoff of the first predetermined physical property to the value for the first predetermined physical property for a candidate compound in the plurality of candidate compounds, and the application of the target, the first cutoff and the second cutoff of the second predetermined physical property to the value for the second predetermined physical property for the candidate compound.

25. The method of any one of claims 1-24, wherein the second plurality of parameters comprises 10,000 or more parameters.

26. The method of any one of claims 1-24, wherein the second plurality of parameters comprises 100,000 or more parameters.

27. The method of any one of claims 1-24, wherein the second plurality of parameters comprises 1 x 106or more parameters.

28. A computer system, comprising one or more processors and memory, the memory storing instructions for performing a method of training a policy model to identify a subset of a latent space comprising a plurality of compounds that each have a value for at least a first predetermined physical property, in one or more predetermined physical properties, that is in a first specified range, wherein the latent space encodes each compound in a combinatorial synthesis library as an undercomplete vector of fixed dimension d in which each entry in the undercomplete vector is a distribution, the method comprising the method of any one of claims 1-27.

29. A computer system, comprising one or more processors and memory, the memory storing instructions for performing a method of training a policy model to identify a subset of a latent space comprising a plurality of compounds that each have a value for at least a first predetermined physical property, in one or more predetermined physical properties, that is in a first specified range, wherein the latent space encodes each compound in a combinatorial synthesis library as an undercompletevector of fixed dimension d in which each entry in the undercomplete vector is a distribution, the method comprising:A) obtaining a specification of a target and a second cutoff for the first predetermined physical property;B) responsive to inputting the specification of the target, the second cutoff for the first predetermined physical property into the policy model, obtaining an action vector of fixed dimension d as output from the policy model by application of the specification of the target, a first cutoff, and the second cutoff for the first predetermined physical property to the policy model in accordance with a first plurality of parameters of the policy model;C) sampling the action vector on a probabilistic basis to identify a plurality of points in the latent space, wherein each point in the plurality of points is of fixed dimension d,D) inputting each point in the plurality of points into a decoder model thereby obtaining a corresponding structure of each candidate compound in a plurality of candidate compounds in accordance with a second plurality of parameters associated with the decoder model, each candidate compound in the plurality of candidate compounds corresponding to a point in the plurality of points;E) determining a value for the first predetermined physical property for each candidate compound in the plurality of candidate compounds using the corresponding structure of the candidate compound;F) acquiring an update to the first plurality of parameters, through application of a loss function to each candidate compound in the plurality of candidate compounds, wherein the loss function includes a reward term, a magnitude of which is determined, at least in part, by an application of the target, the first cutoff and the second cutoff of the first predetermined physical property to the value for the first predetermined physical property for a candidate compound in the plurality of candidate compounds; andG) updating the policy model with the update to the first plurality of parameters thereby training the policy model to sample the latent space to identify the subset of the latent space comprising the plurality of compounds that have the value for the first predetermined physical property that is in the first specified range.

30. A non-transitory computer-readable medium storing one or more computer programs, executable by a computer, for training a policy model to identify a subset of a latent space comprising a plurality of compounds that each have a value for at least a first predetermined physical property, in one ormore predetermined physical properties, that is in a first specified range, wherein the latent space encodes each compound in a combinatorial synthesis library as an undercomplete vector of fixed dimension d in which each entry in the undercomplete vector is a distribution, the computer comprising one or more processors and a memory, the one or more computer programs collectively encoding computer executable instructions that perform a method comprising and one of claims 1-27.

31. A non-transitory computer-readable medium storing one or more computer programs, executable by a computer, for training a policy model to identify a subset of a latent space comprising a plurality of compounds that each have a value for at least a first predetermined physical property, in one or more predetermined physical properties, that is in a first specified range, wherein the latent space encodes each compound in a combinatorial synthesis library as an undercomplete vector of fixed dimension d in which each entry in the undercomplete vector is a distribution, the computer comprising one or more processors and a memory, the one or more computer programs collectively encoding computer executable instructions that perform a method comprising:A) obtaining a specification of a target and a second cutoff for the first predetermined physical property;B) responsive to inputting the specification of the target, a first cutoff, and the second cutoff for the first predetermined physical property into the policy model, obtaining an action vector of fixed dimension d as output from the policy model by application of the specification of the target, the first cutoff, and the second cutoff for the first predetermined physical property to the policy model in accordance with a first plurality of parameters of the policy model;C) sampling the action vector on a probabilistic basis to identify a plurality of points in the latent space, wherein each point in the plurality of points is of fixed dimension cZ;D) inputting each point in the plurality of points into a decoder model thereby obtaining a corresponding structure of each candidate compound in a plurality of candidate compounds in accordance with a second plurality of parameters associated with the decoder model, each candidate compound in the plurality of candidate compounds corresponding to a point in the plurality of points;E) determining a value for the first predetermined physical property for each candidate compound in the plurality of candidate compounds using the corresponding structure of the candidate compound;F) acquiring an update to the first plurality of parameters, through application of a loss function to each candidate compound in the plurality of candidate compounds, wherein the loss functionincludes a reward term, a magnitude of which is determined, at least in part, by an application of the target, the first cutoff and the second cutoff of the first predetermined physical property to the value for the first predetermined physical property for a candidate compound in the plurality of candidate compounds; andG) updating the policy model with the update to the first plurality of parameters thereby training the policy model to sample the latent space to identify the subset of the latent space comprising the plurality of compounds that have the value for the first predetermined physical property that is in the first specified range.

32. A method of training a policy model to identify a subset of a latent space comprising a plurality of compounds that each have at least a first predetermined physical property, in one or more predetermined physical properties, subject to a constraint of maintaining similarity to a test compound, wherein the latent space encodes each compound in a combinatorial synthesis library as an undercomplete real-valued vector of fixed dimension d in which each entry in the undercomplete real- valued vector is a distribution, the method comprising: at a computer system comprising at least one processor and a memory storing at least one program for execution by the at least one processor, the at least one program comprising instructions for:A) obtaining a specification of a cutoff value associated with a maximum reward for the first predetermined physical property;B) responsive to inputting an initial state vector of fixed dimension d that is a latent representation of the test compound into the policy model, obtaining an action vector of fixed dimension d as output from the policy model by application of the initial state vector to the policy model in accordance with a first plurality of parameters of the policy model;C) sampling the action vector on a probabilistic basis to identify a plurality of points in the latent space, wherein each point in the plurality of points is of fixed dimension c / ;D) inputting each point in the plurality of points into a decoder model thereby obtaining a corresponding structure of each candidate compound in a plurality of candidate compounds in accordance with a second plurality of parameters associated with the decoder model, each candidate compound in the plurality of candidate compounds corresponding to a point in the plurality of points;E) determining a value for the first predetermined physical property for each candidate compound in the plurality of candidate compounds using the corresponding structure of the candidate compound;F) acquiring an update to the first plurality of parameters, through application of a loss function to a candidate chemical compound in the plurality of candidate chemical compounds, wherein the loss function includes:(1) a similarity reward term, a magnitude of which is determined by a structural similarity of the candidate compound to the test chemical compound, and(ii) a property reward term, a magnitude of which is determined, at least in part, by a value for the first predetermined physical property that the candidate compound has relative to the specification cutoff value associated with a maximum reward for the first predetermined physical property; andG) updating the policy model with the update to the first plurality of parameters thereby training the policy model to sample the latent space to identify the subset of the latent space comprising compounds that have the predetermined physical property subject to the constraint of maintaining structural similarity to the test compound.

33. The method of claim 32, wherein the first predetermined physical property is compound molecular weight.

34. The method of claim 32, wherein the first predetermined physical property is compound distribution coefficient (logD), compound octanol-water partition coefficient (logP), penalized LogP compound acid dissociation constant (pKA), compound acid dissociation constant (kA), or compound topological polar surface area (TPSA).

35. The method of claim 32, wherein the first predetermined physical property is drug likeness (QED score).

36. The method of claim 32, wherein the first predetermined physical property is permeability.

37. The method of any one of claims 32-36, wherein the combinatorial synthesis library comprises 5,000 compounds.

38. The method of any one of claims 32-36, wherein the combinatorial synthesis library comprises 50,000 compounds.

39. The method of any one of claims 32-36, wherein the combinatorial synthesis library comprises 1 million compounds.

40. The method of any one of claims 32-36, wherein the policy model is a multi-layer perceptron.

41. The method of any one of claims 32-36, wherein the policy model is a deep neural network.

42. The method of any one of claims 32-41, wherein the first plurality of parameters comprises 10,000 or more parameters.

43. The method of any one of claims 32-41, wherein the first plurality of parameters comprises 100,000 or more parameters.

44. The method of any one of claims 32-41, wherein the first plurality of parameters comprises 1 x 106or more parameters.

45. The method of any one of claims 32-44, wherein the sampling C) samples the action vector as a multivariate normal distribution with a mean equal to the action vector and a standard deviation equal to a constant value.

46. The method of any one of claims 32-45, wherein the plurality of candidate compounds comprises 5 or more chemical compounds.

47. The method of any one of claims 32-46, wherein the test compound is an organic compound having a molecular weight of less than 2000 Daltons.

48. The method of any one of claims 32-47, wherein the test compound is an organic compound that satisfies each of the Lipinski rule of five criteria.

49. The method of any one of claims 32-47, wherein the test compound is an organic compound that satisfies at least three criteria of the Lipinski rule of five criteria.

50. The method of any one of claims 32-49, wherein each candidate compound in the plurality of candidate compounds is an organic compound having a molecular weight of less than 2000 Daltons.

51. The method of any one of claims 32-49, wherein each candidate compound in the plurality of candidate compounds is an organic compound that satisfies each of the Lipinski rule of five criteria.

52. The method of any one of claims 32-49, wherein each candidate compound in the plurality of candidate compounds is an organic compound that satisfies at least three criteria of the Lipinski rule of five criteria.

53. The method of any one of claims 32-52, wherein the one or more predetermined physical properties is a plurality of physical properties, the plurality of compounds each have a value for a second predetermined physical property in the plurality of physical properties that is in a second specified range, the obtaining A) further comprises obtaining a specification of a cutoff value associated with a maximum reward for the second predetermined physical property, the determining E) further determines a value for the second predetermined physical property for each candidate compound in the plurality of candidate compounds using the corresponding structure of the candidate compound, and the magnitude of the property reward function is further determined by a value for the second predetermined physical property that the candidate compound has relative to the specification cutoff value associated with a maximum reward for the second predetermined physical property.

54. The method of any one of claims 32-43, wherein the second plurality of parameters comprises 10,000 or more parameters.

55. The method of any one of claims 32-54, wherein the second plurality of parameters comprises 100,000 or more parameters.

56. The method of any one of claims 32-55, wherein the second plurality of parameters comprises 1 x 106or more parameters.

57. The method of any one of claims 32-56, wherein the method further comprises obtaining the initial state vector of fixed dimension d of the test compound as output from an encoder model, comprising a third plurality of parameters, by inputting a string representation of the test compound into the encoder model.

58. The method of claim 57, wherein the third plurality of parameters comprises 10,000 or more parameters.

59. The method of claim 57, wherein the third plurality of parameters comprises 100,000 or more parameters.

60. The method of claim 57, wherein the third plurality of parameters comprises 1 x 106or more parameters.

61. The method of any one of claims 57-60, wherein the encoder model is a first recurrent neural network, and the decoder model is a second recurrent neural network.

62. The method of any one of claims 57-60, wherein the encoder model is a first neural network, and the decoder model is a second neural network.

63. The method of any one of claims 57-62, wherein the string representation is in a SMARTS, a simplified molecular-input line-entry system (SMILES), DeepSMILES, or SELFIES format.

64. A computer system, comprising one or more processors and memory, the memory storing instructions for performing a method of training a policy model to identify a subset of a latent space comprising a plurality of compounds that each have at least a first predetermined physical property, inone or more predetermined physical properties, subject to a constraint of maintaining similarity to a test compound, wherein the latent space encodes each compound in a combinatorial synthesis library as an undercomplete real-valued vector of fixed dimension d in which each entry in the undercomplete real-valued vector is a distribution, the method comprising any the method of any one of claims 32-63.

65. A computer system, comprising one or more processors and memory, the memory storing instructions for performing a method of training a policy model to identify a subset of a latent space comprising a plurality of compounds that each have at least a first predetermined physical property, in one or more predetermined physical properties, subject to a constraint of maintaining similarity to a test compound, wherein the latent space encodes each compound in a combinatorial synthesis library as an undercomplete real-valued vector of fixed dimension d in which each entry in the undercomplete real-valued vector is a distribution, the method comprising: at a computer system comprising at least one processor and a memory storing at least one program for execution by the at least one processor, the at least one program comprising instructions for:A) obtaining a specification of a cutoff value associated with a maximum reward for the first predetermined physical property;B) responsive to inputting an initial state vector of fixed dimension d that is a latent representation of the test compound into the policy model, obtaining an action vector of fixed dimension d as output from the policy model by application of the initial state vector to the policy model in accordance with a first plurality of parameters of the policy model;C) sampling the action vector on a probabilistic basis to identify a plurality of points in the latent space, wherein each point in the plurality of points is of fixed dimension c / ;D) inputting each point in the plurality of points into a decoder model thereby obtaining a corresponding structure of each candidate compound in a plurality of candidate compounds in accordance with a second plurality of parameters associated with the decoder model, each candidate compound in the plurality of candidate compounds corresponding to a point in the plurality of points;E) determining a value for the first predetermined physical property for each candidate compound in the plurality of candidate compounds using the corresponding structure of the candidate compound;F) acquiring an update to the first plurality of parameters, through application of a loss function to a candidate chemical compound in the plurality of candidate chemical compounds, wherein the loss function includes:(1) a similarity reward term, a magnitude of which is determined by a structural similarity of the candidate compound to the test chemical compound, and(ii) a property reward term, a magnitude of which is determined, at least in part, by a value for the first predetermined physical property that the candidate compound has relative to the specification cutoff value associated with a maximum reward for the first predetermined physical property; andG) updating the policy model with the update to the first plurality of parameters thereby training the policy model to sample the latent space to identify the subset of the latent space comprising compounds that have the predetermined physical property subject to the constraint of maintaining structural similarity to the test compound.

66. A non-transitory computer-readable medium storing one or more computer programs, executable by a computer, for training a policy model to identify a subset of a latent space comprising a plurality of compounds that each have at least a first predetermined physical property, in one or more predetermined physical properties, subject to a constraint of maintaining similarity to a test compound, wherein the latent space encodes each compound in a combinatorial synthesis library as an undercomplete real-valued vector of fixed dimension d in which each entry in the undercomplete real- valued vector is a distribution, the one or more computer programs collectively encoding computer executable instructions that perform a method comprising and one of claims 32-63.

67. A non-transitory computer-readable medium storing one or more computer programs, executable by a computer for training a policy model to identify a subset of a latent space comprising a plurality of compounds that each have at least a first predetermined physical property, in one or more predetermined physical properties, subject to a constraint of maintaining similarity to a test compound, wherein the latent space encodes each compound in a combinatorial synthesis library as an undercomplete real-valued vector of fixed dimension d in which each entry in the undercomplete real- valued vector is a distribution, the one or more computer programs collectively encoding computer executable instructions that perform a method comprising:A) obtaining a specification of a cutoff value associated with a maximum reward for the first predetermined physical property;B) responsive to inputting an initial state vector of fixed dimension d that is a latent representation of the test compound into the policy model, obtaining an action vector of fixed dimension d as output from the policy model by application of the initial state vector to the policy model in accordance with a first plurality of parameters of the policy model;C) sampling the action vector on a probabilistic basis to identify a plurality of points in the latent space, wherein each point in the plurality of points is of fixed dimension d,D) inputting each point in the plurality of points into a decoder model thereby obtaining a corresponding structure of each candidate compound in a plurality of candidate compounds in accordance with a second plurality of parameters associated with the decoder model, each candidate compound in the plurality of candidate compounds corresponding to a point in the plurality of points;E) determining a value for the first predetermined physical property for each candidate compound in the plurality of candidate compounds using the corresponding structure of the candidate compound;F) acquiring an update to the first plurality of parameters, through application of a loss function to a candidate chemical compound in the plurality of candidate chemical compounds, wherein the loss function includes:(1) a similarity reward term, a magnitude of which is determined by a structural similarity of the candidate compound to the test chemical compound, and(ii) a property reward term, a magnitude of which is determined, at least in part, by a value for the first predetermined physical property that the candidate compound has relative to the specification cutoff value associated with a maximum reward for the first predetermined physical property; andG) updating the policy model with the update to the first plurality of parameters thereby training the policy model to sample the latent space to identify the subset of the latent space comprising compounds that have the predetermined physical property subject to the constraint of maintaining structural similarity to the test compound.

Citation Information

Patent Citations

  • Generative machine learning systems for drug design

    US20200387831A1

  • System and method for the latent space optimization of generative machine learning models

    US20220358373A1

  • Molecule design

    US20230052677A1

  • Prediction of chemical compounds with desired properties

    US20230335231A1

  • Workflow for generating compounds with biological activity against a specific biological target

    WO2021038420A1