Method, computer system, and program

A computer system using probabilistic autoencoders and generative models addresses inefficiencies in drug discovery by generating compounds with desired properties and predicting interactions, reducing toxicity risks and improving the drug discovery process.

JP2026016526APending Publication Date: 2026-02-03PREFERRED NETWORKS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025178339
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2016-02-03
Filing Date
2025-10-23
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Current drug discovery methods, including high-throughput and virtual screening, are inefficient, costly, and prone to failure due to toxicity issues and side effects, with limited predictive capabilities for compound interactions and properties.

Method used

A computer system utilizing probabilistic autoencoders and generative models to directly generate compound representations that satisfy desired properties and minimize toxicity, incorporating training with reconstruction and regularization errors to optimize the model.

Benefits of technology

The system effectively generates compounds with desired properties and predicts interactions, reducing the risk of toxicity and side effects, and identifies compounds not present in the training dataset, enhancing the drug discovery process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026016526000001_ABST
    Figure 2026016526000001_ABST
Patent Text Reader

Abstract

In various embodiments, the systems and methods described herein relate to generative models.SOLUTION: The generative model may be trained using a machine learning technique with a training set comprising chemical compounds and biological or chemical information related to the chemical compounds. Deep learning architectures may be used. In various embodiments, a generative model is used to generate chemical compounds with desired properties, e.g., activity against a selected target. A generative model may be used to generate chemical compounds that satisfy multiple requirements.SELECTED DRAWING: Figure 2A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to generative machine learning systems for drug design. [Background technology]

[0002] The search for lead compounds with desired properties typically involves high-throughput or virtual screening, methods that are slow, costly, and ineffective. Summary of the Invention [Problem to be solved by the invention]

[0003] In high-throughput screening, compounds from a compound library are examined. However, compound libraries are large, and most of the candidates are not eligible for selection as hit compounds. To minimize the costs associated with this complex approach, some screening methods utilize in silico methods known as virtual screening. However, available virtual screening methods require extensive computational power and can be algorithmically inefficient and time-consuming.

[0004] Furthermore, current hit-to-lead discovery primarily involves exhaustive screening from a vast list of compound candidates. This approach relies on the expectation and hope that a compound with a set of desired properties will be found within the existing list of compounds. Furthermore, even when current screening methods successfully discover lead compounds, this does not mean that these lead compounds can be used as drugs. It is not uncommon for candidate compounds to fail in the later stages of clinical trials. One of the main reasons for failure is toxicity or side effects that do not become apparent until animal or human experiments. Finally, these discovery models are slow and expensive.

[0005] Due to the inefficiencies and limitations of existing methods, there is a need for a drug design method that directly generates candidate compounds with a desired set of properties, such as binding to target proteins.Furthermore, there is a need to generate candidate compounds that are free of toxicity or side effects.Finally, there is a need to predict how candidate compounds interact with off-targets and / or other targets. [Means for solving the problem]

[0006] In a first aspect, methods and systems described herein relate to a computer system for generating compound representations. The system may include a probabilistic autoencoder. The probabilistic autoencoder may include a probabilistic encoder configured to encode a compound fingerprint as a latent variable, a probabilistic decoder configured to decode the latent representation and generate random variables across the values ​​of the fingerprint elements, and / or one or more sampling modules configured to sample from the latent or random variables. The system may be trained by providing compound fingerprints and training labels associated with the compound fingerprints and generating reconstructions of the compound fingerprints, where the training of the system is constrained by a reconstruction error. The reconstruction error may include a negative probability that the encoded compound representation is drawn from the random variables generated by the probabilistic decoder. The system may be trained to optimize, e.g., minimize, the reconstruction error. In some embodiments, the training is constrained by a loss function including the reconstruction error and a regularization error. The probabilistic autoencoder may be trained to learn to approximate an encoding distribution. The regularization error may include a penalty related to the complexity of the encoding distribution. The training may include minimizing the loss function. In some embodiments, the training labels include one or more label elements having predetermined values. In some embodiments, the system is configured to receive target labels including one or more label elements and generate compound fingerprints that satisfy the specified values ​​for each of the one or more label elements. In some embodiments, the training labels do not include the target label. In some embodiments, each compound fingerprint uniquely identifies a compound. In some embodiments, In some embodiments, training further constrains the overall information flow between the probabilistic encoder and the probabilistic decoder. In some embodiments, the probabilistic encoder is configured to provide an output including a pair of a vector of means and a vector of standard deviations. In some embodiments, the sampling module is configured to receive the output of the encoder, define a latent variable based on the output of the encoder, and generate one or more latent representations, where the latent variable is modeled by a probability distribution. In some embodiments, the probability distribution is selected from the group consisting of a normal distribution, a Laplace distribution, an elliptical distribution, a Student's t distribution, a logistic distribution, a uniform distribution, a triangular distribution, an exponential distribution, a reversible cumulative distribution, a Cauchy distribution, a Rayleigh distribution, a Pareto distribution, a Weibull distribution, a reciprocal distribution, a Gompertz distribution, a Gumbel distribution, an Erlang distribution, a lognormal distribution, a gamma distribution, a Dirichlet distribution, a beta distribution, a chi-squared distribution, an F distribution, and variations thereof. In some embodiments, the probabilistic encoder includes an inference model. In some embodiments, the inference model includes a multilayer perceptron. In some embodiments, the probabilistic autoencoder includes a generative model. In some embodiments, the generative model includes a multilayer perceptron. In some embodiments, the system further includes a predictor configured to predict values ​​of selected label elements for the compound fingerprint, hi some embodiments, the label includes one or more label elements selected from the group consisting of bioassay results, toxicity, cross-reactivity, pharmacokinetics, pharmacodynamics, bioavailability, and solubility.

[0007] In another aspect, the systems and methods described herein relate to a training method for generating compound representations. The training method may include training a generative model. Training the training model may include inputting compound fingerprints and associated training labels into the generative model and generating a reconstruction of the compound fingerprint. The generative model may include a probabilistic autoencoder including a probabilistic encoder configured to encode the compound fingerprint as a latent variable, a probabilistic decoder configured to decode the latent representation as a random variable across the values ​​of the fingerprint elements, and / or a sampling module configured to sample from the latent variable to generate the latent representation or sample from the random variable to generate the fingerprint reconstruction. The training labels may include one or more label elements having empirical or predicted values. Training of the system may be constrained by a reconstruction error. The reconstruction error may include a negative likelihood that the encoded compound representation is drawn from the random variables output by the probabilistic decoder. Training may include minimizing the reconstruction error. In some embodiments, training is constrained by a loss function including the reconstruction error and a regularization error. Training may include minimizing the loss function.

[0008] In yet another aspect, the methods and systems described herein relate to a computer system for drug prediction. The system may include a machine learning model, including a generative model. The generative model may be trained with a training dataset including compound fingerprint data and associated training labels, each of which includes one or more label elements. In some embodiments, the generative model includes a neural network having at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more layers of units. In some embodiments, the label elements include one or more elements selected from the group consisting of bioassay results, toxicity, cross-reactivity, pharmacokinetics, pharmacodynamics, bioavailability, and solubility. In some embodiments, the generative model includes a stochastic autoencoder. In some embodiments, the generative model includes a variational autoencoder having a stochastic encoder, a stochastic decoder, and a sampling module. In some embodiments, the stochastic encoder is configured to provide an output including a pair of a vector of means and a vector of standard deviations. In some embodiments, the sampling module is configured to receive the output of the probabilistic encoder, define a latent variable based on the encoder output, and generate one or more latent representations, where the latent variable is modeled by a probability distribution. In some embodiments, the probabilistic decoder is configured to decode the latent representations and generate random variables over the values ​​of the fingerprint elements. In some embodiments, the probability distribution is selected from the group consisting of a normal distribution, a Laplace distribution, an elliptical distribution, a Student's t distribution, a logistic distribution, a uniform distribution, a triangular distribution, an exponential distribution, a reversible cumulative distribution, a Cauchy distribution, a Rayleigh distribution, a Pareto distribution, a Weibull distribution, a reciprocal distribution, a Gompertz distribution, a Gumbel distribution, an Erlang distribution, a lognormal distribution, a gamma distribution, a Dirichlet distribution, a beta distribution, a chi-squared distribution, an F distribution, and variations thereof. In some embodiments, the probabilistic encoder and the probabilistic decoder are trained simultaneously. In some embodiments, the computer system includes GNU. In some embodiments, the generative model further includes a predictor.In some embodiments, the predictor is configured to predict values ​​of one or more label elements for at least a subset of the fingerprint-associated training labels. In some embodiments, the machine learning network is configured to provide an output that includes system-generated compound fingerprints that are not in the training dataset.

[0009] In a further aspect, methods and systems described herein relate to a method for drug prediction. The method may include training a generative model with a training dataset including compound fingerprints and associated training labels, including one or more label elements with empirical or predicted label element values. In some embodiments, the labels include one or more elements selected from the group consisting of bioassay results, toxicity, cross-reactivity, pharmacokinetics, pharmacodynamics, bioavailability, and solubility. In some embodiments, the generative model includes a probabilistic autoencoder. In some embodiments, the generative model includes a variational autoencoder including a probabilistic encoder and a probabilistic decoder and a sampling module. In some embodiments, the method further includes providing an output from the encoder including a pair of a vector of means and a vector of standard deviations for each compound fingerprint in the training dataset. In some embodiments, the probabilistic encoder and the probabilistic decoder are trained simultaneously. In some embodiments, the training includes training a probabilistic encoder to encode compound fingerprints as vectors of means and vectors of standard deviations that define latent variables, deriving latent representations from the latent variables, and training a probabilistic decoder to decode the latent representations as probabilistic reconstructions of the compound fingerprints. In some embodiments, the latent variables are modeled by a probability distribution selected from the group consisting of a normal F-distribution, a Laplace distribution, an elliptical distribution, a Student's t-distribution, a logistic distribution, a uniform distribution, a triangular distribution, an exponential distribution, a reversible cumulative distribution, a Cauchy distribution, a Rayleigh distribution, a Pareto distribution, a Weibull distribution, a reciprocal distribution, a Gompertz distribution, a Gumbel distribution, an Erlang distribution, a lognormal distribution, a gamma distribution, a Dirichlet distribution, a beta distribution, a chi-squared distribution, an F-distribution, and variations thereof. In some embodiments, the training includes optimizing a variational lower bound for the variational autoencoder using backpropagation. In some embodiments, the generative model resides in a computer system having GNU / Linux. In some embodiments, the generative model includes a predictor module.In some embodiments, the method further includes predicting one or more values ​​for label elements associated with one or more compound fingerprints in the training data set. In some embodiments, the method further includes generating an output from the generative model that includes identifying information for compounds not represented in the training set.

[0010] In still further aspects, methods and systems described herein relate to computer systems for generating compound representations. The system may include a probabilistic autoencoder. The system may be trained by inputting a training dataset including compound fingerprints and associated training labels including one or more label elements and generating reconstructions of the compound fingerprints. Training the system may be constrained by a reconstruction error and / or a regularization error. The generated reconstructions may be sampled from a reconstruction distribution. The reconstruction error may include a negative likelihood that the input compound fingerprint is drawn from the reconstruction distribution. Training the system may include training the probabilistic autoencoder to approximate an encoding distribution. The regularization error may include a penalty related to the complexity of the encoding distribution. In some embodiments, the system is configured to generate compound fingerprints that satisfy selected values ​​for one or more label elements. In some embodiments, the training labels do not include the selected values ​​for one or more label elements. In some embodiments, each compound fingerprint uniquely identifies a compound. In some embodiments, the probabilistic autoencoder includes at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more layers. In some embodiments, the computer system may further include a predictor configured to predict values ​​for one or more label elements associated with one or more compound fingerprints in the training dataset. In some embodiments, the label elements include one or more elements selected from the group consisting of bioassay results, toxicity, cross-reactivity, pharmacokinetics, pharmacodynamics, bioavailability, and solubility.

[0011] In yet another aspect, the methods and systems described herein relate to a method for generating compound representations. The method may include training a machine learning model. The training may include inputting a compound fingerprint and an associated label including one or more label elements into the machine learning model and generating a reconstruction of the compound fingerprint. The machine learning model may include a probabilistic autoencoder or a variational autoencoder. In some embodiments, the training is constrained by a reconstruction error and a regularization error. The generated reconstruction may be sampled from a reconstruction distribution. In some embodiments, the reconstruction error includes a negative likelihood that the input compound fingerprint is drawn from the reconstruction distribution. The training may include training the probabilistic autoencoder to approximate an encoding distribution. The regularization error may include a penalty related to the complexity of the encoding distribution.

[0012] In a further aspect, the methods and systems described herein relate to a computer system for drug prediction. The system may include a machine learning model, including a generative model. The machine learning model may be trained with a first training dataset including chemical fingerprint data and an associated set of labels having a first label element, and a second training dataset including chemical fingerprint data and an associated set of labels having a second label element. In some embodiments, the chemical fingerprint data of the first and second training datasets are input to units in at least two layers of a generative network. In some embodiments, labels having a first label element and labels having a second label element are introduced to different parts of the generative network during training. In some embodiments, the first label element represents the activity of a compound associated with the chemical fingerprint in a first bioassay. In some embodiments, the second label element represents the activity of a compound associated with the chemical fingerprint in a second bioassay. In some embodiments, the system is configured to generate representations of compounds that are likely to satisfy a requirement for a specified value for the first label element having a first type and a requirement for a specified value for the second label element. In some embodiments, a high probability is greater than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 95, 98, 99%, or more. In some embodiments, the requirement for a specified value for the first label element includes having a positive result for the first bioassay that is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, 30, 50, 100, 500, 1000, or more standard deviations compared to noise. In some embodiments, the requirement for a specified value for the first label element includes having a positive result for the first bioassay that is at least 10, 20, 30, 40, 50, 100, 200, 500, 1000% greater than the activity of an equimolar concentration of the known compound. In some embodiments, the requirement for a specified value for the first label element includes having a positive result for the first bioassay that is at least 100% greater than the activity of an equimolar concentration of the known compound. In some embodiments, the requirement for a specified value for the first label element includes having a positive result for the first bioassay that is at least 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 25-fold, 50-fold, 100-fold, 200-fold, 300-fold, 400-fold, 500-fold, 1000-fold, 10,000-fold, or 100,000-fold greater than the activity of an equimolar concentration of the known compound. In some embodiments, the requirement for a specified value for the second label element includes having a positive result for the second bioassay that is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, 30, 50, 100, 500, 1000, or more standard deviations compared to noise. In some embodiments, the requirement for a specified value for the second label element includes having a positive result for the second bioassay that is at least 10, 20, 30, 40, 50, 100, 200, 500, or 1000% greater than the activity of an equimolar concentration of the known compound. In some embodiments, the requirement for the specified value of the second label element includes having a positive result for the second bioassay that is at least 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 15-fold, 25-fold, 50-fold, 100-fold, 200-fold, 300-fold, 400-fold, 500-fold, 1000-fold, 10,000-fold, or 100,000-fold greater than the activity of the known compound at an equimolar concentration.

[0013] <Incorporated by reference> All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.

[0014] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings. [Brief explanation of the drawings]

[0015] [Figure 1] FIG. 1 is an explanatory diagram of an autoencoder. [Figure 2A] 1 illustrates an exemplary architecture of a predictor-free multi-component generative model. A generative model with such an architecture may be trained by supervised learning. [Figure 2B] 1 illustrates an example architecture of a multi-component generative model with predictors. A generative model with such an architecture may be trained by semi-supervised learning. [Figure 3] FIG. 10 shows an example for the initial creation of a generated representation of a compound that meets the requirements set by a desired label y. [Figure 4A] 1 provides an exemplary illustration for creating a compound representation generated based on a labeled seed compound. A compound representation x may be generated by using an actual label yD and a desired label y. [Figure 4B] 1 provides an exemplary illustration for creating an unlabeled seed compound. A compound representation x̂ may be generated by using the predicted label y generated by the predictor module and the desired label ŷ. [Figure 5A] 1 depicts an example of an encoder according to various embodiments of the present invention. [Figure 5B] FIG. 2 depicts an example of a decoder according to various embodiments of the present invention. [Figure 6] FIG. 1 depicts an example of a method for training a variational autoencoder, according to various embodiments of the present invention. [Figure 7]FIG. 1 depicts an example of a single-step evaluation and ranking procedure, according to various embodiments of the present invention. [Figure 8] 1 depicts an example of how generated fingerprints and their predicted results are evaluated, according to various embodiments of the present invention; [Figure 9] FIG. 1 depicts an exemplary illustration of a training method for a ranking module. [Figure 10] FIG. 1 depicts an exemplary illustration of a ranking module including a latent representation generator (LRG), a classifier, and an ordering module, according to various embodiments of the present invention. [Figure 11] FIG. 1 depicts an exemplary illustration of the sequential use of an initial generation process and a comparison generation process. [Figure 12] FIG. 1 depicts an exemplary method and system for identification of compound properties that may influence changes in label or label element values. [Figure 13] FIG. 1 depicts a system and method for identification of modifications in a particular compound that may be associated with a desired label or label element value. [Figure 14] FIG. 1 depicts an exemplary illustration of a comparison module using k-medoid clustering. [Figure 15] FIG. 1 depicts an exemplary illustration of a comparison module using k-means clustering. [Figure 16] FIG. 1 is a block diagram of an example computer system capable of performing one or more operations described herein. [Figure 17A] FIG. 1 depicts an example illustration of an alternative configuration of an input layer for fingerprints and labels in a machine learning model, where fingerprints and labels are input into the same layer of the machine learning model. [Figure 17B] FIG. 1 depicts an example illustration of an alternative configuration of an input layer for fingerprints and labels in a machine learning model, where fingerprints and labels are input to different layers of the machine learning model. DETAILED DESCRIPTION OF THE INVENTION

[0016] In various embodiments, the present invention relates to methods and systems that enable the direct generation of compound candidate representations through the use of machine learning and / or artificial intelligence methods. In various embodiments, the methods and systems described herein involve utilizing generative models, deep generative models, directed graphical models, deep directed graphical models, directed latent graphical models, latent variable generative models, nonlinear Gaussian probability networks, sigmoid probability networks, deep autoregressive networks, neural autoregressive distribution estimators, generalized denoising autoencoders, deep latent Gaussian models, and / or combinations thereof. In some embodiments, the generative model utilizes a probabilistic autoencoder, such as a variational autoencoder. A component of a generative model, such as a variational autoencoder, may include a multilayer perceptron implementing a probabilistic encoder and a probabilistic decoder. The encoder and decoder may be trained simultaneously, for example, by using backpropagation.

[0017] The systems and methods described herein may be used to generate novel compounds that were not included in the training dataset used to train the generative model. Furthermore, the methods and systems of the present invention in various embodiments increase the likelihood of identifying one or more compounds with a desired set of properties. In various embodiments, the methods and systems of the present invention include the simultaneous prediction of compound effects and side effects, or the discovery of new uses for existing drugs, commonly referred to as drug repositioning. In various embodiments, references to a "compound" or "generating a compound" relate to uniquely identifying the compound and information related to its generation, but not necessarily to the physical creation of the compound. Uniquely identifying such information may include a chemical formula or structure, a reference code, or any other suitable identifier described herein or known in the art. May include:

[0018] In exemplary embodiments, the desired set of properties for a compound includes one or more of activity, solubility, toxicity, and ease of synthesis. The methods and systems described herein can facilitate prediction of off-target effects, or how a drug candidate will interact with targets other than a selected target.

[0019] While machine learning techniques have been successful in computerized image recognition, the improvements they have provided to date in the field of computerized drug discovery have been modest in comparison. The systems and methods described herein provide a solution that includes generative models that improve predictions about compounds and their activity, effects, side effects, and properties in novel ways. The generative models described herein provide a unique approach by generating compounds according to desired specifications.

[0020] In various embodiments, the methods and systems described herein are provided with compound information typically characterized by a set of molecular descriptors that represent chemical information such as chemical formula, chemical structure, electron density, or other chemical properties. The compound information may include a fingerprint representation of each compound. Additionally, the methods and systems described herein may be provided with labels that include additional information, including biological data, e.g., bioassay results, such as those describing the compound's activity with respect to a particular target, such as a receptor or enzyme. The methods and systems described herein may be trained with a training set that includes pairs of vectors of molecular descriptor values ​​and vectors of label element values. The combination of compound information and label typically includes data regarding the biological and chemical properties of the compound, including, for example, bioassay data, solubility, cross-reactivity, and other chemical characteristics such as hydrophobicity, phase transition boundaries, e.g., y, or any other information that can be used to characterize the structure or function of the compound. Upon training, the systems and methods described herein can output chemical information identifying one or more compounds, such as one or more chemical fingerprints. In some embodiments, the methods and systems described herein can output identifying chemical information for one or more compounds predicted to have desired chemical and / or biological properties. For example, the identified compounds may be predicted to have test results within desired ranges for one or more specified bioassay results, toxicity, cross-reactivity, etc. The methods and systems described herein can, in some cases, output a list of compounds ranked according to their predicted level of having the desired property. The identified compounds may be used as lead compounds or initial compounds in hit-lead research.

[0021] The methods and systems described herein can utilize compounds of a certain size. For example, a generative model, e.g., a deep generative model, in various embodiments, can be trained with and / or generate representations of compounds having molecular weights of less than 100,000, 50,000, 40,000, 30,000, 20,000, 15,000, 10,000, 9,000, 8,000, 7,000, 6,000, 5,000, 4,000, 3,000, 2,500, 2,000, 1,500, 1,250, 1,000, 900, 800, 750, 600, 500, 400, or 300 daltons.

[0022] Some portions of the detailed descriptions which follow are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art.

[0023] All of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise specified, as will be apparent from the following description, descriptions utilizing terms such as "processing" or "calculating" or "computing" or "determining" or "displaying" throughout the description refer to the actions and processes of a computer system or similar electronic computing device that manipulates and transforms data represented as physical (electronic) quantities in the computer system's registers and memory into other data similarly represented as physical quantities in the computer system's memory or registers, or other such information storage, transmission, or display devices.

[0024] The systems and methods of the present invention may include one or more machine learning structures and substructures, such as a generative model implemented in a multilayer perceptron, a probabilistic autoencoder, or a variational autoencoder, and may utilize any suitable learning algorithm described herein or known in the art, for example, but not limited to, backpropagation with stochastic gradient descent to minimize a loss function or backpropagation with stochastic gradient ascent to optimize a variational lower bound. Once a model is trained, it can be used to evaluate new instances of data presented to a computer or computer network for prediction, for example, using a prediction module (or predictor). The prediction module may include some or all of the machine learning structures used during the training phase. In some embodiments, new compound fingerprints may be generated by sampling from the random variables generated by the model.

[0025] In some embodiments, the methods and systems described herein train a probabilistic or variational autoencoder that can then be used as a generative model. In one embodiment, the probabilistic or variational autoencoder is embodied as a multilayer perceptron including at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or more hidden layers. In some cases, the probabilistic or variational autoencoder may include a multilayer perceptron including a probabilistic encoder and a probabilistic decoder. In other embodiments, any of a variety of statistical models may be implemented that can be trained to form a generative model, as described in more detail elsewhere herein. Supervised or semi-supervised training algorithms may be used to train machine learning systems with the specified architecture.

[0026] In a first aspect, the methods and systems described herein relate to a computer system for generating representations of compounds. The system may include a probabilistic or variational autoencoder. The probabilistic or variational autoencoder may include a probabilistic encoder for converting fingerprint data into latent random variables from which latent representations can be sampled, a probabilistic decoder for converting the latent representations into random variables from which samples can be drawn, thereby generating a reconstruction of a compound representation, and a sampling module capable of sampling the latent representations from the latent random variables and / or sampling the compound fingerprints from the random variables. The system may be trained by inputting representations of compounds and their associated labels and generating reconstructions of the compound representations, where the distributions of the compound fingerprints and reconstructions depend on the value of a loss function including a reconstruction error and a regularization error. The reconstruction error may include a negative probability that the input compound representation is drawn from the random variables generated by the probabilistic decoder. The probabilistic autoencoder may be trained to learn to approximate an encoding distribution. The regularization error may include a penalty related to the complexity of the encoding distribution. The system may be trained to optimize, e.g., minimize, the loss function. In some embodiments, the system is trained by further inputting training labels associated with the compounds. is configured to generate compound fingerprints that are likely to satisfy a selected set of desired label element values. In some embodiments, the set of desired label element values ​​does not appear among the labels in the training dataset. In some embodiments, each compound fingerprint uniquely identifies a compound. In some embodiments, the encoder is configured to provide an output including a pair of a vector of means and a vector of standard deviations. The system can define latent random variables based on the encoder output. The latent random variables may be modeled by a probability distribution, such as a normal distribution, a Laplace distribution, an elliptical distribution, a Student's t distribution, a logistic distribution, a uniform distribution, a triangular distribution, an exponential distribution, a reversible cumulative distribution, a Cauchy distribution, a Rayleigh distribution, a Pareto distribution, a Weibull distribution, a reciprocal distribution, a Gompertz distribution, a Gumbel distribution, an Erlang distribution, a lognormal distribution, a gamma distribution, a Dirichlet distribution, a beta distribution, a chi-squared distribution, or an F distribution, or variations thereof. The encoder and / or decoder may include one or more layers of a multilayer perceptron or other type of neural network, such as a recurrent neural network. The system may further include a predictor for predicting a label element value associated with the compound fingerprint. In some embodiments, the label element includes one or more elements selected from the group consisting of bioassay results, toxicity, cross-reactivity, pharmacokinetics, pharmacodynamics, bioavailability, and solubility.

[0027] In another aspect, the systems and methods described herein relate to a method for generating compound representations. The method may include training a generative model. The training may include (1) inputting compound representations and their associated labels and (2) generating a compound fingerprint reconstruction. The generative model may include a probabilistic or variational autoencoder including: a) a probabilistic encoder for encoding fingerprint and label data as latent variables from which a latent representation can be sampled; b) a probabilistic decoder for converting the latent representations into random variables from which a reconstruction of the fingerprint data can be sampled; and c) a sampling module for sampling the latent variables to generate the latent representations or sampling the random variables to generate the fingerprint reconstruction. The system may be trained to optimize, e.g., minimize, a loss function including a reconstruction error and a regularization error. The reconstruction error may include the negative probability that the encoded compound representation is drawn from the random variables output by the probabilistic decoder. The training may include training the variational or stochastic autoencoder to approximate an encoding distribution. The regularization error may include a penalty related to the complexity of the encoding distribution.

[0028] In yet another aspect, the methods and systems described herein relate to computer systems for drug prediction. It is understood that "drug prediction," in connection with various embodiments of the present invention, refers to the analysis of a compound for specific chemical and physical properties. Subsequent activities, such as synthesis, in vivo and in vitro testing, and clinical trials with the compound, are understood to follow in certain embodiments of the present invention, but such subsequent activities are not implied by the term "drug prediction." The system may include a machine learning model, including a generative model. The generative model may be trained with a training dataset that includes compound representations, such as fingerprint data. In some embodiments, the machine learning model includes at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or more layers of units. In some embodiments, the training dataset further includes labels associated with at least a subset of the compounds in the training dataset. The labels may have label elements, such as one or more of the compound's activities and properties, such as bioassay results, toxicity, cross-reactivity, pharmacokinetics, pharmacodynamics, bioavailability, solubility, or any other suitable label element known in the art. The generative model may include a probabilistic autoencoder. In some embodiments, the probabilistic autoencoder comprises at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or more The generative model may include a multi-layer perceptron having layers of units. In some embodiments, the generative model includes a probabilistic autoencoder or a variational autoencoder including a probabilistic encoder, a probabilistic decoder, and a sampling module. The probabilistic encoder may be configured to provide an output including a pair of a vector of means and a vector of standard deviations. The system may define latent random variables based on the encoder output. The latent random variables may be modeled by a probability distribution, such as a normal distribution, a Laplace distribution, an elliptical distribution, a Student's t distribution, a logistic distribution, a uniform distribution, a triangular distribution, an exponential distribution, a reversible cumulative distribution, a Cauchy distribution, a Rayleigh distribution, a Pareto distribution, a Weibull distribution, a reciprocal distribution, a Gompertz distribution, a Gumbel distribution, an Erlang distribution, a lognormal distribution, a gamma distribution, a Dirichlet distribution, a beta distribution, a chi-squared distribution, an F distribution, or variations thereof. The computer system may include GNU / Linux. The generative model may further include a predictor. The predictor may be configured to predict label element values ​​for at least a subset of the compound fingerprints in the training dataset. In some embodiments, the generative model is configured to provide an output that includes a compound representation generated by the model. The representation may be sufficient to uniquely identify the compound. The generated compound may be a compound that was not included in the training dataset, and in some cases may be a compound that has not been previously synthesized or even considered.

[0029] In a further aspect, the methods and systems described herein relate to a method for drug prediction. The method may include training a machine learning model with a training dataset including compound representations and associated label element values ​​representing compound activities or properties for at least a subset of compounds in the training dataset. The machine learning model may include a generative model. In some embodiments, the labels have elements such as bioassay results, toxicity, cross-reactivity, pharmacokinetics, pharmacodynamics, bioavailability, or solubility. The generative model may include a probabilistic autoencoder, such as a probabilistic autoencoder or a variational autoencoder. The probabilistic autoencoder or variational autoencoder may include a probabilistic encoder, a probabilistic decoder, and a sampling module. The method may further include providing an output from the encoder including pairs of a vector of means and a vector of standard deviations. The pairs of the vector of means and a vector of standard deviations may be used to define latent variables. In some embodiments, the method may further include causing the sampling module to derive latent representations from the latent variables. The latent variables may be modeled by a probability distribution such as a normal distribution, a Laplace distribution, an elliptical distribution, a Student's t distribution, a logistic distribution, a uniform distribution, a triangular distribution, an exponential distribution, a reversible cumulative distribution, a Cauchy distribution, a Rayleigh distribution, a Pareto distribution, a Weibull distribution, a reciprocal distribution, a Gompertz distribution, a Gumbel distribution, an Erlang distribution, a lognormal distribution, a gamma distribution, a Dirichlet distribution, a beta distribution, a chi-squared distribution, an F distribution, or variations thereof. In some embodiments, the machine learning model resides in a computer system having a GPU. In some embodiments, the machine learning model includes a predictor module. The method may further include predicting label element values ​​for a subset of the training data using the predictor module. In some embodiments, the method further includes generating an output from the machine learning model including a set of molecular descriptors sufficient to identify the compound. The compound may not be in the training set.

[0030] In yet a further aspect, methods and systems described herein relate to a computer system for generating compound representations. The system may include a probabilistic or variational autoencoder, and the system is trained by inputting a compound representation and generating reconstructions of the compound representation, where the training of the system is constrained by a reconstruction error and / or a regularization error. The generated reconstructions may be sampled from a reconstruction distribution, where the reconstruction error may include a negative probability that the input compound fingerprint is drawn from the reconstruction distribution. The regularization error may include a penalty related to the complexity of the encoding distribution. Label element values ​​associated with the compound may be input to the system at the same point as the compound representation or at a different point; for example, the labels may be input to a decoder of the autoencoder. In some embodiments, the system is configured to generate a compound representation, where the compound is likely to satisfy one or more requirements defined by a set of desired label element values. In some embodiments, the set of desired label element values ​​may not have been part of the training dataset. In some embodiments, each compound fingerprint uniquely identifies a compound. In some embodiments, the training further constrains the overall information flow through the layers of the generative network. In some embodiments, the probabilistic or variational autoencoder comprises a multilayer perceptron having at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or more layers. In some embodiments, the system further comprises a predictor for associating a label with the compound representation. In some embodiments, the label comprises one or more label elements, such as bioassay results, toxicity, cross-reactivity, pharmacokinetics, pharmacodynamics, bioavailability, and solubility.

[0031] In yet another aspect, methods and systems described herein relate to a method for generating compound representations. The method may include training a machine learning model. The training may include (1) inputting a compound representation, such as a fingerprint, into the machine learning model, and (2) generating a reconstruction of the compound representation, e.g., the fingerprint. The machine learning model may include a probabilistic or variational autoencoder. The system may be trained to optimize, e.g., minimize, a loss function including a reconstruction error and a regularization error. The generated reconstruction may be sampled from a reconstruction distribution. The reconstruction error may include a negative likelihood that the input compound fingerprint is drawn from the reconstruction distribution. The training may include training the probabilistic or variational autoencoder to approximate an encoding distribution. The regularization error may include a penalty related to the complexity of the encoding distribution.

[0032] In a further aspect, the methods and systems described herein relate to a computer system for drug prediction. The system may include a machine learning model, including a generative model. The machine learning model may be trained with a first training dataset including a compound representation, such as a fingerprint, and an associated set of labels having values ​​for a first label element, and a second training dataset including a compound representation, such as a fingerprint, and an associated set of labels having values ​​for a second label element. In some embodiments, labels having the first label element and labels having the second label element are introduced into different parts of the generative model during training, for example, into the encoder and decoder, respectively. In some embodiments, the labels having the first label element represent the activity of the compound in a first bioassay. In some embodiments, the labels having the second label element represent the activity of the compound in a second bioassay. In some embodiments, the system is configured to generate representations of compounds that are likely to satisfy requirements for labels having the first label element value and requirements for labels having the second label element value. In some embodiments, a high likelihood is greater than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 95, 98, 99% or more. In some embodiments, the requirement for a first label element includes having a positive result for the first bioassay that is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, 30, 50, 100, 500, 1000 or more standard deviations compared to noise. In some embodiments, the requirement for a first label element includes having a positive result for the first bioassay that is at least 10, 20, 30, 40, 50, 100, 200, 500, 1000% or more compared to the activity of an equimolar concentration of a known compound.In some embodiments, the requirements for the second label element include having a positive result for the second bioassay that is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, 30, 50, 100, 500, 1000, or more standard deviations compared to noise. In some embodiments, the requirements for the second label element include having a positive result for the second bioassay that is at least 10, 20, 30, 40, 50, 100, 200, 500, 1000% greater than the activity of an equimolar concentration of the known compound.

[0033] <Generative Model> In various embodiments, the systems and methods described herein utilize generative models as a core component.

[0034] Generative models according to the methods and systems of the present invention can be used to randomly generate observable data values ​​given the values ​​of one or more hidden parameters. Generative models can be used to directly model data (i.e., model compound observations drawn from a probability density function) or as an intermediate step toward forming a conditional probability density function. Examples of generative models include, but are not limited to, probabilistic autoencoders, variational autoencoders, Gaussian mixture models, hidden Markov models, and restricted Boltzmann machines. Generative models, described in more detail elsewhere herein, typically specify a joint probability distribution over a compound representation, i.e., a fingerprint, and a label associated with the compound.

[0035] As an example, a set of compounds may be represented as x = (x1, x2, , xN), where xi may contain fingerprint representations of the compounds and N is the number of compounds in the set. These compounds may be associated with a set of N labels L = (l1, l2, , lN), where l is a label that may include values ​​of label elements such as the compound's activity, toxicity, solubility, ease of synthesis, or other results in bioassay or predictive studies. A generative model may be constructed under the assumption that these compounds and their associated labels are generated from an unknown distribution D, i.e., D~(xn, ln). Training the generative model may utilize a training method that adjusts the model's internal parameters to model a joint probability distribution p(x, l) given the data examples in a training dataset. After the generative model is trained, it may be used to generate values ​​of x conditioned on the values ​​of l, i.e., x~p(x|l). For example, a generative model trained on a training set of fingerprints and labels can generate representations of compounds that are likely to meet specified label value requirements.

[0036] Autoencoders (collectively referred to as "autoencoders") and variations thereof can be used as components in the methods and systems described herein. Autoencoders such as probabilistic autoencoders and variational autoencoders provide examples of generative models. In various embodiments, autoencoders may be used to implement directed graphical models, as opposed to undirected graphical models such as restricted Boltzmann machines.

[0037] In various embodiments, the autoencoder described herein includes two serialized components: an encoder and a decoder. The encoder can encode input data points as latent variables from which latent representations can be sampled. The decoder can decode the latent representations to generate random variables from which reconstructions of the original input can be sampled. The random variables may be modeled by a probability distribution, such as a normal distribution, a Laplace distribution, an elliptical distribution, a Student's t distribution, a logistic distribution, a uniform distribution, a triangular distribution, an exponential distribution, a reversible cumulative distribution, a Cauchy distribution, a Rayleigh distribution, a Pareto distribution, a Weibull distribution, a reciprocal distribution, a Gompertz distribution, a Gumbel distribution, an Erlang distribution, a lognormal distribution, a gamma distribution, a Dirichlet distribution, a beta distribution, a chi-squared distribution, or an F distribution, or variations thereof. Typically, the dimensionality of the input data and the output reconstruction may be the same.

[0038] In various embodiments, the autoencoders described herein are trained to reproduce their inputs, for example, by minimizing a loss function. Several training algorithms can be used to optimize, e.g., minimize, the reconstruction error and / or regularization error represented by the loss function. Examples of suitable training algorithms are described in more detail elsewhere herein and are otherwise known in the art, including, without limitation, backpropagation with stochastic gradient descent. Furthermore, several methods known in the art, such as dropout, sparse architectures, and denoising, may be used to prevent the autoencoder from overfitting to the training dataset and simply learning the identity function. The term "minimize" as used herein may include minimizing the absolute value of a term.

[0039] A trained autoencoder, such as a trained probabilistic or variational autoencoder, may be used to generate or simulate observable data values ​​by sampling from a modeled joint probability distribution to generate latent representations, and decoding these latent representations to reconstruct the input data points.

[0040] In one embodiment, the weights of the autoencoder are adjusted during training using optimization methods. In one embodiment, the weights are adjusted by optimizing, e.g., minimizing, a loss function using backpropagation with gradient descent. In one embodiment, individual layers of the autoencoder may be pre-trained, and the weights of the entire autoencoder are fine-tuned together.

[0041] In various embodiments, the systems and methods described herein may utilize deep network architectures, including but not limited to deep generative models, probabilistic autoencoders, variational autoencoders, directed graphical models, probabilistic networks, or variations thereof.

[0042] In various embodiments, the generative model described herein includes a probabilistic autoencoder having multiple components. For example, the generative model may have one or more of an encoder, a decoder, a sampling module, and an optional predictor (FIGS. 2A-2B). The encoder may be used to encode a compound representation, e.g., a fingerprint, as an output of a different form, e.g., a latent variable. During training, the encoder must learn an encoding model that specifies a nonlinear mapping of input x to latent variables Z. For example, if the latent variable Z is parameterized as Z = μz(x) + σz(x)εz, where εz = N(0,1), the encoder may output a pair of a vector of means and a vector of standard deviations. The sampling module may draw samples from the latent variable Z to generate the latent representation z. During training, the decoder may learn a decoding model that maps the latent variable Z to a distribution over x; i.e., the decoder may be used to convert the latent representation and label into a random variable X~, from which the sampling module can draw samples to generate the compound fingerprint x~. The latent variables or random variables may be modeled by an appropriate probability distribution function, such as a normal distribution, whose parameters are output by the encoder or decoder, respectively. The sampling module may sample from any appropriate probability distribution, such as a normal distribution, a Laplace distribution, an elliptical distribution, a Student's t distribution, a logistic distribution, a uniform distribution, a triangular distribution, an exponential distribution, a reversible cumulative distribution, a Cauchy distribution, a Rayleigh distribution, a Pareto distribution, a Weibull distribution, a reciprocal distribution, a Gompertz distribution, a Gumbel distribution, an Erlang distribution, a lognormal distribution, a gamma distribution, a Dirichlet distribution, a beta distribution, a chi-squared distribution, an F distribution, or variations thereof, or any other appropriate probability distribution function known in the art. The system may be trained to minimize a reconstruction error, which typically represents the negative likelihood that the input compound xD was drawn from the distribution defined by the random variables generated by the decoder, and / or a normalized error, which typically represents a penalty imposed on the complexity of the model.Without being bound by theory, because the encoding model must approximate the true posterior distribution p(Z|x), which can be intractable, an inference model may be used instead of using direct learning techniques. A variational autoencoder can use an inference model qφ(Z|x) that learns to approximate the true encoding distribution p(Z|x).

[0043] To train a VAE, a variational lower bound may be defined on the likelihood of the data: logpθ(x)=L(θ,φ,x) Here, φ denotes the encoding parameter and θ denotes the decoding parameter. From this definition, L(θ,φ,x)=-DKL(qφ(Z|x)||pθ(Z))+Eq_φ(Z|x)(logpθ(x|Z)) The result is as follows.

[0044] The first right-hand side (RHS) term, which is the Kullback-Leibler (KL) divergence of the approximate encoding model from the prior latent variable Z, can act as a regularization term. The second RHS term is usually called the reconstruction term. The training process can optimize L(θ,φ,x) for both the encoding parameters φ and the decoding parameters θ. The inference model (encoder) qφ(Z|x) is sometimes parameterized as a neural network: qφ(Z|x)=q(Z;g(x,φ)) where g(x) is a function that maps the input x to a latent variable Z, parameterized as Z = μZ(x) + σZ(x)εZ, where εZ = N(0,1) (Figure 5A).

[0045] The generative model (decoder) may be similarly parameterized as a neural network: pθ(x|Z)=p(x;f(Z,θ)) where f(Z) is a function that maps the latent variable Z to a distribution over x (Figure 5B). The decoder output X is X=μx(Z)+σx(Z)εx It may be parameterized as, where εx=N(0,1).

[0046] The inference and generative models may be trained simultaneously by optimizing a variational lower bound using backpropagation with gradient ascent (Figure 6). The optimization of the variational lower bound can serve to minimize a loss function that includes both the reconstruction error and the regularization error. In some cases, the loss function is or includes the sum of the reconstruction error and the regularization error.

[0047] 2A and 2B illustrate the use of generative models in which label information is provided to the model at two or more levels. Furthermore, machine learning models according to various embodiments of the present invention may be configured to accept compound representations and labels at the same layer (FIG. 17A) or different layers (FIG. 17B) of the machine learning model. For example, compound representations may be passed through one or more layers of encoders, and labels associated with each compound representation may be input at a layer after the encoder.

[0048] The systems and methods of the invention described herein can utilize representations of compounds, such as fingerprinting data. Label information associated with portions of a dataset may be missing. For example, for some compounds, assay data may be available that can be used directly in training a generative model. In other cases, label information may not be available for one or more compounds. In certain embodiments, the systems and methods of the invention partially or completely assign label data to compounds and associate it with their The generative model includes a predictor module for associating the fingerprint data with the compound data. In an exemplary embodiment of semi-supervised learning, the training dataset used to train the generative model includes both compounds with experimentally identified label information and compounds with labels predicted by the predictor module (FIG. 2B).

[0049] The predictor may comprise a machine learning classification model. In some embodiments, the predictor is a deep neural network having 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or more layers. In some embodiments, the predictor is a random forest classifier. In some embodiments, the predictor is trained with a training dataset that includes compound representations and their associated labels. In some embodiments, the predictor may have been previously trained with a set of compound representations and their associated labels that is different from the training dataset used to train the generative model.

[0050] Fingerprints that were initially unlabeled for one or more label elements may be associated with label element values ​​for one or more label elements by the predictor. In one embodiment, a subset of the training dataset may include fingerprints without associated labels. For example, compounds that may be difficult to prepare and / or difficult to test may be completely or partially unlabeled. In this case, various semi-supervised learning methods may be used. In one embodiment, a collection of labeled fingerprints is used to train a prediction module. In one embodiment, the predictor implements a classification algorithm trained with supervised learning. After the predictor is fully trained, unlabeled fingerprints may be input to the predictor to generate predicted labels. The fingerprints and their predicted labels are then added to a training dataset that may be used to train a generative model.

[0051] The predictor-labeled compounds may be used to train a first or second generative model. The predictor may be used to assign label element values ​​y to fingerprint feature vectors xD that lack label information. Through the use of predictors, generative models herein may be trained with training datasets that partially contain predicted labels. Once trained, generative models, described in more detail elsewhere herein, may be used to create generated representations of compounds, such as fingerprints. The generated representations of compounds may be created based on various conditions imposed by the desired labels.

[0052] In some embodiments, a generative model is used to generate representations of new compounds that were not presented to the model during the training phase. In some embodiments, a generative model is used to generate compound representations that were not included in the training dataset. In this way, novel compounds that may not be included in the compound database or that may not have been previously considered may be generated. Models trained with training sets that include real compounds may have several advantageous properties. Without being bound by theory, training with examples of real compounds, or drugs that are more likely to act as functional chemicals, can teach a model to generate compounds or compound representations that may possess similar properties with a higher probability than compounds drawn by hand or generated by computer using residual variation, for example.

[0053] Compounds associated with the generated representations may be added to compound databases, used in computational screening methods, and / or synthesized and tested in assays.

[0054] In some embodiments, a generative model is used to generate compounds that aim to resemble a specified seed compound. Compounds similar to the seed may be generated by inputting the seed compound and its associated label into an encoder. The potential representations of the seed compound and the desired label are then input into a decoder. Using the representation of the seed compound as a starting point, the decoder generates random variables from which samples can be drawn. The samples may include fingerprints of compounds that are expected to have some similarity to the seed compound and / or are likely to meet the requirements defined by the desired label.

[0055] In some embodiments, a generative model is used to generate compound representations by specifying a desired label, i.e., a set of desired label element values. Based on the modeled joint probability distribution, the generative model can generate one or more compound representations in which the represented compound is likely to meet the requirements of the specified label element values. In various embodiments, the methods and systems described herein may be used to train a generative model, generate compound representations, or both. The generation phase may follow the training phase. In some embodiments, a first party performs the training phase, and a second party performs the generation phase. The party performing the training phase may provide system parameters determined by the training to a separate computer system owned by the first party, or to a computer system owned by the second party and / or the second party, thereby enabling replication of the trained generative model. Thus, a trained computer system as described herein may refer to a second computer system configured by providing it with parameters obtained by training the first computer system using the training methods described herein, such that the second computer system can reproduce the output distribution of the first system. Such parameters may be transferred to the second computer system in a tangible or intangible form.

[0056] The training phase may involve using labeled fingerprint data to simultaneously train a generative model and a predictor.

[0057] In the generation phase, portions of the computer systems described herein, such as probabilistic decoders, may be used to create generated representations, e.g., fingerprints, of compounds. The systems and methods described herein can generate these representations in a manner that maximizes the probability of a desired outcome, e.g., a bioassay result, for a selected label associated with the generated representation. In some embodiments, the generated representations are first generated by drawing latent representations from a known distribution, such as a standard normal distribution. In some embodiments, comparative techniques are used in the generation phase. For example, a seed compound and its associated label may be input to an encoder, which outputs latent variables from which the latent representations can be sampled. The latent representations and desired labels may then be input together to a decoder. The training algorithms described herein may be adapted to the specific configuration of the generative models utilized within the computer systems and methods described in more detail elsewhere herein. It should be understood that methods known in the art, such as cross-validation, dropout, or noise removal, may be used as part of the training process.

[0058] In some embodiments, the predictor may use a classifier such as a random forest, a gradient boosted decision tree ensemble, or logistic regression.

[0059] A variety of suitable training algorithms can be selected for training the generative models of the present invention, which are described in more detail elsewhere herein. The appropriate algorithm may depend on the architecture of the generative model and / or the task that the generative model is desired to perform. For example, a variational autoencoder may be trained using a combination of variational inference and stochastic gradient ascent. It may be trained to optimize a variational lower bound.

[0060] The regularization constraint may be imposed by various methods: in some embodiments, methods known in the art such as dropout, denoising, or sparse autoencoders may be used.

[0061] <Generation procedure> In various embodiments, the methods and systems described herein are used to generate representations of compounds. These generated representations may not have been part of the training dataset used to train the model. In some embodiments, the compounds associated with the generated representations may be new to the generative model that created them.

[0062] The generated representations and / or associated compounds may be created from a generative model that was never presented with the generated representations and / or associated compounds, hi some embodiments, the generative model was not presented with the generated representations and / or associated compounds during the training phase.

[0063] In some cases, the methods and systems described herein may be used to output generated representations of compounds when creating generative models trained on a training dataset. Thus, information in the training dataset, such as the chemical structures of compounds and their properties, can inform the generation phase and generated representations of compounds.

[0064] In various embodiments, the generative models described herein generate representations of compounds that display activity and are likely to possess the properties specified by a desired label. For example, the desired label may include a specified activity in a particular bioassay test, such as activity against a particular receptor or enzyme. Compounds may be characterized by several molecular descriptors, such as formula, structure, electric density, or other chemical properties, or any other suitable molecular descriptor known in the art. Descriptors related to physical properties as well as compound drawings may be used. For example, the electric field of a ligand resulting from comparative molecular field analysis (CoMFA) may be used. Molecular descriptors may include, but are not limited to, molar refractive index, octynol / water partition coefficient, pKa, the number of atoms of a particular element such as carbon, oxygen, or halogen atoms, atom pair descriptors, the number of bonds of a particular type such as rotatable, aromatic, double, or triple bonds, hydrophilicity and / or hydrophobicity, the number of rings, the sum of positive partial charges on each atom, polarity, hydrophobicity, hydrophilicity, and / or water-accessible surface area, heat of formation, topological connectivity index, topological shape index, electronic topological state index, structural fragment count, surface area, packing density, van der Waals volume, refractive index, chirality, toxicity, topological indices such as Wiener index, Randic branching index, and / or Chi index, descriptors based on three-dimensional representations, etc. This information may be represented as a fingerprint for each compound. The methods and systems described herein train generative models with labels and compound representations to generate compound representations, such as fingerprints, that are expected to have specific properties associated with desired labels, e.g., labels that specify a desired outcome in a particular bioassay. In some embodiments, the generated representations are later used as lead or initial compounds in a hit-read procedure.

[0065] <Candidate generation (initial case)> In the initial case, the generation of candidate compounds is constrained only by the desired label y~. Therefore, initial generation may be used when there are no restrictions on the physical structure of the candidate compounds. Because the generated compounds are constrained only by the desired label y~, initial generation may be more likely to generate novel compounds that may not yet exist in the compound database. Such results may be useful in exploratory drug discovery research.

[0066] In various embodiments, an initial generation method is used utilizing only a sampling module and a decoder. The sampling module can draw samples from a specified probability distribution, which may differ from the probability distribution used to train the generative model. Figure 3 shows an example of initial generation in which the sampling module samples from a standard normal distribution. This generates a latent representation z, which may have no similarity to known compounds. The latent representation z and the desired label y~ may both be input to the decoder. From these inputs, the decoder can generate a random variable X~ across the distribution of molecular descriptors (e.g., fingerprints) that are likely to meet the requirements of the desired label y~. The sampling module then samples from this random variable to generate x~, which may be a fingerprint for the generated candidate compound.

[0067] <Candidate generation (comparison case)> In various embodiments, the systems and methods described herein are utilized to generate a representation, e.g., a fingerprint, of a compound using a seed compound as a starting point. The seed compound may be a known compound for which certain experimental results are known, and the structural characteristics of the generated compound may be expected to show some similarity to those of the seed compound. For example, the seed compound may be an existing drug being repurposed or tested for off-label use, and it may be desirable for the generated candidate compound to retain some of the seed compound's beneficial activities, such as low toxicity and high solubility, but exhibit different activity in other assays, such as binding to a different target, as required by the desired label. The seed compound may also be a compound that has been physically tested to possess a subset of the desired label results, but where improvements in certain other label results, such as reduced toxicity, improved solubility, and / or improved ease of synthesis, are desired. Thus, comparative generation may be used to generate compounds that possess structural similarity to the seed compound but aim to exhibit different label results, such as desired activity, in a specific assay.

[0068] In various embodiments, a representation, such as a fingerprint of a seed compound, and its associated label are input to a generative model, such as a trained probabilistic or variational autoencoder. For example, when a fingerprint of a seed compound and its associated label are input to an encoder, the encoder can output a latent variable Z. From the latent variable Z, a sampling module can draw samples to create a latent representation of the seed compound and its label information. This latent representation and desired label y may be input to a decoder, which can decode them to generate a random variable defined over a space of possible fingerprint values. The sampling module can sample from the random variable to generate the compound representation.

[0069] A generative model or its individual components may be configured to accept a desired label y~ and a latent representation generated based on a seed compound. The original label yD associated with the seed compound and the desired label y~ may differ to varying degrees. In some cases, yD and y~ may differ only with respect to one or more specified aspects, such as with respect to toxicity, but not with respect to other aspects. For example, yD and y~ may be the same for a first bioassay and a second bioassay, but different for a third bioassay. In some embodiments, a seed compound may not have an experimentally determined associated label. In this case, the label yD of the seed compound may be predicted by the prediction module.

[0070] 4A and 4B provide an exemplary illustration for creating a compound representation generated based on a seed compound and associated labels. In this embodiment, both the desired label ŷ of the seed compound and the potential representation z are input to the decoder. According to this embodiment, the decoder outputs a pair of a vector of the mean and a vector of the standard deviation. These vectors are used to generate the compound representation z for the seed compound x. D Similar to (1), a random variable X~ can be defined that models the distribution from which a compound associated with the desired label y~, or possibly an approximate variant of the desired label y~, is likely to be drawn. Samples may be drawn from the random variable X~ to generate a compound representation x~, e.g., in the form of a fingerprint. In various embodiments, a generative network is trained such that the generated compound x~ is likely to have the set of activities and properties specified in the desired label y~.

[0071] In some embodiments, a compound corresponding to the generated representation is chemically prepared. The prepared compound may be tested for a desired property or activity, such as that specified in the label used in the generation phase. The prepared compound may be further tested for additional properties or activities. In some embodiments, the prepared compound may be tested for clinical use, e.g., in multistage animal and / or human use trials.

[0072] <Label Source> Training data may be compiled from compound and associated label information from databases such as PubChem (http: / / pubchem.ncbi.nlm.nih.gov / ). Data may also be obtained from drug screening libraries, combinatorial synthesis libraries, etc. Label elements associated with an assay may include cellular assays and biochemical assays, and in some cases may include multiple related assays, e.g., assays for different families of enzymes. In various embodiments, information about one or more label elements may be obtained from resources such as compound databases, bioassay databases, toxicity databases, clinical records, cross-reactivity records, or any other suitable database known in the art.

[0073] <Fingerprint collection> The compound may be preprocessed to create a representation, e.g., a fingerprint, that can be used in conjunction with the generative model described herein. In some cases, the chemical formula of the compound may be recovered from the representation without degeneracy. In other cases, a representation may map to more than one chemical formula. In still other cases, there may be no identifiable chemical formula that can be inferred from the representation. A nearest neighbor search may be performed in the representation space. The identified neighbors may lead to a chemical formula that can approximate the representation generated by the generative model.

[0074] In various embodiments, the methods and systems described herein utilize fingerprints to represent compounds in the input and / or output of a generative model.

[0075] Various types of molecular descriptors may be used in combination to represent a compound as a fingerprint. In some embodiments, the compound representation including the molecular descriptors is used as input to various machine learning models. In some embodiments, the compound representation includes at least, or at least about, 50, 100, 150, 250, 500, 1000, 2000, 3000, 4000, 5000, or more molecular descriptors. In some embodiments, the compound representation includes less than 10,000, 7,500, 5,000, 4000, 3000, 2000, 1000, 500, 250, 150, 200, or 50 molecular descriptors.

[0076] Molecular descriptors are used for all chemicals in all assays and / or threshold combinations. The data may be normalized over the entire composite.

[0077] A compound fingerprint typically refers to a sequence of molecular descriptor values ​​that contain information about the chemical structure of a compound (e.g., in the form of a connectivity table). Thus, a fingerprint can be a shorthand expression that identifies the presence or absence of some structural or physical characteristics in the original chemistry of a compound.

[0078] In various embodiments, fingerprinting includes hash-based or dictionary-based fingerprinting. Dictionary-based fingerprinting relies on a dictionary. A dictionary typically refers to a set of structure fragments used to determine whether each bit in a fingerprint sequence is "on" or "off." Each bit in a fingerprint can represent one or more fragments that must be present in the main structure for that bit to be set in the fingerprint.

[0079] Some fingerprinting applications may use a "hash coding" technique. Thus, the fragments present in the molecule may be "hash coded" to fingerprint bit positions. Hash-based fingerprinting may allow all of the fragments present in the molecule to be encoded in the fingerprint. However, hash-based fingerprinting may cause several different fragments to set the same bit, which may lead to ambiguity.

[0080] Generating a representation of a compound as a fingerprint may be achieved by using publicly available software suites from various vendors (see, for example, www.talete.mi.it / products / dragon_molecular_descriptor_list.pdf, www.talete.mi.it / products / dproperties_molecular_descriptors.htm, www.moleculardescriptors.eu / softwares / softwares.htm, www.dalkescientific.com / writings / diary / archive / 2008 / 06 / 26 / fingerprint_background.html, or vega.marionegri.it / wordpress / resources / chemical-descriptors).

[0081] <Method> A key advantage of the present invention is the ability to discover drugs that may have fewer side effects. The generative models described herein may be trained by including in a training dataset compound activity for specific assays where specific results are known to cause adverse and / or toxic reactions in humans or animals. Thus, the generative model may be taught the relationship between compound representations and beneficial and undesired effects. In the generation phase, desired labels y~ input into the decoder can identify desired compound activity in assays associated with beneficial and / or undesired side effects. The generative model can then generate representations of compounds that simultaneously satisfy both beneficial and toxic / side effect requirements.

[0082] By simultaneously meeting desired outcomes for beneficial effects and unwanted side effects, the methods and systems described herein enable more efficient searches in the early stages of the drug discovery process, potentially reducing the number of clinical trials that fail due to unacceptable side effects of test drugs, which may lead to a reduction in both the duration and cost of the drug discovery process.

[0083] In some embodiments, the methods and systems described herein are used to find new targets for already existing compounds. A generative network can create a generated representation for a compound based on a desired label, where the compound is known to have another effect. Thus, a generative model trained with multiple label elements can respond to the use of a generation phase by inputting desired labels for different effects to generate a representation for a compound known to have a first effect, effectively identifying a second effect. Thus, a generative model may be used to identify a second label for an existing compound. Reusing clinically tested compounds can potentially lower risk during clinical research, and furthermore, compounds determined in this way can be particularly valuable because their efficacy and safety can be demonstrated effectively and inexpensively.

[0084] In some embodiments, the generative models herein may be trained to learn values ​​for types of label elements in a non-binary manner. The generative models herein may be trained to recognize higher or lower levels of compound effect for a particular label element. Thus, the generative models may be trained to learn the level of efficacy and / or the level of toxicity or side effects for a given compound.

[0085] The methods and systems described herein are particularly powerful in generating representations of compounds, including compounds not presented in the model and / or not previously existing compounds, thereby expanding compound libraries. Additionally, various embodiments of the present invention also facilitate traditional drug screening processes by allowing the output of generative models to be used as input datasets for virtual or experimental screening processes.

[0086] In various embodiments, the generated representations relate to compounds that have similarity to compounds in the training dataset. Similarity may include various aspects. For example, a generated compound may have a high degree of similarity to a compound in the training dataset, but may be much more likely to be chemically synthesizable and / or chemically stable than a compound in the training dataset to which it is similar. Furthermore, a generated compound may be similar to a compound in the training dataset, but may have a much higher likelihood of having a desired effect and / or being free of undesired effects than existing compounds in the training dataset.

[0087] In various embodiments, the methods and systems described herein generate compounds or their representations taking into account their ease of synthesis, solubility, and other practical considerations. In some embodiments, the generative model is trained using label elements, which may include solubility or synthesis mechanism. In some embodiments, the generative model is trained using training data including synthesis information or solubility levels. Desired labels related to these factors may be used in the generation phase to increase the likelihood that the generated compound representation will relate to a compound that behaves according to the desired solubility or synthesis requirements. In various drug discovery applications, multiple candidate fingerprints may be generated. The collection of generated fingerprints can then be used to synthesize actual compounds that can be used in high-throughput screening. Prior to compound synthesis and HTS, it is useful to evaluate whether the generated fingerprints have desired assay results and / or structural properties. The generated fingerprints may be evaluated (in comparative generation) based on their predicted results and their similarity to the seed compound. If the generated fingerprints have the desired properties, they may be ranked based on their druglikeness.

[0088] Additional system modules can be introduced into these procedures: a comparison module may be used to compare two fingerprints or two sets of assay results; a ranking module may be used to rank members of a set of fingerprints by drug-likeness score; a classifier may be used to classify compound fingerprints by assigning drug-likeness scores; and an ordering module may be used to order a set of scored fingerprints.

[0089] In various embodiments, the methods and systems of the present invention may be used to evaluate the predicted outcomes of generated compounds and / or rank generated compounds. In various embodiments, the predicted assay outcomes of generated fingerprints are compared to desired assay outcomes. Fingerprints with predicted outcomes that match the desired assay outcomes may be ranked for further consideration, for example, by a drug-likeness score.

[0090] 7 depicts an example of a single-step evaluation and ranking procedure according to various embodiments of the present invention. The generated representation x may be created according to various methods described herein, e.g., by initial generation or comparative generation. The generated representation x, e.g., a representation in the form of a fingerprint or related compounds, may be input to a trained predictor module. (The predictor module may have been trained, e.g., during a semi-supervised learning process for unlabeled data.) The predictor module may output a predicted set y of assay results for the generated representation x.

[0091] The predicted assay result ŷ and the desired assay result ŷ may be input to a comparison module (FIG. 7). The comparison module may be configured to compare the predicted result with the desired result. If the comparison module determines that the predicted result is the same as the desired result, x̂ may be added to the set of unranked candidates U; otherwise, x̂ may be rejected. The unranked set may be ranked by a ranking module, as described in more detail elsewhere herein.

[0092] In various embodiments, the methods and systems of the present invention may be used to evaluate generated representations, for example fingerprints generated via comparative generation.

[0093] In comparative generation, a seed compound may be used to generate a new fingerprint similar to the seed. Following the comparative generation process, an evaluation step may be used to determine whether the generated fingerprint is sufficiently similar to the seed. In this embodiment, a comparison module may be used to compare corresponding parameters of two fingerprints, typically the generated representation and the fingerprint of the seed compound. If a threshold or threshold similarity of the identity parameters is achieved, the two fingerprints may be marked as sufficiently similar.

[0094] 8 depicts an example of a method for evaluating generated fingerprints and their predicted results, according to various embodiments of the present invention. Thus, the generated representation x and the associated seed compound representation xD are input to a comparison module. The comparison module first compares x and x for similarity. D It may be constructed to compare x with x. D If the comparison module determines that x~ is sufficiently similar to x, then x~ may be retained. Otherwise, x~ may be rejected.

[0095] In various embodiments, the retained generated representation x̂ may be input to a predictor module, as described in more detail elsewhere herein. The predictor module may be used to output a predicted label ŷ. A comparison module may be used to compare the predicted label ŷ with a desired label ŷ (the desired label ŷ is the seed compound representation x̂). D (It may have been used to create the generated representation during comparison generation with ŷ). For a generated representation x~, if the comparison module finds sufficient similarity between ŷ and y~, x~ may be added to an unranked candidate set U. The unranked set U may be ranked by a ranking module. The ranking module may output a ranked set R containing the generated representations.

[0096] The systems and methods described herein, in various embodiments of the invention, utilize a ranking module, which may be configured to have several functions, including assigning a drug-likeness score to each fingerprint and ranking the collection of fingerprints according to their drug-likeness scores.

[0097] A common existing method for assessing a compound's drug-likeness is to confirm its compliance with Lipinski's Rule of Five. Additional factors, such as the logarithm of the partition coefficient (logP) and the molar refractive index, may also be used. However, simple filtering methods, such as whether a compound's logP and molecular weight fall within a certain range, may only allow for classification analysis that assigns a pass or fail value. Furthermore, in some cases, standard drug-likeness characteristics may not provide sufficient discriminatory power to accurately assess a compound. (For example, the highly successful drugs Lipitor and Singulair both failed to pass two or more of Lipinski's rules and would have been rejected by a simple filtering process.)

[0098] In some embodiments, the desired ranking of compounds may be achieved by a ranking module described herein. Ranking modules according to various embodiments of the present invention evaluate compound representations, such as fingerprints, based on their latent representations, rather than relying on filtering standard drug-likeness features. Without being bound by theory, the latent representation of a compound's fingerprint represents a high-level abstraction and non-linear combination of features that can provide a more accurate description of a compound's behavior than standard drug-likeness features can provide.

[0099] 9 depicts an exemplary illustration of a training method for a ranking module. In various embodiments, an autoencoder is trained on a large set of compound representations. A latent representation generator (LRG) can form the first part of the autoencoder in a similar position to the encoder. The LRG can be used to generate latent representations (LRs) of the compounds. The latent representations may be input to a classifier. The classifier may be trained using supervised learning. The training dataset for the classifier may include labeled drug and non-drug compounds. The classifier may be trained to output a continuous score representing the drug-likeness of the compound.

[0100] 10 depicts an exemplary illustration of a ranking module including an LRG, a classifier, and an ordering module, according to various embodiments of the present invention. Members of an unranked set of compound representations may be input to a latent representation generator (LRG), and the latent representations may be input to a classifier. The classifier may be configured to provide a drug-likeness score for each potential representation. The compound representations and / or related compounds may be ordered, for example, from highest drug-likeness score to lowest drug-likeness score. The ranking module may be used to provide a compound representation, e.g., a fingerprint, and / or a ranked set of compounds as output.

[0101] In various embodiments of the present invention, the systems and methods described herein relate to exploring novel compound space through initial generation and comparative generation. According to various embodiments, initial generation and comparative generation may be utilized in sequence. The systems and methods described herein may thus be used to generate novel compounds, or representations, e.g., fingerprints, that satisfy a particular set of assay results. Similar compounds within the representation space around the compound representation may be explored using the systems and methods described herein. For example, an initial compound representation may be generated using a process of initial generation or comparative generation with a desired label, and one or more generated representations may be output. The compound space around the generated representation may then be explored around these initial representations. According to various embodiments, initial generation and comparative generation may be used in sequence.

[0102] FIG. 11 depicts an exemplary illustration of sequentially using initial generation and comparison generation. Such a combination may be used to explore compound space around an initial compound associated with a desired label. Thus, based on a desired assay result y, a fingerprint x may be generated using initial generation. Previously unknown compounds may be prioritized by applying a filter through the use of a comparison module. The comparison module may compare x to a database of known compounds. If the comparison module determines that x is already present in the database of known compounds, x may be flagged for rejection. If the comparison module determines that x is a previously unknown compound, x may be input into a predictor. The predictor may generate a predicted assay result y for x.

[0103] A new representation x+ may be generated by using representation x~ and its predicted assay result y^ as a seed for comparison generation. A predictor may be used to generate a predicted assay result y+ for x+. A comparison module may be used to determine whether y+ is the same as or similar to the desired assay result y~. If identity or sufficient similarity is found, x+ may be marked for retention. The retained representation may be added to a set U of unranked candidates. Any desired number of fingerprints x+ may be generated from the initial seeds of x~ and y^ by repeated applications of comparison generation.

[0104] The unranked set U of candidate representations may be input to a ranking module, which may output a ranked set R of compound representations and / or related compounds.

[0105] In various embodiments, the systems and methods described herein may be used to identify compound properties that can affect the outcome of a particular assay. Without being bound by theory, a small number of specific structural features may be modifications that alter the performance of a compound in a particular assay. In various embodiments, the systems and methods described herein provide a process for identifying candidate modifications associated with the performance of a compound in a particular assay. The identified candidate modifications may be used as a starting point for matched molecular pair analysis (MMPA).

[0106] In an exemplary embodiment, two generation processes, e.g., two initial generation processes, are performed utilizing different seed labels. In one, the desired label y~ is used as a positive seed. In the other, the opposite label y* is used as a negative seed. For example, if y~ is a single binary assay result, the negative seed y* may be the opposite result for that assay. Without being bound by theory, using a single assay result may introduce unnecessarily large variability in the resulting generated fingerprint. To reduce variability, a vector of label elements may be used as the positive seed y~. For example, if y~ is composed of a vector of label element values, e.g., In the assay results of , y* may differ from y~ by one label element value.

[0107] Thus, in various embodiments, two sets of compound representations, A and B, may be generated from the two generation processes. Set A may include compounds generated from the positive seed y~. Set B may include compounds generated from the negative seed y*. The two sets of compound representations may be input to a comparison module. The comparison module may be configured to identify compound representation parameters that are most likely to account for differences in the labels or label elements of interest. The comparison module is described in further detail elsewhere herein.

[0108] In some embodiments, two or more initial generation processes, each using a different label, may be used to generate multiple sets of compounds in a manner similar to that described above for the embodiment with two generation processes. These sets may be analyzed to identify significant variations in the compound representations that may be associated with different label values.

[0109] In various embodiments, the systems and methods described herein may be used to search for mutations associated with a desired label element value for a particular compound, i.e., mutations in a particular compound that may be responsible for a particular label element value. In some embodiments, the method is implemented by running two comparative generation processes using the same seed compound representation but different target labels or label element values. The two comparative generation processes may be run in parallel, and two sets of compound representations may be generated. The comparison module may be used to identify specific structural differences between the representations generated with positive results and the representations generated with negative results (Figure 13).

[0110] The generated representations may first be evaluated by their similarity to the seed compound. If they are sufficiently similar, a predictor module may be used to determine a predicted label or label element value for each representation. The predicted label or label element value may be compared to the target label or label element value (Figure 13).

[0111] The comparison-generation process may be performed iteratively. The resulting candidate-generated representations may be grouped into two sets, A and B, with desired cardinality. Members of A may be compared to members of B by a comparison module. The comparison module may identify homogeneous and heterogeneous structural variants between the two sets. The comparison module is described in more detail in the Examples below and elsewhere herein. These structural variants may be used as starting points for further analysis via MMPA.

[0112] In some embodiments, three or more comparative generation processes are used to generate representations using different labels for each process. As described above for the embodiment with two generation processes, multiple sets of compounds may be generated. These sets may be analyzed to identify significant variations in the compound representations that may be associated with different label values.

[0113] In various embodiments, the systems and methods described herein utilize a comparison module. The comparison module may be configured to have a single function or multiple functions. For example, the comparison module may combine two functions into one module, such as (1) determining whether two vectors of labels or two compound representations are similar or identical, and (2) comparing two sets of compound representations to identify parameters that are most likely to cause changes in specified labels or label element values. In other embodiments, the comparison module may have a single function or three or more functions.

[0114] In some embodiments, the comparison module is configured to perform a comparison of two objects for similarity or identity. The comparison may include a simple pairwise comparison for similarity or identity, in which corresponding elements of two objects, such as two vectors of assay results or two fingerprints, are compared. A threshold, such as a user-specified threshold, may be used to determine whether two objects pass or fail the comparison. In some embodiments, the systems and methods described herein may be used to set the threshold, for example, by determining a threshold that results in a viable grouping of a training set of objects.

[0115] In some embodiments, the comparison module is configured to perform the comparison on latent representations output by a latent representation generator (LRG). The LRG may be used to encode compound representations, such as fingerprints, as latent representations. The resulting distributions of latent representations may be compared, and a determination of similarity or identity may be made.

[0116] In some embodiments, the comparison module is configured to compare sets of objects for the identification of significant compound variations. For example, when comparing two sets of fingerprints, several methods may be used to identify significant compound variations.

[0117] In some embodiments, the comparison module uses a linear model to identify important parameters. Without being bound by theory, interaction terms can be added to the model to address the possibility that interactions between parameters may contribute to differences in label or label element values, such as differences in particular assay results, toxicity, side effects, or other label elements described in more detail herein, or any other suitable label element known in the art.

[0118] In some embodiments, the comparison module is configured to utilize the Gini coefficient as a measure of inequality in the populations. The Gini coefficient may be calculated for one, some, or all parameters of the objects by calculating the average difference between all possible pairs of objects divided by the average size. Without being bound by theory, a large Gini coefficient for a parameter tends to indicate a high degree of inequality in that parameter between members of set A and members of set B. In various embodiments, a desired number of parameters with the largest Gini coefficients may be selected as parameters most likely to be associated with a change in label or label element value, e.g., assay result. Selection may select the top 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more parameters. In some embodiments, selection selects parameters with Gini coefficients above a threshold level or with a certainty above the threshold associated with a change in label or label element value.

[0119] In some embodiments, a classification tree may be used in parallel with the Gini coefficient calculation. The parameter with the largest Gini coefficient may be selected to be the root of the classification tree. The remainder of the classification tree may be learned, for example, by top-down induction. A desired number of important parameters may be identified by observing the behavior of the tree at an appropriate level.

[0120] If the cardinality of the two sets of fingerprints is low, the Gini coefficient may be calculated directly. Without further elaboration, as the cardinalities of sets A and B become large, direct calculation of the Gini coefficient may become difficult or impractical due to combinatorial explosion. The systems and methods described herein may be configured to utilize methods that reduce the number of required pairwise comparisons between A and B, for example, by applying clustering methods. Thus, the parameter Gini coefficient may be calculated by pairwise comparisons between the centroids of the clusters resulting from clustering the members of A and the members of B.

[0121] Without being bound by theory, because compound representations have a large number of parameters, e.g., in the thousands or more, the dimensionality may make directly clustering the members of A and B infeasible. Representations of sets A and B in a space with thousands of dimensions may be very sparse. A large number of data points may be required to achieve statistically significant clustering in the compound representation space. In various embodiments, the systems and methods of the present invention can address these issues by utilizing alternative clustering methods. In some embodiments, the methods and systems of the present invention are used to cluster vectors containing latent representations of members of A and B. These latent representations may be of lower dimensionality. Clustering the latent representations may be even more advantageous because the latent representations can incorporate nonlinear combinations of parameters of members of A and B. This ability may, in some cases, provide the latent representations with superior ability to explain the behavior of compounds or specific features thereof, such as specific chemical residues.

[0122] In various embodiments, the systems and methods of the present invention are used to cluster compound representations by performing clustering of related latent representations. For example, the systems and methods of the present invention may be used to calculate the Gini coefficient in the latent representation space using k-medoid clustering.

[0123] FIG. 14 depicts an exemplary illustration of a comparison module using k-medoid clustering. Thus, latent representations may be generated for members of sets A and B. For example, a latent representation generator (LRG) may be used to encode members of sets A and B as latent representations to form latent representation sets AL and BL, respectively. A clustering method, such as k-medoid clustering, may be applied to the members of the latent representation sets. Following clustering, centroids of the clustered sets may be extracted to form centroid sets AC and BC of latent representations. Without being bound by theory, because the centroids in some clustering methods, such as k-medoid clustering, are actual members of the original dataset, in applying such clustering methods, sets AC and BC are expected to contain latent representations of members of the original sets A and B. Compound representations corresponding to members of AC and BC can be searched to form two sets of fingerprints, AF and BF. The cardinality of AF and BF may be significantly lower than the cardinality of the original sets A and B. Members of sets AF and BF may be used to identify compound modifications that may cause changes in label or label element values, such as assay results.

[0124] In some cases, the systems and methods of the present invention may be used to calculate the Gini coefficient using k-means clustering in the latent representation space. Figure 15 depicts an exemplary illustration of a comparison module using k-means clustering. Thus, members of sets A and B may be encoded as latent representations, as may be the case with the k-medoids method. For example, a latent representation generator (LRG) may be used to encode members of sets A and B as latent representations to form latent representation sets AL and BL, respectively. k-means clustering may be applied to the members of the latent representation sets. The centroids resulting from the k-means clustering are the weights of the latent representations. The centroids may be extracted to form centroid sets AC and BC. Without being bound by theory, the members of centroid sets AC and BC may often not be encoded latent representations that correspond to some members of the original sets A and B. However, the members of the centroid sets may be decoded to generate corresponding members in a compound representation space. For example, a latent representation decoder module (LRD) may be used to generate compound representations, e.g., fingerprints, that correspond to the centroids, which may be grouped within sets AF and BF, respectively.

[0125] 9 depicts the training of an autoencoder on a large set of compound representations in an exemplary embodiment. A latent representation decoder (LRD) can form the second part of the autoencoder in a similar position to the decoder. That is, during training of the autoencoder, the decoder can learn to regenerate the original compound representations from the latent representations.

[0126] The generated representations in AF and BF may be of relatively low cardinality compared to the original sets A and B. The members of the generated representations in AF and BF may be used to identify important compound transformations.

[0127] In various embodiments, the systems and methods described herein handle inputs of different composition or length, e.g., labels with different label elements and / or different numbers of label elements. For example, during training, different compounds in the training set may have labels of different lengths. Well-known drugs may have more assay results than new compounds. Additionally, during the generation phase, the desired labels y are compared to the labels y used to train the model. D It may be shorter than that.

[0128] In various embodiments, a masking module, such as one utilizing a probabilistic mask, may be used to homogenize different objects, e.g., different labels, with respect to length and / or composition. In some cases, methods similar to dropout can be used to make a probabilistic or variational autoencoder robust to missing values.

[0129] In various embodiments, the probabilistic mask is generated by subtracting the training label y from the training label y prior to training. D A masking module may be used to generate a masked version of a desired label. For example, a masking module may be configured to process various labels before inputting them into a generative model. If two labels have different numbers of label element values, a masking module may be used to add zero values ​​to all label elements that are missing. Additionally, probabilistic masking may be used to randomly zero out label element values ​​during training. By training a generative model in this manner, the model may initially be able to handle training labels and desired labels that may have different numbers of label elements.

[0130] An exemplary embodiment of the masking module operates on assay results having a binary outcome. The assay results can be encoded as label element values ​​of -1 for inactivity and 1 for activity. The masking module can apply a probabilistic mask to each label element value in the training dataset. For a mask, the label is y D = (m1y1,m2y2,...), where yi is the unmasked label element, mi is the mask for yi, and mi takes values ​​of 0 or 1. For training, the values ​​of mi may be set randomly, or they may be set according to the empirical probability that the corresponding label element value is not present.

[0131] If miyi=0, then for the forward pass in backpropagation, no correction may be necessary because a value of 0 may not contribute to the activation of the next layer. To avoid propagating errors to nodes with missing input values ​​during the backward pass, input nodes with missing values ​​may be flagged and disconnected during the backward pass. This training method can enable generative models to handle labels of different lengths during the training and generation process.

[0132] <Computer System> The present invention also relates to apparatus for performing the operations herein. This apparatus may be specially constructed for the required purposes, or it may comprise a general-purpose computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored on a computer-readable storage medium such as any type of disk, including, but not limited to, floppy disks, optical disks, CD-ROMs, and magneto-optical disks, read-only memory (ROM), random-access memory (RAM), EPROM, EEPROM, magnetic or optical cards, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.

[0133] The descriptions presented herein are not inherently related to any particular computer or other apparatus. In addition to general-purpose systems, more specialized apparatuses may be constructed to practice various embodiments of the present invention. Additionally, the present invention is not described with reference to any particular programming language. It will be appreciated that a variety of programming languages ​​may be used to implement the teachings of the present invention as described herein. A machine-readable medium includes any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer). For example, machine-readable media include read-only memory ("ROM"), random-access memory ("RAM"), magnetic disk storage media, optical storage media, flash memory devices, electrical, optical, acoustic, or other forms of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.), and the like.

[0134] FIG. 16 is a block diagram of an exemplary computer system capable of performing one or more operations described herein. With reference to FIG. 16, the computer system may include an exemplary client or server computer system. The computer system may include a communication mechanism or bus for communicating information and a processor coupled to the bus for processing information. The processor may include a microprocessor, but is not limited to a microprocessor such as, for example, a Pentium, PowerPC, or Alpha. The system further includes a random access memory (RAM) or other dynamic storage device (referred to as main memory) coupled to the bus for storing information and instructions to be executed by the processor. The main memory may also be used to store temporary variables or other intermediate information during execution of instructions by the processor. In various embodiments, the methods and systems described herein utilize one or more graphical processing units (GPUs) as processors. The GPUs may be used in parallel. In various embodiments, the methods and systems of the present invention utilize a distributed computing architecture having multiple processors, such as multiple GPUs.

[0135] The computer system may also include a read-only memory (ROM) and / or other static storage device coupled to the bus for storing static information and instructions for the processor, and a data storage device, such as a magnetic or optical disk and its corresponding disk drive. The data storage device is coupled to the bus for storing information and instructions. In some embodiments, the data storage device may be located at a remote location, for example, in a cloud server. The computer system may further be coupled to a display device, such as a cathode ray tube (CRT) or liquid crystal display (CD), coupled to the bus for displaying information to a computer user. An alphanumeric input device, including alphanumeric and other keys, may also be coupled to the bus for communicating information and command selections to the processor. A further user input device is a cursor controller, such as a mouse, trackball, trackpad, stylus, or cursor direction keys, coupled to the bus for communicating directional information and command selections to the processor and for controlling cursor movement on the display. Another device that may be coupled to the bus is a hardcopy device, which may be used to print instructions, data, or other information on a medium, such as paper, film, or a similar type of medium. Additionally, audio recording and playback devices such as speakers and / or microphones may optionally be coupled to the bus for audio interfacing with the computer system. Another device that may be coupled to the bus is a wired / wireless communication capability for communication to a telephone or handheld palm device.

[0136] It should be noted that any or all of the components of the system and associated hardware may be used in the present invention, however, it will be appreciated that other configurations of computer systems may include some or all of the devices. [Example]

[0137] <Input data for the encoder during training> In one example, the data may be a compound representation (x), such as a fingerprint, that includes a feature vector of molecular descriptors. D ), and the label associated with the represented compound (y D ) to the encoder. The input pair to the encoder is IE = (xi D , yi D ) and xi D is the number of dimensions dim_xi D is a real-valued vector with yi D is the corresponding xi D This shows the label data for D Number of dimensions of dim_xi D may be fixed across the entire training dataset. D The elements of y can be scalars or vectors, possibly with any dimension. D The label element values ​​in can be continuous or binary.

[0138] According to the explanation in this example, x has dimension 10 D , and y containing a single label element value D , then example input data could be: x D =(1.2,-0.3,1.5,4.3,-2.9,1.3,-1.5,2.3,10.2,1.1), y D =3, The input to the encoder is IE=((1.2,-0.3,1.5,4.3,-2.9,1.3,-1.5,2.3,10.2,1.1),3) is. [Example]

[0139] <Encoder output during training> An exemplary output structure for the encoder is described. Given an IE = (xi D , yi D), the encoder outputs a pair of real-valued vectors of the mean μE,i and standard deviation σE,i, denoted as OE=(μE,i,σE,i)=((μE,i,1,···,μE,i,d),(σE,i,1,···,σE,i,d)). The dimensions of the vectors μE and σE are the same in this example. However, the dimensions of the vectors μE and σE are different depending on the dim_xi D , or dim_xi D + dim_yi D The OEs may differ from the OEs. The OEs are provided by the encoder in a deterministic manner. For a given set of IEs and encoder parameters, a single OE pair is provided. For a dimensionality of 4, exemplary outputs of the encoder are given by μE = (1.2, -0.02, 10.5, 0.2) and σE = (0.4, 1.0, 0.3, 0.3). [Example]

[0140] <Creating the latent variable Z during the training process> In this example, the mean and standard deviation output by the encoder define a latent variable Z = (N(μE,i,1,σE,i,1), ,N(μE,i,d,σE,i,d)), where μE,i and σE,i are vectors output by the encoder and N denotes the normal distribution. For example, if the encoder output includes μE = (1.2, -0.02, 10.5, 0.2) and σE = (0.4, 1.0, 0.3, 0.3), the sampling module can define the latent random variable Z = (N(1.2, 0.4), N(-0.02, 1.0), N(10.5, 0.3), N(0.2, 0.3)). [Example]

[0141] <Creating latent representations using the sampling module during the training process> An exemplary sampling module draws a sample from a probability distribution, such as that defined by a latent variable Z and a random variable X, or multiple samples from a set of probability distributions. In this example, the sampling module can draw samples from the latent variable Z to generate a latent representation z having the same dimensionality as the latent variable Z. In this example, a single latent representation z is drawn from the latent variable Z. For Z=(N(1.2,0.4),N(−0.02,1.0),N(10.5,0.3),N(0.2,0.3), an exemplary latent representation vector z is z=(0.9,−0.1,10.1,0.1). If desired, the sampling module can draw multiple latent representations z from a single latent variable Z. [Example]

[0142] <Input to the decoder during training (ID)> In this example, the decoder generates an ordered pair (z, y D ), where z is a latent representation sampled from the latent random variable Z, and y D is a label. In this example, the label y D is the input feature vector x D is the same as the label associated with y. D is input twice in the training process, once to the encoder and once to the decoder. For example, ID might contain the pair ((0.9,-0.1,10.1,0.1),3).

[0143] The input layers of both the encoder and decoder are configured to be able to receive both the fingerprint and its associated label. During comparison generation, this configuration facilitates the use of two different input labels: the original label y D is input to the encoder and the desired label y is input to the decoder. [Example]

[0144] <Decoder output during training> In this example, the decoder generates a real-valued vector of means μD,i and a real-valued vector of standard deviations σ D,i Generates as output a pair: OD = (μ D,i ,σ D,i )=((μ D,i ,1···,μ D,i,d ),(σ D,i,1 ,···,σ D,i,d )) In this example, the vector μ D and σ D The dimension of is the feature vector x input to the encoder. D For example, dim_xi D = 10, the decoder will D =(1.2,-0.3,1.5,4.3,-2.9,1.3,-1.5,2.3,10.2,1.1), μ D =(1.1,-0.2,1.1,3.9,-3.5,0.1,-2.0,1.9,9.3,1.0) and σ D =(0.1,0.3,0.2,0.5,1.0,0.5,1.0,0.2,0.1,1.0) can be output.

[0145] From the decoder output, the latent variable X~ is given by X~=(N(μ D,i,1 ,σ D,i,1 ),···,N(μ D,i,d ,σ D,i,d )), and μ D,i and σ D,i is the vector output by the decoder. For example, μ D =(1.1,-0.2,1.1,3.9,-3.5,0.1,-2.0,1.9,9.3,1.0) and σ D =(0.1,0.3,0.2,0.5,1.0,0.5,1.0,0.2,0.1,1.0), then X~=(N(1.1,0.1),N(-0.2,0.3),···,N(1.0,1.0)). The sampling module can then draw a sample x from X~, where x is the generated representation of the compound. [Example]

[0146] <Sampling latent representation z from a standard normal distribution in the initial generation step> This example concerns the initial generation process. In this example, latent representations z are drawn from a standard normal distribution N(0,1) by a sampling module. A single desired label y~ is used. For each compound representation to be generated by the model, a separate latent representation z is drawn from N(0,1). For example, if the user wishes to generate two compound representations, two separate latent representations z~ are drawn. z1 and z2 is drawn from N(0,1). If the dimension of z is 4, the sampling module, in one example, z1 =(0.2,-0.1,0.5,0.1) and z2 =(0.3,0.1,0,-0.3) can be derived. [Example]

[0147] <Input to the decoder in the initial generation process> In this example, the decoder is input with the latent representations z previously sampled from N by the sampling module, as well as the desired labels y~. The labels y~ may be specified by the user according to the desired properties and activities of the compounds represented by the generated fingerprint. The desired labels y~ are a subset of the label elements used to train the model, i.e., the labels y~. D It must contain the desired values ​​for the label elements contained in y. D If the masking module has fewer label elements than y, then the masking module can assign a value of 0 to the missing label elements of y before inputting y to the decoder. DIt may contain one or more values ​​of label elements that differ from the value of the corresponding label element in N. It is possible to draw multiple samples z from N to generate multiple x~ with a single desired label y~. It is also possible to generate more than one compound representation from a single latent representation z by inputting several pairs consisting of z and different desired labels y~ to the decoder and generating more than one random variable X~. [Example]

[0148] <Decoder output in the initial generation step> In this example, the decoder generates a real-valued vector of means μ D ~ and real-valued vector of standard deviations σ D ~ pair (μ D ~, σ D In this example, the vector μ D ~ and σ D The dimension of ~ is the feature vector x, which is the fingerprint used to train the model. D For example, x D If the dimension of is 10, the decoder will, in one example, D ~=(1.1,-0.2,1.1,3.9,-3.5,0.1,-2.0,1.9,9.3,1.0) and σ D Outputs ~=(0.1,0.3,0.2,0.5,1.0,0.5,1.0,0.2,0.1,1.0). [Example]

[0149] <Construction of random variable X~ in the initial generation procedure> From the decoder output, the probability variable X~ is X~=(N(μ D、i、1 ,σ D、i、1 ),···,N(μ D、i、d ,σ D,i,d )), and μ D,i and σ D,i is the vector output by the decoder. For example, μ D =(1.1,-0.2,1.1,3.9,-3.5,0.1,-2.0,1.9,9.3,1.0) and σD =(0.1,0.3,0.2,0.5,1.0,0.5,1.0,0.2,0.1,1.0), then X~=(N(1.1,0.1),N(-0.2,0.3),···,N(1.0,1.0)). [Example]

[0150] <In the initial generation process, we generate a representation x~ by sampling from the random variable X~> To generate the compound representation x~, the sampling module draws samples from the random variable X~. Defining X~ so that its dimensions are the same as the dimensions of the fingerprint feature vector used to train the model may allow the dimensions of the representation x~ to be the same as the dimensions of the fingerprint feature vector. If desired, multiple compound representations may be sampled from the random variable X~. For example, if the random variable X~=(N(1.1,0.1),N(-0.2,0.3),...,N(1.0,1.0)), four samples may be drawn from X~, resulting in four compound representations in one example. x1 ~=(1.0,-0.1,···,3.0), x2 ~=(1.2,-0.5,···,1.8), x3 ~=(1.0,-0.1,···,0.5), and x4 This results in ~=(0.9,0.3,···,1.1). [Example]

[0151] <Encoder input and output in the comparison generation procedure> In this example, the inputs to and outputs from the encoder are of the same type as those used in Examples 1 and 2 during training of the encoder and decoder. For example: x D =(1.2,-0.3,1.5,4.3,-2.9,1.3,-1.5,2.3,10.2,1.1), y D =3, μ E =(1.2,-0.02,10.5,0.2), and σE =(0.4,1.0,0.3,0.3) is.

[0152] However, whereas in Examples 1 and 2 the inputs to and outputs from the encoder are used to train a generative model, in this example they are used in the process of generating novel compound representations. [Example]

[0153] <Construction of latent variable Z and sampling of latent enemy expression z in the comparative generation procedure> In this example, the same procedure is used to define a latent variable Z and sample from Z to create a latent representation z as was used in Examples 3 and 4 above.

[0154] for example: μ E =(1.2,-0.02,10.5,0.2), σ E =(0.4,1.0,0.3,0.3), Z=(N(1.2,0.4),N(-0.02,1.0),N(10.5,0.3),N(0.2,0.3)), and z=(0.9,-0.1,10.1,0.1) is.

[0155] However, while in Examples 3 and 4, latent variables Z and latent representations z were used to train a generative model, in this example they are used in the process of generating compound representations. If desired, multiple latent representations z may be derived from the latent variables Z. [Example]

[0156] <Decoder input and output in the comparison and generation procedure> In this example, the same procedure is used to construct both the input to the decoder and the output of the decoder as was used in Examples 8 and 9. For example: ID=(z,y~), OD=(μ D ~,σ D ~), μD~=(1.1,-0.2,1.1,3.9,-3.5,0.1,-2.0,1.9,9.3,1.0), and σ D ~=(0.1,0.3,0.2,0.5,1.0,0.5,1.0,0.2,0.1,1.0) is.

[0157] As in Examples 9, 10, and 11, the output of the decoder is used to generate compound representations. However, whereas in Example 8 the latent representation z is drawn from a standard normal distribution, in this example it is drawn from latent variables Z, which are the distributions of the seed compounds x. D and a latent variable for its associated label yD. The sampling module draws samples from the latent variable Z to generate a latent representation z. One or more latent representations z may be drawn from the latent variable Z and paired with one or more desired labels y~ in various combinations to generate multiple outputs from the decoder. [Example]

[0158] <Construction of random variables X~ and sampling of compound representations x~ in comparative generation procedures> In this example, the same procedure as used in Examples 10 and 11 is used to define a random variable X̂ and generate a compound representation x̂ by sampling from X̂. For example: X~=(N(1.1,0.1),N(-0.2,0.3),...,N(1.0,1.0)), x1~=(1.0,-0.1,···,3.0), x2~=(1.2,-0.5,···,1.8), x3~=(1.0,-0.1,···,0.5), and x4~=(0.9,0.3,...,1.1) is.

[0159] In the initial generation process described in Example 11, the random variable X~ is created solely from essentially random potential representations and the desired label y~. Therefore, compounds identified by the generated compound representation x~ are only expected to have activities and properties that fit the requirements of the desired label y~. However, in this Example 15, the random variable X~, and therefore the compound representation x~, is generated solely from the potential representations of the designated seed compound x~. D and its associated label y D Therefore, in the comparative generation procedure of this embodiment, the generated compound representation x~ is generated from the seed compound x D It can be expected to both retain some salient aspects of and have activities and properties that match the requirements of the desired label y. [Example]

[0160] <Evaluation of the predicted results of the generated compounds and subsequent ranking> In this example, the predicted assay results of the generated fingerprints are compared to the desired assay results, and fingerprints with predicted results that match the desired assay results are then ranked by drug-likeness score.

[0161] After generation of a fingerprint x~, e.g., via initial generation or comparative generation, x~ is input to a trained predictor module. (The predictor module may have been trained, e.g., during a semi-supervised learning process for unlabeled data.) The predictor module outputs a predicted set y^ of assay results for the generated fingerprint x~.

[0162] The predicted assay result ŷ and the desired assay result ŷ are input into a comparison module (FIG. 7). If the predicted result is the same as the desired result, x̂ is added to the set of unranked candidates U; otherwise, x̂ is rejected. The unranked set is then ranked by a ranking module, for example, as described in Example 18. [Example]

[0163] <Evaluation of fingerprints generated through comparison generation> In this example, fingerprints generated using the comparative generation process are evaluated for similarity to a seed compound and for having a label similar to a desired label. In the comparative generation procedure exemplified above, a seed compound is used to generate a new fingerprint that resembles the seed. Once the fingerprint is generated, a further evaluation step is used to determine whether the generated fingerprint is sufficiently similar to the seed. A comparison module is used to compare corresponding parameters of the two fingerprints. If a threshold or threshold similarity of the identity parameters is achieved, the two fingerprints are marked as sufficiently similar.

[0164] After generating the fingerprint x~, the seed compounds x~ and x D Both are input to the comparison module. D If x~ is sufficiently similar to , then x~ is retained; otherwise, x~ is rejected. If retained, x~ is input to the predictor module, and a predicted label y^ is provided by the prediction module. A comparison module is used to compare the predicted label y^ with the desired label y~. If the predicted label y^ is sufficiently similar or the same as the desired label y~, then x~ is added to an unranked candidate set U. The unranked set of fingerprints is then ranked by a ranking module to output a ranked set R. [Example]

[0165] <Training the Ranking Module and Applying the Ranking Module> In this example, the ranking module is trained to rank the generated representations x. The generated representations may have been filtered by other modules, such as a comparison module, before entering the ranking module. In this example, the ranking module has two functions: (1) assigning a drug-likeness score to each fingerprint, and (2) ranking the set of fingerprints according to their drug-likeness scores.

[0166] The ranking module is configured to evaluate the fingerprints based on their latent representations.

[0167] First, an autoencoder is trained on a large set of compound fingerprints. After training, the first half of the autoencoder, LRG, is used to generate latent representations of compounds (Figure 9). The latent representations are input to a classifier, which is trained using supervised learning. The training dataset includes approximately 2,500 FDA-approved drugs, all with the class label Drug, and a large set of other non-drug compounds, all with the label Not Drug. The classifier outputs a continuous score representing the drug-likeness of the compound. To apply the ranking module, members of the unranked set of generated compound fingerprints are input to a latent representation generator (LRG), and then the generated latent representations are input to the classifier. Each compound receives a drug-likeness score from the classifier. The compounds are then ordered from highest to lowest score. The final output is a ranked set of candidate compound fingerprints. It is a combination. [Example]

[0168] Sequential application of initial generation and comparative generation to explore novel compound space For a particular set of assay results, it may be desirable to generate a new compound that meets those results and then search for similar compounds in the space around the original compound. For this application, initial generation and comparative generation may be used in sequence.

[0169] Based on the desired assay result \(y^*\), a fingerprint \(x^*\) is generated using an initial generation (Figure 11). To identify compounds not previously known, the comparison module compares \(x^*\) with a database of known compounds. If \(x^*\) already exists within the database, \(x^*\) is rejected. If \(x^*\) is a previously unknown compound, \(x^*\) is input into a predictor to generate a predicted assay result \(y^\wedge\).

[0170] Next, the fingerprint \(x^*\) and its predicted assay result \(y^\wedge\) are used as a seed for comparative generation. A new fingerprint \(x^+\) is generated, along with its predicted assay result \(y^+\). The comparison module then determines whether \(y^+\) is the same as the desired assay result \(y^*\). If so, \(x^+\) is retained and added to a set of unranked candidates. Any desired number of fingerprints \(x^+\) may be generated from the initial seeds of \(x^*\) and \(y^\wedge\) by repeated application of comparative generation.

[0171] After a desired number of candidates are generated and collected as a set \(U\) of unranked candidate fingerprints, the unranked set is input into a ranking module, which outputs a ranked set \(R\).

Example

[0172] <QSAR Analysis - Part I: Identification of Compound Characteristics Likely to Affect the Results of a Specific Assay> This method is used to identify compound characteristics that may cause a particular assay result. This method provides a way to identify candidate mutations, i.e., specific structural characteristics that change the performance of a compound in a particular assay. These may then be used as a starting point for matched molecular pair analysis (MMPA).

[0173] In this example, two initial generation processes are executed in parallel. On the one hand, the desired assay result y~ is used as the positive seed. On the other hand, the opposite assay result y* is used as the negative seed. If y~ is a single binary assay result, the negative seed y* is the opposite result of that assay. To reduce the variation in the resulting generated fingerprints, a vector of assay results may be used as the positive seed y~. In this case, for only one result of the target assay, y* is different from y~.

[0174] Two sets A and B of compound fingerprints are generated. A contains compounds generated from the positive seed y~, and B contains compounds generated from the negative seed y*. After generating the desired number of members for each set, the two sets are input into a comparison module. The comparison module identifies the fingerprint parameters that are most likely to cause differences in the target assay results. Exemplary comparison modules are described in more detail in later examples and elsewhere in this specification.

Example

[0175] <QSAR Analysis - Part II: Exploration of Variations Regarding Desired Results for Specific Compounds> In this example, a method for exploring variations in specific compounds that can cause specific assay results is described. In this method, two comparison generation processes are repeatedly executed in parallel to generate two sets of fingerprints (Figure 13). These processes use the same seed compound but each uses a different set of target assay results, for example, positive target y~ and negative target y*, where y~ and y* differ only in a single assay result. A comparison module is used to identify specific structural differences between the fingerprints generated with the positive target and those generated with the negative target.

[0176] The generated fingerprints are first evaluated by their similarity to the seed compound. If the comparison module finds them sufficiently similar to the seed compound, a predictor is used to provide a predicted assay result for each generated fingerprint. The predicted assay results are then checked for similarity or identity with the corresponding target assay results y~ and y*, respectively.

[0177] The comparison-generation process is performed as many times as necessary to generate two sets of candidate fingerprints, A and B, with the desired cardinality, where A contains generated fingerprints created with the positive target y~ and B contains generated fingerprints created with the negative target y*. Members of A are compared to members of B using a comparison module. The comparison module is configured to identify homogeneous and heterogeneous structural variants within the two sets. These structural variants can then be used as starting points for further analysis via MMPA. [Example]

[0178] <Comparison module> This example describes a comparison module with two functions: (1) determining whether two objects, e.g., two vectors of assay results or two fingerprints, are similar or identical, and (2) identifying fingerprint parameters that are most likely to cause variation in a particular assay result by comparing two sets of fingerprints.

[0179] A. Comparing two objects for similarity A simple pairwise comparison for similarity compares corresponding elements of two objects, e.g., two vectors of assay results or two fingerprints, and a user-specified threshold is set to determine whether the two objects pass or fail the comparison.

[0180] A second method for comparing two fingerprints is to encode the fingerprints as latent representations using a Latent Representation Generator (LRG). Then, the corresponding distributions of the latent representations are compared and a determination of similarity is made.

[0181] B. Comparison of a Set of Objects for Identifying Important Compound Transformations When comparing two sets of fingerprints, several methods are used to identify important transformations of compounds. One simple method is to use a linear model to identify important parameters. For example, interaction terms can be added to the model to address the possibility that interactions between parameters caused changes in assay results.

[0182] A second method involves the use of the Gini coefficient. The Gini coefficient is calculated for each parameter by dividing the average of the differences between all possible pairs of fingerprints by the average size. The parameter with the largest Gini coefficient is selected as the parameter most likely to be related to changes in assay results.

[0183] In an extension of this method, a classification tree is used. The parameter with the largest Gini coefficient is selected to be the root of the classification tree. The rest of the classification tree is learned by top-down induction. Then, by observing the behavior of the tree at appropriate levels, the desired number of important parameters is identified.

[0184] If the cardinality of the two sets of fingerprints is low, the Gini coefficient may be calculated directly. In some cases, a clustering method is applied to reduce the number of pairwise comparisons required between A and B. Then, the Gini coefficient of the parameters is calculated by pairwise comparison between the centroids of A and B.

Examples

[0185] <Calculation of the Gini Coefficient Using k-Medoid Clustering> In this example, the comparison module is configured to utilize the clusters of the potential representations of sets A and B. First, a Latent Representation Generator (LRG) is used (FIG. 14) to encode the members of sets A and B as potential representations to form sets AL and BL, respectively. Next, K-medoid clustering is applied to the members of sets A L and B L . Following the clustering, the centroids of the potential representations of sets A C and B C are extracted to form the centroid sets of the potential representations of sets A F and B F . Fingerprints corresponding to the members of A C and B C are retrieved to form two sets of fingerprints of A F and B F . Next, the members of A F and B F are used to identify compound modifications that may cause changes in assay results or other label element values.

Example

[0186] <Calculation of Dini coefficient using k - means clustering> In this example, k - means clustering is used instead of K - medoid clustering in the method described in Example 23. As in the K - medoid method, the members of sets A and B are encoded as potential representations. K - means clustering is applied to the set of potential representations. The centroids resulting from the k - means clustering are decoded as fingerprints using a Latent Representation Decoder Module (LRD) and stored in respective sets A F and B F . Sets A F and B F are used to identify important compound modifications associated with changes in labels or label element values.

[0187] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the invention. It is understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. The following claims define the scope of the invention, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby.

Claims

1. A method executed by at least one processor, comprising: the at least one processor: generating latent representations by inputting information about the compound into a first machine learning model; predicting a property of the compound by inputting the latent representation of the compound into a classifier; training the classifier based on the results of the prediction; method.

2. the at least one processor: generating latent representations of other compounds by inputting information of the other compounds into the first machine learning model; generating reconstruction information of information of the other compound by inputting the latent representation of the other compound into a second machine learning model; training the first machine learning model and the second machine learning model so that an error between the information on the other compound and the reconstruction information on the other compound is small; The method of claim 1.

3. the first machine learning model is an encoder; the second machine learning model is a decoder. The method of claim 2.

4. The dataset used to train the classifier includes information on drugs and non-drugs.

4. The method according to any one of claims 1 to 3.

5. generating the latent representation by inputting information about the compound into the first machine learning model, the at least one processor: acquiring latent variables by inputting information about the compound into the first machine learning model; generating the latent representation by sampling using the latent variables; Including, 5. The method according to any one of claims 1 to 4.

6. The latent variable is expressed using any one of a normal distribution, a Laplace distribution, an elliptical distribution, a Student's t distribution, a logistic distribution, a uniform distribution, a triangular distribution, an exponential distribution, a reversible cumulative distribution, a Cauchy distribution, a Rayleigh distribution, a Pareto distribution, a Weibull distribution, a reciprocal distribution, a Gompertz distribution, a Gumbel distribution, an Erlang distribution, a log-normal distribution, a gamma distribution, a Dirichlet distribution, a beta distribution, a chi-squared distribution, or an F-distribution. The method of claim 5.

7. A method executed by at least one processor, comprising: the at least one processor: generating latent representations by inputting compound information into a trained first machine learning model; predicting properties of the compound by inputting the latent representations into a trained classifier; method.

8. the first machine learning model is an encoder; The method of claim 7.

9. generating the latent representation by inputting information about the compound into the trained first machine learning model, the at least one processor: obtaining latent variables by inputting information about the compound into the first machine learning model that has already been trained; generating the latent representation by sampling using the latent variables; Including, The method according to claim 7 or claim 8.

10. The latent variable is expressed using any one of a normal distribution, a Laplace distribution, an elliptical distribution, a Student's t distribution, a logistic distribution, a uniform distribution, a triangular distribution, an exponential distribution, a reversible cumulative distribution, a Cauchy distribution, a Rayleigh distribution, a Pareto distribution, a Weibull distribution, a reciprocal distribution, a Gompertz distribution, a Gumbel distribution, an Erlang distribution, a log-normal distribution, a gamma distribution, a Dirichlet distribution, a beta distribution, a chi-squared distribution, or an F-distribution.

10. The method of claim 9.

11. the at least one processor: The latent representation and the label information of the compound are input to the classifier to predict the properties of the compound.

11. The method according to any one of claims 7 to 10.

12. The label information includes drug or non-drug information. The method of claim 11.

13. the properties include the drug-likeness of the compound; 13. The method of any one of claims 1 to 12.

14. The classifier outputs a continuous score representing the drug-likeness of the compound. The method of claim 13.

15. the compound information includes at least one of a molecular descriptor or a fingerprint representation of the compound; 15. The method of any one of claims 1 to 14.

16. The information on the compound includes information on the chemical structure of the compound.

15. The method of any one of claims 1 to 14.

17. the at least one processor: ranking the compounds based on the properties; 17. The method of any one of claims 1 to 16.

18. the at least one processor: clustering the compounds based on the latent representations; 18. The method of any one of claims 1 to 17.

19. at least one processor; The at least one processor executes the method of any one of claims 1 to 18. Computer system.

20. causing at least one processor to perform the method of any one of claims 1 to 18; program.