System and method for query-based random access to a virtual chemical combinatorial synthesis library

A graph-based generative model efficiently navigates and queries combinatorial synthesis libraries, addressing scalability and computational challenges, ensuring chemical validity and accessibility, thereby facilitating effective access to large non-enumerated chemical spaces.

JP2025520025APending Publication Date: 2025-07-01ATOMWISE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024567507
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-16
Filing Date
2023-05-16
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing methods for navigating and querying combinatorial synthesis libraries face challenges with scalability, chemical validity, and synthetic accessibility, particularly in non-enumerable chemical spaces, as they rely on exhaustive enumeration and struggle with long autoregressive chains and computational inefficiencies.

Method used

A system and method using a graph-based generative model that encodes molecular queries, allowing for query-based random access to large combinatorial synthesis libraries by learning a hierarchy of keys across the library components, minimizing autoregression, and enabling efficient parallelization, thus overcoming scalability and computational complexity issues.

Benefits of technology

The system provides efficient navigation and random access to vast non-enumerated compound libraries, ensuring chemical validity and synthetic accessibility with reduced computational complexity and parameter count, making it suitable for large molecular graphs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025520025000001_ABST
    Figure 2025520025000001_ABST
Patent Text Reader

Abstract

A system and method for querying a combinatorial synthesis library that includes a plurality of compounds, represents a plurality of reaction types, where each reaction type is mapped to a plurality of reactants, and each reactant is mapped to a plurality of synthons accepts a query in the form of a single graph into a molecular encoder model, thereby obtaining a query vector. The query vector is input into a reaction query generator model, thereby obtaining a first reaction type and a first plurality of reactants. A synthon is determined for each reactant by inputting the reactant into a synthon query generator model. Thus, a set of synthons is determined, each corresponding to a reactant among the first plurality of reactants. A molecular structure in the combinatorial synthesis library is identified that includes the set of synthons arranged according to synthesis rules associated with the first reaction type.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 342,574, filed May 16, 2022, entitled "Graph Generative Model for Combinatorial Synthesis Libraries", which is incorporated herein by reference.

[0002] The present disclosure generally relates to screening combinatorial synthesis libraries for target analogs.

Background Art

[0003] Virtual high - throughput screening (vHTS)

[47] has received significant support in early - stage drug discovery, at least in part due to make - on - demand chemical libraries that utilize combinatorial synthesis constructs. These combinatorial synthesis libraries (CSLs) enable access to a very broad chemical space from a set of very small chemically - accessible building blocks that can be combined according to known synthetic routines. In recent years, these libraries have grown from millions to billions and now to trillions of compounds [22, 33, 44, 58]. As a result, virtual chemical libraries are rapidly approaching sizes that exceed what can be explicitly enumerated, presenting new challenges for virtual screening. For example, the Enamine Readily Accessible (REAL) library

[20] utilizes off - the - shelf molecular building blocks and parallel synthesis, enabling lead times of about a few weeks, heralding an era in which the waiting time between in - silico and in - vitro high - throughput screening is increasingly shortened.

[0004] As a result of the combinatorial explosion enabled by such structures, early-stage drug discovery has now "crossed the Rubicon" into a non-enumerable realm. This poses new challenges for in-silico hit finding and optimization, which rely on screening explicitly enumerated compounds. These methods are not suitable for non-enumerable settings as they increase proportionally with the number of compounds.

[0005] Virtual high-throughput screening and enumeration. Often, the first step in a vHTS campaign is to prepare a compound library for later use [1,17]. Compound sampling and scoring techniques have been developed

[18] , but these approaches nevertheless rely on comprehensively accessible libraries. An exception is the virtual synthon hierarchical enumeration screening (V-SYNTHES) approach

[44] , which exploits the modularity of parallel synthesis libraries. However, by design, V-SYNTHES does not permit query-based random access. On the other hand, SpaceMACS

[45] and SpaceLight [4] can provide query-based access to modular libraries by decomposing queries into fragments and matching them to synthons in the library by similarity search. In parallel with these efforts, machine learning has received significant attention in vHTS for predicting activity scores based on docked conformations [15,43,57], for predicting activity scores by giving ligands and proteins separately (undocked) [38,56], and for improving or completely replacing classical molecular docking with machine learning approaches [39,51,52].

[0006] Deep learning approaches molecular generation. De novo drug design is playing an increasingly important role in the identification of novel chemical entities in drug discovery campaigns [10,36,50,54]. Two dominant neural network-based paradigms for molecular generation are text-based and graph-based generative models. Initial research in text-based generative models (also referred to as chemical language models) applied recurrent neural networks to SMILES strings [16,46]. These methods were very promising and stimulated interest in molecular generation within the ML community, but they do not guarantee the generation of valid SMILES strings. To improve validity, approaches leveraging the syntactic constraints of the SMILES notation have been proposed [12,29], and separately, the recently proposed SELFIES notation [28,37] guarantees validity and is seeing increasing adoption as such. However, in both cases, modeling using such text-based representations of chemical substances has known drawbacks (e.g., subjectivity, perhaps similar molecular structures with large edit distances).

[0007] For some applications, it is interesting to utilize generative models that can be adapted to molecular databases and enable navigation of these databases via the adapted models. The ability of language models to adapt to molecular databases has been investigated in prior work [2] applying deep language models to GDB-13 [7], which is a database of 975,000,000 compounds formed by completely enumerating molecules up to 13 atoms of the element types C, N, O, S, and Cl according to simple chemical stability and synthetic feasibility rules. The authors found that training on 0.1% of the entire library allowed the model to cover approximately 70% of the compounds in the GDB-13 library. Furthermore, the language model the authors trained was generating compounds that did not satisfy the GDB-13 constitution in approximately 15% of the cases.

[0008] Graph generation models have recently received significant attention as an alternative to their text-based counterparts. The earliest of these models focused on generating graphs of a fixed size in a single shot

[48] , or autoregressively generating graphs of arbitrary size one atom or bond at a time [32, 42, 60, 35]. These approaches also struggle to reliably generate chemically valid molecules and encounter difficulties with large molecular graphs.

[0009] To address both points, fragment-based graph generation models have been proposed and are gaining popularity [23, 24, 25, 27]. These models have the advantage of decomposing molecules into valid subcomponents and ensuring chemical validity by explicitly prohibiting actions that would result in invalid combinations of fragments. Such explicit validity checks can be performed for all actions at the expense of additional computation. Other text-based and graph-based generation models tend to struggle with large molecular graphs due to the long autoregressive chains required to generate them, but fragment-based graph generation models require an autoregressive length on the order of the number of fragments containing the molecule. This can be significant when the fragments themselves contain many atoms.

[0010] However, due to the general difficulty of autoregressive graph generation, several problems remain. Unlike text-based models where the autoregressive order is less ambiguous (e.g., tokens are typically decoded in a left-to-right order), graphs do not have such a canonical node order, which presents challenges for graph-based autoencoders [32, 59]. Furthermore, while they require shorter autoregressive chains than their counterparts, existing fragment-based graph generation models still require an increasing autoregressive length in terms of the overall size of the molecule because they cannot effectively parallelize autoregressive decoding.

[0011] Fragment-based graph generation models and SELFIES-based language models each address the issue of chemical validity, but there is a separate issue of synthetic accessibility. Prior research has cast doubt on the synthetic feasibility of compounds proposed by many existing generation models

[13] , which, if not properly addressed, can limit the practicality of these models in drug discovery applications. Subsequent research has attempted to ameliorate these drawbacks by (i) including explicit penalties for synthetic inaccessibility via a scoring function

[19] , (ii) restricting the model to fragments from known compounds [34, 40, 53], or (iii) inducing a bias towards simple and known synthetic routes [8, 9, 21].

[0012] In view of the above background, what is needed in the art is a scalable approach for navigating CSL using query-based random access. SUMMARY OF THE INVENTION

[0013] The present disclosure addresses the above need in the art. A system and method for querying a combinatorial synthesis library that includes a plurality of compounds and represents a plurality of reaction types, where each reaction type is mapped to a plurality of reactants and each reactant is mapped to a plurality of synthons. The system receives a query in the form of a single graph into a molecular encoder model, thereby obtaining a query vector. The query vector is input into a reaction query generator model, thereby obtaining a first reaction type and a first plurality of reactants. The synthons are determined for each reactant by inputting the reactant into a synthon query generator model. Thus, a set of synthons is determined, each corresponding to a reactant among the first plurality of reactants. The molecular structure in the combinatorial synthesis library is identified that includes the set of synthons arranged according to the synthesis rules associated with the first reaction type.

[0014] In one embodiment, the molecular encoder model is a graph-based generative model that utilizes the structure of CSL to provide efficient navigation of the associated chemical space. The model learns a hierarchy of keys across the components of the library and uses these keys to process queries for retrieval. The encoder processes the molecular graph and returns a query vector as output, which the decoder uses to retrieve molecules from the CSL through an efficient sequence of query-key comparisons that utilize the hierarchical structure of the CSL, requiring minimal autoregression and enabling efficient parallelization. The graph-based generative model in such an embodiment functions as a "neural database" and provides random access to an extremely large non-enumerated compound library. Thus, this model provides molecules that are effectively and cost-effectively available. Further, this model overcomes the challenges posed by long autoregressive chains in compound generation and improves scalability to large molecular graphs. Further, this model reduces the number of parameters by a factor of 10 compared to equivalent methods and greatly improves the computational complexity for searching the CSL.

[0015] A system and method for querying a combinatorial synthesis library are provided. The combinatorial synthesis library includes a plurality of compounds and represents a plurality of reaction types. Each respective reaction type of the plurality of reaction types has a corresponding mapping to a corresponding plurality of reactants. Each respective reactant of each corresponding plurality of reactants has a corresponding mapping to a corresponding plurality of synthons.

[0016] In some such embodiments, the plurality of reaction types includes one or more reaction types, and the combinatorial synthesis library includes one or more compounds for each reaction type of the plurality of reaction types.

[0017] The molecular query is input into a molecular encoder model. The query is graph-based. The molecular encoder model includes a message-passing neural network including a plurality of message-passing layers collectively including a first plurality of parameters. Inputting the graph-based query into the molecular encoder model results in a query vector as an output of the molecular encoder model by applying the first plurality of parameters to the graph-based query.

[0018] In some such embodiments, the graph-based query is a single graph including a plurality of nodes and a plurality of edges, and each node among the plurality of nodes is connected by at least one edge among the plurality of edges to another type of node among the plurality of nodes.

[0019] In some such embodiments, each respective node among the plurality of nodes is associated with (i) a corresponding element type among a plurality of element types, (ii) a node degree among a plurality of node degrees, (iii) a hybridization among a plurality of hybridizations, (iv) a number of bound hydrogens, (v) a formal charge from a set of formal charges, and (vi) a binary representation of aromaticity.

[0020] In some such embodiments, each respective bond among the plurality of bonds is associated with (i) a bond type, (ii) a binary representation of conjugation, (iii) a binary of a representation of whether each respective bond is within a ring, and (iv) a representation of stereochemistry.

[0021] In some such embodiments, the graph, which can be any graph, represents a single-molecule compound, and such single-molecule compound is present in a combinatorial synthesis library.

[0022] In some such embodiments, the graph represents a weighted complex of a first graph of a first molecular compound and a second graph of a second molecular compound.

[0023] In some such embodiments, the graph represents a weighted complex of a plurality of graphs of a second plurality of compounds, and the second plurality of compounds have a common property.

[0024] In some such embodiments, the common property is a Tanimoto distance less than a threshold with respect to each other compound of the second plurality of compounds.

[0025] In some such embodiments, the common property is a binding coefficient with respect to a macromolecular target that is less than a threshold.

[0026] The query vector is input into a reaction query generator model that includes a second plurality of parameters. This results in an output from the reaction query generator model 116 of a first reaction type among a plurality of reaction types by applying the second plurality of parameters to the query vector.

[0027] In some such embodiments, the reaction query generator model is a two-layer perceptron with intermediate ReLU activation.

[0028] In some such embodiments, the output from the reaction query generator model is used to identify a first reaction key type among a plurality of reaction key types through a first query key lookup, and each reaction key type among the plurality of reaction key types represents a synthetic reaction that can be used to synthesize one or more compounds within a combinatorial synthesis library.

[0029] The corresponding synthons are determined for each respective reactant among the first plurality of reactants corresponding to the first reaction type by inputting each respective reactant 86 into a synthon query generator model that includes a third plurality of parameters, thereby applying the third plurality of parameters to each respective reactant and obtaining the corresponding synthon as an output from the synthon query generator model, thereby obtaining a set of synthons, each synthon in the set of synthons corresponding to a reactant among the first plurality of reactants.

[0030] In some such embodiments, the synthon query generator model is a two-layer perceptron with intermediate ReLU activation.

[0031] In some such embodiments, the output from the synthon query generator model is used to identify a synthon key for the corresponding synthon through a second query key lookup.

[0032] In some such embodiments, the first plurality of parameters includes 100,000 parameters, the second plurality of parameters includes 5,000 parameters, and the third plurality of parameters includes 5,000 parameters.

[0033] In some such embodiments, the first plurality of reactants includes three or more reactants, and the corresponding mapping for the corresponding plurality of synthons for the reactants among the three or more reactants includes ten or more synthons.

[0034] The molecular structure (compound) is identified in a combinatorial synthesis library that includes a set of synthons arranged according to, for example, synthesis rules related to the first reaction type.

[0035] In some such embodiments, the plurality of compounds in the combinatorial synthesis library includes one billion or more compounds, and the identified molecular structure is any one of the one billion or more compounds that satisfy the query.

[0036] In some such embodiments, the plurality of compounds in the combinatorial synthesis library includes one trillion compounds or more, and the identified molecular structure is any one of the one trillion or more compounds that satisfy the query.

[0037] In the drawings, embodiments of the systems and methods of the present disclosure are shown by way of example. It should be clearly understood that the description and drawings are for purposes of illustration only and are not intended as a definition of the limitations of the systems and methods of the present disclosure.

Brief Description of the Drawings

[0038]

Figure 1A

Figure 1B

Figure 2A

Figure 2B

Figure 2C

Figure 2D

Figure 3A

Figure 3B

Figure 3C

Figure 3D

Figure 3E

Figure 3F

Figure 3G

Figure 3H

Figure 4A

Figure 4B

Figure 5A

Figure 5B

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

[0039] Like reference numerals refer to corresponding parts throughout the drawings.

DETAILED DESCRIPTION OF THE INVENTION

[0040] Reference is now made in detail to embodiments illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. It will be apparent to one of ordinary skill in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. All patents and publications mentioned herein are incorporated by reference in their entirety.

[0042] Definitions. As used herein, the terms "administer", "administration", or "administering" refer to (1) providing, giving, dosing, and / or prescribing, either by or under the direction of a healthcare provider or a duly authorized agent thereof, in accordance with the present disclosure, and / or (2) placing, causing to be ingested, or causing to be consumed in a mammal in accordance with the present disclosure.

[0043] As used herein, the terms "co-administer", "co-administering", "administered in combination with", "administering in combination with", "simultaneously", and "concurrently" include administering two or more active pharmaceutical ingredients to a subject such that both the active pharmaceutical ingredients and / or their metabolites are present in the subject at the same time. Co-administration includes simultaneous administration in separate compositions, administration at different times in separate compositions, or administration in a composition in which two or more active pharmaceutical ingredients are present. Simultaneous administration in separate compositions and administration in a composition in which both agents are present are preferred.

[0044] The terms "active pharmaceutical ingredient" and "drug" include the compounds described herein, and any pharmaceutically acceptable analogs, derivatives, salts, solvates, hydrates, co-crystals, or prodrugs thereof. The terms "active pharmaceutical ingredient" and "drug" may also include the compounds described herein, and any pharmaceutically acceptable analogs, derivatives, salts, solvates, hydrates, co-crystals, or prodrugs thereof that bind to a target molecule.

[0045] The term "in vivo" refers to events occurring within the body of a subject.

[0046] The term "in vitro" refers to events occurring outside the body of a subject. In vitro assays can include cell-based assays in which live or dead cells are employed, and may also include cell-free assays in which intact cells are not employed.

[0047] As used herein, the term "if" can be interpreted to mean "when" or "upon" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "when determined" or "when [described condition or event] is detected" can be interpreted to mean "when determined" or "in response to determination" or "when [described condition or event] is detected", or "when [described condition or event] is detected", depending on the context.

[0048] The terms "effective amount" or "therapeutically effective amount" refer to an amount of a compound or combination of compounds described herein sufficient to achieve the intended application, including but not limited to the treatment of a disease. The therapeutically effective amount can vary depending on the intended application (in vitro or in vivo), or the subject and disease state being treated (e.g., the subject's weight, age and gender), the severity of the disease state, the mode of administration, etc., which can be readily determined by one of ordinary skill in the art. This term also applies to the dose that induces a specific response in a target cell. The specific dose varies depending on the specific compound selected, the dosing regimen to be followed, whether the compound is administered in combination with other compounds, the timing of administration, the tissue to which it is administered, and the physical delivery system by which the compound is carried.

[0049] "Therapeutic effect", as the term is used herein, encompasses therapeutic and / or prophylactic benefits. Prophylactic effects include delaying or eliminating the onset of a disease or condition, delaying or eliminating the development of symptoms of a disease or condition, slowing, stopping, or reversing the progression of a disease or condition, or any combination thereof.

[0050] The term "pharmaceutically acceptable salt" refers to salts derived from various organic and inorganic counterions known in the art. Pharmaceutically acceptable acid addition salts can be formed with inorganic acids and organic acids. Preferred inorganic acids from which the salts can be derived include, for example, hydrochloric acid, hydrobromic acid, sulfuric acid, nitric acid, and phosphoric acid. Preferred organic acids from which the salts can be derived include, for example, acetic acid, propionic acid, glycolic acid, pyruvic acid, oxalic acid, maleic acid, malonic acid, succinic acid, fumaric acid, tartaric acid, citric acid, benzoic acid, cinnamic acid, mandelic acid, methanesulfonic acid, ethanesulfonic acid, p-toluenesulfonic acid, and salicylic acid. Pharmaceutically acceptable base addition salts can be formed with inorganic bases and organic bases. Inorganic bases from which the salts can be derived include, for example, sodium, potassium, lithium, ammonium, calcium, magnesium, iron, zinc, copper, manganese, and aluminum. Organic bases from which the salts can be derived include, for example, primary, secondary, and tertiary amines, substituted amines including naturally occurring substituted amines, cyclic amines, and basic ion exchange resins. Specifically, isopropylamine, trimethylamine, diethylamine, triethylamine, tripropylamine, ethanolamine, etc. are included. In some embodiments, the pharmaceutically acceptable base addition salts are selected from ammonium salts, potassium salts, sodium salts, calcium salts, and magnesium salts. The term "cocrystal" refers to a molecular complex derived from some cocrystal formers known in the art. Unlike salts, cocrystals typically do not involve hydrogen transfer between the cocrystal and the drug, but instead involve intermolecular interactions such as hydrogen bonding, aromatic ring stacking, or dispersion forces between the cocrystal former and the drug in the crystal structure.

[0051] "Pharmaceutically acceptable carrier" or "pharmaceutically acceptable excipient" is intended to include any and all solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic and absorption delaying agents, and inert ingredients. The use of such pharmaceutically acceptable carriers or pharmaceutically acceptable excipients for active pharmaceutical ingredients is well known in the art. Its use in the therapeutic compositions of the present disclosure is contemplated, except when any conventional pharmaceutically acceptable carrier or pharmaceutically acceptable excipient is incompatible with the active pharmaceutical ingredient. Additional active pharmaceutical ingredients, such as other drugs disclosed herein, can also be incorporated into the described compositions and methods.

[0052] For example, when ranges are used herein to describe physical or chemical characteristics such as molecular weight or chemical formula, all combinations and sub-combinations of the range and specific embodiments therein are intended to be included. The use of the term "about" when referring to a number or numerical range means that the recited number or numerical range is an approximation within the range of experimental variability (or within the range of statistical experimental error), and thus the number or numerical range can vary. The variation is typically 0% to 15%, or 0% to 10%, or 0% to 5% of the recited number or numerical range. The term "comprising" (and related terms such as "comprise", "comprises", "having", or "including") includes embodiments such as an embodiment of any composition, method, or process "consisting of" or "consisting essentially of" the recited features.

[0053] As used interchangeably herein, the terms "classifier" or "model" refer to a machine learning model or algorithm.

[0054] In some embodiments, the "model" is supervised machine learning. Non-limiting examples of supervised learning algorithms include, but are not limited to, logistic regression, neural networks, support vector machines, naive Bayes algorithms, nearest neighbor algorithms, random forest algorithms, decision tree algorithms, boosted tree algorithms, polynomial logistic regression algorithms, linear models, linear regression, gradient boosting, hybrid models, hidden Markov models, Gaussian NB algorithms, linear discriminant analysis, or any combination thereof. In some embodiments, the classifier is a multi-class classifier algorithm. In some embodiments, the model is a two-stage Stochastic Gradient Descent (SGD) model. In some embodiments, the model is a deep neural network (e.g., a deep and wide sample-level classifier).

[0055] Neural network. In some embodiments, the model is a neural network (e.g., a convolutional neural network and / or a residual neural network). Neural network algorithms, also known as artificial neural networks (ANNs), include convolutional and / or residual neural network algorithms (deep learning algorithms). A neural network can be a machine learning algorithm that can be trained to map an input data set to an output data set, and a neural network comprises an interconnected group of nodes organized into multiple layers of nodes. For example, a neural network architecture can include at least an input layer, one or more hidden layers, and an output layer. A neural network can include any total number of layers and any number of hidden layers, and the hidden layers function as trainable feature extractors that enable mapping a set of input data to an output value or a set of output values. As used herein, a deep learning algorithm (DNN) can be a neural network that includes multiple hidden layers, e.g., two or more hidden layers. Each layer of a neural network can include a number of nodes (or "neurons"). A node can receive an input coming directly from either input data or the output of nodes in a previous layer and can perform a particular operation, e.g., a summation operation. In some embodiments, the connections from the input to the nodes are associated with parameters (e.g., weights and / or weight coefficients). In some embodiments, a node is an input, x iThe sum of the products of all pairs of the neural network's weighting coefficients, bias values, and thresholds, or other computational parameters may be calculated. In some embodiments, the weighted sum is offset by a bias b. In some embodiments, the output of a node or neuron may be gated using a threshold function or activation function f, which may be a linear or non-linear function. The activation function can be, for example, a rectified linear unit (ReLU) activation function, a leaky ReLU activation function, or a saturated hyperbolic tangent, identity, binary step, logistic, arcTan, soft sign, parametric rectified linear unit, exponential linear unit, softPlus, bent identity, softExponential, sinusoid, sine, Gaussian, or sigmoid function, or any combination thereof.

[0056] The weighting coefficients, bias values, and thresholds, or other computational parameters of the neural network can be "taught" or "learned" during a training phase using one or more sets of training data. For example, the parameters can be trained using input data from the training data set and gradient descent or backpropagation such that the output value(s) calculated by the ANN match the examples included in the training data set. The parameters can be obtained from a backpropagation neural network training process.

[0057] Any of a variety of neural networks may be suitable for use in accordance with the present disclosure. Examples can include, but are not limited to, feedforward neural networks, radial basis function networks, recurrent neural networks, residual neural networks, convolutional neural networks, residual convolutional neural networks, or any combination thereof. In some embodiments, machine learning utilizes pre-trained and / or transfer-learned ANNs or deep learning architectures. Convolutional and / or residual neural networks can be used in accordance with the present disclosure.

[0058] For example, a deep neural network classifier comprises an input layer, a plurality of individually parameterized (e.g., weighted) convolutional layers, and an output scorer. The parameters (e.g., weights) of each of the convolutional layers as well as the input layer contribute to a plurality of parameters (e.g., weights) associated with the deep neural network classifier. In some embodiments, at least 100 parameters, at least 1000 parameters, at least 2000 parameters, or at least 5000 parameters are associated with the deep neural network classifier. Therefore, since the deep neural network classifier cannot be solved mentally, it is necessary to use a computer. In other words, in such embodiments, when an input to the classifier is provided, the classifier output needs to be determined using a computer rather than mentally. For example, reference is made to Krizhevsky et al., 2012, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems 2, Pereira, Burges, Bottou, Weinberger, eds., pp. 1097-1105, Curran Associates, Inc., Zeiler, 2012 “ADADELTA: an adaptive learning rate method,” / 1212.5701, and Rumelhart et al., 1988, “Neurocomputing: Foundations of research,” ch. Learning Representations by Back-propagating Errors, pp. 696-699, Cambridge, MA, USA: MIT Press, each of which is hereby incorporated by reference.

[0059] Neural network algorithms including convolutional neural network algorithms suitable for use as classifiers are disclosed, for example, in Vincent et al., 2010, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion,” J Mach Learn Res 11, pp. 3371-3408, Larochelle et al., 2009, “Exploring strategies for training deep neural networks,” J Mach Learn Res 10, pp. 1-40, and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of which is incorporated herein by reference. Further exemplary neural networks suitable for use as classifiers are disclosed in Duda et al., 2001, Pattern Classification, Second Edition, John Wiley & Sons, Inc., New York, and Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, each of which is incorporated herein by reference in its entirety. Further exemplary neural networks suitable for use as classifiers are also described in Draghici, 2003, Data Analysis Tools for DNA Microarrays, Chapman & Hall / CRC, and Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, each of which is incorporated herein by reference in its entirety.

[0060] As used herein, the term "parameter" refers to any coefficient in an algorithm, model, regressor, and / or classifier, or similarly, any value of an internal or external element (e.g., weight and / or hyperparameter) that can affect (e.g., modify, adjust, and / or tune) one or more inputs, outputs, and / or functions in the algorithm, model, regressor, and / or classifier. For example, in some embodiments, a parameter can refer to any coefficient, weight, and / or hyperparameter that can be used to control, modify, adapt, and / or tune the behavior, learning, and / or performance of an algorithm, model, regressor, and / or classifier. In some cases, a parameter is used to increase or decrease the effect of an input (e.g., feature) on an algorithm, model, regressor, and / or classifier. As a non-limiting example, in some embodiments, a parameter is used to increase or decrease the effect of a node (e.g., a neural network) that includes one or more activation functions. The assignment of a parameter to a particular input, output, and / or function is not limited to any one paradigm for a given algorithm, model, regressor, and / or classifier, but can be used in any suitable algorithm, model, regressor, and / or classifier architecture for the desired performance. In some embodiments, a parameter has a fixed value. In some embodiments, the value of a parameter is manually and / or automatically adjustable. In some embodiments, the value of a parameter is modified by a validation and / or training process for an algorithm, model, regressor, and / or classifier (e.g., by an error minimization and / or backpropagation method). In some embodiments, the algorithms, models, regressors, and / or classifiers of the present disclosure include a plurality of parameters.In some embodiments, the plurality of parameters are n parameters, where n ≥ 2, n ≥ 5, n ≥ 10, n ≥ 25, n ≥ 40, n ≥ 50, n ≥ 75, n ≥ 100, n ≥ 125, n ≥ 150, n ≥ 200, n ≥ 225, n ≥ 250, n ≥ 350, n ≥ 500, n ≥ 600, n ≥ 750, n ≥ 1,000, n ≥ 2,000, n ≥ 4,000, n ≥ 5,000, n ≥ 7,500, n ≥ 10,000, n ≥ 20,000, n ≥ 40,000, n ≥ 75,000, n ≥ 100,000, n ≥ 200,000, n ≥ 500,000, n ≥ 1×106, n ≥ 5×10. 6 , or n ≥ 1×10 7 . Thus, the algorithms, models, regressors, and / or classifiers of the present disclosure cannot be implemented mentally. In some embodiments, n is between 10,000 and 1×107, between 100,000 and 5×106, or between 500,000 and 1×106. In some embodiments, the algorithms, models, regressors, and / or classifiers of the present disclosure operate in a k-dimensional space, where k is a positive integer greater than or equal to 5 (e.g., 5, 6, 7, 8, 9, 10, etc.). Thus, the algorithms, models, regressors, and / or classifiers of the present disclosure cannot be implemented mentally.

[0061] To avoid ambiguity, in this specification, specific features (e.g., integers, characteristics, values, uses, diseases, formulas, compounds, or groups) described in connection with a particular aspect, embodiment, or example of the present disclosure are intended to be understood as applicable to any other aspect, embodiment, or example described herein, unless inconsistent. Thus, such features may be used in combination with any of the definitions, claims, or embodiments defined herein, as appropriate. All features disclosed in this specification (including any appended claims, abstract, and drawings), and / or all steps of any method or process so disclosed, may be combined in any combination, except combinations in which at least some of the features and / or steps are mutually exclusive. The present disclosure is not limited to any details of any disclosed embodiment. The present disclosure extends to any novel one, or novel combination, of the features disclosed in this specification (including any appended claims, abstract, and drawings), or any novel one, or any novel combination, of the steps of any method or process so disclosed.

[0062] Furthermore, as used herein, the term "about" means that dimensions, sizes, formulations, parameters, shapes, and other quantities and characteristics are not exact and need not be exact, but may be approximate and / or larger or smaller, as desired, reflecting tolerances, conversion factors, rounding, measurement errors, and equivalents, and other factors known to those of skill in the art. Generally, dimensions, sizes, formulations, parameters, shapes, or other quantities or characteristics are "about" or "approximately", whether or not so expressly stated. Note that embodiments of very different sizes, shapes, and dimensions may employ the described arrangements.

[0063] Furthermore, the transitional terms "comprising", "consisting essentially of", and "consisting of", when used in the appended claims, define the claims in terms of whether additional claim elements or steps not recited, in their original or amended forms, are excluded from the claims where they exist. The term "comprising" is intended to be inclusive or open-ended and does not exclude any additional unrecited element, method, step, or material. The term "consisting of" excludes any element, step, or material not specified in the claims, and in the latter case, excludes impurities ordinarily associated with the specified material(s). The term "consisting essentially of" limits the claims to the specified element, step, or material(s), and those that do not substantially affect the basic and novel characteristics of the claimed invention. All embodiments of the present invention can alternatively be more specifically defined by any of the transitional phrases "comprising", "consisting essentially of", and "consisting of".

[0064] Figures 1A and 1B show a computer system 100 for querying a combinatorial synthesis library. Referring to FIGS. 1A and 1B, in an exemplary embodiment, computer system 100 comprises one or more computers. For purposes of illustration in FIGS. 1A and 1B, computer system 100 is represented as a single computer including all of the functions of the disclosed computer system 100. However, the present disclosure is not so limited. The functions of computer system 100 may be distributed among any number of networked computers and / or may exist in each of several networked computers and / or virtual machines. Those skilled in the art will understand that a variety of different computer topologies are possible for computer system 100 and that all such topologies are within the scope of the present disclosure.

[0065] Referring to FIGS. 1A and 1B with the above in mind, computer system 100 includes one or more central processing units (CPUs) 64, one or more graphics processing units (GPUs) 74, a network or other communication interface 76, a user interface 68 (e.g., including an optional display 70 and an optional keyboard 72 or other form of input device), a memory 58 (e.g., random access memory, persistent memory, or a combination thereof), one or more magnetic disk storage and / or persistent devices 60 optionally accessed by one or more controllers 62, one or more communication buses 12 for interconnecting the foregoing components, and a power supply 66 for powering the foregoing components. Unless the components of memory 58 are persistent, the data in memory 58 can be seamlessly shared with non-volatile memory 60 using known computing techniques such as caching. Memory 60 can include mass storage located remotely from central processing unit(s) 64. In other words, some of the data stored in memory 58 and / or memory 60 can actually be outside of computer system 100, but can be electronically accessed by computer system 100 via network interface 76 over the Internet, an intranet, or other form of network or electronic cable and hosted on a computer. In some embodiments, computer system 100 utilizes models implemented from memory associated with one or more graphical processing units 74 to improve the speed and performance of the system. In some alternative embodiments, computer system 100 utilizes models executed from memory 58 rather than memory associated with the graphical processing unit.

[0066] The memory 58 of computer system 100 includes ● An optional operating system 78 including procedures for handling various basic system services, and ● A combinatorial synthesis library query model 80 for querying a combinatorial synthesis library, ● A combinatorial synthesis library 82 containing a plurality of compounds 92, wherein the combinatorial synthesis library represents a plurality of reaction types 82, each such reaction type is indexed by a reaction type key 84 and has a corresponding mapping to a plurality of reactants, and each reactant 86 of each corresponding plurality of reactants has a corresponding mapping to a plurality of synthons 88, each such synthon is indexed by a synthon key 90, such that each respective compound 92 is indexed by a reaction type key 94, and this is used to determine reactants 96 and synthon keys 98 for each respective compound, the combinatorial synthesis library 82; ● A molecular query 101 used to query a combinatorial synthesis library 82, wherein the molecular query is in the form of a graph 102 including a plurality of nodes 102 and a plurality of edges 104, each node of the plurality of nodes represents an atom, and each node is connected to another node of the plurality of nodes by at least one edge of the plurality of edges, and the edge represents a covalent bond between two nodes, the molecular query 101; ● A molecular encoder model 110 optionally including a message passing neural network including a plurality of message passing layers 112 collectively including a first plurality of parameters 114, wherein the molecular encoder model outputs a query vector by applying the first plurality of parameters 114 to the graph 102 in response to receiving the molecular query 101 in the form of the graph 102, the molecular encoder model 110; ● A reaction query generator model 116 including a second plurality of parameters 118, wherein the reaction query generator model 116 provides a first reaction type 82-1 of the plurality of reaction types 82 as an output from the reaction query generator model 116 by applying the second plurality of parameters 118 to the query vector in response to receiving the query vector, the reaction query generator model 116; ● A synthon query generator model 120 including a third plurality of parameters 122, wherein the synthon query generator model 120, in response to receiving an identification of a reactant, applies the third plurality of parameters 122 to each respective reactant to provide, as an output from the synthon query generator model 120, a corresponding synthon 88, and stores the synthon query generator model 120.

[0067] In some implementations, one or more of the data elements or modules identified above in the computer system 100 are stored in one or more of the aforementioned memory devices and correspond to a set of instructions for performing the above functions. The identified data, modules, or programs (e.g., sets of instructions) above need not be implemented as separate software programs, procedures, or modules, and thus, various subsets of these modules can be combined in various implementations or otherwise rearranged. In some implementations, the memory 60 (and optionally the memory 58) optionally stores a subset of the modules and data structures identified above. Further, in some embodiments, the memory 60 (and optionally the memory 58) stores additional modules and data structures not described above.

[0068] Since a system for querying a combinatorial synthesis library has been disclosed, a method for performing such a query is detailed with reference to FIG. 2 and discussed below.

[0069] Referring to block 200, in some embodiments, a system and method for querying a combinatorial synthesis library 82 are provided. The combinatorial synthesis library 82 includes a plurality of compounds 92 and represents a plurality of reaction types 82. Each respective reaction type 82 of the plurality of reaction types has a corresponding mapping to a corresponding plurality of reactants 86. Each respective reactant 86 of each corresponding plurality of reactants has a corresponding mapping to a corresponding plurality of synthons 88.

[0070] One aspect of the present disclosure focuses on combinatorial synthetic chemical libraries (CSLs), which are accessible through a combination of a smaller set of readily available building blocks and a set of pre-described, generally successful multi-component synthesis rules, resulting in a vast expansion of chemical space (e.g., more than 10 8 over, more than 10 9 over, more than 10 10 over, or more than 10 11 over compounds). A diagram of such a configuration is shown in Figure 3A. In some embodiments, the CSL is composed of a set of multi-dimensional synthesis tables. In some such embodiments, each synthesis table describes a multi-component chemical reaction that can be represented by an exponent

Number

[0071] Following

[44] , the term "R group" or simply "group" is used to refer to a placeholder group in a chemical reaction,

Number

Number

Number

[0072] Each R group spans a number of molecular building blocks, perhaps a large number, called synthons that can be utilized in the corresponding reaction. A synthon is

Number

Number

[0073] For convenience, in some embodiments, the notation σ:R→P(S) is used to represent a function that returns the set of synthons belonging to a particular R group r

Number

[0074] Product

Number

Number

Number

Number

[0075] Thus, a synthon-based library D ≡ (T, R, S, f, Ψ, σ) is completely characterized by its reaction T, R groups R, and synthons S, together with the synthesis rule f, the mapping Ψ of the reaction to the R groups, and the mapping σ of the R groups to the synthons. Due to the combinatorial nature of such constructs, they can be used to construct libraries that span a vast region of chemically accessible space where compounds can be obtained (i) with high probability, (ii) at reasonable cost, and (iii) with short lead times.

[0076] Using the language of probability, the distribution over X induced by D can be described through the following factorization.

Number

Number

Number

[0077] As described, all valid (t,u) pairs in D are equally likely under p. Note that if all products in D can be reached following only a single synthetic route, p(x|D) is a uniform distribution over the portion of X accessible by D.

[0078] Referring to block 202, in some embodiments, the plurality of reaction types 82 includes 20 or more reaction types, and the combinatorial synthesis library 82 includes 100 or more compounds 92 for each reaction type 82 of the plurality of reaction types. In some embodiments, the plurality of reaction types 82 includes 20, 30, 40, 50, 60, 70, 80, 90, 100, 500, 1000, 2000, 3000, 5000, 10,000 or more reaction types, and the combinatorial synthesis library 82 includes 100, 200, 300, 400, 500, 1000, or 10,000 or more compounds 92 for each reaction type 82 of the plurality of reaction types.

[0079] Referring to block 204, the molecular query 101 is input to the molecular encoder model 110. The query is graph-based. In some embodiments, the molecular encoder model 101 includes a message-passing neural network including a plurality of message-passing layers 112 collectively including a first plurality of parameters 114. Inputting the graph-based query to the molecular encoder model 110 results in a query vector as the output of the molecular encoder model by applying the first plurality of parameters 114 to the graph-based query.

[0080] In some embodiments, the Molecular Encoder model 101

Number

Number

Number

Number

[0081]

[0082] Note that in some embodiments, the Molecular Encoder model 110 takes as input a molecular graph and can thus generate queries for compounds not in the library D. This is useful in some embodiments for finding analogs (i.e., compounds that can be purchased from a catalog and are chemical analogs of the query molecule) by a catalog.Referring to block 206, in some embodiments, a graph-based query is a single graph 102 that includes a plurality of nodes 102 and a plurality of edges 104, and each node 102 of the plurality of nodes is connected to another node of the plurality of nodes by at least one edge 104 of the plurality of edges. In some embodiments, the single graph 102 represents a single molecular compound, each node of the plurality of nodes represents an atom of the single molecular compound, and each edge of the plurality of edges represents a covalent bond between two atoms within the single molecular compound. In some embodiments, the plurality of nodes represents each non-hydrogen atom of the molecular compound. In some embodiments, the plurality of nodes represents each atom of the molecular compound. In some embodiments, the plurality of edges represents each covalent bond of the molecular compound.

[0083] In some such embodiments, each respective node of the plurality of nodes is associated with any one, two, three, four, or more of the following characteristics: (i) the corresponding element type of the plurality of element types, (ii) the node degree of the plurality of node degrees, (iii) the hybridization of the plurality of hybridizations, (iv) the number of bonded hydrogens, (v) the formal charge from the set of formal charges, and (vi) the binary representation of aromaticity.

[0084] Referring to block 208, in some embodiments, each respective node 102 of the plurality of nodes 102 is associated with (i) the corresponding element type of the plurality of element types, (ii) the node degree of the plurality of node degrees, (iii) the hybridization of the plurality of hybridizations, (iv) the number of bonded hydrogens, (v) the formal charge from the set of formal charges, and (vi) the binary representation of aromaticity.

[0085] Referring to block 210, in some embodiments, each respective bond 104 of the plurality of bonds is associated with (i) the bond type, (ii) the binary representation of conjugation, (iii) the binary of the indication of whether each respective bond is within a ring, and (iv) the indication of stereochemistry.

[0086] In some embodiments, the atoms and bonds of a molecular compound are represented as a graph as a set of binary features. In some embodiments, these binary features are described in Tables 1 and 2, respectively. In an exemplary implementation, the node dimension and the edge dimension are each set to 64. Thus, in this example, the atom embedding consists of 50 × 64 = 3,200 parameters, and the bond embedding consists of 12 × 64 = 768 parameters. [Table 1] [Table 2]

[0087] In some embodiments, a single graph represents a query molecular compound as a set of atomic features and a set of bond features. For example, in some embodiments, a single graph represents a query molecular compound as a set of atomic features and a set of bond features of Tables 1 and 2, respectively. In some such embodiments, the set of atomic features includes element type, node degree, hybridization, chirality, bound hydrogen, formal charge, aromaticity, as shown in Tables 1 and 2, and the set of bond features includes bond time, conjugation, intra-ring, and stereochemistry. In some such embodiments, each non-hydrogen atom in the query molecular compound is represented by 500, 1000, 1500, 2000, 2500, 3000, or more than 3000 parameters in the set of atomic features, and each covalent bond in the molecular compound is represented by 100, 200, 300, 400, 500, 600, 700, or more than 700 additional parameters in the set of bond features.

[0088] Referring to block 212, in some embodiments, a single graph 102 represents a single molecular compound present in a combinatorial synthesis library. In some embodiments, a single graph 102 represents a single molecular compound that is not present in the combinatorial synthesis library.

[0089] Referring to block 214, in some embodiments, a single graph 102 represents a weighted complex of a first graph of a first molecular compound and a second graph of a second molecular compound. In some embodiments, a single graph 102 represents a complex of two or more molecular compounds.

[0090] Referring to block 216, in some embodiments, a single graph 102 represents a weighted complex of a plurality of graphs of a second plurality of compounds, and the second plurality of compounds have a common property. Referring to block 218, in some embodiments, the common property is a Tanimoto distance less than a threshold with respect to each other compound of the second plurality of compounds. Referring to block 220, in some embodiments, the common property is a binding coefficient to a macromolecular target that is less than a threshold.

[0091] Referring to block 220, a query vector z (e.g., represented by the above formula) is input into a reaction query generator model 116 that includes a second plurality of parameters 118. This results in an output from the reaction query generator model 116 of a first reaction type among a plurality of reaction types by applying the second plurality of parameters 118 to the query vector.

[0092] In some embodiments according to block 220, the query of block 204

Number

[0093] In some embodiments, the reaction query generator model 116 is utilized. In some embodiments, the reaction query generator model 116 can be represented by the following operations. q T =ReactionQueryGenerator(z) (11) In some such embodiments, ReactionQueryGenerator

Number

Number

[0094] Thus, referring to block 224, in some embodiments, the output from the reaction query generator model 116 is used to identify a first reaction key type 84 among a plurality of reaction key types through a first query key lookup, and each reaction key type among the plurality of reaction key types represents a synthetic reaction that can be used to synthesize one or more compounds 92 within the combinatorial synthesis library 82. In some such embodiments, the output from the reaction query generator model 116 is the reaction key type t. In other embodiments, the output from the reaction query generator model 116 is a probability distribution of key types p(t|z,D), such as derived from equation (12), from which a specific reaction key type t is obtained by sampling the probability distribution.

[0095] Referring to block 222, in some embodiments, the reaction query generator model 116 is a two-layer perceptron (not counting the input and output layers) with intermediate ReLU activation. In some embodiments, the reaction query generator model 116 is an N-layer perceptron (MLP), where N is a positive integer greater than or equal to 2 (e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10), representing the number of hidden layers in the model. In some embodiments, the MLP is a class of feed-forward artificial neural network (ANN) that includes at least three node layers: an input layer, one or more hidden layers, and an output layer. In such embodiments, except for the input nodes, each node is a neuron that uses a non-linear activation function. In some embodiments, the activation function is the rectified linear unit ReLU activation. In some embodiments, the activation function is the rectified linear unit (ReLU) activation function, the leaky ReLU activation function, or other functions such as the saturated hyperbolic tangent, identity, binary step, logistic, arcTan, soft sign, parametric rectified linear unit, exponential linear unit, softPlus, bent identity, softExponential, sine, sinusoid, Gaussian, or sigmoid function, or any combination thereof. Further disclosure regarding suitable MLPs that can function as the reaction query generator model 116 can be found in Vang-mata ed., 2020, Multilayer Perceptrons: Theory and Applications, Nova Science Publishers, Hauppauge, New York. In some embodiments, the reaction query generator model 116 has the architecture of any of the models disclosed in the above-defined section.

[0096] Given the reaction t sampled from block 220, the required R groups (via Ψ) are revealed, and further, the synthons eligible for each group (via σ) are revealed. Thus, referring to block 226, in some embodiments, for each reactant 86, the corresponding synthon 88 inputs the respective reactant into a synthon query generator model 120 that includes a third plurality of parameters 122, thereby applying the third plurality of parameters 122 to the respective reactant to obtain, as an output from the synthon query generator model 120, the corresponding synthon, and determining, for each reactant among the first plurality of reactants corresponding to the first reaction type, from among the corresponding plurality of synthons mapped to the respective reactants 86. This results in a set of synthons, and each synthon in the set of synthons corresponds to a reactant among the first plurality of reactants. In other words, the synthon query generator model 120 is implemented to decode the synthon tuple for reaction t. In some embodiments, the synthon query generator model 120 implements the operation SynthonQueryGenerator.

Number

Number

Number

Number

[0097] Referring to block 228, in some embodiments, the synton query generator model 120 is a two-layer perceptron with intermediate ReLU activation. In some embodiments, the synton query generator model 120 is an N-layer perceptron (MLP), where N is a positive integer greater than or equal to 2 (e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10), representing the number of hidden layers in the model. In some embodiments, the MLP is a class of feed-forward artificial neural network (ANN) that includes at least three node layers: an input layer, one or more hidden layers, and an output layer. In such embodiments, except for the input nodes, each node is a neuron that uses a non-linear activation function. In some embodiments, the activation function is the rectified linear unit ReLU activation. In some embodiments, the activation function is the rectified linear unit (ReLU) activation function, the leaky ReLU activation function, or other functions such as the saturated hyperbolic tangent, identity, binary step, logistic, arcTan, soft sign, parametric rectified linear unit, exponential linear unit, softPlus, bent identity, softExponential, sine curve, sine, Gaussian, or sigmoid function, or any combination thereof. Further disclosure regarding suitable MLPs that can function as the synton query generator model 120 in some embodiments of the present disclosure can be found in Vang-mata ed., 2020, Multilayer Perceptrons: Theory and Applications, Nova Science Publishers, Hauppauge, New York. In some embodiments, the synton query generator model 120 has the architecture of any of the models disclosed in the above-defined section.

[0098] Referring to block 230, in some embodiments, the output from the synton query generator model 120 is used to identify the synton key 80 for the corresponding synton through a second query key lookup.

[0099] Referring to block 232, in some embodiments, the first plurality of parameters 114 (of the molecular encoder model 110) includes 100,000 parameters, the second plurality of parameters 118 (of the reaction query generator model 116) includes 5,000 parameters, and the third plurality of parameters 122 (of the synthon query generator model 120) includes 5,000 parameters. In some embodiments, the first plurality of parameters 114 (of the molecular encoder model 110) includes 10,000, 50,000, 100,000, or 1×10 6 parameters, the second plurality of parameters 118 (of the reaction query generator model 116) includes 1,000, 2,000, 3,000, 4,000, 5,000, or 10,000 parameters, and the third plurality of parameters 122 (of the synthon query generator model 120) includes 1,000, 2,000, 3,000, 4,000, 5,000, or 10,000 parameters.

[0100] Referring to block 234, in some embodiments, the first plurality of reactants 86 includes three or more reactants, and the corresponding mapping for the corresponding plurality of synthons for the reactants among the three or more reactants includes ten or more synthons. In some embodiments, the first plurality of reactants 86 includes 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more reactants. In some embodiments, the corresponding mapping for the corresponding plurality of synthons for the reactants in the first plurality of reactants 86 includes 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, or 100 or more synthons.

[0101] Referring to block 236, the molecular structure (compound) 92 is identified in a combinatorial synthesis library 82 that includes a set of synthons 88 arranged according to a synthesis rule f associated with the first reaction type 82. As used herein, the synthesis rule f generates a compound x from a reaction and a synthon tuple pair (t; u) (e.g., x := f(t; u)).

[0102] Referring to block 238, in some embodiments, the plurality of compounds in combinatorial synthesis library 82 includes more than one billion compounds, and the identified molecular structure is any one of the more than one billion compounds that satisfy the query. Referring to block 240, in some embodiments, the plurality of compounds in combinatorial synthesis library 82 includes more than one trillion compounds, and the identified molecular structure is any one of the more than one trillion compounds that satisfy the query. In some embodiments, combinatorial synthesis library 82 includes 1×10 8 、1×10 9 、1×10 10 、1×10 11 、1×10 12 、or 1×10 13 compounds.

[0103] Combinatorial synthesis library variational autoencoder.

[0104] Next, consider the task of retrieving product x from library D, which amounts to finding reaction t and synthon tuple u that satisfy x = f(t, u). This can be cast as an inference problem of finding p(t, u|x, D). The latent variable model of x leads to the following variational formulation.

Number

[0105] Here, p(z) represents the prior distribution of the latent variable. First, reaction t is selected, and then, conditioned on t, the synthon

Number

Number

Number

[0106] This encodes molecule x into the latent space z ~ q(z|x), first decodes the reaction type t ~ p(t,z,D), and then, conditioned on the sampled reaction, for i = 1, … k t (where t is a positive integer greater than or equal to 2) for the radical

Number

Number

[0107] The resulting latent variable model can be referred to as a combinatorial synthesis library variational autoencoder. Figures 3B - 3H show a step - by - step depiction of the method. The three main modules that make up this autoencoder are (i) the library encoder, (ii) the molecule encoder, and (iii) the molecule decoder. Module (ii) forms the basis of q(z|x), and modules (i) and (iii) form the basis of p(t,u|z,D).

[0108] Library encoder. In some embodiments, the combinatorial synthesis library D ≡ (T,R,S,f,ψ,σ) is hierarchically organized, with synthons S at the bottom of the hierarchy, R - groups R in the middle, and reactions T at the top. Here, we describe a strategy for learning the relevant hierarchy of representations that describe the library at these three levels of resolution in an end - to - end fashion. These representations can then be used when retrieving results from the library given a query. The library encoder is shown in Figures 3C - 3E.

[0109] Starting from the bottom of the hierarchy with Synthon S, a graph neural network is used to learn the representation of each synthon in a fully inductive manner, SynthonEncoder,

Number

[0110] Moving up the hierarchy, a deep set neural network is used to represent R of the R group by summarizing the representations of the synthons belonging to a particular base. Formally, RGroupEncoder:

Number

[61] , the R group encoder has the form of RGroupEncoder

Number

Number

[0111] In some embodiments, φ and ρ are parameterized as a model other than a multi-layer perceptron. For example, in some embodiments, φ and ρ are parameterized as any of the models disclosed in the definition section for the "model" disclosed above. In some embodiments, for example, a more performant set-vector neural network such as a set transformer

[31] is used. However, the advantage of using a multi-layer perceptron is that, for all

Number

Number

[0112] In some embodiments, to represent a reaction, another deep set neural network, ReactionEncoder:

Number

[0113] Putting these together, the representations cascade as follows.

Number

Number

Number

[0114] In some embodiments, likelihood factorization is considered, such that given the molecular representation z, the molecular decoder first decodes the reaction type and then proceeds by separately decoding one synthon for each group, conditioned on the reaction type. Thus, in some implementations, the key vector is used for each reaction as well as for each synthon to compare with the associated query vector. Thus, the ReactionKeyGenerator used in some embodiments has a function to generate a key vector when given a reaction representation

Number

Number

Number

Number

[0115] In some embodiments, each of these key generators is parameterized as some other form of model. For example, in some embodiments, each of these key generators is parameterized as any of the models disclosed in the above definition section under the defined "model".

[0116] Exemplary training. For large CSLD*, encoding the entire library in each iteration of the training loop may require an excessive amount of GPU memory. To overcome this, in some embodiments, a mini-batch strategy is used where a random subset

Number

[0117] Algorithm 1 - Training Procedure: input full library D*, library subsampler p(D|D*), batch size N, model parameters θ, KL divergence weight β≥0, optimization selection while not reaching the topping criterion do # Prepare mini-batch Sample library D~p(D|D*) Sample mini-batch of reaction-synthon chain pairs, (t n, u n ) ~ p(t, u|D) (for n = 1…N) Form the corresponding product through synthesis, x n = f(t n , u n )(for n = 1…N) # Encode the library According to Equation (5), each

Number

Number

Number

Number

Number

Number

Number

Number

Number

Number

Number

Number

Number

Number

Number

Mathematics

Mathematics

Mathematics

Mathematics

Mathematics

[0118] Posterior density estimation. Given the trained generative model p θ (x|z), the product can also be sampled via (x,z)~p θ (x|z)p(z) (discarding z) (Reference 3). However, due to the batch sampling strategy outlined in this specification, this generally does not well match the uniform distribution over products in D due to the bias introduced by the batch sampling strategy discussed above (first sampling the reactions uniformly and then, conditioned on the reactions, sampling the products uniformly).

[0119] This can be corrected using importance weighting during the training phase. However, in some embodiments, another approach is implemented by utilizing the following post - density estimation strategy

[14] . A number of products are sampled from the target distribution x ~ p(x|D), and these products are encoded through the molecular encoder

Number

Number

[0120] In some implementations, the products are then,

Number

[0121] Computational complexity, scalability, and efficiency. In this section, the computational complexity, scalability, and efficiency of some embodiments of the systems and methods of the present disclosure are described.

[0122] First, the library D can be encoded with a complexity of O(|S| + |R| + |T|), where it is noted that the constant depends on the complexity of the synthon, R groups, and reaction encoder. Nevertheless, this is logarithmic compared to simply encoding each product in D, which has a complexity of O(|D|).

[0123] Of more note is the computational complexity of the molecular decoder. To clarify, consider a simplified D consisting of a single k-component reaction. M i shall denote the number of synthons for the groups i = 1, …, k. Simply, the nearest neighbor look-up in D has [Number] a complexity of. On the other hand, in some embodiments, the systems and methods of the present disclosure perform look-ups independently using the synthons within each R group, [Number] achieving a complexity - logarithmic improvement. Thus, the structure of the molecular decoder of the present disclosure is well - suited for very large - scale combinatorial libraries of interest in early - stage drug discovery.

[0124] Another advantage of the disclosed decoding strategy is that it depends minimally on autoregression. In fact, in some embodiments, only a single step of autoregression is required, regardless of the size of the generated graph (exactly 2 autoregressive lengths). Thus, the disclosed systems and methods scale to molecular graphs of widely variable size that follow combinatorial synthesis constructs.

[0125] Furthermore, the disclosed systems and methods are guaranteed to generate chemically valid and synthetically accessible molecular graphs without performing explicit validity checks. This is comparable to prior work where the validity of each candidate action is verified at each step of the autoregression and invalid actions are excluded from the selection set. Chemoinformatics libraries such as RDKit

[30] have efficient C++ implementations for these checks, but nevertheless significantly increase the runtime. Additionally, these models have been shown to generate invalid molecular graphs at a significantly high rate when no explicit validity checks are performed

[23] .

[0126] Final note. The disclosed combinatorial synthesis library variational autoencoder is a new graph-based generative model for the navigation of combinatorial synthesis libraries. The disclosed model utilizes minimal autoregression, enables efficient generation of large molecular graphs, and improves scalability. The compounds generated by the disclosed model are chemically valid and cost-effectively accessible. In some embodiments, the disclosed model is a neural database that provides random access to non-exhaustive libraries. In the following experiments, the ability of the disclosed model to model ultra-large and realistic on-demand libraries is demonstrated, opening the way to more scalable strategies for the exploration of non-exhaustive chemical libraries for early-stage drug discovery.

[0127] In some embodiments, the disclosed synthon lookups in the decoder scale linearly with the number of synthons in the R group, which can present a challenge as the library continues to add more synthons per R group. In some embodiments, this is mitigated using a more scalable query-key design [11,26]. Further, in some embodiments, due to the presence of significant R group symmetry (such as in polymers), some embodiments of the decoder include modifications to break parity, and some such modifications may not allow the same convenient parallelization. Finally, softmax has limitations in mapping from real-valued potentials to probabilities selected from

[55] . Thus, in some embodiments, instead of using softmax, a sparse or alternative recognition softmax variant is used.

[0128] Synthon Encoder. In some embodiments, the SynthonEncoder is a message-passing graph neural network that takes as input node, edge, and graph features and generates updates to these features after each round of message passing in a residual fashion. There is a single message-passing layer below.

[0129] Prerequisites. The graph

Number

Number

Number

[0130] Edge model. The edge model updates the edge features as follows. Edge

Number

Number

Number

Number

[0131] Node model. Node

Number

Number

Number

Number

[0132] Graph model. The graph features are updated by accumulating the messages from each node in V as follows. The graph

Number

Number

[0133]

Number

Number

Number

[0134] Summarizing the above. The message passing neural network applies a sequence of message passing layers as outlined above. In one implementation form according to the present disclosure, the node, edge, and graph feature dimensions are set to 64, and four message passing layers are utilized. For the input graph features, in some embodiments, a zero vector is used. The final graph features serve as a synthon representation, which is used when generating a cascade representation for the synthon keys as well as the R groups and reactions. In total, one embodiment of the synthon encoder according to the present disclosure is described by 152,832 parameters.

[0135] R group encoder. In some embodiments according to the present disclosure, the R group encoder (RGroupEncoder) follows the design outlined in DeepSets, where the synthon representations are each processed separately by an MLP and pooled together (here, average pooling is used to enable the network to focus on the characteristics of the distribution of the synthons belonging to the R group), and the result is further processed by another MLP. In some embodiments according to the present disclosure, both MLPs are set as two-layer networks having a ReLU activation between two linear operations. In some embodiments, all dimensions are 64, and thus, the RGroupEncoder utilizes a total of 16,640 parameters. In some alternative embodiments, any of the models disclosed in the above definition section is used to construct the R group encoder.

[0136] Reaction Encoder. In some embodiments, the ReactionEncoder follows the same design as the RGroupEncoder, except that sum pooling is used instead of average pooling to allow the network to focus on multiple sets of R groups in the reaction. In some embodiments, all dimensions are 64, and thus the ReactionEncoder also utilizes a total of 16,640 parameters. In some alternative embodiments, any of the models disclosed in the above definition section are used to construct the reaction encoder.

[0137] Synthon Key Generator. In some embodiments, the SynthonKeyGenerator is an MLP that generates a synthon key from a synthon representation. In some embodiments, linear layers are used. In some embodiments, both the input and output dimensions are set to 64, and thus in such embodiments, the SynthonKeyGenerator utilizes a total of 4,160 parameters. In some alternative embodiments, any of the models disclosed in the above definition section are used to construct the synthon key encoder.

[0138] Reaction Key Generator. In some embodiments, the ReactionKeyGenerator is an MLP that generates a reaction key from a reaction representation. In some embodiments, linear layers are used. In some embodiments, both the input and output dimensions are set to 64, and in some embodiments, the ReactionKeyGenerator utilizes a total of 4,160 parameters. In some alternative embodiments, any of the models disclosed in the above definition section are used to construct the reaction key generator.

[0139] Molecular Encoder. In some embodiments, the Molecular Encoder generates molecular queries and utilizes the same message passing design as the Synthon Encoder. In some such embodiments, the node, edge, and graph features are all set to 64, and 4 layers of message passing are used. However, in some embodiments, unlike the Synthon Encoder, the Molecular Encoder has an additional variational linear layer that generates conditional mean and conditional log variance vectors from the graph features generated by the final message passing round. In some embodiments, the variational linear layer takes 64-dimensional graph features as input and generates a 128-dimensional output, which is split into a mean part and a log variance part. In some embodiments, the Molecular Encoder utilizes a total of 152,832 + 8,320 = 161,152 parameters. In some alternative embodiments, any of the models disclosed in the above definition section are used to construct the molecular encoder.

[0140] Molecular Query Processing Network. In some embodiments, the molecular queries are regularized to be close to the prior probability p(z) as a result of the VAE objective. To enable the decoder to utilize these features better, in some embodiments, the molecular queries are processed by an MLP before being passed to the reaction and synthon query generators. For simplicity, in some embodiments, a simple 2-layer MLP with intermediate ReLU activation, where all dimensions are 64 and which results in 8,320 parameters, is used. In some embodiments, any of the models disclosed in the above definition section are used instead of such an MLP.

[0141] Reaction Query Generator. In some embodiments, the ReactionQueryGenerator is an MLP that generates reaction queries from molecule queries. In some embodiments, a two-layer network with intermediate ReLU activation is used. In some embodiments, all dimensions are set to 64, and thus the total number of parameters of the ReactionQueryGenerator utilizes 8320. In some embodiments, any of the models disclosed in the above definition section is used instead of such an MLP for the reaction query generator.

[0142] Synthon Query Generator. In some embodiments, the SynthonQueryGenerator is an MLP that generates synthon queries from molecule queries, reaction representations, and R-group representations. In general, in some embodiments, these three feature types are concatenated when a common dimensionality of 64 is used throughout to keep the implementation simple, and in some embodiments, a sum is used instead of concatenation (which can be shown to be equivalent to concatenation with additional constraints on the weight matrix of the subsequent linear layer). Similar to the ReactionQueryGenerator, in some embodiments, a two-layer MLP with intermediate ReLU activation results in a total of 8320 parameters for the SynthonQueryGenerator. In some alternative embodiments, any of the models disclosed in the above definition section is used instead of such an MLP for the synthon query generator.

Examples

[0143] Example 1 - Comparison of the disclosed model with JT-VAE and RationaleRL. Data. To demonstrate the capabilities of the disclosed system and method with respect to the types of real-world combinatorial synthesis catalogs employed in today's large-scale hit discovery programs, an experimental example was conducted using the Enamine REadily AccessibLe (REAL) library, which consists of 340K synthons and over 1000 reactions. The reactions in REAL range from 2 to 4 components, and the number of synthons per R group can range from single digits to tens of thousands. In total, the REAL library describes the chemical space of over 16 billion commercially available compounds [4], which can be obtained at low cost on the order of 3 - 4 weeks. Advantageously, using this library with the disclosed system and method, exemplary experiments can be reproduced and further research in the machine learning community regarding combinatorial synthesis libraries can be facilitated.

[0144] Training. During training, subsets of the library were sampled as follows. Of the approximately 1300 reaction types in the REAL database, 20 reactions were first randomly and uniformly sampled, and then 100 products per reaction, including the relevant synthons in the library subset, were sampled. Thus, these library subsets each describe approximately 300K - 1.5M compounds, which is significantly smaller than the full library of 16 billion compounds. See Algorithm 1 above for details.

[0145] Testing. For test-time inference, decoding was for the full library of 16 billion compounds. This constitutes a test-time distribution shift with respect to training, but it was observed that the disclosed system and method generalize well to the full library without modification. For completeness, the analysis of the test-time distribution shift is shown in Figure 8.

[0146] As described in Algorithm 1, for ease of handling, in each training iteration, a small subset of the full library is sampled. In particular, the library subsampler uniformly utilized the first sample across the reactions and subsequently randomly sampled a fixed number of products in each reaction. The synthons associated with these sampled products include the synthons of the library subset. In the training of the CSLVAE model used in this experiment, 20 reactions and 100 products per reaction were sampled, which results in a mini-batch of 2000 products. As described herein, these associated library subsets each describe a chemical space of approximately 300K - 1.5M compounds. However, during testing, decoding is performed with the full library of 16B compounds, which constitutes a fairly dramatic test time distribution shift. Figure 8 shows the magnitude of the test time distribution shift with respect to the reconstruction quality (measured by the average likelihood). Starting by sampling the five library subsets as described above, 2000 compounds are depicted for each subset. Each of these five drawings is represented by a different line in the figure. These library subsets are then extended by first including all the synthons across the 20 sampled reactions, calculating the average reconstruction likelihood for the same 2000 products according to the extended library, adding all the synthons for 100 additional randomly sampled reactions, then adding all the synthons for yet another 100 randomly sampled additional reactions, and finally further increasing the library by including all the synthons and all the reactions. At each step, the library subset describes an increasingly large chemical space. It is observed that the average likelihood decreases in a nearly linear logarithmic fashion with respect to the library size as a result of this distribution shift.

[0147] Reconstruction and generation of molecules. One embodiment of the disclosed system and method (CSLVAE) was compared against two existing molecular graph generation models, namely JT-VAE

[23] and RationaleRL

[25] . All three models were trained from scratch on the Enamine REAL library.

[0148] In JT-VAE, the molecular graph is represented by a junction tree over chemical fragments. Decoding proceeds by first generating the junction tree in a depth-first manner, placing fragments at each node, and then orienting the fragments to match at the bond points. On the other hand, RationaleRL takes an initial rationale as input. The goal of the decoder is to complete the molecule autoregressively (one graph edit per step). In this example, a product is taken from the library, all synthons except one are removed, and the resulting graph is treated as the starting rationale. Thus, RationaleRL is tasked with generating the missing synthon in this example.

[0149] Table 3 summarizes the key findings from this exercise. First, note that the implementation forms of the disclosed systems and methods utilized have approximately 10 times fewer parameters than the two alternative forms considered, due to the inductiveness of the library encoder. All three methods achieve 100% chemical validity, but the disclosed systems and methods achieve this result without explicit validity checking. The average likelihood is calculated by taking the average of the reconstruction likelihoods for each compound over a large number of products sampled from the library. This is a measure of how well the model can reconstruct the complete molecular graph (e.g., on average, how likely it is to reproduce the query molecule through the decoder), and can also be roughly interpreted as a measure of coverage / reachability (e.g., what percentage of the library can be faithfully covered). Finally, the experimental examples highlight the challenges faced by existing graph generation models when applied to ultra-large combinatorial synthesis libraries, namely that they struggle to reliably generate compounds within the library. For JT-VAE, less than 1 in 34 compounds were found in REAL. On the other hand, RationaleRL has the advantage that although library completion occurs only about half the time (see Figure 7), starting rationales are provided in the form of compounds from REAL with all but one synthon removed. In contrast, the disclosed systems and methods guarantee that the results are intentionally from the library.

Table 3

[0150] Dataset preparation. The Enamine REAL library is a combinatorial synthesis library of approximately 1300 reaction types in the scope of 2-4 component reactions, along with approximately 340K synthons. Overall, the REAL library describes a chemical space of over 16B make-on-demand compounds. In two baseline models of this comparison, using the code provided by the authors on the full REAL library (both in the vocabulary generation step and in writing the products to disk) resulted in memory issues. Therefore, the two baselines were trained and evaluated on a subset of the full REAL library. The memory limitation is not faced by CSLVAE (the model of the present disclosure), which was trained on and evaluated on the full library. Thus, RationaleRL and JT-VAE can be compared item by item (each being trained on the same subset of REAL), while the results of CSLVAE are larger and reflect training on the more extensive and diverse full REAL library. RationaleRL and JT-VAE were not developed for the purpose of exploring combinatorial synthesis libraries, and their use in this comparison serves as an attempt to directly compare the disclosed method (CSLVAE) with the application of existing graph generation models to combinatorial synthesis libraries without modification.

[0151] To construct the data on which RationaleRL and JT-VAE were trained, 1300 reaction types were ranked by the number of products included in each reaction, and 50 of the medium-scale reactions were selected. In total, these reactions describe a chemical space of 125M compounds. Products were sampled from each reaction such that all synthons were represented, resulting in a training set of approximately 500K compounds.

[0152] Details of RationaleRL. The pre-training phase of RationaleRL, which trains a graph-based variational autoencoder that attempts to reconstruct a molecular graph from a rationale, was utilized (as the name suggests, it does not require RL and is part of a fine-tuning phase not utilized here). Given a starting rationale and a complete molecule, the decoder of RationaleRL autoregressively completes the molecule atom-by-atom and bond-by-bond. To form the starting rationale, take the product from REAL and remove all but one synthon. Thus, RationaleRL is tasked with completing the missing synthon given a complete molecule (as input to the encoder) and a starting rationale (as input to the decoder along with the latent code). In Figure 7, as the number of distinct reactions during training increases, the probability that the completions generated by RationaleRL result in products from the library drops sharply, and after 50 reactions are included, less than half of all completions result in products in the REAL library. These experiments are based on code provided by the authors and can be found at https: / / github.com / wengong-jin / multiobj-rationale.

[0153] Details of JT-VAE. JT-VAE is a graph-based variational autoencoder that generates molecules following the scaffold of a fragment tree structure. Unlike RationaleRL, JT-VAE autoregressively samples chemical fragments rather than atoms or bonds. Before training, the vocabulary generation step described in the JT-VAE paper was applied to products sampled from a subset of REAL used in the baseline experiments, resulting in a total of 325 fragments. These experiments are based on code provided by the authors and can be found at https: / / github.com / wengong-jin / icml18-jtnn.

[0154] Details of the disclosed model (CSLVAE). During training, an annealing schedule for β

[49] was utilized (see Algorithm 1), starting with β = 0 and incrementing by 1e-5 every 2000 iterations, with a maximum value of β being 1. Training received a total of 200K iterations, during which CSLVAE saw a total of 2000 * 200K = 400M compounds (not 400M unique compounds due to the batch sampling strategy). Thus, by the time training stopped, CSLVAE had seen less than 2.5% of the complete REAL library.

[0155] Example 2 - Latent Space Visualization. The latent space learned by the disclosed system and method was qualitatively inspected. Of interest was to verify whether the proposed model learned a latent space that changed relatively smoothly across the covered chemical space (e.g., a small perturbation to a query induced only minor edits in the resulting molecular graph). Two types of checks were performed: latent space interpolation and local neighborhood visualization.

[0156] Figure 4A includes an example of interpolation in the latent space. Specifically, molecular queries for the starting compound (top left) and the target compound (bottom right) were interpolated in raster order. Molecules directly adjacent to the starting and target compounds are the relevant reconstructions. The products were decoded with respect to the complete REAL library of 16 billion compounds. Below each molecule in Figure 4A is its Tanimoto similarity (5) to the starting and target molecules, respectively. It was observed that the interpolation traversed regions of chemical space where the similarity decreased (increasing reference) gradually with respect to the starting (target reference) compound.

[0157] Figure 4B visualizes the latent space around products randomly sampled from the REAL library. Following prior work

[29] , two random directions around the molecular query (central compound) were sampled and a random 2D plane was formed in the high-dimensional latent space by decoding the products obtained using the argmax decision rule. When the movement in the latent space is small (e.g., one synthon at a time, modification to smaller functional groups), it is observed that the latent space is smooth in the sense that the molecule gradually deforms with only minor edits, and the molecular scaffold is generally locally conserved.

[0158] Example 3 - Retrieving analogs by autoencoding. The disclosed systems and methods were used to find analogs of query compounds in large-scale CSLs. In FIGS. 5A and 5B, the disclosed model of the present disclosure is presented using two classes of molecules, namely those in the library (FIG. 5A) and those not in the library (FIG. 5B). Each query molecule in FIG. 5A (upper left) and each query molecule in FIG. 5B (upper left) are encoded 15 times and probabilistically decoded using the disclosed model (e.g., shown in FIGS. 1A, 1B, and 3). Thus, given a molecular query, a conditional random completion is generated from the decoder. In both cases, autoencoding searched for highly similar compounds as measured by Tanimoto similarity. In the example within the library, the model was able to successfully search for the query compound (third row, second column in FIG. 5A). Additionally, many examples are observed where autoencoding returns compounds that have the same reaction type as the query molecule but differ by one or two synthons. For the example outside the library, the disclosed systems and methods found skeletally related compounds with high Tanimoto similarity, demonstrating the utility of this approach for rapid analog search in large (enumeration-impossible) libraries.

[0159] Example 4 - Posterior density estimation As a way to force random samples from the CSLVAE, the use of posterior density estimation for the latent codes corresponding to the products uniformly sampled from the library is disclosed above to more tightly track uniformly randomly sampling from the library. Algorithm 2, seen in Figure 9, shows an overview of this procedure.

[0160] Experiments were conducted to demonstrate that Algorithm 2 achieves the intended results. The experiments proceed as follows. 10,000 molecules are randomly and uniformly drawn from the library and processed as a training set for the posterior density estimator. Three density estimators of increasing expressiveness are considered: Multivariate Normal (MVNormal), Mixture of 5 Normals (MoG-5), and Mixture of 10 Normals (MoG-10). A standard approach of sampling from an isotropic multivariate normal prior distribution was compared. Thus, in this experiment, four alternative schemes are used to sample the latent codes that will later be decoded into compounds from the library. Then, another 10,000 molecules are randomly and uniformly sampled from the library again and processed as a reference set for comparison. [Table 4]

[0161] According to

[41] , the aforementioned set of generated molecules is compared to a reference set of molecules with respect to the following computable molecular properties, namely synthetic accessibility (SA), quantitative estimate of drug-likeness (QED), molecular weight (MW), and logarithm of the octanol-water partition coefficient (logP): In particular, the Jensen-Shannon distance (square root of the Jensen-Shannon divergence) is calculated between the distribution of these properties on the reference set and each set in question. Table 4 summarizes the results of this exercise. Three items worthy of note are: (a) the compounds sampled to train the reference compound and the posterior density estimator have low divergence across various properties, (b) pre-sampling generates compounds with high divergence with respect to uniform sampling across the library, and (c) using a more expressive density estimator with respect to the latent code can better match the distribution of the latent code with respect to the training set (which is interchangeable with the reference set), resulting in increasingly lower divergence across various properties with respect to the reference set.

[0162] Example 5 - Comparison with Existing Analogue Enumeration Approaches In Example 5, an attempt was made to compare the analog generation ability of CSLVAE with that of Arthor, a state-of-the-art commercial similarity search tool for synthetic libraries developed by NextMove Software. Arthor performs analog enumeration using a custom ECFP4 bitvector representation of molecules and returns compounds with high Tanimoto similarity according to this fingerprint. For a given query compound, the top 100 analogs returned by Arthor from REAL were enumerated. Similarly, each query compound was encoded with the CSLVAE encoder to generate the corresponding 100 probabilistic decodings. For all analogs, their RDKit ECFP4 Tanimoto similarity was calculated using the query compound, and the top 1 analog was retained. The distribution of the Tanimoto similarity of the top 1 analogs thus returned was compared between Arthor and CSLVAE. As a control, a simple random baseline policy was used in which 100 compounds were randomly sampled from REAL and the top 1 analog was selected again based on the RDKit ECFP4 Tanimoto similarity. Since both CSLVAE and the random baseline constitute probabilistic policies, this procedure was repeated 30 times for each query compound, and the average top 1 Tanimoto similarity was adopted.

[0163] For the query compounds, 24 out of 51 new drugs approved by the FDA in 2021 were used, excluding drugs that do not meet the conventional set of small molecule criteria (e.g., monoclonal antibodies), and the compounds used in this experiment are shown in Figure 10. Figure 11 shows box plots of the top 1 Tanimoto similarity for Arthor, CSLVAE, and the random baseline.

[0164] CSLVAE finds more distant ECFP4 analogs compared to Arthor (the gold standard for fingerprint similarity searches), yet it can identify analogs for unknown novel drugs in a conventional manner. In this regard, it is typical to use a Tanimoto similarity threshold in the range of 0.3 to 0.35 to indicate whether a pair of molecules can be considered analogs.

[0165] Of particular note, CSLVAE is not nearly as resource-intensive as Arthor, which is specifically designed for fast and efficient fingerprint-based analog enumeration and requires a proper infrastructure and setup. For these reasons, it is difficult to perform a direct apples-to-apples comparison of the computational requirements. Nevertheless, the following relevant information was obtained from this experiment. The CSLVAE experiment was run on a machine equipped with an NVIDIA Tesla K80 GPU and an Intel Xeon E5-2686 CPU. On average, sampling 100 analogs for a given query using CSLVAE was completed in 11.41 seconds in this setup (encoding and decoding). Furthermore, all parameters and buffers of the trained CSLVAE model (including representations and keys for synthons, R groups, and reactions) required only 170 MB of memory. In comparison, the Arthor experiment was run in a distributed manner on 100 pods each having 8 CPUs, and enumerating the top 100 analogs required a total of 32.46 seconds, with an approximate total CPU time of 25,968 seconds. Additionally, Arthor needs to invest a significant amount of memory and storage to perform analog enumeration, and the inventors' in-house setup uses 3 GB per shard and makes a 64 GB RAM requirement for each Arthor worker.

[0166] The ability of CSLVAE to represent large CSLs with significantly fewer resources and perform analog retrieval with a significantly improved execution time (perhaps three orders of magnitude faster) makes use of parallel synthon lookup, thereby requiring a few keys on the order of the number of synthons in the library rather than the number of products in the library, and furthermore enabling a kind of similarity search in time that is (not linear but) logarithmic in the number of products in the library, due to its decoding strategy.

[0167] Example 6 - Encoder Transfer to Molecular Property Prediction The CSLVAE training objective can be regarded as a kind of contrastive pretext task that attempts to align the representation of a given molecule with the representation corresponding to the retrieval instructions within the CSL. Therefore, it is natural to consider whether the encoder of the trained CSLVAE model can demonstrate good transfer performance in a target prediction task. To investigate, an MLP trained with CSLVAE was compared with those trained with molecular fingerprints (ECFP4 and ECFP6) for the molecular property prediction task. The octanol - water partition coefficient (logP) and the quantitative estimate of drug - likeness (QED) [6] were used as the targets for prediction.

[0168] For this experiment, a dataset was constructed by randomly and uniformly sampling 100K compounds from REAL and splitting the examples into training, validation, and test folds using an 80-10-10 split. For each compound, in addition to its ECFP4 and ECFP6 fingerprints, its CSLVAE query was extracted as its descriptor. The training fold was used to separately fit an MLP to each of these descriptors to predict the logP and QED scores of the molecules. The iteration achieving the lowest validation RMSE was selected and its test RMSE was recorded. This was repeated 5 times and the average test RMSE and standard deviation were reported. To demonstrate the extent to which CSLVAE learns molecular features to predict such molecular properties of out-of-domain compounds, this exercise was repeated on a dataset of 250K molecules from ZINC. The results of this exercise are summarized in Table 5. [Table 5]

[0169] The results of this experiment confirm that the latent space learned by CSLVAE can be successfully used to predict quantities such as logP and QED, and in particular, in the in-domain case, it performs better than predictors that fit chemical fingerprints. However, in the out-of-domain case, the predictor fit for the ECFP6 fingerprint functions significantly better than the predictor fit for the CSLVAE query, suggesting that the features learned by CSLVAE may lack some relevant predictive information for input molecules that are significantly different from the CSL for which it was trained.

[0170] Overview of the CSLVAE architecture used in Example 7 - Example. Table 6 summarizes the number of parameters used in an exemplary CSLVAE architecture for the examples of the present disclosure. [Table 6]

[0171] References. [1]Atanu Acharya, Rupesh Agarwal, Matthew B Baker, Jerome Baudry, Debsindhu Bhowmik, Swen Boehm, Kendall G Byler, SY Chen, Leighton Coates, Connor J Cooper, et al. Supercomputer-based ensemble docking drug discovery pipeline with application to COVID-19. Journal of Chemical Information and Modeling, 60(12):5832-5852, 2020. [2]Josep Arus-Pous, Thomas Blaschke, Silas Ulander, Jean-Louis Reymond, Hongming Chen, and Ola Engkvist. Exploring the GDB-13 chemical space using deep generative models. Journal of Cheminformatics, 11(1):1-14, 2019. [3]David Bajusz, Anita Racz, and Karoly Heberger. Why is Tanimoto index an appropriate choice for fingerprint-based similarity calculations? Journal of Cheminformatics, 7(1):1-13, 2015. [4]Louis Bellmann, Patrick Penner, and Matthias Rarey. Topological similarity search in large combinatorial fragment spaces. Journal of Chemical Information and Modeling, 61(1):238-251, 2020. [5]Andreas Bender and Robert C Glen.Molecular similarity:a key technique in molecular informatics.Organic and Biomolecular Chemistry,2(22):3204-3218,2004. [6]G Richard Bickerton,Gaia V Paolini,Jeremy Besnard,Sorel Muresan,and Andrew L Hopkins.Quantifying the chemical beauty of drugs.Nature Chemistry,4(2):90-98,2012. [7]Lorenz C Blum and Jean-Louis Reymond.970 million druglike small molecules for virtual screening in the chemical universe database GDB-13.Journal of the American Chemical Society,131(25):8732-8733,2009. [8]John Bradshaw,Brooks Paige,Matt J Kusner,Marwin Segler,and Jose Miguel Hernandez- Lobato.A model to search for synthesizable molecules.Advances in Neural Information Processing Systems,32,2019. [9]John Bradshaw,Brooks Paige,Matt J Kusner,Marwin Segler,and Jose Miguel Hernandez-Lobato.Barking up the right tree:an approach to search over molecule synthesis DAGs.Advances in Neural Information Processing Systems,33:6852-6866,2020.

[10] Hongming Chen.Can generative-model-based drug design become a new normal in drug discovery? Journal of Medicinal Chemistry,65(1):100-102,2021.

[11] Krzysztof Choromanski,Valerii Likhosherstov,David Dohan,Xingyou Song,Andreea Gane,Tamas Sarlos,Peter Hawkins,Jared Davis,Afroz Mohiuddin,Lukasz Kaiser,et al.Rethinking attention with performers.arXiv preprint arXiv:2009.14794,2020.

[12] Hanjun Dai,Yingtao Tian,Bo Dai,Steven Skiena,and Le Song.Syntax-directed variational autoencoder for structured data.arXiv preprint arXiv:1802.08786,2018.

[13] Wenhao Gao and Connor W Coley.The synthesizability of molecules proposed by generative models.Journal of Chemical Information and Modeling,60(12):5714-5723,2020.

[14] Partha Ghosh,Mehdi SM Sajjadi,Antonio Vergari,Michael Black,and Bernhard Schoelkopf.From variational to deterministic autoencoders.arXiv preprint arXiv:1903.12436,2019.

[15] Pawe? Gniewek,Bradley Worley,Kate Stafford,Henry van den Bedem,and Brandon Anderson.Learning physics confers pose-sensitivity in structure-based virtual screening.arXiv preprint arXiv:2110.15459,2021.

[16] Rafael Gomez-Bombarelli,Jennifer N Wei,David Duvenaud,Jose Miguel Hernandez-Lobato,Benjamin Sanchez-Lengeling,Dennis Sheberla,Jorge Aguilera-Iparraguirre,Timothy D Hirzel,Ryan P Adams,and Alan Aspuru-Guzik.Automatic chemical design using a data-driven continuous representation of molecules.ACS Central Science,4(2):268-276,2018.

[17] Christoph Gorgulla,Andras Boeszoermenyi,Zi-Fu Wang,Patrick D Fischer,Paul W Coote,Krishna M Padmanabha Das,Yehor S Malets,Dmytro S Radchenko,Yurii S Moroz,David A Scott,et al.An open-source drug discovery platform enables ultra-large virtual screens.Nature,580(7805),2020.

[18] David E Graff,Eugene I Shakhnovich,and Connor W Coley.Accelerating high-throughput virtual screening through molecular pool-based active learning.Chemical Science,12(22):7866-7881,2021.

[19] Ryan-Rhys Griffiths and Jose Miguel Hernandez-Lobato.Constrained bayesian optimization for automatic chemical design using variational autoencoders.Chemical Science,11(2):577-586,2020.

[20] Oleksandr O Grygorenko,Dmytro S Radchenko,Igor Dziuba,Alexander Chuprina,Kateryna E Gubina,and Yurii S Moroz.Generating multibillion chemical space of readily accessible screening compounds.Iscience,23(11),2020.

[21] Julien Horwood and Emmanuel Noutahi.Molecular design in synthetically accessible chemical space via deep reinforcement learning.ACS Omega,5(51):32984-32994,2020.

[22] John J Irwin and Brian K Shoichet.Docking screens for novel ligands conferring new biology:Miniperspective.Journal of Medicinal Chemistry,59(9):4103-4120,2016.

[23] Wengong Jin,Regina Barzilay,and Tommi Jaakkola.Junction tree variational autoencoder for molecular graph generation.International Conference on Machine Learning,2018.

[24] Wengong Jin,Regina Barzilay,and Tommi Jaakkola.Hierarchical generation of molecular graphs using structural motifs.International Conference on Machine Learning,pages 4839-4848,2020.

[25] Wengong Jin,Regina Barzilay,and Tommi Jaakkola.Multi-objective molecule generation using interpretable substructures.International Conference on Machine Learning,2020.

[26] Nikita Kitaev, Łukasz Kaiser,and Anselm Levskaya.Reformer:the efficient transformer.arXiv preprint arXiv:2001.04451,2020.

[27] Xiangzhe Kong,Zhixing Tan,and Yang Liu.Graphpiece:Efficiently generating high-quality molecular graph with substructures.arXiv preprint arXiv:2106.15098,2021.

[28] Mario Krenn,Florian Haese,AkshatKumar Nigam,Pascal Friederich,and Alan Aspuru-Guzik.Self-referencing embedded strings (selfies):A 100% robust molecular string representation.Machine Learning:Science and Technology,1(4),2020.

[29] Matt J Kusner,Brooks Paige,and Jose Miguel Hernandez-Lobato.Grammar variational autoencoder.International Conference on Machine Learning,2017.

[30] Greg Landrum.RDKit:Open-source cheminformatics,2006.

[31] Juho Lee,Yoonho Lee,Jungtaek Kim,Adam Kosiorek,Seungjin Choi,and Yee Whye Teh.Set transformer:A framework for attention-based permutation-invariant neural networks.International Conference on Machine Learning,pages 3744-3753,2019.

[32] Qi Liu,Miltiadis Allamanis,Marc Brockschmidt,and Alexander Gaunt.Constrained graph variational autoencoders for molecule design.Advances in Neural Information Processing Systems,31,2018.

[33] Jiankun Lyu,Sheng Wang,Trent E Balius,Isha Singh,Anat Levit,Yurii S Moroz,Matthew J O’Meara,Tao Che,Enkhjargal Algaa,Kateryna Tolmachova,et al.Ultra-large library docking for discovering new chemotypes.Nature,566(7743):224-229,2019.

[34] Krzysztof Maziarz,Henry Jackson-Flux,Pashmina Cameron,Finton Sirockin,Nadine Schneider,Nikolaus Stiefl,Marwin Segler,and Marc Brockschmidt.Learning to extend molecular scaffolds with structural motifs.arXiv preprint arXiv:2103.03864,2021.

[35] Rocio Mercado,Tobias Rastemo,Edvard Lindeloef,Guenter Klambauer,Ola Engkvist,Hongming Chen,and Esben Jannik Bjerrum.Graph networks for molecular design.Machine Learning:Science and Technology,2(2):025023,2021.

[36] Joshua Meyers,Benedek Fabian,and Nathan Brown.De novo molecular design and generative models.Drug Discovery Today,26(11):2707-2715,2021.

[37] AkshatKumar Nigam,Robert Pollice,Mario Krenn,Gabriel dos Passos Gomes,and Alan Aspuru-Guzik.Beyond generative models:superfast traversal,optimization,novelty,exploration and discovery (stoned) algorithm for molecules using selfies.Chemical Science,12(20):7079-7090,2021.

[38] Hakime Oeztuerk,Arzucan Oezguer,and Elif Ozkirimli.DeepDTA:deep drug-target binding affinity prediction.Bioinformatics,34(17):i821-i829,2018.

[39] Joseph M Paggi,Julia A Belk,Scott A Hollingsworth,Nicolas Villanueva,Alexander S Powers,Mary J Clark,Augustine G Chemparathy,Jonathan E Tynan,Thomas K Lau,Roger K Sunahara,et al.Leveraging nonstructural data to predict structures and affinities of protein-ligand complexes.Proceedings of the National Academy of Sciences,118(51),2021.

[40] Pavel Polishchuk.Control of synthetic feasibility of compounds generated with CReM.Journal of Chemical Information and Modeling,60(12):6074-6080,2020.

[41] Daniil Polykovskiy,Alexander Zhebrak,Benjamin Sanchez-Lengeling,Sergey Golovanov,Oktai Tatanov,Stanislav Belyaev,Rauf Kurbanov,Aleksey Artamonov,Vladimir Aladinskiy,Mark Veselov,et al.Molecular sets (MOSES):a benchmarking platform for molecular generation models.Frontiers in Pharmacology,11:1931,2020.

[42] Mariya Popova,Mykhailo Shvets,Junier Oliva,and Olexandr Isayev.Molecularrnn:Generating realistic molecular graphs with optimized properties.arXiv preprint arXiv:1905.13372,2019.

[43] Matthew Ragoza,Joshua Hochuli,Elisa Idrobo,Jocelyn Sunseri,and David Ryan Koes.Protein- ligand scoring with convolutional neural networks.Journal of Chemical Information and Modeling,57(4):942-957,2017.

[44] Arman A Sadybekov,Anastasiia V Sadybekov,Yongfeng Liu,Christos Iliopoulos-Tsoutsouvas,Xi-Ping Huang,Julie Pickett,Blake Houser,Nilkanth Patel,Ngan K Tran,Fei Tong,et al.Synthon-based ligand discovery in virtual libraries of over 11 billion compounds.Nature,601(7893):452-459,2022.

[45] Robert Schmidt,Raphael Klein,and Matthias Rarey.Maximum common substructure searching in combinatorial make-on-demand compound spaces.Journal of Chemical Information and Modeling,2021.

[46] Marwin HS Segler,Thierry Kogej,Christian Tyrchan,and Mark P Waller.Generating focused molecule libraries for drug discovery with recurrent neural networks.ACS Central Science,4(1):120-131,2018.

[47] Brian K Shoichet.Virtual screening of chemical libraries.Nature,432(7019):862-865,2004.

[48] Martin Simonovsky and Nikos Komodakis.Graphvae:Towards generation of small graphs using variational autoencoders.International Conference on Artificial Neural Networks,2018.

[49] Casper Kaae Sonderby,Tapani Raiko,Lars Maaloe,Soren Kaae Sonderby,and Ole Winther.Ladder variational autoencoders.Advances in Neural Information Processing Systems,29,2016.

[50] Tiago Sousa,Joao Correia,Vitor Pereira,and Miguel Rocha.Generative deep learning for targeted compound design.Journal of Chemical Information and Modeling,61(11):5343-5361,2021.

[51] Kate Stafford, Brandon M Anderson, Jon Sorenson, and Henry van den Bedem. AtomNet PoseRanker: Enriching ligand pose quality for dynamic proteins in virtual high-throughput screens. Journal of Chemical Information and Modeling, 62(5):1178-1189, 2022.

[52] Hannes Staerk, Octavian-Eugen Ganea, Lagnajit Pattanaik, Regina Barzilay, and Tommi Jaakkola. Equibind: Geometric deep learning for drug binding structure prediction. arXiv preprint arXiv:2202.05146, 2022.

[53] Kosuke Takeuchi, Ryo Kunimoto, and Juergen Bajorath. R-group replacement database for medicinal chemistry. Future Science OA, 7(8):FSO742, 2021.

[54] Xiaochu Tong, Xiaohong Liu, Xiaoqin Tan, Xutong Li, Jiaxin Jiang, Zhaoping Xiong, Tingyang Xu, Hualiang Jiang, Nan Qiao, and Mingyue Zheng. Generative models for de novo drug design. Journal of Medicinal Chemistry, 64(19):14011-14027, 2021.

[55] Kenneth E Train. Discrete choice methods with simulation. Cambridge University Press, 2009.

[56] Masashi Tsubaki, Kentaro Tomii, and Jun Sese. Compound-protein interaction prediction with end-to-end learning of neural networks for graphs and sequences. Bioinformatics, 35(2):309 - 318, 2019.

[57] Izhar Wallach, Michael Dzamba, and Abraham Heifets. AtomNet: a deep convolutional neural network for bioactivity prediction in structure-based drug discovery. arXiv preprint arXiv:1510.02855, 2015.

[58] Wendy A Warr, Marc C Nicklaus, Christos A Nicolaou, and Matthias Rarey. Exploration of ultra-large compound collections for drug discovery. Journal of Chemical Information and Modeling, 62(9):2021 - 2034, 2022.

[59] Robin Winter, Frank Noe, and Djork-Arne Clevert. Permutation-invariant variational autoencoder for graph-level representation learning. Advances in Neural Information Processing Systems, 34, 2021.

[60] Jiaxuan You, Bowen Liu, Zhitao Ying, Vijay Pande, and Jure Leskovec. Graph convolutional policy network for goal-directed molecular graph generation. Advances in Neural Information Processing Systems, 31, 2018.

[61] Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Russ R Salakhutdinov, and Alexander J Smola. Deep Sets. Advances in Neural Information Processing Systems, 30, 2017.

[0172] Conclusion For purposes of explanation, the foregoing description has been presented with reference to specific implementations. However, the above exemplary considerations are not intended to be exhaustive or to limit implementations to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The implementations were chosen and described in order to best explain the principles and their practical application, thereby enabling others skilled in the art to best utilize the implementations and various implementations with various modifications as are suited to the particular use contemplated.

Claims

1. A computer system for querying a combinatorial synthesis library comprising a plurality of compounds, wherein the combinatorial synthesis library represents a plurality of reaction types, each of the plurality of reaction types has a corresponding mapping to a corresponding plurality of reactants, each of the reactants of each corresponding plurality of reactants has a corresponding mapping to a corresponding plurality of synthons, and the computer system comprises one or more central processing units, one or more graphics processing units, wherein each of the one or more graphics processing units comprises 100 or more cores, a memory addressable by the one or more central processing units, the memory storing at least one program for being at least partially implemented by the one or more graphics processing units, (A) inputting a query, which is a single graph, into a molecular encoder model, the molecular encoder model comprising a message passing neural network comprising a plurality of message passing layers collectively comprising a first plurality of parameters, thereby obtaining a query vector by applying the first plurality of parameters to the single graph, (B) inputting the query vector into a reaction query generator model comprising a second plurality of parameters, thereby obtaining, as an output from the reaction query generator model, a first reaction type among the plurality of reaction types by applying the second plurality of parameters to the query vector, (C) For each of the respective reactants of the first plurality of reactants corresponding to the first reaction type, input the respective reactant into a synthon query generator model including a third plurality of parameters, thereby applying the third plurality of parameters to the respective reactant, and obtaining, as an output from the synthon query generator model, the corresponding synthon, and determining from among the corresponding plurality of synthons mapped to the respective reactant, thereby determining a set of synthons, each of the synthons in the set of synthons corresponding to a reactant among the first plurality of reactants. (D) A computer system including instructions for identifying a molecular structure in the combinatorial synthesis library including the set of synthons arranged according to a synthesis rule associated with the first reaction type. The computer system according to claim 1, wherein the reaction query generator model is a two-layer perceptron having intermediate ReLU activation. The computer system according to claim 1, wherein the synthon query generator model is a two-layer perceptron having intermediate ReLU activation. The first plurality of parameters includes 100,000 parameters. The second plurality of parameters includes 5,000 parameters. The computer system according to claim 1, wherein the third plurality of parameters includes 5,000 parameters. The single graph includes a plurality of nodes and a plurality of edges. The computer system according to claim 1, wherein each of the plurality of nodes is connected to another node among the plurality of nodes by at least one of the plurality of edges. Each of the plurality of nodes (i) a corresponding element type among a plurality of element types, (ii) a node degree among a plurality of node degrees, (iii) a hybridization among a plurality of hybridizations, (iv) the number of bound hydrogens, (v) a formal charge from among a set of formal charges, and (vi) associated with a binary representation of aromaticity. The computer system according to claim 5. Each of the plurality of bonds (i) a bond type ​ ​ ​ ​ ​ ​ (ii) conjugate binary representation, (iii) binary of the indication of whether each of the respective bonds is within the ring, and (iv) the computer system according to claim 5 or 6, associated with the indication of stereochemistry.

8. The computer system according to any one of claims 1 to 7, wherein the plurality of reaction types includes 20 or more reaction types, and the combinatorial synthesis library includes 100 or more compounds for each reaction type among the plurality of reaction types.

9. The computer system according to any one of claims 1 to 8, wherein the first plurality of reactants includes 3 or more reactants, and the corresponding mapping for the corresponding plurality of synthons for the reactants among the 3 or more reactants includes 10 or more synthons.

10. The output from the reaction query generator model is used to identify a first reaction key among a plurality of reaction keys through a first query key lookup, The computer system according to any one of claims 1 to 9, wherein each reaction key among the plurality of reaction keys represents a synthetic reaction that can be used to synthesize one or more compounds in the combinatorial synthesis library.

11. The computer system according to any one of claims 1 to 10, wherein the output from the synthon query generator model is used to identify a synthon key for the corresponding synthon through a second query key lookup.

12. The computer system according to any one of claims 1 to 11, wherein the single graph represents a single molecular compound present in the combinatorial synthesis library.

13. The computer system according to any one of claims 1 to 11, wherein the single graph represents a weighted complex of a first graph of a first molecular compound and a second graph of a second molecular compound.

14. The single graph represents a weighted complex of a plurality of graphs of a second plurality of compounds, The computer system according to any one of claims 1 to 11, wherein the second plurality of compounds have a common property.

15. The computer system according to claim 14, wherein the common property is a Tanimoto distance less than a threshold value for each of the other compounds among the second plurality of compounds.

16. The computer system according to claim 14, wherein the common characteristic is a binding coefficient to a macromolecular target that is less than a threshold value.

17. The computer system according to any one of claims 1 to 16, wherein the plurality of compounds includes one billion or more compounds, and the molecular structure output by the identifying (D) is any one of the one billion or more compounds that satisfy the query.

18. The computer system according to any one of claims 1 to 16, wherein the plurality of compounds includes one trillion or more compounds, and the molecular structure output by the identifying (D) is any one of the one trillion or more compounds that satisfy the query.

19. The computer system according to any one of claims 1 to 18, wherein the single graph represents a query molecular compound as a set of atomic features and a set of bonding features.

20. The set of atomic features includes element type, node degree, hybridization, chirality, bound hydrogen, formal charge, aromaticity, The computer system according to claim 19, wherein the set of bonding features includes bond time, conjugation, in-ring, and stereochemistry.

21. Each non-hydrogen atom in the query molecular compound is represented by 2000 or more parameters in the set of atomic features, and each covalent bond in the molecular compound is represented by 500 or more parameters in the set of bonding features. The computer system according to claim 20.

22. A method for querying a combinatorial synthesis library containing a plurality of compounds, wherein the combinatorial synthesis library represents a plurality of reaction types, each of the plurality of reaction types has a corresponding mapping to a corresponding plurality of reactants, each of the reactants in each corresponding plurality of reactants has a corresponding mapping to a corresponding plurality of synthons, the method is a computer system, one or more central processing units, one or more graphics processing units, wherein each of the one or more graphics processing units comprises 100 or more cores, one or more graphics processing units, A memory addressable by the one or more central processing units, the memory storing at least one program for being at least partially implemented by the one or more graphics processing units, in a computer system comprising: (A) a query, the query being a single graph, inputting the query into a molecular encoder model, the molecular encoder model including a message passing neural network including a plurality of message passing layers collectively including a first plurality of parameters, thereby obtaining a query vector by applying the first plurality of parameters to the single graph; (B) inputting the query vector into a reaction query generator model including a second plurality of parameters, thereby obtaining, as an output from the reaction query generator model, a first reaction type among the plurality of reaction types by applying the second plurality of parameters to the query vector; (C) for each respective reactant among the first plurality of reactants, inputting the respective reactant into a synthon query generator model including a third plurality of parameters, thereby obtaining the corresponding synthon as an output from the synthon query generator model by applying the third plurality of parameters to the respective reactant, and determining from among the corresponding plurality of synthons mapped to the respective reactants, thereby determining a set of synthons, each synthon in the set of synthons corresponding to a reactant among the first plurality of reactants; (D) identifying a molecular structure in the combinatorial synthesis library including the set of synthons arranged according to a synthetic rule associated with the first reaction type. A method comprising instructions for implementing the method.

23. A computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by a computer system comprising one or more central processing units, one or more graphics processing units, each of the one or more graphics processing units having 100 or more cores, and a memory, cause the computer system to query a combinatorial synthesis library, the combinatorial synthesis library representing a plurality of reaction types, each of the plurality of reaction types having a corresponding mapping to a corresponding plurality of reactants, each of the reactants of each of the corresponding plurality of reactants having a corresponding mapping to a corresponding plurality of synthons, the query of the combinatorial synthesis library being performed by the one or more graphics processing units, (A) inputting a query, which is a single graph, into a molecular encoder model, the molecular encoder model including a message-passing neural network including a plurality of message-passing layers collectively including a first plurality of parameters, thereby obtaining a query vector by applying the first plurality of parameters to the single graph; (B) inputting the query vector into a reaction query generator model including a second plurality of parameters, thereby obtaining, as an output from the reaction query generator model, a first reaction type among the plurality of reaction types by applying the second plurality of parameters to the query vector; (C) For each respective reactant among the first plurality of reactants, input the respective reactant into a synthon query generator model that includes a third plurality of parameters, thereby applying the third plurality of parameters to the respective reactant, and as an output from the synthon query generator model, obtain the corresponding synthon by obtaining the corresponding synthon from among the corresponding plurality of synthons mapped to the respective reactant, thereby determining a set of synthons, wherein each synthon in the set of synthons corresponds to a reactant among the first plurality of reactants; and (D) identifying a molecular structure in the combinatorial synthesis library that includes the set of synthons arranged according to synthesis rules associated with the first reaction type; A computer-readable storage medium that is at least partially implemented by a method including.

24. A computer system for querying a combinatorial synthesis library containing a plurality of compounds, The combinatorial synthesis library represents a plurality of reaction types, Each respective reaction type among the plurality of reaction types has a corresponding mapping to a corresponding plurality of reactants, Each respective reactant among each corresponding plurality of reactants has a corresponding mapping to a corresponding plurality of synthons, and the computer system includes: one or more processing units; a memory addressable by the one or more processing units, the memory storing at least one program for execution by the one or more processing units, the at least one program including: (A) A query, wherein the query is an arbitrary graph, input the query into a molecular encoder model, wherein the molecular encoder model includes a first plurality of parameters, thereby applying the first plurality of parameters to the arbitrary graph to obtain a query vector, (B) Input the query vector into a reaction query generator model that includes a second plurality of parameters, and thereby apply the second plurality of parameters to the query vector to obtain, as an output from the reaction query generator model, a first reaction type among the plurality of reaction types. (C) For each respective reactant among the first plurality of reactants corresponding to the first reaction type, input the respective reactant into a synthon query generator model that includes a third plurality of parameters, and thereby apply the third plurality of parameters to the respective reactant to obtain, as an output from the synthon query generator model, the corresponding synthon, and determine from among the plurality of corresponding synthons mapped to the respective reactants, thereby determining a set of synthons, wherein each synthon in the set of synthons corresponds to a reactant among the first plurality of reactants. (D) A computer system comprising instructions for identifying molecular structures in the combinatorial synthesis library that include the set of synthons according to synthesis rules associated with the first reaction type.

25. The computer system according to claim 24, wherein the reaction query generator model is a two-layer perceptron having intermediate ReLU activation.

26. The computer system according to claim 24, wherein the synthon query generator model is a two-layer perceptron having intermediate ReLU activation.

27. The first plurality of parameters includes 100,000 parameters. The second plurality of parameters includes 5,000 parameters. The computer system according to claim 24, wherein the third plurality of parameters includes 5,000 parameters.

28. The arbitrary graph includes a plurality of nodes and a plurality of edges. The computer system according to claim 24, wherein each node among the plurality of nodes is connected to another node among the plurality of nodes by at least one edge among the plurality of edges.

29. Each node among the plurality of nodes (i) A corresponding element type among a plurality of element types. (ii) A node degree among a plurality of node degrees. (iii) hybridization among a plurality of hybridizations, (iv) number of hydrogen bonds, (v) formal charge from a set of formal charges, and (vi) the computer system according to claim 28, associated with a binary representation of aromaticity. (Claim 30) Each respective bond among the plurality of bonds is (i) bond type, (ii) binary representation of conjugation, (iii) binary of a representation of whether or not the respective bond is within a ring, and (iv) the computer system according to claim 28 or 29, associated with a representation of stereochemistry. (Claim 31) The plurality of reaction types includes 20 or more reaction types, and the combinatorial synthesis library includes 100 or more compounds for each reaction type among the plurality of reaction types. The computer system according to any one of claims 24 to 30. (Claim 32) The first plurality of reactants includes three or more reactants, and the corresponding mapping for the corresponding plurality of synthons for the reactants among the three or more reactants includes 10 or more synthons. The computer system according to any one of claims 24 to 31. (Claim 32) The output from the reaction query generator model is used to identify a first reaction key among a plurality of reaction keys through a first query key lookup, Each reaction key among the plurality of reaction keys represents a synthetic reaction that can be used to synthesize one or more compounds in the combinatorial synthesis library. The computer system according to any one of claims 24 to 32. (Claim 33) The output from the synthon query generator model is used to identify a synthon key for the corresponding synthon through a second query key lookup. The computer system according to any one of claims 24 to 32. (Claim 34) The arbitrary graph represents a single molecular compound present in the combinatorial synthesis library. The computer system according to any one of claims 24 to 33. (Claim 35) The arbitrary graph represents a weighted complex of a first graph of a first molecular compound and a second graph of a second molecular compound. The computer system according to any one of claims 24 to 33. (Claim 36) The arbitrary graph represents a weighted complex of a plurality of graphs of a second plurality of compounds, The computer system according to any one of claims 24 to 33, wherein the plurality of arbitrary compounds have common characteristics.

37. The computer system according to claim 36, wherein the common characteristic is a Tanimoto distance less than a threshold value with respect to each of the other compounds among the second plurality of compounds.

38. The computer system according to claim 36, wherein the common characteristic is a binding coefficient with respect to a macromolecular target that is less than a threshold value.

39. The computer system according to any one of claims 24 to 38, wherein the plurality of compounds include 1 billion or more compounds, and the molecular structure output by the identifying (D) is any one of the 1 billion or more compounds that satisfy the query.

40. The computer system according to any one of claims 24 to 38, wherein the plurality of compounds include 1 trillion or more compounds, and the molecular structure output by the identifying (D) is any one of the 1 trillion or more compounds that satisfy the query.

41. The computer system according to any one of claims 24 to 40, wherein the arbitrary graph represents a query molecular compound as a set of atomic features and a set of bonding features.

42. The set of atomic features includes element type, node degree, hybridization, chirality, bound hydrogen, formal charge, aromaticity, The computer system according to claim 41, wherein the set of bonding features includes bond time, conjugation, intra-ring, and stereochemistry.

43. Each non-hydrogen atom in the query molecular compound is represented by 2000 or more parameters in the set of atomic features, and each covalent bond in the molecular compound is represented by 500 or more parameters in the set of bonding features. The computer system according to claim 42.

44. The computer system according to any one of claims 24 to 43, wherein the molecular encoder model includes a message passing neural network including a plurality of message passing layers that collectively include the first plurality of parameters.

45. A method for querying a combinatorial synthesis library containing a plurality of compounds, The combinatorial synthesis library represents a plurality of reaction types, Each of the plurality of reaction types has a corresponding mapping to a corresponding plurality of reactants, Each of the reactants of each corresponding plurality of reactants has a corresponding mapping to a corresponding plurality of synthons, The method is implemented in a computer system comprising one or more processing units, and a memory addressable by the one or more processing units, the memory storing at least one program for execution by the one or more processing units, the at least one program comprising (A) a query, the query being an arbitrary graph, inputting the query into a molecular encoder model comprising a first plurality of parameters, thereby obtaining a query vector by applying the first plurality of parameters to the arbitrary graph; (B) inputting the query vector into a reaction query generator model comprising a second plurality of parameters, thereby obtaining, as an output from the reaction query generator model, a first reaction type of the plurality of reaction types by applying the second plurality of parameters to the query vector; (C) for each reactant of the first plurality of reactants, inputting the respective reactant into a synthon query generator model comprising a third plurality of parameters, thereby obtaining the corresponding synthon as an output from the synthon query generator model by applying the third plurality of parameters to the respective reactant, determining from among the corresponding plurality of synthons mapped to the respective reactant, thereby determining a set of synthons, each synthon of the set of synthons corresponding to a reactant of the first plurality of reactants; (D) identifying a molecular structure in the combinatorial synthesis library comprising the set of synthons arranged according to a synthetic rule associated with the first reaction type. A method comprising instructions for performing a method. Claim 45 A computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions that, when executed by a computer system comprising one or more processing units and a memory, cause the computer system to perform a method of querying a combinatorial synthesis library, the combinatorial synthesis library represents a plurality of reaction types, each reaction type of the plurality of reaction types has a corresponding mapping to a corresponding plurality of reactants, each reactant of each corresponding plurality of reactants has a corresponding mapping to a corresponding plurality of synthons, the combinatorial synthesis library (A) Inputting a query, which is an arbitrary graph, into a molecular encoder model including a first plurality of parameters, thereby obtaining a query vector by applying the first plurality of parameters to the arbitrary graph; (B) Inputting the query vector into a reaction query generator model including a second plurality of parameters, thereby obtaining a first reaction type of the plurality of reaction types as an output from the reaction query generator model by applying the second plurality of parameters to the query vector; (C) For each reactant of the first plurality of reactants, inputting the respective reactant into a synthon query generator model including a third plurality of parameters, thereby obtaining the corresponding synthon as an output from the synthon query generator model by applying the third plurality of parameters to the respective reactant, and determining from among the corresponding plurality of synthons mapped to the respective reactant, thereby determining a set of synthons, each synthon in the set of synthons corresponding to a reactant of the first plurality of reactants; (D) Identifying a molecular structure in the combinatorial synthesis library that includes the set of synthons arranged according to a synthesis rule associated with the first reaction type. A computer-readable storage medium that is queried by a method including the above steps.