Order-aware string-based molecular representation for conditional molecule generation

O-SMILES addresses the issue of inconsistent traversal orders in SMILES by maintaining a consistent order between source and target molecules, enhancing the accuracy and efficiency of machine learning models in conditional molecule generation tasks.

WO2025148028A1PCT designated stage expired Publication Date: 2025-07-17MICROSOFT TECHNOLOGY LICENSING LLC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/072086
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-12
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Current molecular representations, such as SMILES, are not order-aware, leading to challenges in training machine learning models for conditional molecule generation due to differences in traversal order and syntax, which complicates the learning of chemical relationships between molecules.

Method used

An order-aware SMILES (O-SMILES) technique that maintains consistent traversal order between source and target molecules, using a root node and atom mapping to generate string-based representations, reducing reliance on decoder inference and minimizing edit distance.

Benefits of technology

Improves the accuracy and efficiency of machine learning models in tasks like retrosynthesis prediction and molecular property improvement by ensuring that differences in string-based molecular representations are primarily due to structural changes, rather than arbitrary traversal orders.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024072086_17072025_PF_FP_ABST
    Figure CN2024072086_17072025_PF_FP_ABST
Patent Text Reader

Abstract

This disclosure introduces a novel method and system for the representation of molecules to be used by computer systems for conditional molecule generation. These techniques leverage the traversal order of a source molecule, using it as a guide to determine a traversal order of a target molecule. This approach eliminates the need for a machine learning decoder to infer the atomic order, simplifying the process of decoder generation and reducing the risk of overfitting. This strategy eases the training of the decoder and enhances the generalization of the model. Experiments conducted on retrosynthesis prediction, forward synthesis prediction, and molecular property prediction tasks show that using this technique for generating string-based representations of molecules results in more accurate predictions.
Need to check novelty before this filing date? Find Prior Art

Description

ORDER-AWARE STRING-BASED MOLECULAR REPRESENTATION FOR CONDITIONAL MOLECULE GENERATIONBACKGROUNDEffectively representing molecules is a fundamental part of chemical research, with applications in areas such as drug discovery and material design. Two common types of molecular representations are molecular graphs and string-based sequences. Molecular graphs consist of atoms as nodes and chemical bonds as edges between the nodes. They are powerful tools for describing the topological structure of a molecule and are widely used for visualization. The most common string-based sequence is Simplified Molecular Input Line Entry System (SMILES) which is a way of representing molecules and reactions using a linear notation, facilitating their storage and input into computer systems.Conditional molecule generation, a significant application in computer-aided drug discovery, involves generating a target molecule based on a conditional input. This could include forward synthesis prediction, retrosynthesis prediction, and molecular property enhancement. The inherent vastness of the search space for all possible transformations makes conditional molecule generation a formidable challenge. Traditional methods rely on human-designed rules and chemical knowledge of experts. Computer-aided planning programs have also been available for many years. However, the increase in chemical reaction rules has made manually hard-coding chemical rules into computer systems impractical and inefficient. More recently, people have begun to explore fully data-driven approaches that rely on artificial intelligence and machine learning, eliminating the need for expert knowledge or manual rule coding.However, it can be challenging to represent molecules in a way that is readily interpretable by current machine learning models. Therefore, there is a need for a molecular representation that can be readily understood and efficiently processed by machine learning models. This disclosure is made with respect to these and other considerations.SUMMARYThis disclosure pertains to a technique for order-aware string-based molecular representation specifically designed for conditional molecule generation using machine learning. This disclosure provides a novel molecular representation method. The molecular representation method uses text strings to represent molecules and is based on SMILES; however, it may be used with other string-based representations besides SMILES. The molecular representation of this disclosure may be referred to as order-aware SMILES or “O-SMILES. ”The standard canonicalization algorithm used by SMILES to generate a string-based representation of a molecule is not order aware. Canonical SMILES generates string-based representations of each molecule independently without consideration of any relationship between molecules in a chemical reaction. Due to the rules used when generating string-based representations of molecules with SMILES, two molecules with very similar structures may be represented by dissimilar text strings. This presents a problem when using SMILES strings for training a machine learning model to perform conditional molecule generation. The complexity of the chemical rules defined in the canonical SMILES syntax influences the text string representing a molecule and this makes it challenging for a decoder to autoregressively learn the relationship between a source and target molecule used for training. This is because the machine learning model is attempting to learn the SMILES syntax rules as well as the molecular and chemical changes in order to understand the relationship between the source SMILES and target SMILES.In order to generate a text-string representation there is a traversal order of the molecule that specifies the order the atoms are represented in the text string. A change in traversal order can result in different text strings being generated for the same molecule. The O-SMILES technique is different from canonical SMILES, and other techniques for string-based molecular representations, because it attempts to maintain as much as possible the same traversal order between a source and a target molecule. String-based molecular representations of source and target molecules generated in this order-aware matter can be used to train a machine learning model for conditional molecule generation. O-SMILES utilizes the traversal order of a source molecule as a prior and follows it to traverse the target molecule, thus eliminating the need for a decoder to infer the predicted atomic order. Because the traversal order is consistent, differences in the text strings used for training come primarily from actual differences in chemical structure. Removing confounding differences introduced solely by the technique for generating string-based molecular representations can ease the decoder generation and potentially reduce overfitting.This technique maintains the same traversal order by using the same root node for both the source and the target molecule and assigning a unique order to the atoms in the source molecule. The string-based molecular representation of the target molecule is then determined by mapping the same atoms in the source molecule to the target molecule according to this order. Correspondence between atoms in the source and target molecules can be identified using an atom mapper. Use of the same root node and the same traversal order for generating the SMILES representations of a source molecule and a target molecule makes the two SMILES strings more similar than they would be if generated by a different technique. This similarity can be measured by the edit distance between the source molecule string and the target molecule string.String-based molecular representations of source molecules and target molecules created with this technique are then used to train a machine learning model for conditional molecule generation. Due to the similarity of the string-based molecular representations, the machine learning model can learn the chemical and molecular relationships without being confounded by differences introduced solely by how the SMILES strings are generated. Once trained, the machine learning model is used to perform conditional molecule generation tasks such as retrosynthesis prediction, forward reaction prediction, and molecular property improvement.Features and technical benefits other than those explicitly described above will be apparent from a reading of the following Detailed Description and a review of the associated drawings. This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. The term “techniques, ” for instance, may refer to system (s) , method (s) , computer-readable instructions, module (s) , algorithms, hardware logic, and / or operation (s) as permitted by the context described above and throughout the document.BRIEF DESCRIPTION OF THE DRAWINGSThe Detailed Description is described with reference to the accompanying figures. In the figures, the left-most digit (s) of a reference number identifies the figure in which the reference number first appears. The same reference numbers in different figures indicate similar or identical items. References made to individual items of a plurality of items can use a reference number with a letter of a sequence of letters to refer to each individual item. Generic references to the items may use the specific reference number without the sequence of letters.FIG. 1 is a diagram showing the use of a machine learning model to perform conditional molecule generation.FIG. 2 is a diagram comparing traversal orders of molecules used for generating string-based representations with canonical SMILES, R-SMILES, and O-SMILES.FIG. 3 is a diagram of a pipeline for generating string-based molecular representations for a source molecule and a target molecule using the techniques of this disclosure.FIG. 4 is a flowchart of a method for generating string-based molecular representations for a source molecule and a target molecule model that are used to train a machine learning model.FIG. 5 is a flowchart of a method for using a machine learning model to perform conditional molecule generation.FIG. 6 is a computer architecture diagram illustrating an illustrative computer hardware and software architecture for a computing device capable of implementing aspects of the techniques and technologies presented herein.DETAILED DESCRIPTIONThis disclosure provides a technique to generate string-based molecular representations of molecules that are well-suited for training machine learning models. The molecules may be part of a labeled training dataset such as a collection of chemical reactions that include reactants and products. The techniques of this disclosure generate the string-based molecular representations in an order aware manner so that structural similarities between a pair of molecules in the training data are reflected in the order of characters in the string-based molecular representations.A training dataset containing this type of string-based molecular representation can be used to train a machine learning model to perform conditional molecule generation. Conditional molecule generation is a method used in drug discovery and molecular design where a machine learning model learns a distribution of molecular structures with given properties. It involves generating one or more target molecules based on one or more source molecules and a conditional input. The conditional input could be specific properties and characteristics that are desired for the target molecule. Examples of conditional molecule generation include forward synthesis prediction, retrosynthesis prediction, and molecular property improvement.Forward synthesis prediction involves predictions of reaction outcomes (product) with a given set of substrates (reactants and reagents) . Retrosynthesis predicts possible reactants for a product molecule. The process involves transforming a product molecule into simpler precursor structures, regardless of any potential reactivity or interaction with reagents. Each precursor material is examined using the same method, and this procedure is repeated until simple or commercially available structures are identified. Molecular property improvement aims to improve a specific molecular property by modifying a molecule under certain similarity constraints. For example, a source molecule is a molecule with a poor value for a property and a target molecule is a molecule with a better value for that property. Given a source molecule, the task of molecular property improvement is to output a different, target molecule with a better molecular property that has a similarity to the source molecule above some similarity threshold.In conditional molecule generation, there is a pair of molecules < GS, GT >, where the source molecule GS and the target molecule GT are from the source domain S and the target domain T, respectively. For example, GS and GT are the product and reactants in a retrosynthesis prediction, and for molecular property improvement, GS and GT are the molecules with poor and good properties respectively. Conditional molecule generation aims to learn a mapping function f: S → T that can generate the target molecule GT conditioned on the input molecule GS. Thus, the name “conditional molecule generation. ” When using text strings to represent the molecules, e.g., GS: X = [x1, x2, ..., x|S|] , GT : Y = [y1, y2, ..., y|T|] , where xi / yj is the tokens in the text strings of source / target molecules, sequence-to-sequence modeling can be used to train a machine learning model. The training objective is formulated to minimize:where Pθ denotes the parameterized conditional probability with parameter θ to be learned. In the inference stage, the trained machine learning model can be used to autoregressively generate the target string-based molecular representation. With strings as the representations of molecules, conditional molecule generation can be formulated as a sequence-to-sequence translation problem analogous to those handled by natural language processing. Methods or models from natural language processing can be adapted for conditional molecule generation.A machine learning model can be trained to perform any type of conditional molecule generation with an appropriate training dataset. The training dataset includes source molecule and target molecule pairs specific to the particular type of conditional molecule generation. The training dataset will include products and reactants for forward synthesis prediction and retrosynthesis prediction. For molecular property improvement, the training dataset may contain multiple molecules each labeled with values for a particular property. A training pair is not necessarily limited to only two molecules; there may be multiple source and / or multiple target molecules such as multiple reactant molecules that can be used to synthesize a single product molecule. The accuracy of the machine learning model is based on relationships learned from the training data.The training data is, of course, not the actual molecules themselves but rather representation of those molecules in a format that can be understood by a machine learning model. One contribution provided by this disclosure is a technique for generating string-based molecular representations in which the chemical relationships are more readily understood by machine learning models. Machine learning models trained with the string-based molecular representations of this disclosure are more accurate and provide better predictions than machine learning models trained with prior types of string-based molecular representations.FIG. 1 is a diagram showing the use of a machine learning model 100 to perform conditional molecule generation. The machine learning model 100 may be any type of machine learning model that can perform sequence-to-sequence translation. In some implementations, the machine learning model 100 is an encoder-decoder-based model. The decoder may be an autoregressive decoder. An autoregressive decoder is a model that generates output sequences based on the input representations. It predicts each token conditioned on the previous tokens. An autoregressive decoder generates the elements of the output sequence one-by-one until the decoder decides that the sequence is ready and then it generates a final token. This means that the decoder uses its own estimated output at time t as the input for the next time step. The next input for the decoder is chosen based on the decoder’s parameters. This chosen input, also known as the input vector, is then added to the existing sequence of inputs. This process is repeated in an autoregressive manner, meaning each new input depends on the previous ones. Machine learning models that include autoregressive decoders perform well at language generation tasks.One class of machine learning model 100 that may be used is a transformer-based model. The transformer architecture and model is described in Vaswani, Ashish, et al. “Attention is all you need. ” Advances in neural information processing systems, 30. "Curran Associates, Inc (2017) : 5998-6008. The transformer model as described in Vaswani used without modification to the architecture is referred to as a “vanilla transformer” model. The transformer is an autoregressive model, where the last predicted token is taken as input for predicting the next token. The cross-attention represents the correlation between tokens from input strings and tokens of output strings.At inference, an input molecule 102 is provided to the machine learning model 100. Depending on the specific task that machine learning model 100 is trained to perform, the input molecule 102 may be a product, a reactant, or a molecule with a poor molecular property. The input molecule 102, or more precisely data representing the input molecule 102 may be provided in any form. For example, a user may indicate a common name, a trade name, International Union of Pure and Applied Chemistry (IUPAC) name, select the molecule from a menu or list, or the like. In some implementations, the input molecule 102 is provided as a molecular graph 104. Alternatively, the input molecule 102 identified in another format (e.g., by name) may be converted into a molecular graph 104 using known techniques.A molecular graph 104 is a graph-based method of depicting a molecule as set of vertices and edges where the vertices represent atoms and the edges represent chemical bonds. The vertices are labeled with the types of the corresponding atoms, and the edges are labeled with the types of bonds. This provides a convenient way to visualize and analyze the structure of a molecule. Molecular graphs are powerful tools for describing the topological structure of a molecule and are widely used for visualization purposes. However, the molecular graph 104 is not necessarily displayed or visualized to a user. The molecular graph 104 may be used to generate a string-based molecular representation of the input molecule 102. However, if a string-based molecular representation of the input molecule 102 has been previously generated or is already known, that representation may be obtained or looked up without using the molecular graph 104.An input string-based molecular representation 106 is the way that the input molecule 102 is represented when provided to the machine learning model 100. The string-based molecular representation uses linear representation to describe the structure of the input molecule 102. This string-based representation, like natural language text, provides an ordered sequence of characters in which each character has a meaning and the order of the characters with respect to each other also conveys meaning. The input string-based molecular representation 106 can be processed by the machine learning model 100 similar to a string of natural language text and divided into tokens using known techniques.The input string-based molecular representation 106, and all other string-based molecular representations mentioned in this disclosure, may be any type of linear notation or string-based sequences that represents a molecular structure as a string of characters. Currently, the most frequently used representation is SMILES, but there are other representations known to those of ordinary skill in the art such as DeepSMILES, SMARTS, SMIRKS, and SELFIES. Any of these, or other formats yet to be developed, may be used with the techniques of this disclosure.The machine learning model 100 is trained to generate an output string-based molecular representation 108 based on the input string-based molecular representation 106 and potentially other inputs or commands provided by the user. The additional inputs or commands may vary depending on the conditional molecule generation task that the machine learning model 100 is trained to perform. For example, if the conditional molecule generation task is molecular property improvement, the machine learning model 100 may also receive instructions that provide a target value for the property (i.e., how much it should be improved) and a molecular similarity threshold indicating how much the input molecule 102 can be changed.In the inference stage, the machine learning model 100 is used to autoregressively generate the output string-based molecular representation 108. The accuracy of the output string-based molecular representation 108 representing an output molecule (s) 110 with the desired relationship to the input molecule 102 is based on the training of the machine learning model 100. The output string-based molecular representation 108 is the same format of representation as the input string-based molecular representation 106. That is, if the input string-based molecular representation 106 is SMILES then the output string-based molecular representation 108 is also a SMILES string.The output string-based molecular representation 108 may be converted to a different representation of the output molecule (s) 110 such as a common name, skeletal structure, or any other format. The output string-based molecular representation 108 may also be provided directly to a user without translation or conversion. In some implementations, the output string-based molecular representation 108 may be provided to other computer-based processing, for example through an application program interface (API) , without being directly presented to a user.FIG. 2 compares traversal orders of molecules used for generating string-based molecular representations with canonical SMILES, R-SMILES, and O-SMILES. FIG. 2 shows, for the same set of products and reactants, illustrative traversal orders of molecular graphs that would be used to generate string-based representations for three different versions of SMILES. Individual atoms in the molecular graphs are labeled with numerals. These numerals provide unique identifiers for each atom. The sequence of the numerals also indicates the traversal order used by the algorithms that generate the respective SMILES strings. Thus, the traversal order starts at the atom labeled 0, proceeds to the atom labeled 1, and so forth through all of the atoms in a molecular graph.Techniques for generating canonical SMILES strings are documented and vary depending on the specific software used. SMILES is generated by a depth-first traversal of the molecular graph, a molecule can have multiple valid SMILES representations, which leads to the existence of multiple correct output SMILES for a given input SMILES. The one-to-many mapping between input SMILES and output SMILES renders synthesis prediction extremely challenging as the computational model learns only the chemical rules for chemical reactions but also the SMILES syntax for SMILES string validity. Several canonicalization methods can be adopted to generate canonical SMILES that ensures a one-to-one mapping between molecules and SMILES. However, these methods are designed for creating representations of individual molecules without considering the relationship between product and reactant molecules.In some implementations, the Morgan algorithm is used to assign unique, sequential atom numbering for every atom in a given molecule. It operates in two phases by first enumerating all atom numbering obeying certain rules and then iteratively eliminating assignments until only one remains. The initial value assigned to each atom accounts for the following atomic invariants: number of heavy atom connections; number of non-hydrogen bonds; atomic number; sign of charge; absolute charge; number of attached hydrogens. The Morgan algorithm is described in H.L. Morgan “The Generation of a Unique Machine Description for Chemical Structures-A Technique Developed at Chemical Abstracts Service. ” J. of Chem. Documentation 1965 5 (2) , 107-113.The CANGEN algorithm is a technique that may be used to generate a SMILES string. The iterative refinement process in the CANGEN algorithm involves assigning initial values to atoms based on atomic properties, updating these values based on neighboring atom values, and repeating this process until a stable state is reached, resulting in unique identifiers for each atom. The CANGEN algorithm is described in David Weininger, Arthur Weininger, and Joseph L. Weininger “SMILES. 2. Algorithm for generation of unique SMILES notation, ” J. of Chem. Info. and Computer Sci. 1989 29 (2) , 97-101. Both the Morgan algorithm and the CANGEN algorithm can be used for canonical smiles generation.Canonical SMILES generation is not order aware as illustrated in FIG. 2. The starting atom 0 is an oxygen in the product and a nitrogen in the reactants. The traversal order of the product and the reactants is very different even though the structures are similar. The difference in traversal orders results in generation of products / reactant SMILES strings that are different from each other not only due to difference in molecular structure but also due to the arbitrary order in which the algorithm traverses the molecular graphs. Thus, the use of canonical SMILES for generating string-based molecular representation to train machine learning models creates training data that is challenging for a model to learn.R-SMILES is a modification of canonical SMILES that utilizes the same root (i.e., starting atom) for both the products and reactants molecules, reducing the edit distance between the respective string-based representations and facilitating the input-output mapping process. R-SMILES as described in Zipeng Zhong et al., “Root-aligned SMILES: a tight representation for chemical reaction prediction. ” Chem. Sci., 2022, 13, 9023-9034. However, even when sharing the same root, it is possible for the traversal order to be different. FIG. 2 shows an example in which the traversal order on the ring structures in the product and reactants differs even though node 0 and node 1 represent the same atoms. The difference in traversal orders can, like canonical SMILES, result in different string-based molecular representations with a large edit distance even though the product and reactants are structurally similar.O-SMILES generated according to the techniques of this disclosure is a molecule representation method that aligns not only the root node but also the neighbor node to make the input and output SMILES more similar, further reducing the edit distance by a large margin. FIG. 2 shows the same traversal order not only for the starting notes but for all portions of the products and reactants that have the same structure. O-SMILES is specially designed for conditional molecule generation, where the string-based molecular representation of the target molecule is decided by the traversal order of the input molecule. Specifically, this approach utilizes the canonical algorithm and root alignment of R-SMILES but extends them by first assigning a unique order to the atoms in the input SMILES sequence. The input molecule will be the product in retrosynthesis or the reactants in forward synthesis. The target SMILES is then determined by mapping the same atoms in the input molecule to the output molecule according to this order. Hence, O-SMILES can eliminate the issues discussed in the above cases and training a machine learning model is easier because the one-one mapping between input and output molecules is maintained. When trained on O-SMILES strings, the machine learning model is primarily learning the chemical relationships between the products and reactants. This minimizes the need for the model to learn the complex syntax of canonical algorithms.The iterative refinement process of the canonicalization algorithm used to generate canonical SMILES (and R-SMILES after the root atom) can pose a challenge for an autoregressive decoder to infer the atomic order because identical chemical structures can be represented by different strings. O-SMILES improves upon other representation techniques because it maintains sequential dependency. Autoregressive decoders generate sequences one element at a time, and each prediction depends on the previous one. In the context of molecular generation, this means that the order in which atoms are added to the molecule can significantly impact the final structure. However, the iterative refinement process in algorithms like CANGEN does not inherently have a specific order for the atoms.The technique of O-SMILES is able to generate string-based molecular representations with lower computational complexity. The iterative nature of both the refinement process and the autoregressive decoding for canonical SMILES and R-SMILES can lead to increased computational complexity. This can make the process slower and more resource-intensive, particularly for larger molecules. Thus, the use the O-SMILES can reduce processor cycles and energy consumption during the processes of generating string-based molecular representations and training a machine learning model. It additionally results in more accurate predictions by the machine learning model.FIG. 3 is a diagram of the overall framework and an example of O-SMILES generation for retrosynthesis prediction. Because of the random choices in molecular graph traversal, there are a large number of valid SMILES sequences for a given molecule. For conditional molecule generation, the traversal order between a source molecule 300 and the corresponding target molecule 302 should be roughly a one-one mapping so that the model learning can be easier with this prior knowledge. Thus, this technique maintains the randomness in SMILES generation, while leveraging the one-one mapping property retained in the source and target molecules to ensure the uniqueness of the label sequence.In this example, the source molecule 300 (GS) is a product of a synthesis reaction between target molecules 302 (GT) that are reactants which can be used to synthesize the product. The current way for conditional molecule generation is to sequentially generate the tokens in string-based molecular representations. However, in supervised learning, a sequence label needs to be assigned for each training sample, and the valid traversal order is arbitrary. In addition to adopting a canonicalization algorithm, the techniques in this disclosure also assign a unique sequence label for each training sample according to its source molecule 300. In this way, how to traverse the target molecule 302 is determined by its input, not an external algorithm or a black box tool.First, a source neighbor priority queue 304 is randomized for each atom in the source molecule 300. The source neighbor priority queue 304 is a random ordering of the atoms in the source molecule 300 that also identifies adjacent atoms for each atom in the source molecule 300. The source neighbor priority queue 304 is generated by first assigning each atom in the source molecule 300 a unique identitywhere | M |is the number of atoms in the source molecule.For all but the simplest molecules, there are multiple source neighbor priority queues 304 (Q) that can be generated. A random neighbor priority queue can be generated for each atom with all its neighbor nodeswhereThe priority queue may be initialized randomly, rather than provided by external tools. Then, one atom from the source molecule 300 is selected as a root node. The root node may be selected randomly.Second, starting from this root node and source neighbor priority queue 304 of each node, a depth-first search (DFS) algorithm is applied to obtain a source traversal order 306 (LS) . Each source neighbor priority queue 304 can result in a different source traversal order 306. DFS is an algorithm used for traversing or searching tree or graph data structures. The function of the source neighbor priority queue 304 is that when visiting node i, the more front the neighbor node is in the queue, the higher priority of it being visited. However, to form a valid SMILES (or other string-based molecular representation) later, the source neighbor priority queue 304 may not be completely correct, e.g., when the current node lies on a ring system, the priority of its branch neighbor should be higher than the neighbor in the same ring according to the rules for generating SMILES. Thus, an order generated by the source neighbor priority queue 304 may undergo post-processing with the rules for generating a valid string-based molecular representation (such as SMILES) to obtain the source traversal order 306 (LS) . Last, the source neighbor priority queue 304 Q is updated with the source traversal order 306 to ensure the validness for generation of the source string-based molecular representation 316.Third, according to the source traversal order 306 for the source molecule 300 the source neighbor priority queue 304 is translated from the source side to the target side. In an implementation, an atom mapper 308 may be used to make this translation. The unique source string-based molecular representation 316 generated for the source molecule 300 is used as the label for the target molecule 302. This creates a target neighbor priority queue 310 for the structure of the target molecule 302 that is informed by the identity of corresponding atoms which may be, but are not necessarily, identified by the atom mapper 308.Atom mapping is a process that identifies the correspondence between the atoms of reactants and products in a chemical reaction. It is a one-to-one correspondence, also known as atom-atom mapping, that remains unchanged even when chemical bonds are rearranged during a reaction. Atom maps convey the complete information necessary to disentangle the mechanism, i.e., the bond rearrangement, of a chemical reaction because they unambiguously identify the bonds that differ between reactant and product molecules.The atom mapper 308 is leveraged so that the unchanged parts between the source molecule 300 and target molecule 302 constitute a one-one mapping. The atom mapper 308 translates the source neighbor priority queue 304 from the source side to the target side. In one implementation, the queue order according to the mapping information is maintained for those atoms that are common in both the source molecule 300 and the target molecule 302; atoms that exist in the source molecule 300 and not in the target molecule 302 are discarded; atoms that exist in the target molecule 302 and not in the source molecule 300 are placed at the end of the target neighbor priority queue 310.Any existing or later developed atom mapper 308 may be used. The atom mapper 308 is a tool used in computational chemistry to map the atoms of reactants to the corresponding atoms of products in a chemical reaction. There are various open source and commercial atom mappers known to those of ordinary skill in the art including, but not limited to, RXNMapper, Indigo, ChemAxon, RDTool which is part of RDKit, and NextMove. Like the source neighbor priority queue 304, there are multiple target neighbor priority queues 310 that can be generated from the target molecule 302.Next, the DFS algorithm is again applied to the target neighbor priority queue 310 in the same way as before to obtain a target traversal order 312 (LT) . A separate target traversal order 312 may be generated for each of the target neighbor priority queues 310.Finally, a string generator 314, such as, but not limited to, an existing SMILES generator, is used to generate a source string-based molecular representation 316 from the source traversal order 306 (LS) and a target string-based molecular representation 318 from the target traversal order 312 (LT) . For example, to generate SMILES strings, RDKit (available at rdkit. org) may be used by replacing a traversal order generated by the Morgan algorithm with the source traversal order 306 or the target traversal order 312. Each pair of string-based molecular representations generated from the same traversal order may be used as a training pair. Because there may be multiple possible source traversal orders 306 and possible target traversal orders 312, the string generator 314 may generate multiple different source string-based molecular representations 316 as well as multiple different target string-based molecular representations 318. Note that for each source molecule 300 and target molecule 302 pair, due to possible randomness in selecting the root node and in initialization of the source neighbor priority queue 304 of the source molecule 300, the training pairs are not unique. Without being bound by theory, it is believed that these various traversal orders could make the decoder machine learning model learn more about the molecular graph from the expanded string-based molecular representation space and increase the generalization ability of the model.One advantage of this technique is data argumentation. This technique for generating string-based molecular representations can also be viewed as a simple and effective data augmentation method that considers molecular graph information between the input and output molecule graphs. This technique can conduct data augmentation from two levels, the random root and the sampled atom order which can generate multiple input-output pairs as training data by enumerating different atoms as the root node and by following different traversal orders of the atoms based on random neighbor priority queues. Generating multiple different strings representing a molecule (where each is paired with another string generated using the same starting node and traversal order) can improve the generalization of a machine learning model because there is more training data and the data includes multiple alternative string-based representations of the same chemical structure.FIG. 4 is a flow diagram of an illustrative method 400 for generating string-based molecular representations for a source molecule and a target molecule model that are used to train a machine learning model. Method 400 may be performed using the techniques illustrated in FIG. 3 and the system illustrated in FIG. 6 below.At operation 402, an indication of a source molecule is received. The indication of the source molecule may be any type of representation of the source molecule such as, but not limited to, a molecular graph representing the source molecule. It may also be a name of the source molecule or selection of the source molecule, database or list. The indication of the source molecule may be provided by a user or it may be provided by another software program.At operation 404, a source neighbor priority queue is built that captures neighbor relationships between atoms in the source molecule. The source neighbor priority queue assigns each atom in the source molecule a unique identity and identifies neighboring atoms for each atom in the source molecule. In some implementations, the atoms in the source neighbor priority queue are ordered according to a random priority order.At operation 406, one atom in the source molecule is selected as a root node. The root node may be selected according to a set of rules. For example, it may be selected based on the type of atom or its connectivity to other atoms. It may also be selected randomly.At operation 408, a source traversal order of the source molecule is generated. The source traversal order may be generated by performing a DFS through the source neighbor priority queue starting from the root node. The DFS traverses the source neighbor priority queue by a specific order which becomes the source traversal order. The source traversal order as initially generated by DFS may be modified based on rules for generating a valid string-based molecular representation. For example, if the string-based molecular representation is SMILES, rules for traversing a molecular graph such as dealing with the order of atoms in a ring may be applied to modify the traversal order generated by the DFS algorithm. This is done so that the ultimate traversal order conforms with any rules used to generate a valid string-based molecular representation.At operation 410, an indication of a target molecule is received. Indication of the target molecule may be received in the same way the indication of the source molecule was received at operation 402. Thus, in some implementations, the indication of the target molecule may be received as a molecular graph representing the target molecule. The target molecule and the source molecule are a pair of labeled data used for training. The relationship between the target molecule and the source molecule depends on the specific conditional molecule generation task that is being trained. For example, the target molecule may be a product and the source molecule may be a reactant when training for retrosynthesis prediction.At operation 412, a target neighbor priority queue of the target molecule is built based on the source traversal order of the source molecule generated at operation 408 and atom correspondence between the source molecule and the target molecule. Atom correspondence is identification of atoms that are the same in the source molecule and the target molecule after molecular arrangements. In some labeled reaction data such as USPTO-50K atom correspondence for products and reactants is provided. However, if the atom correspondence is not provided in the training data it may be identified using any one of many readily available atom mappers by techniques that are known to those of ordinary skill in the art. Building the target neighbor priority queue may include maintaining a same queue order for atoms in the target molecule that have a one-to-one correspondence with atoms in the source molecule, omitting atoms from the source molecule that are not present in the target molecule, and including atoms that are present in the target molecule but not in the source molecule at the end of the target neighbor priority queue.At operation 414, the atom in the target molecule that corresponds to the root node from the source molecule is selected as the root node for the target molecule. The atom mapper may be used to identify the root node in the target molecule.At operation 416, a target traversal order in the target molecule is generated. The target traversal order may be generated by performing a DFS through the target neighbor priority queue starting from the root node. The technique for generating the target traversal order may be identical to the technique used for generating the source traversal order at operation 408. Thus, the target traversal order may also be based on rules for generating a valid string-based molecular representation.At operation 418, a source string-based molecular representation is generated by a string generator using the source traversal order. The specific technique for generating the source string-based molecular representation may be the same as used by an existing technique represented molecular structures of strings such as SMILES. The difference is that the standard process for determining a traversal order to the molecular structure to generate the string is replaced by the source traversal order generated according to this method 400. The source string-based molecular representation of the source molecule is a string of characters that represents the structure of the source molecule.At operation 420, a target string-based molecular representation is generated by the string generator using the target traversal order. The target string-based molecular representation is generated by the same techniques used to generate the source string-based molecular representation. Because the traversal order from the source molecule is carried over to that of the target molecule, differences between the source string-based molecular representation and the target string-based molecular representation will be due to differences in the molecules not due to arbitrary traversal order differences.At operation 422, a machine learning model is trained with the source string-based molecular representation and the target string-based molecular representation as a training pair. The machine learning model may be a model that includes an autoregressive decoder such as, but not limited to, a transformer model. The machine learning model is typically trained on a large set of training data that contains multiple training pairs. These multiple training pairs can come from multiple labeled reactions. Additional training pairs can come from data augmentation techniques available with this method 400.A single source molecule and target molecule pair can be used to generate multiple pairs of source and target string-based molecular representations. There are multiple possible source neighbor priority queues that can be generated from the source molecule. The number generally increases with the complexity of the source molecule. Further variation can be introduced by using different root nodes. Using DFS, a different source traversal order can be created for each of the different source neighbor priority cues. Thus method 400 can generate a plurality of source traversal orders of the source molecule that represent alternative traversal orders through the source molecule.Similarly, there can be multiple target neighbor priority cues. From these, DFS may be used to generate a plurality of target traversal orders of the target molecule that represent alternative traversal orders through the target molecule. Each of the plurality of source traversal orders is paired with one of the plurality of target traversal orders based on atom correspondence. The atom correspondence may be included in the training data or identified by the atom mapper. This ensures that each source traversal order and target traversal order is paired such that the traversal orders are the same to the extent possible given the structure of the molecules.Now, with multiple source traversal orders and multiple target traversal orders it is possible to generate a plurality of paired source string-based molecular representations and target string-based molecular representations for the source molecule and the target molecule. Each pair of these string-based molecular representations is generated from a different traversal order. This provides multiple training pairs from a single pair of molecules that can be used to train the machine learning model.Once trained, this machine learning model may be used at inference by providing it an indication of an input molecule and receiving from it an indication of an output molecule. The indication of the input molecule is converted to the same string-based molecular representation used for training the machine learning model if it is not already in that format. As used herein, source molecule and target molecule refer to molecules used for training a machine learning model. Input molecule and output molecule refer to molecules provided to and received from the machine learning model when that model is used to generate a predicted molecular structure.FIG. 5 is a flow diagram of an illustrative method 500 for using a machine learning model to perform conditional molecule generation. Method 500 may be performed using the techniques illustrated in FIG. 1 or the system illustrated in FIG. 6 below.At operation 502, an indication of input molecule is received. The indication of the source molecule may be any type of representation of the source molecule such as, but not limited to, a molecular graph representing the source molecule. It may also be a name of the source molecule or selection of the source molecule, database or list. The indication of the source molecule may be provided by a user or it may be provided by another software program.If the conditional molecule generation task is retrosynthesis prediction, then the input molecule is a product. If the conditional molecule generation task is for synthesis protection, then the input molecule is a reactant. If the conditional molecule generation task is molecular property improvement, then the input molecule is a source molecule with a first value for a molecular property.At operation 504, a string-based molecular representation of input molecule is provided to machine learning model. If the indication of the input molecule received at operation 502 is in another format, it is converted to the string-based molecular representation that was used for training the machine learning model. For example, if the machine learning model was trained on SMILES, then the indication of the input molecule is provided as a SMILES string. The machine learning model may include an autoregressive decoder and, in some implementations, may be a transformer model. The machine learning model may include a tokenizer that divides the string-based molecular representation into tokens using any one of a number of known techniques.The machine learning model is trained on a plurality of source molecule and target molecule pairs represented as string-based molecular representations. For each of the source molecule and target molecule pairs, a target traversal order of a target molecule is determined by a source traversal order of a source molecule. The same root node is used in the target traversal order and the source traversal order, and the target traversal order is the same as the source traversal order for all atoms that share a one-to-one correspondence between the source molecule and the target molecule. In some implementations, the atoms that share the one-to-one correspondence between the source molecule and the target molecule are identified as such in the training data. However, in other implementations the one-to one correspondence is identified by use of an atom mapper. Thus, portions of the target molecule that are structurally the same as the source molecule are traversed in the same order.In some implementations, the source traversal order is generated by building a source neighbor priority queue that captures neighbor relationships between atoms in the source molecule. A DFS is then performed through the source neighbor priority queue thereby generating the source traversal order. Also, in such implementations the target traversal order is generated by building a target neighbor priority queue that captures neighbor relationships between atoms in the target molecule. A DFS is then performed through the target neighbor priority queue thereby generating the target traversal order.The source traversal order is then used to generate a source string-based molecular representation of the source molecule. Similarly, target traversal order is used to generate a target string-based molecular representation of the target molecule. Because the traversal orders are the same to the extent possible, this minimizes an edit distance between the source string-based molecular representation and the target string-based molecular representation. In other words, differences between the two string representations are due to structural differences in the molecules not due to arbitrary differences in traversal orders used for generating the string-based molecular representations. Using string-based molecular representations generated in this way to train the machine learning model affects the weights of neural networks within the machine learning model and the structure of the model itself.At operation 506, an indication of an output molecule is generated by the machine learning model. If implemented as a transformer model, the machine learning model may autoregressively generate a string-based representation of the output molecule. If the conditional molecule generation task is molecular retrosynthesis prediction, then the output molecule is a reactant. If the conditional molecule generation task is forward synthesis prediction, then the output molecule is a product. If the conditional molecule generation task is molecular property improvement, then the output molecule is a molecule with more than a threshold similarity to the source molecule and a second value for the molecular property. The second value represents an improvement or a better value for the property than the first value associated with the input molecule.Indication of output molecule is generated by the machine learning input string-based molecular representation. However, this string-based representation may be converted into one or more other formats, using known techniques, such as into a skeletal structure or chemical name. The indication of the output molecule may be provided to a user or sent to another software program.FIG. 6 shows details of an example computing system 600 for a device, such as a computer or a server configured as part of a cloud-based platform, capable of executing computer instructions (e.g., a module or a component described herein) . The computer architecture 600 illustrated in FIG. 6 includes one or more processor (s) 602, a system memory 604, including a random-access memory 606 ( “RAM” ) and a read-only memory ( “ROM” ) 608, and a system bus 610 that couples the memory 604 to the processors (s) 602. The processor (s) 602 may also comprise or be part of a processing system. In various examples, the processor (s) 602 of the processing system are distributed. Stated another way, one processor (s) 602 of the processing system may be located in a first location (e.g., a rack within a datacenter) while another processor (s) 602 of the processing system is located in a second location separate from the first location.Processing unit (s) , such as processor (s) 602, can represent, for example, a CPU-type processing unit, a GPU-type processing unit, a field-programmable gate array (FPGA) , another class of digital signal processor (DSP) , or other hardware logic components that may, in some instances, be driven by a CPU. For example, illustrative types of hardware logic components that can be used include Application-Specific Integrated Circuits (ASICs) , Application-Specific Standard Products (ASSPs) , System-on-a-Chip Systems (SOCs) , Complex Programmable Logic Devices (CPLDs) , and the like.A basic input / output system containing the basic routines that help to transfer information between elements within the computer architecture 600, such as during startup, is stored in ROM 608. The computer architecture 600 further includes a mass storage device 612 for storing an operating system 614, application (s) 616, modules / components 618, and other data described herein. The operating system 614, application (s) 616, and modules / components 618 may comprise computer-executable instructions implemented by the processor (s) 602. Examples of module / components 618 include a machine learning model 100, an atom mapper 308, and a string generator 314.The mass storage device 612 is communicatively connected to processor (s) 602 through a mass storage controller connected to the bus 610. The mass storage device 612 provides non-volatile storage for the computer architecture 600. It should be appreciated by those skilled in the art that the mass storage device 612 can be any available computer-readable storage medium or communications medium that can be accessed by the computer architecture 600. The mass storage device 612 is a type of memory. Anything shown as stored in the mass storage device 612 may alternatively be stored on another computing device such as one accessible via the network 620.Computer-readable media can include computer-readable storage media and / or communication media. Computer-readable storage media can include one or more of volatile memory, nonvolatile memory, and / or other persistent and / or auxiliary computer storage media, removable and non-removable computer storage media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data. Thus, computer storage media includes tangible and / or physical forms of media included in a device and / or hardware component that is part of a device or external to a device, including RAM, static random-access memory (SRAM) , dynamic random-access memory (DRAM) , phase-change memory (PCM) , ROM, erasable programmable read-only memory (EPROM) , electrically erasable programmable read-only memory (EEPROM) , flash memory, compact disc read-only memory (CD-ROM) , digital versatile disks (DVDs) , optical cards or other optical storage media, magnetic cassettes, magnetic tape, magnetic disk storage, magnetic cards or other magnetic storage devices or media, solid-state memory devices, storage arrays, network-attached storage, storage area networks, hosted computer storage or any other storage memory, storage device, and / or storage medium that can be used to store and maintain information for access by a computing device.In contrast to computer-readable storage media, communication media embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave, or other transmission mechanism. As defined herein, computer-readable storage medium does not include communication medium. That is, computer-readable storage media does not include communications media and thus excludes media consisting solely of a modulated data signal, a carrier wave, or a propagated signal, per se.According to various configurations, the computer architecture 600 may operate in a networked environment using logical connections to remote computers through a network 620. The computer architecture 600 may connect to the network 620 through a network interface unit 622 connected to the bus 610. An I / O controller 624 may also be connected to the bus 610 to control communication in input and output devices.It should be appreciated that the software components described herein may, when loaded into the processor (s) 602 and executed, transform the processor (s) 602 and the overall computer architecture 600 from a general-purpose computing system into a special-purpose computing system customized to facilitate the functionality presented herein. The processor (s) 602 may be constructed from any number of transistors or other discrete circuit elements, which may individually or collectively assume any number of states. More specifically, the processor (s) 602 may operate as a finite-state machine, in response to executable instructions contained within the software modules disclosed herein. These computer-executable instructions may transform the processor (s) 602 by specifying how the processor (s) 602 transitions between states, thereby transforming the transistors or other discrete hardware elements constituting the processor (s) 602.Examples(1) Retrosynthesis PredictionThe retrosynthesis prediction task, aims to predict the reactants (target) for a product molecule (source) . The retrosynthesis experiment used a widely used benchmark, called USPTO-50K, which is extracted from the United States Patent and Trademark Office (USPTO) literature. The USPTO-50k consists of 50K reaction pairs of which the reaction type and the atom mapping information are available. Thus, for this dataset an atom mapper was not needed to generate the training data. For a fair comparison with prior works, the data was the same as used in Hanjun Dai, Chengtao Li, Connor Coley, Bo Dai, and Le Song. Retrosynthesis prediction with conditional graph logic network. Advances in Neural Information Processing Systems, 32, 2019. This dataset split the training, validation, and test data into 80%, 10%and 10%in advance. Experiments were conducted in two settings where the reaction type is known or not.Datasets created to have string-based molecular representations generated according to the techniques of this disclosure will be compatible with any model architecture designed for sequence-to-sequence translation. In this example, the vanilla Transformer architecture was used. Specifically, both the encoder and decoder have 8 Transformer layers, with embedding dimension 256, feed-forward layer dimension 1024, and attention head 4. The dropout before the residual connection is 0.3 and the weight of label smoothing is 0.1. The optimization algorithm is Adam with learning rate 0.0005 and invert_sqrtlearning rate scheduler

[0036] . The Adam algorithm is described in Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv: 1412.6980, 2014.This method is evaluated by the top-k exact match accuracy (or top-k accuracy) , whether one of the top-k predicted reactants is exactly the same as the ground-truth, where k ∈ {1, 3, 5} . For each product molecule, the root node and neighbor priority queue are randomized, then sorted by their scores, and filtered the last top-k reactant set as the prediction. This is a template-tree method and is compared to other template-free methods which include Transformer with canonical SMILES, MEGAN, GTA, Dual-TF, Chemformer, and R-SMILES.The experimental results on USPTO-50K are shown in Table 1 with reaction types known and unknown. Compared with state-of-the-art template-tree methods, the techniques of this disclosure achieve better results. Specifically, compared with canonical SMILES (denoted as C-SMILES) and R-SMILES, the difference of O-SMILES is that this technique changes the model input and label in a systematic way (i.e., by generating the string-based molecular representations in an order-aware manner) . This change alone resulted in improvements which shows that forcing the decoder to learn external canonical rules is a source of error and demonstrates that it is effective to let the decoder infer the atomic order from the source molecule.Table 1. Performance comparison on molecular retrosynthesis. The results are reported based on the average performance over five runs.(2) Forward Reaction PredictionThe techniques of this disclosure can be applied to forward reaction prediction in the same way as a retrosynthesis prediction because this technique maps the atom order between source and target molecules, no matter source / target is product / reactants or the reverse. The techniques for performing forward reaction prediction are the same as those described in Zhong et al. Thus, the comparison is only to R-SMILES. Using this technique, which differs from R-SMILES by using the same traversal order after the root node, results in marked improvement.Table 2. Performance comparison on forward reaction prediction.(3) Molecular Property ImprovementFor molecular property improvement, given an input molecule x, the task is to output a different molecule y with better molecular property and to prevent the model from ignoring the input molecule x and generating an arbitrary molecule, the output molecule y must be above a similarity threshold, i.e., sim (x, y) > δ. Because the input and output molecules must be similar to some extent, for this task, RDKit was used to find the maximum common substructure between the pair of molecules. Then the unchanged, common substructure is used for atom mapping information when building O-SMILES representations.This example considers three molecular properties. The first is the penalized logP score (or logP) , which measures the solubility and synthetic accessibility of a compound. The task is about to improve this property such that logP (y) > logP (x) . We conduct on this property with two settings where the similarity threshold δ = {0.4, 0.6} , which are referred to as LogP0.4 and LogP0.6, respectively. The second is the qualitative estimation of the drug-likeness score (QED) , which quantifies the drug-likeness of a compound. The task is to improve a molecule with a QED score from the lower range [0.7, 0.8] to the higher range [0.9, 1.0] . The similarity threshold for this task is δ = 0.4. The third is the DRD2 score, which evaluates the biological activity against dopamine type 2 receptor of a compound with a property prediction model. The task is to improve an inactivated molecule (i.e., DRD2 < 0.5) to be active (i.e., DRD2 > 0.5) . The similarity threshold for this dataset is δ = 0.4.The vanilla transformer is again used as the basic machine learning architecture. The number of transformer layers is 6, the embedding dimension is 256, the attention head is 4, the feed-forward layer dimension is 1024, and the dropout is 0.3. For evaluation, in each task the root node and neighbor priority queue are randomized and then the top-k predictionsare filtered as model output. In all tasks, k = 20. For LogP0.4 and LogP0.6 datasets, the average highest property improvement is used as the evaluation metric. Specifically, compound yi is selected as the final prediction which gives the highest property improvement, i.e., i = argmax (logP (yi) ) , i = {1, ···, k} and satisfies sim (x, yi) ≥ δ, and then calculate the average number over the test set. For QED and DRD2 datasets, the average success rate is used as the evaluation metric. Specifically, for each test data, the model is successful only if one of all k predictions satisfies all the similarity and property improvement constraints. The similarity function is defined as func (x, y) = 1 - Dist (x, y) , where Dist (·, ·) is the Tanimoto distance over Morgan fingerprints of two molecules. Techniques for determining the similarity between two molecules, such as Tanimoto distance, are known to those of ordinary skill in the art.The results for molecular property improvement compared to the following pre-existing methods: JT-VAE, which generates molecular graphs in two phases by exploiting valid subgraphs as components; CG-VAE, which incorporates hard domain-specific constraints into molecular generation; GCPN, a general graph convolutional network-based model for goal-directed graph generation through reinforcement learning; MMPA, a matched molecular pair analysis platform; JTNN, a junction tree encoder-decoder framework; C-SMILES, a Transformer model, with the input and output all canonical SMILES; HierG2G, which generates molecules in a hierarchical encoder-decoder model; BT4MolGen, which utilizes a backward model and unlabeled data for better molecule generation; and R-SMILES.The results are shown in Table 3 below. From the table, it is apparent that techniques of this disclosure outperforms canonical SMILES (C-SMILES) on all four datasets, especially on the QED and DRD2 datasets, e.g., QED score 71.9 with C-SMILES and 94.5 with O-SMILES, where the source molecule and target molecule lie in two different regions, showing the effectiveness of this method for molecular property improvement. Additionally, the method of this disclosure also performs sophisticated model generation methods such as HierG2G and BT4MolGen with lower training costs and simpler training strategy. The lower training costs mean that a model can be trained using fewer processor cycles and less energy. Finally, compared to R-SMILES, maintaining the same traversal order between source and target molecules eases the decoder learning and improves the generation performance.Table 3. Performance comparison on molecular property improvement. The results are reported based on the average performance over five runs.Illustrative EmbodimentsThe following clauses described multiple possible embodiments for implementing the features described in this disclosure. The various embodiments described herein are not limiting nor is every feature from any given embodiment required to be present in another embodiment. Any two or more of the embodiments may be combined together unless context clearly indicates otherwise. As used herein in this document “or” means and / or. For example, “A or B” means A without B, B without A, or A and B. As used herein, “comprising” means including all listed features and potentially including addition of other features that are not listed. “Consisting essentially of” means including the listed features and those additional features that do not materially affect the basic and novel characteristics of the listed features. “Consisting of” means only the listed features to the exclusion of any feature not listed.Clause 1 A method of training a machine learning model (100) to perform conditional molecule generation, the method comprising:receiving an indication of a source molecule (300) ;building a source neighbor priority queue (304) that captures neighbor relationships between atoms in the source molecule;randomly selecting one atom in the source molecule as a root node;generating a source traversal order (306) of the source molecule;receiving an indication of a target molecule (302) ;building a target neighbor priority queue (310) of the target molecule based on the source traversal order of the source molecule and atom correspondence between the source molecule and the target molecule;generating a target traversal order (312) of the target molecule by performing a DFS through the target neighbor priority queue starting from the root node;generating a source string-based molecular representation (316) of the source molecule using the source traversal order;generating a target string-based molecular representation (318) of the target molecule using the target traversal order; andtraining the machine learning model with the source string-based molecular representation and the target string-based molecular representation as a training pair, the machine learning model comprising an autoregressive decoder.Clause 2 The method of clause 1, wherein the indication of the source molecule is a molecular graph representing the source molecule and the indication of the target molecule is a molecular graph representing the target molecule.Clause 3 The method of clause 1 or 2, wherein the source neighbor priority queue assigns each atom in the source molecule a unique identity, identifies neighboring atoms for each atom in the source molecule, and orders the atoms in the source neighbor priority queue according to a random priority order.Clause 4 The method of any of clauses 1–3, wherein the one atom in the source molecule that is the root node is selected randomly.Clause 5 The method of any of clauses 1–4, wherein the source traversal order is generated by performing a depth-first search (DFS) through the source neighbor priority queue starting from the root node.Clause 6 The method of any of clauses 1–5, wherein generating the source traversal order and the target traversal order are further based on rules for generating a valid string-based molecular representation.Clause 7 The method of any of clauses 1–6, wherein the atom correspondence between the source molecule and the target molecule is identified by an atom mapper.Clause 8 The method of any of clauses 1–7, wherein building the target neighbor priority queue comprises:maintaining a same queue order for atoms in the target molecule that have a one-to-one correspondence with atoms in the source molecule,omitting atoms from the source molecule that are not present in the target molecule, andincluding atoms that are present in the target molecule but not in the source molecule at the end of the target neighbor priority queue.Clause 9 The method of any of clauses 1–8, wherein the target traversal order of the target molecule is generated by performing a DFS through the target neighborhood priority queue starting from the root node.Clause 10 The method of any of clauses 1–9, further comprising:generating a plurality of source traversal orders of the source molecule that represent alternative traversal orders through the source molecule;generating a plurality of target traversal orders of the target molecule that represent alternative traversal orders through the target molecule, wherein each of the plurality of source traversal orders is paired with one of the pluralities of target traversal orders based on atom correspondence;generating a plurality of paired source string-based molecular representations and target string-based molecular representations for the source molecule and the target molecule, wherein each pair of string-based molecular representations is generated from a different traversal order; andtraining the machine learning model with the plurality of paired source string-based molecular representations and target string-based molecular representations as labeled training pairs.Clause 11 The method of clause 10, wherein the atom correspondence is identified by the atom mapper.Clause 12 The method of any of clauses 1–11, wherein the machine learning model is a transformer model.Clause 13 The method of any of clauses 1–12, further comprising providing an indication of an input molecule to the machine learning model and receiving from the machine learning model an indication of an output molecule.Clause 14 Computer readable media storing instructions that, when executed by a processor, cause the processor to perform the method of any of clauses 1–13.Clause 15 A system comprising a processor and a memory, the memory storing instructions that when executed by the processor cause the system to perform the method of any of clauses 1–13.Clause 16 A system for training a machine learning model (100) to perform conditional molecule generation, the system comprising:a processor (602) ;memory (612) storing instructions that, when executed by the processor, cause the system to perform operations comprising:receiving an indication of a source molecule (300) ;generating a source traversal order (306) of the source molecule;receiving an indication of a target molecule (302) ;generating a target traversal order (312) of the target molecule based on the source traversal order of the source molecule that maintains a same traversal order for a portion of the target molecule that has a one-to-one atom correspondence with a portion of the source molecule;generating a source string-based molecular representation (316) of the source molecule using the source traversal order;generating a target string-based molecular representation (318) of the target molecule using the target traversal order; andtraining the machine learning model with the source string-based molecular representation and the target string-based molecular representation as a training pair, the machine learning model comprising an autoregressive decoder.Clause 17 The system of clause 16, wherein generating the source traversal order of the source molecule comprises:building a source neighbor priority queue that captures neighbor relationships between atoms in the source molecule,selecting one atom in the source molecule as a root node, andgenerating the source traversal order by performing a DFS through the source neighbor priority queue starting from the root node.Clause 18 The system of clause 16 or 17, wherein generating the target traversal order of the target molecule comprises:building a target neighbor priority queue of the target molecule based on the source traversal order of the source molecule and atom correspondence between the source molecule and the target molecule, andgenerating the target traversal order by performing a DFS through the target neighbor priority queue starting from a root node that was used to generate the source traversal order.Clause 19 The system of clause 18, wherein building the target neighbor priority queue comprises:maintaining a same queue order for atoms in the target molecule that have a one-to-one correspondence with atoms in the source molecule,omitting atoms from the source molecule that are not present in the target molecule, andincluding atoms that are present in the target molecule but not in the source molecule at the end of the target neighbor priority queue.Clause 20 The system of any of clauses 16–19, wherein, generating the source traversal order and the target traversal order are further based on rules for generating a valid string-based molecular representation.Clause 21 A method for performing conditional molecule generation, the method comprising:receiving an indication of an input molecule (102) ;providing an input string-based molecular representation (106) of the input molecule to a machine learning model (100) that comprises an autoregressive decoder and is trained on a plurality of source molecule and target molecule pairs represented as string-based molecular representations,wherein, for each of the source molecule and target molecule pairs, a target traversal order (312) of a target molecule (302) is determined by a source traversal order (306) of a source molecule (300) such that the same root node is used in the target traversal order and the source traversal order and the target traversal order is the same as the source traversal order for all atoms that share a one-to-one correspondence between the source molecule and the target molecule,wherein, the source traversal order is used to generate a source string-based molecular representation (316) of the source molecule and the target traversal order is used to generate a target string-based molecular representation (318) of the target molecule thereby minimizing an edit distance between the source string-based molecular representation and the target string-based molecular representation; andgenerating, by the machine learning model, an indication of an output molecule (110) .Clause 22 The method of clause 21, wherein the conditional molecule generation is molecular retrosynthesis prediction, the input molecule is a product, and the output molecule is a reactant.Clause 23 The method of clause 21, wherein the conditional molecule generation is molecular property improvement, the input molecule is a source molecule with a first value for a molecular property, and the output molecule has more than a threshold similarity to the source molecule and a second value for the molecular property.Clause 24 The method of any of clauses 21–23, wherein the machine learning model comprises a transformer model that during inference uses the machine learning model to autoregressively generate a string-based molecular representation of the output molecule.Clause 25 The method of any of clauses 21–24, wherein the source traversal order is generated by building a source neighbor priority queue that captures neighbor relationships between atoms in the source molecule and, starting from the root node, performing a DFS through the source neighbor priority queue.Clause 26 The method of any of clauses 21–25, wherein the target traversal order is generated by building a target neighbor priority queue that captures neighbor relationships between atoms in the target molecule and is ordered based on the source traversal order of the source molecule and atom correspondence between the source molecule and the target molecule and, starting from the root node, performing a DFS through the target neighbor priority queue.Clause 27 The method of any of clauses 21–26, wherein the atoms that share the one-to-one correspondence between the source molecule and the target molecule are identified by an atom mapper.Clause 28 Computer readable media storing instructions that, when executed by a processor, cause the processor to perform the method of any of clauses 21–27.Clause 29 A system comprising a processor and a memory, the memory storing instructions that when executed by the processor cause the system to perform the method of any of clauses 21–27.ConclusionWhile certain example embodiments have been described, including the best mode known to the inventors for carrying out the invention, these embodiments have been presented by way of example only, and are not intended to limit the scope of the inventions disclosed herein. Thus, nothing in the foregoing description is intended to imply that any particular feature, characteristic, step, module, or block is necessary or indispensable. Indeed, the novel methods and systems described herein may be embodied in a variety of other forms; furthermore, various omissions, substitutions and changes in the form of the methods and systems described herein may be made without departing from the spirit of the inventions disclosed herein. Skilled artisans will know how to employ such variations as appropriate, and the embodiments disclosed herein may be practiced otherwise than specifically described. The accompanying claims and their equivalents are intended to cover such forms or modifications as would fall within the scope and spirit of certain of the inventions disclosed herein.The terms “a, ” “an, ” “the” and similar referents used in the context of describing the invention are to be construed to cover both the singular and the plural unless otherwise indicated herein or clearly contradicted by context. The terms “based on, ” “based upon, ” and similar referents are to be construed as meaning “based at least in part” which includes being “based in part” and “based in whole, ” unless otherwise indicated or clearly contradicted by context. The terms “portion, ” “part, ” or similar referents are to be construed as meaning at least a portion or part of the whole including up to the entire noun referenced.It should be appreciated that any reference to “first, ” “second, ” etc. elements within the Summary and / or Detailed Description is not intended to and should not be construed to necessarily correspond to any reference of “first, ” “second, ” etc. elements of the claims. Rather, any use of “first” and “second” within the Summary, Detailed Description, and / or claims may be used to distinguish between two different instances of the same element (e.g., two different sensors) .In closing, although the various configurations have been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended representations is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claimed subject matter.Furthermore, references have been made to publications, patents and / or patent applications throughout this specification. Each of the cited references is individually incorporated herein by reference for its particular cited teachings as well as for all that it discloses.

Claims

1.A method of training a machine learning model (100) to perform conditional molecule generation, the method comprising:receiving an indication of a source molecule (300) ;building a source neighbor priority queue (304) that captures neighbor relationships between atoms in the source molecule;selecting one atom in the source molecule as a root node;generating a source traversal order (306) of the source molecule;receiving an indication of a target molecule (302) ;building a target neighbor priority queue (310) of the target molecule based on the source traversal order of the source molecule and atom correspondence between the source molecule and the target molecule;generating a target traversal order (312) of the target molecule;generating a source string-based molecular representation (316) of the source molecule using the source traversal order;generating a target string-based molecular representation (318) of the target molecule using the target traversal order; andtraining the machine learning model with the source string-based molecular representation and the target string-based molecular representation as a training pair, the machine learning model comprising an autoregressive decoder.2.The method of claim 1, wherein the indication of the source molecule is a molecular graph representing the source molecule and the indication of the target molecule is a molecular graph representing the target molecule.3.The method of claim 1, wherein the source neighbor priority queue assigns each atom in the source molecule a unique identity, identifies neighboring atoms for each atom in the source molecule, and orders the atoms in the source neighbor priority queue according to a random priority order.4.The method of claim 1, wherein the source traversal order is generated by performing a depth-first search (DFS) through the source neighbor priority queue starting from the root node.5.The method of claim 1, wherein building the target neighbor priority queue comprises:maintaining a same queue order for atoms in the target molecule that have a one-to-one correspondence with atoms in the source molecule,omitting atoms from the source molecule that are not present in the target molecule, andincluding atoms that are present in the target molecule but not in the source molecule at the end of the target neighbor priority queue.6.The method of claim 1, further comprising:generating a plurality of source traversal orders of the source molecule that represent alternative traversal orders through the source molecule;generating a plurality of target traversal orders of the target molecule that represent alternative traversal orders through the target molecule, wherein each of the plurality of source traversal orders is paired with one of the plurality of target traversal orders based on atom correspondence;generating a plurality of paired source string-based molecular representations and target string-based molecular representations for the source molecule and the target molecule, wherein each pair of string-based molecular representations is generated from a different traversal order; andtraining the machine learning model with the plurality of paired source string-based molecular representations and target string-based molecular representations as labeled training pairs.7.The method of claim 1, wherein the machine learning model is a transformer model.8.The method of claim 1, further comprising providing an indication of an input molecule to the machine learning model and receiving from the machine learning model an indication of an output molecule.9.A system for training a machine learning model (100) to perform conditional molecule generation, the system comprising:a processor (602) ;memory (612) storing instructions that, when executed by the processing unit, cause the system to perform operations comprising:receiving an indication of a source molecule (300) ;generating a source traversal order of the source molecule (306) ;receiving an indication of a target molecule (302) ;generating a target traversal order (312) of the target molecule based on the source traversal order of the source molecule that maintains a same traversal order for a portion of the target molecule that has a one-to-one atom correspondence with a portion of the source molecule;generating a source string-based molecular representation (316) of the source molecule using the source traversal order;generating a target string-based molecular representation (318) of the target molecule using the target traversal order; andtraining the machine learning model with the source string-based molecular representation and the target string-based molecular representation as a training pair, the machine learning model comprising an autoregressive decoder.10.The system of claim 9, wherein generating the source traversal order of the source molecule comprises:building a source neighbor priority queue that captures neighbor relationships between atoms in the source molecule,selecting one atom in the source molecule as a root node, andgenerating the source traversal order by performing a DFS through the source neighbor priority queue starting from the root node.11.The system of claim 9, wherein generating the target traversal order of the target molecule comprises:building a target neighbor priority queue of the target molecule based on the source traversal order of the source molecule and atom correspondence between the source molecule and the target molecule, andgenerating the target traversal order by performing a DFS through the target neighbor priority queue starting from a root node that was used to generate the source traversal order.12.The system of claim 11, wherein building the target neighbor priority queue comprises:maintaining a same queue order for atoms in the target molecule that have a one-to-one correspondence with atoms in the source molecule,omitting atoms from the source molecule that are not present in the target molecule, andincluding atoms that are present in the target molecule but not in the source molecule at the end of the target neighbor priority queue.13.The system of claim 9, wherein, generating the source traversal order and the target traversal order are further based on rules for generating a valid string-based molecular representation.14.A method for performing conditional molecule generation, the method comprising:receiving an indication of an input molecule (102) ;providing an input string-based molecular representation (106) of the input molecule to a machine learning model (100) that comprises an autoregressive decoder and is trained on a plurality of source molecule and target molecule pairs represented as string-based molecular representations,wherein, for each of the source molecule and target molecule pairs, a target traversal order (312) of a target molecule (302) is determined by a source traversal order (306) of a source molecule (300) such that the same root node is used in the target traversal order and the source traversal order and the target traversal order is the same as the source traversal order for all atoms that share a one-to-one correspondence between the source molecule and the target molecule,wherein, the source traversal order is used to generate a source string-based molecular representation (316) of the source molecule and the target traversal order is used to generate a target string-based molecular representation (318) of the target molecule thereby minimizing an edit distance between the source string-based molecular representation and the target string-based molecular representation; andgenerating, by the machine learning model, an indication of an output molecule (110) .15.The method of claim 14, wherein the conditional molecule generation is molecular retrosynthesis prediction, the input molecule is a product, and the output molecule is a reactant.16.The method of claim 14, wherein the conditional molecule generation is molecular property improvement, the input molecule is a source molecule with a first value for a molecular property, and the output molecule has more than a threshold similarity to the source molecule and a second value for the molecular property.17.The method of claim 14, wherein the machine learning model comprises a transformer model that during inference uses the machine learning model to autoregressively generate a string-based molecular representation of the output molecule.18.The method of claim 14, wherein the source traversal order is generated by building a source neighbor priority queue that captures neighbor relationships between atoms in the source molecule and, starting from the root node, performing a DFS through the source neighbor priority queue.19.The method of claim 14, wherein the target traversal order is generated by building a target neighbor priority queue that captures neighbor relationships between atoms in the target molecule and is ordered based on the source traversal order of the source molecule and atom correspondence between the source molecule and the target molecule and, starting from the root node, performing a DFS through the target neighbor priority queue.20.The method of claim 14, wherein the atoms that share the one-to-one correspondence between the source molecule and the target molecule are identified by an atom mapper.