Molecular structure transformer for property prediction

The Transformer model with encoder and decoder networks addresses inefficiencies in predicting molecular properties by accurately representing structural information, enabling efficient ionic liquid selection for plastic depolymerization.

JP2026021368APending Publication Date: 2026-02-10X DEVELOPMENT LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025178009
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-11-29
Filing Date
2025-10-22
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing techniques for predicting molecular properties, particularly in the context of ionic liquid selection for depolymerizing plastics, are inefficient due to the high-dimensional space of experimental conditions and the inability of current molecule representation methods to capture complete structural information.

Method used

A computer-implemented method using a Transformer model with an encoder and decoder network to transform multidimensional embedding spaces, incorporating bond string and position representations, to accurately predict molecular properties and reaction outcomes.

Benefits of technology

Enhances the prediction of molecular properties and reaction outcomes by capturing complete structural information, facilitating more efficient recycling of plastics through ionic liquid selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026021368000001_ABST
    Figure 2026021368000001_ABST
Patent Text Reader

Abstract

To better predict characteristics of ionic liquid molecules and / or results of reactions involving the ionic liquid molecules.SOLUTION: The computer-implemented method may include accessing a multi-dimensional embedding space that supports associating embeddings of molecules with predicted values of a given property of the molecules. The method may also include identifying one or more points of interest in the embedded space based on the prediction values. Each of the one or more points of interest may include a set of coordinate values in the multi-dimensional embedded space and may be associated with a corresponding predicted value of the given characteristic. The method may further include generating, for each of the one or more points of interest, a structural representation of the molecule by transforming, using the decoder network, the set of coordinate values included in the point of interest. The method may include outputting, for each of the one or more points of interest, a result identifying a structural representation of the molecule corresponding to the point of interest.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a nonprovisional application of and claims the benefit of U.S. Provisional Application No. 63 / 264,640, filed November 29, 2021, U.S. Provisional Application No. 63 / 264,641, filed November 29, 2021, U.S. Provisional Application No. 63 / 264,642, filed November 29, 2021, and U.S. Provisional Application No. 63 / 264,643, filed November 29, 2021, the entire contents of which are incorporated herein by reference in their entirety for all purposes. [Background technology]

[0002] A problem in chemistry is predicting the specific properties of a new molecule. Predicting molecular properties is useful for identifying new molecules for recycling. Chemical recycling aims to break down plastic waste into the monomer components from which it was created, enabling a circular economy in which polymers are produced from chemically recycled plastics instead of relying on non-renewable inputs derived from petroleum. Plastic recycling can involve converting waste plastics (polyethylene terephthalate (PET), polylactic acid (PLA)) into their monomer components (bis(2-hydroxyethyl) terephthalate (BHET), lactate) to replace virgin plastics derived from oil. Ionic liquids (ILs) are a highly tunable class of chemicals that have shown promising capabilities for depolymerizing plastics, but it is unclear how to navigate the large ionic liquid design space to improve reaction yields.

[0003] Selecting a specific ionic liquid for use in depolymerization is a challenging task. First, given the number of ionic liquid candidates that exist and the different reaction conditions, it is infeasible to experimentally characterize the properties of all ionic liquids under appropriate conditions. More specifically, ionic liquids consist of a tunable selection of cation and anion molecules, resulting in a high-dimensional space for selecting experimental parameters. For example, in the National Institute of Standards & Technology (NIST) ILThermo database, there are 1,652 binary ILs with 244 cations and 164 anions. Combinatorially, this means that there are 38,364 additional new ILs generated from the NIST database alone. Selecting a specific IL under a sampling of experimental conditions (e.g., investigating three solvents, five ratios of ionic liquid to solvent, three temperatures, and three reaction times) results in a highly complex reaction space containing over 5,400,000 different reaction conditions. In typical experimental design, domain knowledge and literature review are required to reduce the search space, but this process is expensive and does not lend itself to evaluating the complete design space.

[0004] Therefore, being able to better predict the properties of ionic liquid molecules and / or the outcomes of reactions involving ionic liquid molecules may facilitate more efficient recycling.

[0005] One approach to generating these predictions is to use machine learning to translate the representation of a new molecule into a prediction. However, machine learning requires that the molecule be represented by a set of numbers (e.g., via characterization, fingerprinting, or embedding).

[0006] However, existing techniques for numerically representing molecules are unable to capture the complete structural information of a molecule: rather, structural information is either completely ignored or only partially represented. Summary of the Invention

[0007] Some embodiments may include a computer-implemented method. The method may include accessing a multidimensional embedding space that supports associating an embedding of a molecule with a predicted value of a given property of the molecule. The method may also include identifying one or more interest points in the multidimensional embedding space based on the predicted value. Each of the one or more interest points may include a set of coordinate values ​​in the multidimensional embedding space, may convey spatial information of atoms or bonds in the molecule, and may be associated with a corresponding predicted value of the given property. The method may further include, for each of the one or more interest points, generating a structural representation of the molecule by transforming the set of coordinate values ​​included in the interest point using a decoder network. Training the decoder network may include learning to convert locations in the embedding space into outputs representing molecular structural properties. Training the decoder network may be performed at least partially concurrently with training the encoder network. The method may include outputting, for each of the one or more interest points, a result identifying a structural representation of the molecule corresponding to the interest point.

[0008] In some embodiments, training the encoder network may involve learning to transform partial or complete bond string and position (BSP) representations of molecules into positions in the embedding space, where each BSP representation may identify the relative positions of atoms connected by bonds within the represented molecule.

[0009] In some embodiments, training an encoder network may involve learning to transform partial or complete molecular graph representations of molecules into positions in an embedding space, where each molecular graph representation can identify bond angles and distances within the represented molecule.

[0010] In some embodiments, the decoder network and the encoder network may be trained by training a Transformer model that uses self-attention. The Transformer model may include a decoder network and an encoder network.

[0011] In some embodiments, the decoder network and the encoder network may be trained by training a transformer model that includes an attention head.

[0012] In some embodiments, the method may include training a machine learning model including an encoder network and a decoder molecule by accessing a set of supplemental training elements. Each of the set of training elements may include a representation of a structure of a corresponding given molecule. The training may further include, for each supplemental training element in the set of supplemental training elements, masking at least a portion of the representation to obscure at least a portion of the structure of the corresponding given molecule. The training may include training the machine learning model to predict the obscured at least a portion of the structure.

[0013] In some embodiments, training the encoder network may further include fine-tuning the encoder network to convert positions in space into predictions corresponding to values ​​of given properties.

[0014] In some embodiments, each BSP representation of a molecule used to train the encoder network may include a set of coordinates for each of the atoms connected by bonds in the represented molecule, and may further identify each of the atoms connected by bonds in the represented molecule.

[0015] In some embodiments, the BSP representation of the molecules may be used to train an encoder network to identify the bond type for each of at least some of the bonds in each molecule.

[0016] In some embodiments, the format of the structural representation identified in the results may differ from the BSP representation.

[0017] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium that includes instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more of the methods disclosed herein.

[0018] In some embodiments, a computer program product is provided that is tangibly embodied in a non-transitory machine-readable storage medium and includes instructions configured to cause one or more data processors to perform some or all of one or more of the methods disclosed herein.

[0019] The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions to exclude any equivalents of the features shown and described, or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, while the invention as claimed has been specifically disclosed by embodiments and optional features, it will be understood that modifications and variations of the concepts disclosed herein may be effected by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims. [Brief explanation of the drawings]

[0020] The present disclosure is described in conjunction with the accompanying drawings. [Figure 1A] 1 shows a diagrammatic representation of a process to produce the ciprofloxacin SMILES string: C1CC1N2C=C(C(=O)C3=CC(=C(C=C32)N4CCNCC4)F)C(=O)O according to an embodiment of the present invention. [Figure 1B] 1 shows a diagrammatic representation of a process to produce the ciprofloxacin SMILES string: C1CC1N2C=C(C(=O)C3=CC(=C(C=C32)N4CCNCC4)F)C(=O)O according to an embodiment of the present invention. [Figure 1C]1 shows a diagrammatic representation of a process to produce the ciprofloxacin SMILES string: C1CC1N2C=C(C(=O)C3=CC(=C(C=C32)N4CCNCC4)F)C(=O)O according to an embodiment of the present invention. [Figure 1D] 1 shows a diagrammatic representation of a process to produce the ciprofloxacin SMILES string: C1CC1N2C=C(C(=O)C3=CC(=C(C=C32)N4CCNCC4)F)C(=O)O according to an embodiment of the present invention. [Figure 1E] 1 shows a diagrammatic representation of a process to produce the ciprofloxacin SMILES string: C1CC1N2C=C(C(=O)C3=CC(=C(C=C32)N4CCNCC4)F)C(=O)O according to an embodiment of the present invention. [Figure 2] 1 illustrates an exemplary process for constructing a BSP representation of a single molecule from a SMILES string, according to an embodiment of the present invention. [Figure 3] 1 illustrates an exemplary process for constructing a BSP representation of a reactant / reagent from a SMILES string, according to an embodiment of the present invention. [Figure 4] 1 illustrates an exemplary process for generating token embeddings that are input to a model (e.g., a Transformer model) according to an embodiment of the present invention. [Figure 5] Figure 1 shows a depiction of masked bond model training in the context of molecular property prediction, according to an embodiment of the present invention. The input molecule is shown at the bottom, with masked bonds in black. This is mapped to a BSP representation, replacing the masked bonds with special mask tokens, and fed into an encoder network. The model returns predictions for bonds at each point, and ultimately learns to fill in the masked locations accurately. [Figure 6] 1 illustrates an exemplary use of an encoder network and a decoder network to convert representations between structural identifiers and an embedding space, according to an embodiment of the present invention. [Figure 7] FIG. 1 shows a diagram of a message passing graph convolutional layer according to an embodiment of the invention. [Figure 8]In atomic masking, nodes are randomly masked and a GNN is trained to predict the correct label of the masked nodes. Figure based on Hu et al., arXiv:1905.12265v3. [Figure 9] This shows how, for context prediction, a subgraph is a K-hop neighborhood around a selected central node, where K is the number of GNN layers. Figure based on Hu et al., arXiv:1905.12265v3. [Figure 10] FIG. 1 illustrates the purpose of latent representation of molecules according to an embodiment of the present invention. [Figure 11] 1 illustrates the calculation of relative position in terms of angle and distance according to an embodiment of the present invention. [Figure 12A] 1 illustrates message generation and message aggregation according to an embodiment of the present invention. [Figure 12B] 1 illustrates message generation and message aggregation according to an embodiment of the present invention. [Figure 13] 1 illustrates read attention aggregation according to an embodiment of the present invention. [Figure 14] 1 shows an overview of a model including an encoder network used for molecular property prediction according to an embodiment of the present invention. [Figure 15] Figure 1 illustrates a reaction prediction model according to an embodiment of the present invention. The main component is a transformer-based architecture that operates on BSP input and returns a SMILES string that predicts the reaction products. [Figure 16A] 10 shows an example of a correctly predicted reaction that demonstrates understanding of a model of the active site of a reaction, according to an embodiment of the present invention. [Figure 16B] 10 shows an example of a correctly predicted reaction that demonstrates understanding of a model of the active site of a reaction, according to an embodiment of the present invention. [Figure 17A] Examples of incorrectly predicted responses are shown. [Figure 17B] Examples of incorrectly predicted responses are shown. [Figure 18]10 shows the attention weights of the fourth head of the encoder network at its third layer according to an embodiment of the present invention. [Figure 19] 1 illustrates traversal of a molecular graph in a depth-first search to build a combined string representation, according to an embodiment of the present invention. [Figure 20] 1 illustrates a directional variational transformer model according to an embodiment of the present invention. [Figure 21] 1 illustrates the relationship between molecules, embeddings, and property predictions according to an embodiment of the present invention. [Figure 22] 1 illustrates a process for identifying ionic liquids for depolymerizing compounds according to an embodiment of the present invention. [Figure 23] 1 shows ionic liquid cations generated from probing the buried space around a molecule, according to an embodiment of the present invention. [Figure 24] 1 shows an ionic liquid cation generated from probing the buried space between two molecules, according to an embodiment of the present invention. [Figure 25] 1 illustrates the interplay between Bayesian optimization and benchtop experiments according to an embodiment of the present invention. [Figure 26] 1 illustrates a Bayesian optimization process according to an embodiment of the present invention. [Figure 27] 1 shows an exemplary process flow chart associated with ionic liquid-based depolymerization optimization in accordance with an embodiment of the present invention. [Figure 28A] 1 shows the results of applying Bayesian optimization to minimizing the enthalpy of mixing according to an embodiment of the present invention. [Figure 28B] 1 shows the results of applying Bayesian optimization to minimizing the enthalpy of mixing according to an embodiment of the present invention. [Figure 29] 1 is an exemplary architecture of a computing system implemented as some embodiments of the present disclosure.

[0021] In the accompanying drawings, similar components and / or features may have the same reference numerals. Furthermore, various components of the same type may be distinguished according to the reference numeral by a prime and a second numeral that distinguishes between the similar components. When only a first numeral is used herein, the description is applicable to any one of the similar components having the same first numeral, regardless of the second numeral. DETAILED DESCRIPTION OF THE INVENTION

[0022] I. Overview The embedding framework can map individual molecules to embeddings in a high-dimensional space in which structurally similar molecules map closer to each other. These representations can be processed using a molecular property prediction model, and novel molecules can be identified in the space corresponding to representations from some seed set of subjects. These embeddings can be fed as input to a model that estimates specific thermodynamic properties that can be used to predict a molecule's ability to degrade a specific polymer. Molecules with unfavorable properties can be filtered out, and the search can be expanded around promising candidates, ultimately returning a small set of molecules (e.g., ionic liquids) predicted to efficiently depolymerize plastics. Candidate molecules can be processed by a Bayesian optimization system that recommends new experiments, learns from their results, and recommends further experiments until convergence on optimal reaction performance. Bayesian optimization can also be performed over the learned embedding space using the embedding framework.

[0023] II. Machine Learning for Generating Multidimensional Representations of Molecules Accurately representing molecules is key to using models to predict molecular properties, design new molecules with desired properties, or predict chemical reaction outputs. Existing approaches to representing molecules include two categories: property-based and model-based.

[0024] A property-based fingerprint is a collection of features that describe different aspects of a molecule. For example, a molecule can be represented by a vector that describes the number of each type of atom it contains, as shown for methanol below: [ka]

[0025] An exemplary property-based fingerprint for methanol may include a count of each atom in the molecule.

[0026] Another example of a property-based fingerprint is the Morgan fingerprint. Morgan fingerprints (sometimes known as extended connectivity fingerprints, or ECFPs) use limited structural information to construct a vector representation of a molecule. In particular, Morgan fingerprints only partially capture the structure of a molecule, limited by not accounting for the three-dimensional orientation of atoms. And while fingerprints capture some details of a molecule's structure, they are fundamentally limited by the availability of chemical data, since each property included in the fingerprint must be included for every molecule in the dataset. In general, there is a lack of experimental chemical data suitable for machine learning. Note that Morgan fingerprints do not contain explicit property information beyond an approximate encoding of the molecular graph, allowing them to be easily applied to any molecule, contributing to their widespread use.

[0027] Model-based fingerprints rely on machine learning to generate these vector representations and include two classes: deep neural networks (DNNs) and graph neural networks (GNNs). GNNs capture molecular structure by operating directly on molecular graphs but are computationally limited in their ability to capture long-range interactions within molecules. A molecular graph (i.e., a chemical graph) is a representation of the structural formula of a molecule. The graph can contain vertices corresponding to atoms and edges corresponding to bonds. DNNs can be more flexible, but generally treat molecules as text by using string representations as input. The most common of these string representations are SMILES and, to a lesser extent, SELF-referencing Embedded Strings (SELFIES). These representations are typically obtained by traversing the molecular graph in a depth-first search (i.e., an algorithm for visiting all nodes in the graph) and using tokens to represent ring and branched structures.

[0028] Figures 1A-1E illustrate the process of generating a SMILES string for a particular molecule. Figure 1A shows an example molecular graph. Figure 1B shows where the ring structure has been broken so that the molecule can be written as a string. Figure 1C shows highlighting of various components of the molecule. Figure 1D shows the SMILES string corresponding to the highlighting in Figure 1C. Figure 1E shows another SMILES string corresponding to the molecular graph in Figure 1A.

[0029] Certain approaches have represented molecules as text and applied techniques from the field of natural language processing (NLP) to, for example, predict products given reactants. However, while string representations are flexible enough to describe any molecule, they may not be able to capture the molecule's rich three-dimensional structure. For example, in Figure 1E, the fluorine atom F is at a significant spatial distance from the carboxylic acid group C(=O)O, yet is nearly adjacent in the SMILES string. Furthermore, as shown in Figures 1D and 1E, a single molecule can be represented by more than one SMILES string.

[0030] The limited information content of these string representations may explain why previous NLP-inspired models perform poorly for trait prediction tasks.

[0031] Embodiments described herein include encoding a molecule into an embedding space. The embedding space can convey spatial information of atoms or bonds within a molecule. For example, an encoder network can convert a partial or complete bond string and position (BSP) representation, which may include atom coordinates, into a position within the embedding space. As another example, an encoder network can convert a molecular graph representation of a molecule into a position within the embedding space. The molecular graph representation can include angles and distances of atoms or bonds within the molecule, possibly relative to other atoms.

[0032] II.A. Combined String and Positional Molecular Representations Thus, in some embodiments of the present invention, three-dimensional information of a molecule can be represented through a bond string and position (BSP) molecular representation, which simultaneously captures both the chemical configuration (bond string) and three-dimensional structure (bond positions) of any molecule. A BSP molecular representation can be generated using (for example) the structure optimization method of the RDKit, which can identify the three-dimensional coordinates of each atom in the molecule. Other models capable of identifying three-dimensional coordinates can also be used. For example, a connectivity table can be converted into a distance boundary matrix, which can be smoothed using a triangular boundary smoothing algorithm. The smoothed boundary matrix can be used to identify an adapted random distance matrix that can be embedded in three dimensions to identify the three-dimensional coordinates of each atom in the molecule. A coarse force field and boundary matrix can be used to fine-tune the atom coordinates. As another example, instead of using a coarse force field and boundary matrix to refine the coordinates, torsion angle preferences from the Cambridge Structural Database can be used to fine-tune the coordinates. For example, an experimental torsion basic knowledge distance geometry (ETKDG) approach can be used to identify the three-dimensional coordinates of each atom in the molecule.

[0033] Each bond in a molecule can be represented as follows: <first atom><bond type><second atom> (e.g., "C10O" for a carbon atom bonded to an oxygen atom via a single bond), and its corresponding bond position is represented by [<coordinates of first atom>, <coordinates of second atom>]. This representation does not require tokens to specify branches and rings, since this information is inherently present in the coordinates of each bond. That is, the three-dimensional structure of the molecule can be included directly in the model input, instead of requiring the model to learn this structure from a SMILES string.

[0034] Figure 2 shows an example of a BSP representation. The top of Figure 2 shows a SMILES string over a molecular graph. Table 204 shows the BSP representation. The first row of table 204 shows a representation of the bonds as string tokens. The entries in the same column below the bond representation show the coordinates of the first atom and second atom of each bond. The coordinate of the first atom is denoted by "a" and the coordinate of the second atom is denoted by "b". x, y, z refer to the three-dimensional coordinate system.

[0035] II.B. Representation of Reactant / Reagent Arrays In the BSP representation, bond positions directly capture positional information. Therefore, in contrast to standard transformer-based models, there is no need to use a separate token-level positional embedding to identify bond positions. However, to distinguish between different molecules in a single reaction, a static "molecule position" embedding can be used to indicate which molecule in the reactant / reagent sequence the bond corresponds to. Thus, the unique position of any bond in a reaction sequence can be defined by its bond position and its molecular position.

[0036] Figure 3 shows an example of constructing a BSP representation of reactants / reagents from SMILES strings. The bond strings, shown in the left column, list the bonds in each molecule. The molecular positions in the middle column indicate which molecule the bond belongs to, allowing the model to distinguish one molecule from another. The third column, the bond vector, contains the coordinates of the bond in three-dimensional space.

[0037] II.C. Transformer Model for Generating Molecular Fingerprints The BSP representation of a molecule can be used as input to an encoder network to convert the representation into an embedding representation in an embedding space. The encoder network can be pre-trained by training a machine learning model including the encoder network to perform a given task. The machine learning network can include a Transformer network including a Bidirectional Encoder Representations from Transformers (BERT) model. The given task can include predicting properties of masked bonds in a molecule. For example, a BERT model can be trained to predict missing bond tokens from an incomplete initial bond string representation of a molecule.

[0038] The dimensionality of the embedding space may be less than the dimensionality of the BSP representation. The embedding space may be a high-dimensional embedding space having at least 3 dimensions, at least 5 dimensions, at least 10 dimensions, at least 20 dimensions, at least 30 dimensions, or at least 50 dimensions. Alternatively or additionally, the embedding space may have fewer than 5 dimensions, fewer than 10 dimensions, fewer than 20 dimensions, fewer than 30 dimensions, fewer than 50 dimensions, or fewer than 70 dimensions. Within the embedding space, structurally similar molecules may be separated by short distances, while molecules lacking structural similarity may be separated by long distances. The BSP representation input to the Transformer model may include one, two, or three of the following embeddings:

[0039] 1. The standard learnable embedding used for each token in the combined string, 2. A bond position embedding, obtained by projecting the bond positions onto the same dimensions as the bond embedding, which encodes the positions of the bonds within the molecule; and 3. Molecular position embedding, which is the position embedding used in transformer-based models. All bonds belonging to the same molecule have the same molecular position embedding.

[0040] For example, Figure 4 shows a token embedding (which can then be fed into a Transformer model) that includes a combination of all three of the listed embeddings. The bond embedding is determined from the bond string. The bond position embedding is obtained from the bond vector using a neural network layer (e.g., MLP [Multi-Layer Perception]). The molecule position embedding is obtained from the molecule position. Item 404 shows the static sinusoidal embedding, which is a constraint vector that helps distinguish different molecules in this embodiment. The bond embedding, bond position embedding, and molecule position embedding make up the token embedding.

[0041] II.D. Pretraining the Transformer Model Pre-training a Transformer model as a variational autoencoder can generate fingerprints such that more structurally similar molecules have fingerprints that are closer to each other. These fingerprints can then be used for a diverse range of tasks, from thermodynamic property prediction and toxicity classification, to achieve state-of-the-art performance. This model may outperform some other models in property prediction.

[0042] Bond string and position (BSP) molecular representations can directly identify information about the complete three-dimensional structure of a molecule. BSP molecular representations can be used to train machine learning models (e.g., transformer-based models). For example, a model can be trained to predict "missing" (or "masked") bonds artificially removed from each representation based on the rest of the representation. That is, given the location of an unknown bond, the model is trained to predict the correct bond token by examining neighboring bonds in three-dimensional space.

[0043] The training dataset can include 3D representations of molecules. For example, the training dataset can include unique molecules from the MIT_USPTO dataset, which contains hundreds of thousands of chemical reactions excised from US patents for a total of approximately 600k unique molecules. Because a single molecule can have multiple conformers, the 3D representation of each molecule is not unique, so multiple molecular representations with different coordinates can be generated. This acts as a data augmentation routine and helps reduce overfitting for downstream tasks.

[0044] Figure 5 shows the overall process of training a masked binding model. Some of the tokens from the binding string representation of a molecule can be selected (e.g., randomly selected) and replaced with the [MASK] token. The corresponding binding positions of all mask tokens can be left unchanged. The masked binding strings and binding positions can be fed into an encoder network (e.g., a BERT encoder) of a transformer model. The model loss is then calculated using only the predictions at each masked position.

[0045] The masked input BSP representation and predicted unmasked BSP representation for the example of Figure 5 are as follows: Input: C10O N10O [MASK]C10N [MASK]C10C C15C C15C C15C C10O [MASK]C15N C15C Output: C10O N10O C10N C10N C20O C10C C15C C15C C15C C10O C15N C15N C15C The Transformer model can include an encoder network and a decoder network. Thus, pre-training the Transformer model can include training the encoder network to learn how to transform BSP molecular representations into an embedding space, and training the decoder network to learn how to transform data points in the embedding space into the corresponding BSP molecular representation or another representation that identifies the structure of the molecule, such as a Simplified Molecular-Input Line-Entry System (SMILES) representation.

[0046] Figure 6 shows an example of these transformations (while simplifying the dimensionality in the embedding space). In the example shown, the embedding of benzene (point 604) is separated from the embeddings of methanol (point 608) and ethanol (point 612). The example shown also shows a decoder (via point 616 and the line leading to point 616) that transforms a given data point in the embedding space into a predicted molecule (corresponding to isopropyl alcohol).

[0047] II.E. Graph Neural Networks for Generating Molecular Fingerprints Rather than generating molecular fingerprints using an encoder network trained within a Transformer model, fingerprints can be generated using a graph neural network (GNN). Molecules can be interpreted as molecular graphs, with atoms as nodes and bonds as edges. Under such a representation, a GNN can be used to obtain an embedding of the molecule. A typical GNN may include multiple graph convolution layers. To update node features, the graph convolution layers can aggregate the features of neighboring nodes. There are many variations of graph convolution. For example, message passing layers can be particularly expressive and enable the incorporation of edge features that are important for molecular graphs.

[0048] Figure 7 shows a representation of the message passing layer. Each node can collect messages from its neighbors. A node can receive messages from Xi and X j The edge information E ij Messages containing X j From X i The message value may be sent to ij Then, the message M ij can be aggregated using a permutation invariant function such as the mean or sum. The aggregated messages are

number

number

[0049] II.F.GNN Pre-Training Similar to Transformers, GNNs can be pre-trained on unlabeled molecules. Two methods for GNN pre-training include atom masking and context prediction.

[0050] In atom masking, some nodes are selected (e.g., using random or pseudorandom selection techniques) and replaced with mask tokens. Then, a GNN can be applied to obtain the corresponding node embeddings. Finally, a linear model is applied on top of the embeddings to predict the labels of the masked nodes. Figure 8 shows the method for atom masking pre-training. "X" indicates a masked node in the molecular graph. A GNN is used to obtain the identities of the masked nodes.

[0051] In context prediction, for each node v, the neighborhood graph and context graph of v can be defined as follows. The K-hop neighborhood of v includes all nodes and edges that are at most K hops away from v in the graph. This is motivated by the fact that a K-layer GNN aggregates information over the K-hop neighborhood of v, and thus the nodes that embed h(K)v depend on nodes that are at most K hops away from v. The context graph of node v represents the graph structure surrounding the neighborhood of v. The context graph can be described by two hyperparameters r1 and r2, and the context graph can represent a subgraph that is between r1 hops and r2 hops away from v (i.e., a ring with width r2 - r1). A constraint r1 < K can be implemented so that some nodes are shared between the neighborhood graph and the context graph, and these nodes can be called context anchor nodes. The constraint may include that K is 2, 3, 4, 5, 6, 7, 8, 9, or 10. These anchor nodes provide information about how the neighborhood graph and context graph can be connected to each other.

[0052] Figure 9 shows the molecular graph and subgraph of the K-hop neighborhood around the central node. K is the number of hops allowed between a given node and its adjacent nodes, and if two nodes are connected by a path with K or fewer edges, they are considered adjacent nodes. It is set to 2 in the figure. The context is defined as the surrounding graph structure between r1 hops and r2 hops from the central node, and r1 = 1 and r2 = 4 are used in the figure. Figure 9 shows a subgraph of the K-hop neighborhood that includes nodes within 2 hops of the central node. Figure 9 also shows the context subgraph created by extracting the portion of the molecule that falls between r1 and r2. Box 904 represents the embedding vector of the K-hop neighborhood subgraph, and box 908 represents the embedding vector of the context graph. Dots represent vector dot product similarity.

[0053] Table 1 shows the performance metrics of GNNs generated using different datasets, evaluation metrics, and pretraining types. In general, the performance metrics associated with atomic mask pretraining slightly outperformed those associated with context prediction pretraining. [Table 1]

[0054] II.G. Directional Variational Transformer To predict molecular properties, a molecular graph can be represented as a fixed-size latent vector so that the molecule can be reconstructed from the latent vector as a SMILES string. The size of the latent vector is a hyperparameter and can be determined empirically. This latent vector can then be used to predict the molecular properties. FIG. 10 shows a molecular graph 1004 as an example. The molecular graph can be represented using BSP. The molecular graph 1004 can have a latent representation 1008. The latent representation 1008 can be used for property prediction 1012. Furthermore, a SMILES string 1016 can also be reconstructed from the latent representation 1008.

[0055] One encoder-decoder architecture that can be used to generate latent representations for feature prediction is the Directional Variational Transformer (DVT). In a DVT, the encoder network may be a graph-based module. The graph-based module may include a graph neural network (GNN). The graph-based module may consider the distance between atoms and the spatial direction from one atom to another. DimeNet (github.com / gasteigerjo / dimnet) is an example of a GNN that considers both atoms and the spatial direction from one atom to another. Like DimeNet, a DVT can embed messages passed between atoms rather than the atoms themselves. The message-passing layer is described in Figure 7. The decoder network may take the latent representation from the encoder network as input and then generate a SMILES representation.

[0056] Selected differences between DVT and Variational Transformers (VTs) (e.g., the model trained as a Variational Autoencoder in Section II.D) include:

[0057] 1. The DVT model uses relative positions based on angles and distances instead of positions based on absolute coordinates.

[0058] 2. Unlike VT, DVT does not perform global attention. Instead, DVT focuses only on bonds that either share a common atom or neighboring bonds within a certain threshold distance. The threshold distance can include bonds within a maximum number of hops (e.g., all nodes at most two edges away) or a physical distance (e.g., 5A).

[0059] 3. DVT uses a separate readout node that aggregates the molecular graph to generate a fixed-size latent vector.

[0060] The DVT model can be used in place of another Transformer model. For example, the DVT model may be pre-trained and trained similarly to the variational autoencoder model described in Section II.D. Furthermore, aspects of the Transformer model described using the DVT model can also be applied to other Transformer models.

[0061] II.G.1. Relative Position Generation For the DVT encoder network, the input dataset can represent a molecule in a way that identifies a sequence of bonds with the positions of its atoms in 3-D space. The positions of the atoms can be used to calculate the relative distance and angle between two bonds.

[0062] Figure 11 shows the relative positions in terms of angles and distances. Figure 11 shows four atoms: A1, A2, A3, and A4. A1 and A2 are connected by a bond, and A3 and A4 are connected by a bond. B1(A1,A2) represents the bond between atoms A1 and A2, and B2(A3,A4) represents the bond between atoms A3 and A4. The distances between atoms and bonds are shown in the figure. The distance between two bonds is represented by [d1,d2], where dn = d(x,y) is the Euclidean distance between atoms x and y. For d1, x is A2 and y is A4. For d2, x is A2 and y is A3. The angle between two bonds is calculated as [a1,a2], where a1 = a(x,y,z) is the angle between the xy line and the yz line. For example, for a1, x is A1, y is A2, and z is A3. For a2, x is A1, y is A2, and z is A4.

[0063] The degree of an atom in a bond can be defined by its degree of occurrence during a depth-first search through the molecular graph of the molecule. Because molecular graphs created from standard SMILES are unique, the model can learn to generalize the degree. Generalizing the degree can refer to the model learning how to generate a standard ordering of SMILES during training, such that when a BSP of a molecule not seen during training is input, the model can output the standard SMILES of that molecule. Once the degree of an atom in a bond is fixed, a second atom can be selected to calculate distances and angles. For example, Figure 11 shows the distance and angle from atom A2 rather than atom A1, even though both A1 and A2 are in the same bond. Alternatively, the distance and angle can be calculated from the first atom in the bond (e.g., A1). The model's performance can remain high regardless of whether the angle and distance calculations are performed using the first atom in the bond or the second atom in the bond.

[0064] II.G.2. Encoder A graph representation of a molecule may be input to the encoder network. The graph representation may identify spatial relationships between atoms and / or spatial properties related to bonds between atoms. In some embodiments, the graph representation may include a representation of the molecule in two dimensions. In some embodiments, the graph representation may include a representation of the molecule in three dimensions. The encoder network may generate a fixed-size latent vector as output. The encoder network may include multiple heads. For example, the encoder network may include two heads: a graph attention head and a read attention head. Other heads that may be used include a standard attention head or an OutputBlock head, similar to those used in DimeNet.

[0065] a.Graph Attention Head The graph attention head performs attention-based message passing between nodes to provide relative angle and distance information. As shown in FIGS. 12A and 12B, message passing may include two steps. In step 1, nodes may send messages to their neighboring nodes, and a message from node i to node j is a combination of node i's embedding and node i's relative position with respect to node j, as shown in FIG. 12A. The message includes the angle from E_t to another node (e.g., A_t i) and the distance from E_t to another node (e.g., D_t i). In some embodiments, every node may send messages to its neighboring nodes. The number of message passing steps may be the same as the number of graph attention layers in the encoder network. For example, the number of graph attention layers may be 2, 3, 4, 5, 6, 7, 8, 9, 10, or more. In step 2, each node may aggregate its incoming messages using an attention mechanism, as shown in FIG. 12B. Each node may generate a key-value pair (i.e., k and v) for every incoming message. Each node can also generate a query vector (i.e., q) from its own embedding. Using the query and key, each node can calculate an attention score by performing a vector dot product between the query vector and the key vector for all incoming messages. The updated embedding of the target node may be generated based on a weighted average of the value vectors of all incoming edges. An example of a weight is W([A_ti, D_ti]), which is a vector containing directional information between bonds t and i in Figure 12A. If the attention is between bonds that share a common atom, A_ti and D_tj are scalar values ​​representing a single angle and distance value. If the attention is between bonds within a certain threshold distance, A_ti and D_tj are vectors of size 2 and can be calculated as shown in Figure 11.

[0066] In this example, attention scores are used as weights. For example, a set of embeddings may be generated, each of which represents the bond angles or distances in one or more molecules (e.g., such that the set of embeddings corresponds to a single ionic liquid molecule or a combination of a single ionic liquid molecule and a target molecule). For each particular embedding, a set of key-value-query pairs is generated, corresponding to the pairing between the particular embedding and each embedding in the set of embeddings. The attention mechanism (see Vasawni et al., "Attention is all you need," 31 st A method for embedding a bond angle or bond distance (such as that described in the Conference on Neural Information Processing Systems (2017)) can be used to determine how much to weight embeddings of different bond angles or bond distances when generating an updated embedding corresponding to a given bond angle or bond distance.

[0067] b. Read attention head The read attention head may be a separate head used in the encoder network. The read head can aggregate all node embeddings to generate a fixed-size latent vector using the attention mechanism. A read node may be used to aggregate all nodes. The read node may be a single node that is connected to all other nodes but is excluded from the message passing mechanism. The read node is R in Figure 13. The read node may serve as a query vector and may have embeddings learned during training. Key-value pairs (i.e., k and v) from all nodes in the graph may be generated. The query and keys may be used to calculate an attention score for all nodes in the graph. If the attention score is a weight, a weighted aggregation of all value vectors may be performed.

[0068] II.G.3. Decoder The fixed-size latent vectors generated by the encoder network can be input to a decoder network, which can generate a SMILES representation of the molecule as output. The decoder network can be trained to learn to convert data points in the embedding space into the SMILES representation.

[0069] III. Fine-tuned networks for molecular property prediction Because the Transformer model includes an encoder network, pre-training the Transformer model can result in the encoder network being pre-trained (so that it can convert the BSP representation of the molecule into a representation in the embedding space).

[0070] The encoder network can then be fine-tuned so that the BSP representation can be converted into predicted values ​​for a particular property of a sample of molecules (e.g., viscosity, density, solubility in a given solvent, activity coefficient, or enthalpy). For example, a classifier or regressor can be attached to the output of the encoder network, and the classifier or regressor can be fine-tuned for a particular task. Figure 14 shows an example model including an encoder network containing a linear layer with an activation function (e.g., a Softmax function). This linear layer can be replaced with a classifier or regressor and fine-tuned to convert the embedding values ​​produced by the encoder network into predicted property values.

[0071] The model shown in FIG. 14 further includes a multi-head attention layer. The multi-head attention layer may use query, key, and value matrices to determine the degree to which to "attend" to each of a set of elements in a received dataset (e.g., a BSP representation of a molecule) when processing another of the set of elements. More specifically, each element may be associated with three matrices: query, key, and value. For a given element, a query may be defined to capture which type of key should be addressed. Thus, the query may be compared (via dot product) with the key of the same element and the key of each other element to generate weights for the pair of elements. These weights may be used to generate a weighted average of the element values. The multi-head attention layer may perform this process multiple times using different "heads." Each head may address different types of data (e.g., data from nearby elements, rare values, etc.). The model shown in FIG. 14 further includes an "add&norm" layer. This residual connection performs element-wise summation and layer normalization (across feature dimensions).

[0072] IV. Reaction Prediction Although a fine-tuned encoder network (comprising an encoder network and regressors, classifiers, and / or activation functions) can generate property predictions for individual molecules, encoder networks do not provide a principled way to process mixtures of molecules, such as ionic liquids composed of distinct cations and anions. Therefore, separate reaction prediction models can be generated that use the same BSP representation described above but employ different encoder / decoder model architectures.

[0073] An exemplary architecture of a reaction prediction model is shown in Figure 15. The model includes an encoder network 1504 and a decoder network 1508. The bond positions, bond embeddings, and molecule positions are inputs to the encoder network 1504. The output of the encoder network 1504 is input to the decoder network 1508. The embeddings of the two molecules are also input to the decoder network 1508. The output of the model in Figure 15 is a probability for the output of the chemical reaction.

[0074] Reaction prediction can be treated as a machine transformation problem between the BSPs of reactants and reagents and a product representation (e.g., a SMILES, SELFIES, or BSP representation). A SMILES representation of the product may be advantageous over a BSP representation because it is important to convert the BSP representation into a human-readable chemical formula. A SMILES representation of the product may be even more advantageous than a SELFIES representation because it may be more difficult for the model to infer ring and branch sizes when using the SELFIES representation.

[0075] The training task for a separate reaction prediction model can be learning to predict products from reactants. Figures 16A and 16B show some successful examples of the model. Each row contains a reactant and a reagent, then a ground truth product and a predicted product formed. The predicted product matches the ground truth product. Figures 17A and 17B show some failure examples of the model. In Figures 17A and 17B, the predicted product does not match the ground truth product. One observed failure mode of the model is copying one of the molecules from the reactants or reagents instead of performing the reaction. This suggests that the failure mode is predicting no reaction from the reactants and reagents.

[0076] IV.A. Illustrative Results of Response Prediction Models IV.A.1. DeepChem Task To evaluate the BSP representation for molecular property prediction, a BERT masked bond model was trained on molecules from the MIT_USPTO and STEREO_USPTO datasets. The STEREO_USPTO dataset includes stereochemistry (i.e., bond orientation in the molecule) in its SMILES description, which makes overall encoding / decoding more difficult. Using STEREO_USPTO, the model predicts both graph structure and bond orientation. The BSP input to the model was constructed as described in Figure 2. Because only one molecule exists in each sample, molecular position was not used. A classifier or regressor was then fitted on top of the pre-trained BERT model and fine-tuned for downstream classification or regression tasks. Results for 10 different datasets from the DeepChem evaluation suite are shown in Table 2. Different tasks represent different properties of molecules. For example, in the Tox21 dataset, 12 tasks means that each molecule has 12 properties for the model to predict. Validation scores for the USPTO dataset are shown in Table 3. [Table 2] [Table 3]

[0077] Table 2 shows that different datasets and tasks can yield high RMSE or AUC evaluation metrics. Table 2 shows the results for predicting molecular properties. Table 3 shows that the validation accuracy with the MIT_USPTO dataset is approximately 86% and with STEREO_USPTO is approximately 66%. The validation scores for the USPTO dataset are for predicting the product of a reaction.

[0078] IV.A.2. Prediction of Ionic Liquid Properties A pure binary ionic liquid (IL) is composed of two molecules, i.e., a cation and an anion. A mixed ionic liquid can have more than two components. Because masked language models are trained on single molecules, it is not useful to obtain the embedding of an ionic liquid. Therefore, the encoder network of the reaction prediction model can be used to obtain the embedding of an ionic liquid. To distinguish between pure binary ILs and mixed ILs, the component number was used as an additional input after obtaining the embedding. Table 4 shows exemplary validation scores for two properties (density and viscosity). The number of components refers to the number of ionic liquids. For example, if the number of components is 2, there are two ionic liquids containing two cations and two anions. [Table 4]

[0079] IV.A.3. Analysis of Attention Weights To understand the model's ability to associate different bonds in three-dimensional space, we visualized the attention weights of a reaction prediction model trained on the MIT_USPTO dataset. There are three types of attention: self-attention to reactants / reagents, self-attention to products, and attention between reactants / reagents and products.

[0080] The reactant / reagent self-attention weights were extracted from the model's encoder module, which takes the BSP representation as input. Figure 18 shows the attention weights of the fourth head of the encoder in its third layer. As shown, the model learned to find neighboring bonds in three-dimensional space using bond locations evident from the diagonal nature of the attention map. Bond strings were generated by traversing the molecular graph, as shown in Figure 19. Thus, some of the neighboring bonds appeared far apart in the bond string sequence, resulting in the wing-like pattern seen in the attention map.

[0081] IV.B. Example Results Using Directional Variational Transformers FIG. 20 shows an exemplary architecture of a reaction prediction model using DVT. The main component is a graph-based encoder network 2004. A molecular graph can be input to the encoder network 2004. The encoder network 2004 can then output a fixed-size latent vector. The encoder network 2004 may use relative positions based on angles and distances instead of positions based on absolute coordinates. The decoder network 2008 can be a decoder network from a Transformer model (e.g., as shown in FIG. 15). The fixed-size latent vector can be input to the decoder network 2008. The decoder network 2008 can then generate a SMILES representation of the molecule. As shown in FIG. 20 but not in FIG. 15, the encoder network 2004 is configured to receive angle and distance identifications as input. In addition, the encoder network 2004 (shown in FIG. 20) is not configured to receive bond positions or molecular positions (e.g., coordinates in BSP) as input. However, in some embodiments, the molecular graph may be interpreted as a type of BSP representation in which the nodes of the graph are connected string tokens. The joint position vectors (coordinates) can be converted into angle and distance values, which are then passed to the graph attention head.

[0082] Table 5 shows the performance data for the variational transformer (VT) model and the directional variational transformer (DVT) model. DVT had higher smoothness test results than VT for 100,000 samples. The smoothness test indicates the percentage of randomly generated latent embeddings that yield valid molecules when decoded using the decoder model. A higher smoothness test value indicates a better, smoother latent space. VT had higher reconstruction accuracy than DVT. Reconstruction accuracy is the percentage of molecules from the validation dataset that are perfectly reconstructed by the model. A higher reconstruction accuracy indicates better performance of the encoder / decoder model. DVT required fewer average iterations than VT to find the target ionic liquid. In this experiment, latent embeddings were obtained for all ILs, and then discrete Bayesian optimization was performed to find how many iterations it took to find the IL with the lowest viscosity score. Fewer iterations are desirable. DVT required fewer dimensions to compress the ionic liquid. In this experiment, latent embeddings are obtained for all ILs, and then the embedding vectors are compressed to fewer dimensions using PCA so that 99% of the variance is preserved. DVTs can represent molecules in a lower dimensional space than VTs, which allows for greater computational efficiency. Table 5 shows that DVTs can be advantageous over VTs in the number of iterations required to find the target ionic liquid and the number of dimensions required to compress the ionic liquid. [Table 5]

[0083] V. Searching for candidate molecules using computer simulations and variational autoencoders Material / compound selection is a persistent and long-standing problem in many areas of materials science / chemical synthesis, which is primarily time- and resource-constrained, given the lack of widespread use of high-throughput experiments (HTE). One exemplary type of material selection is identifying ionic liquids that efficiently and effectively depolymerize specific polymers (e.g., to facilitate recycling).

[0084] Some embodiments relate to using artificial intelligence and computing, at least in part, to replace and / or supplement wet-lab approaches to select materials / compounds that are suitable for a given use case in a manner that requires a relatively small amount of time and / or a relatively small amount of resources. While experimentally evaluating millions of ionic liquid options would take years or perhaps decades, AI and computational models provide the ability to do so on a practical timescale.

[0085] In some embodiments, a database can be generated containing predicted liquid viscosity and solubility properties for each of a set of ionic liquids. Solubility properties can include enthalpy of mixing and activity coefficients. Solubility properties can be associated with a given ionic liquid and a specific compound (e.g., a specific polymer). These predicted properties can be generated by running simulations, including COSMO-RS (Conductor-like Screening model for Real Solvents) simulations based on quantum / thermodynamic methods. These simulations can allow for screening and / or filtering of one or more existing ionic liquid libraries for specific IL and IL / polymer solution properties. Molecular dynamics and density functional theory (DFT) calculations can also be performed. COSMO-RS simulations post-process quantum mechanical calculations to determine the chemical potential of each chemical species in solution, from which other thermodynamic properties (e.g., enthalpy of mixing and / or activity coefficients) can be determined. The quantum mechanical and / or thermodynamic methods can include COSMO-RS or DFT.

[0086] Although running a COSMO-RS simulation may be faster than running a wet-lab experiment, COSMO-RS simulations can use a screening charge density as input. Because this charge density is obtained from computationally time-consuming density functional theory (DFT) calculations, this can slow down COSMO-RS simulations for compounds that do not have pre-calculated DFT results.

[0087] Running the COSMO-RS simulation may include: Enumerate ionic liquid combinations by taking ionic liquid pairs (cations and anions) from a database and evaluating all possible cation-anion combinations. Databases can include the COSMO-RS ionic liquid database ADFCRS-IL-2014, which contains 80 cations and 56 anions with pre-calculated density functional theory (DFT) results (corresponding to 4,480 different ionic liquids) and / or COSMObase, which contains 421 cations and 109 anions with pre-calculated DFT results (corresponding to 45,889 unique ionic liquids).

[0088] Solute (polymer) identification: The solute can be associated with stored pre-calculated DFT results for the central monomer of the trimer, and potentially also for several central monomers of the oligomer for some polymers.

[0089] To replicate the example experimental setup for depolymerization using ionic liquids, set up a COSMO-RS calculation using specific mole fractions (e.g., 0.2 for polymer, 0.4 for ionic liquid cation, 0.4 for ionic liquid anion, and 0.8 for ionic liquid when cation and anion are used together). Temperature and pressure can be set or kept to the values ​​used by default in COSMO-RS calculations.

[0090] ● Obtain the solution mixing enthalpy and activity coefficient of the polymer in the ionic liquid solution from the COSMO-RS output and store the results in a database.

[0091] The predictions can be used to perform a simple screening of IL cation and anion pairs to select an incomplete subset of ionic liquids with the lowest predicted solution mixing enthalpies and activity coefficients. The viscosity of the incomplete subset of ionic liquids at depolymerization temperatures and pressures can be verified to be reasonable for mass transfer (e.g., directly from the ILThermo database or from predictive models based on transformer embedding). The subset can be further filtered based on the depolymerization temperatures and pressures that can be identified.

[0092] Thus, in some embodiments, these simulations may be performed on only an incomplete subset of the ionic liquid population, and one or more property prediction models may be used to predict solubility properties for other ionic liquids. Each of the property prediction model(s) may be trained and / or fitted using representations of the ionic liquids in the incomplete subset and corresponding values ​​for a given property. For example, a first model may be defined to predict the enthalpy of mixing (corresponding to a particular compound [e.g., a particular polymer]), a second model may be defined to predict the activity coefficient (corresponding to a particular compound), and a third model may be defined to predict the IL viscosity. The representation of the ionic liquid may include a BSP representation of the ionic liquid. The model may be trained on an embedding space to relate molecular embeddings to property values. Each of the property prediction model(s) may include a regression model.

[0093] One or more generative models can be used to explore the embedding space (e.g., where the initial molecular representations are transformed using an encoder network) and interpolate between the representations of candidate ionic liquids to discover desired enhanced IL and IL / polymer solution properties for the polymer degradation reaction. That is, the generative model can explore additional molecules and map the embeddings of these molecules to values ​​of one or more properties. The generative model(s) can include (for example) a regression model. The generative model can be used to generate continuous predictions across the embedding space.

[0094] For example, Figure 21 shows how individual points in the embedding space may correspond to a given chemical structure and predicted properties (which may be predicted using a reconstruction decoder). In an embodiment, a generative model may refer to a model that produces an output and identifies molecules (e.g., ionic liquids) not yet known to be useful in a particular context (e.g., depolymerization). The predicted properties may be identified using another model (e.g., a regression model on top of the generative model that converts a given location in the embedding space to a predicted property).

[0095] In some cases, a single model (e.g., a single generative model with a regression task) can be configured to generate multiple outputs for any given location in the embedding space, where the multiple outputs correspond to multiple properties. For example, a single model may generate predicted mixing enthalpy, activity coefficients, and viscosity for a given location in the embedding space. In some cases, multiple models are used, where each model generates a prediction corresponding to a single property. For example, a first model can predict mixing enthalpy, a second model can predict viscosity, and a third model can predict activity coefficients.

[0096] In the embedded space representation shown in FIG. 21 , data points 2104, 2108, and 2112 are shown. These data points are associated with measured values ​​of properties, while data point 2116 is not associated with a measured property value. In the example shown, a benzene molecule (represented by data point 2104) is predicted to have unfavorable properties, while a methanol molecule (represented by data point 2112) and an ethanol molecule (represented by data point 2108) are predicted to have favorable properties for a particular polymer degradation reaction. After predicting the properties in the embedded space using a regression model, a promising molecular representation (represented by data point 2116) is identified. Data point 2116 is close to data points 2108 and 2112 in the embedded space. Assuming that the molecules represented by data points 2108 and 2112 have favorable properties, the proximity of data point 2116 suggests that the molecule represented by data point 2116 also has favorable properties. "Promising molecules" can correspond to locations in the embedding space where the conditions associated with each predicted property are met.

[0097] For example, a promising molecule can be associated with a predicted viscosity below or above a viscosity threshold, a predicted activity coefficient of a polymer in a solution of the promising molecule below an activity coefficient threshold, and a predicted enthalpy of mixing for a solution of the promising molecule and polymer below a mixing enthalpy threshold. That is, the predictions generated by the generative model(s) can be used to identify locations within the embedded space that correspond to desired properties of interest (e.g., predicted viscosities below / above a predetermined viscosity threshold, enthalpies of mixing below a predetermined enthalpy threshold, and / or activity coefficients below a predetermined activity coefficient threshold).

[0098] As another example, a score may be generated for each of some or all of the ionic liquids represented by positions in the embedding space. The score may be defined as a weighted average based on predicted target properties. The values ​​of the predicted target properties may be normalized. A higher score may indicate a more promising molecule. By way of example, the score may be defined as a weighted average of the predicted normalized viscosity (including logarithmic viscosity), or the negative of the predicted normalized viscosity (including logarithmic viscosity), the negative of the predicted normalized enthalpy of mixing, and / or the negative or positive of the predicted normalized activity coefficient (including logarithmic activity coefficient). One or more promising ionic liquids may be identified as corresponding to the n highest scores (e.g., as corresponding to the highest score or as corresponding to any of the top 10 scores) or as corresponding to a score above an absolute threshold. Alternatively, one or more promising ionic liquids may be identified as corresponding to the n lowest scores (e.g., as corresponding to the lowest score or as corresponding to any of the bottom 10 scores) or as corresponding to a score below an absolute threshold.

[0099] In some cases, a variational autoencoder is used to transform representations (e.g., BSP representations) of data points to predict the molecular structure and / or identity corresponding to the data points. The variational autoencoder may be configured with an encoder network that converts a representation of a molecule (e.g., a BSP representation) into an embedding spatial distribution (e.g., the mean and standard deviation of the embedding representation) and a reconstruction decoder network configured to convert the representation sampled from the embedding spatial distribution back into an initial spatial representation of the molecule. In some examples, the reconstruction decoder network is configured to generate a SMILES representation of the molecule instead of a BSP representation (because a SMILES representation may be more interpretable to a human user). To calculate the loss for a given prediction, the initial BSP representation may be converted to a corresponding initial SMILES representation. The penalty may be scaled based on the difference between the initial SMILES representation and the decoder-predicted representation. Alternatively, the generated SMILES representation may be converted to a BSP representation, and then the penalty may be scaled based on the difference between the initial BSP representation and the final BSP representation.

[0100] The trained decoder network from the variational autoencoder can then be used to convert interest points (e.g., associated with predictions that satisfy one or more conditions) into SMILES representations, thereby identifying molecules. Interest points may correspond (for example) to local maxima or absolute maxima of given predicted values, local minima or absolute minima of given predicted values, a given local maxima or absolute maxima, or a score that depends on multiple predicted values. For example, scores may be defined to be negatively correlated with the mixing enthalpy prediction, the activity coefficient prediction, and the viscosity prediction, and the interest point may correspond to the highest score.

[0101] 22 illustrates an exemplary workflow for identifying candidate molecules for a reaction (e.g., depolymerization) using computer simulation and a variational autoencoder. In step 2204, a data store containing one or more properties and structures of molecules is created and / or accessed. For example, the one or more properties may include viscosity, solubility, enthalpy, and / or activity coefficients. To identify these values, a COSMO-RS calculation in step 2208 may be accessed corresponding to a compound that already has pre-calculated quantum mechanical (DFT) calculations.

[0102] In step 2212, one or more models (including, for example, generative models and regression models) can be defined to relate molecular representations to predicted molecular properties across space (using the properties and molecular representations, such as the BSP representations described above). Each model can be specific to a polymer system to explore the space of various molecules corresponding to ionic liquids that can be used to depolymerize the polymer. The property values ​​predicted by the model can influence whether a solution of a given type of molecule (associated with the independent variable position in the embedded space) is predicted to depolymerize the polymer.

[0103] In step 2216, the model(s) can thus generate, for each location in the embedded space, one or more predicted properties of the molecule (or solution of molecules) corresponding to the space.

[0104] Generative and regression models can be used to identify one or more regions (e.g., one or more locations, one or more areas, and / or one or more volumes) in the embedded space that correspond to desired properties. What constitutes a "desirable" property can be defined based on input from a user and / or default settings. For example, a user may be able to adjust a threshold for each of one or more properties within an interface. The interface can update to indicate the number of molecules that meet the threshold criteria.

[0105] In step 2220, a decoder network (trained in a corresponding variational autoencoder 2218) can be used to convert data points in one or more regions of the embedding space into a structural representation of the molecule (e.g., a SMILES or BSP representation of the molecule). Thus, the decoder network of the variational autoencoder can be trained to reliably convert data points from the embedding space into a space that unambiguously conveys molecular structure. Given that the output from the generative model and regression model(s) can identify locations in the embedding space(s) that have properties of interest, one or more candidate molecules of interest can be identified using the variational autoencoder.

[0106] Properties of one or more candidate molecules may then be experimentally measured in step 2224. Such experiments may confirm their selection as molecules of interest, or may be used to update the embedding space, one or more transformations, and / or one or more selection criteria.

[0107] Example results for VA variational autoencoders A transformer-based variational autoencoder model was trained on a large dataset of chemical molecules containing both charged species (i.e., ionic liquid cations and anions) and uncharged species. The autoencoder model was then used to generate new ionic liquid cations by exploring the space near known ionic liquid cations.

[0108] Figure 23 shows the exploration of the space around a molecule to generate new ionic liquid cations. The starting ionic liquid cation is on the left and is labeled "Charged Start." Random Gaussian noise was added to the embedding. Four ionic liquid cations were then decoded from the embedding. The SMILES strings for each cation are listed below the four resulting ionic liquid cations. All four decoded outputs are charged molecules, although this is not always the case due to the probabilistic nature of the model. Figure 23 shows that a transformer-based variational autoencoder model can be used to generate additional ionic liquid cations from a single ionic liquid cation.

[0109] Figure 24 shows the exploration of the space between two molecules by linear interpolation. The two molecules are the ionic liquid cation in the upper left, labeled "charge start," and the ionic liquid cation in the lower right, labeled "charge end." Again, due to the stochastic nature of the model, not all intermediate molecules need to correspond to charged species. Figure 24 shows seven ionic liquid cations generated in the space between the charge start and charge end. The SMILES strings are displayed below the generated cations. Other interpolation techniques, such as SLERP (spherical linear interpolation), can also be utilized. Figure 24 demonstrates that additional ionic liquid cations can be generated by exploring the space between two ionic liquid cations using a transformer-based variational autoencoder model.

[0110] Similar findings can be made for ionic liquid anions, and by combining different generated cations and anions, new ionic liquids can be generated.

[0111] VI. Optimization of ionic liquid-based depolymerization Chemical recycling involves a reaction process in which waste materials (e.g., plastic waste) are converted back into their molecular components, which can then be used as fuel or feedstock through chemical recycling. Optimized chemical reactions are selected by attempting to maximize conversion (amount of dissolved plastic relative to the amount of starting plastic), yield (amount of monomer relative to the amount of starting plastic), and selectivity (amount of monomer relative to the total amount of reaction products). These decomposition reactions are complex because they rely on intricate coupling between chemical and physical processes spanning multiple time and length scales, meaning that design of experiments (DOE) is expensive and cannot be easily generalized across chemical space.

[0112] Chemical reaction optimization involves maximizing a function (e.g., a utility function) that depends on a set of reaction parameters. While various optimization algorithms have been developed (e.g., convex methods for local minima such as gradient descent, conjugate gradient methods, and BFGS, and non-convex black-box function optimization methods for global optimization such as systematic grid search), they can require many evaluations of the function, making them unsuitable for expensive processes such as laboratory chemical reactions, where each evaluation requires running a new benchtop experiment.

[0113] An alternative approach for chemical reaction optimization uses the technique of Bayesian optimization via Gaussian processes (GP). This technique can achieve good accuracy with limited evaluation. However, Bayesian optimization via Gaussian processes has poor scalability to high parameter dimensions and is difficult to achieve with N 3 It is computationally expensive because it scales computationally as N, where N is the number of training data points used to seed the Bayesian optimization.

[0114] Some embodiments disclosed herein include techniques for Bayesian optimization via Gaussian processes while reducing dimensionality and requiring only a limited amount of training (seed) data. Chemical and experimental space can then be efficiently navigated to identify ionic liquids with favorable properties.

[0115] Applying Bayesian optimization to depolymerization reactions using ionic liquids can be challenging. The number of possible ionic liquids can be large. Ionic liquids contain cations and anions, and there may be thousands of possible cations and anions, or hundreds of commercially available cations and anions. The interactions between the ionic liquid and the polymer being depolymerized may not be fully understood. The mechanism by which the ionic liquid degrades the polymer may not be known. Therefore, Bayesian optimization may be more difficult to apply to depolymerization than in other contexts. For example, depolymerization may be more complex than pharmaceutical reactions, in which a desired drug (or a molecule that may possess some of the desired properties of a drug) is expected to bind to a certain portion of a target (e.g., a pocket in a protein). As a result, Bayesian optimization of depolymerization may be more difficult to reach a convergent solution than Bayesian optimization of pharmaceutical reactions.

[0116] VI.A. Dimensionality Reduction of Ionic Liquids Bayesian optimization can be improved by using a reduced-dimensional space. The structure and / or other physical properties of ionic liquids can be represented by a high-dimensional space (e.g., a BSP or SMILES representation). A reduced-dimensional space with similar molecules located close to each other can aid Bayesian optimization. Matrices can be used to represent the high-dimensional structure of ionic liquids. Dimensionality reduction techniques can then process matrices corresponding to various ionic liquids to identify new sets of dimensions. The new sets of dimensions may capture at least 80%, 85%, 90%, 95%, or 99% of the variability of the dependent variable with the higher-dimensional space. Principal component analysis (PCA) is one possible dimensionality reduction technique. Given a set of points in a high-dimensional space, PCA finds the directions in the space along which the points vary most. By selecting a small number of these directions and projecting the points along them, a lower-dimensional description of each point can be obtained. Within this space, molecules with fingerprint similarity (e.g., structural similarity and / or thermodynamic property similarity) are then located near each other, and molecules lacking fingerprint similarity are located far from each other. This new space may include a relatively small number of dimensions (e.g., less than 30, less than 20, less than 15, less than 10, less than 8, or less than 5). This new space may include any of the embedding spaces described herein, including embedding spaces derived from BSP representations or GNNs (e.g., DVTs).

[0117] As an example, a training set can be defined to include representations of binary ionic liquids with known viscosities. For each binary ionic liquid represented in the training set, the cations and anions can be converted to SMILES, BSP, or other suitable representations. The representations can then be converted to descriptors with a fixed length, such as having less than 10,000 values, less than 7500 values, less than 5000 values, or less than 2500 values. A single feature matrix representation of the ionic liquid can be generated by concatenating the anion and cation descriptors. The feature matrix may be normalized. The feature matrix can then be decorrelated by calculating pairwise correlations between columns. If the correlation analysis detects a correlation between two columns (e.g., by detecting a correlation coefficient above an upper threshold or below a lower threshold), one of the columns can be removed from the feature matrix and stored separately. The dimensionality of the feature matrix can then be reduced to aid in Bayesian optimization. Using dimension reduction techniques (e.g., component analysis such as principal component analysis [PCA]), each feature matrix and other feature matrices can be converted into a set of components and their respective weights (e.g., explained variance). A given number (e.g., a predetermined number such as 4, 5-10, 10-15, 15-20, or 20-30) of components associated with the highest weights (e.g., explained variance) can be used as a dimension-reduced representation of the ionic liquid.

[0118] In some embodiments, the dimensionality of ionic liquids can be reduced by using the embedding spaces described herein instead of PCA. Ionic liquids can be represented using various descriptors or in BSP-based embeddings, which are descriptors that map similar molecules close to each other. Dimensionality can be reduced using techniques such as PCA.

[0119] VI.B. Bayesian Optimization of Reaction Conditions Using Discrete Sampling Within the reduced-dimensional space, Bayesian optimization via Gaussian processes can be used to find individual locations within the space where experimental data are particularly useful in characterizing how a property (e.g., an output such as conversion or yield) varies across the space. For example, each location within the space may correspond to an ionic liquid, and a use case may include determining the conversion and yield of a depolymerization reaction.

[0120] A function is constructed to estimate how the output varies across a reduced-dimensional space using an initial training dataset and data determined by a Gaussian process (GP). The initial training dataset may include measurements of a given property for each of a plurality of ionic liquids. Locations in the reduced-dimensional space corresponding to the plurality of ionic liquids are identified, and a Gaussian process prior is defined using the properties and locations (in the reduced-dimensional space) of the plurality of ionic liquids. A posterior distribution is then formed by calculating the likelihood of an experimental outcome at each of one or more locations by updating the Gaussian process prior with the locations and properties corresponding to the seed experiments. The posterior distribution and a specified acquisition function can be used to construct a utility function, which is a concrete form of the acquisition function in which all optimization hyperparameters are set. A location in the space associated with the highest value of the utility function can be selected, and the ionic liquid and other reaction parameters corresponding to that location can be identified for the next experiment (to collect measurements of the given properties of the ionic liquid and reaction parameters). After experimental data is collected, the posterior distribution can be updated with the measurements, the utility function can be updated, and a different location can be identified for further experimental data collection.

[0121] Figure 25 illustrates the interaction between Bayesian optimization and benchtop experiments. The initial output of the Bayesian optimizer 2504 may provide values ​​for the parameters shown in box 2508. These parameters may include the identity of the ionic liquid (e.g., cation and anion), temperature T, ionic liquid fraction, and solvent fraction. A benchtop (or other scale) experiment 2512 may be performed. The benchtop experiment can provide a yield value (box 2516). This yield is related to the identity of the ionic liquid, temperature T, ionic liquid fraction, and solvent fraction. The yield is used to seed the Bayesian optimizer 2504, and the process is repeated until the Bayesian optimizer 2504 converges or an optimization constraint (e.g., number of iterations) is satisfied. The optimizer can navigate through the reduced dimensionality in the learned embedding created by the model to generate a multidimensional representation of the molecule. The model can recommend entirely new molecules based on the embedding space. The embedding space of the variational autoencoder can be used to seed the optimization routine.

[0122] Figure 26 shows the discrete sampling, x obs The Bayesian optimization process is shown through a GP iterative loop using x. The thicker lines 2604, 2606, 2608, and 2610 indicate loops that start with a seed set of experimental results 2612 that are used to construct the posterior distribution 2616. The green dots (e.g., point 2620) represent the x obs The red points (e.g., point 2624) are all seed experiments (x s ) The trained utilities for all seed experiments (shown in graph 2628) are obs The next experiment at star 2632 is x probe is x probe has the highest utility, so x obs x probe If x has been previously proposed, the posterior distribution 2616 does not change, so the optimization has converged to a solution (box 2636). probeIf , has not been previously proposed, the loop continues by following the thicker lines 2604, 2606, 2608, and 2610. Discrete sampling can be applied in Bayesian optimization even when the values ​​of the parameters can be continuous.

[0123] First, a Gaussian process prior is constructed. A Gaussian process (GP) is a prior that is generated by any finite set of N points {x n ∈X} n=1 N R N It is defined by the property that it induces a multivariate Gaussian distribution on the black-box function f(x) (box 2640). N We assume a GP above, but the prior distribution is x obs It is constructed using only a specific finite set of points called x n =x obs and x obs is discrete. The set of points x obs contains experimentally evaluated points and provides experimental data. N describes the design space and can include parameters such as solvent type, solvent concentration, solvent ratio, and reaction temperature.

[0124] Second, a posterior distribution is constructed (e.g., posterior distribution 2616). The seed experiment is s ,y s} N s=1 and x s and y s is the case of N (i.e., x s is the set of experimental conditions, and y s are the known experimental results for {x s ,y s} N can be thought of as a training set. The black-box function f(x) is s ~Normal(f(x s ),ν) and ν are observable x sA posterior distribution is formed on the function, assuming it is drawn from a GP prior where x is the variance of the noise introduced by . Seed experiments need to be supplied to construct the posterior distribution. Seed experiments are a subset of observations (x s ⊂x obs )

[0125] Third, the next experiment is determined. Then, using the posterior distribution with the specified acquisition function, we calculate the utility function, U(x s ) In an embodiment, the acquisition function may be a GP upper confidence bound (UCB) that minimizes regret over the course of the optimization.

number

[0126] κ corresponds to the degree of exploration versus exploitation (low kappa indicates exploitation, high kappa corresponds to exploration). θ is a hyperparameter of the GP regressor. The following experiment is proposed via proxy optimization:

number

[0127] Fourth, x probe Given a black box function, f(x probe ) as in the example given here. probe f(x probe ) during testing. probe ) is obtained. As in the second step above, the posterior distribution is probe ) may be updated to take into account the results of the next experiment. The third step of determining the next experiment may be repeated for the updated posterior distribution. Then, the new f(x probe ) we can repeat the fourth step of evaluating the black-box function x probex s If σ is within the set of σ, the optimization is considered to have converged and the Bayesian optimization loop may terminate.

[0128] VI.C. Exemplary Methods FIG. 27 is a flow chart of an exemplary process 2700 associated with ionic liquid-based depolymerization optimization.

[0129] At block 2710, process 2700 may include accessing a first dataset including a plurality of first data elements. Each of the plurality of first data elements may characterize a depolymerization reaction. Each first data element may include an embedded representation of a structure of a reactant of the depolymerization and a reaction property value that characterizes the reaction between the reactant and a particular polymer. The embedded representation of the structure of the reactant may be identified as a set of coordinate values ​​within an embedding space. For example, the embedding space may be encoded from a SMILES or BSP representation and / or any embedding space described herein. The embedding space may capture at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% of the relative variance.

[0130] In some embodiments, the embedding space can use principal components determined from PCA. SMILES, BSP, or other representations can be converted into descriptors that can provide characteristic information based on structure. As an example, a moderated descriptor with a fixed length of 1,000-1,500, 1,500-2,000, or 2,000-3,000 can be used. The process feature matrix can then be reduced to remove duplicate and empty entries. In some embodiments, a value of 0 may be assigned to empty entries. A feature matrix can be obtained. The feature matrix can be normalized. The feature matrix can be decorrelated by removing columns with pairwise correlation coefficients greater than a predetermined threshold (e.g., 0.4, 0.5, 0.6, 0.7, 0.8, or 0.9) and / or less than the negative of the predetermined threshold. PCA can be used on the resulting feature matrix. PCA may result in 3, 4, 5, 5-10, 10-15, 15-20, 20-30, or more than 30 principal components (PCs).

[0131] The embedded representation may include properties of reactants in a set of reactants in the depolymerization reaction. The properties of the reactants may include viscosity, activity coefficients, bond types, enthalpies of formation, heats of combustion, or properties derived therefrom. The reactants in the plurality of first data elements may include ionic liquids, including any described herein. In some embodiments, the reactants may also include solvents, including any described herein, and / or the polymer to be depolymerized. A suitable database, such as the National Institute of Standards and Technology (NIST) database for binary ionic liquids with known viscosities, can be queried to provide property information.

[0132] The first data set may be seed data for the Bayesian optimization. The first data elements may include variables associated with describing the depolymerization reaction.

[0133] The reaction property values ​​can characterize an output of the depolymerization reaction. The output can include yield, amount of product, product conversion, selectivity, and / or profit. The output can be unknown before conducting the experiment using the reaction inputs, but can be known after conducting the experiment. The computing device can access a first dataset.

[0134] In some embodiments, the plurality of first data elements may include reaction input values ​​characterizing operating conditions of the depolymerization reaction. For example, reaction input variables may include time, temperature, ratio (e.g., ionic liquid to solvent), quantity, cost, and / or pressure. The plurality of first data elements may also include a representation of a solvent, which may or may not be embedded.

[0135] In some embodiments, the first dataset may be generated from an experiment using candidate molecules determined as described herein, for example, the candidate molecules may be determined using an embedding space and a variational autoencoder as described herein.

[0136] At block 2720, process 2700 may include constructing a prediction function for predicting reaction property values ​​from the embedded representation of the reactant structures. Constructing the prediction function may use the first dataset. For Bayesian optimization, the prediction function may be an objective function, a black-box function, a surrogate function, or a Gaussian process prior. The function may have several set points for the first input dataset with a probability distribution for the second input dataset, similar to the graph shown for the posterior distribution 2616 in FIG. 26. The prediction function may be constructed, at least in part, using training data corresponding to the set of molecules selected using Bayesian optimization. A computing device may construct the function.

[0137] Constructing the prediction function may include estimating reaction characteristic values ​​and / or reaction inputs not present in the first dataset. These reaction characteristic values ​​and / or reaction input values ​​not present in the first dataset may be continuous or discrete parameters (e.g., time, temperature, rate, amount, embedding space). The estimated reaction characteristic values ​​and / or reaction input values ​​may be determined by a Gaussian process.

[0138] The one or more specific points may be predefined, discrete points. For example, the specific points may not include any values ​​in the embedding space, but instead may be limited to only some values ​​in the embedding space. These values ​​in the embedding space may correspond to molecules (e.g., ionic liquids) that are physically present in the field inventory or available to be tested. Predefined may refer to one or more specific points that are determined before the prediction function is constructed or the utility function is evaluated. In some embodiments, one or more of the one or more specific points is not the same as any set of coordinate values ​​in the embedding space for the plurality of first data elements.

[0139] At block 2730, process 2700 may include evaluating a utility function. The utility function may convert a given point in the embedding space into a utility metric that represents the extent to which identifying an experimentally derived reactant property value for the given point is predicted to improve the accuracy of the reactant property value. The utility function may be evaluated by evaluating an acquisition function. The acquisition function may minimize regret over the course of the optimization. The acquisition function may be a GP upper confidence bound (UCB). The acquisition function may be any acquisition function described herein. The utility function may be a specific form of an acquisition function with all optimization hyperparameters set. The utility function may include parameters related to the degree of exploration and exploitation. The utility function may be any utility function described herein. Graph 2628 illustrates the utility function. A computing device may evaluate the utility function.

[0140] At block 2740, the process 2700 can identify one or more particular points in the embedding space as corresponding to a high utility metric based on the utility function. For example, in Figure 26, the graph 2628 shows a utility function with a star 2632 indicating a maximum value. The value associated with that maximum value can be one of the one or more particular points.

[0141] At block 2750, process 2700 may include outputting, by the computing device, results for each particular point of the one or more particular points, identifying a reactant corresponding to the particular point or a reactant structure corresponding to the particular point. The results may be displayed to a user.

[0142] In some embodiments, identifying one or more particular points further includes identifying one or more response characteristic values ​​as corresponding to high utility metrics. The output may include an experimental procedure including the one or more response characteristic values. For example, the output may include reaction conditions. In some embodiments, process 2700 may include conducting an experiment using the identified response characteristic values ​​and / or reaction input values.

[0143] In some embodiments, process 2700 includes accessing an inventory data store containing quantities of reactants. For example, the inventory data store may include quantities of ionic liquid or solvent. The quantities may be adjusted using one or more specific points to determine an adjusted quantity. For example, the one or more specific points may correspond to using n kilograms of ionic liquid A, and process 2700 includes subtracting n kilograms of A from the quantity of A in the inventory data store. Process 2700 may include comparing the adjusted quantity to a threshold value. The threshold value may be zero or some amount that reflects the minimum amount readily available for the experiment. If the adjusted quantity is less than the threshold value, an order may be output for an additional amount of reactant. The order may be a message sent from a computing device to another computing device that manages purchasing of the reactant.

[0144] In some embodiments, process 2700 may include determining that one or more particular points are equivalent to one or more coordinate values ​​of the set of coordinate values ​​in the embedding space of the plurality of first data elements. Process 2700 may further include outputting a message indicating that the one or more particular points represent a converged solution. The computing device may terminate the Bayesian optimization. For example, the utility function may not be evaluated again.

[0145] In some embodiments, the first data set may include points previously identified using a utility function, and an experiment may be performed using the points previously identified using the utility function to determine at least one response characteristic value in the plurality of first data elements.

[0146] In some embodiments, the method may include updating the prediction function using a second data set. The second data set may include a plurality of second data elements. The second data elements may include the same data elements as the first data elements. A portion of the plurality of second data elements may be determined from performing an experiment using reactants or reactant structures identified by the output results.

[0147] Process 2700 may include additional implementations, such as any single implementation or any combination of implementations described below and / or with respect to one or more other processes described elsewhere herein.

[0148] Although Figure 27 illustrates example blocks of process 2700, in some implementations, process 2700 may include additional, fewer, different, or differently arranged blocks than those illustrated in Figure 27. Additionally or alternatively, two or more of the blocks of process 2700 may be performed in parallel.

[0149] Embodiments may also include depolymerization products resulting from conducting experiments using reactants and reaction property values ​​and / or reaction input values ​​corresponding to one or more particular points identified by process 2700.

[0150] Embodiments may also include a method of conducting an experiment, which may include conducting an experiment using reactants and / or reaction input values ​​corresponding to one or more particular points identified by process 2700.

[0151] Embodiments may also include methods of obtaining reactants including an ionic liquid and / or a solvent, where the identity of the reactant or solvent and / or the amount of the reactant or solvent may be determined by the reactant and / or reaction input values ​​corresponding to one or more particular points identified by process 2700.

[0152] Embodiments may also include reactants and / or solvents obtained after being identified by process 2700.

[0153] VI.D. Exemplary Implementations A method for reducing the dimensionality of ionic liquids for Bayesian optimization is described. Additionally, three examples are provided that use the reduced dimensionality and apply a discrete sampling approach to Bayesian optimization via GP. The enthalpy of mixing is minimized across mole fraction and chemical space. Furthermore, the Bayesian optimization method is tested in actual depolymerization experiments, and the results show that the process works well to predict polylactic acid (PLA) conversion and yield.

[0154] VI.D.1. Dimension reduction To illustrate an example of dimensionality reduction, we can access physical property information (e.g., viscosity) for ionic liquids. For example, we queried the National Institute of Standards and Technology (NIST) for all binary ionic liquids (ILs) with known viscosity using pyilt2report (wgserve.de / pyilt2 / pyilt2report.html) (accessed September 15, 2021).

[0155] The cations and anions in the ionic liquid were converted to SMILES representations by querying a lookup table such as the National Institutes of Health's (NIH) Power User Gateway (PUG) web interface.

[0156] The SMILES representations were converted to mordred descriptors (github.com / mordred-descriptor / mordred) (accessed September 15, 2021), each of which was a fixed-length vector of 1,826. Anion mordred descriptors were appended to the cation mordred descriptors of binary ILs. NaN (not a number) values ​​were filled with zeros. A rectangular feature matrix of shape 117 × 3,652 was created. The length of 3,652 is twice the fixed length of 1,826 as a result of appending vectors for both cations and anions. 117 is the number of binary ILs in the dataset.

[0157] The feature matrix can be reduced in size. For each of the 3,652 columns, their corresponding mean was subtracted and divided by n-1 (e.g., 116). This normalization can avoid columns that appear significant simply because their magnitude is higher than other columns. The feature matrix was decorrelated by calculating the correlation between pairs of columns and removing columns with correlation coefficients >0.5 or <-0.5. The feature matrix included column indices that were duplicates based on high correlation, retaining one of each duplicate column. Each column was taken in turn, and the column's correlation with all other columns was calculated. Any of these other columns with high correlations were removed. The process of analyzing correlations and removing columns was repeated for the remaining columns. In this example, after decorrelation, the feature matrix had a shape of 117 x 26.

[0158] Principal component analysis (PCA) was performed on this processed feature matrix. The first four principal components (PCs) captured more than 95% of the relative variance. These four PCs were used for Bayesian optimization.

[0159] In summary, four pieces of information were collected that were used to facilitate the transformation of chemical space: (1) column means in the two-way IL dataset for use in normalizing the columns, (2) two-way samples accessible in the database, (3) highly correlated columns from the combined cation and anion descriptors, and (4) principal components from the feature matrix after decorrelation.

[0160] The reduced dimensionality provided by the PC was used in the Bayesian optimization process. All observations (e.g., x in Figure 26) obs and x s ) were stored in arrays called rxnData and rxnDataResults. Columns corresponding to the four PCs (ilPC1, ilPC2, ilPC3, ilPC4) are included in rxnData. In addition, rxnData had more variables than just the selection of ILs, such as the choice of solvent. More generally, the columns of rxnData used in Gaussian Processes (GPs) are used in Bayesian optimization. Similarly, rxnDataResults can have two or more target columns (e.g., yield and conversion). The first target value was used for training.

[0161] The open-source BayesianOptimization package, available via the MIT license (github.com / fmfn / BayesianOptimization) (accessed September 15, 2021), can be used. The package was adapted to accommodate custom inputs for the array.

[0162] A Gaussian process prior was constructed as described herein. A Matern kernel with v = 2.5 was used, which is a standard choice for representing smoothly varying (i.e., twice differentiable) functions. A posterior distribution is constructed (e.g., as described in Posterior Distribution 2616). The next experiment is determined. The posterior distribution, along with a specified acquisition function, is used to find a utility function, U(x s) was constructed. The acquisition function used is the GP upper confidence bound (UCB), which minimizes regret over the course of the optimization:

number

[0163] VI.D.2. Minimizing the Enthalpy of Mixing over Mole Fraction and Temperature Bayesian optimization was applied to determine the lowest enthalpy of mixing for ionic liquid and solvent pairs. The PC of the ionic liquid determined above was used for the Bayesian optimization. The lowest enthalpy of mixing indicates the most favorable mixing. Data on the enthalpy of mixing for all binary mixtures in the NIST database was collected. The data represent 214 IL and solvent pairs, each pair having varying amounts of data on the enthalpy of mixing as a function of solvent mole fraction and temperature. To first order, the enthalpy of mixing can be described as a canonical solution mixing model: H mix =Ωx A x B , where Ω is the mixing parameter, and if Ω>0, mixing is unfavorable, and if Ω<0, mixing is favorable. A +x B = 1, so to first order, the relationship between enthalpy of mixing versus mole fraction is quadratic.

[0164] Enthalpy minimization for Ω<0 is not as interesting as for Ω>0 and has therefore been studied (the minimum occurs at x=0 or x=1.0). Figures 28A and 28B show the enthalpy minimization for solvent mole fraction x obs H throughout mix We show the results of applying Bayesian optimization to minimize

[0165] Figures 28A and 28B show two sampling approaches, both with the same utility function. The x-axis represents mole fraction. The y-axis of the top graph represents the black-box function. The dashed lines 2804 and 2808 represent the mean of the black-box function. The shaded teal areas (e.g., areas 2812 and 2816) represent 3σ for the black-box function. The observed values ​​x s (e.g., points 2820 and 2824) are red, and the observation x obs (e.g., points 2828 and 2832) are green. Yellow stars (e.g., stars 2836 and 2850) are points that maximize the utility function, x probe is.

[0166] In Figure 28A, the minimization was performed over continuous space, evidenced by the smoothly varying evaluation of f(x) for all values ​​of mole fraction x. At the end of 10 iterations, the Bayesian optimization did not converge to a solution, as the utility function suggested the next experiment (i.e., the mole fraction corresponding to star 2836) was far from the global minimum.

[0167] In Figure 28B, the minimization was performed over discrete space, as evidenced by the evaluation of f(x) for only the given information (i.e., the ground truth shown by the green and red points), and converged to a global minimum of x = 0.2052 after eight iterations.

[0168] In this example, mixing enthalpy as a black-box function was successfully demonstrated. A global minimum of mixing enthalpy was found for Bayesian optimization using discrete sampling. On the other hand, by the end of 10 iterations in the continuous case, a mole fraction of x ~ 0.60 was suggested for the next experiment, which was far from the global minimum of x = 0.2052 reached in the discrete case after 8 iterations. Thus, discrete optimization was more efficient and less expensive than continuous optimization. This result was surprising because mole fraction is a continuous variable, not a discrete one, and it was not expected that considering mole fraction as a discrete variable would result in more efficient and less expensive optimization than considering mole fraction as a continuous variable.

[0169] VI.D.3. Minimizing the Enthalpy of Mixing over Chemical Space A Bayesian optimization via GP across different molecules and chemistries was performed. The PCs of the ionic liquids determined above were used in the Bayesian optimization. The enthalpy of mixing is minimized across the entire chemical space of the new ionic liquid and solvent pair using only a discrete acquisition function. The black-box function in this case is f(x solvent ,T,IL,solvent)=H mix where x solvent is the mole fraction of solvent, T is the temperature, IL is the ionic liquid, solvent is the solvent, and H mix is the enthalpy of mixing. Black box functions cannot be written analytically, even to first order.

[0170] The size of the seed experiment given to seed the Bayesian optimization was considered a "bundle." Each "bundle" is a set of experiments for a given ionic liquid and solvent pair versus H. mix The bundle contained all the mole fraction and temperature data for the process 2700. The bundle contained multiple first data elements (in this example, ionic liquid, solvent, mole fraction) and one or more associated output values ​​(in this example, H mix) may contain information similar to that of the ionic liquid and solvent pair. Separating bundles by ionic liquid and solvent pair rather than by mole fraction is a realistic setup for seed experiments because the cost of purchasing the ionic liquid used with the solvent is much greater than the cost of producing different mole fractions of solvent relative to the IL. In other words, when an IL is purchased or made, thermodynamic data is readily available or obtainable for a range of mole fractions.

[0171] Table 6 shows the average results of 100 trials searching for the minimum enthalpy found in five single-addition experiments given a particular bundle size. Seed bundle sizes were varied from 1 to 7. The average across 100 trials is shown along with the average difference, or improvement in minimization, between the starting minimum enthalpy of mixing (within the bundle) and the final enthalpy of mixing. The larger the difference, the better the model suggested a successful experiment.

[0172] Larger bundle sizes provided the model with more chemical and thermodynamic data to suggest experiments. The average minimum enthalpy is expected to decrease with increasing initial seed bundle size because more data is available with more bundles before any Bayesian optimization is performed.

[0173] Interestingly, in addition to the decrease in the average minimum enthalpy, the mean difference (improvement based on Bayesian optimization suggestions) also increased with larger bundle sizes, as shown in Table 6. Thus, more chemical information provided as seed data suggested reactions with greater improvements.

[0174] Impressively, using four bundles and five additional experiments, the model explored across chemical and thermodynamic space and found conditions that yielded enthalpies of mixing lower than 97% of all other binary enthalpies of mixing in the NIST database (n=4672). [Table 6]

[0175] VI.D.4. Maximizing PLA Conversion and Yield Bayesian optimization was applied to the depolymerization of polylactic acid (PLA). The PC of the ionic liquid determined above was used for Bayesian optimization. The conversion rate (C) and yield (Y) were the targets of optimization. The conversion rate is the amount of total product relative to the amount of starting plastic, and the yield is the amount of target monomer relative to the amount of starting plastic. The black-box model is as follows:

number

[0176] Experiments on depolymerization of PLA (n obs = 94) were collected from a literature review. These selected experiments are obs The Bayesian optimization algorithm functioned as a metric for the Bayesian algorithm. Because the experiments represent a biased population (successful experiments are published, unsuccessful experiments are not), Bayesian optimization is not subject to a realistic observation space. Therefore, Bayesian optimization performance was compared to a baseline random draw scenario. A pure Bayesian approach (one seed experiment and five Bayesian optimization steps) was compared to a pure random approach (six randomly selected experiments), averaged over 100 trials.

[0177] Each test involved picking nSeed (number of seed experiments) randomly drawn from all 94 curated experiments. Then, nExp (additional experiments) were proposed and "executed" (searched in rxnDataResults in this case). The total number of experiments was kept constant. At the end of six experiments, the maximum conversion rate or yield was checked. If the conversion rate (yield) was greater than 95% (85%), the test was successful.

[0178] Table 7 shows the results of 100 trials to predict conversion and yield. The 100 trial average from Bayesian optimization (nSeed / nExp=1 / 5) was compared to a random draw (nSeed / nExp=6 / 0). Conversion results are on the left and yield results are on the right. The % success rate is the average likelihood that the maximized reaction will be greater than 95% for conversion and greater than 85% for yield. The difference is the average difference between the maximum value of nSeed and the maximum value for the entire process (nSeed, nExp). There are no replicates in nSeed / nExp=6 / 0, so the difference is 0%. [Table 7]

[0179] Based on these results, the pure Bayesian approach was the most successful approach (100% success rate for conversion rate and 99% success rate for yield), and the Bayesian model outperformed random draw despite the biased dataset.

[0180] VI.D.5. Depolymerization of Mixed Waste Bayesian optimization techniques are applied to reactions for depolymerizing mixed waste (such as PET / PLA mixtures), contaminated waste (a real-world challenge expected to diminish conversion rates and yields), and a highly overlooked waste stream (black plastic) that has yet to be addressed in academic research. The embedded space of ionic liquids described herein is used in optimizing depolymerization reactions.

[0181] VII. System Environment 29 is an example architecture of a computing system 2900 implemented as some embodiments of the present disclosure. Computing system 2900 is only one example of a suitable computing system and is not intended to suggest any limitation as to the scope of use or functionality of the present disclosure. Neither should computing system 2900 be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in computing system 2900.

[0182] 29, computing system 2900 includes a computing device 2905. Computing device 2905 can reside on a network infrastructure, such as in a cloud environment, or can be a separate, independent computing device (e.g., a service provider's computing device). Computing device 2905 can include a bus 2910, a processor 2915, a storage device 2920, a system memory (hardware devices) 2925, one or more input devices 2930, one or more output devices 2935, and a communication interface 2940.

[0183] The bus 2910 enables communication between components of the computing device 2905. For example, the bus 2910 may be any of several types of bus structures including a memory bus or memory controller, a peripheral bus, and a local bus that uses any of a variety of bus architectures to provide one or more wired or wireless communication links or paths for transferring data and / or power to, from, or between various other components of the computing device 2905.

[0184] Processor 2915 may be one or more processors, microprocessors, or specialized special-purpose processors that include processing circuitry operable to interpret and execute computer-readable program instructions, such as program instructions for controlling the operation and performance of one or more of various other components of computing device 2905 to carry out the functions, steps, and / or performance of the present disclosure. In certain embodiments, processor 2915 interprets and executes the processes, steps, functions, and / or operations of the present disclosure that may be operatively carried out by computer-readable program instructions. For example, processor 2915 may search for, e.g., import and / or otherwise obtain or generate, ionic liquid properties, encode molecular information into an embedding space, decode points in the embedding space into molecules, construct a prediction function, and evaluate a utility function. In embodiments, information obtained or generated by processor 2915 may be stored in storage device 2920.

[0185] The storage device(s) 2920 may include, but are not limited to, removable / non-removable, volatile / non-volatile computer-readable media, such as non-transitory machine-readable storage media, such as magnetic and / or optical recording media and their corresponding drives. The drives and their associated computer-readable media provide storage of computer-readable program instructions, data structures, program modules, and other data for operation of the computing device 2905 in accordance with different aspects of the present disclosure. In an embodiment, the storage device 2920 may store an operating system 2945, application programs 2950, ​​and program data 2955 in accordance with aspects of the present disclosure.

[0186] The system memory 2925 may include one or more storage media, including, for example, a non-transitory machine-readable storage medium such as flash memory, permanent memory such as read-only memory (“ROM”), semi-permanent memory such as random access memory (“RAM”), any other suitable type of non-transitory storage component, or any combination thereof. In some embodiments, the input / output system 2960 (BIOS), containing the basic routines that help to transfer information between various other components of the computing device 2905, such as during start-up, may be stored in ROM. Additionally, data and / or program modules 2965, such as at least a portion of the operating system 2945, program modules, application programs 2950, ​​and / or program data 2955, are accessible to and / or presently operated on by the processor 2915 and may be contained in RAM. In an embodiment, the program modules 2965 and / or application programs 2950 may include, for example, processing tools for identifying and annotating spectral data, metadata tools for adding metadata to data structures, and one or more encoder networks and / or encoder-decoder networks for predicting spectra, which provide instructions for execution by the processor 2915.

[0187] The one or more input devices 2930 may include one or more mechanisms that allow an operator to input information into the computing device 2905, including, but not limited to, a touchpad, a dial, a click wheel, a scroll wheel, a touch screen, one or more buttons (e.g., a keyboard), a mouse, a game controller, a trackball, a microphone, a camera, a proximity sensor, a light detector, a motion sensor, a biometric sensor, and combinations thereof. The one or more output devices 2935 may include one or more mechanisms that output information to an operator, such as, but not limited to, an audio speaker, headphones, an audio line-out, a visual display, an antenna, an infrared port, haptic feedback, a printer, or combinations thereof.

[0188] The communication interface 2940 may include any transceiver-like mechanism (e.g., a network interface, a network adapter, a modem, or a combination thereof) that allows the computing device 2905 to communicate with remote devices or systems, such as mobile devices, or other computing devices, such as servers in a networked environment, such as a cloud environment. For example, the computing device 2905 may be connected to remote devices or systems via one or more local area networks (LANs) and / or one or more wide area networks (WANs) using the communication interface 2940.

[0189] As described herein, computing system 2900 may be configured to train an encoder-decoder network to predict characteristic spectral features from structural representations of materials obtained as structural strings. In particular, computing device 2905 may perform tasks (e.g., processes, steps, methods, and / or functions) in response to processor 2915 executing program instructions contained in a non-transitory machine-readable storage medium, such as system memory 2925. The program instructions may be loaded into system memory 2925 from another computer-readable medium (e.g., a non-transitory machine-readable storage medium), such as data storage device 2920, or from another device via communication interface 2940 or a server within or outside a cloud environment. In embodiments, an operator may interact with computing device 2905 via one or more input devices 2930 and / or one or more output devices 2935 to facilitate performance of tasks according to aspects of the present disclosure and / or achieve the end result of such tasks. In additional or alternative embodiments, hardwired circuitry may be used in place of or in combination with program instructions to perform tasks, e.g., steps, methods, and / or functions consistent with different aspects of the present disclosure. Thus, the steps, methods, and / or functions disclosed herein may be performed with any combination of hardware circuitry and software.

[0190] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium including instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium, including instructions configured to cause one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein.

[0191] The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions to exclude any equivalents of the features shown and described, or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, while the invention as claimed has been specifically disclosed by embodiments and optional features, it will be understood that modifications and variations of the concepts disclosed herein may be effected by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims.

[0192] The description provides only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the description of preferred exemplary embodiments provides an enabling description for those skilled in the art to practice various embodiments. It will be understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope of the appended claims.

[0193] Specific details are given in the following description to provide a thorough understanding of the embodiments. However, it will be understood that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order to avoid obscuring the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.

[0194] The terms "a," "an," or "the" are intended to mean "one or more" unless specifically indicated to the contrary. The use of "or" is intended to mean "inclusive or" rather than "exclusive or" unless specifically stated to the contrary. A reference to a "first" element does not necessarily imply a second element. Additionally, a reference to a "first" or "second" element does not limit the referenced elements to a particular location unless expressly stated. The term "based on" is intended to mean "based at least in part on."

[0195] The claims may be drafted to exclude any element that may be optional. As such, this statement is intended to serve as a precondition for using exclusive terminology such as "solely," "only," and the like in connection with the recitation of claim elements or the use of a "negative" limitation.

[0196] Where a range of values ​​is provided, unless the context clearly dictates otherwise, it is understood that each intervening value, to the tenth of the unit of the lower limit, between the upper and lower limit of that range is also specifically disclosed. Each smaller range between any stated or intervening value in a stated range and any other stated or intervening value in that stated range is encompassed within an embodiment of the disclosure. The upper and lower limits of these smaller ranges may independently be included or excluded, and each range in which either, either, or both limits are included in the smaller range is also encompassed within the disclosure, subject to any specifically excluded limits in the stated range. When a stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure.

[0197] All patents, patent applications, publications, and descriptions mentioned in this specification are incorporated herein by reference in their entirety to disclose and describe the methods and / or materials in connection with which the publications are cited, as if each individual publication or patent was specifically and individually indicated to be incorporated by reference. None are admitted to be prior art.

Claims

1. 1. A computer-implemented method comprising: accessing a multidimensional embedding space that supports relating embeddings of molecules to predicted values ​​of given properties of said molecules; identifying one or more interest points in the multidimensional embedding space based on the predicted values, each of the one or more interest points comprising: a set of coordinate values ​​in the multidimensional embedded space; conveying spatial information of atoms or bonds within the molecule; identifying the one or more points of interest associated with a corresponding predicted value of the given characteristic; for each of the one or more interest points, generating a structural representation of the molecule by transforming the set of coordinate values ​​contained in the interest point using a decoder network, wherein training of the decoder network includes learning to transform positions in the embedding space into outputs representing molecular structural properties, and wherein the training of the decoder network is performed at least partially concurrently with the training of an encoder network; for each of the one or more points of interest, outputting a result identifying the structural representation of the molecule corresponding to the point of interest; 11. A computer-implemented method comprising:

2. 2. The method of claim 1, wherein training the encoder network comprises learning to transform partial or complete bond string and position (BSP) representations of molecules into positions in the embedding space, each BSP representation identifying the relative positions of atoms connected by bonds within the represented molecule.

3. 3. The method of claim 2, wherein each BSP representation of the molecule used to train the encoder network includes a set of coordinates for each of the atoms connected by the bond in the represented molecule, further identifying each of the atoms connected by the bond in the represented molecule.

4. 4. The method of claim 2 or 3, wherein the BSP representations of the molecules are used to train the encoder network to identify bond types for each of at least some bonds in each molecule.

5. The method of any one of claims 2 to 4, wherein the format of the structural representation identified in the result is different from the BSP representation.

6. 2. The method of claim 1 , wherein training the encoder network comprises learning to transform partial or complete molecular graph representations of molecules into positions in the embedding space, each molecular graph representation identifying bond angles and distances within the represented molecule.

7. 6. The method of claim 1, wherein the decoder network and the encoder network are trained by training a Transformer model that uses self-attention, the Transformer model comprising the decoder network and the encoder network.

8. The method of claim 1 or 6, wherein the decoder network and the encoder network are trained by training a Transformer model that includes an attention head.

9. a machine learning model including the encoder network and the decoder network, accessing a set of supplemental training elements, each of the set of training elements including a representation of the structure of a corresponding given molecule; for each supplemental training element in the set of supplemental training elements, masking at least a portion of the representation to obscure at least a portion of the structure of the corresponding given molecule; 9. The method of claim 1, further comprising: training the machine learning model to predict the obscured at least part of the structure.

10. 10. The method of claim 1, wherein training the encoder network further comprises fine-tuning the encoder network to convert positions in the space into predictions corresponding to values ​​of the given property.

11. 1. A system comprising: one or more data processors; a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform the method of any one of claims 1 to 10.

12. A computer program product tangibly embodied in a non-transitory machine-readable storage medium, the computer program product comprising instructions configured to cause one or more data processors to perform the method of any one of claims 1 to 10.