Method of predicting molecular identity
A transformer-based machine learning framework processes raw NMR spectra to predict molecular structure and substructures end-to-end, addressing the combinatorial challenge of molecular structure elucidation with high accuracy and efficiency, reducing the search space by orders of magnitude.
Patent Information
- Application Number
- PCT/US2025/040957
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-12
- Filing Date
- 2025-08-06
- Publication Date
- 2026-02-19
AI Technical Summary
Existing methods for elucidating molecular structure from one-dimensional nuclear magnetic resonance (NMR) spectra face significant challenges due to the combinatorial explosion of possible molecules as the number of constituent atoms increases, requiring extensive preprocessing and prior knowledge, and are inefficient and error-prone.
A transformer-based machine learning framework that processes raw 1D1H and13C NMR spectra without prior information, integrating a multitask model to predict molecular structure and substructures end-to-end, using minimal preprocessing and leveraging a pretrained transformer for substructure-to-structure task initialization.
The framework achieves high accuracy in predicting the exact molecular structure up to 69.6% within the first 15 predictions for molecules with up to 19 heavy atoms, significantly reducing the search space by up to 11 orders of magnitude, and provides a rapid and automated tool for structure elucidation.
Smart Images

Figure US2025040957_19022026_PF_FP_ABST
Abstract
Description
METHOD OF PREDICTING MOLECULAR IDENTITY CROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority under 35 U.S.C. §119 from U.S. Application No. 63 / 682,183, filed August 12, 2024, which is incorporated herein by reference for all purposes. BACKGROUND
[0002] The background description provided here is for the purpose of generally presenting the context of the disclosure. Information described in this background section, as well as aspects of the description that may not otherwise qualify as prior art at the time of filing, are neither expressly nor impliedly admitted as prior art against the present disclosure.
[0003] The present disclosure is concerned with the elucidation of the structure of chemical compounds. The present disclosure provides a method, an artificial neural network, and a training protocol for the provided neural network to elucidate the chemical structure from the spectroscopic data of the chemical compounds. SUMMARY
[0004] To achieve the foregoing and in accordance with the purpose of the present disclosure, a computer-implemented method for generating a prediction of a molecular formula and molecular structure of an unknown chemical compound is provided. Nuclear magnetic resonance (NMR) spectral data of the unknown chemical compound comprising at least one of a13C NMR spectrum and a1H NMR spectrum corresponding to the unknown chemical compound is received as input. Thespectral data is provided to a machine learning framework pretrained on a plurality of spectra-structure pairs, wherein each pair comprises (i) a NMR spectral representation and (ii) a corresponding molecular structure or molecular formula. The NMR spectral data corresponding to the unknown chemical compound is processed by the machine learning model to generate a list of one or more candidate molecular formulas and corresponding candidate molecular structures. The list of candidate molecular formulas and corresponding candidate molecular structures for the unknown chemical compound is provided as output.
[0005] In another manifestation, a system for generating a prediction of a molecular formula and molecular structure of an unknown chemical compound is provided. A non-transient computer readable medium stores instructions that, when executed by the one or more processors, cause the system to perform operations, comprising receiving, as input, nuclear magnetic resonance (NMR) spectral data of the unknown chemical compound comprising at least one of a13C NMR spectrum anda1H NMR spectrum corresponding to the unknown chemical compound, the NMR spectral data to a machine learning framework pretrained on a plurality of spectra-structure pairs, wherein each pair (i) a NMR spectral representation and (ii) a corresponding molecular structure or molecular formula, processing, by the machine learning model, the NMR spectral data corresponding to the unknown chemical compound to generate a list of one or more candidate molecular formulas and corresponding candidate molecular structures, and outputting a list of candidate molecular formulas and corresponding candidate molecular structures for the unknown chemical compound
[0006] These and other features of the present invention will be described in more detail below in the detailed description of the disclosure and in conjunction with the following figures. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The present disclosure is illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings and in which like reference numerals refer to similar elements and in which:
[0008] FIG. 1 illustrates a full end-to-end multitask structure elucidation model, which takes in the1H and / or13C nuclear magnetic resonance (NMR) spectra and predicts both the substructures in the and the molecule’s structure.
[0009] FIG. 2A illustrates the transformer and the best multitask model test accuracy as a function of the problem size.
[0010] FIG. 2B illustrates the distribution of Tanimoto similarities of the best incorrect predictions relative to the target molecule after removing correct predictions and invalid SMILES.
[0011] FIG. 3 shows examples of molecules that were correctly or incorrectly predicted in an embodiment.
[0012] FIG. 4 shows examples of molecules that were correctly or incorrectly predicted in an embodiment.
[0013] FIG. 5 shows the distribution of model predictions on a test set.
[0014] FIG. 6 is a diagram of the multitask workflow during training.
[0015] FIG. 7 shows a learning curve for the training of the substructure-to-structure model.
[0016] FIG. 8 shows two different train-validation-test splittings of a 3M dataset.
[0017] FIG. 9 shows training and validation loss curves for different combinations of spectra
[0018] FIG. 10(a) – 10(b) illustrate the deconstruction of a molecule into its constituent substructures.
[0019] FIG. 11(a) – 11(b) illustrate another deconstruction of a molecule into its constituent substructures.
[0020] FIG. 12(a) – 12(b) illustrate another deconstruction of a molecule into its constituent substructures.
[0021] FIG. 13 illustrates an example of a set of predictions by the substructure-to-structure transformer which contains the correct molecule.
[0022] FIG. 14 illustrates another example of a set of predictions by the substructure-to- structure transformer which contains the correct molecule.
[0023] FIG. 15 illustrates the performance of the substructure-to-structure transformer when using Morgan fingerprints, an alternative encoding of molecular structure to the substructures. On the left-hand side, the darker bars correspond to fingerprint-to-structure and the lighter bars correspond to substructure-to-structure.
[0024] FIG. 16 illustrates the performance of the complete spectrum-to-structure multitask model when the transformer is pretrained using Morgan fingerprints as an alternative encoding of molecular structure. On the left-hand side, the darker bars correspond to using the fingerprint-to- structure transformer as the pretrained model, and the lighter bars correspond to using the substructure-to-structure transformer as the pretrained model.
[0025] FIG. 17 is a schematic of how a Morgan fingerprint is constructed for a given molecule.
[0026] FIG. 18 illustrates training and validation loss curves for the multitask model with simulated training data and experimental validation / test data. The divergence of the training loss on simulated data and validation loss on experimental data shows clear overfitting of the model to the simulated spectra.
[0027] FIG. 19 is a high level schematic view of a system that may be used in an embodiment.
[0028] In the drawings, like reference numerals are sometimes used to designate like structural elements. It should also be appreciated that the depictions in the figures are diagrammatic and not to scale. DETAILED DESCRIPTION OF ILLUSTRATED EMBODIMENTS
[0029] Rapid determination of molecular structures can greatly accelerate workflows across many chemical disciplines. However, elucidating structure using only one-dimensional (1D) nuclear magnetic resonance (NMR) spectra, the most readily accessible data, remains an extremely challenging problem because of the combinatorial explosion of the number of possible molecules asthe number of constituent atoms is increased. Some embodiments provide a multitask machine learning framework that predicts the molecular structure (formula and connectivity) of an unknown chemical compound solely based on its 1D1H and / or13C NMR spectra. First, we show how a transformer architecture can be constructed to efficiently solve the task, traditionally performed by chemists, of assembling large numbers of molecular fragments into molecular structures. Integrating this capability with a convolutional neural network, we build an end-to-end model for predicting structure from spectra that is fast and accurate. We demonstrate the effectiveness of this framework on molecules with up to 19 heavy (non-hydrogen) atoms, a size for which there are trillions of possible structures. Without relying on any prior chemical knowledge, such as the molecular formula, we show that our approach predicts the exact molecule 69.6% of the time within the first 15 predictions, reducing the search space by up to 11 orders of magnitude.
[0030] 1D1H and13C NMR spectroscopy are workhorse methods in structure elucidation, since they are both easy to perform and provide a wealth of information about molecular composition, connectivity, and stereochemistry. However, interpretation of NMR spectra remains time-intensive and error-prone, often requiring high levels of chemical expertise and prior knowledge that constrains the space of possible structures. For small molecules or structures that are composed of a known subset of chemical building blocks, the measured spectra can be compared with databases of spectra, obtained either from previous experiments or forward prediction methods, with the hope of obtaining a match. However, the number of possible molecules suffers from a combinatorial explosion as the number of constituent atoms increases. For example, for organic compounds made up of up to 11 heavy (non-hydrogen) atoms (C, N, O, S, and / or halogen), there are ∼26 million molecules consistent with the basic bonding rules of chemistry (not including stereoisomers), but by 17 heavy atoms, this number grows to ∼166 billion possible structures, and by 21 heavy atoms, exceeds 20 trillion. Addressing this formidable challenge to unlock unsupervised structure elucidation of organic molecules would remove a key bottleneck in chemical research and could be paired with automated synthetic workflows to create closed-loop discovery platforms.
[0031] We and others have reported machine learning (ML) approaches to structure elucidation based on identifying molecular substructures (functional groups, small fragments) from NMR spectral data and using this information to assemble a molecular structure or based on the task of selecting which of a set of provided candidate molecules corresponds to a given NMR spectrum and assigning the peaks in that provided spectrum to specific atoms. While ML models have proven tobe remarkably effective in identifying substructures, converting substructure information to a molecular structure is much more challenging. Beam searching over possible structures or building the structure in an atom-by-atom manner are viable for small systems but quickly lose effectiveness as the number of atoms increases because of the combinatorial scaling of the problem size.
[0032] Rather than predicting fragments as an intermediate step, recent work has sought to directly predict molecular structure from spectra in an end-to-end fashion, using deep learning architectures such as graph convolutional neural networks and transformers. For molecules with up to 13 heavy atoms, a top-10 accuracy of 78.5% was achieved with a transformer model using as inputs the molecular formula and infrared (IR) spectra that were preprocessed into a sequence of integers representing the intensity. On systems of up to 35 heavy atoms, a similar approach with transformers yielded a top-10 accuracy of 86.6% using the molecular formula and1H NMR and13C NMR spectra preprocessed into strings containing information about peak shifts, splittings, and multiplicities. For larger organic systems ranging from a few dozen to over a hundred heavy atoms, it was shown that one can obtain a top-10 accuracy of 94.2% when feeding the molecular formula,13C NMR shifts, and a Simplified Moleculary Input Line Entry System (SMILES) string representation of a large fragment of the target molecule into a transformer pretrained on 360 million molecules. These studies suggest that end-to-end frameworks are a promising approach to address the combinatorial problem when extensive preprocessing is applied to the spectra and sufficient chemical information is provided (e.g., the molecular formula). However, in many cases, information such as the molecular formula is not readily available, and preprocessing spectra, such as assigning peak shifts, splittings, and multiplicities, is burdensome and introduces biases. It is, therefore, critical to develop a framework that can accurately and efficiently predict the molecular structure of an unknown chemical compound using spectral data with minimal preprocessing, spectral data with automated preprocessing without manual assignment of peak shifts, splittings, or multiplicities, or raw spectral data alone. Here, raw spectral data refers to either the raw data signal acquired by the spectrometer, which is often called the free induction decay (FID), or the Fourier transformed version of the FID, which is typically an x-y plot of intensity in arbitrary units vs chemical shift in units of parts per million (ppm).
[0033] Some embodiments provide a transformer-based ML framework for solving the most challenging version of the structure elucidation problem, applying minimal preprocessing to the input NMR spectral data of at least one of1H and13C NMR spectra and using no other priorinformation such as molecular formula or molecular fragments. First, some embodiments train a transformer model to solve the problem of constructing the molecular structure (i.e., the molecular formula and the connectivity of the constituent atoms) when given only information about the presence or absence of a set of 957 very simple predefined molecular substructures (≤7 atoms). We show that a transformer architecture recovers the exact molecular structure with high accuracy, succeeding 93.2% of the time within the first 15 predictions when tested on molecules with up to 19 heavy atoms. Next, some embodiments integrate this pretrained transformer into a multitask model that predicts both molecular substructure and molecular structure. This end-to-end model inputs only 1D1H NMR and13C NMR spectra and yields the correct molecular structure 69.6% of the time within the first 15 predictions when tested on simulated NMR spectral data for molecules with up to 19 heavy atoms. This model is thus capable of rapidly constraining the massive chemical search space (>2 trillion possibilities) using just routinely available NMR spectra, providing a complementary tool to other existing structure elucidation, reaction prediction, and retrosynthesis frameworks. Results and discussion
[0034] An overview of our structure elucidation framework is shown in FIG. 1. The top row of FIG. 1 illustrates a full end-to-end multitask structure elucidation model, which takes in the1H and / or13C NMR spectra and predicts both the molecular substructures in the molecule and the molecule’s structure. The framework uses an end-to-end approach, without the explicit prediction of the molecular substructures as an intermediate step employed in our previous work to avoid the significant loss of information when compressing the latent features of the model to the low- dimensional representation afforded by the molecular substructure arrays.
[0035] The top of FIG. 1 illustrates an overview of the full multitask structure elucidation workflow and the bottom of FIG. 1 illustrates the substructure-to-structure workflow. Weights from a transformer pretrained on the substructure-to-structure task are used to initialize the multitask model. Specific details regarding the transformer model architecture and multitask model architecture used in an embodiment are described below.
[0036] The bottom row of FIG. 1 shows our approach to the substructure-to-structure problem in an embodiment, where molecular structures, represented by SMILES strings, are formed out ofmolecular fragments, represented by the substructure arrays. As will be shown below in an embodiment, pre-training on the task of structure elucidation from molecular substructures substantially improves the accuracy of the full multitask structure elucidation model. Hence, we first demonstrate how a transformer can be trained on this task before showing how this can be integrated into the complete end-to-end multitask framework.
[0037] A precise schematic detailing the dimensions of the layers used in an embodiment is shown in FIG. 6. FIG. 6 is a diagram of the multitask workflow during training, with the substructure-to-structure transformer (left box) feeding into the overall multitask architecture (right box). E refers to the model embedding dimension, T is the length of the target SMILES string in tokens, and N is the number of encoder / decoder layers in the transformer. The transformer (left box) uses the PyTorch Transformer class. Conv1D is a one dimensional convolution using the PyTorch Conv1d layer, MLP (a multi-layer perceptron) refers to a feed-forward neural. 1 max pooling is performed using the PyTorch MaxPool1d layer, embedding is done using the PyTorch Embedding layer, and Norm refers to layer normalization.
[0038] For the substructure-to-structure transformer (FIG. 6, left box), we use an encoder decoder transformer as implemented in PyTorch. Table 1 details the exact architecture of the final model used for the results in the main text. All models used absolute positional encoding based on sinusoidal functions of different frequencies as described by the following equations: ^^^^^^,^^^ = sin^^^^ / 10000^^ / ^^^^^^^where pos is the position, i is the dimension, and dmodelis the model embedding dimension. Positional encodings are added to the embeddings for both the molecular substructures and the SMILES tokens. Table 1: Substructure-to-structure architectural parameters, descriptions, and their values. Parameter Description Value d model Embedding dimension of the model 128 dim feedforward Hidden layer dimension of the feed forward neural network within the transformer blocks 1024 source size Total number of possible token values for embedding substructure sequences 958src pad token The index used for padding source sequences to the same length 0target size Total number of possible token values for embedding SMILES token sequences 24 tgt pad token The index used for padding target sequences to the same length 21 num encoder layers The number of encoder layers 6num decoder layers The number of decoder layers 6nhead The number of heads used in multihead attention 8activation The activation function used for intermediate encoder / decoder layers reludropout The probability for a particular element of an input tensor to be randomly set to 0 0.1 layer norm eps The constant used for numerical stability in layer normalization 1E-5 Table 2: Spectrum-to-substructure encoder architectural parameters, descriptions, and their values.Parameter Description Valuelayer norm eps The constant used for numerical stability in layer normalization 1E-5
[0039] For the spectrum-to-structure and spectrum-to-substructure multitask model, the architecture of the convolutional embedding for the1H NMR and embedding for the13C NMR is shown on the right side box of FIG. 1. Theof the encoder-decoder transformer component is the same as in Table 1, and Table 2 describes the architecture of the encoder component used for molecular substructure elucidation in the final multitask model. An important step in obtaining the molecular substructure profile from the encoder is aggregating the information from the higher-dimensional raw encoder output into the lower-dimensional format of the molecular substructure profiles. To this end, we adapt a method called sequence pooling. Given an input sequence x0 and a function f parameterized by a neural network, sequence pooling performs the following operations in order: ^= ^^ ^ ∈ $×&×'^ ! #^(^ = softmax^)^^ & $×*×&^^ ^ ∈ #+ = squeeze^^(^^^^ ∈ #$×'where N is the batch size, T is the sequence length, E is the embedding dimension, and g(·) ∈ RE×1 is a learnable linear transformation. This can be understood as attending across the sequence dimension of the data after processing it through the encoder and assigning importance weights to each element in the sequence before aggregation.Substructure-to-structure prediction
[0040] We first consider the problem of deducing a molecular structure from knowledge of the presence or absence of a set of molecular substructures, which we define as small fragments of a molecule with defined bonding relationships. This task is inspired by the way chemists commonly interpret NMR spectra, using peaks to identify fragments and assembling those fragments into candidate structures. Manually converting a relatively small number of molecular substructures into a molecular structure is already challenging even for an experienced chemist, and it quickly becomes impractical when confronted with dozens or hundreds of molecular substructures, which is necessary for accommodating the complexity of structure space, as molecular substructures only provide information about smaller local environments within a larger molecule. Hence, devising an automated strategy for structure elucidation from these molecular fragments can itself help accelerate structure determination.
[0041] Here, we treat the substructure-to-structure problem as a language translation problem, where we translate a sequence of substructures into a sequence of tokens that can be combined to form a molecule’s SMILES string, recovering the molecular connectivity. The SMILES string encodes all the relevant structural information for a molecule and can be used to obtain the corresponding molecular formula and corresponding molecular structure. A molecular structure can also be converted into its corresponding SMILES string. Our approach is outlined in the bottom row of FIG. 1, where we use a transformer model architecture for this task. Autoregressive transformer models have been shown to excel at translation tasks owing to the strong inductive bias provided by the multihead attention mechanism, which captures global correlations and uses that information to generate the target sequence in a conditioned manner. This strength is critical in a chemical context because the overall connectivity of a molecule arises from considering all molecular fragments collectively and many fragments may or may not overlap depending on the molecular context. Here, we adapt a full encoder-decoder transformer architecture for the substructure-to-structure task.
[0042] The inputs to the substructure-to-structure transformer are a binary vector with each entry corresponding to the presence or absence of a specific molecular substructure and a start token. Using only this information, the transformer generates a SMILES string token by token, which it continues to build until it terminates upon the prediction of the stop token.
[0043] To train the transformer on this task, we started with a set of ∼143 thousand molecules from the SpectraBase database containing only C, N, and O as the non-hydrogen atoms andcombined it with a set of ∼3 million molecules randomly sampled from the GDB-17 dataset to create a final dataset of ∼3.1 million molecules. After canonicalizing all the SMILES strings using RDKit, the molecular substructure vector for each molecule was computed using RDKit’s molecular substructure match functionality against the set of 957 predefined molecular substructures used in our previous work, represented as Smiles Arbitrary Target Specification (SMARTS) strings. For the target that the transformer uses in decoding, the canonicalized SMILES were tokenized using a regular expression, and an alphabet was determined from all unique tokens across all SMILES to ensure that all SMILES can be represented during training and testing.
[0044] In an embodiment used to generate the data for training the substructure-to-structure transformer model, we started with a set of 142894 SMILES strings containing only C, N, O, and H atoms from the SpectraBase dataset. We canonicalized all the SMILES strings using RDKit and removed all stereochemical information from the strings, including designation of double bond stereochemistry. We combined this set of 142894 SMILES with randomly sampled SMILES from the GDB-17 dataset to create a final dataset of 3116791 SMILES strings. Using the set of 957 substructures from our previous work, we constructed the substructure arrays for this dataset by performing a substructure search, generating a binary vector of length 957 for each molecule where a “1” means a predefine molecular substructure is present and a “0” means a predefined molecular substructure is absent. SMILES strings were tokenized using a regular expression.
[0045] FIG. 2A illustrates transformer and the best multitask model test accuracy as a function of the problem size. The problem size is determined by extrapolating an exponential fit to the number of molecules in GDB-9, GDB-11, GDB-13, and GDB-17, and the plot begins with the number of possible structures for 10 heavy (non-hydrogen) atoms. The darker bars denote the substructure-to-structure, and the lighter bars denote spectrum-to-structure. FIG. 2B illustrates distribution of Tanimoto similarities of the best incorrect predictions relative to the target molecule after removing correct predictions and invalid SMILES.
[0046] During training, the dataset was split into a training, validation, and test set using a random split, with 80% of the data used for training, 10% used for validation, and 10% used for testing. To produce a more compact representation of the binary substructure arrays for the embedding, an integer array was created which lists the index (1 to 957) of only the present (non- zero) substructures for each molecule. Each unique SMILES token was also assigned to an integer index and used for embedding. Right padding was applied to ensure all substructure vectors andSMILES vectors in a batch were of the same length before being passed into the transformer. The padding was masked during training.
[0047] In an embodiment, the substructure-to-structure transformer model was trained using a batch size of 32 with a constant learning rate of 1×10−5and a weight decay of 1×10−5using the Adaptive Moment Estimation (ADAM) optimizer implemented in PyTorch. The loss function for the substructure-to-structure model was a cross-entropy loss between the model’s output probability distribution and the correct sequence of tokens for the target SMILES string. The dataset, composed of the 3116791 SMILES strings and substructure arrays, was partitioned using a random splitting with 80% used for training, 10% used for validation, and 10% used for testing. Early stopping was used to prevent overfitting by monitoring the loss on the validation set. The model was optimized for 312 epochs, and the final model selected was that with the lowest validation loss. The final substructure-to-structure model was trained for 308 epochs, at which point the minimum validation loss was attained. The learning curve for the training of the substructure-to-structure model is shown in FIG. 7.
[0048] The reason that the validation loss is consistently lower than the training loss is because of the use of dropout as a regularization technique when training the transformer model. During forward passes through the model when in the training stage, dropout randomly samples certain elements of the input tensor passed through the layer and sets the elements to zero with some probability p. This means that only a subset of the network’s neurons are used when computing the training loss in each forward pass. During validation, dropout layers are disabled and all neurons within the model are used (and all neurons are also used when evaluating the test set with the final fit model). Thus, the training loss is computed using a subset of the model’s neurons, whereas the validation loss is computed over the full model, leading to a lower validation loss than training loss. To demonstrate that the expected behavior (train loss less than or equal to validation loss) is obtained when not using dropout layers we have re-evaluated the full model’s (i.e. with all neurons used) loss on the train and validation set from 10 saved checkpoints and show that in each case, the validation loss is higher than the training loss. These results are shown in Table 3. Furthermore, by enabling dropout layers during validation and refitting the transformer model, the model’s validation loss is consistently higher than the training loss over the first 100 epochs, as expected. This is shown in FIG. 8 for two different train-validation-test splittings of the 3M dataset.
[0049] Table 3: Training and validation loss from checkpoints re-evaluated on the training set with all neurons. The validation loss is observed to be consistently higher than the training loss when all neurons of the model are used to evaluate the training set. Epoch Validation Loss Training loss Validation - Training 301 0.02750 0.02300 +0.00449 302 0.02747 0.02295 +0.00452 303 0.02743 0.02289 +0.00454 304 0.02739 0.02294 +0.00445 306 0.02739 0.02270 +0.00468 307 0.02748 0.02296 +0.00452 308 0.02729 0.02269 +0.00460 309 0.02735 0.02273 +0.00463 310 0.02751 0.02278 +0.00473 311 0.02750 0.02291 +0.00459
[0050] To test the trained transformer model for the substructure-to-structure task, top-k random sampling was used to generate SMILES from a given substructure array with k = 5, and 15 predictions were generated for each input, comprising the set of predicted candidate molecular formulas and corresponding candidate molecular structures, which can be obtained directly from the candidate SMILES strings. Using these 15 generated candidate SMILES strings, the prediction was considered correct if one of these molecules, upon canonicalization, matched the exact canonical SMILES of the actual molecule. In some embodiments, the generated candidate molecular formulas and corresponding candidate molecular structures, which can be obtained directly from the candidate SMILES strings, can be a ranked list of candidate molecular formulas and candidate molecular structures where the ranking is determined by some metric. On the test set, the model achieved an accuracy of 93.2%. Examples of molecules correctly predicted by the transformer model are shown at the top section of FIG. 3, demonstrating that the model can deduce complex molecular structures using only substructures as input. These example molecules require the transformer to assemble between 23 and 65 substructures into the correct molecule. The difficulty of this task is illustrated in FIGS. 10-12, which show molecules that were correctly predicted by thesubstructure-to-structure model and the substructures that were provided to it. The test set accuracy as a function of the number of heavy atoms is shown as the darker bars in FIG. 2A for up to 19 heavy atoms. Notably, the accuracy of the substructure-to-structure model shows little variation as the problem size grows. For example, at 10 heavy atoms, where the problem size is ∼2 million, the accuracy is 96.0% and only drops to 92.8% by 17 heavy atoms, where it has increased to ∼200 billion. However, the accuracy drops below 80% for the molecules with 18 or 19 heavy atoms. Although this sudden decrease could be explained in part by the problem size entering the trillions for molecules of that size, a more likely explanation might be that GDB-17 only contains molecules with up to 17 heavy atoms, while the SpectraBase dataset, which does contain molecules with 18 and 19 heavy atoms, is much smaller. Hence, whereas 59% in our training data come from molecules with 17 heavy atoms, only 1% come from molecules with 18 or 19 heavy atoms, meaning that the model has had fewer opportunities to learn how to predict systems of that size. However, the fact that the model retains reasonable accuracy for these larger molecules suggests that the approach is generalizable. Since computing the substructure vectors is facile using RDKit, the set of substructures could also easily be expanded to cover fragments more commonly encountered in larger molecules, which would be expected to further improve performance.
[0051] FIG. 3 (Top) shows examples of molecules that were correctly predicted by the transformer model. The molecules shown have between 23 to 65 substructures, and examples of correctly predicted molecules and their constituent substructures are shown in FIGS. 10-12. FIG. 3 (Bottom) shows examples of molecules with incorrect predictions from the transformer model. The number beneath each predicted molecule is the Tanimoto similarity between the prediction and target.
[0052] To better understand the failure modes of the substructure-to-structure model, we examine the distribution of Tanimoto similarities computed between the model’s best incorrect prediction and the correct molecule in cases where the correct molecule was not predicted.
[0053] This is shown for the test set as the solid line in FIG. 2B. Overall, 89.6% of the best incorrect predictions have a similarity to the correct molecule that is greater than or equal to 0.50, with an average similarity of 0.68. To see how these similarity scores correlate to the molecular structure of the incorrect predictions, the bottom section of FIG. 3 shows some examples target molecules and incorrect predictions at varying degrees of Tanimoto similarity. We see that even for incorrect predictions at similarities near 0.50, the predictions already contain many of the desiredfunctional groups and bonding motifs of the target, albeit with incorrect connectivity. As the similarity increases, the connectivity and structure of the predictions improve, and at similarities near 0.80 and above, the incorrect predictions are very close to the target, often differing by a few atoms or bonds. This is encouraging as it shows that the transformer is capable of learning how to map the binary substructure representation to real molecular fragments, leading to the recovery of a significant portion of the target molecule scaffold, even in cases where the exact correct molecule is not predicted. Considering both the high proportion of best incorrect predictions above 0.50 and the high mean similarity, these results show that even when the model returns incorrect predictions, the molecules generated can still provide meaningful structural information about the target molecule. Spectrum-to-structure and spectrum-to-substructure prediction using a multitask workflow
[0054] We now focus on the more challenging problem of directly elucidating the molecular structure and substructures from the1H and13C NMR spectra, which we refer to as spectrum-to- structure and spectrum-to-prediction, respectively. Our goal is to design a model capable of predicting a molecule’s structure and the substructures it contains using only the spectral data as inputs, operating in an end-to-end fashion, i.e. taking the spectrum and directly predicting the structure it corresponds to and the substructures within that molecule with no intermediate steps. The workflow for our approach is shown in the top row of FIG. 1. As we will emphasize further below an important component of training this multitask spectrum-to-structure model is the initialization of the transformer using the weights obtained from training on the substructure-to- structure task (as indicated by the arrow 108 in FIG. 1 from the bottom workflow to the multitask model on the top).
[0055] For training and testing the spectrum-to-structure model, we used1H and / or13C NMR simulated NMR spectra for the ∼143 thousand set of molecules frompredicted using MestreNova, using the same train-validation-test split of the molecules as on the substructure-to- structure task. To effectively utilize all portions of the spectral signal, we took a minimal preprocessing approach to the spectra to retain as much information as possible. Specifically, since the1H NMR contains detailed splittings and rapidly varying features, we processed the spectrum by reducing the dimensionality of the data by linearly interpolating the spectrum of 32768 values down to a grid of 28000 values corresponding to a shift range of -2 to 12 ppm with a resolution of 0.0005 ppm. To assist the numerical stability when fitting the multitask model whileretaining the relative intensities of all peaks, all intensities were normalized to between 0 and 1 by dividing each spectrum by the intensity of its highest peak. Normalization did not have a statistically significant impact on model accuracy. For the13C NMR spectra, because the intensities are not generally reliable indicators of the relative number of nuclei and the peak shapes are not informative because peaks are typically decoupled, we processed the spectra by extracting the chemical shifts and binning them into 80 bins spanning the shift range of 3.42 to 231.3 ppm.
[0056] In an embodiment used to generate data for training the spectrum-to-structure and spectrum-to-substructure multitask model, we converted the set of 142894 SMILES strings from SpectraBase into two-dimensional mol files using Open Babel and then computed the1H and13C NMR spectra using MestreNova version 14.2.
[0057] 1H NMR spectra were computed with a line width of 0.75 Hz using deuterated chloroform as the solvent. Labile protons were excluded from the calculation of the spectra. The spectra were output as grids of 32768 intensities which were then interpolated down to grids of 28000 values spanning a ppm shift range of -2 to 12 ppm with a resolution of 0.0005 ppm. Each spectrum was normalized by dividing by the intensity of its highest peak, thereby normalizing all intensities to between 0 and 1 to assist with model stability.13C NMR spectra were computed with a line width of 1.50 Hz in proton decoupled mode. Chemical shifts were extracted from the spectrum and binned into 80 bins spanning the shift range of 3.42 to 231.3 ppm.
[0058] During training, we passed the processed1H NMR spectra through a one-dimensional convolutional neural network (CNN) to down sample the signal into a lower-dimensional space. For the processed13C NMR spectra, since it is a binary vector, we embedded it into a dense vector representation. If both spectra were being used as input to the spectrum-to-structure model, then their features were concatenated. If only one spectrum was being used as input, then only the features for that spectrum were retained. These features were then passed through a full encoder- decoder transformer for structure elucidation, which outputs SMILES strings, and a transformer encoder for substructure elucidation, which outputs substructure probability arrays as shown schematically on the top row of FIG. 1. In some embodiments, the substructures can be further ranked by their probabilities of being present or by some other metric.
[0059] In an embodiment, the four spectrum-to-structure and spectrum-to-substructure multitask models, which correspond to different conditions of input data and whether a pretrained transformer was used, shown in Table 2 of the main text were trained using the following procedure. Themodels were trained using a constant learning rate of 1 × 10−5and a batch size of 64 without any weight decay.
[0060] FIG. 8 shows training and validation loss curves over the first 100 epochs for two transformer models trained with a different splitting of the data and dropout enabled during validation. We see that in each case, the validation loss closely tracks the training loss and exceeds the training loss by the end of 100 epochs.
[0061] The loss function being optimized for each multitask model is a weighted sum of the cross entropy loss for the SMILES prediction and a binary cross entropy loss for the substructure prediction. Given a prediction-target pair for the SMILES prediction (ysmi,yˆsmi) and a prediction-target pair for the substructure prediction (ysub,yˆsub), the total loss iswhere α and β were
[0062] The SMILES strings, substructure arrays, and spectra was partitionedusing the same train-validation-test split as in the substructure-to-structure task. For each of the four models, early stopping was used to prevent overfitting by monitoring the loss on the validation set. The total number of epochs that each of the four models was trained for is shown in Table 5. The final model selected for each set of conditions was that with the lowest validation loss, with the number of epochs at which that loss was reached shown in Table 4 for each of the four models. The learning curves for the training of the spectrum-to-structure and spectrum-to-substructure multitask models are shown in FIG. 9. By comparing the learning curves, we note that using a pretrained transformer not only improves the structure elucidation accuracy, as described in the main text, but also accelerates convergence relative to training the multitask model from scratch.
[0063] Table 4 shows the number of epochs each multitask model was trained to attain the minimum validation loss. Data used Pretrained Transformer Total number of epochs Final model trained epoch 13C NMR Only Yes 779 261 1H NMR Only Yes 521 288 1H +13C NMR Yes 414 250 1H +13C NMR No 465 456Effect of spectrum normalization on model performance
[0064] Part of the data processing for the1H NMR spectra described in the main text is the normalization of the intensities of each such that the maximum intensity is 1 and theminimum intensity is 0. This feature was done to assist with the numerical stability when training the model. To investigate if this normalization has any effect on the multitask model’s performance, we trained 5 versions of the multitask models (1H NMR +13C NMR with the pretrained 3M transformer) with normalization of the1H NMR and 5 versions of themultitask model without normalization. All models the same of 142894 molecules withthe same training-validation-test split, with the only variation being the random seed used for initialization of the model. Table 5 shows the structure elucidation accuracy for all 10 training runs. The normalization had no impact on the substructure F1score, with each model obtaining an F1score of 0.86. Averaging across the 5 runs for each condition, we get an average accuracy of 70.2±0.84% with normalization and 71.4±0.90% without normalization. Thus, the average accuracy without normalization is within 1.2% of the average accuracy with normalization, which is within statistical fluctuations (i.e., 1.5 standard deviations) of the models arising from different initializations. Table 5: Structure accuracy of multitask models trained with and without normalization of the1H NMR.Structure Accuracy (%) With Structure Accuracy (%) Without Normalization Normalization 69.6 72.1 69.2 70.1 69.8 72.7 70.8 71.0 71.5 71.2 Early stopping on experimental data
[0065] As a test of the multitask model’s performance on experimental data, we used a set of 310 experimental spectra from our previous work to perform validation and testing of a multitaskmodel trained on simulated data. The experimental1H NMR spectra were processed as described above, while the experimental13C NMR spectra were binned using only 40 bins instead of 80 bins. To ensure consistency, we trained a new model a reprocessed version of the training and validation set where the13C NMR had only 40 bins instead of 80.
[0066] We trained a multitask model (1H NMR +13C NMR with the pretrained 3M transformer) using simulated data. Experimental data was used for the validation and test set, with 214 spectra in the validation set and 106 in the test set. During training, the validation loss on the experimental data was monitored and used to perform early stopping to mitigate overfitting. The learning curve is shown in FIG. 18. The model overfits very quickly to the simulated data, with the validation loss on the experimental data rapidly increasing after the first 30 epochs. The lowest validation loss was obtained at epoch 27, with a structure elucidation accuracy of 33.0% and a substructure F1 score of 0.60. While this is lower than the numbers reported in the main text on the simulated data (69.6% and 0.86, respectively), this emphasizes that significant improvements might be attainable if sufficient quantities of experimental data were used to train the model instead of only simulated data.
[0067]
[0070] Table 6 shows test set structure elucidation accuracy and substructure elucidation F1 score of the multitask model as a function of the type of data used and weight initialization for the transformer component. Table 6 shows the structure elucidation accuracy on the test set as a function of both the spectral data used (1H and / or13C NMR) and if weights from a transformer pretrained on the substructure-to-structure task were used; in other cases, randomly initialized weights were used. Structure elucidation accuracy was tested in the same way as for the substructure-to-structure task, where a target was considered correctly predicted if its canonical SMILES string appeared within the set of 15 predictions upon canonicalization. The highest accuracy of 69.6% (Table 6, third row) is obtained when using both the1H and13C NMR spectra with a pretrained transformer for structure elucidation. Without using the pretrained transformer, the structure elucidation accuracy drops significantly from 69.6% to 53.3%, demonstrating the importance of the pretraining step to the success of the overall multitask framework. This result suggests that pre-training on the substructure-to-structure task helps the model learn a chemically relevant latent space that translates well when integrated into the multitask framework. In some embodiments, a substructure has no more than 7 atoms, and the pretraining uses at least 50 substructures. In some embodiments, a substructure has no more than 7 atoms, and the pretraininguses at least 100 substructures.
[0068] Table 6 Data Used Pretrained Transformer Structure Accuracy (%) Substructure F1Score 13C NMR Only Yes 22.0 0.69 1H NMR Only Yes 59.6 0.81 1H +13C NMR Yes 69.6 0.86 1H +13C NMR No 53.3 0.86
[0069] To investigate the relative importance of the1H and13C NMR as inputs, we trained two versions of the model using either1H or13C NMR as the sole input. The accuracy of the model is significantly better with only1H NMR (59.6%) than with only13C NMR (22.0%), reflecting the fact that1H NMR contains far more information about the molecule’s connectivity relative to13C NMR, thus playing a more critical role in structure elucidation. However, the best accuracy is only obtained when combining both spectra, showing that they are complementary to each other in terms of the information they each provide. We also note that using1H NMR alone with the pretrained transformer gives a greater structure prediction accuracy than obtained using both1H and13C NMR data without the pretrained transformer, highlighting further the advantages ofapproach.
[0070] FIG. 2A shows in lighter bars the structure elucidation accuracy of the spectrum-to- structure multitask model as a function of the number of heavy atoms. In comparison to the substructure-to-structure model, there is a more systematic decrease in accuracy with the size of the molecule, decreasing from 77.5% at 10 heavy atoms to 52.0% at 19 heavy atoms. However, this decrease in accuracy is dwarfed by the increase in the problem size that grows combinatorially and thus increases over 5 orders of magnitude over that range of molecule sizes. Hence, not only does the model’s accuracy decay extremely slowly compared to the growth of the number of possible structures, the model is also able to achieve this impressive scaling without any reliance on the molecular formula, molecular weight, or other information about the system that would otherwise constrain the problem. The top section of FIG. 4 shows some of the molecules that were correctly predicted, emphasizing that the multitask model can elucidate the structures of molecules solely from their1H and13C NMR spectra across a wide range of chemical motifs that would be extremely challenging without additional information. These correctly predicted molecules were, in each case,one of the 15 predictions generated by the model. FIGS. 13 and 14 illustrate the other predictions made for two of these molecules, demonstrating that even the incorrect predictions are structurally similar to the correct prediction and therefore also a useful starting point for further analysis.
[0071] For the molecules where the spectrum-to-structure model was unable to predict the correct molecule, we can again use Tanimoto similarities to assess how close the predicted molecule was. The distribution of Tanimoto similarities of the best incorrect predictions for the spectrum-to- structure task is shown as a dashed line in FIG. 2B. From this, we see that compared to the substructure-to-structure task, the distribution of similarity scores is shifted to lower values, with only 64.7% of incorrect predictions having a similarity of 0.50 or greater with a mean similarity of 0.56. The lower similarity scores relative to the substructure-to-structure task reflect the fact that direct structure elucidation from only the NMR spectra is a considerably more challenging inverse problem that is less constrained than constructing the structure from a set of substructures. However, the bottom section of FIG. 4 shows that the molecules with the highest Tanimoto similarity obtained from incorrect predictions of the multitask model still contain many of the chemical motifs expected in the target compound, and so still provide insight into the system that is valuable when deducing the molecular connectivity.
[0072] The multitask model also outputs a prediction for the target molecule’s substructure profile, which can be interpreted as the probability of a given molecular fragment being present in the molecule based on its NMR spectra. These profiles, which provide additional information as to which of the 957 substructures are present in the system, are useful in cases where the model does not arrive at the correct structure since they could be used by a chemist to infer other possible structures of the molecule. To evaluate the substructure elucidation accuracy, one must be careful to take into account that most molecules only contain a relatively small fraction of the total 957 substructures and therefore the profiles are dominated by 0’s (which indicate the absence of a substructure). This leads to a highly imbalanced classification problem where it is insufficient to only use accuracy as a performance metric. To address this issue, Table 6 shows the F1score that balances the need to account for true and false positives as well as negatives. From this, we see that for this spectrum-to-substructure task, while the performance improvements in the F1 score of using both1H and13C NMR spectra over either individually are retained, the pretraining of the transformer has a negligible impact on the substructure elucidation. This is perhaps expected since although the multitask model produces both a prediction of structure and substructures in an end-to-end fashion, the pretrained transformer weights arise from the task of substructure-to-structure prediction, which is a downstream task from substructure prediction.
[0073] To quantify the model’s predictive performance on the spectrum-to-substructure task beyond measures like the F1 score, we can examine the distribution of the model’s predicted probabilities. FIG. 5 illustrates a distribution of true / false positives and true / false negatives as a function of the probability predicted by the multitask model using both1H and13C NMR of the test set as inputs and a pretrained transformer. A decision boundary of used to distinguish betweenpositives and negatives. In cases where the model predicts a low probability of a substructure being present (<0.1), the prediction is 99.8% accurate. Conversely, when the model predicts a high probability of a substructure being present (>0.9), the prediction is 96.3% accurate. The slightly lower accuracy in the case of predicting positive substructures arises from the significant imbalance of positives to negatives, since in the training data each molecule only contains a small fraction of the 957 substructures, and so there are far more true negatives than true positives in the dataset. The uncertainty of the model increases when moving away from the extremes of the predicted probability towards the decision boundary of 0.5 between positive and negative predictions; however, only 2.6% of substructure predictions have a probability within the range of 0.1 < p < 0.9, so the predicted probability from the model is a strong indicator of the presence or absence of a substructure the majority of the time. Using Morgan fingerprints as an alternative substructure encoding in an embodiment
[0074] One of the key limitations of the outlined approach is the restrictive nature of the original 957 set of substructures, which only enumerated motifs containing C, N, O, and H atoms and was intended to differentiate small molecules with no more than 10 heavy atoms. The above results demonstrated that this set was still effective at describing systems up to 19 heavy atoms in size, but expanding to larger systems with greater elemental coverage will require a more flexible and tunable representation of substructures.
[0075] To overcome these limitations, we present results using Morgan fingerprints. Unlike the pre-defined set of substructures, Morgan fingerprints dynamically generate binary vector encodings of molecules by applying a hashing algorithm to environments of a specified radius. An illustration of this is shown in FIG. 17. Not only are Morgan fingerprints fast to compute, but they scale easily to different element compositions and are tunable in the sense that both the size of the fingerprintand the radius of the encoded environments are user-selected parameters.
[0076] Using the Morgan fingerprints, we trained a new transformer on the substructure-to- structure task with the same dataset and train-validation-test split of 3M smiles, but this time feeding in the Morgan fingerprints rather than the substructure vectors. The results for this model are shown in FIG. 15, where we compare the performance of the fingerprint transformer model against the substructure transformer model using the error as a function of the system size in terms of the number of heavy atoms. From the results, it is clear that the fingerprint transformer performs better than the substructure transformer on recovering the exact molecular structure. Using this pretrained fingerprint transformer, we next trained the model on the spectrum-to-structure task, the results of which are presented in FIG. 16. While the use of the fingerprint transformer did not improve the results on the spectrum-to-structure task, it did not affect the performance of the model negatively, suggesting that using the morgan fingerprints is a sufficiently general strategy for expanding the size and element coverage of our workflow.
[0077] In some embodiments, a raw1H NMR spectrum is used. In some embodiments, a raw spectrum is a spectrum without any1processing. In some embodiments, a raw H NMR spectrum may be a1H NMR spectrum that is not processed other than being normalized before using the raw1Hspectral data to generate a tokenized NMR spectral representation, wherein the that encodes spectral features comprising at least one of peak positions, intensities, and coupling constants as a structured sequence. In some embodiments, a raw1H NMR spectrum may be a1H NMR spectrum that is not processed other than being normalized. The raw1H NMR spectralis then encoded into a latent spectral representation. In some embodiments, a 1H NMR spectrum may be the Free Induction Decay signal that comes from a NMR that is then encoded using tokenization or a neural network to generate a tokenized or latent NMR spectral representation, respectively. In some embodiments, the13C NMR spectrum is also raw. In some embodiments, the13C NMR spectrum is not raw, but instead has been preprocessed so that information extracted from a spectrum is encoded using tokenization to generate a tokenized NMR spectral representation or encoded using a neural network to generate a latent NMR spectral representation. When data is not raw, the data is subjected to a subjective manual interpretation, such as peak assignment. The ability of an embodiment to use raw data allows for a faster and more consistent result with less manual labor and without requiring supervision. In some embodiments, data from the13C NMR spectrum provided to generate atokenized NMR spectral representation is not raw since in such embodiments only binary peak information derived from the13C NMR spectrum is provided to generate a tokenized spectral representation. In some embodiments, the full spectrum of a13C NMR spectrum is provided as raw data to generate a tokenized NMR spectral representation or a latent NMR spectral representation through the use of a neural network. In some embodiments, a raw13C NMR spectrum may be the Free Induction Decay signal that comes from an NMR spectrometer that is then encoded using tokenization or a neural network to generate a tokenized or latent NMR spectral representation, respectively.
[0078] To provide an example of a system for generating a prediction of the molecular formula and atomic connectivity of an unknown chemical compound in an embodiment, FIG. 19 is a high level block diagram showing a computer system 1900 that is suitable for implementing an embodiment. The computer system may have many physical forms, ranging from an integrated circuit, a printed circuit board, and a small handheld device, up to a huge supercomputer. The computer system 1900 includes one or more processors 1902, and further can include an electronic display device 1904 (for displaying graphics, text, and other data), a main memory 1906 (e.g., random access memory (RAM)), storage device 1908 (e.g., hard disk drive), removable storage device 1910 (e.g., optical disk drive), user interface devices 1912 (e.g., keyboards, touch screens, keypads, mice or other pointing devices, etc.), and a communication interface 1914 (e.g., wireless network interface). The communication interface 1914 allows software and data, such as a spectrum, to be transferred between the computer system 1900 and external devices via a link. The system may also include a communications infrastructure 1916 (e.g., a communications bus, cross- over bar, or network) connected to the aforementioned devices / modules.
[0079] Information transferred via communications interface 1914 may be in the form of signals such as electronic, electromagnetic, optical, or other signals capable of being received by communications interface 1914, via a communication link that carries signals and may be implemented using wire or cable, fiber optics, a phone line, a cellular phone link, a radio frequency link, and / or other communication channels. With such a communications interface, it is contemplated that the one or more processors 1902 might receive information from a network or might output information to the network in the course of performing the above-described method steps. Furthermore, method embodiments may execute solely upon the processors or may execute over a network, such as the Internet, in conjunction with remote processors that share a portion ofthe processing.
[0080] The term “non-transient computer readable medium” is used generally to refer to media such as main memory, secondary memory, removable storage, and storage devices, such as hard disks, flash memory, disk drive memory, CD-ROM, and other forms of persistent memory, and shall not be construed to cover transitory subject matter, such as carrier waves or signals. Examples of computer code include machine code, such as one produced by a compiler, and files containing higher level code that are executed by a computer using an interpreter. Computer readable media may also be computer code transmitted by a computer data signal embodied in a carrier wave and representing a sequence of instructions that are executable by a processor.
[0081] In some embodiments, the link connects a spectrometer 1924 to the communications interface 1914, allowing for raw spectra to be fed from the spectrometer 1924 to the computer system 1900 without any manual processing. In some embodiments, the system comprises one or more processors and memory that stores instructions. The instructions, when executed by the one or more processors, cause the system to perform operations comprising receiving, as input, nuclear magnetic resonance (NMR) spectral data comprising at least one of a 13C NMR spectrum and a 1H NMR spectrum corresponding to the unknown chemical compound, preprocessing the received NMR spectral data to generate a tokenized spectral representation, wherein the tokenization encodes spectral features comprising at least one of peak positions, intensities, and coupling constants as a structured sequence, providing the tokenized spectral representation to a machine learning model pretrained on a plurality of tokenized spectral-structural pairs (spectra-structure pairs), wherein each pair comprises (1) a tokenized NMR spectral representation and (2) a corresponding molecular structure and connectivity, processing, by the machine learning model, the tokenized spectral representation of the unknown compound to generate a list of one or more candidate molecular formulas and associated atomic connectivity structures, and outputting the list of predicted candidate molecular formulas and connectivity structures for the unknown chemical compound. In some embodiments, candidate molecular structures are used to predict NMR spectra. The predicted NMR spectra for the candidate molecular structures are compared to the input NMR spectral data in order to evaluate prediction accuracy. Conclusion
[0082] In summary, some embodiments provide a multitask machine learning approach fordirect structure and substructure elucidation from1H and13C NMR spectra that leverages pretraining an encoder-decoder transformer on the related task of substructure-to-structure elucidation. By integrating this transformer into our multitask framework, we have shown that our end-to-end model is able to predict the structure correctly 69.6% within the first 15 predictions using only the1H and13C NMR as input for systems of up to 19 heavy atoms. Furthermore, our end-to-end model can simultaneously predict which substructures are present in amolecule with an accuracy of 96.3% when the predicted probability is above 0.9 and which substructures are not present with 99.8% accuracy when the predicted probability is below 0.1, with only 2.6% of the predicted substructure probabilities not falling within those ranges. What is particularly remarkable is that while the problem size (number of possible molecules that can be constructed consistent with basic chemical bonding rules) grows combinatorially with the number of heavy atoms from 10 to 19 heavy atoms our multitask model shows only a 25.5% decrease in accuracy over a range in which the number of possible molecules increases by 5 orders of magnitude. This suggests that our approach to the inverse problem of spectrum-to-structure is scalable to larger chemical systems. Our model thus provides an efficient avenue for direct spectrum-to-structure elucidation from1H and13C NMR spectra without dependence on any prior chemical information such as the molecular formula or molecular fragments.
[0083] Some embodiments elucidate even larger molecules with an extended range of elements and the prediction of stereochemistry. Such embodiments expand the set of substructures used to pretrain the transformer with substructures containing additional elements and larger, more complicated chemical motifs, such as larger ring systems or protecting groups. Relative stereochemistry could be incorporated within our existing multitask framework by training on SMILES representations containing the characters specifying stereocenters and cis- and trans- double bond configurations, and the corresponding NMR spectra of molecules. This could be achieved by identifying all the stereocenters in the current training set and enumerating all possible stereoisomers and using these to augment the training set. In some embodiments, an alternative approach to tackling the stereochemistry problem is to predict composition and connectivity as shown here and then do forward prediction of the spectra of different stereoisomers of candidate structures and compare to the input spectra to identify the most likely stereoisomer.
[0084] Although in this work we have concentrated on absolute structure prediction with no chemical knowledge of the compound beyond its NMR spectra, this framework could also be easilyadapted for cases where some information about the chemical system is available, such as in the case of reaction prediction or retrosynthesis where the starting materials of a reaction are known, and hence the problem is considerably more constrained. We also envision this technology to be used in a complementary manner with other existing structure elucidation, reaction prediction, or retrosynthesis frameworks as a way to rapidly detect structural or substructural changes in a reaction pathway or as a way to quickly constrain the number of possible candidate structures to a point where more refined techniques can be applied tractably.
[0085] Our approach opens up new possibilities in the field of ML-driven structure elucidation by introducing a fast and efficient structure elucidation framework that can operate in an unsupervised manner without relying on additional knowledge. Some embodiments can make a full prediction of the structure and substructures of an input1H and13C NMR spectra for a system of 19 heavy atoms in under 3 seconds, even on standardhardware (a single core of an AMD Ryzen 73700X 8-core processor). This framework thus has the potential to provide a highly accessible technology to greatly accelerate characterization and chemical discovery at levels ranging from high school chemical education to industrial research settings.
[0086] Acquiring large datasets of high-quality experimental NMR spectra in the form needed to train our model (i.e., original first input delay (FID) files) currently poses a challenge but is essential to creating a model that encodes the nuances of real spectra that are not fully captured by the simulated spectra used to train the current model. For example, the accuracy of the framework drops from 69.6% to 33.0% when used to predict the structures from 106 experimental1H and13C NMR spectra of molecules employed in our previous work, emphasizing the potential accuracy improvements that could be achieved on experimental spectra if more training data were available. In some embodiments, a community-driven effort may aid in curating a large database of experimental NMR spectra, which will help the model learn to predict structure from spectra collected under a variety of conditions. This synergistic effort of assisting community chemical characterization efforts in tandem with improving the prediction model provides an opportunity to greatly improve the accuracy of this tool while supporting the chemistry community.
[0087] While this invention has been described in terms of several preferred embodiments, there are alterations, permutations, modifications, and various substitute equivalents which fall within the scope of this invention. It should also be noted that there are many alternative ways of implementing the methods and apparatuses of the present invention. It is therefore intended that thefollowing appended claims be interpreted as including all such alterations, permutations, modifications, and various substitute equivalents as fall within the true spirit and scope of the present invention. As used herein, the phrase “A, B, or C” should be construed to mean a logical (“A OR B OR C”), using a non-exclusive logical “OR,” and should not be construed to mean ‘only one of A or B or C. Each step within a process may be an optional step and is not required. Different embodiments may have one or more steps removed or may provide steps in a different order. In addition, various embodiments may provide different steps simultaneously instead of sequentially.
Claims
CLAIMS What is claimed is:
1. A computer-implemented method for generating a prediction of a molecular formula and molecular structure of an unknown chemical compound, the method comprising: a) receiving, as input, nuclear magnetic resonance (NMR) spectral data of the unknown chemical compound comprising at least one of a13C NMR spectrum and a1H NMR spectrum corresponding to the unknown chemical compound; b) providing the NMR spectral data to a machine learning framework pretrained on a plurality of spectra-structure pairs, wherein each pair comprises (i) a NMR spectral representation and (ii) a corresponding molecular structure or molecular formula; c) processing, by the machine learning model, the NMR spectral data corresponding to the unknown chemical compound to generate a list of one or more candidate molecular formulas and corresponding candidate molecular structures; and d) outputting the list of candidate molecular formulas and corresponding candidate molecular structures for the unknown chemical compound.
2. The method, as recited in claim 1, further comprising before step b ) preprocessing the received NMR spectral data to generate: (i) a tokenized spectral representation for the13C NMR spectrum wherein the tokenization encodes spectral features comprising of at least one of peak positions, intensities, and coupling constants of the13C NMR spectrum as a structured sequence or (ii) a latent spectral representation for the1H NMR spectrum wherein the latent spectral representation is generated through a neural network that takes in raw1H NMR spectral data of the1H NMR spectrum; and wherein the providing the NMR spectral data to the machine learning framework pretrained on a plurality of spectra-structure pairs, provides the tokenized spectral representation or latent spectral representation to a machine learning framework pretrained on a plurality of spectra- structure pairs, wherein each pair comprises (i) at least one of a tokenized spectral representation for the13C NMR spectrum and a latent spectral representation for the1H NMR and (ii) a corresponding molecular structure.
3. The method as recited in claim 2, before step a pretraining the machine learning model.
4. The method, as recited in claim 3, wherein the pretraining trains on a plurality of spectra- structure pairs, wherein each pair comprises (i) at least one of a tokenized spectral representation forthe13C NMR spectrum and a latent spectral representation for the1H NMR and (ii) a corresponding molecular structure.
5. The method, as recited in claim 3, wherein thesimulated NMR spectral data.
6. The method, as recited in claim 3, wherein the pretraining comprises inputting a binary vector indicating a presence or absence of a plurality of predefined molecular substructures.
7. The method, as recited in claim 6, wherein the pretraining uses at least 50 substructures.
8. The method, as recited in claim 3, wherein the pretraining uses Morgan fingerprints.
9. The method, as recited in claim 2, wherein the one or more candidate molecular formulas and candidate molecular structures are represented as SMILES strings.
10. The method, as recited in claim 2, wherein the preprocessing the received NMR spectral data to generate a tokenized spectral representation uses a convolutional neural network.
11. The method, as recited in claim 1, wherein the NMR spectral data of the unknown compound is at least one of one-dimensional13C NMR spectra and1H NMR spectra.
12. The method, as recited in claim 1, further comprising a ranked list ofsubstructures of the unknown compound.
13. The method, as recited in claim 12, wherein the outputting the ranked list of substructures of the unknown chemical compound is generated from the prediction of the substructures of the unknown compound.
14. The method, as recited in claim 12, wherein the substructures correspond to fragments of the unknown chemical compound with defined bonding relationships and can be represented as SMARTS strings.
15. The method, as recited in claim 1, wherein the NMR spectral data of the unknown compound comprises raw1H NMR spectra.
16. The method, as recited in claim 1, further comprising predicting the NMR spectra for the candidate molecular structures and comparing the predicted NMR spectra for the candidate molecular structures to the input NMR spectral data.
17. A system for generating a prediction of a molecular formula and molecular structure of an unknown chemical compound, the system comprising: a) one or more processors; and b) a non-transient computer readable medium storing instructions that, when executed by the one or more processors, cause the system to perform operations, comprising:i) receiving, as input, nuclear magnetic resonance (NMR) spectral data of the unknown chemical compound comprising at least one of a13C NMR spectrum and a1H NMR spectrum corresponding to the unknown chemical compound; ii) providing the NMR spectral data to a machine learningpretrained on a plurality of spectra-structure pairs, wherein each pair comprises (i) a NMR spectral representation and (ii) a corresponding molecular structure or molecular formula; iii) processing, by the machine learning model, the NMR spectral data corresponding to the unknown chemical compound to generate a list of one or more candidate molecular formulas and corresponding candidate molecular structures; and iv) outputting a list of candidate molecular formulas and corresponding candidate molecular structures for the unknown chemical compound.
Citation Information
Patent Citations
Molecular structure determination from NMR spectroscopy
WO2010026418A1
Deep imitation learning for molecular inverse problems
WO2021091883A1
Method for structure elucidation
WO2023135113A1