Structure elucidation method

JP2025502880A5Pending Publication Date: 2026-01-16BOEHRINGER INGELHEIM INT GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024541806
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-01-12
Filing Date
2023-01-10
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing methods, such as NMR spectroscopy, struggle to directly determine the geometric or molecular structure of unknown compounds from their measurement spectra, requiring expert intervention and being time-consuming.

Method used

A fully automated method using machine learning models, specifically graph neural networks and residual neural networks, to generate candidate structures and predictive spectra from molecular formulas, allowing for rapid and reliable structural elucidation.

Benefits of technology

Enables fast, efficient, and expert-free structural elucidation of unknown compounds by generating and comparing predictive spectra with measured spectra, improving success rates and reducing reliance on human expertise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention relates to a method for elucidating the structure of an unknown compound from a measured spectrum of a sample, the method comprising at least one machine learning model, in particular a first machine learning model for generating a structure of the compound and / or a second machine learning model for generating a predicted spectrum from the structure.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a method for structure elucidation for elucidating the structure or molecular structure of an unknown compound from a measured spectrum of a sample, as well as to a data processing apparatus, a computer program product and a computer readable storage medium. [Background technology]

[0002] In chemistry and pharmacy, it is often important to analyze the chemical composition of a sample or to analyze the compounds contained in a sample. For example, the sample may be or have a drug that is to be analyzed for potential impurities. In another example, the sample may be or have a newly synthesized compound that is to be verified. Several methods are available for analyzing samples, including nuclear magnetic resonance (hereinafter abbreviated as NMR) spectroscopy, mass spectrometry, infrared spectroscopy, Raman spectroscopy, and X-ray crystallography.

[0003] In organic chemistry and pharmacology, NMR spectroscopy is one of the most used methods to identify compounds of a sample. Although this method has many advantages, it is not possible to directly infer the geometric or molecular structure of the measured compound from the NMR spectrum.

[0004] Although NMR spectroscopy experiments typically provide a complete description of the hydrocarbon backbone of an organic molecule, when the measured spectrum does not match known spectra of known compounds, it can be very difficult to make a perfect one-to-one assignment of observed features in the spectrum (especially chemical shifts) to individual atomic nuclei in the sample under investigation.

[0005] For example, a typical task in the manufacture of a pharmaceutical or drug or other chemical product is to verify that the desired pharmaceutical or drug or chemical product has in fact been synthesized or manufactured, and / or to verify whether impurities are present in a sample, and if so, to determine the impurities, by comparing a measured spectrum, in particular an NMR spectrum, of the sample with the expected spectrum of compounds contained or expected in the sample.

[0006] If the measured spectrum matches the expected known spectrum of the compound contained in the sample, it is very easy to confirm the presence of the compound of interest in the sample. However, it is possible that an unknown compound is present in the sample, for example, if an impurity is present or if the synthesis does not result in the desired compound. In this case, it can be very difficult to assign the measured spectrum to the compound.

[0007] In the context of this disclosure, the term "structure elucidation" refers to the process or method of determining the geometric or molecular structure of compounds contained in a sample from the measured spectrum of the sample.

[0008] The challenge of structure elucidation is to find the molecular structure that best matches the measured spectra, especially the NMR spectra, of a sample. However, structure elucidation is a very laborious and time-consuming task, and usually requires experts with a lot of experience and expertise. [Prior art documents] [Non-patent literature]

[0009] [Non-Patent Document 1] J.Chem.Inf.Model.2012,52,7,1757-1768(Published on May 15, 2012) [Non-Patent Document 2] "Generative Adversarial Nets" by Ian J. Goodfellow et al., Advances of Neural Information Processing Systems 27 (NIPS 2014) [Non-Patent Document 3] "Graph attention networks" by Velicovic et al. https: / / arxiv.org / abs / 1710.10903 [Non-Patent Document 4] He et al., "Deep residual learning for image recognition," Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770-778 Summary of the Invention [Problem to be solved by the invention]

[0010] Therefore, a fully automated method for structure elucidation is desirable.

[0011] It is an object of the present invention to provide a method for structure elucidation that is fully automated, rapid and / or reliable. [Means for solving the problem]

[0012] The above object is solved by a method according to claim 1, a data processing device according to claim 13, a computer program product according to claim 14 or a computer-readable storage medium according to claim 15. Advantageous developments are the subject matter of the dependent claims.

[0013] In particular, the present invention relates to a method for elucidating the structure of an unknown compound from the measured spectrum of a sample. The term "structure" refers in particular to the molecular structure of a compound or its molecules and is further defined below.

[0014] In the method according to the invention, structures of candidate compounds are generated and predicted spectra are generated from these generated structures. The predicted spectra are compared to the measured spectra. In particular, based on the comparison, one of the predicted spectra is selected and the structure corresponding to the selected predicted spectrum is determined as the structure of the unknown compound.

[0015] According to a first aspect, a first machine learning model generates a structure of a compound, and a second machine learning model generates a predicted spectrum from the structure generated by the first machine learning model. This makes the structure elucidation very fast, reliable and efficient. In particular, the method of structure elucidation can be fully automated or at least largely automated, and the need to consult an expert for structure elucidation can be eliminated.

[0016] According to another aspect, which can be implemented independently, a machine learning model (hereinafter referred to as a first machine learning model) generates structures of candidate compounds, and the first machine learning model is trained to generate realistic structures from molecular formulas and / or empirical formulas. This contributes to fast, reliable and efficient structure elucidation. In particular, at least a part of the structure elucidation can be automated, i.e., possible compounds or candidate compounds can be provided, in other words, candidates for unknown compounds can be identified.

[0017] According to another aspect, which can be implemented independently, a machine learning model (hereinafter referred to as the second machine learning model) generates a predicted spectrum from a molecular structure, and the second machine learning model has a residual neural network. This contributes to fast, reliable and efficient structure elucidation. In particular, the structure elucidation can be at least partially automated. In particular, the candidate compound or the spectrum of the candidate compound can be automatically generated.

[0018] Advantageously, the first machine learning model comprises and / or uses one or more artificial neural networks, preferably one or more graph neural networks, in particular generative adversarial networks. The use of graph neural networks or generative adversarial networks has proven to be particularly suitable and efficient for generating structures.

[0019] Preferably, the first machine learning model is trained using a first training data set. Preferably, the first machine learning model is trained for the generation of realistic structures. In particular, the first machine learning model is trained for the generation of realistic structures from molecular formulas and / or empirical formulas. The first machine learning model preferably comprises a first training data set, and / or the first training data set preferably forms part of the first machine learning model. By training the first machine learning model using the first training data set, it can be achieved that the first machine learning model generates only chemically and / or physically possible structures. In particular, if the first training data set is sufficiently large and diverse, the first machine learning model can be preferably trained to generate novel structures, i.e. structures not included in the first training data set.

[0020] The first training data set preferably comprises structures of a plurality of real molecules. In this way, the first machine learning model can be effectively trained to generate only realistic structures. The generation of chemically and / or physically impossible structures is preferably avoided or at least significantly reduced.

[0021] It is advantageous if the first machine learning comprises a generator and preferably a discriminator and / or to train the first machine learning model with the generator and preferably the discriminator. The generator is in particular an object that generates structures from molecular formulas and / or empirical formulas and / or is trained to generate structures. The discriminator is in particular an object that is trained to distinguish and / or to distinguish between real structures, in particular structures from the first training data set, and other or artificial structures, in particular structures generated by the generator. In particular during training, the generator is fed with the results and / or decisions of the discriminator, and the discriminator is fed with both structures, in particular structures from the first training data set and structures generated by the generator. In this way, the generator and the discriminator are trained mutually.

[0022] In this way, on the one hand, the generator can be trained to generate only realistic and / or physically and / or chemically possible structures, while on the other hand, the discriminator is thus trained to learn what "real" molecules look like and to distinguish between real molecules and artificial molecules or structures generated by the generator.

[0023] In particular, the generator and the discriminator are trained mutually, such that after training, a generator is obtained that produces realistic and / or chemically and / or physically possible molecules or structures.

[0024] The discriminator is preferably used only in training and / or not in the actual structure elucidation and / or not in the application stage. Here, the term "actual structure elucidation" refers specifically to the step of generating a molecule after training of the first machine learning model has been completed.

[0025] The first machine learning model, in particular the generator, is preferably trained to generate and / or generate multiple structures from a given molecular and / or empirical formula, which can improve the success rate of the structure elucidation method.

[0026] The first machine learning model, in particular the generator, preferably has and / or uses a random noise generator and / or a random variable. The random noise generator or the random variable can in particular achieve or guarantee that a number of different structures are generated from one given molecular formula and / or empirical formula. In particular, the use of the random noise generator or the random variable ensures the diversity of the structures generated. This contributes to a high success rate of structure elucidation.

[0027] The discriminator is preferably trained to distinguish between real structures, particularly from the first training data set, and structures generated by the generator, which allows to efficiently train the generator for the generation of realistic and / or chemically and / or physically possible structures.

[0028] The second machine learning model preferably comprises and / or uses one or more artificial neural networks. Particularly preferably, the second machine learning model comprises and / or uses a graph neural network, a graph attention network and / or a residual neural network. The use of an artificial neural network, in particular a graph attention network and / or a residual neural network, has proven to be advantageous in the generation of predicted spectra, in particular NMR spectra, from a given structure. This allows for a fast and / or reliable generation of predicted spectra, in particular predicted spectra with high accuracy. In particular, the use of a graph neural network and / or a residual neural network allows the number of external training samples required to be reduced and spectra, in particular NMR spectra, of molecules of unlimited size to be calculated. Furthermore, the generation of predicted spectra can be parallelized. Surprisingly, due to the use of a graph neural network, a graph attention network and / or a residual neural network, the structure elucidation method and / or the second machine learning model of the present invention are able to predict different chemical shifts of diastereotopic protons in NMR spectra.

[0029] The second machine learning model is preferably trained using a second training dataset, which preferably includes data that is different from the first training dataset and / or has a different structure and / or different information than the data from the first training dataset.

[0030] In particular, the second training data set comprises structures labelled with associated spectral features. In other words, for every structure in the second training data set, the second training data set (additionally) comprises information on the features that this structure will have when its spectrum is measured. Preferably, the spectra are NMR spectra and / or the spectral features are chemical shifts, in particular 1 H and / or13 C chemical shifts.

[0031] In generating the predicted spectrum, it is preferred that for each spectral feature, in particular each chemical shift of the spectrum, an expectation or mean value and a corresponding measure of dispersion, in particular the standard deviation, are calculated, which makes it possible to predict different spectral features, in particular chemical shifts, of diastereotopic protons.

[0032] Preferably, the measured spectrum is an NMR spectrum. In particular, a spectrum, in particular an NMR spectrum, of a sample is measured. However, measuring a spectrum is not an essential feature of the method. It is also possible to carry out the method without an explicit step of measuring a sample, for example, if the sample has already been measured prior to and / or independently of the method and / or if the measured spectrum exists as a data set. Thus, the step of measuring the spectrum of the sample can be separate from the method of the invention, and is preferably a step preceding it.

[0033] The molecular formula and / or empirical formula is preferably determined by measuring the mass spectrum of a sample.

[0034] The method according to the invention is preferably a computer-implemented method, thus allowing the method to be at least partially or fully automated.

[0035] According to another aspect, the invention relates to a data processing device comprising means for implementing the method.

[0036] According to another aspect, the invention relates to a computer program product comprising instructions which, when executed by a computer, cause the computer to carry out the method.

[0037] According to another aspect, the invention relates to a computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to perform the method.

[0038] A "molecular structure" in the sense of the present disclosure is preferably the geometrical structure of a molecule, in particular the geometrical and / or three-dimensional arrangement of the atoms of a molecule.

[0039] The "empirical formula" of a compound in the sense of the present disclosure is preferably the simplest integer ratio of atoms present in the compound. An empirical formula does not refer to the arrangement or number of atoms. In particular, the total number of atoms of a given compound cannot be inferred from its empirical formula. As a simple example of this concept, the empirical formula of sulfur monoxide (SO) is simply SO, as is the empirical formula of disulfur dioxide (SO2). Thus, sulfur monoxide and disulfur dioxide have the same empirical formula (SO), but sulfur monoxide has only one sulfur atom and one oxygen atom, whereas disulfur dioxide has two sulfur atoms and two oxygen atoms.

[0040] A "molecular formula" in the sense of this disclosure preferably indicates the number of each type of atom in the molecule of a compound. The molecular formula is the same as the empirical formula of a molecule that has only one atom of a particular type. In other cases, the molecular formula can have a larger number. In the above example, the molecular formula of sulfur monoxide is SO, which is the same as the empirical formula. In the case of disulfur dioxide, the molecular formula is S2O2, which is different from the empirical formula SO.

[0041] The empirical and molecular formulas do not provide any information regarding the molecular structure or geometric structure of a molecule.

[0042] A "compound" in the sense of the present disclosure is preferably a chemical compound composed of multiple identical molecules.

[0043] The "structure" of a compound or its molecules in the sense of the present invention is preferably the molecular structure of the compound or its molecules, in other words the two-dimensional and / or three-dimensional arrangement of the individual atoms of the molecule. The term "structure" therefore in particular denotes the arrangement of the atoms or nuclei of the compound or its molecules. The structure or molecular structure can also exist or be represented as a formula representing the geometrical / molecular structure, in particular a structural formula or a skeletal formula, etc. For example, the structural formula of ethanol (C2H6O) is: [ka] This can be read as follows.

[0044] A "candidate compound" in the sense of the present disclosure is preferably a compound that is a candidate for an unknown compound in a sample whose structure is unknown and / or to be elucidated. In other words, a candidate compound is a compound that may or is likely to be an unknown compound. In particular, the empirical formula and / or molecular formula of the candidate compound is the same as the empirical formula and / or molecular formula of the unknown compound, which has or can be measured, in particular by performing mass spectrometry of the sample and / or unknown compound.

[0045] A "measured spectrum" in the sense of the present disclosure is preferably the result of a measurement of a sample, for example by NMR spectroscopy, mass spectrometry, infrared spectroscopy, Raman spectroscopy, X-ray crystallography or similar spectroscopic methods. However, particularly preferably, the spectrum is an NMR spectrum. A measured spectrum is preferably present as and / or represented by a set of data or data points, in particular digital data. The term "measured spectrum" is used in particular to distinguish a measured spectrum from an expected spectrum, which is explained below.

[0046] A "predicted spectrum" in the sense of the present disclosure is preferably a spectrum that is not actually measured and / or is (artificially) generated, in particular by a machine learning model and / or other computer program, module and / or algorithm. In particular, the predicted spectrum is generated by a second machine learning model. The predicted spectrum preferably exists as and / or is represented by a set of data or data points, in particular digital data. In particular, the predicted spectrum has the same data type and / or data structure as the measured spectrum.

[0047] A "nuclear magnetic resonance spectrum" (hereinafter abbreviated as "NMR spectrum") in the sense of the present disclosure is preferably a spectrum measured by NMR spectroscopy or a predicted spectrum having the same data type or data structure as a measured NMR spectrum. In particular an NMR spectrum comprises or consists of a number of features. These features are in particular peaks in the spectrum. In the context of NMR, the spectral features and / or peaks are in particular denoted as "chemical shifts". The chemical shift is the resonance frequency of a nucleus in a magnetic field relative to a reference. The chemical shift is preferably expressed in ppm.

[0048] A "machine learning model" in the sense of the present disclosure preferably comprises and / or uses a machine learning algorithm and one or more (training) datasets. In other words, a machine learning model is preferably the output of a machine learning algorithm executed on one or more (training) datasets. In particular, a machine learning model stands for something learned by a machine learning algorithm. An algorithm is in particular a procedure executed on one or more (training) datasets to create a machine learning model. An algorithm may in particular be an artificial neural network, in particular preferably a graph neural network.

[0049] A "training dataset" in the sense of the present disclosure is preferably a dataset used to train a machine learning model and / or a machine learning algorithm.

[0050] An "artificial neural network" in the sense of the present disclosure is preferably a computer learning system that uses a network of functions that can understand a data input in one form and transform it into a desired output, usually in another form. A neural network is composed of at least two layers, preferably three or more layers. In particular, an artificial neural network has an input layer, an output layer, and one or more hidden layers, i.e. layers between the input layer and the output layer. Each layer contains one or more units called “neurons.” The concept of artificial neural networks is inspired by the brain and the way it learns.

[0051] A "graph neural network" in the sense of the present disclosure is preferably an artificial neural network that has and / or makes use of and / or can be directly applied to graphs.

[0052] A "graph" in the sense of the present disclosure is preferably a data structure consisting of nodes and edges. In the context of artificial neural networks, the nodes of the graph can represent entities such as different people, and the edges of the graph can represent relationships or links between the nodes, such as personal relationships between people. In another example, the nodes can represent the nuclei of molecules, and the edges can represent chemical bonds between the nuclei. In the implementation of machine learning models, the graph can be represented as a matrix, in particular an adjacency matrix.

[0053] A "feature vector" in the sense of the present disclosure is preferably an array or vector assigned to a node. A feature vector has one or more elements that contain information about the node.

[0054] An "embedding" in the sense of the present disclosure is preferably a low-dimensional space or vector into which a high-dimensional vector can be transformed. Embeddings can in particular be learned. In particular, node embeddings are particularly low-dimensional vector representations of the nodes in a graph.

[0055] The above-mentioned aspects and features of the present invention and the aspects and features of the present invention that will become apparent from the claims and the following description can in principle be implemented independently of one another, but can also be implemented in any combination or order.

[0056] Further aspects, advantageous aspects, features and characteristics of the present invention will become apparent from the claims and the following description of preferred embodiments with reference to the drawing, in which FIG. 1 shows a schematic representation of the method according to the invention.

[0057] The structure elucidation method according to the invention is shown diagrammatically in Figure 1. The method is in particular a method for elucidating a structure 1, in particular a geometric and / or molecular structure 1, of an unknown compound 2 from a measured spectrum 3 of a sample 4.

[0058] The method is preferably used to discover and / or determine contamination of a sample 4. In this case, the unknown compound 2 is a contaminant. The sample 4 in this case is preferably a drug or pharmaceutical agent.

[0059] Preferably, the method is a computer-implemented method.

[0060] First, a brief overview of the method is given, followed by a detailed description of the different method steps. [Brief description of the drawings]

[0061] [Figure 1] 1 illustrates diagrammatically the structure elucidation method according to the present invention; DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0062] overview First, a spectrum 3 of the sample 4 is preferably measured. Hereinafter, this spectrum 3 is referred to as a measured spectrum 3.

[0063] Although the presence of a measured spectrum 3 of sample 4 is a necessary requirement for carrying out the method, the actual step of measuring spectrum 3 is not an essential feature of the method. In particular, the step of measuring spectrum 3 of sample 4 can be carried out separately from, and in particular prior to, the method.

[0064] The measured spectrum 3 is particularly preferably an NMR spectrum of the sample, but in principle it may also be a spectrum 3 measured by other spectroscopic methods than NMR.

[0065] The sample 4 preferably is or comprises a drug or pharmaceutical agent, but can be any sample 4 that can be analysed by spectroscopy, in particular NMR. The sample 4 preferably comprises at least one active ingredient.

[0066] The sample 4 preferably comprises an unknown compound 2. The unknown compound 2 gives rise to features in the measured spectrum 3 that cannot be assigned to known or expected compounds (hypotheses) contained in the sample 4. The presence of features in the measured spectrum 3 that cannot be assigned to known and / or expected compounds therefore suggests that the unknown compound 2 is contained in the sample 4. This may be due to the presence of impurities in the sample 4 and / or due to a synthesis that did not or only did not yield the expected or desired compound to be synthesized.

[0067] Sample 4 may be a sample from which unknown unknown compound 2 has been isolated and / or enriched.

[0068] The measured spectrum 3 preferably comprises one or more spectral features 5. The spectral features 5 are preferably peaks in the measured spectrum 3. In particular, in the case of an NMR spectrum, the spectral features 5 are chemical shifts.

[0069] Preferably, the molecular formula and / or empirical formula 6 of the unknown compound 2 is determined. This can be done, for example, by performing mass spectrometry of the sample 4 and / or of the particularly isolated and / or enriched unknown compound 2. The step of determining the molecular formula and / or empirical formula 6 of the unknown compound 2 is not an essential feature of the method of the invention, but preferably precedes it.

[0070] For structure elucidation, a structure 1 of a candidate compound is preferably generated. In particular, a number of different structures 1 of the candidate compound are generated from one or the same molecular formula and / or empirical formula 6. This is preferably done by a machine learning model 7 (hereinafter, specifically referred to as the first machine learning model 7).

[0071] From the generated structures 1, preferably predicted spectra 8 are generated. In particular, exactly one predicted spectrum 8 is generated for each generated structure 1. This is preferably done by a machine learning model 9 (hereinafter, specifically referred to as the second machine learning model 9).

[0072] The terms “first” and “second” machine learning model do not imply any hierarchy between the machine learning models 7 and 9, but serve to distinguish the machine learning models. Thus, the prefixes “first” and “second” may be omitted. Thus, the first machine learning model 7 may be referred to as machine learning model 7, and the second machine learning model 9 may be referred to as machine learning model 9.

[0073] In particular, it is also possible that the method comprises / uses only one machine learning model, i.e. only the (first) machine learning model 7 or only the (second) machine learning model 9.

[0074] The first machine learning model 7 and / or the second machine learning model 9 preferably (each) comprises an algorithm, in particular a machine learning algorithm, and / or a training data set. Preferably, the first machine learning model 7 comprises a different algorithm than the second machine learning model 9 and / or the first machine learning model 7 comprises a different training data set than the second machine learning model 9. Preferably, the algorithm of the first machine learning model 7 and / or the second machine learning model 9 is an artificial neural network, in particular a graph neural network.

[0075] The predicted spectrum 8 is preferably of the same type and / or has the same data structure as the measured spectrum 3. For example, if the measured spectrum 3 is an NMR spectrum, then the predicted spectrum 8 is also an NMR spectrum. In particular, the (only) difference between the predicted spectrum 8 and the measured spectrum 3 is that the measured spectrum 3 is or has actually been measured using a sample 4, whereas the predicted spectrum 8 is artificially generated and / or calculated, in particular by a second machine learning model 9. The different terms "measured spectrum 3" and "predicted spectrum 8" mainly only serve two purposes: to distinguish between the spectrum 3 measured using a sample and the spectrum 8 artificially generated and / or calculated.

[0076] Structure 1 is in particular a molecular structure, or in other words the geometrical, two-dimensional and / or three-dimensional structure of a molecule 1. Structure 1 is preferably present as or represented by data defining the relative positions of atoms or nuclei within a molecule, but may also be present as or represented by a structural formula, a skeletal formula or other suitable data and / or formula.

[0077] Structure 1 is preferably calculated or generated starting from the molecular formula and / or empirical formula 6.

[0078] After generating the predicted spectrum 8, the predicted spectrum 8 is preferably compared to the measured spectrum 3. This is preferably done automatically and / or performed by a computer.

[0079] Then, preferably, one of the predicted spectra 8 is selected. This is preferably done based on a comparison of the predicted spectrum 8 with the measured spectrum 3. In particular, the predicted spectrum 8 that best matches the measured spectrum 3 is selected.

[0080] The selection of a predicted spectrum 8 is, inter alia, a determination of which of the predicted spectra 8 best matches the measured spectrum 3. The predicted spectra 8 that are not selected are preferably discarded / rejected.

[0081] Finally, the structure 1 corresponding to the selected predicted spectrum 8, i.e. in particular the structure 1 from which the selected predicted spectrum 8 was generated, is preferably determined as the structure 1 of the unknown compound 2. This is preferably performed automatically and / or computationally.

[0082] In particular, all steps of the method are performed automatically and / or by a computer and / or the method is a fully automated and / or computer-implemented method.

[0083] It is a preferred embodiment of the present invention that a first machine learning model 7 is used to generate structures 1 of candidate compounds and a second machine learning model 9 is used to generate predicted spectra 8 from the structures 1 generated by the first machine learning model 7. In other words, the structure elucidation method preferably utilizes two different machine learning models 7, 9 having different learning algorithms and / or being differently trained, i.e. trained with different training data sets and / or trained for different tasks. In particular, the task of generating structures 1 or candidate compounds is separated from the task of generating predicted spectra 8. This allows for effective and / or efficient training and can make both the generation of potential structures 1 and the generation of predicted spectra 8 of (candidate) compounds more efficient, faster or more reliable.

[0084] The general idea of ​​the above embodiment, i.e., performing structure elucidation by first generating structures 1 of candidate compounds and then generating predicted spectra 8 of these structures 1, avoids a fundamental problem in structure elucidation, i.e., directly inferring molecular structure from measured spectra 3, and thus allows structure elucidation to be performed more quickly and efficiently.

[0085] First machine learning model According to a preferred embodiment, which can also be carried out independently, the generation of the candidate compound structure 1 is carried out by a machine learning model, in particular a first machine learning model 7.

[0086] The first machine learning model 7 is preferably trained to generate realistic structures 1. Realistic structures in this sense are in particular physically and / or chemically possible structures. The structures 1 are preferably generated from molecular and / or empirical formulas 6.

[0087] In particular, the first machine learning model 7 is trained before using the first machine learning model 7 for the actual structure elucidation. In other words, the method according to the invention preferably comprises a training or training phase of the first machine learning model 7 and a phase of applying the first machine learning model 7.

[0088] During the training or training phase, the first machine learning model 7 preferably learns how to generate realistic structures 1, in particular starting from molecular formulas and / or empirical formulas 6. After training or once the training phase is finished, the first machine learning model 7 is used or can be used to generate structures 1 of candidate compounds, in particular to elucidate structures 1 of unknown compounds 2 from measured spectra 3 of samples 4.

[0089] The application or adaptation phase is preferably the phase after the training phase is completed, in other words the phase where the first machine learning model 7 is used to generate structures 1 of candidate compounds and / or for the actual structure elucidation.

[0090] The first machine learning model 7 is preferably trained using a training dataset 10, hereinafter referred to as the first training dataset 10. Preferably, the first machine learning model 7 comprises the first training dataset 10 and / or the first training dataset 10 forms part of or a component of the first machine learning model 7.

[0091] The first machine learning model 7 is trained in particular to generate realistic structures 1. In other words, the aim of training the first machine learning model 7 is to achieve or ensure that the structures 1 generated by the first machine learning model 7 are physically and / or chemically possible. This is achieved in particular by selecting an appropriate first training dataset 10.

[0092] The first training data set 10 preferably comprises a plurality of real molecule or chemical compound structures 1. The structures 1 preferably exist as and / or are represented by digital data. The real molecule or chemical compound structures 1 may exist or be provided, for example, as data defining or including the relative geometric positions of the atoms or nuclei of the molecule or chemical compound and / or as a formula representing the structure 1, such as a structural or skeletal formula.

[0093] The first training data set 10 is preferably a database or is obtained from a database. A preferred example of the first training data set 10 and / or database is the ZINC database available under https: / / zinc.docking.org. The ZINC database is described in detail in J. Chem. Inf. Model. 2012, 52, 7, 1757-1768 (published May 15, 2012).

[0094] The first training data set 10 preferably comprises structures 1 of real molecules or compounds, in particular molecular formulas and / or empirical formulas 6. Particularly preferably, for every molecule or compound in the first training data set 10, the first training data set 10 comprises a structure 1 and a molecular formula and / or empirical formula 6 of the respective molecule or compound.

[0095] The first machine learning model 7 is trained in particular to generate structures 1 from molecular formulas and / or empirical formulas 6. In other words, the molecular formulas and / or empirical formulas 6 are preferably used as input for or constitute the input of the first machine learning model 7. The machine learning model 7 then generates one or more structures 1 from the input molecular formulas and / or empirical formulas 6. The generated structures 1 preferably constitute the output of the first machine learning model 7. This is also depicted diagrammatically in FIG.

[0096] The structure 1 generated by the first machine learning model 7 may exist or be provided, for example, as data defining or including the relative geometric positions of atoms or nuclei of a molecule or compound, and / or as a formula representing the structure 1, such as a structural or skeletal formula.

[0097] Preferably, the first machine learning model 7 comprises and / or uses one or more artificial neural networks, preferably one or more graph neural networks. Particularly preferably, the artificial or graph neural network is a generative adversarial network. The use of such a network has proven to be particularly advantageous for the generation of the structure 1.

[0098] Generative adversarial networks are described in particular in the paper "Generative Adversarial Nets" by Ian J. Goodfellow et al. (2014), https: / / arxiv.org / abs / 1406.2661, also published in Advances of Neural Information Processing Systems 27 (NIPS 2014).

[0099] The first machine learning model 7 , in particular an artificial neural network, preferably comprises a generator 11 and preferably a discriminator 12 .

[0100] Preferably, the generator 11 is a machine learning model and / or comprises an artificial neural network, in particular a graph neural network. Preferably, the generator 11 comprises and / or uses a first training data set 10.

[0101] Preferably, the discriminator 12 is a machine learning model and / or comprises an artificial neural network, in particular a graph neural network. The discriminator 12 preferably comprises and / or uses a first training data set 10.

[0102] The first training data set 10 is preferably used for training both the generator 11 and the discriminator 12. In other words, during training, the first training data set 10 is preferably used as input for the generator 11 and as input for the discriminator 12.

[0103] Preferably, the generator 11 and the discriminator 12 are separate machine learning models. Preferably, the generator 11 and the discriminator 12 have different or separate machine learning algorithms and / or have or use the same training data set 10.

[0104] The generator 11 and the discriminator 12 preferably form a pair of generative adversarial networks. Thus, the generator 11 and the discriminator 12 preferably compete with each other in the form of a game, in particular a zero-sum game, where the gain of one agent is the loss of the other agent.

[0105] Particularly preferably, the generator 11 is formed by a generative model G and / or the discriminator 12 is formed by a discriminative model D as described in the above-cited paper "Generative Adversarial Nets" by Goodfellow et al.

[0106] The first machine learning model 7 is preferably trained using a generator 11 and a discriminator 12. The generator 11 and the discriminator 12 are preferably trained mutually.

[0107] The generator 11 preferably generates the structure 1, in particular starting from the molecular formula and / or empirical formula 6.

[0108] During and / or during training, the structures 1 generated by the generator 11 are preferably presented or fed to a discriminator 12. The task of the discriminator 12 during training is to distinguish between the structures 1 generated by the generator 11 and structures 1 of real molecules, which are preferably taken from a first training data set 10.

[0109] In particular, the discriminator 12 is trained to learn the differences between the structures 1 generated by the generator 11 and the actual structures 1. This is done in particular by comparing the structures 1 generated by the generator 11 with the actual structures 1, in particular taken from the first training data set 10, and giving feedback to the discriminator 12 if its decisions are corrected.

[0110] The real structures 1 are in particular structures 1 of molecules that exist in reality, for example structures 1 of molecules that have been previously synthesized or isolated. The real structures 1 are in particular structures 1 that are contained in the first training data set 10.

[0111] Preferably, the decisions of the discriminator 12 are in turn presented to the generator 11, so that among other things the generator 11 receives feedback on the generated structures 1 and learns how the real structures 1 "look" like. In this way the generator 11 is preferably learned or trained to generate realistic and / or physically and / or chemically possible structures 1.

[0112] The goal of the (mutual) training of the generator 11 and the discriminator 12 is for the generator 11 to become so good at generating structures 1 that all structures 1 generated by the generator 11 are realistic and / or physically and / or chemically possible. In other words, the generator 11 preferably learns to “fool” the discriminator 12.

[0113] Once training is complete, the structures 1 generated by the generator 11 are preferably indistinguishable from real structures 1 or from structures 1 from the first training data set 10 .

[0114] The discriminator 12 is preferably used only in the training phase and / or not after the training phase. The discriminator 12 is preferably not used in the application phase and / or in the actual structure elucidation and / or is used (only) before the application phase.

[0115] The first machine learning model 7 preferably generates structure(s) 1 from the molecular formula and / or empirical formula 6 in one algorithmic iteration, in particular in only one or exactly one algorithmic iteration.

[0116] The first machine learning model 7 and / or the generator 11 preferably comprises and / or uses a fully connected layer.

[0117] The first machine learning model 7, in particular the generator 11, preferably uses or has a fully connected input graph. The structure 1 generated or to be generated by the first machine learning model 7 is preferably represented by a graph, in particular a fully connected graph.

[0118] Preferably, the first machine learning model 7, in particular the generator 11, is generated and / or trained to generate multiple and / or different structures 1 from a given or identical molecular formula and / or empirical formula 6. This is done, in particular, by using random variables and / or a random noise generator. In other words, the first machine learning model 7, in particular the generator 11, is preferably generated and / or trained to generate structures 1 using a random noise generator and / or random variables. This can ensure the diversity of the generated structures 1.

[0119] Preferably, the atoms or nuclei of structure 1 are represented by nodes of the graph, and the bonds or chemical bonds between the atoms or nuclei are represented by edges of the graph.

[0120] Preferably, a feature vector is assigned to each node of the graph or each kernel of the generated structure 1.

[0121] The feature vector preferably comprises data identifying the atoms or nuclei represented by each node. Preferably, this data is the element symbol (e.g., O for oxygen, C for carbon, N for nitrogen, etc.) and / or atomic number (e.g., 8 for oxygen, 6 for carbon, 7 for nitrogen, etc.) of the atoms or nuclei represented by each node. The atomic number is in particular the number of protons contained in the nucleus or atom. In other words, an element of the feature vector preferably includes the element symbol and / or atomic number represented by each node. However, other suitable data for identifying the atoms or nuclei represented by each node may also be used.

[0122] Preferably, the feature vector comprises, in particular, data identifying the atom or nucleus represented by each node, a random variable, in other words, preferably one element of the feature vector contains a random variable, which preferably ensures the randomness of the model.

[0123] The random variables in the feature vector are preferably generated by a random noise generator and / or according to a probability distribution, for example a uniform probability distribution or a Gaussian probability distribution.

[0124] It is therefore particularly preferred that each node is assigned a feature vector, each node representing one atom or nucleus of the structure 1, the feature vector having at least or exactly two elements, one element containing data identifying the atom or nucleus represented by the respective node, in particular its atomic number and / or element symbol, and one element containing a random variable.

[0125] However, a feature vector may also have more than two elements. In particular, the feature vector may comprise, in addition to elements including data identifying the atoms or nuclei represented by each node and random variables, further features, such as information about the nodes and / or molecules, in particular chemical information.

[0126] For example, further features, particularly chemical information, regarding the nodes and / or molecules may be information about the molecular sum formula and / or possible neighbors of the atom or nucleus represented by the respective node, information about the chemical structure, such as the ring and / or specific functional groups that the atom or nucleus represented by the respective node forms part of, information about preferred bond types and / or preferred binding partners, information about the number of valence electrons, etc.

[0127] The further information is preferably represented by and / or contained in one or more elements of the feature vector.

[0128] The following table shows an example of a graphical representation, specifically a feature vector, for a dicyanoketene molecule with the molecular formula C4N2O. [Table 1]

[0129] In particular, the use of a random noise generator and / or random variables in the feature vector ensures randomness of the model and enables the first machine learning model 7 and / or generator 11 to generate multiple and / or different structures 1 from one molecular and / or empirical formula 6. Providing further features, in particular chemical, nodal and / or molecular information, can aid in the generation of realistic structures 1 and in particular can make the training and / or generation of structures 1 faster and / or more efficient and / or more reliable.

[0130] In particular, from a given molecular formula and / or empirical formula 6, multiple and / or different structures 1 are generated by the first machine learning model 7, in particular the generator 11, by applying the machine learning model 7 or generator 11 multiple times to the same molecular formula and / or empirical formula 6, in each instance using different random variables, in particular different random variables in the feature vector. In particular, the use of different random variables in the feature vector preferably leads to different generated structures 1, even if the molecular formula and / or empirical formula 6 is the same in each instance.

[0131] As an alternative to the use of random noise generators and / or random variables, randomness of the model can also be ensured or achieved in other ways.

[0132] For example, instead of using random variables in the feature vector, it is possible to use a feature vector with only one element that contains data identifying the atom or nucleus represented by each node, and randomize the edges or connections between the nodes. Thus, in such an approach, the input graph is not fully connected.

[0133] Second Machine Learning Model According to another preferred embodiment, generating the predicted spectrum 8 from the structure 1 is performed by a machine learning model, in particular a second machine learning model 9, although this can also be performed independently.

[0134] The second machine learning model 9 preferably comprises and / or uses an artificial neural network, preferably a graph neural network, in particular a graph attention network and / or a residual neural network, also known as ResNet.

[0135] A graph attention network is, among other things, an artificial neural network with a graph attention layer. Graph attention networks are described, among other things, in the paper "Graph attention networks" by Velicovic et al. (2018), available at https: / / arxiv.org / abs / 1710.10903.

[0136] The graph attention layer, in particular,

number

number

number

number

number

number

number

number

[0137] Residual neural networks, or ResNets, are notably described in the paper "Deep residual learning for image recognition" by He et al., 2015, available at https: / / arxiv.org / abs / 1512.03385, and also in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770-778.

[0138] The residual neural network is preferably an artificial neural network having one or more residual blocks. The residual block is a block with two layers, where x is the input to the residual block and / or the first layer of the residual block, y is the output of the residual block, and F(x) is the output of the second layer of the residual block, and the output of the residual block is y=F(x)+x. Thus, the residual block has an input x and an output y=F(x)+x. In this way, the residual block is realized by adding the input x to the output F(x). Further mathematical details for implementing this concept are described in the above-cited paper "Deep residual learning for image recognition" (2015) by He et al.

[0139] The second machine learning model 9 preferably comprises a spectrum generator 13. The spectrum generator 13 preferably comprises or is formed by an artificial neural network, in particular a graph neural network and / or a residual neural network.

[0140] In particular, the second machine learning model 9 and / or the spectrum generator 13 comprises and / or uses a combination of a graph attention network and a residual neural network. Particularly preferably, the second machine learning model 9, in particular the spectrum generator 13, comprises and / or uses a graph attention network with a head of a residual neural network.

[0141] The second machine learning model 9 and / or the spectrum generator 13 preferably comprises and / or is trained with a training dataset 14 (hereinafter specifically referred to as the second training dataset 14).

[0142] In particular, the second machine learning model 9 is trained before using the second machine learning model 9 for the actual structure elucidation. In other words, the method according to the invention preferably comprises a training or training phase of the second machine learning model 9 and an application or application phase of the second machine learning model 9.

[0143] During the training phase, the second machine learning model 9 preferably learns how to generate a predicted spectrum 8, in particular starting from a given structure 1. After training, or once the training phase is finished, the second machine learning model 9 is used or can be used to generate a predicted spectrum 8 from the structure 1, in particular to elucidate the structure 1 of the unknown compound 2 from the measured spectrum 3 of the sample 4.

[0144] The application or application stage is preferably the stage after the training stage is completed, in other words the stage of using the second machine learning model 9 to generate predicted spectra 8 of candidate compounds.

[0145] The second machine learning model 9 preferably comprises and / or is formed by a spectrum generator 13 and a second training data set 14. The spectrum generator 13 preferably is or comprises an algorithm, in particular a machine learning algorithm.

[0146] The spectrum generating unit 13 preferably generates the spectrum 8 or the spectral features 5, in particular the chemical shifts, in one shot.

[0147] The second training data set 14 preferably comprises structures 1 labelled with associated spectral features 5 .

[0148] The spectral feature 5 is preferably a chemical shift, especially when the measured spectrum 3 and / or the predicted spectrum 8 are NMR spectra. Particularly preferably, the spectral feature 5 is 1 H and / or 13 C chemical shifts. Thus, for all of the given structure 1 contained in the second training data set 14, 1 H and / or 13 Preferably, the C atoms or nuclei are labelled with the relevant spectral features 5 or chemical shifts.

[0149] The second machine learning model 9, in particular the spectrum generator 13, is preferably trained to generate and / or generate a predicted spectrum 8 from the structure 1. In particular, the spectrum 8 is generated from the structure 1 generated by the first machine learning model 7 and / or the structure 1 generated by the first machine learning model 7 is used as input for the second machine learning model 9, in particular the spectrum generator 13.

[0150] In particular, for every structure 1 one, in particular exactly or only one, predicted spectrum 8 is calculated or generated by the second machine learning model 9 and / or the spectrum generator 13 .

[0151] The predicted spectrum 8 preferably constitutes the output of a second machine learning model 9, also depicted diagrammatically in FIG.

[0152] In the second machine learning model 9 and / or in the spectrum generator 13, the structure 1 is preferably equipped with and / or labelled with chemical information, in particular chemical information relating to each atom or nucleus of the given structure 1. The chemical information is in particular information relating to the chemical state of the atom or nucleus, i.e. in particular information relating to the atom or nucleus itself and to the bonds surrounding and / or with other atoms or nuclei of the structure 1. This chemical information makes it possible to predict or generate a predicted spectrum 8 and / or spectral features 5 of the structure 1.

[0153] The chemical information relating to the atoms or nuclei preferably comprises one or more of the following characteristics: (i) the atomic number of the atom or nucleus and / or other appropriate data to identify the atom or nucleus; (ii) valency or number of binding partners; (iii) aromaticity, in particular when an atom or nucleus is part of an aromatic structure; (iv) s, sp, sp 2 , sp 3 , sp 3 d or sp 3 d 2 etc. hybrid state; (v) formal charge; (vi) a predefined valence or number of valence electrons; (vii) Information regarding rings, particularly when atoms or nuclei are members of a ring, and / or information regarding the size of the ring, e.g. the number of atoms or nuclei forming the ring.

[0154] The structure 1 is preferably represented in a graph in the second machine learning model 9 and / or in the spectrum generator 13. Preferably, atoms or nuclei of the structure 1 are represented by nodes of the graph.

[0155] Preferably, each node or atom or nucleus is assigned a feature vector, which preferably comprises chemical information of the atom or nucleus represented by the node.

[0156] The feature vector is preferably a vector having one or more features, in particular, the feature vector has one or more, preferably all, of the above features (i) to (vii).

[0157] The features are preferably one-hot encoded, which facilitates computational efficiency and / or fast calculation or generation of the predicted spectrum 8.

[0158] The second machine learning model 9 and / or spectrum generator 13 preferably uses a graph attention network to generate the predicted spectrum 8. Preferably, for each given structure 1, the second machine learning model 9 and / or spectrum generator 13 generates a predicted spectrum 8 for the structure 1 and / or a node embedding comprising the information necessary to compute the predicted spectrum 8 for the structure 1.

[0159] Thus, in other words, the predicted spectrum 8 is preferably represented by and / or encoded in the form of a node embedding. The node embedding is in particular an abstract representation of the predicted spectrum 8 and / or is not human readable. In other words, the node embedding preferably includes all information regarding the predicted spectrum 8 and / or all information necessary to display or calculate the predicted spectrum 8, but the information is not included in the node embedding in a form that can be directly understood, read or interpreted by a human. In particular, in the node embedding, the predicted spectrum 8 is not included or represented in the same way as the measured spectrum 3.

[0160] Preferably, a further machine learning algorithm, in particular a residual neural network, is used to generate or calculate the predicted spectrum 8 from the node embeddings. In particular, the further machine learning algorithm or the residual neural network outputs the predicted spectrum 8 in the same representation or data type as the measured spectrum 3.

[0161] For example, if the measured spectrum 3 is present in the form of a diagram or a data table, the generated predicted spectrum 8, in particular the predicted spectrum 8 generated or output by a further machine learning algorithm or a residual neural network, is also present in the form of a diagram or a data table, respectively.

[0162] It is therefore particularly preferred that the second machine learning model 9 and / or the spectrum generator 13 comprises and / or uses a graph attention network and a residual neural network for generating the predicted spectrum 8 from the structure 1, where starting from the structure 1, node embeddings are generated by the graph attention network, the node embeddings being preferably a representation or encoding of the predicted spectrum 8, and the predicted spectrum 8 is then generated from the node embeddings by the residual neural network.

[0163] The second machine learning model 9 or the spectrum generator 13 is preferably capable of predicting the (different) spectral features 5, in particular the chemical shifts, of the diastereotopic protons.

[0164] In particular, for each spectral feature 5, in particular a chemical shift of the predicted spectrum 8, an expectation or mean value and a corresponding measure of dispersion, in particular a standard deviation, is calculated, in particular by the second machine learning model 9 and / or the spectrum generator 13. The expectation or mean value is preferably the position of the spectral feature 5 in the predicted spectrum 8. For example, the expectation or mean value is the position of the chemical shift, in other words the "ppm value" of the chemical shift.

[0165] In particular, this procedure of calculating both the expected mean value and the corresponding measure of variance was found to lead to the ability of the second machine learning model 9 to predict the (distinct) spectral features 5, in particular the chemical shifts, of diastereotopic protons. That is, the prediction of the mean value or position of the spectral features 5 or chemical shifts was found to be very accurate, and typically the corresponding measure of variance or standard deviation was found to be zero. However, in the case of diastereotopic protons, the calculated standard deviation is greater than zero.

[0166] The spectral feature 5 or chemical shift of the diastereotopic protons is preferably calculated by adding and subtracting the corresponding measures of dispersion or standard deviations from the expected or average values. Preferably, the spectral feature 5 or chemical shift of one of the diastereotopic protons is the sum of the expected or average value and the corresponding measures of dispersion, and the spectral feature 5 or chemical shift of the other of the diastereotopic protons is the difference between the expected or average value and the corresponding measures of dispersion.

[0167] For example, if the expectation or mean value of two diastereotopic protons is 4.42 ppm and the corresponding measure of dispersion or standard deviation is 0.18 ppm, then the spectral feature 5 or chemical shift of one of the diastereotopic protons is calculated to be 4.60 ppm (=4.42 ppm+0.18 ppm) and the calculated position of the second diastereotopic proton is 4.24 ppm (=4.42 ppm-0.18 ppm).

[0168] The individual aspects of the present invention can be implemented independently of one another, but can also be implemented in any desired combination and / or order. [Explanation of symbols]

[0169] 1. Structure 2 Unknown compound 3. Measured spectrum 4. Sample 5. Spectral characteristics 6 Molecular formula and / or empirical formula 7 (First) Machine Learning Model 8. Predicted Spectrum 9 (Second) Machine Learning Model 10 (First) Training Data Set 11. Generator 12 Discriminator 13 Spectral Generator 14 (second) training data set

Claims

1. A method for elucidating the structure (1) of an unknown compound (2) from a measured spectrum (3) of a sample (4), comprising: A structure (1) of a candidate compound is generated, and a predicted spectrum (8) is generated from the generated structure (1); The predicted spectrum (8) is compared with the measured spectrum (3); One of the predicted spectra (8) is selected, and the structure (1) corresponding to the selected predicted spectrum (8) is determined as the structure (1) of the unknown compound (2); a) a first machine learning model (7) generates a structure (1) of the candidate compound, and a second machine learning model (9) generates the predicted spectrum (8) from the structure (1) generated by the first machine learning model (7); and / or b) a first machine learning model (7) generates a structure (1) of the candidate compound, the first machine learning model (7) being trained to generate a realistic structure (1) from a molecular formula and / or an empirical formula (6); and / or c) a second machine learning model (9) generates the predicted spectrum (8) from the generated structure (1), the second machine learning model (9) having a residual neural network; method.

2. 2. The method of claim 1, wherein the first machine learning model (7) comprises one or more artificial neural networks, preferably graph neural networks, in particular a pair of generative adversarial networks.

3. 3. The method according to claim 1 or 2, wherein the first machine learning model (7) is trained using a first training data set (10), preferably for the generation of the realistic structures (1), in particular from molecular formulas and / or empirical formulas (6), the first training data set (10) preferably comprising structures (1) of a plurality of real molecules.

4. 3. The method according to claim 1 or 2, wherein the first machine learning model (7) is trained using a generator (11) and a discriminator (12), preferably the generator (11) and the discriminator (12) are trained mutually and / or the discriminator (12) is used only in training and / or is not used in the actual structure elucidation and / or application phase.

5. 3. The method according to claim 1 or 2, wherein the first machine learning model (7), in particular the generator (11), is generated and / or trained to generate a plurality of structures (1) from a given molecular formula and / or empirical formula (6), in particular using a random variable and / or a random noise generator.

6. 5. The method of claim 4, wherein the discriminator (12) is trained to distinguish between actual structures (1) from the first training data set (10) and structures (1) generated by the generator (11).

7. 3. The method according to claim 1 or 2, wherein the second machine learning model (9) comprises one or more artificial neural networks, preferably a graph neural network, a graph attention network and / or a residual neural network.

8. The second machine learning model (9) is trained using a second training data set (14), which preferably includes relevant spectral features (5), preferably chemical shifts, in particular 1 H and / or 13 3. The method of claim 1 or 2, comprising the structure (1) labeled with C chemical shifts.

9. 3. The method according to claim 1 or 2, wherein for each spectral feature (5) of the predicted spectrum (8), in particular a chemical shift, an expectation or mean value and a corresponding measure of dispersion, in particular a standard deviation, are calculated.

10. 3. The method according to claim 1 or 2, wherein the measured spectrum (3) is an NMR spectrum and / or a spectrum, in particular an NMR spectrum, of the sample (4) is measured.

11. 3. The method according to claim 1 or 2, wherein the molecular formula and / or the empirical formula (6) is determined by measuring the mass spectrum of the sample (4).

12. The method of claim 1 or 2, wherein the method is a computer-implemented method.

13. A data processing device comprising means for carrying out the method according to claim 1 or 2.

14. A computer program product comprising instructions that, when executed by a computer, cause the computer to carry out the method according to claim 1 or 2.

15. A computer-readable storage medium comprising instructions that, when executed by a computer, cause the computer to perform the method of claim 1 or 2.