Method for classifying physical, chemical and / or physiological properties of molecules
Patent Information
- Application Number
- EP2023742060
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-13
- Filing Date
- 2023-07-13
- Publication Date
- 2025-05-21
AI Technical Summary
Current methods for selecting molecules with desired physical, chemical, or physiological properties are inefficient and labor-intensive, particularly in predicting molecular smells, due to unknown structure-odor relationships and the need for extensive experimental testing.
A method using a mathematical model to classify molecules based on structural patterns and their probabilities of belonging to specific classes, allowing for the preselection of molecules with desired properties, followed by experimental confirmation.
This approach significantly reduces the need for extensive experimental testing, saving time and resources by enabling the identification of molecules with specific properties and providing insights into structural pattern-class relationships, thereby improving the efficiency of molecular design and selection.
Smart Images

Figure IMGF000006_0001 
Figure IMGF000006_0002 
Figure IMGF000008_0001
Abstract
Description
[0001] Methods for classifying physical, chemical and / or physiological properties of molecules
[0002] The invention relates to a method for selecting molecules with a desired physical, chemical, and / or physiological property from a group of molecules, wherein a classification according to a chemical, physical, and / or physiological property of a molecule is carried out using a mathematical model. This allows molecules with the desired property to be selected from the group of molecules. For this selection of molecules, experimental confirmation is then carried out to determine whether the molecules actually exhibit the desired physical, chemical, and / or physiological property.Furthermore, the use of the method according to the invention for selecting at least one molecule with a desired chemical, physical and / or physiological property from a group of molecules and for identifying the influence of structural patterns in molecules on at least one chemical, physical and / or physiological property of molecules is described.
[0003] Molecules exhibit chemical, physical, and physiological properties. While physical properties can be quantified by measuring underlying physical quantities, chemical properties can be quantified by measuring an underlying chemical quantity during the reaction of a molecule with another substance. The physical properties of a molecule include, for example, the molecule's color. Water solubility, on the other hand, is considered a chemical property of a molecule. Molecules also exhibit physiological properties. These include physical and chemical properties of substances from the perspective of their perceptibility or their effect on the environment. Examples of this are the smell and taste of a molecule.
[0004] Chemical, physical, and physiological properties are of great interest for a wide range of applications. Physiological properties describe properties of molecules that affect the organism of living beings. According to the invention, this includes properties such as the taste or smell of molecules. Furthermore, according to the invention, this also includes the classification of molecules into permitted and prohibited chemicals in cosmetics and personal care. This is regulated by the application authorization according to Articles Regulation, Annex II - Restricted Substances, Annex II of the European Chemicals Agency (ECHA). The taste of molecules directly appeals to the human sense of taste and thus decisively influences eating behavior and, in particular, which foods are perceived as pleasant or unpleasant.The taste evoked by molecules is therefore of great importance, especially in the food industry.
[0005] Smell is one of the five human senses and plays an important role in daily life. For example, the smell of food influences our eating behavior [1], and smells in threatening situations influence human memory of such situations [2]. In addition to their importance to humans, smells also play an important role in business, particularly in the food and cosmetics industries, where the development of new flavors and the identification of odor-active molecules are essential. When developing new odorants, a predictive approach during molecular design is required to reduce the space of candidate molecules from virtually all to a promising set of structures.
[0006] Although much progress has been made in odor prediction in recent years [3, 4, 5, 6], unfortunately, little is still known about the relationship between a molecule's structure and its odor, so chemists cannot be provided with a "toolbox" to design molecular structures with a specific odor in mind [7, 8]. Furthermore, there is disagreement about the dimensionality of odor space [9, 10]. To derive the rather vague property of odor from objectively measurable or calculable molecular properties, a relationship between physicochemical parameters and odor can be used. Using this approach and principal component analysis (PCA), Khan et al. predicted the pleasantness of the odor of molecules and identified it as one of the dimensions of human odor perception
[0011] , in agreement with other studies
[0012] .
[0007] To predict a specific odor, Keller et al. investigated the performance of 22 different machine learning models in predicting 19 odor descriptors. They used physiochemical properties such as atom type, functional groups, or topological and geometric information. The models successfully predicted eight of the 19 descriptors considered. The authors looked for correlations between features and descriptors and found significant correlations between sulfur-containing molecules and the descriptors "garlic" and "burnt." Based on the good performance of the linear models, the authors concluded that there is a linear, summative effect of the features on odor perception
[0013] ,
[0014] .
[0008] Shang et al. investigated various combinations of feature generation models and machine learning algorithms to predict the odor of molecules from ten possible descriptors. They applied the models in GC / O (gas chromatography analysis with olfactometric detection). With an accuracy of 97.08%, the Support Vector Machine (SVM) achieved the best results in prior feature selection using Boruta
[0015] . However, when predicting aroma molecules that were not included in the model building, the accuracy dropped to 70% [6]. The models used features calculated with the chemoinformatics software Dragon for odor prediction. These features are also used by Snitz et al.used to predict the odor of odor mixtures [5]. Training a deep autoencoder
[0016] also enabled the extraction of features that can be used as an alternative to using features generated by Dragon. Tran et al. developed the autoencoder DeepNose to extract molecular features. DeepNose features achieved equally good results in predicting odor perceptions compared to Dragon features [3].
[0009] While the models used are promising and useful in their own right, they employ a multitude of different features that do not provide deep insight into the mechanism of prediction. Due to their opaque nature, the state-of-the-art models function more like a "black box," thus still lacking knowledge about structure-odor relationships.
[0010] This means that in both industry and science, sensory-trained experts must smell molecules to determine their odor. Due to largely unknown structure-odor relationships, the trial-and-error approach prevails in the development of flavorings or the identification of odor-active molecules. This is very time-consuming, labor-intensive, and therefore uneconomical.
[0011] It is equally desirable to be able to derive other physical, chemical or physiological properties of a molecule from its structure.
[0012] Based on the prior art, it is therefore the object of the invention to provide a method by which molecules with a desired physical, chemical or physiological property can be selected from a given set of molecules without having to examine all molecules with regard to the desired property using experimental methods.
[0013] For this purpose, the invention provides a method for selecting molecules with a desired physical, chemical and / or physiological property from a group of molecules, comprising the steps
[0014] • Providing a group of Ok molecules by a user, where ke N;
[0015] • Providing a classification according to a chemical, physical and / or physiological property of a molecule, comprising C, classes, where N;
[0016] • Providing a mathematical model for classification, wherein the mathematical model describes relationships G,j between a structural pattern and a class, in particular by probabilities that a structural pattern Fj of a molecule belongs to a class C, or that a molecule of a class C, has a structural pattern Fj;
[0017] • Selection of a weighting function a,j for the mathematical model by a user;
[0018] • Assignment of all Ok molecules into the C, classes of classification by the mathematical model, wherein the mathematical model comprises the steps: a) Determination and storage of Fj structural patterns of the chemical structure of each of the Ok molecules assigned to the respective molecule, where N each; b) Assignment of the probability Gy to each structural pattern Fj of a molecule for each class C, and calculation of the influence ,j according to the formula for each structural pattern Fj of a molecule for each class C,; c) calculation of a score P iik for each molecule Ok, using for each class C,, whereby the influences / y of all structural patterns Fj contained in a molecule Ok are summed up for each class C, d) assignment of each molecule to the class C with the highest score Bk for the respective molecule;
[0019] • Display and / or output of the molecules assigned to the classes of the classification and optionally the associated point values P iik , the associated influences / y and the structural patterns p;
[0020] • Selection of molecules assigned to the class with the desired physical, chemical and / or physiological property;
[0021] • Experimental confirmation of the physical, chemical and / or physiological property of at least some of the selected molecules by a user; and / or verification and / or identification of the relationship between at least one structural pattern Fj and a class C by a user.
[0022] Furthermore, the use of the method according to the invention for selecting at least one molecule with a desired chemical, physical and / or physiological property from a group of molecules and for identifying the influence of structural patterns in molecules on at least one chemical, physical and / or physiological property of molecules is described.
[0023] Detailed description
[0024] According to the present invention, a group of O kMolecules provided by a user, where ke N. In one embodiment of the present invention, between 20 and 1000 molecules are provided, preferably between 20 and 800 molecules are provided, particularly preferably between 20 and 300 molecules are provided. In this case, provided means that the structural formulas of the molecules are present and thus made available. This is possible, for example, by providing the molecules in the SMILES structural code, which encodes structural patterns as SMARTS [17, 18, 19]. In addition, however, it is possible to have each of the molecules available as a substance for experimental confirmation at a later point in time.
[0025] According to the invention, a classification according to a chemical, physical and / or physiological property of a molecule, comprising C, classes, where N is provided.
[0026] In one embodiment of the invention, the classification is selected from structure-based properties of molecules, in particular from the group comprising odor, taste, color, toxicity, water solubility, and permitted and prohibited chemicals in cosmetics and personal care. In a particularly preferred embodiment, the classification is based on the odor of the molecules.
[0027] A classification comprises several classes; for example, the water solubility classification comprises the classes hydrophilic and hydrophobic. The toxicity classification comprises the classes toxic and non-toxic. The color classification can include different colors as classes, for example blue, red, yellow, green. The taste classification accordingly includes different tastes, such as bitter, sour, sweet, salty, and umami. The odor classification preferably includes odors such as 'woody, resinous', 'floral', 'fruity, non-citrus', 'medicinal', 'perfumed', 'light', 'heavy', 'sweet', 'aromatic', 'fragrant', 'obnoxious' as classes. The odor classification particularly preferably includes the odors 'woody, resinous', 'floral', 'fruity, non-citrus', 'medicinal', 'perfumed' as classes.
[0028] Furthermore, a mathematical model for the provided classification is provided by a user. According to the invention, the mathematical model has probabilities Gy that a structural pattern Fj of a molecule belongs to a class C, or that a molecule of a class C, has a structural pattern Fj. The mathematical model was previously trained using a training data set for the selected classification. A training data set comprises 0 / molecules for which it is known which class C, classification they are assigned to, where 1, 1 G N. In a further embodiment, a molecule can be assigned to several classes C. The creation of the mathematical model is explained later in the description.
[0029] According to the invention, a weighting function a for the mathematical model is selected by a user. A suitable weighting function a,j is selected from the group of statistical measures, such as tf-idf functions, normalization functions, equally weighted functions. The tf and idf values are calculated using the training data set and the formulas generally known to those skilled in the art
[0026] .
[0030] Subsequently, all Ok molecules are assigned to the C classes of the classification by the mathematical model. The mathematical model has the following steps: a) Determining and storing Fj structural patterns of the chemical structure of each of the Ok molecules assigned to the respective molecule, where N each; b) Assigning the probability Gij to each structural pattern Fj of a molecule for each class C, and calculating the influence / y according to the formula for each structural pattern F, of a molecule for each class C,; c) Calculating a score Piik for each molecule Ok , using for each class C,, where the influences / y of all structural patterns Fj contained in a molecule Ok are summed up for each class C, d) assignment of each molecule to the class C with the highest score P iik for the respective molecule;
[0031] In step a), structural patterns Fj of the chemical structure of each of the Ok molecules are determined. These are stored and assigned to the respective molecule. In the process, all structural patterns of the Ok molecules are determined. Structural patterns that do not occur in the training data set are assigned an influence / y of zero and are thus not considered in the method. According to the inventive method, each structural pattern Fj of a molecule is assigned a probability Gij for each class C. The respective probability Gij is known from the mathematical model for each structural pattern Fj for a given classification. The influence / ,j is calculated according to the formula h,j — a i,j ' j,j (1) for each class C, calculated. a ifj corresponds to the previously selected weighting function. The weighting function allows for additional information about the relationship between structural pattern and class to be taken into account, such as selectivity and specificity using the tf-idf function. A score P is then calculated. iik for each molecule O k according to the formula
[0032] Pi,k = ^iFjEO k h,j (?) for each class C. According to the formula, the influences ,j of all structural patterns Fj present in a molecule O k are added up for each class C. According to the invention, a point value P iik for a molecule. The molecule is then assigned to class C, the classification that has the highest score P iikfor the respective molecule. According to the invention, the molecule is therefore assigned to at least one class C. In one embodiment of the invention, the molecule is assigned to several classes. This occurs when the highest point value P iik is the same for several classes. The assignment is then made to classes C, for which the highest point value P iik was determined.
[0033] In a further embodiment of the present invention it is provided that if the point values of a molecule are the same for all classes C, this molecule is marked as unpredictable. This can occur, for example, if a molecule consists entirely of structural patterns that do not occur in the training dataset and are therefore each assigned an influence of zero.
[0034] The mathematical model therefore makes it possible to assign the molecules to classes C, the classification. The mathematical model is based on the assumption that each structural pattern has a specific influence on a class and that a structural pattern-class relationship exists. The present invention thus enables a sorting of the O k Molecules are classified into the classes of the classification. By applying the mathematical model, a pre-selection of molecules is made that are contained in the provided group of molecules and exhibit the desired physical, chemical, or physiological property.
[0035] This allows a user to subject a smaller selection of OK molecules to further experimental investigations in order to identify molecules with the desired physical, chemical, or physiological properties. Advantageously, it is no longer necessary, as was previously the case, to subject all OK molecules to experimental investigations; instead, the molecules with the highest scores in a specific classification class, and thus with a desired physical, chemical, and / or physiological property, can be subjected to experimental confirmation. This experimental confirmation determines whether a molecule actually exhibits the physical, chemical, and / or physiological property it should have according to the classification.
[0036] For example, if molecules from a group that have the odor "floral" are to be filtered out, the mathematical model for odor classification is applied, and the molecules assigned to the "floral" class are then subjected to experimental confirmation. It is advantageous to start with the molecule that has the highest score P iik in this class. Subsequently, further molecules in this class can be investigated experimentally, whereby these are advantageously arranged in a sequence according to descending point values P iik are experimentally investigated. In one embodiment, only the molecule with the highest score in a class is experimentally investigated. In a further embodiment of the present invention, all molecules are experimentally investigated whose score P iiknot more than 50%, preferably not more than 30%, particularly preferably not more than 10% of the highest point value P iik in this class. In a further embodiment of the present invention, all molecules of a class of classification are experimentally investigated.
[0037] According to the invention, the molecules Ok are therefore displayed and / or output in association with the classes of the classification. In one embodiment, the molecules are displayed and / or output in such a way that the molecules are arranged in descending order according to their point value P iik in a class C, starting with the molecule with the highest score P iik In one embodiment, the associated score P iikand / or the associated influences ,j and / or the associated structural pattern Fj are displayed and / or output. Subsequently, the molecules assigned to the class with the desired physical, chemical, and / or physiological property are selected.
[0038] As already described, this is followed by experimental confirmation of the physical, chemical, and / or physiological properties of at least some of the selected molecules by a user. This experimental verification simultaneously verifies the classification of the molecule by a user. The type of experimental confirmation depends on the classification used. The following table provides a non-exhaustive overview of common experimental methods that can be used to verify the physical, chemical, and physiological properties of molecules. All other common experimental methods known to the person skilled in the art are equally applicable.
[0039] In one embodiment of the present invention, a user further checks and / or identifies the relationship between at least one structural pattern Fj and a class C. This advantageously makes it possible to gain insight into the structural pattern-class relationship. Physical, chemical, and / or physiological properties of molecules can thus be traced back to specific structural patterns of the molecules.
[0040] The present invention thus enables significant savings in personnel and technical effort, since all molecules Ok of a given group no longer need to be experimentally investigated in order to select at least one molecule of a specific class and thus with a specific physical, chemical, or physiological property. By applying the mathematical model, a selection of molecules is made, and the subsequent experimental confirmation can be carried out specifically with this selection of molecules. This saves time and money compared to state-of-the-art methods. Furthermore, it is not necessary to have all molecules available as substances for experimental investigations, which saves additional costs.
[0041] According to the invention, a mathematical model is used which comprises the probability G,j for defined structural patterns for defined classes C, of a classification or a molecule of a class C, has a structural pattern Fj.
[0042] For this purpose, the mathematical model is trained according to the invention by means of a training data set for a selected classification, whereby a training data set comprising 0 / molecules of known class C is specified, where I, ie N. Training in this context means nothing other than that the probabilities G t j = for defined structural patterns for defined classes C, based on a given data set or that the probabilities G tj = Pr( |F7) can be calculated that a molecule of a class C has a structural pattern Fj . The structural patterns of the molecules are known for the data set, as is the class of classification into which the respective molecules belong. In one embodiment of the present invention, a molecule can also be assigned to multiple classes.
[0043] The method for training the mathematical model comprises the following steps: i. Determining and storing Fj structural patterns of the chemical structure of each molecule assigned to the respective molecule, each N ; ii. Calculating the probability Gy that a structural pattern Fj of a
[0044] Class C, where G t j = Pr(F 7 0C | er
[0045] Calculation of the probability G that a molecule of a class C has a structural pattern Fj, where G t j = Pr(c F7) =
[0046] In step i, the structural patterns Fj of each molecule are determined. A structural pattern is a partial fragment of the molecule's chemical structure. Not all structural components of the molecules need to be used; a prior feature selection can be performed using an algorithm or statistical values.
[0047] For example, the determination of the structural pattern Fj of a molecule can be implemented using so-called fingerprint algorithms. One such fingerprint algorithm is the RDKit topology fingerprint [20, 21]. Furthermore, the Dragon software
[0022] and graph convolutional neural networks
[0023] are known for determining molecular structures. A new method considers molecules as graphs and converts nodes and edges of the graphs into a vector, allowing molecules to be represented purely structure-based
[0024] .
[0048] In one embodiment of the present invention, not all structural patterns that occur in a group of molecules are used in the method according to the invention. In this case, the structural patterns determined and stored in step a) of the method according to the invention represent a selection from a larger number of structural patterns. The selection can be performed, for example, by an algorithm, an idf weighting, or a tf-idf weighting. An algorithm can, for example, select based on the minimum number of molecules that exhibit a structural pattern or based on correlations between different structural patterns.
[0049] For each structural pattern Fj, a probability G,j is then calculated that a structural pattern belongs to a class C. The probability G,j is determined using the formula calculated.
[0050] Alternatively, for each structural pattern Fj, a probability G,j is calculated that a molecule of class C, has a structural pattern Fj. The probability Gy is calculated using the formula calculated.
[0051] The present invention can be used to select molecules with a desired chemical, physical, and / or physiological property from a group of molecules. Furthermore, the present invention can be used to identify the influence of structural patterns in molecules on at least one chemical, physical, and / or physiological property of molecules.
[0052] In a particularly preferred embodiment, the method according to the invention is used to determine the odor of a molecule, or to select from a group of molecules those molecules that exhibit a specific odor. In this case, the classification is the odor, and the classes are individual odors, such as 'floral' and 'medicinal'.
[0053] Advantageously, the method according to the invention also provides insight into the structural pattern-odor relationship. Since the method calculates an influence j for each structural pattern in the form of a quantitative value for each class and thus for each odor, by comparing these influences, structural patterns can be identified that appear to have a strong impact on a specific odor. The structural patterns can therefore also be ranked according to their influence on a specific odor.
[0054] The invention is explained in more detail below with reference to 2 figures and 3 embodiments.
[0055] Figure 1 shows a sequence of the method according to the invention;
[0056] Figure 2 shows results of the method according to the invention.
[0057] Figure 1 shows a sequence of the method according to the invention, which is described in more detail in embodiment 2.
[0058] Figure 2 shows results of the method according to the invention, in which a mathematical model was implemented with different weighting functions a,j and with and without selection of the structural patterns.
[0059] Example 1 - Odor determination
[0060] A mathematical model was trained using a group of five molecules as an example for classifying odors into the two classes "floral" and "medicinal." This means that structural patterns F were determined for all molecules. For each of the five molecules, its class membership(s) was known. Using this information, the probabilities Gij for each structural pattern Fj were calculated. Figure 1 lists the five molecules for the training dataset. The molecules are represented in the structure code SMILES, and the structural patterns are coded as SMARTS. For the sake of clarity, the three structural patterns [CX4H3], [CX4], and dccccd were shown as examples. For each of the five molecules, its classification "floral" or "medicinal" was known. Structural patterns with the value 1.0 in the table occur in the respective molecule, and structural patterns with the value 0.0 do not occur in the respective molecule.
[0061] From the training data set, the probabilities Gij for each of the three structural patterns were calculated for the class 'floral' and for the class 'medicinal' using formula (3).
[0062] From a group of 10 molecules, those molecules were to be filtered out that exhibit a "floral" odor. The inventive procedure for this is explained in more detail below using one of the 10 molecules as an example. For the molecule CCOCOCC, it was determined which of the structural patterns from the training data set occurred in it. Furthermore, an equal weighting was specified as the weighting function a,j, so that all weighting factors were 1. The influences Z^ were then calculated for all structural patterns according to formula (1). The results for both classes for all three structural patterns are shown in Figure 1. The molecule CCOCOCC only exhibits the structural patterns [CX4H3], [CX4], so the influences of these structural patterns in both classes were summed according to formula (2). This resulted in a score of P, <=1.67 for the class "floral" and a score of P, / <=1.50 for the class "medicinal." The molecule was then assigned to class .All other 9 molecules were classified according to the same principle. Three molecules were assigned to the class "floral" and seven molecules to the class "medicinal." These three molecules were subsequently selected.
[0063] Of the three molecules in the "floral" class, the molecule CCOCOCC had the highest score. Due to the manageable number of molecules assigned to the "floral" class, all three molecules were subsequently investigated experimentally. Substances consisting of each of the three molecules were examined by a person trained in odor perception, and it was found that all three molecules could also be assigned to the "floral" class in experimental confirmation.
[0064] Embodiment 2 - Validation of the Mathematical Model The method according to the invention was carried out using a group of 64 molecules. The 64 molecules were assigned to the odor classification with the classes "floral," "medicinal," "woody, resinous," "disgusting," "fruity, non-lemony," and "perfumed." For training the model, 63 of the 64 molecules were used, each of which had a known class. A mathematical model for the odor classification was created. The class of the remaining molecule was then calculated using the mathematical model. Various weighting functions a,j and / or various selections of structural patterns were used for this purpose. The following table in Figure 2 presents the results. The accuracy when estimating the odor of a molecule is 21.35%.This means that the method according to the invention can classify the odor of molecules with at least twice the accuracy compared to estimation alone. The most accurate classification results using the mathematical model were obtained when a,j was a tf-idf weighting. The accuracy was over 65%. When calculating the accuracy, all molecules that could not be classified were considered "false."
[0065] For two of the molecules, no classification could be calculated. One was for hexanol, because it only exhibited structural patterns that occur in all classes. And for thiophene, which, in turn, only exhibited structural patterns that occur exclusively in this molecule of the 64 molecules. Therefore, the mathematical model could not provide probabilities for these structural patterns.
[0066] Example 3 - Permitting the use of chemicals in the cosmetics and personal care industry
[0067] The method according to the invention was used to predict the use approval of chemicals in cosmetics and personal care. A dataset consisting of 800 molecules (400 with and 400 without a use approval) and 500,047 structural fragments was used to train the mathematical model. The mathematical model, using the tf-idf-weighted conditional probability Pr(Cj|Fj), was able to replicate with an accuracy of over 85% whether molecules in the training dataset had a use approval. The mathematical model was used to predict the use approval for 200 additional molecules (100 with and 100 without a use approval). The results were compared with the FCM and Articles Regulation, Annex 11 - Restricted Substances, Annex 11 of the European Chemicals Agency (ECHA). Only 11 molecules were incorrectly classified as permitted. The overall accuracy was 81%.Thus, the process according to the invention can significantly save labor and personnel costs in the synthesis of chemicals for cosmetics and personal care by focusing more on predicted permitted substances.
[0068] literature
[0069] [1] a) LG Fine, CE Riera, Frontiers in physiology 2019, 10, 1151 ; b) P
[0070] Morquecho-Campos, K. de Graaf, S. Boesveldt, Food quality and preference 2020, 85, 103959.
[0071] [2] JE Taylor, H. Lau, B. Seymour, A. Nakae, H. Sumioka, M. Kawato, A. Koizumi, Frontiers in Neuroscience 2020, 14, 255.
[0072] [3] NB Tran, DR Kepple, SA Shuvaev, AA Koulakov, International Conference on Machine Learning 2019, 6305.
[0073] [4] a) A. Keller, RC Gerkin, Y. Guan, A. Dhurandhar, G. Turu, B. Szalai, JD Mainland, Y. Ihara, CW Yu, R. Wolfinger, Science 2017, 355, 820; b) H. Li, B. Panwar, GS Omenn, Y. Guan, Gigascience 2018, 7, gix127.
[0074] [5] K. Snitz, A. Yablonka, T. Weiss, I. Frumin, RM Khan, N. Sobel, PLoS computational biology 2013, 9, e1003184.
[0075] [6] L. Shang, J. Liu, Y. Tomiura, K. Hayashi, Analytical chemistry 2017 , 89, 11999.
[0076] [7] a) CS Sell, Angewandte Chemie International Edition 2006, 45, 6254; (b) M. Genva, T. Kenne Kemene, M. Deleu, L. Lins, M.-L. Fauconnier, International journal of molecular sciences 2019, 20, 3018.
[0077] [8] KJ Rossiter, Chemical reviews 1996, 96, 3201.
[0078] [9] K. Kaeppler, F. Mueller, Chemical senses 2013, 38, 189.
[0079]
[0010] R. Kumar, R. Kaur, B. Auffarth, AP Bhondekar, PloS one 2015, 10, e0141263.
[0080]
[0011] RM Khan, C.-H. Luk , A. Flinker , A. Aggarwal , H. Lapid , R. Haddad , and N. Sobel .
[0081]
[0012] a) M. Zarzo, Journal of Sensory Studies 2008, 23, 354; (b) A. Koulakov, BE Kolterman, A. Enikolopov, D. Rinberg, Frontiers in Systems Neuroscience 2011 , 5, 65.
[0082]
[0013] MB Course, WR Rudnicki, J Stat Softw 2010, 36, 1.
[0083]
[0014] A. Keller, e-Neuroforum 2003, 9, 121.
[0084]
[0015] MB Course, A. Jankowski, WR Rudnicki, Fundamental Informaticae 2010, 101, 271.
[0085]
[0016] Hinton GE, Salakhutdinov RR, Science 2006, 313,504.
[0086]
[0017] D. Weininger, Journal of Chemical Information and Computer Sciences 1988, 28, 31.
[0087]
[0018] Daylight Chemical Information Systems, Inc., "3. SMILES - A Simplified Chemical Language", can be found under https: / / www.daylight.com / dayhtml / doc / theory / theory.smiles.html, 2019,
[0088]
[0019] Daylight Chemical Information Systems, Inc., "4. SMARTS - A Language for Describing Molecular Patterns", can be found under https: / / www.daylight.com / dayhtml / doc / theory / theory.smarts.html, 2019.
[0020] https: / / doi.Org / 10.1186 / S13321 -020-00445-4
[0089]
[0021] https: / / www.rdkit.org / UGM / 2012 / Landrum_RDKit_UGM.Fingerprints.Final.pptx.pdf
[0090]
[0022] http: / / www.talete.mi.it / products / dragon_molecular_descriptor_list.pdf
[0091]
[0023] https: / / ai.googleblog.com / 2019 / 10 / learning-to-smell-using-deep-learning.html
[0024] arXiv:1910.10685v2
[0092]
[0025] Chris Manning and Hinrich Schütze, Foundations of Statistical Natural Language Processing, MIT Press. Cambridge, MA: May 1999
[0093]
[0026] Heiner Strickenschmidt, Ontotogies: Concepts, Technologies and Applications, Springer Verlag, 2009
Claims
Claims 1. A method for selecting molecules with a desired physical, chemical and / or physiological property from a group of molecules, comprising the steps • Providing a group of Ok molecules by a user, where ke N; • Providing a classification according to a chemical, physical and / or physiological property of a molecule, comprising C, classes, where N; • Providing a mathematical model for classification, wherein the mathematical model describes relationships G,j between a structural pattern and a class, in particular by probabilities that a structural pattern Fj of a molecule belongs to a class C, or that a molecule of a class C, has a structural pattern Fj; • Selection of a weighting function a,j for the mathematical model by a user; • Assignment of all Ok molecules into the C, classes of classification by the mathematical model, wherein the mathematical model comprises the steps: a) Determination and storage of Fj structural patterns of the chemical structure of each of the Ok molecules assigned to the respective molecule, where N each; b) Assignment of the probability Gij to each structural pattern Fj of a molecule for each class C, and calculation of the influence / ,j according to the formula for each structural pattern Fj of a molecule for each class C,; c) calculation of a score P iik for each molecule Ok, using for each class C,, where the influences ,j of all structural patterns Fj contained in a molecule Ok are summed up for each class C, d) assignment of each molecule to the class C with the highest score P iik for the respective molecule; • Display and / or output of the molecules assigned to the classes of the classification and optionally the associated point values P iik , the associated influences / ,7and the structural patterns Ff, • Selection of molecules assigned to the class with the desired physical, chemical and / or physiological property; • Experimental confirmation of the physical, chemical and / or physiological properties of at least some of the selected molecules by a user; and / or verification and / or identification of the relationship between at least one structural pattern Fj and a class C by a user. Method according to claim 1, characterized in that the display and / or output of at least some of the molecules is carried out in such a way that the molecules are arranged in descending order according to their point value P iik in a class C, starting with the molecule with the highest score P iik. Method according to one of the preceding claims, characterized in that the mathematical model is trained by a training data set for the selected classification, wherein a training data set comprising 0 / molecules of known class C is specified, where I, ie N, comprising the steps i. Determining and storing Fj structural patterns of the chemical structure of each molecule assigned to the respective molecule, each N ; ii. Calculating the probability Gy that a structural pattern Fj belongs to a class C, where Calculation of the probability G that a molecule of a class C has a structural pattern Fj, where G t j = Pr(c F7) = Method according to claim 3, characterized in that the Fj structural patterns determined in step i) are selected by an algorithm, an idf weighting, or a tf-idf weighting. Method according to one of the preceding claims, characterized in that the classification is selected from the group of structure-based properties of molecules, in particular from the group containing odor, taste, color, water solubility, toxicity, and permitted and prohibited chemicals in cosmetics and personal care. Method according to one of the preceding claims, characterized in that the weighting function a,j is selected from the group of statistical measures, such as tf-idf functions, normalization functions, and equally weighted functions. Method according to one of the preceding claims, characterized in that all molecules are experimentally investigated whose score P iiknot more than 50%, preferably not more than 30%, particularly preferably not more than 10% of the highest point value P iik deviates in this class. Use of the method according to one of claims 1 to 7, characterized in that the method is used to select at least one molecule with a desired chemical, physical and / or physiological property from a group of molecules. Use of the method according to one of claims 1 to 7, characterized in that the method is used to identify the influence of structural patterns in molecules on at least one chemical, physical and / or physiological property of molecules.