Computer olfaction model
By training a machine learning model to fit the affinity between odor molecules and olfactory receptors, a comprehensive database is created, solving the problem of inaccurate odor prediction in existing technologies and achieving accurate olfactory perception prediction of odor molecules in human subjects.
Patent Information
- Application Number
- CN202480045097.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-03
- Filing Date
- 2024-07-02
- Publication Date
- 2026-02-03
AI Technical Summary
Existing technologies lack precise methods and automated means to analyze and reliably measure the association between the structure of a substance and its olfactory perception induced in subjects.
By providing a database containing the three-dimensional structure of olfactory receptors (ORs) and the structure of odor molecules, a machine learning model is trained to fit the affinity between odor molecules and olfactory receptors, creating a comprehensive database to predict the olfactory sensation of odor molecules in human subjects.
It enables accurate prediction of olfactory perception of odor molecules in human subjects, and can predict the odor experience of odor molecules in subjects based on their structure and properties, thus improving the accuracy and automation of odor prediction.
Smart Images

Figure CN121464484A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a computer-implemented method for predicting molecular odors, comprising: providing a first database containing three-dimensional (3D) structures and / or chemical structures of one or more olfactory receptors (ORs); training a machine learning model or artificial intelligence (AI) based on data contained in the first database; providing a second database containing (i) (3D) structures of one or more odorant molecules and (ii) corresponding textual descriptions of the odor sensations induced by the one or more odorant molecules in human subjects; training a machine learning (ML) model or artificial intelligence (AI) based on data contained in the second database; and fitting the (3D) structures of one or more odorant molecules to the (3D) structures of one or more ORs to predict the odor of the one or more odorant molecules. The method involves: determining an affinity score for each fitted combination of an odor molecule with its corresponding olfactory receptor, wherein fitting includes using and training an ML model or AI, and wherein the affinity score indicates the degree of fit between the corresponding odor molecule and the OR; creating a comprehensive third database or third data embedding structure / layer containing data from a first database, data from a second database, and the affinity scores determined for each fitted combination of the corresponding odor molecule with its corresponding olfactory receptor; providing a (3D) structure of at least a first target odor molecule not included in the second database; and predicting, by machine learning or AI, the odor sensation induced by at least the first target odor molecule in human subjects, based on the data contained in the third database or third data embedding structure / layer. Background Technology
[0002] Odor perception is a highly complex process. After an odorant enters the nasal cavity, it is recognized by olfactory sensory neurons (OSNs), which contain various types of olfactory receptors (ORs). Subsequently, the neural signals from the OSNs are transmitted to the olfactory cortex of the brain for processing (Fleischer et al., 2009).
[0003] Generally speaking, the olfactory receptor (OR) group includes odorant receptors (OR), vomeronasal receptors (V1R and V2R), trace amine-associated receptors (TAAR), formyl peptide receptors (FPR), and guanylate cyclase GC-D, which mainly belong to the G protein-coupled receptor protein (GPCR) family (Fleischer et al., 2009).
[0004] Existing technologies indicate that the three-dimensional structure of odorants and optionally their physicochemical properties (such as boiling point, vapor pressure, and molecular polarity) can influence the olfactory experience (odor) perceived by a subject. The association between the chemical structure and properties of such odorants and the induced biological response is known as the quantitative structure-activity relationship (QSAR) model (Dearden, 1994; Hau and Connell, 1998).
[0005] However, existing technologies still lack precise methods and automated means to analyze and reliably measure the relationship between the structure of a substance and its olfactory perception induced in subjects.
[0006] Given the existing technology, there is still an urgent need in the field to provide improved methods for predicting molecular odors or olfactory perception. Summary of the Invention
[0007] In view of the prior art, the technical problem to be solved by the present invention is to provide alternative and / or improved methods for predicting molecular odor or olfactory perception.
[0008] This problem is solved by the features described in the independent claim. Preferred embodiments of the invention are provided in the dependent claims.
[0009] The first aspect of this invention relates to a computer-based method for predicting molecular odor (odor / olfactory sensation), comprising the following steps: a. Provide a first database containing information about the structure (preferably three-dimensional (3D) structure and / or chemical structure) of one or more olfactory receptors (ORs), preferably, wherein the database contains information about the 3D structure of the ligand (or agonist) binding site of the one or more olfactory receptors; b. Train a machine learning model or AI (receptor encoder) based on the data contained in the first database; c. Provide a second database containing (i) the (3D) structure of one or more odor (odorant) molecules and preferably (ii) a corresponding textual description of the odor sensation induced by said one or more odor molecules in human subjects; d. Train a machine learning (ML) model or artificial intelligence (AI) (odorant encoder) based on the data contained in the second database; d.1. Optionally provide interaction data between one or more olfactory receptor molecules from a first database and one or more odor molecules from a second database; e. Fitting the (3D) structure of one or more odor molecules to the (3D) structure of one or more olfactory receptors (ORs), preferably to the ligand binding sites of the one or more ORs, thereby determining an affinity score for each fitted combination of the one or more odor molecules and the corresponding olfactory receptors. The fitting process includes using and training an ML model or AI (paired interaction encoder), optionally, wherein the use and / or training of the ML model or AI also takes into account the interaction data provided in step d.1. The affinity score (a numerical value) indicates the degree of fit between the corresponding odor molecule and the olfactory receptor (OR), preferably the degree of fit with the ligand binding site of the corresponding OR, and preferably indicates the ability (i) of the corresponding odor molecule to bind or interact with the corresponding olfactory receptor and / or (ii) to activate the corresponding olfactory receptor (and / or the intensity of such interaction). f. Create a comprehensive third database or third data embedding structure / layer containing data from the first database, data from the second database, and affinity scores determined for each fitted (preferably fitted in step e) combination of the corresponding odor molecule and the corresponding olfactory receptor; g. Provide the (3D) structure (and / or any other structural and / or sequence information) of at least the first target odor molecule not included in the second database; and h. Based on data contained in a third database or a third data embedding structure / layer, predict, via machine learning or AI, the odor sensation induced by at least the first target odor molecule in human subjects.
[0010] In an embodiment of the method of the present invention, the data contained in the first database, the data contained in the second database, and the affinity scores (obtained in step e) for each fitted combination of the corresponding odor molecule and the corresponding olfactory receptor, together with optionally other data, are all stored in a third database or a third data embedding structure / layer.
[0011] In the implementation scheme, after fitting one or more odor molecules to one or more ORs, the obtained data may be normalized and / or scaled, for example, by normalizing the data of all or more analyzed ORs using multiple methods (non-limiting examples such as t-scores).
[0012] In one implementation, pre-training is performed, including training the olfactory receptor encoder and / or the odor molecule encoder separately, preferably in conjunction with the paired interaction encoder layer (e.g., in the implementations of steps b, d, and f above).
[0013] In the implementation scheme, pre-training includes learning latent spatial embeddings and / or intermolecular relationships using data from one or more of a first database, a second database, or a third database, or other data (e.g., ligand-OR interaction data and / or universal ligand-receptor and / or ligand-protein interaction data).
[0014] In an embodiment of the method of the present invention, step h: predicting the odor sensation induced by at least the first target odor molecule in a human subject is achieved by using one or more of the machine learning models or AIs trained in steps b, d, and / or e above, based on the (3D) structure (and / or any other structural and / or sequence information) of at least the first target odor molecule to predict: - At least one odor molecule contained in a third database or a third data embedding structure / layer, wherein the odor molecule has the highest similarity to at least a first target odor molecule in terms of affinity score to one or more olfactory receptors contained in the third database or the third data embedding structure / layer; and / or - At least the odor sensation induced by the first target odor molecule in human subjects, wherein the odor sensation preferably corresponds to one or more textual descriptions of the odor sensation induced by the odor molecule in human subjects contained in a third database or a third data embedding structure / layer.
[0015] In an embodiment of the method of the present invention, step h: predicting the odor sensation induced by at least the first target odor molecule in a human subject includes: (I) Using the ML model or AI (odorant encoder) trained in step d, compare the (3D) structure (and / or any other structural and / or sequence information) of at least one or more of the first target odor molecule and the odorant molecules contained in the second or third database; (II) Using the ML model or AI (paired interaction encoder) trained in step e, fit the (3D) structure (and / or any other structural and / or sequence information) of at least the first target odor molecule to the (3D) structure of one or more olfactory receptors (ORs), preferably to the ligand binding site of the OR, thereby determining the affinity score of each fitted combination of at least the first target odor molecule with the corresponding OR; (III) Compare one or more of the affinity scores determined in step (II) for each fitted combination of at least the first target odor molecule with one or more olfactory receptors (ORs) and the affinity scores of one or more ORs for each combination of one or more odor molecules contained in a third database or a third data embedding structure / layer; and (IV) Accordingly, identify at least one odor molecule contained in a third database or a third data embedding structure / layer, wherein the odor molecule has the highest similarity score to at least a first target odor molecule in terms of affinity for one or more ORs. The corresponding textual description of the odor sensation induced in human subjects by the one or more odor molecules identified in step (IV), contained in the third database or the third data embedding structure / layer, indicates the odor sensation induced in human subjects by at least the first target odor molecule.
[0016] In the implementation plan, at least the first target odor molecule is chemically synthesized based on the (3D) structure provided in step g above.
[0017] In an implementation, predicting the odor sensation induced in human subjects by at least a first target odor molecule includes ultimately generating and / or displaying (e.g., via a graphical user interface) text output containing a description of the odor sensation induced in human subjects as predicted in step h.
[0018] In its implementation, the present invention includes a computer model for predicting molecular odors, wherein the molecular odor refers to a subject's (subjective) olfactory experience (odor sensation / olfactory perception), or the (subjective) odor sensation produced in the subject. Although human subjects' descriptions of smell or odor are subjective, the method of the present invention preferably facilitates the prediction of average human descriptions of a particular odor, or facilitates expert predictions of a particular odor.
[0019] In the embodiments, the databases (such as a first database, a second database, and / or a third database) contain information about the 3D structure of ligand binding sites for one or more olfactory receptors. In the embodiments, the databases (such as a first database, a second database, and / or a third database) (also) contain information about the DNA, RNA, and / or amino acid sequences of ORs and / or olfactory molecules, and when training machine learning models or AI (receptor encoders) based on the data contained in said databases, only the DNA, RNA, and / or amino acid sequence data of the OR and / or olfactory molecule are considered. In the embodiments, the fitting steps and / or odor perception prediction also consider only or simultaneously the said DNA, RNA, and / or amino acid sequence data. In the embodiments, the term "structural data" can also be understood to include DNA, RNA, and / or amino acid sequence data / information.
[0020] In a preferred embodiment, the first database contains structural data (three-dimensional (3D) structures and / or chemical structures) of one or more olfactory receptors (ORs) (preferably multiple ORs). In this embodiment, the first database contains structural data of preferably more than two, three, four, or five, and more preferably at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 1000, 1500, or 2000 olfactory receptors (and / or receptors involved in odorant perception (olfactory perception) and / or auxiliary (or transport) factors involved in odorant perception).
[0021] In a preferred embodiment, the second database contains (i) the (3D) structures of one or more (preferably multiple) odor (odorant) molecules and (ii) corresponding textual descriptions of the odor sensations induced by the corresponding odor molecules in human subjects. In this embodiment, the second database contains structural data of preferably more than two, three, four, or five, and more preferably at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 1000, 1500, or 2000 odor (odorant) molecules, along with corresponding textual descriptions of the odor sensations induced by the corresponding odor molecules in human subjects.
[0022] In embodiments, it is advantageous if the first database and / or the second database each contain data on a variety of molecules (receptors / proteins / molecules), preferably more than 100 molecules, because in these embodiments, the method of the present invention focuses not only on the absolute numerical value of the affinity between the odorant molecule and the receptor and / or cofactor, but also on the ratioal patterns, similarities, differences, and / or correlations of the affinity between the odorant and different receptors and / or cofactors. For example, this method may be useful when the odorant molecule exhibits a specific “affinity pattern” (binding pattern) to a particular receptor structure (e.g., related or similar structures of two or more receptors) and / or cofactor structure (e.g., related or similar structures of two or more factors). In embodiments, machine learning algorithms can learn from a particular odorant structure and its affinity score “pattern” to a particular receptor or molecular structure to predict the binding or affinity patterns of new, unknown (or only new / unknown to the model) molecules, thereby predicting the odor or olfactory sensation / perception induced in subjects.
[0023] In a preferred embodiment, step e: fitting the (3D) structure of one or more odor molecules to the (3D) structure of one or more olfactory receptors (ORs), including fitting one or more or various or a large number of odor molecules to various or a large number of olfactory receptors (ORs). In a preferred embodiment, the target odor molecule is also fitted / compared to various or a large number of odor molecules (step (I)), and / or subsequently fitted to various or a large number of olfactory receptors (step (II)).
[0024] In other words, in embodiments, the method of the present invention aims to identify patterns arising from the measured affinity (affinity scores) of a scent agent to different receptors and / or cofactors. In the context of the present invention, preferably, not only the actual numerical values of the measured affinity scores can be used, but also the ratio patterns occurring between affinity scores of different receptors and / or cofactors can be used to make the different predictions described herein. Non-limiting examples may include: a particular scent agent must have low affinity scores to several specific receptors to induce the perception of an orange blossom scent, while having high affinity scores to several specific receptors induces the perception of a rose scent.
[0025] In some implementations, affinity scores include or constitute multidimensional numerical values, such as vectors and / or tensors.
[0026] In some implementations, using AI or ML models to predict molecular odors has the following advantages: the trained ML model or AI (e.g., including and / or employing one or more encoders and / or neural networks) is not only able to predict molecular odors based on the measured affinity score for ORs, but also by taking into account other data, such as the (predicted) conformational (changes) induced by the interaction of receptors with odor molecules. This allows the AI or ML model (e.g., including a deep neural network-based encoder) to determine the different olfactory sensations induced by the two odor molecules in the subject, even if the AI or ML model has determined that the two odor molecules have the same affinity / affinity score for one or more ORs (from a first database). This is because, for example, they bind to the receptor with comparable affinity but produce (or induce) different conformations, such as changing the overall energy level of the interaction.
[0027] Therefore, in embodiments, the method of the present invention further includes determining the most similar pattern (the highest overlap of affinity score patterns (for a specific receptor / factor)) between the corresponding target odor molecule and one or more odor molecules in a second or third database. Thus, in embodiments, the method of the present invention determines one or more odor molecules in the second or third database that have the most similar receptor binding patterns or behaviors or the highest overlap of affinity scores for a specific or all receptors. In some embodiments, the compared affinity scores do not need to be exactly the same for a specific receptor, but should have the same "trend" (correlation) in their affinity score values. Depending on the specific embodiment and data, affinity score values considered to have the "same trend" may have affinity score differences of 0-1%, 0-5%, 0-10%, 0-15%, 0-20%, or 0-30% (e.g., where related affinity scores indicate strong, moderate, or weak binding / affinity, i.e., showing similar affinity patterns).
[0028] In other words, in this embodiment, the method of the present invention aims to find / match odor molecules in a second or third database that have (highly) similar, comparable, or identical affinity (affinity score) to olfactory receptors / factors in the first database for a target odor molecule. In this embodiment, once an odor molecule with the most similar affinity (affinity score) to the target odor molecule is identified from the second or third database, the odor of the target odor molecule can be predicted using textual descriptions of the odor of the "similarly bound" odor molecule in the second or third database. In this embodiment, this process can be performed by one or more ML models or AIs trained on data from the first, second, and / or third databases and optionally other data or data embeddings.
[0029] In some embodiments, the step of “providing a first database” includes: (firstly) determining the three-dimensional (3D) structure and / or chemical structure of said one or more olfactory receptors (ORs) (and / or optionally one or more receptors involved in odor perception and / or their auxiliary (or transport) factors involved in odor perception) from a data containing one or more nucleic acid sequences (e.g., mRNA, RNA, and / or DNA sequences) and / or protein (amino acid) sequences and / or SMILES strings or other suitable data containing one or more olfactory receptors (ORs) (and / or optionally one or more receptors involved in odor perception and / or their auxiliary (or transport) factors involved in odor perception), thereby preferably providing a first database containing said structures. In embodiments, the structure can be determined by computational methods (e.g., computer simulation structure prediction), mathematical calculations, and / or chemical methods (e.g., crystallographic methods).
[0030] In the implementation, a machine learning (ML) method or model or AI is used, preferably trained on data based on a first database, a second database, and / or a third database or a third embedding structure. In the implementation, to save computational resources and / or because some ML models will achieve better results with a shorter (selected) feature list, prediction and / or training are performed using only a selection or subset of, for example, the first, second, and / or third databases, containing only 100, 50, or 25 minimally (e.g., structurally) relevant receptors / factors.
[0031] In the implementation scheme, machine learning (e.g., neural networks, neural networks) is used to automate the comparison of databases, for example, using statistical correlation analysis. In the implementation scheme, before making a prediction, a machine learning model or AI (e.g., a neural network or deep neural network) is trained separately or in parallel on a first database, a second database, and / or a third database, such that the AI learns the affinity (scoring) patterns between the odor molecule structures and receptor (binding site) structures contained in the databases. For example, the machine learning model / AI is trained only on molecule / receptor structures and their corresponding affinities, such that, for example, the model knows how to exclude similar receptors / proteins and / or select irrelevant receptors / proteins from the first or third database, thereby reducing the number of proteins / receptors used for prediction in the first or third database.
[0032] In the implementation plan, textual descriptions of olfactory perceptions (induced in subjects) that can be perceived by human subjects, such as fragrance, odor, or odor, may be expressed using terms such as "musky," "sweet," "floral," "fruity," "rotten," "spicy," "fresh," "intense," "lemon," "citrus," "limonene," "musky," "rose," or "rose top note," "orange blossom middle note," and "lemon base note." Those skilled in the art are familiar with various fragrance description methods used by perfumers and in the relevant fields.
[0033] In the implementation, the predicted odor sensation induced by at least a first target odor molecule includes (text) output describing the relevance and intensity of the odor sensation induced by the target molecule (e.g., “The scent of this new molecule reminds me of limonene, which is 85% similar to it, and linalool, which is 75% similar to it,” or “It reminds me of a citrus scent in a women’s perfume that is 86% similar to it”).
[0034] In the implementation plan, similarity can be expressed as a numerical score between 0 and 10, between 0 and 100, or as a percentage between 0% and 100%, such as "similarity reaches 80%" or "affinity score difference 0.5".
[0035] In the embodiments, the affinity score may have a value between 0.001 and 1, 0 and 10, 0 and 100, -100 and 100, -50 and 50, -10 and 10, or -1 and 1, wherein the score value is proportional (related) to the binding strength between two molecules (e.g., an odor molecule and its binding site on another molecule (e.g., a receptor or factor). Therefore, in the embodiments, the affinity score may have a value of -100, -50, -10, -5, -3, -2, -1, -0.5, -0.05, -0.1, -0.01, 0.01, 0.05, 0.1, 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, or any value in between. In one implementation, strong / high affinity may be indicated by an affinity score of 1, while weak / low affinity may be indicated by an affinity score of 0.1. In other implementations, strong affinity may be indicated by an affinity score of 100, while weak affinity may be indicated by an affinity score of 1, and so on.
[0036] In some implementations, higher-dimensional "affinity scores" can be measured, such as numerical vectors of length 1-10, 10-100, 10-500, or 50-1000. Such implementations are suitable for situations where, for example, AI or ML models learn more complex correlation / interaction patterns between ORs and ligands (odor molecules), such as patterns that cannot be described by a single numerical score.
[0037] In some embodiments employing machine learning methods to determine binding (affinity) patterns, the machine learning model preferably learns / trains how to identify specific affinity / binding patterns between a molecular structure and one or more receptor (and / or cofactor) structures. Therefore, in embodiments, the machine learning model / AI is preferably trained or learned how to identify, for a target odor molecule, one or more odor molecules from a second or third database that have the most similar affinity patterns or binding behaviors to one or more receptors (and / or cofactors), thereby determining the odor, olfactory perception, or olfactory sensation of the target odor molecule. In embodiments, the machine learning model / AI is preferably trained or learned to determine its odor, olfactory perception, or olfactory sensation solely based on the structure of the target odor molecule, even without determining or comparing affinity scores between the odor molecule and receptors (and / or cofactors).
[0038] In the implementation scheme, step g(I) determines the specific receptor-affinity signature of the target odor molecule and one or more odor molecules in a second or third database, thereby identifying one or more odor molecules from the second or third database that have the highest similarity to the receptor affinity pattern of the target odor molecule. In the implementation scheme, this determination is achieved through statistical methods: by calculating probabilities and / or likelihoods, and / or by analyzing which receptors are consistently associated with each other, for example, molecules with structure / domain Z bind to class X and class Y receptors with 80% likelihood.
[0039] In embodiments, the method of the present invention is based on the structural information of a receptor (e.g., a G protein-coupled receptor (GPCR)), preferably three-dimensional structural information (preferably 3D chemical structure), which participates in the olfactory process: including but not limited to olfactory receptors (ORs) and other receptors optionally involved in the perception / recognition of olfactory molecules (odors), and optionally also including helper / transporter proteins (e.g., OBP2A, OBP2B, etc.) that transport odor molecules (e.g., hydrophobic and / or volatile small molecules or odorants) to the receptor. In embodiments, the method of the present invention considers not only the structural information of the OR and / or olfactory molecules, but also the DNA, RNA, and / or amino acid sequence data encoding the OR and / or olfactory molecules. In embodiments, the method of the present invention considers only the DNA, RNA, and / or amino acid sequence data encoding the OR and / or olfactory molecules, without considering any (other) structural information.
[0040] In the implementation, the term "olfactory receptor (OR)" also includes one or more (other) receptors and / or their cofactors involved in the perception of odorants.
[0041] In the implementation scheme, at least one olfactory receptor system is selected from the group consisting of: odorant receptor (OR), formyl peptide receptor (FPR), trace amine-associated receptor (TAAR), vomeronasal receptor (V1R and V2R), and guanylate cyclase GC-D (these are primarily protein-coupled receptors (GPCRs)).
[0042] In one embodiment, at least one cofactor is an odorant-binding protein, such as OBP2A or OBP2B. In another embodiment, at least one cofactor is an odorant-binding protein, a lipoprotein, or any other carrier protein for hydrophobic molecules, such as odor molecules or pheromones.
[0043] In this implementation, the (3D) structure of the binding sites of all relevant receptors involved in odor perception can be determined. This can be achieved, for example, using experimental crystallographic information, computer simulations or any machine learning model based on experimental data, the genetic code, any combination of the above methods, or other methods.
[0044] In the implementation scheme, the (3D) structure of the molecule or its parts, such as the agonist or ligand binding site and / or cofactor of the receptor (e.g., the receptor in the first database), is determined using experimental crystallographic information, computer simulations, AI and / or machine learning (ML) models trained on experimental data, crystallographic information, gene sequence information of the corresponding olfactory receptor, or any combination thereof.
[0045] In the implementation scheme, the second and / or third databases contain a curated dataset of odor molecules, including their chemical (3D) structures and textual descriptions of their odors (preferably provided by human experts or otherwise collected).
[0046] In the implementation scheme, the first database and / or the third database contain a curated dataset of three-dimensional (3D) (chemical) structures of one or more olfactory receptors (ORs) and / or optionally one or more receptors involved in odorant perception (olfactory perception) and / or auxiliary (or transport) factors involved in odorant perception.
[0047] In one embodiment, an affinity score can be assigned to each odor molecule contained in the second database, which indicates / classifies / grades / evaluates the binding strength and / or fit of the odor molecule with each receptor (or its binding site / pocket) or only a subset thereof contained in the first database. In other embodiments, instead of determining the affinity of the first odor molecule with all receptors, the method of the present invention uses statistical tools and machine learning (ML) tools to identify the dependencies and / or correlations between the chemical (3D) structures of different receptors in the first database, thereby scoring / grading the affinity / binding strength and / or fit of the corresponding odor molecule with a simplified list of receptors (e.g., by setting a threshold for the correlation coefficient between different receptors, scoring only the least relevant receptors in the first database, such as 10, 15, 20, 25, 30, 35, 40, 45, 50, or 100, or 10-500, 10-250, 10-100, 10-50, or 10-25, rather than scoring all of the receptors contained therein).
[0048] In an implementation scheme, for example in step e, the dependencies and / or correlations between different olfactory receptors and / or odor molecules are identified by using statistical methods and / or machine learning-based methods. Affinity scores are determined only for subsets of odor molecules and / or olfactory receptors contained in the first, second, and / or third databases. Optionally, affinity scores are assigned to 25-100 or 25-200 least relevant receptors and / or odor molecules by setting a threshold for the correlation coefficients between different receptors and / or odor molecules.
[0049] In implementation schemes, affinity scores can be determined using the most commonly used computer physics simulation methods, such as the well-known software tool AutoDock Vina. Many other software solutions using similar or different algorithms are currently available. Alternatively, ML algorithms (e.g., Gnina, "A deep learning framework for molecular docking") or algorithms combining statistical tools with ML tools and simulation elements (e.g., algorithms, models, or software) can be employed.
[0050] In implementations, for example in step e, affinity scores are determined using computer physical simulations, ML algorithms, and / or a combination of statistical methods and ML methods, as well as algorithms that compute (3D or interaction) simulation elements (e.g., algorithms, models, or software).
[0051] In the implementation, the AI or ML model can learn more complex correlations between odor molecules and OR interactions (e.g., affinity and / or conformational changes during interaction), rather than measuring “simple” affinity scores, and then use this correlation to predict the odor (olfactory sensation) induced by the target odor molecule.
[0052] In one implementation, after determining the corresponding affinity score, the method of the present invention can predict the olfactory sensation / perception (odor) of a molecule that has never been encountered before, based on existing data and a predictive model. In another implementation, this prediction can be achieved using any number of statistical tools and machine learning tools and algorithms, including but not limited to: simple statistical methods that compare the affinity of the unknown molecule with known molecules in the dataset and calculate a simple distance metric. In yet another implementation, the nearest neighbor molecule can be selected to find the highest statistical consistency among odor-related text descriptions.
[0053] In the implementation, methods that outperform and / or are more accurate than basic statistical methods may be employed, such as any number of ML algorithms, including but not limited to (deep) neural networks trained on the relevant data and / or any database or public database described herein (such as PubChem or any database of publicly available SMILES strings).
[0054] The databases described herein may be the first, second, or third database as described herein, any public database (such as PubChem or any database that publishes SMILES strings), and / or databases containing information / data on the chemical composition, chemical structure, RNA sequence, DNA sequence, peptide sequence, SMILES strings, and / or corresponding olfactory perception (odor) of commercially available perfumes, natural products (such as flowers, fruits, foods, and beverages) and any other analytical reports describing the odorant components of natural or artificial odorant articles, or any combination thereof. In other words, in embodiments, the databases (such as the first, second, and / or third databases) contain data on the structure (preferably three-dimensional (3D) structure and / or chemical structure) of one or more olfactory receptors (ORs) and / or odorant molecules, and optionally contain amino acid sequences and / or SMILES strings of one or more receptors and / or nucleic acid sequences (e.g., RNA and / or DNA sequences) respectively encoding said olfactory receptors (ORs) and / or odorant molecules.
[0055] In the context of this invention, the terms "information" and "data" are used interchangeably.
[0056] In the embodiments, the second and / or third databases further contain vapor pressure and / or boiling point data of the odor molecules contained therein. In the embodiments, when determining odor molecules with the most similar affinity scoring pattern, the boiling point and / or vapor pressure data may also be considered. In the embodiments, the second and / or third databases further contain data on vapor pressure, boiling point, molecular weight, partition coefficient (logP), topological polar surface area, number of hydrogen bond donors and acceptors, number of aromatic rings and rotatable bonds, and / or other molecular properties of the odor molecules contained therein. In the embodiments, when determining odor molecules with the most similar affinity scoring pattern, the boiling point, vapor pressure, molecular weight, partition coefficient (logP), topological polar surface area, number of hydrogen bond donors and acceptors, number of aromatic rings and rotatable bonds, and / or other molecular properties may also be considered.
[0057] In the implementation scheme, the statistical method in step g includes: comparing the determined affinity scores of the target odor molecule with one or more olfactory receptors with the affinity scores of one or more odor molecules contained in a second database, and calculating a (simple) distance metric; and / or wherein the machine learning method in step g includes: determining the nearest neighbor odor molecule to find the highest statistical consistency between the affinity scores of odor molecules in the second database and the first target odor molecule, thereby predicting the odor sensation induced by the target odor molecule in human subjects based on the odor-related text descriptions of the nearest neighbor odor molecules in the second database.
[0058] In the implementation scheme, at least the first target odor molecule comprises a mixture of at least the first and second target odor molecules. And wherein in step g, in addition to providing (i) the (3D) structure of at least the first and second target odor molecules, (ii) the relative percentages of at least the first and second target odor molecules in the vapor of the mixture and / or (iii) the statistical probability of each odor molecule interacting with olfactory receptors in the nasal cavity of a human subject, and wherein step h includes: predicting the odor sensation induced in a human subject by the mixture of at least the first and second target odor molecules.
[0059] In some such embodiments, the relative percentages of at least the first and second odor molecules in the vapor of the mixture are determined by taking into account both the relative percentages of at least the first and second odor molecules in the mixture and the corresponding vapor pressure of the mixture.
[0060] In embodiments where at least one target odor molecule is a mixture of at least two different odor molecules, the relative percentage of each molecule in the vapor of the mixture is determined by taking into account the relative percentage of the different odor molecules in the mixture and / or the corresponding vapor pressure of the mixture.
[0061] The method includes using the prediction model described above (first aspect of the invention) and combining it with information about the molecules that make up the mixture (e.g., orange essential oil, Chanel No. 5 perfume, etc.).
[0062] By calculating the relative abundance of each molecule in the mixture and its vapor pressure, the relative abundance of each molecule in the vapor of the mixture and the statistical probability of interaction between each component and olfactory / odor receptors in the nasal cavity can be calculated. Active transport of odor-binding proteins (OBPs) can be used, but is not required. By combining this prediction with the affinity score of each molecule and using the same method (1) described above for each molecule, the odor of the mixture at any given moment from the start of exposure to complete drying can be predicted.
[0063] In embodiments where the target odor molecule is a mixture of two or more odor (target) molecules, the vapor pressure (or boiling point) of each target odor molecule is determined (based on data or using a mathematical or computer model). In embodiments, the affinity score of each target odor molecule in the mixture may be determined sequentially or in parallel. In embodiments, the final affinity pattern is corrected, normalized, or calculated given the differences in vapor pressure (or boiling point) between target odor molecules in the mixture. For example, if a first target odor molecule has a lower vapor pressure (boiling point), fewer of these molecules will reach the nasal cavity or the subject compared to other molecules with higher vapor pressures (boiling points), resulting in a lower proportion of the final odor perceived by the subject.
[0064] In an embodiment, the second and / or third databases further contain vapor pressure and / or boiling point data for the odor molecules contained therein. In an embodiment, the boiling point and / or vapor pressure data are also considered when determining odor molecules with the most similar affinity scoring patterns. In an embodiment, this consideration includes subsequent or parallel normalization of the odor of the mixture, for example, by weighting the effect of each target odor molecule in the mixture (e.g., in percentage form).
[0065] In one specific, non-limiting, exemplary embodiment, each odor molecule Mi in the mixture is represented by its embedding vector Oi. This embedding vector can be obtained through any of the foregoing embodiments or any other embodiment of the present invention. In the weighting step, the summation weights can be derived in various ways. For the purposes of this embodiment, this document adopts the commonly used definition in the art, namely, applying relative concentration as the summation weight.
[0066] Subsequently, the relative concentration Ci of each odor molecule Mi in the air mixture was determined. This can be achieved through theoretical calculations or experimental measurements, such as headspace analysis using techniques like gas chromatography (GC), mass spectrometry, or carbon-13 nuclear magnetic resonance (NMR). In the complexation embedding step of the mixture, the complexation embedding vector Omix of the odor mixture was calculated as a weighted sum of the individual embedding vectors. The weights were defined based on the relative concentrations of the molecules.
[0067] Subsequently, the formula for calculating the composite embedding vector is: O mix = ∑wi • Oi Among them O mix represents the olfactory embedding vector of the mixture; wi represents the relative concentration of molecule Mi; Oi represents the olfactory embedding vector of molecule Mi.
[0068] In embodiments of the invention, a machine learning model / algorithm or artificial intelligence may be trained on one or more databases, such as a first database, a second database, and / or a third database, before predicting the olfactory experience (odor) of unknown or unevaluated molecules. These databases contain: three-dimensional (3D) structures of one or more olfactory receptors (ORs), other receptors involved in olfaction, and optionally helper / transporter proteins (e.g., proteins that transport odor molecules to receptors involved in odor perception); and / or (3D) structures of one or more odor molecules and corresponding textual descriptions of the odor sensation induced by said one or more odor molecules in human subjects; and / or numerical data describing the affinity (e.g., affinity scores) of odor molecules to binding sites of receptors (e.g., ORs) involved in olfactory perception (such as odor).
[0069] In embodiments of the invention, a machine learning model / algorithm or artificial intelligence may be trained to evaluate / analyze and / or identify the three-dimensional (chemical) structure of a molecule (e.g., provided in a database as described herein) before predicting the olfactory experience (odor) of an unknown or unassessed molecule, and to compare and / or fit the three-dimensional (chemical) structure of the molecule with the three-dimensional (chemical) structures of other molecules, thereby generating / predicting an affinity score or a numerical calculation of fitting / binding ability in the embodiments. For example, the model described herein may be trained to fit the three-dimensional (chemical) structure of an odor molecule (e.g., from a second database or newly predicted according to the invention) to a three-dimensional binding site / pocket of an olfactory receptor or other receptor involved in olfactory perception, and / or, for example, to compare the newly predicted molecule with the three-dimensional (chemical) structure of one or more known odor molecules (e.g., from a second database).
[0070] In the implementation, the model can use one or more molecular characterizations (e.g., the structure of the target odor molecule) to “directly” predict odor-related text descriptions and / or similar odor molecules (structurally neighboring molecules), such as molecules contained in a second or third database, without pre-calculating any affinity for the newly predicted molecule and / or fitting any molecule / receptor structure.
[0071] In the implementation, machine learning (or AI) can be used to predict and / or train based on the following input data: (“classical”) alphanumeric representations of chemical structures (e.g., H2O), numerical vectors, data containing atomic coordinates and chemical bonds (e.g., XML), Simplified Molecular Linear Input Specification (SMILES) strings, InChI or InChIKey, graphs (e.g., chemical structure diagrams containing atoms and interatomic chemical bonds), and / or pixel grids. In the implementation, such graphs can be generated procedurally, such as MolView, PubChem, or similar programs / databases, or, for example, as data containing atomic coordinates and interatomic chemical bonds (XML).
[0072] In the context of this invention, any machine learning-based model (AI) or algorithm is conceived and applicable, such as artificial intelligence (AI), machine learning (ML) algorithms or models, neural networks (NN), deep neural networks (DNN), convolutional neural networks (CNN), graph neural networks (GNN), random forest-based prediction models, recurrent neural networks, supervised learning, unsupervised learning, attention algorithms, statistical diffusion models, Transformers, LLM, GPT, and / or any combination thereof or any similar techniques. Machine learning-based models may include one or more layers. For example, a CNN may include two or more layers. Machine learning-based models may also involve one or more feature selection and / or filtering steps. Machine learning-based models may also involve adversarial learning. In embodiments, training the machine learning-based model may include reinforcement learning.
[0073] In specific implementations, the method or computer implementation of the present invention may include one or more of the following steps: An embodiment of the method of the present invention is or includes an advanced process comprising one or more computational models (e.g., ML models or AI) designed to simulate human olfaction using deep learning techniques. An embodiment of this model can be used to predict the olfactory signature spectrum of novel molecules using a master dataset of olfactory receptor (OR) structures and a sub-dataset of odor molecules.
[0074] In the implementation scheme, the model can be pre-trained using publicly available experimental receptor-ligand interaction data and / or structural data of olfactory receptors (ORs) and odorant molecules, and optionally “fine-tuned” to simulate the affinity of molecules for a range of (large) ORs, thereby creating a unique odor pattern.
[0075] In the implementation, the neural network architecture of the model (e.g., an ML model or AI) can be described by several key units or steps and specific training steps or stages.
[0076] In an implementation scheme, the method of the present invention first includes a general protein-ligand interaction pre-training step or phase on the adopted ML model or AI, which may be performed, for example, using a first database and a second database according to the present invention, and optionally also using a third database or embedding disclosed herein.
[0077] In an implementation, during the initial training of the corresponding ML model or AI, the method of the present invention includes training one or more specific encoders (encoder units), such as encoders for receptor proteins and / or (potential) ligands (such as odorant molecules).
[0078] For example, in an embodiment, a specific encoder, such as a receptor encoder, an odorant encoder, and / or a pairing interaction encoder, is generated in the following steps of the method of the present invention: step b: training a machine learning model or AI (receptor encoder) based on data contained in a first database; step d: training a machine learning (ML) model or artificial intelligence (AI) (odorant encoder) based on data contained in a second database; and / or step f: fitting the (3D) structure of one or more odor molecules to the (3D) structure of one or more olfactory receptors (ORs), wherein the fitting includes using and training an ML model or AI (pairing interaction encoder).
[0079] In the context of artificial intelligence and machine learning methods, encoders or encoding units typically (without limiting the scope of this invention) transform and / or embed input data in the initial stages of the generation process. In embodiments of the method of this invention, such encoders preferably embed data on olfactory receptor (OR) proteins and / or odorant molecules (e.g., small molecules) into a unified latent space, preferably derived from a first, second, and / or third database, which models their (3D) and / or chemical structures, physical properties, and / or chemical interactions, for example preferably from the first and second databases according to the invention and optional other data resources. Typically, the latent space is a (compressed) representation of the input data, where each dimension corresponds to a specific property or feature. Such pre-training steps preferably provide an efficient method for simultaneously characterizing ORs and odorant molecules, capturing key interaction features needed in subsequent steps of the method of this invention.
[0080] Figure 1 A non-limiting example of this process is disclosed, which illustrates the architecture and pre-training implementation of a general protein-ligand interaction model.
[0081] In one implementation, a specific (or specially trained or fine-tuned) odor receptor encoder is generated. For example, this can be achieved using a pre-trained receptor encoder and / or data contained in a first and / or second and / or third database or embedding, preferably containing detailed data on the odor receptor (e.g., its chemical structure, crystal structure, ligand (agonist) binding site, potential ligands, chemical interactions, and / or associated cell signaling pathways, etc.). In another implementation, a series (large number) of such specially trained receptor encoders are generated and trained such that, for example, one or more different, or even each, olfactory receptor (OR) (initially derived from the first database) is characterized by a “specific odor receptor encoder.” In yet another implementation, each encoder in this series is fine-tuned to efficiently predict the affinity of ligands for a specific olfactory receptor (OR), thereby enhancing the model’s predictive ability for odorant interactions.
[0082] In this implementation, a "odor encoder" layer training step is then performed. In this implementation, the constructed series of specific olfactory receptor encoders are preferably used in conjunction with a (comprehensive) database of odorant (and non-odorant) molecules (such as a second or third database or a third data embedding structure / layer) to train the final "odor encoder" layer. In this implementation, this layer generates unique odor embeddings for molecules by simulating the interaction between molecules and the trained olfactory receptor encoder (OR encoder), thereby effectively mapping molecules into a high-dimensional "odor space".
[0083] In implementations, training steps, such as step b (receptor encoder), step d (odorant encoder), and / or step e (pairing interaction encoder), include or constitute (pre)training and / or data collection or data embedding, which may aid in subsequent prediction and / or evaluation of odor molecules and their interactions (intensity / degree) with olfactory receptors and / or their similarity to other odor molecules considered to induce specific olfactory perceptions in subjects. In implementations, pre-training and / or data collection steps include collecting or constructing (comprehensive) datasets, databases, or data embeddings and / or learning specific relationships, characteristics, and dependencies of various types of data concerning one or more olfactory receptors (ORs) and / or odorant molecules and optionally their interactions, such as DNA sequences, crystallographic data, (3D) structural data, chemical structural data, chemical interaction and / or signaling pathway data, ligand binding sites, and / or predicted structures. In alternative implementations, a database, dataset, or embedding containing one or more of the aforementioned data concerning one or more olfactory receptors (ORs) and / or odorant molecules and optionally their interactions may be provided. In another alternative implementation, one or more pre-trained ML models or AIs trained on such databases, datasets, or embeddings may be provided.
[0084] In the implementation, the pre-training and / or data collection steps include one or more ML models or AIs learning from one or more (comprehensive) datasets or databases (e.g., a first database, a second database, and / or a third database) containing various types of data about one or more odor / odorant molecules and / or olfactory receptors and optionally their interactions, such as DNA sequences, crystallographic data, (3D) structural data, chemical structural data, chemical interaction and / or signaling pathway data, ligand binding sites, and / or predicted structures.
[0085] In one implementation, the pre-training and / or data collection steps include using a database of (preferably experimental) crystallographic (3D) information (e.g., publicly available data) about odor molecules, olfactory receptors, and / or ligand-receptor pairs / interactions. In another implementation, a first, second, and / or third database contains (preferably experimental) crystallographic 3D information (e.g., publicly available data) about odor molecules, olfactory receptors, and / or ligand-receptor pairs / interactions.
[0086] In the implementation scheme, the (3D) structure of the olfactory receptor and / or odor molecule, for example as disclosed in the first or second database respectively, is determined using experimental crystallographic information, computer simulations, artificial intelligence (AI) and / or machine learning (ML) models trained on experimental data, crystallographic information, gene sequence information of the corresponding olfactory receptor and / or odor molecule, or any combination thereof.
[0087] In the implementation, pre-training and / or data collection includes training the olfactory receptor encoder and / or the odor / odorant molecule (ligand) encoder, respectively, preferably in conjunction with the pairing interaction encoder layer (e.g., in the implementations of steps b, d, and f above).
[0088] In the implementation scheme, pre-training includes learning latent space embeddings and / or intermolecular relationships using known molecular binding (interaction) data.
[0089] In the implementation, the "paired interaction encoder" layer / unit employs a Transformer-based architecture, preferably equipped with a cross-attention mechanism to capture the complex relationships (chemical interactions) between OR atoms (e.g., atoms located at their ligand binding sites) and their odor / odorant molecule ligands. In the implementation, this preferably enables all three encoding units to efficiently update their representations through a multi-head self-attention layer (preferably further processed by a feedforward neural network, normalization, and dropout layer).
[0090] In the implementation, the second training step / stage includes constructing / generating one or more specific olfactory receptor encoders. The OR structure is preferably encoded using pre-trained receptor encoding units (encoders), thereby creating a series (large number) of preset embeddings. In the implementation, the series of preset embeddings may allow the receptor encoder to be omitted in subsequent training stages / steps and / or inference stages / steps, so that only the olfactory receptor encoder is used for the final prediction, such as the final prediction of the odor or structure of a new target olfactory molecule.
[0091] In the implementation, olfactory encoder training is performed. The ligand encoder and pairing interaction encoding units are "frozen" (reserved, fixed, or locked), while the downstream olfactory encoder layer is preferably trained on a diverse small molecule dataset (its size is preferably limited to those that can interact with ORs) by randomly masking, for example, up to 37% of the pairing encodings and predicting the olfactory encodings generated by the complete network. This adapts the network to ORs rather than the more generalized receptor. Through a self-distillation process, the model can learn from its high-confidence predictions, thereby enhancing model robustness.
[0092] In one implementation, an olfactory alignment step is performed, which includes aligning a (pre)trained encoder with a second and / or a third dataset or embedding containing odorant molecules with detailed odor descriptions. In another implementation, known odorant molecules are encoded by a (“frozen,” locked, reserved) pre-trained odorant encoder and paired with each pre-encoded OR embedding during the alignment process, and preferably subsequently analyzed via an olfactory encoding pipeline. In yet another implementation, an additional layer is also trained to predict olfactory descriptions based on the olfactory embeddings.
[0093] In the implementation, the final layer of the olfactory encoder is in a non-frozen state, thereby enabling rapid and efficient alignment with odor sensations induced in human subjects (e.g., perfumer perception).
[0094] In one implementation, the final prediction of the odor sensation induced by at least a first target odor molecule in a human subject includes an odorant (odor molecule) encoding step or stage. In another implementation, the new target odorant molecule is characterized by its structural features and embedded into a preferably previously generated high-dimensional vector space using a pre-trained (preferably frozen) odorant molecule (ligand) encoder (“odorant encoder”). In yet another implementation, for the final prediction of the target molecule's odor, each encoder and / or model is preferably “frozen” (locked, so that no further training is performed).
[0095] In this implementation, the odorant encoder consists of several independent transducer encoder layers. Preferably, the transducer encoder layers process molecular features and update their representations, preferably through a multi-head self-attention mechanism during the pre-training phase / step.
[0096] In the implementation, the transformer typically undergoes self-supervised learning, preferably including a prior (unsupervised) pre-training step.
[0097] In the implementation, the training steps / stages include a pairing interaction encoder step or stage. In another implementation, the characterization of the odorant molecule is combined with a pre-encoded OR characterization via a cross-attention mechanism. This feature (“interaction block”) preferably uses a triangular attention mechanism to enhance the learning of cross-pairing relationships, thereby optimizing the molecular characterization.
[0098] In the implementation, the training steps / stages include an affinity prediction step or stage. In the implementation, the learned interaction representations are projected onto the affinity matrix via linear transformations and activation functions. In the implementation, a hybrid density network (MDN) models the affinity distribution using learned statistical potential parameters, thereby aiding in the ranking of molecular affinities.
[0099] In embodiments, the method of the present invention performs an "inference" step or stage to ultimately predict the odor sensation induced by at least a first target odor molecule in human subjects. In embodiments, the inference step or stage includes odorant encoding and / or affinity simulation steps or stages. In embodiments, only the structural features of new molecules are processed using an odorant encoder (layer). In embodiments, the odorant encoder block preferably updates the molecular characterization through independent transformer layers. In embodiments, the embedding of a new target odorant molecule (e.g., a small molecule) is paired with some or all (pre-computed) individual OR embeddings via a paired interaction encoder. In embodiments, this step is followed by processing with a feedforward neural network, normalization, and dropout layers to efficiently create odor spatial embeddings.
[0100] In this implementation, the prediction maps (new) target odor molecules to a high-dimensional “odor space” (“odor space” can be considered as a specific implementation of, for example, chemical space), where each molecule is associated, for example, based solely on its simulated olfactory effect. This preferably enables the identification of (structurally diverse) (target) odor molecules that elicit similar olfactory responses or effects in subjects.
[0101] In one implementation, using an "odor space" embedding, molecules can be analyzed and compared based on their predicted olfactory feature spectra using known (state-of-the-art) algorithms. In another implementation, this embedding can also be transferred to an additional network layer to align the latent space with a second database, thereby pairing perfumer descriptions with the learned olfactory embeddings.
[0102] Figure 2 A non-restrictive example of this workflow has been published.
[0103] In the implementation plan, a specific training method may be used, which includes one or more of the following steps.
[0104] In the implementation scheme, the final layer of the olfactory encoder is in a non-frozen state, thereby achieving rapid and efficient alignment with the perfumer's perception (olfactory perception).
[0105] The inventors unexpectedly discovered that the method of this invention provides excellent detection accuracy. In exemplary technical experiments, the initial parts of the system (up to all units of the olfactory encoder) were validated against known benchmarks and achieved industry-leading performance in docking tasks. The method of this invention exhibits a low error range and high accuracy in predicting olfactory descriptors, with a success rate exceeding 95% based on the validation dataset.
[0106] In the implementation scheme, the method of the present invention can be applied to a variety of task scenarios, including but not limited to: molecular virtual screening and all embodiments described herein, namely odor “telephone”, VR system odor synthesizer, olfactory pleasure sensor system, perfume formulation reconstruction, natural product approximation simulation, and computer simulation process for the generation and development of novel fragrance and flavor components.
[0107] In summary, the method of this invention utilizes industry-leading deep learning technology and extensive pre-training to accurately simulate human olfaction. By encoding molecules into a high-dimensional odor space, it provides a powerful tool for fragrance design, odor detection, and other olfactory applications, and can be applied, for example, to any and all of the application examples detailed herein.
[0108] In some implementations, the method for predicting molecular odor includes the following steps: a. Provide a first database containing structures of one or more olfactory receptors (ORs) and / or optionally one or more receptors involved in odorant perception (olfactory perception) and / or auxiliary (or transport) factors involved in odorant perception, preferably three-dimensional (3D) structures and / or chemical structures; b. Determine the (3D) structure of ligand binding sites and / or accessory (or transport) factors involved in odorant perception of one or more olfactory receptors and / or optionally odorant perception (olfactory perception) contained in the first database; c. Provide a second database containing (i) the (3D) structure of one or more odor molecules (odorants) and (ii) corresponding textual descriptions of the odor sensations (olfactory perceptions) induced by said one or more odor molecules in (human) subjects; d. Fit the (3D) structure of one or more odor molecules to the (3D) structure of the ligand binding site of one or more (olfactory) receptors (determined by step b), and / or optionally to the (3D) structure of one or more receptors involved in odorant perception and / or auxiliary (or transport) factors involved in odorant perception, thereby determining the affinity score of each fitted combination of the one or more odor molecules with the one or more olfactory receptors and / or optionally with the one or more receptors involved in odorant perception and / or auxiliary (or transport) factors involved in odorant perception. The (high) affinity score indicates the (high) fit of one or more odor molecules to the ligand binding site of the corresponding (olfactory) receptor (or cofactor), and preferably indicates the ability of one or more odor molecules to bind to or interact with the (olfactory) receptor (or cofactor) and / or activate the olfactory receptor; and e. Optionally, a third database is created, which contains data from the first database, data from the second database, and affinity scores determined for each fitted combination of one or more odor molecules with the corresponding (olfactory) receptors (or cofactors).
[0109] In some embodiments, the present invention relates to a computer-based method for predicting molecular odors, comprising the following steps: a. Provide a first database containing the three-dimensional (3D) (chemical) structures of one or more olfactory receptors (ORs) and / or optionally one or more receptors involved in odorant perception (olfactory perception) and / or auxiliary (or transport) factors involved in odorant perception; b. Determine the (3D) structure of agonist (or ligand) binding sites of one or more olfactory receptors and / or optional odorant sensing (olfactory perception) receptors contained in the first database and / or auxiliary (or transport) factors involved in odorant sensing; c. Provide a second database containing (i) the (3D) structure of one or more odor molecules and (ii) corresponding textual descriptions of the odor sensations induced by said one or more odor molecules in human subjects; d. Fit the (3D) structure of one or more odor molecules to the agonist (or ligand) binding sites of one or more olfactory receptors contained in a first database and / or optionally to the (3D) structures of one or more receptors and / or accessory (or transporter) factors involved in odor perception (olfactory perception), thereby determining the affinity score of each fitted combination of the one or more odor molecules with the corresponding (olfactory) receptor / factor. The (high) affinity score indicates the (high) fit of one or more odor molecules to the agonist (or ligand) binding site of the corresponding olfactory receptor, and preferably indicates the ability of the one or more odor molecules to bind to or interact with the (olfactory) receptor (or cofactor) and / or activate the olfactory receptor. e. Optionally, a third database is created, which contains data from the first database, data from the second database, and affinity scores determined for each fitted combination of one or more odor molecules with the corresponding olfactory receptors (or cofactors); f. Providing the (3D) structure of at least one first target odor molecule not included in the second database; and g. Predicting the odor sensation induced by the target odor molecule in human subjects, including: (I) Fitting the (3D) structure of at least one first target odor molecule to the (3D) structure of the binding site of one or more olfactory receptors (and / or cofactors) (for which an affinity score has been determined in step d), thereby determining the affinity score of each fitted combination of the at least one first target odor molecule with the corresponding (olfactory) receptor (or cofactor). (II) Compare one or more of the affinity scores measured for at least one first target odor molecule in step g with the affinity scores of one or more odor molecules contained in the third database and / or measured in step d, thereby selecting from the second or third database one or more odor molecules with the most similar pattern (highest similarity) affinity score or the most similar affinity score pattern to one or more (olfactory) receptors (or cofactors) measured in step g. Based on one or more affinity scores assigned to the corresponding olfactory receptor and odor molecule combination in step d, and textual descriptions of the odor sensation induced by the corresponding odor molecule in human subjects in the second and / or third databases, the odor sensation (olfactory perception / odor) induced by the target odor molecule in human subjects is predicted. This prediction includes the use of statistical methods, and / or the use and / or training of machine learning algorithms (or models).
[0110] In the implementation scheme, the computer-based method for predicting molecular odor includes the following steps: a. Provide a first database containing the three-dimensional (3D) (chemical) structures of one or more olfactory receptors (ORs) and / or optionally one or more receptors involved in odorant perception (olfactory perception) and / or auxiliary (or transport) factors involved in odorant perception; b. Determine the (3D) structure of agonist (or ligand) binding sites of one or more olfactory receptors and / or optional odorant sensing (olfactory perception) receptors contained in the first database and / or auxiliary (or transport) factors involved in odorant sensing; c. Provide a second database containing (i) the (3D) structure of one or more odor molecules and (ii) corresponding textual descriptions of the odor sensations induced by said one or more odor molecules in human subjects; d. Fit the (3D) structure of one or more odor molecules to the agonist (or ligand) binding sites of one or more olfactory receptors contained in a first database and / or optionally to the (3D) structures of one or more receptors and / or accessory (or transporter) factors involved in odor perception (olfactory perception), thereby determining the affinity score of each fitted combination of the one or more odor molecules with the corresponding (olfactory) receptor / factor. The (high) affinity score indicates the (high) fit of one or more odor molecules to the agonist (or ligand) binding site of the corresponding olfactory receptor, and preferably indicates the ability of the one or more odor molecules to bind to or interact with the (olfactory) receptor (or cofactor) and / or activate the olfactory receptor. e. Create a third database containing data from the first database, data from the second database, and affinity scores determined for each fitted combination of one or more odor molecules with the corresponding olfactory receptors (or cofactors); f. Train machine learning models or AI based on data contained in a third database; g. Providing the (3D) structure of at least one first target odor molecule not included in the second database; and h. Using the machine learning model or AI trained in step f to predict the odor sensation induced by the target odor molecule in human subjects, including: predicting one or more similar odor molecules contained in a third database (which preferably have an affinity pattern similar to olfactory receptors and / or cofactors), and / or predicting the odor sensation (olfactory perception / odor) induced by the target odor molecule in human subjects, wherein the odor sensation (olfactory perception / odor) preferably corresponds to one or more textual descriptions of human odor sensations contained in the third database.
[0111] In another aspect or embodiment, the present invention relates to a computer-based method for predicting the odor of a mixture of odor molecules.
[0112] In an implementation, for example, in step g, in addition to providing (i) the (3D) structure of at least the first and second target odor molecules, (ii) the relative percentages of at least the first and second target odor molecules in the vapor of the mixture and / or (iii) the statistical probability of each odor molecule interacting with olfactory receptors in the nasal cavity of a human subject, and wherein predicting odor sensation (e.g., step h) includes predicting the odor sensation induced in a human subject by the mixture of at least the first and second target odor molecules.
[0113] Implementations of the method of this invention may employ, include, or constitute an olfactory pleasure sensor system. Generally, the hedonic olfactory model refers to a framework for classifying odors based on the pleasantness or unpleasantness of odor perception, reflecting the inherent biological responses of humans to different scents. This model utilizes the natural tendency of the human olfactory system to distinguish odors generally considered "good" or "bad," which often correspond to broader biological and evolutionary cues. The olfactory pleasure system plays a crucial role in guiding human behavior by labeling expected and unexpected things. This includes distinguishing between health and disease, cleanliness and filth, and nutrition and spoilage.
[0114] In the context of this invention, the inventors have developed several versions of human olfactory models, as described herein and in embodiments (e.g., NN and classical molecular docking models). Thereby, the inventors have discovered that embedding model knowledge into human biology through the use of olfactory receptors yields highly desirable results: automatically aligning the model's mathematical odor space with human olfactory hedonism. This inherent gradient can be detected in high-dimensional mathematical olfactory spaces using known statistical dimensionality reduction techniques or other clustering algorithms. In other words, without any additional alignment (e.g., using human-based labels or perfumer descriptions), the model possesses a strong built-in tendency to distinguish between "good" and "bad" odors.
[0115] In this implementation, a dedicated physical sensor hedonic system can leverage the principles of hedonic olfaction models to create advanced sensors for various industrial and medical applications. These sensors can be designed to detect and evaluate odors based on the pleasantness or unpleasantness of the perceived odor, without needing to identify the exact compound causing the odor. This method simplifies the detection process while ensuring robust and reliable odor assessment.
[0116] Possible applications of such models include, but are not limited to, the pharmaceutical industry (disease detection), such as early detection of diseases through breath or body odor analysis. For example, a sensor is designed to identify "unpleasant" odors in human breath known to be associated with various disease states. The advantages of such methods are non-invasive early diagnosis and continuous monitoring.
[0117] The document also envisions food storage quality management, such as monitoring the freshness and safety of stored foods. For example, a sensor could detect signs of spoilage in food storage facilities by identifying unpleasant odors associated with bacterial growth or chemical degradation. The advantages of such methods could include ensuring food safety, reducing waste, and improving inventory management.
[0118] Further applications of this type of model could include monitoring the cleanliness of public restrooms, such as maintaining the cleanliness and hygiene of public restrooms. For example, a sensor system could continuously monitor unpleasant odors in the air of public restrooms and automatically alert maintenance personnel when cleaning is required. The advantages of this approach are: improved user experience, ensuring high hygiene standards, and reduced maintenance costs.
[0119] Implementations of the method of this invention can also be applied to cell culture control, such as monitoring the quality and state of cell cultures in research and production environments. For example, a sensor can detect changes in the odor characteristics of cell cultures, indicating contamination or suboptimal growth conditions. The advantages of such methods are: ensuring culture viability, improving research results, and preventing contamination.
[0120] The advantage of the pleasure sensor system according to embodiments of the present invention lies in its simplicity and robustness, because such systems are based on detecting overall odor characteristics rather than specific molecular components, making them simpler and more reliable. For example, instead of identifying specific spoilage compounds in food, the sensor detects the overall unpleasant odor associated with spoilage.
[0121] Furthermore, such systems offer broad detection capabilities, as they can identify negative conditions (unpleasant odors) across a wide range of applications without requiring detailed knowledge of the exact molecules involved. For example, the sensor could alert to an overall unpleasant odor in a public restroom, indicating poor cleanliness regardless of the specific cause. Another advantage lies in the cost-effectiveness of such systems: simplified design and broad detection capabilities reduce production and operating costs. For instance, a single sensor type can be used across multiple applications, reducing the need for specialized equipment. Moreover, real-time monitoring is possible: such systems / models provide continuous, real-time monitoring and immediate alerts, enabling rapid action on detected problems. For example, a restroom maintenance alert could be issued immediately upon detecting an unpleasant odor.
[0122] Ultimately, this system achieves high versatility because the sensor system can be customized for specific use cases in various industries such as pharmaceuticals, food storage, public health, and biotechnology. For example, sensitivity settings can be customized for different environments to ensure accurate detection. In summary, the dedicated physical sensor hedonic system demonstrates the practical utility of the hedonic olfactory model according to embodiments of the present invention in various industrial and medical scenarios. By focusing on detecting overall odor characteristics associated with pleasant or unpleasant odors, these sensors provide a simple, reliable, and cost-effective solution for quality, safety, and hygiene maintenance in a variety of applications. The ability to monitor in real time and provide immediate alerts further enhances its value, making it an indispensable tool in modern industry and medicine.
[0123] On the other hand, the present invention relates to a generative model for discovering novel olfactory or aromatic molecules and mixtures thereof. Using the methods described herein (e.g., including machine learning (ML) models) along with any machine learning model capable of generating three-dimensional molecular structures, novel odor molecules, unlike any previously unseen, can be generated from odor descriptions. Therefore, in embodiments, the present invention relates to a computer-implemented method comprising training at least one machine learning model or AI to predict the three-dimensional (chemical) structure of a molecule capable of inducing a desired olfactory experience or perception (odor) in a human subject by binding to binding sites of one or more olfactory receptors or other receptors involved in odor perception. In embodiments, the method may include the steps of: predicting or generating the three-dimensional (chemical) structure of a molecule based on machine learning, predicting binding (e.g., strength or fit) to binding sites of one or more olfactory receptors or other receptors involved in odor perception, and / or subsequently predicting the odor sensation (olfactory perception / odor) induced by the predicted molecule in a human subject, such as odor sensation (olfactory perception) data induced in a human subject based on a specific three-dimensional (chemical) molecular structure, such as data disclosed / included in a second database.
[0124] In an embodiment, the present invention relates to a computer-implemented method for predicting the (3D) structure of a target odor molecule capable of inducing a target odor sensation in human subjects, the method being based on the binding affinity of the target odor molecule to at least one olfactory receptor, and comprising the following steps: a. Provide a first database containing the structures (preferably three-dimensional (3D) structures and / or chemical structures) of one or more olfactory receptors (ORs), preferably wherein the database contains data on the 3D structures of ligand (or agonist) binding sites of one or more olfactory receptors contained in the first database; b. Train a machine learning model or AI (receptor encoder) based on the data contained in the first database; c. Provide a second database containing (i) the (3D) structure of one or more odorant molecules and (ii) corresponding textual descriptions of the odor sensations induced by said one or more odorant molecules in human subjects; d. Train a machine learning (ML) model or artificial intelligence (AI) (odorant encoder) based on the data contained in the second database; d.1 Optionally provide interaction data between one or more olfactory receptor molecules from a first database and one or more odorant molecules from a second database; e. Fitting the (3D) structure of one or more odor molecules to the (3D) structure of one or more olfactory receptors (ORs), preferably to the ligand binding sites of the one or more ORs, thereby determining an affinity score for each fitted combination of the one or more odor molecules and the corresponding olfactory receptors. The fitting process includes using and training an ML model or AI (paired interaction encoder), optionally, wherein the use and / or training of the ML model or AI also takes into account the interaction data provided in step d.1. The affinity score (a numerical value) indicates the degree of fit between the corresponding odor molecule and the olfactory receptor (OR), preferably the degree of fit with the ligand binding site of the corresponding OR, and preferably indicates the ability of the corresponding odor molecule (i) to bind or interact with the corresponding olfactory receptor and / or (ii) to activate the corresponding olfactory receptor. f. Create a comprehensive third database or third data embedding structure / layer that contains data from the first database, data from the second database, and affinity scores determined for each fitted combination of the corresponding odor molecule and the corresponding olfactory receptor; And further includes the following steps: G. Select a text description of the target odor sensation induced by the first odor molecule from a third database or a third data embedding structure / layer; H. Determining one or more olfactory receptors with the highest affinity score for the first odor molecule from a third database or a third data embedding structure / layer; and I. Based on data from a third database or a third data embedding structure / layer, predict the (3D) structure of target odor molecules not included in the third database. The predicted (3D) structure is modeled / predicted and fitted with a high affinity score to the (3D) structure of the ligand binding site of one or more olfactory receptors selected in step H. The high affinity score of the predicted (3D) structure of the target odor molecule indicates the ability of the target odor molecule to bind or interact with the ligand binding site of the corresponding one or more olfactory receptors, and / or The ability to activate one or more olfactory receptors This induces an odor sensation similar to that of the first odor molecule in human subjects.
[0125] In an implementation of the method, the prediction method used in step H is implemented by an AI and / or machine learning (ML) model, preferably a gradient-free ML model, which preferably includes a genetic algorithm for generating molecules and / or optimizing the prediction toward a given numerical vector, or It is implemented by a gradient-based ML model, which preferably includes a generative adversarial network (GAN) or a diffusion probability model.
[0126] In an implementation of the method, the AI and / or ML model is or includes one or more of the machine learning models or AIs trained in steps b, d, and f.
[0127] In an embodiment of the method, the predicted (3D) structure of the target odor molecule is used to chemically synthesize the predicted molecular structure.
[0128] In some other embodiments, the present invention relates to a computer-based method for predicting molecular odors, comprising the following steps: a. Provide a first database containing structures of one or more olfactory receptors (ORs) and / or optionally one or more receptors involved in odorant perception (olfactory perception) and / or auxiliary (or transport) factors involved in odorant perception, preferably three-dimensional (3D) structures and / or chemical structures; b. Determine the (3D) structure of agonist (or ligand) binding sites of one or more olfactory receptors and / or optional odorant sensing (olfactory perception) receptors contained in the first database and / or auxiliary (or transport) factors involved in odorant sensing; c. Provide a second database containing (i) the (3D) structure of one or more odor molecules (odorants) and (ii) corresponding textual descriptions of the odor sensations (olfactory perceptions) induced by said one or more odor molecules in (human) subjects; d. Fitting the (3D) structure of one or more odor molecules to the agonist (or ligand) binding sites of one or more (olfactory) receptors and / or optionally to the (3D) structure of one or more receptors and / or accessory (or transporter) factors involved in odor perception (olfactory perception), thereby determining the affinity score of each fitted combination of the one or more odor molecules with the corresponding olfactory receptor / factor. The affinity score preferably indicates (correlated / proportional) the fit of one or more odor molecules to the agonist (or ligand) binding site of the corresponding olfactory receptor, and preferably indicates the ability of the one or more odor molecules to bind to or interact with the olfactory receptor and / or activate the olfactory receptor. e. Optionally, a third database is created, which contains data from the first database, data from the second database, and affinity scores determined for each fitted combination of one or more odor molecules with their respective olfactory receptors; F. Select odor sensations induced by the second target odor molecule from the third database; G. Determine one or more (olfactory) receptors with the highest affinity scores for the second target odor molecule from a third database; and H. Based on data from a third database, predict the (3D) structure of a third (new) target odor molecule, wherein the predicted (3D) structure is modeled / predicted to fit the (3D) structure of the agonist (or ligand) binding site of the olfactory receptor selected in step G with a (high) affinity score (preferably equivalent to or higher than the affinity score of the second target odor molecule to the corresponding receptor), wherein the affinity score of the predicted (3D) structure of the third target odor molecule preferably indicates the ability of the third (new) target odor molecule to bind to or interact with and / or activate the olfactory receptor, and wherein the third target odor molecule is not included in the third database.
[0129] Each feature disclosed in this document (above and below) within a particular (computer implementation) method context is to be regarded as applicable to any other (computer implementation) method disclosed herein (above and below) with the same degree of disclosure.
[0130] In the implementation scheme, the prediction method used in step H is a machine learning (ML) method, preferably: i. A gradient-free ML method, preferably comprising a genetic algorithm for generating molecules and / or optimizing predictions toward a given numerical vector; or ii. Gradient-based ML methods, preferably generative adversarial networks or diffusion probability models.
[0131] In the implementation scheme, the predicted (3D) structure of the third target odor molecule is used to chemically synthesize the predicted molecule structure.
[0132] For a given odor description, the affinity list can be calculated / predicted by reversing the process described in the first aspect of the invention.
[0133] In implementation schemes, the methods of the present invention may include the use of simple statistical tools, artificial intelligence, machine learning (ML) algorithms or models, including neural networks (NN), deep neural networks (DNN), convolutional neural networks (CNN), graph neural networks (GNN), Transformers, LLM, GPT, random forest-based prediction models, recurrent neural networks, supervised learning, unsupervised learning, attention algorithms, statistical diffusion models, and / or any combination thereof or similar techniques. Machine learning-based models may include one or more layers. For example, a CNN may include two or more layers. Machine learning-based models may also include a feature selection step.
[0134] A complete or partial list of new affinities, or any other characterization constructed using statistical or ML tools (i.e., different embeddings of affinity and other data into numerical vectors), can then be used to predict new molecules.
[0135] This prediction can be made using a gradient-free method, such as using a genetic algorithm to generate molecules and optimize them against a given numerical vector. Alternatively, it can be made using any other machine learning method, such as gradient-based or gradient-free methods. Non-limiting examples include GANs, diffusion probability models, autoencoders, etc.
[0136] Use any form of characterization of the odorant molecule, such as SMILES (Sharma et al., 2021), molecular fingerprints, or three-dimensional (3D) models such as .pdb files. These molecules can then be synthesized in an organic chemistry lab and used in any commercial product due to their odor properties.
[0137] In the implementation scheme, as an alternative to or supplement to the first database, public databases such as the Olfactory Receptor Database (ORDB; Yale University) and / or the PubChem database may be used.
[0138] In the implementation plan, as an alternative to or supplement to the second database, public databases such as the AromaDb database (Central Institute of Medicinal and Aromatic Plants, Indian Council of Scientific and Industrial Research (CSIR)), the SuperScent database (Charlit Hospital, Germany), and / or the PubChem database may be used.
[0139] In another aspect, the present invention relates to an apparatus, device, or system. In one embodiment, the present invention relates to a data processing (computing) apparatus, device, or system comprising components for performing the steps of the method of the present invention. In another embodiment, the present invention relates to a data processing (computing) apparatus, device, or system comprising a processor adapted to / configured to perform the method of the present invention.
[0140] In another aspect, the present invention relates to a computer program comprising instructions which, when executed by a computer, cause the computer to perform the steps of the method of the present invention.
[0141] In another aspect, the present invention relates to a computer-readable (storage) medium or data carrier comprising instructions which, when executed by a computer, cause the computer to perform the steps of the method of the present invention.
[0142] The inventors believe that the implementation schemes of the method of the present invention are applicable to a variety of task scenarios, including but not limited to: molecular virtual screening and all embodiments described below, namely, odor “telephone”, VR system odor synthesizer, olfactory pleasure sensor system, perfume formulation reconstruction, natural product approximation simulation, and computer simulation process for the generation and development of novel fragrance and flavor components.
[0143] The implementation of the method of the present invention can be applied to a variety of task scenarios, including but not limited to: molecular virtual screening and all embodiments described below, namely, odor “telephone”, VR system odor synthesizer, olfactory pleasure sensor system, perfume formulation reconstruction, natural product approximation simulation, and computer simulation process for the generation and development of novel fragrance and flavor components.
[0144] The embodiments and features of the present invention described in relation to odor (odor perception) prediction methods and (3D) structure prediction methods for target odor molecules capable of inducing target odor perception in human subjects should be considered applicable to every other aspect of this disclosure, such that features characterizing one method can be used to characterize another, and vice versa. The various aspects of the present invention are unified, benefit from, based on, and / or interconnected by the common and unexpected discovery of predicting molecular odor based on molecular structure (and vice versa), preferably using computational methods such as ML and / or AI for prediction.
[0145] All cited patent and non-patent documents are incorporated herein by reference in their entirety.
[0146] In the context of this invention, "subject" is preferably a human subject, but in embodiments it may also be any other subject, such as a mammal, animal or other organism.
[0147] In the context of embodiments of the present invention, the terms “odor”, “smell”, “olfactory perception”, “olfactory sensation”, “odor sensation” or “odor sensation induced in human subjects” are used interchangeably.
[0148] In the context of the embodiments of this invention, the terms "odor molecule" and "odorant molecule", as well as "olfaction", "odor" and "odorant" are used interchangeably.
[0149] In the context of the embodiments of this invention, the terms "odor receptor", "odorant receptor" and "olfactory receptor" are used interchangeably.
[0150] In the context of embodiments of the present invention, "odor molecule" refers to any odorant or molecule that has the ability to induce olfactory sensation in a subject, such as hydrophobic, hydrophilic and / or volatile small molecules, preferably by binding to a corresponding receptor, such as an olfactory receptor (OR) or any other receptor involved in odor perception.
[0151] Olfactory receptors can be obtained from the database https: / / senselab.med.yale.edu / ORDB / .
[0152] In the embodiments described herein, when referring to olfactory receptors, it may also refer to other receptors involved in odorant perception (olfactory perception) and / or auxiliary (or transport) factors involved in odorant perception. In other words, in the embodiments, the term "olfactory receptor (OR)" also includes one or more (other) receptors ("other" means receptors that are "usually" not called olfactory receptors) and / or their auxiliary factors involved in odorant perception.
[0153] Olfactory receptors (ORs), also known as odorant receptors, are chemical receptors expressed in the cell membranes of olfactory receptor neurons. They are involved in detecting odorous compounds (called odorants), thereby triggering the sense of smell. Activated olfactory receptors initiate nerve impulses, transmitting information about odors to the brain. These receptors belong to class A of the rhodopsin-like family of G protein-coupled receptors (GPCRs). Based on their structure and location, olfactory receptors are divided into several receptor families, including odorant receptors (ORs), formyl peptide receptors (FPRs), guanylate cyclase GC-D, vomeronasal receptors (V1R and V2R), and trace amine-associated receptors (TAARs).
[0154] OR can be distinguished from other GPCRs by several conserved amino acid motifs; these motifs include the LHTPMY motif in the first intracellular loop, the most characteristic MAYDRYVAIC motif at the end of the transmembrane (TM) domain 3 (TM3), the very short SY motif at the end of TM5, the FSTCSSH fragment at the beginning of TM6, and PMLNPF in TM7 (Fleischer et al., 2009).
[0155] In the embodiments described herein, any receptor involved in odor perception, i.e., a receptor to which odor molecules (odorant molecules or odorants) can bind and induce olfactory sensation in a subject, may be referred to as an "olfactory receptor." In the embodiments described herein, the first database may also contain the structures of any receptors involved in odor perception, i.e., the structures of receptors to which odor molecules (odorants) can bind and induce olfactory sensation in a subject. In this document, the terms "odor molecule," "odorant molecule," or "odorant" are used interchangeably.
[0156] In the embodiments described herein, the first database may also include molecules involved in odor perception (“helper molecules” or “cofactors”), i.e., molecules that can bind to and transport odor molecules (odorants) to corresponding receptors, thereby inducing olfactory sensation in the subject when the odorant binds to the receptor. Such “helper molecules” / “cofactors” may include, but are not limited to, odorant-binding protein molecules such as OBP2A, OBP2B, and lipocalcinin 1 (LCN1).
[0157] Odor-binding proteins (OBPs) are typically small (between 10 and 30 kDa) soluble proteins that are highly concentrated in the nasal mucus of vertebrates, such as mammals or, preferably, humans. Because of their affinity for odor molecules and pheromones, OBPs are thought to play a role in olfactory perception.
[0158] A universal nomenclature system has been developed for the olfactory receptor family, which forms the basis for the symbols assigned to the genes encoding these receptors by the Human Genome Organization (HUGO). The naming format for each member of the olfactory receptor family is "ORnXm", where "OR" is the family identifier (olfactory receptor superfamily); "n" is an integer representing the family (e.g., 1-56), whose members share more than 40% sequence identity; and "X" is a letter representing the subfamily (A, B, C...) (typically with more than 60% DNA sequence identity). "M" is an integer representing a single family member (subtype). For example, OR1A1 represents the first subtype in subfamily A of olfactory receptor family 1.
[0159] Members of the same olfactory receptor subfamily (which typically share more than 60% DNA sequence identity) may recognize structurally similar odorant molecules.
[0160] Two main classes of human olfactory receptors have been identified: Class I (fish-like receptors) OR family 51-56 and Class II (tetrapod-specific receptors) OR family 1-13. Class I receptors are thought to be specifically used to detect hydrophilic odorants, while Class II receptors detect more hydrophobic compounds.
[0161] VN1R1 is a putative human pheromone receptor expressed in the human olfactory mucosa.
[0162] In this document, the terms “olfactory experience,” “odor sensation,” or “olfactory perception” are preferably used interchangeably.
[0163] In the context of this invention, odor molecules or odorants are preferably hydrophobic and / or volatile small molecules that can induce olfactory (odor) sensation or perception in subjects (preferably human subjects).
[0164] In the embodiments described herein, the subjects are preferably mammalian subjects, such as animals or humans, and more preferably human subjects. In the embodiments, the method of the present invention can also be used for subjects other than humans, such as animals (e.g., rodents, dogs, or cats).
[0165] In embodiments, the term "fitting," such as fitting the (3D) structure of one or more odor molecules to the (3D) structure of one or more olfactory receptors (ORs), preferably to the ligand binding sites of the one or more ORs, may refer to or include (computational / computer simulation) modeling and / or computation of chemical parameters, physical forces, and / or other interaction parameters that indicate specific (chemical / physical / atomic attraction / attraction interaction) interactions between two molecules (e.g., ligand and receptor, preferably with the active site of the receptor and / or the ligand binding site). In this context, the term "fitting combination" in embodiments may refer to the pairing / combination of a corresponding odor molecule with a corresponding olfactory receptor after such fitting, docking, modeling, and / or corresponding computation (e.g., in step e of the method of the present invention). In some embodiments, fitting may be performed by suitable software (such as AutoDock Vina or PyMOL) and / or by machine learning models or AI, for example, based on (training / learning) certain structural and / or interaction data of ligand and receptor molecules (e.g., from a training dataset / database).
[0166] Machine learning can generally be categorized into supervised learning, unsupervised learning, and reinforcement learning.
[0167] Machine learning generally refers to algorithms that use training data to build computational models that can make predictions or decisions independently (mainly without external instructions / commands).
[0168] In embodiments, the methods of the present invention may include the use of simple statistical tools, artificial intelligence (AI) and / or machine learning (ML) algorithms or models, preferably including neural networks (NN), deep neural networks (DNN), convolutional neural networks (CNN), artificial neural networks (ANN), graph neural networks (GNN), Transformer, LLM, GPT, random forest-based prediction models, recurrent neural networks, supervised learning, unsupervised learning, attention algorithms, and statistical diffusion models and / or any combination thereof or similar techniques. In the embodiments described herein, the terms "artificial intelligence (AI)" and "machine learning (ML) model" are used interchangeably.
[0169] Random (decision) forest is an ensemble learning method (using multiple learning algorithms to improve prediction performance) suitable for tasks such as regression or classification, which builds multiple decision trees during training (supervised learning).
[0170] Deep learning is a type of machine learning based on neural networks with multiple layers, such as deep neural networks (DNN), convolutional neural networks, recurrent neural networks, deep belief networks, deep reinforcement learning, and transformers. The learning methods can be supervised, unsupervised, or semi-supervised.
[0171] Artificial neural networks (ANNs) are often described as neural networks inspired by the biological neural networks of the brain.
[0172] Generative Adversarial Networks (GANs) are machine learning frameworks in which two neural networks compete against each other in a zero-sum game, where the gain of one network is the loss of the other. GANs are capable of generating new data with the same statistical characteristics as the initially provided training dataset. GANs are preferably able to learn in an unsupervised manner.
[0173] Typically, a Convolutional Neural Network (CNN) consists of interconnected layers, including, for example, an input layer (which receives input data), no hidden layers, one or more hidden layers, and an output layer (which produces the final result / output). In a CNN, hidden layers comprise one or more layers that perform convolutions.
[0174] Graph Neural Networks (GNNs) are a type of artificial neural network that can be represented as a graph (a data structure containing interconnected nodes and edges).
[0175] Generally, embedding, or the “embedding process,” is the process by which ML models or AI transform “real-world” data into machine-understandable or processable data. Embedding typically refers to the process of converting “real-world” (input) data into numerical representations, where the input data is transformed into complex mathematical representations that reflect the relationships and inherent properties between the data points contained in the “real-world” input data. In the embodiments described herein, input data (e.g., derived from a first, second, and / or third database) is embedded into machine-readable, understandable, and processable data structures for further use by the AI or ML model. For example, in embodiments described herein, data generated by the method of the present invention may be stored in a third database or embedded with numerical representations that can be read and processed by the AI or ML model, such as a third data embedding structure / layer.
[0176] Diffusion probability models are developed based on the principles of nonequilibrium thermodynamics. Therefore, these models create a series of "diffusion steps" (Markov chains = stochastic models describing a sequence of possible events, where the probability of each event depends only on the state reached by the previous event), first adding random noise to the initial dataset, and then learning how to denoise the data again. Diffusion models typically consist of three main components: a forward process, a reverse process, and a sampling procedure.
[0177] Feature Subset Selection (FSS) is a machine learning method that uses only a subset of available features for machine learning. FSS is necessary in part because it is technically impossible to include all features, or because of the discriminative problems that arise when there are many features but only a small dataset, or to avoid model overfitting (see the bias-variance dilemma).
[0178] Feature selection is typically a preferred step in machine learning. It refers to the process of selecting a subset of relevant features (variables or predictors) for model building. This technique is often used to simplify models that are easier for researchers / users to interpret, improve data compatibility with different learning model classes, shorten training time, and avoid dimensionality defects. Feature selection can also be used to encode inherent symmetries in the input space. Feature selection can also be called variable selection, attribute selection, or variable subset selection. Feature selection can be used to ignore irrelevant or redundant data. Another different approach is feature extraction. Feature extraction creates new features as a function of the original features, while feature selection returns a subset of features. Feature selection techniques can be applied to datasets containing multiple features and with relatively small amounts of data or samples. Feature selection algorithms can be viewed as a combination of various search techniques that search for new feature subsets by applying evaluation metrics that score different sets of features.
[0179] Typically, the dissociation constant KD is negatively correlated with affinity. The KD value may be related to molecular concentration (number of molecules), for example, the molecular concentration (number of molecules) required to bind to another molecule. Therefore, the lower the KD value (lower concentration), the higher the affinity of the molecule for, for example, a receptor binding site. Attached Figure Description
[0180] The invention will now be further described with reference to the accompanying drawings. These drawings are not intended to limit the scope of the invention, but rather to represent preferred embodiments of various aspects of the invention in order to more fully illustrate the invention described herein.
[0181] Figure 1 Architecture and pre-training of a general protein-ligand interaction model.
[0182] Figure 2 Embedding small molecules into a high-dimensional olfactory space.
[0183] Figure 3 AutoDock Vina software was used to dock and score all OR-odorant molecule pairs.
[0184] Figure 4 A translation model trained using an encoder model and dataset 2, capable of bidirectional conversion between olfactory embeddings and odor descriptor types provided in dataset 2. For example, natural language translation.
[0185] Figure 5 The process shown begins with the perfumer's specific requirements for fragrance molecules (1). The perfumer creates a fragrance description (2). This description is converted into an embedding (4) using a translation model (3). This olfactory embedding is used to guide the molecule generation algorithm (5). The generation model returns a list of molecule options and their olfactory embeddings (6). The olfactory embeddings are converted into a human-understandable description with the help of a translation model (not shown in the figure) and returned to the perfumer. Based on the provided description, the perfumer decides (7) whether to send certain molecules to the laboratory for synthesis (8), or to optimize the description and perform the generation process again. Detailed Implementation
[0186] The present invention will now be further described through specific embodiments. These embodiments are not intended to limit the scope of the invention, but rather represent preferred embodiments of various aspects of the invention to more fully illustrate the invention described herein.
[0187] The embodiment of the method of this invention is an advanced computational model designed to simulate human olfaction using deep learning techniques. The embodiment of the model utilizes a main dataset of olfactory receptor (OR) structures and a secondary dataset of odorant molecules to predict the olfactory signature spectrum of novel molecules. The model is pre-trained using publicly available experimental receptor-ligand interaction data and fine-tuned to simulate the affinity of molecules for a range of ORs, thereby creating unique odor patterns.
[0188] Example 1 In one implementation, the neural network architecture of the model of the present invention is implemented as follows.
[0189] The computational model (“neural network-nose” model) is trained and used in different stages.
[0190] General Protein-Ligand Interaction Pre-training Phase In the initial phase, the model primarily focuses on training protein encoder and ligand encoder units. These units embed receptor proteins and small molecules into a unified latent space that models their 3D structure, physical properties, and chemical interactions. This pre-training phase provides an efficient method for simultaneously characterizing receptors and ligands, capturing key interaction features needed in subsequent phases.
[0191] Figure 1 An overview of the architecture and pre-training of a general protein-ligand interaction model is shown.
[0192] Odor Receptor Specific Encoder Training: Using Dataset 1 (a detailed dataset of odor receptors) and the encoders pre-trained in the first stage, the model trained a series of odor receptor specific encoders. Each encoder in this series was fine-tuned to efficiently predict ligand affinity for specific ORs, thereby enhancing the model's ability to predict odorant interactions.
[0193] Odor encoder layer training: A series of specific OR encoders are combined with a comprehensive database of odorant (and non-odorant) molecules to train the final odor encoder layer. This layer generates unique odor embeddings for molecules by simulating the interaction between molecules and the trained OR encoders, thereby effectively mapping molecules into a high-dimensional "odor space".
[0194] Network architecture The subsequent training phase includes the following steps.
[0195] 1. Pre-training and Data Collection A comprehensive dataset (referred to as Dataset 1) is compiled, which includes olfactory receptor (OR) DNA sequences, crystallographic data, or predicted structures.
[0196] A database containing experimental crystallographic 3D information on ligand-receptor pairs (publicly available) was used.
[0197] The receptor encoder and ligand encoder are trained in conjunction with the pairing interaction encoder layer. These units employ a transformer-based architecture and a cross-attention mechanism to capture the complex relationships between receptor atoms and their ligands. This allows all three encoding units to efficiently update their representations through a multi-head self-attention layer (followed by a feedforward neural network, normalization, and dropout layers).
[0198] 2. Receptor encoding The olfactory receptor (OR) structure is then encoded using pre-trained receptor encoding units, creating a series of preset embeddings that allow the protein encoder to be omitted in subsequent training and inference phases.
[0199] 3. Odor agent coding Odorant molecules are characterized by their structural features and embedded into the same high-dimensional vector space using a pre-trained ligand encoder.
[0200] The odorant encoder block consists of several independent transformer encoder layers, which process molecular features and update their representations through a multi-head self-attention mechanism during the pre-training phase.
[0201] 4. Paired Interaction Encoder Next, the characterization of odorant molecules is combined with the characterization of pre-encoded olfactory receptors (ORs) through a cross-attention mechanism.
[0202] This interaction block uses a triangular attention mechanism to enhance the learning of cross-pairing relationships, thereby optimizing molecular characterization.
[0203] 5. Affinity Prediction The learned interaction representations are projected onto the affinity matrix through linear transformation and activation functions.
[0204] Hybrid density networks (MDNs) use learned statistical potential parameters to model the affinity distribution, thereby aiding in the ranking of molecular affinities.
[0205] Reasoning stage 1. Odor agent coding and affinity simulation: In this embodiment, only the odorant encoder (encoder block) is used to process the structural features of the new molecule.
[0206] The odorant encoder block updates the molecular characterization through independent transformer layers.
[0207] Subsequently, the odorant molecule (e.g., small molecule) embeddings are paired with all (pre-computed) individual olfactory receptor (OR) embeddings using a paired interaction encoder, and then processed by a feedforward neural network, normalization, and dropout layers to effectively create odor spatial embeddings.
[0208] This model maps corresponding molecules into a high-dimensional "odor space," where the correlation between molecules is entirely based on their simulated olfactory effects. This makes it possible to identify molecules with different structures but that elicit similar olfactory responses.
[0209] Using odor spatial embedding, molecules can be analyzed and compared based on predicted olfactory feature spectra using known algorithms. In an implementation, the embedding can also be transferred to an additional network layer to align the latent space with a second database (database #2), thereby pairing the perfumer's description of the olfactory sensations induced in subjects with the learned olfactory embedding.
[0210] Figure 2 An overview of embedding odorant molecules (e.g., small molecules) into a high-dimensional olfactory space is shown.
[0211] Training methods 1. Pre-training The receptor coding unit, ligand coding unit, and pairing interaction coding unit were pre-trained on a large receptor-ligand interaction dataset to learn the general characteristics of ligand-G protein-coupled receptor (GPCR) interactions.
[0212] Pre-training preferably focuses on learning latent space embeddings and intermolecular relationships using known combined data.
[0213] 2. Olfactory encoder training The ligand encoder and pairing interaction encoding units are "frozen" (locked; not learned), and the downstream olfactory encoder layer is trained on a diverse dataset of small molecules (limited to those that can interact with ORs) by randomly masking up to 37% of the pairing encodings and predicting the olfactory encodings generated by the full network. This adapts the network to ORs rather than the more generalized receptor. Through a self-distillation process, the model learns from its high-confidence predictions, thereby enhancing model robustness.
[0214] 3. Olfactory Alignment The fully trained system was then aligned with dataset two, which contains odorant molecules with detailed odor descriptions.
[0215] Known molecules are encoded via a frozen, pre-trained odorant encoder, paired with each pre-encoded OR embedding, and processed via an olfactory encoding pipeline. Additional layers are also trained to predict olfactory descriptions based on the olfactory embeddings.
[0216] The final layer of the olfactory encoder is in a non-frozen state, thus enabling rapid and efficient alignment with the perfumer's perception.
[0217] Summarize This embodiment demonstrates that the prediction method of the present invention has high detection accuracy. The initial part of the system (up to all units of the olfactory encoder) has been validated against known benchmarks and has achieved industry-leading performance in docking tasks. The method of the present invention exhibits a low error range and high accuracy in predicting odor descriptors, with a success rate exceeding 95% based on the validation set from the second dataset (dataset 2).
[0218] in conclusion This invention utilizes industry-leading deep learning technology and extensive pre-training to accurately simulate human olfaction. By encoding molecules into a high-dimensional odor space, it provides a powerful tool for fragrance design, odor detection, and other olfactory applications, and can be applied, for example, to any and all of the application examples detailed herein.
[0219] The methods and implementations of the present invention used in this paper are applicable to a variety of task scenarios, including but not limited to: molecular virtual screening and all embodiments described below, namely, odor “telephone”, VR system odor synthesizer, olfactory pleasure sensor system, perfume formulation reconstruction, natural product approximation simulation, and computer simulation process for the generation and development of novel fragrance and flavor components.
[0220] Example 2 One embodiment of the model in this invention employs a classic docking method without incorporating neural networks (NN), artificial intelligence (AI), or machine learning (ML) methods. This embodiment is entirely based on mature molecular docking technology. Numerous commercial and open-source software suites currently provide access to such algorithms. For the purposes of this embodiment, we will use the freely available AutoDock Vina to build a simple model with excellent predictive capabilities.
[0221] Figure 3 An exemplary overview of docking and scoring of all OR-odorant molecule pairs using AutoDock Vina software is shown.
[0222] method docking phase 1. Data Collection Phase / Steps In the first step, the method of the present invention acquires or provides a first dataset (dataset 1) containing three-dimensional (3D) structures of olfactory receptors (ORs). These structures may be derived from experimental crystallographic data, computational predictions, or a combination thereof.
[0223] Subsequently, a second dataset (dataset 2) is assembled or provided, which contains the (3D) structures of odor molecules and textual descriptions of the odor sensations they induce in human subjects. Alternatively or additionally, the second database may also contain any other high-quality odor descriptors created by professional perfumers.
[0224] 2. Molecular docking First, AutoDock Vina is used to dock each odor molecule in the second dataset (dataset 2) with each olfactory receptor (OR) in the first dataset (dataset 1). Then, the binding affinity score for each molecule docking interaction is calculated. These affinity scores indicate the goodness of fit between the odor molecule and the OR binding site. Finally, various methods can be used to normalize and scale the data. In this embodiment, a t-score method is used to normalize all receptors.
[0225] 3. Affinity Scoring Vector For each molecule in the second dataset (dataset 2), its normalized affinity score list can be viewed as embedding that odor molecule into a high-dimensional olfactory space in which molecules with similar odors are closer to each other.
[0226] Odor Prediction Similarity analysis 1. Distance measurement in olfactory space Although those skilled in the art would expect that the “curse of dimensionality” would severely impair the ability to use distance metrics such as the L2 norm, the inventors’ experiments show that, thanks to biological characteristics, the data retains a very stable structure that allows for high success rates even when using these simplest distance metrics.
[0227] Therefore, in this embodiment of the invention, a simple distance metric is used to compare molecules based on their olfactory feature spectra. This involves calculating the Euclidean distance between the affinity score vectors of different odor molecules. Subsequently, molecules with similar odor feature spectra are identified by finding the molecule with the smallest distance in the olfactory space.
[0228] 2. Odor perception prediction For a new target odor molecule, it is first docked (fitted) with all the affinity scores (ORs) in Dataset 1 using AutoDock Vina to generate its affinity score vector. This vector is then compared with the affinity score vectors of molecules in Dataset 2 to identify the best match. Finally, based on the olfactory description of the best-matching molecule in Dataset 2, the odor sensation induced by the new molecule is predicted.
[0229] The implementation scheme of this invention (“Relational Nose” model) was tested based on the validation portion of dataset 2, and the results showed that this type of model can achieve a high accuracy (76.6%) in predicting the odor description of these “unseen” odorants.
[0230] The embodiments of this invention (“relational nose” model) are applicable to a variety of task scenarios, including but not limited to: molecular virtual screening and all embodiments described below, namely, odor “telephone”, VR system odor synthesizer, olfactory pleasure sensor system, perfume formulation reconstruction, natural product approximation simulation, and computer simulation process for the generation and development of novel fragrance and flavor components.
[0231] in conclusion The classical molecular docking relation nose model utilizes classical docking techniques to create high-dimensional embeddings of molecules in olfactory space. By evaluating the docking interactions between odor molecules and a series of ORs, this model provides a robust framework for predicting olfactory sensations and discovering new odorant molecules based on binding characteristic spectra. This method maintains high accuracy and reliability without requiring complex AI or ML methods.
[0232] Example 3 Embedding of odor mixtures The two implementation schemes analyzed in Examples 1 and 2 above (the "neural network-nose" model and the molecular docking "relational nose" model) both introduce a novel concept for characterizing the olfactory feature spectrum of odor mixtures. This concept revolves around creating composite patterns or embeddings that capture the combined effects of multiple odor molecules. The embedding vector of an odor mixture is derived from a weighted sum of the individual embedding vectors of its constituent molecules. This method can accurately and scalably characterize complex olfactory sensations.
[0233] method 1. Characterization of single molecules First, each odor molecule Mi in the mixture is represented by its embedding vector Oi. This embedding vector can be obtained through any of the embodiments in Example 1 or 2 above (neural network-nose model, classical molecular docking relationship nose model) or any other implementation of the present invention.
[0234] 2. Weighting In this step, the summation weights can be derived in various ways. For the purposes of this embodiment, this document adopts the commonly used definition in the art, namely, using relative concentration as the summation weight.
[0235] Subsequently, the relative concentration Ci of each odor molecule Mi in the air mixture was determined. This can be achieved through theoretical calculations or experimental measurements, such as headspace analysis using techniques like gas chromatography (GC), mass spectrometry, or carbon-13 nuclear magnetic resonance (NMR).
[0236] 3. Composite embedding of mixtures First, calculate the composite embedding vector O of the odor mixture. mix This is a weighted sum of the individual embedding vectors. The weights are defined based on the relative concentrations of the molecules.
[0237] Subsequently, the formula for calculating the composite embedding vector is: O mix =∑Wi·Qi Where Omix represents the olfactory embedding vector of the mixture; wi represents the relative concentration of molecule Mi; and Oi represents the olfactory embedding vector of molecule Mi.
[0238] Example 4 In another embodiment, the olfactory sensation can be captured, transmitted, and reproduced at a remote location using an embodiment of the present invention.
[0239] This process may include the following steps: 1. Measure the input odor The odor of a sample is measured using a sensor system. A common method is headspace gas chromatography-mass spectrometry (GC-MS), which can identify and quantify volatile compounds present in the air surrounding the sample.
[0240] 2. Encode the detected molecules. Each molecule detected by the sensor system is encoded using a neural network-nose model, a classical molecular docking relation-nose model, or any other embodiment of the invention. This step involves generating an embedding vector for each molecule to capture its olfactory properties.
[0241] 3. Generate Mixture Mode The individual embedding vectors of the detected molecules are combined to form a mixture pattern. This operation is achieved using the aforementioned weighted summation method: E_mixture = ∑(C_i * E_i) Where C_i represents the relative concentration of molecule M_i; E_i represents the embedding vector of molecule M_i.
[0242] 4. Transport Mixture Mode Composite embedded vectors or hybrid patterns (E_mix) can be stored or transmitted to the receiving end via a communication medium. This may involve sending data over the Internet, a local area network, or any other suitable communication channel.
[0243] 5. Recreate the sense of smell At the receiving end, the system has a pre-defined list of aroma components that can be used to reproduce olfactory sensations. The system uses generative AI methods or other optimization algorithms (which can be gradient-based or gradient-free) to generate mixtures of available components to optimize the similarity between the reproduced mixture and the source pattern E_ mixture.
[0244] An example workflow is as follows: 1. Measurement Headspace GC-MS was used to analyze the fresh oranges. Several key volatile molecules, such as limonene, myrcene, and octanal, were detected.
[0245] 2. Encoding Each molecule is encoded as an embedding vector using a neural network-nose model. For example: E_limonene, E_myrcene, E_octanal.
[0246] Mixed mode Determine the relative concentrations of molecules and calculate the mixture pattern: E_mixture = (C_limonene * E_limonene) + (C_myrcene * E_myrcene) + (C_octanal * E_octanal) 3. Transmission Transmit the composite mixture mode E_mix to a remote location.
[0247] 4. Reproduce At the destination, the system uses its list of available aroma components and applies a generative AI algorithm to adjust the concentrations of synthetic limonene, myrcene, and octanal to recreate the olfactory sensation of a fresh orange.
[0248] in conclusion This odor "telephone" embodiment demonstrates the potential of digital olfactory communication. By encoding, transmitting, and reproducing odor signature spectra, it opens up new possibilities for remote sensory experiences. Whether for quality control in the food industry, remote diagnostics, or immersive virtual reality environments, this method utilizes advanced olfactory modeling techniques to bridge the gap between remote locations.
[0249] Example 5 VR System Scent Synthesizer This embodiment extends the capabilities of the present invention by synthesizing odors based on contextual information, such as textual descriptions of VR scenes, to provide an immersive virtual reality (VR) experience.
[0250] An example workflow is as follows: 1. Contextual Input Instead of starting with actual odor measurements, the system receives textual descriptions or other forms of contextual input related to the VR scene. For example, a VR scene might describe a bustling city filled with the aromas of fresh bread, flowers, and spices.
[0251] 2. Generating modes via synthesizer The synthesizer generates olfactory patterns using external control (e.g., textual descriptions). This can be achieved through natural language processing (NLP) and generative AI models, or by generating patterns during the VR experience using any other tools, or by pre-generating patterns and “playing them back” during the VR session. Text input: "A bustling city filled with the aroma of fresh bread, flowers and spices." NLP: The system parses text input and identifies key aroma components: fresh bread, flowers, and spices.
[0252] Generative AI: The system uses generative models to create embedding vectors for the expected olfactory experience.
[0253] 3. Recreate the sense of smell: The system uses its list of aroma components to reproduce the synthesized olfactory pattern. Furthermore, it adjusts the concentration of each available component using generative AI or optimization algorithms to match the expected olfactory embedding pattern. The goal is to ensure that the reproduced scent closely resembles the requests of the VR application.
[0254] in conclusion A VR system odor synthesizer exemplifies how olfactory technology can enhance virtual experiences. By generating and reproducing odors based on contextual input such as text descriptions, the system significantly improves the immersive quality of VR environments. This capability can be applied to various fields, including games, virtual tourism, training simulations, and therapeutic environments, making virtual experiences more realistic and engaging.
[0255] Example 6 Fragrance Formula Reconstruction Reconstructing a perfume formula by replacing a single ingredient without significantly altering the perceived scent is traditionally a complex task requiring experienced perfumers. However, this process can be automated using the model implementation scheme and auxiliary calculation method described in this invention. An exemplary workflow is as follows: 1. Generate the original recipe mixture mode: Similar to the olfactory “telephone” example, an olfactory pattern of the original perfume formula is first generated. This pattern characterizes the overall fragrance profile formed by the combined effects of the individual components and their respective proportions.
[0256] Example: Suppose an original perfume formulation contains ingredients A, B, C, D, and E in specific proportions. Generate an olfactory pattern O_original to capture the fragrance characteristic spectrum of these ingredient combinations. Obtain the olfactory embedding pattern of the mixture through experimental methods (e.g., headspace analysis of the current perfume mixture) or by computing the target olfactory pattern (e.g., weighted summation of ingredient patterns, where the weight of each ingredient is proportional to its concentration and relative vapor pressure in the formulation).
[0257] 2. Remove unwanted components: Identify and remove the components that need to be replaced. This may involve factors such as raw material supply, cost, or regulatory changes.
[0258] Example: Component D needs to be removed. The resulting formula now consists of components A, B, C, and E.
[0259] The refactoring process can also introduce additional constraints, such as using only natural ingredients, ensuring stability in specific environments (e.g., highly alkaline detergents), or keeping costs within a specific range. Example: Suppose the refactored perfume formula must contain only natural ingredients and be below a specified cost. Generative models consider these constraints when optimizing the new formula.
[0260] 3. Minimize the additions and modifications to approximate the original model: The goal is to approximate the original olfactory pattern O_original as closely as possible by adjusting the remaining components and any potential additions. This will be achieved using generative AI models or optimization algorithms to find the optimal combination and ratio of remaining and additions that best match O_original.
[0261] Example: Introduce potential alternative components of D (e.g., F and G), and adjust the proportions of A, B, C, E, F and G to form a new pattern that approximates the original O_O_new.
[0262] An example workflow is as follows: 1. Generate the original recipe mixture mode: Original formula: Ingredients A, B, C, D, E Original proportions: 20%, 25%, 15%, 30%, 10% O_Original: The olfactory pattern generated based on the above ingredients and proportions.
[0263] 2. Remove unwanted components: Undesirable ingredient: D Remaining ingredients: Ingredients A, B, C, E 3. Approximate simulation of the original model: Potential alternative ingredients: F, G Optimization process: Adjust the ratio of A, B, C, and E and introduce F and G to create O_new.
[0264] Objective: Ensure that O_new is as close as possible to O_original.
[0265] New formula: with optimized proportions of ingredients A, B, C, E, F, and G.
[0266] in conclusion Using AI and computational methods can significantly simplify the fragrance formulation reconstruction process. By generating olfactory patterns of the original formulation, removing unwanted ingredients, and optimizing the remaining components, the new formulation's fragrance closely approximates the original. This process not only improves efficiency but also reduces reliance on manual intervention. Furthermore, constraints such as ingredient type, stability requirements, and cost can be seamlessly integrated into the optimization process, ensuring that the new formulation meets all necessary conditions while maintaining the desired fragrance profile.
[0267] Example 7 Natural Product Approximate Simulation Another task similar to the aforementioned perfume formulation reconstruction task in terms of implementation method is to use a combination of perfume ingredients to simulate the fragrance characteristic spectrum of natural products (e.g., flowers, fruits, or natural environments). However, the difference is that in this case, the original model originates from a natural source rather than an existing perfume formulation.
[0268] An exemplary workflow for solving this task using the implementation scheme of the present invention is as follows: 1. Generate natural source patterns The first step involves capturing the olfactory patterns of natural products. This can be achieved by analyzing volatile compounds released from natural sources using techniques such as gas chromatography-mass spectrometry (GC-MS).
[0269] Example: Assume the natural product is rose. Generate an olfactory pattern O_natural that characterizes the aroma spectrum of rose.
[0270] 2. Using perfume ingredients to approximate the natural process. The goal is to simulate the natural olfactory pattern O_Natural as closely as possible using available fragrance ingredients. This involves employing generative AI models or optimization algorithms to find the optimal combination and ratio of fragrance ingredients that best match O_Natural.
[0271] Example: Using synthetic or natural ingredients such as geranium oil, citronellol, nerol, and phenylethyl alcohol, a new model that approximates the natural state is formed.
[0272] 3. Apply additional restrictions if necessary: The simulation process can also introduce additional constraints, such as using only natural ingredients, ensuring stability in specific environments (e.g., temperature, pH), or keeping costs within a specific range.
[0273] Example: Assume that the simulated product must contain only natural ingredients and remain stable in laundry detergent. Generative models take these constraints into account when optimizing new formulations.
[0274] in conclusion The process of simulating the fragrance profile of natural products using perfume ingredients involves capturing olfactory patterns from natural sources, identifying key compounds, and optimizing the combination of perfume ingredients to match natural scents. This process can be run efficiently using generative AI models or optimization algorithms and can introduce additional constraints such as ingredient type, stability requirements, and cost. The result is a formulation that highly replicates the natural scent while meeting all specified conditions.
[0275] Example 8 In the implementation scheme, the method of the present invention can be used in a computer simulation process for the generation and development of novel flavor and fragrance components.
[0276] Overview The practicality of computer simulation processes in the generation and development of novel fragrance and flavor components is based on any of the aforementioned human olfactory system models. This model and its corresponding development process offer a transformative approach to the traditionally complex and costly process of creating new aroma molecules. By utilizing the aforementioned models and datasets, the system simplifies the process, making it more efficient, cost-effective, and environmentally friendly.
[0277] An exemplary workflow using the embodiments of the present invention is as follows: 1. Database #2 A database containing molecular and olfactory descriptions created by perfumers.
[0278] 2. Encoder Model Any variation of the aforementioned human olfactory model, including but not limited to the neural network-nose model and the classic molecular docking-nose model.
[0279] 3. Translation Model Function: Converts olfactory codes into scent descriptions that humans can understand.
[0280] Mechanism: Using a multimodal deep neural network, olfactory encoding and human language descriptions of fragrance are embedded in the same latent space.
[0281] 4. Generative Model Function: To generate new molecular structures that meet the expected olfactory and other properties.
[0282] Mechanism: Guided by olfactory coding of encoder models and additional criteria such as molecular size, atom type, syntheticity and production cost.
[0283] 5. Iterative Evaluation and Optimization Process: Selected molecules generated by the system can be synthesized in the laboratory, evaluated by perfumers, and then fed back into the system for continuous improvement.
[0284] Figure 4An example of such a translation model is shown, which is trained using an encoder model and dataset 2 and can perform bidirectional translation between olfactory embeddings and odor descriptor types provided by dataset 2 (e.g., natural language translation).
[0285] The more detailed workflow is as follows: 1. Initial fragrance description: The process begins with the perfumer providing a detailed description of the desired fragrance. This can be done in natural language (e.g., "a fresh floral scent with a hint of citrus") or in any other structured manner. The description includes the fragrance notes, scent type, and its relationship to other chemical components. Positive descriptions include "the scent of vanillin," while negative descriptions include "used to neutralize, mask, and eliminate the unpleasant odors produced by cigarette smoke."
[0286] 2. Convert to olfactory embedding: The translation model converts these odor descriptors / descriptions into olfactory embeddings (a type of numerical vector) suitable for guiding the generative model.
[0287] 3. Molecular formation: The generative model uses this olfactory embedding to generate candidate molecules. In addition to evaluating the aroma signature spectra of these molecules, other properties (e.g., molecular size, syntheticity, cost-effectiveness) can be used as additional signals for the generative algorithm.
[0288] 4. Perfumer Evaluation: Perfumers review candidate molecules and their corresponding fragrance descriptions generated by the translation model. Suitable candidate molecules are then selected for synthesis.
[0289] 5. Synthesis and Evaluation: Selected molecules are synthesized in a chemical laboratory, and then perfumers evaluate their olfactory quality, stability, and compatibility with other ingredients.
[0290] 6. Feedback Loop: Insights gained from the evaluation can be used to optimize the model. This iterative process ensures that the accuracy and reliability of the system are continuously improved.
[0291] The advantages of the above method are reflected in the following aspects: Increased efficiency: This implementation scheme significantly reduces the number of laboratory experiments required to identify usable aroma molecules, thereby saving time and resources; Cost-effectiveness: This method minimizes the need for large-scale chemical synthesis and testing, thereby reducing R&D costs; Environmental benefits: This method reduces the consumption of chemical reagents, thereby reducing the ecological footprint of fragrance and flavor development; Improved precision: The method of this invention provides high control over the generated molecules, ensuring that they conform to specific olfactory characteristic spectra and other expected properties; and Innovation: This method can create novel fragrance and flavor molecules that may not be discovered by traditional methods.
[0292] These methods can be applied to the fragrance industry to create new perfume ingredients with unique aroma profiles. For example, they can be used to develop novel floral scents that are both appealing and environmentally friendly. Additionally, these methods can be used to design new flavor molecules for food and beverages. For example, they can be used to create cost-effective alternatives to expensive spices such as cardamom. Furthermore, these methods can be applied to regulatory compliance to develop alternatives to banned or restricted fragrance molecules. For example, they can be used to generate alternatives to molecules such as lily aldehyde and neolily aldehyde, which are restricted due to environmental and health concerns.
[0293] in conclusion The implementation scheme for generating and developing novel flavor and fragrance components using computer simulation processes represents a breakthrough in this field. By leveraging the power of artificial intelligence and a deep understanding of the human olfactory system, this method holds the promise of revolutionizing the creation of novel flavor and fragrance molecules. Its significant advantages in efficiency, cost, environmental impact, and innovation make it a valuable tool for the flavor and fragrance industry.
[0294] References Fleischer J, Breer H, Strotmann J. Mammalian olfactory receptors. Front Cell Neurosci. 2009 Aug 27;3:9. doi: 10.3389 / neuro.03.009.2009. PMID:19753143; PMCID: PMC2742912. John C. Dearden, Quantitative structure-activity relationships (QSAR) and odour, Food Quality and Preference, Volume 5, Issues 1-2, 1994, Pages 81-86. Hau, K.M. and Connell, D.W. (1998), Quantitative Structure-ActivityRelationships (QSARs) for Odor Thresholds of Volatile Organic Compounds(VOCs). Indoor Air, 8: 23-33. https: / / doi.org / 10.1111 / j. 1600-0668.1998.t01 -3-00004.X
Claims
1. A computer-based method for predicting molecular odors, characterized in that, Includes the following steps: a. Provide a first database containing information about the structure of one or more olfactory receptors OR, said structure preferably being a three-dimensional (3D) and / or chemical structure; b. Train a machine learning model or AI receptor encoder based on the data contained in the first database; c. Provide a second database containing (i) the 3D structures of one or more odor molecules and (ii) corresponding textual descriptions of the odor sensations induced by said one or more odor molecules in human subjects; d. Train a machine learning (ML) model or an artificial intelligence (AI) odor encoder based on the data contained in the second database; e. Fit the 3D structure of one or more odor molecules to the 3D structure of one or more olfactory receptors (ORs), preferably to the ligand binding sites of one or more ORs, thereby determining the affinity score for each fitted combination of the one or more odor molecules and the corresponding olfactory receptors. The fitting process includes using and training an ML model or AI-paired interaction encoder. The affinity score indicates the degree of fit between the corresponding odor molecule and the olfactory receptor OR, preferably the degree of fit with the ligand binding site of the corresponding OR, and preferably indicates the ability of the corresponding odor molecule to (i) bind or interact with the corresponding olfactory receptor and / or (ii) activate the corresponding olfactory receptor. f. Create a comprehensive third database or third data embedding structure that includes data from the first database, data from the second database, and affinity scores determined for each fitted combination of the corresponding odor molecule and the corresponding olfactory receptor; g. Provide the 3D structure of at least a first target odor molecule not included in the second database; and h. Based on the data contained in the third database or third data embedding structure, predict the odor sensation induced by the at least first target odor molecule in human subjects using machine learning or AI.
2. The computer implementation method according to claim 1, characterized in that, Step h: Predicting the odor sensation induced by the at least first target odor molecule in human subjects, based on the 3D structure of the at least first target odor molecule, using one or more of the machine learning models or AI trained in steps b, d, and / or e according to claim 1. - At least one odor molecule included in the third database or third data embedding structure, wherein the odor molecule has the highest similarity score to one or more olfactory receptors included in the third database or third data embedding structure with the at least first target odor molecule; and / or - The odor sensation induced in human subjects by the at least first target odor molecule, wherein the odor sensation preferably corresponds to one or more textual descriptions of the odor sensation induced in human subjects by the odor molecule contained in the third database or third data embedding structure.
3. The computer implementation method according to claim 1 or 2, characterized in that, Step h: predicting the odor sensation induced by the at least first target odor molecule in human subjects, including: (I) Using the ML model or AI odorant encoder trained in step d, compare the 3D structures of at least the first target odor molecule and one or more of the odorant molecules contained in the second database or the third database; (II) Using the ML model or AI pairing interaction encoder trained in step e, the 3D structure of the at least first target odor molecule is fitted to the 3D structure of one or more of the olfactory receptors (ORs), preferably to the ligand binding site of the OR, thereby determining the affinity score of each fitted combination of the at least first target odor molecule and the corresponding OR. (III) Compare one or more of the affinity scores determined in step (II) for each fitted combination of the at least first target odor molecule with the one or more olfactory receptor ORs with the affinity scores of each combination of the one or more ORs with the one or more odor molecules contained in the third database or third data embedding structure; and (IV) Accordingly, identify at least one odor molecule contained in the third database or third data embedding structure, wherein the odor molecule has the highest similarity score to the one or more ORs with the at least first target odor molecule. The corresponding textual description of the odor sensation induced in human subjects by the one or more odor molecules identified in step (IV), contained in the third database or third data embedding structure, indicates the odor sensation induced in human subjects by the at least first target odor molecule.
4. The method according to any one of the preceding claims, characterized in that, Finally, the at least first target odor molecule is chemically synthesized based on the 3D structure provided in step g according to claim 1.
5. The method according to any one of the preceding claims, characterized in that, The at least one olfactory receptor is selected from the group consisting of: olfactory receptor OR, formyl peptide receptor FPR, guanylate cyclase GC-D, vomeronasal receptors V1R and V2R, and trace amine-associated receptor TAAR.
6. The method according to any one of the preceding claims, characterized in that, The 3D structure of olfactory receptors and / or odor molecules is determined using the following information: experimental crystallographic information, computer simulations, artificial intelligence (AI) and / or machine learning (ML) models trained on experimental data, gene sequence information of the corresponding olfactory receptors and / or odor molecules, or any combination thereof.
7. The method according to any one of the preceding claims, characterized in that, In step e, the dependencies and / or correlations between different olfactory receptors and / or odor molecules are identified by using statistical methods and / or machine learning-based methods. Affinity scores are determined only for a subset of the odor molecules and / or olfactory receptors contained in the first database, the second database, and / or the third database. Optionally, affinity scores are assigned to 25-100 least relevant receptors and / or odor molecules by setting a threshold for the correlation coefficients between different receptors and / or odor molecules.
8. The method according to any one of the preceding claims, characterized in that, In step e, the affinity score is determined using computer physical simulation, ML algorithms, and / or algorithms that combine statistical methods with ML methods and computational simulation elements.
9. The method according to any one of the preceding claims, characterized in that, The at least first target odor molecule comprises a mixture of at least first and second target odor molecules, and wherein in step h, in addition to providing (i) the 3D structure of the at least first and second target odor molecules, (ii) the relative percentage of each of the at least first and second target odor molecules in the vapor of the mixture and / or (iii) the statistical probability of each odor molecule interacting with olfactory receptors in the nasal cavity of a human subject, and wherein step h comprises: predicting the odor sensation induced by the mixture of the at least first and second target odor molecules in a human subject.
10. The method according to the preceding claims, characterized in that, The relative percentage of the at least first and second odor molecules in the vapor of the mixture is determined by further considering the relative percentage of the at least first and second odor molecules in the mixture and / or the corresponding vapor pressure of the mixture.
11. A computer-implemented method for predicting the 3D structure of a target odor molecule capable of inducing a target odor sensation in human subjects, the method being based on the binding affinity of the target odor molecule to at least one olfactory receptor, characterized in that, Includes steps a to f of the method according to claim 1, and further includes the following steps: g. Select a text description of the target odor sensation induced by the first odor molecule from the third database or the third data embedding structure; h. Determine one or more olfactory receptors with the highest affinity score for the first odor molecule from the third database or third data embedding structure; and i. Based on the data from the third database or the third data embedding structure, predict the 3D structure of the target odor molecule that is not included in the third database. The predicted 3D structure is modeled / predicted and fitted with a high affinity score to the 3D structure of the ligand binding site of the one or more olfactory receptors selected in step h, wherein the high affinity score of the predicted 3D structure of the target odor molecule indicates the ability of the target odor molecule to bind or interact with the ligand binding site of the corresponding one or more olfactory receptors, and / or The ability to activate one or more olfactory receptors This induces an odor sensation similar to the first odor molecule in human subjects.
12. The method according to claim 11, characterized in that, The prediction method used in step h is implemented by an AI and / or machine learning (ML) model, preferably a gradient-free ML model. The gradient-free ML model preferably includes a genetic algorithm for generating molecules and / or optimizing the prediction towards a given numerical vector, or... It is implemented by a gradient-based ML model, which preferably includes a generative adversarial network (GAN) or a diffusion probability model.
13. The method according to claim 12, characterized in that, The AI and / or ML model is or includes one or more of the machine learning model or AI trained in steps b, d and e according to claim 1.
14. The method according to claims 11-13, characterized in that, The predicted 3D structure of the target odor molecule was used to chemically synthesize the predicted molecular structure.