Computer-implemented method for providing at least one fragrance molecule representation for fragrance or flavor product production
A computer-implemented method using a machine learning model to generate fragrance molecule representations addresses the challenge of predicting odorant receptor activity, enabling the creation of fragrance and flavor products with desired sensory characteristics by associating ingredients with receptor activation and modulation.
Patent Information
- Application Number
- PCT/EP2025/068513
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-04-21
- Filing Date
- 2025-06-30
- Publication Date
- 2026-01-08
AI Technical Summary
Existing fragrance and flavor formulation technologies lack efficient methods to predict odorant receptor activity and interactions, limiting the ability to create new ingredients and formulations that achieve desired sensory experiences.
A computer-implemented method using a machine learning model to generate fragrance molecule representations by determining odorant receptor activity targets, allowing for the association of ingredients with receptor activation and modulation, and enabling the creation of desired tonalities and intensities in fragrance or flavor products.
Enables the prediction of odorant receptor activity and modulation, facilitating the development of fragrance and flavor products with desired sensory characteristics without requiring specific knowledge of receptor activation, and allowing for the quick exploration of molecular scaffolds.
Smart Images

Figure EP2025068513_08012026_PF_FP_ABST
Abstract
Description
[0001] DESCRIPTION
[0002] TITLE OF THE INVENTION: COMPUTER-IMPLEMENTED METHOD FOR PROVIDING AT LEAST ONE FRAGRANCE MOLECULE REPRESENTATION FOR FRAGRANCE OR FLAVOR PRODUCT PRODUCTION
[0003] TECHNICAL FIELD OF THE INVENTION
[0004] The present invention relates to a computer-implemented method for providing at least one fragrance molecule representation, representing a materializable fragrance or flavor molecule for fragrance or flavor product production, to a computer-implemented method to train a machine learning model to provide at least one fragrance molecule representation, and to a corresponding computer program product, computer-readable storage medium and device.
[0005] The present invention is applicable to fragrance and flavor formulation in, for example, the context of perfumery, cosmetics, and comestible applications.
[0006] BACKGROUND OF THE INVENTION
[0007] Following the initial identification of odorant receptors in 1991 , much has been done to advance the understanding of olfaction and how humans perceive volatile chemicals. At the periphery of the human olfactory system, the main olfactory epithelium lining the nasal cavity harbors several million olfactory sensory neurons (OSNs), each expressing a single member from a family of approximately 400 intact odorant (olfactory) receptor (OR) genes. Each OR can be activated, inhibited, or modulated by several volatile compounds while, in turn, a single volatile compound can also activate, inhibit, or modulate several ORs. This results in a combinatorial OSN activation code for each volatile chemical, and mixture thereof. Each receptor is stochastically expressed in OSNs in a monogenic and monoallelic fashion, yielding one OSN type for each OR allele present in the genome. Depending on the frequency of gene choice for each OR, the subset of the total OSN population dedicated to each OSN type varies. The OSN-type subpopulations each provide a discrete channel of information about the molecular identity, or identities, and concentrations of volatile compounds with which they are in contact at any given time. The level of activity induced in each OSN of each OSN type varies according to the concentration of volatile compounds to which its OR is receptive, as well as according to the binding and activation parameters associated with each OR-compound pair. Further, some OR-ligand pairs induce or enhance OSN activity, while others reduce or prevent OSN activity. This can occur competitively, where compounds compete to bind an OR, or non-competitively, where more than one compound can bind an OR simultaneously. The resulting combinatorial logic of these approximately 400 input channels (i.e. the peripheral olfactory system) is transformed by subsequent, downstream, neurological networks and results in our sense of smell. The fragrance and flavor (F&F) industry is constantly in search of new ingredients, novel perfumery and flavor applications, improved sensory experiences, and compounds that are more stable, biodegradable and non-toxic.
[0008] Because an odorant or aroma molecule may interact with several olfactory receptors (ORs), it is often difficult to infer the odor or aroma quality of volatile compounds based on their chemical structures alone and, correspondingly, of mixtures of such volatile compounds. Conversely, because each OR may interact with several odorant and aroma molecules possessing different chemical structures and evoke different sets of qualities, it has often been difficult to infer the nature of the information encoded by an OR.
[0009] It is widely known that predicting the perception of an ingredient, which corresponds to one or more fragrance molecule, in itself or of an ingredient formula, which corresponds to a sum of ingredients, is extremely difficult. This difficulty results from non-linear interactions between fragrance molecules, on one side, and from significant biological perception variations on another side. Furthermore, today, while it is clear that the psychophysical perception of an ingredient or formula is linked to the activation and modulation of odorant receptors located in the nose of animals, the ingredients to odorant receptor interactions are mostly unknown. What is particularly unknown is why a particular molecule triggers a particular odorant receptor.
[0010] Today, there exists no efficient way of predicting odorant receptor activity in link with a particular odorant, fragrance molecule, flavor ingredient or formula.
[0011] More generally, as there exists no efficient way of predicting odorant receptor activity, volatile compound odor; fragrance and flavor formulation capabilities available today are limited.
[0012] SUMMARY OF THE INVENTION
[0013] The present invention is intended to remedy all or part of these disadvantages.
[0014] To this effect, according to a first aspect, the present invention aims at a computer- implemented method for providing at least one fragrance molecule representation, representing a materializable fragrance molecule for fragrance or flavor product production, which comprises: a step of determining at least one value representative of an odorant receptor activity target caused by the material exposition of said odorant receptor to the fragrance molecule associated with a fragrance molecule representation, a step of generating at least one fragrance molecule representation by a trained machinelearning model using as input value representative of an odorant receptor activity target, and a step of providing at least one of the generated at least one fragrance molecule representation.
[0015] Such provisions allow, based on empirical measurements, to determine underlying rules which associate ingredients with odorant receptor activation. Such underlying rules may then be used to generate fragrance molecules which can be integrated into fragrance or flavor products. It should be noted that while the above aspect refers to a fragrance or aroma molecule, such an aspect can also be applied to an ingredient comprising at least two fragrance or aroma molecules or to a formula comprising at least two ingredients.
[0016] In particular embodiments, during the step of generating, a value representative of a quantity is generated for at least one fragrance molecule represented by a fragrance molecule representation.
[0017] Such provisions allow for the association of an absolute or relative quantity of a particular molecule, ingredient or formula in order to achieve the activation of at least one odorant receptor.
[0018] In particular embodiments, the method object of the present invention further comprises a step of defining at least one determined (desired or undesired) tonality value and / or intensity, the step of determining at least one value representative of an odorant receptor activity target being defined as a function of at least one said defined determined tonality value and / or intensity.
[0019] Such provisions allow for a user to determine the type of smell desired for the molecule, ingredient or formula, without the user requiring knowing which odorant receptors to specifically activate. Formulas with multiple ingredients generate receptor modulation (inhibition or enhancement) that affect the value and / or intensity due to modulated receptor activity. Symmetrically, a user can define an undesired tonality value, and the molecule, ingredient or formula generated can be generated to fit as little as possible with this undesired tonality.
[0020] In particular embodiments, the method object of the present invention further comprises a step of determining receptor to perceptual descriptor associations, which comprises the steps of:
[0021] - collecting a set of at least one perceptual descriptor,
[0022] - for at least one collected perceptual descriptor, screening at least one olfactory receptor, represented by an olfactory receptor representation in a computing system, determining the impact of an ingredient on at least one olfactory receptor, and providing an association of at least one collected perceptual descriptor to at least one odorant receptor and / or to at least one ingredient, said at least one association being used at least during the step of determining.
[0023] Such provisions allow for the association of perceptual descriptors to odorant receptors.
[0024] In particular embodiments, the step of generating further generates a scaffold representation, representing a materializable molecular scaffold, associated to at least one generated fragrance molecule representation, said method further comprising: a step of selecting a scaffold representation and a step of regenerating at least one fragrance molecule representation by a trained machinelearning model using as input value representative of an odorant receptor activity target and the selected scaffold representation, and
[0025] - the step of providing further providing said regenerated fragrance molecule representation.
[0026] Such provisions allow for the quick exploration of generated fragrance molecules by filtering molecular scaffolds of interest. In particular embodiments, the method object of the present invention comprises the steps of: providing an original set of: exemplar fragrance molecule representations, said set of exemplar fragrance molecule representations being representative of materialized fragrance molecules, and at least one associated fragrance molecule feature value, exemplar odorant receptor representation, representative of physically existing odorant receptors, and at least one associated odorant receptor feature value, exemplar odorant receptor activation values for said at least one said exemplar fragrance molecule representation and at least one exemplar odorant receptor representation, representative of empirically measured activation values caused by the material exposition of said exemplar odorant receptor to said exemplar fragrance molecule, and training a machine-learning model using the training set, wherein the machine-learning model is trained to associate at least one feature value with an odorant receptor activation value, wherein said trained machine-learning model is used during the step of generating.
[0027] Such provisions allow for the constitution of the trained machine-learning model used.
[0028] In particular embodiments, the method object of the present invention comprises a step of recording, in a database, empirically measured odorant receptor activation values caused by the material exposition of an odorant receptor, corresponding to an exemplar odorant receptor representation, to a fragrance molecule, corresponding to an exemplar fragrance molecule representation, said database being used during the step of providing an original set.
[0029] In particular embodiments, the method object of the present invention comprises: a step of materially exposing an odorant receptor, corresponding to an exemplar odorant receptor representation, to a fragrance molecule, corresponding to an exemplar fragrance molecule representation and a step of measuring odorant receptor activation values as a result of the material exposition. In particular embodiments, at least one fragrance molecule feature value is representative of: at least one perceptual descriptor associated with said fragrance molecule, a molecular fingerprint of said fragrance molecule, a measurable or measured physico-chemical value for said fragrance molecule, a machine-learning representation embedding, a three-dimensional conformer ensemble representation, a mixtures representation, or an analytical spectrum of said fragrance molecule. Machine-learning representation embeddings expose “hidden” relations between molecules that the human may not detect.
[0030] Three-dimensional conformer ensemble representations are useful as what specific conformer of a molecule may bind to a receptor is unknown. In some instances, the same molecule may bind to different receptors in different conformations. Thus, a conformer ensemble can help identify the set of conformers that may be important for binding.
[0031] Mixture representations are similar to conformers and are useful as an ingredient may contain multiple ingredients (that could be difficult to isolate individually) such that it is unknown which specific ingredient (or multiple ingredients) bind to the receptor.
[0032] Analytical spectrum of molecules may be useful to capture extremely complex mixtures like plant extracts (well over 100s of molecules that can be difficult to isolate).
[0033] In particular embodiments, at least one fragrance molecule feature value is representative of a three-dimensional point cloud representation associated with at least one electrostatic potential value.
[0034] In particular embodiments, at least one perceptual descriptor is associated with at least one intensity weighting value.
[0035] An intensity value can correspond to a number on a scale or to a binary value, for example.
[0036] In particular embodiments, the method object of the present invention comprises, prior to the step of providing an original set: a step of generating a conformer for at least one exemplar fragrance molecule representation, a step of computing of a three-dimensional point cloud for at least one conformer and of at least one electrostatic potential value for at least one point of the volumetric point cloud.
[0037] In particular embodiments, the method object of the present invention comprises a step of converting at least one three-dimensional point cloud into a graph structure representation, said graph structure representation being used as at least one fragrance molecule feature value.
[0038] In particular embodiments, at least one odorant receptor feature value is representative of: a multi-state three-dimensional model of at least one exemplar odorant receptor, a machine learning latent representation of at least one exemplar odorant receptor, and / or an odorant receptor activation level as a function of fragrance molecule concentration values.
[0039] Multi-state three-dimensional models allow for better disambiguation between activation and modulation of the receptor. Machine learning latent representation is a quick, convenient way to represent a receptor.
[0040] In particular embodiments, the method object of the present invention comprises a step of sending a digital command representative of an instruction of materialising at least one fragrance ingredient corresponding to at least one fragrance ingredient representation provided.
[0041] Such provisions allow for the materialization of a generated molecule. In particular embodiments, the method object of the present invention comprises a step of materialising at least one fragrance ingredient corresponding to at least one fragrance ingredient representation provided.
[0042] Such provisions allow for the materialization of a generated molecule.
[0043] According to a second aspect, the present invention aims at a computer-implemented method to train a machine learning model to provide at least one fragrance ingredient representation, representing a materializable fragrance ingredient, which comprises the steps of: providing an original set of: exemplar fragrance molecule representations, said set of exemplar fragrance molecule representations being representative of materialized fragrance molecules, and at least one associated fragrance molecule feature value, exemplar odorant receptor representation, representative of physically existing odorant receptors, and at least one associated odorant receptor feature value, exemplar odorant receptor activation values for said at least one said exemplar fragrance molecule representation and at least one exemplar odorant receptor representation, representative of empirically measured activation values caused by the material exposition of said exemplar odorant receptor to said exemplar fragrance molecule, and training a machine-learning model using the training set, wherein the machine-learning model is trained to associate at least one feature value with an odorant receptor activation value.
[0044] According to a third aspect, the present invention aims at a computer program product which comprises instructions which upon execution by a computer cause the computer to execute the method object of the present invention.
[0045] According to a fourth aspect, the present invention aims at a computer-readable storage medium storing programming instructions which upon execution by a computer cause the computer to execute the method object of the present invention.
[0046] According to a fifth aspect, the present invention aims at a device for providing at least one fragrance molecule representation, representing a materializable fragrance molecule for fragrance or flavor product production, which comprises: means of determining at least one value representative of an odorant receptor activity target caused by the material exposition of said odorant receptor to the fragrance molecule associated with a fragrance molecule representation, means of generating at least one fragrance molecule representation by a trained machinelearning model using as input value representative of an odorant receptor activity target, and means of providing at least one of the generated at least one fragrance molecule representation. According to a sixth aspect the present invention aims at computer-implemented method for providing at least one odorant receptor representation, representing a physiological odorant receptor, which comprises: a step of defining a set of at least two perceptual descriptor representation and, for each said perceptual descriptor representation, a value representative of a relative weight of said descriptor in relation to at least one other defined descriptor, a step of predicting at least one activated odorant receptor representation by a trained machine-learning model using as input value the defined set and corresponding relative weight values, and a step of providing at least one of the predicted at least one activated odorant receptor representation.
[0047] In particular embodiments, the method object of the present invention further comprises at least one of the steps of: providing an original set of:
[0048] - weighted perceptual descriptors for individual fragrance compounds, odorant receptor activation values, classified into high or low, for said individual fragrance compounds or molecules, rebalancing the dataset to introduce a bias in favor of either the class of high or low odorant receptor activation values, training at least one random forest machine learning model to predict high / low odorant receptor activation from the weighted perceptual descriptors of said compounds as a function of the rebalanced original set, downstream of a step of predicting, determining the importance of perceptual descriptor features, as a function of the SHAP values for the predicted values, and / or assessing the contribution of a perceptual descriptor to at least one odorant receptor activation as a function of aggregated SHAP values across compounds or molecules for said perceptual descriptor.
[0049] The second to sixth aspects of the present invention exhibit the same advantages as the related first aspect.
[0050] BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Other advantages, purposes and particular characteristics of the invention shall be apparent from the following non-exhaustive description of at least one particular embodiment of the present invention, in relation to the drawings annexed hereto, in which:
[0052] Figure 1 represents, schematically, a succession of steps of a particular embodiment of the method subject of the present invention,
[0053] Figure 2 represents, schematically, a computer system with which an embodiment of the method subject of the present invention can be implemented, Figure 3 represents, schematically, a succession of steps of a particular embodiment implementing electrostatic point clouds,
[0054] Figure 4 represents, schematically, different steps of the generation of conformers and conversion to electrostatic point clouds,
[0055] Figure 5A represents, schematically, a succession of steps of a particular embodiment comprising encoding of electrostatic point clouds and training,
[0056] Figure 5B represents, schematically, example illustrations for visualizing the conversion of electrostatic point clouds to density-augmented graphs as well as the relation between positive and negative pairs for contrastive learning,
[0057] Figure 6A represents, schematically, an example model of a particular embodiment comprising property prediction from electrostatic point clouds,
[0058] Figure 6B represents, schematically, model results comparing various performance metrics on various tasks of a particular embodiment against a standard 2D-GNN,
[0059] Figure 7A represents, schematically, a succession of steps of a particular embodiment comprising scaffold hopping capabilities,
[0060] Figure 7B represents, schematically, an example model of a particular embodiment comprising electrostatic point cloud-conditioned molecule generation,
[0061] Figure 7C represents, schematically, an illustration of different scaffolds with similar electrostatic point clouds,
[0062] Figure 8 represents, schematically, a succession of steps of a particular embodiment of the method comprising the training of multi-modal embedding spaces,
[0063] Figure 9 represents, schematically, a succession of steps of a particular embodiment of the method of the present invention comprising the use of variational autoencoders and density estimates to represent perceptual space,
[0064] Figure 10 represents, schematically, a particular embodiment of the method of the present invention regarding the use of multiple instance learning,
[0065] Figure 11 represents, schematically, an example machine-format of the three-dimensional structure of a molecule,
[0066] Figure 12 represents, schematically, example inputs, outputs, a descriptor map and sources regarding a preprocessing step of descriptor maps,
[0067] Figure 13 represents, schematically, a succession of steps of a particular embodiment of the method to generate odorant receptor representations object of the present invention,
[0068] Figure 14 represents, schematically, a succession of steps of a particular embodiment of the method object of the present invention,
[0069] Figure 15 represents results of SHAP analysis of an example random forest model, wherein an anti-correlation can be seen for green with OR7A17,
[0070] Figures 16A to 16D represent an example receptor and the probability distributions of various descriptors, more particularly figures 16A to 16D represent an example of a convergence onto a tonality, a tonality with no real difference from the screening set, and a tonality with an "anticorrelation" (as the activity index cutoff is increased, the probability of the descriptor appearing drops to 0), and
[0071] Figure 17 represents, schematically, an example model of a particular embodiment comprising of a transformer layer with attention bias.
[0072] DETAILED DESCRIPTION OF THE INVENTION
[0073] This description is not exhaustive, as each feature of one embodiment may be combined with any other feature of any other embodiment in an advantageous manner. Also, various inventive concepts may be embodied as one or more methods, of which an example has been provided. The acts performed as part of the method may be ordered in any suitable way. Accordingly, embodiments may be constructed in which acts are performed in an order different than illustrated, which may include performing some acts simultaneously, even though shown as sequential acts in illustrative embodiments.
[0074] The indefinite articles ‘a’ and ‘an’, as used herein in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean ‘at least one’.
[0075] The phrase ‘and / or’, as used herein in the specification and in the claims, should be understood to mean ‘either or both’ of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with ‘and / or’ should be construed in the same fashion, i.e. ‘one or more’ of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the ‘and / or’ clause whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to ‘A and / or B’, when used in conjunction with open-ended language such as ‘comprising’ can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.
[0076] As used herein in the specification and in the claims, ‘or’ should be understood to have the same meaning as ‘and / or’ as defined above. For example, when separating items in a list, ‘or’ or ‘and / or’ shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one, of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as ‘only one of’ or ‘exactly one of’, or, when used in the claims, ‘consisting of’, will refer to the inclusion of exactly one element of a number or list of elements. In general, the term ‘or’ as used herein shall only be interpreted as indicating exclusive alternatives (i.e. ‘one or the other but not both’) when preceded by terms of exclusivity, such as ‘either,’ ‘one of,’ ‘only one of’, or ‘exactly one of’. ‘Consisting essentially of,’ when used in the claims, shall have its ordinary meaning as used in the field of patent law.
[0077] As used herein in the specification and in the claims, the phrase ‘at least one’, in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase ‘at least one’ refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, ‘at least one of A and B’ (or, equivalently, ‘at least one of A or B’, or, equivalently ‘at least one of A and / or B’) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.
[0078] In the claims, as well as in the specification above, all transitional phrases such as ‘comprising,’ ‘including,’ ‘carrying,’ ‘having,’ ‘containing,’ ‘involving,’ ‘holding,’ ‘composed of’, and the like are to be understood to be open-ended, i.e. , to mean including but not limited to. Only the transitional phrases ‘consisting of’ and ‘consisting essentially of’ shall be closed or semi-closed transitional phrases, respectively.
[0079] In a general manner, the terms ‘representation’ or ‘digital representation’ refer to any bijective digital representation of a physical item, such as a molecule. Such a representation may correspond to, for example, an entry in a database. A representation may refer to a label representative of the name, chemical structure or internal reference of an ingredient, for example.
[0080] In the context of the present description, the term “materialized” is intended as existing outside of the digital environment of the present invention. ‘Materialized’ may mean, for example, readily found in nature or synthesized in a laboratory or chemical plant. In any event, a materialized fragrance molecule representation presents a tangible reality. The terms ‘to be materialized or ‘materializable refer to the act of materialization of a fragrance ingredient associated with a representation.
[0081] As used herein, the terms “means of inputting” refer to, for example, a keyboard, mouse and / or touchscreen adapted to interact with a computing system in such a way to collect user input. In variants, the means of inputting are logical in nature, such as a network port of a computing system configured to receive an input command transmitted electronically. Such input means may be associated with a GUI (Graphic User Interface) shown to a user or an API (Application programming interface). In other variants, the means of inputting may be a sensor configured to measure a specified physical parameter relevant for the intended use case. Examples of means of inputting are disclosed in regard to figure 2.
[0082] As used herein, the terms “computing system”, “computer”, or “computer system” designate any electronic calculation device, whether unitary or distributed, capable of receiving numerical inputs and providing numerical outputs by and to any sort of interface, digital and / or analog. Typically, a computing system designates either a computer executing software having access to data storage or a client-server architecture wherein the data and / or calculation is performed at the server side while the client side acts as an interface. Examples of such computing systems are disclosed in regard to figure 2.
[0083] As used herein, the terms “fragrance molecule” designate any molecule, preferably volatile, presenting a fragrance capacity.
[0084] As used herein, the term “formula” designates a liquid, solid or gaseous assembly of at least two ingredients. It should be noted that a formula corresponds to at least two ingredients, each ingredient corresponding to at least one fragrance molecule.
[0085] As used herein, the term “composition” refers to the olfactory perception of a formula. A composition is defined by the perceivability of at least one tonality.
[0086] As used herein, "olfactory tonality" or "tonality" means an organoleptic property of a volatile molecule initiated by the activation of an OR, which activity is routed through the olfactory bulb and processed by the central nervous system to produce a specific experience in the subject. Said experience is of a nature that may be described by the subject linguistically, often with reference to everyday objects with which the olfactory experience is associated. Additional modifiers are typically required to capture the specificity of the tonality. Everyday objects that possess a characteristic “smell” typically elicit the experience of multiple separable tonalities when evaluated by a trained expert such as a perfumer. The experience of a single tonality is entirely dependent on the specific OR activated in the subject, such that a single OR produces the perception of a single tonality and the multitude of tonalities produced by many everyday objects is determined by the set of ORs activated by the set of volatile molecules they emit. Though the mechanism differs in many important ways, the perception of an olfactory tonality may be considered analogous to the perception of a single discernable musical note through the auditory system, discernable even when multiple notes are played at once. Similarly, where the same musical note may be described in multiple ways (F-sharp, G-flat, etc.), so the same olfactory tonality may be described in multiple ways.
[0087] Non-limiting examples of an olfactory tonality include earthy, coconut, celery, hay etc. (see Example 1 in US 62 / 911 ,096). For example, OR11 A1 agonists all share a common earthy tonality even as they elicit an overall set of distinct tonalities (see Example 2 in US 62 / 911 ,096). A volatile molecule may have greater than one tonality. For example, beta ionone has violet and blonde woody tonalities arising from the activation of OR5A1 and OR7A17, respectively (see Example 3 in US 62 / 911 ,096).
[0088] As used herein, the terms “perceived intensity” or “perceived psychophysical intensity” refer to the relationship between concentration of an ingredient and a quantitative value representing the intensity of perception of said ingredient or of a tonality of said ingredient.
[0089] As used herein, the terms “perceptual descriptor” refers to any word used to describe a smell or taste. Such a perceptual descriptor may be “fruity”, “vanilla”, or “woody” for example. Such perceptual descriptors may be associated with numerical values named “descriptor weights” which are ratings of how strong a perceptual descriptor is perceived for a molecule. The term perceptual descriptor may correspond to the term tonality, but may also include valence and other characteristics, for example.
[0090] Such a perceived intensity can be obtained by the measurement of the psychophysical intensity of an ingredient, for example, empirically by registering an input representative of perceived intensity by a panel of users. Such an input can be registered via a human-machine interface of any kind. Such input is stored into a registry. Preferably, this step of measurement is performed for various given gas-phase concentrations of the fragrance ingredient.
[0091] Such a perceived intensity is typically modelled by a sigmoid curve which corresponds to a sigmoid fit function of the measured individual values obtained from the panel of users, the sigmoid curve being represented by the following non-exhaustive parameters:
[0092] - theta, half-maximal intensity concentration,
[0093] - Imax, which refers to the maximum perceived intensity for a fragrance or flavor ingredient
[0094] - ODT, which refers to the “odor detection threshold” and which corresponds to the minimum gas phase concentration for which the tonality is perceived, and
[0095] - a curve parameter, which refers to a mathematical value which represents the slope of the sigmoid curve at concentration theta.
[0096] As used herein, a "fragrance" or “flavor” refers to the olfactory perception resulting from the sum of odorant receptor(s) activation, enhancement and inhibition (when present) by at least one volatile molecule. Accordingly, by way of illustration and by no means intending to limit the scope of the present disclosure, a "fragrance" results from the olfactory perception arising from the sum of a first volatile molecule that activates an OR associated with a coconut tonality, a second volatile molecule that activates an OR associated with a celery tonality, and a third volatile molecule that inhibits an OR associated with a hay tonality.
[0097] The inventors have found that the activation of one OR triggers the perceivability of one tonality in one-to-one relationship so that a composition, being an assembly of at least one tonality, can be encoded by a set of corresponding activated ORs, potentially with an associated activation level.
[0098] As used herein, a "note" or "olfactory note" or "perfumery note" identifies a tonality category. For example, floral notes include lily of the valley and violet tonalities.
[0099] The term "olfactory receptor" refers to one or more members of a family of G protein- coupled receptors (GPCRs) that are expressed in olfactory cells. This includes chemosensory GPCRs, including olfactory receptors (such as ORs (odorant receptors), TAARs (for trace amine- associated receptors), and so on). “OR” or “odorant receptor” polypeptides are further considered as such if they pertain to the 7-transmembrane-domain G protein-coupled receptor superfamily encoded by a single approximately 1 kb long exon and exhibit characteristic olfactory receptorspecific amino acid motifs. The seven domains are called "transmembrane" or "TM" domains TM I to TM VII connected by three "internal cellular loop" or "IC" domains IC I to IC III, and three "external cellular loop" or "EC" domains EC I to EC III. The motifs and the variants thereof are defined as, but not restricted to, the MAYDRYVAIC motif (SEQ ID NO: 7) overlapping TM III and IC II, the FSTCSSH motif (SEQ ID NO: 8) overlapping IC III and TM VI, the PMLNPFIY motif (SEQ ID NO: 9) in TM VII as well as three conserved C residues in EC II, and the presence of highly conserved GN residues in TM I [Zhang, X. & Firestein, S. Nat. Neurosci. 5, 124-133 (2002); Malnic, B., et al. Proc. Natl. Acad. Sci. U. S. A. 101 , 2584-2589 (2004)]. OSNs (olfactory sensory neurons) can also be identified on the basis of morphology or by the expression of proteins specifically expressed in olfactory cells. OR family members may have the ability to act as receptors for olfactory signal transduction.
[0100] An "agonist of an OR" or "agonist compound" refers to a volatile compound or a ligand that binds to an OR, activates the OR and induces an olfactory receptor transduction cascade.
[0101] An "antagonist of an OR" or "antagonist compound" refers to a volatile compound or a ligand that binds to an OR and reduces the measured activity of an olfactory receptor in the presence of an agonist.
[0102] An “inhibitor” refers to a compound or ligand that binds to an OR and that reduces the activity of the receptor through an olfactory receptor transduction cascade exposed to the agonist compound when compared to the activity of the receptor exposed to the agonist compound in the absence of the inhibitor. In one aspect, the compound that is not an agonist binds to a site on the receptor different from the site where the agonist compound binds.
[0103] An “enhancer” refers to a compound or ligand that binds to an OR and that enhances the activity of the receptor through an olfactory receptor transduction cascade exposed to the agonist compound when compared to the activity of the receptor exposed to the agonist compound in the absence of the enhancer. In one aspect, the compound that is not an agonist binds to a site on the receptor different from the site where the agonist compound binds.
[0104] "Potency" refers to the measure of the activity of a receptor induced by the binding of an agonist (odorant) in terms of amount or concentration required to produce a given activity level. It indicates the sensitivity of the receptor to different agonist concentrations and is often assessed by calculating the EC50 (the agonist concentration necessary to achieve half-maximal receptor activity).
[0105] "Efficacy" refers to the measure of the activity level by which the OR responds to a given agonist and is obtained by measuring the activation span between constitutive (baseline in the absence of an agonist) and agonist induced activity.
[0106] The “receptive field" of an odorant receptor refers to the range of compounds activating the receptors. It describes the set of distinct volatile compounds able to bind to the receptor, trigger a transduction cascade within a cell, and hence transform the chemical stimulus to an electrical signal via the olfactory sensory neuron projecting to the brain.
[0107] As used herein, the terms “fragrance ingredient” designates a perfuming ingredient, a flavor ingredient, an aroma ingredient, a perfumery carrier, a flavor carrier, a perfumery adjuvant, a flavor adjuvant, a perfumery modulator, flavor modulator. Preferably, such a fragrance or fragrant chemical compound is volatile. Such an ingredient may be a natural ingredient.
[0108] By “perfuming ingredient” it is meant here a compound, which is used in a perfuming preparation or a composition to impart a hedonic effect. In other words, such an ingredient, to be considered as being a perfuming one, must be recognized by a person skilled in the art as being able to impart or modify in a positive or pleasant way the odor of a composition, and not just as having an odor.
[0109] The nature and type of the perfuming ingredients do not warrant a more detailed description here, which in any case would not be exhaustive, the skilled person being able to select them on the basis of his general knowledge and according to the intended use or application and the desired organoleptic effect. In general terms, these perfuming co-ingredients belong to chemical classes as varied as alcohols, lactones, aldehydes, ketones, esters, ethers, acetates, nitriles, terpenoids, nitrogenous or sulphurous heterocyclic compounds and essential oils, and said perfuming coingredients can be of natural or synthetic origin. Said perfuming ingredients are in any case listed in reference texts such as the book by S. Arctander, Perfume and Flavor Chemicals, 1969, Montclair, New Jersey, USA, or its more recent versions, or in other works of a similar nature, as well as in the abundant patent literature in the field of perfumery. It is also understood that said perfuming ingredients may also be compounds known to release in a controlled manner various types of perfuming ingredients also known as properfume or profragrance.
[0110] By “perfumery carrier” it is meant here a material which is practically neutral from a perfumery point of view, i.e., that does not significantly alter the organoleptic properties of perfuming ingredients. Said carrier may be a liquid or a solid.
[0111] As liquid carrier one may cite, as non-limiting examples, an emulsifying system, i.e., a solvent and a surfactant system, or a solvent commonly used in perfumery. A detailed description of the nature and type of solvents commonly used in perfumery cannot be exhaustive. However, one can cite as non-limiting examples, solvents such as butylene or propylene glycol, glycerol, dipropyleneglycol and its monoether, 1 ,2,3-propanetriyl triacetate, dimethyl glutarate, dimethyl adipate 1 ,3-diacetyloxypropan-2-yl acetate, diethyl phthalate, isopropyl myristate, benzyl benzoate, benzyl alcohol, 2-(2-ethoxyethoxy)-1 -ethanol, tri-ethyl citrate or mixtures thereof, which are the most commonly used. For the compositions which comprise both a perfumery carrier and a perfumery base, other suitable perfumery carriers than those previously specified, can be also ethanol, water / ethanol mixtures, glycerol, limonene or other terpenes, isoparaffins such as those known under the trademark Isopar® (origin: Exxon Chemical) or glycol ethers and glycol ether esters such as those known under the trademark Dowanol® (origin: Dow Chemical Company), or hydrogenated castors oils such as those known under the trademark Cremophor® RH 40 (origin: BASF), esters and emollients such as the trademark Cetiol®, vegetable oils, essential oils.
[0112] Solid carrier is meant to designate a material to which the perfuming composition or some element of the perfuming composition can be chemically or physically bound. In general, such solid carriers are employed either to stabilize the composition, or to control the rate of evaporation of the compositions or of some ingredients. The use of solid carriers is of current use in the art and a person skilled in the art knows how to reach the desired effect. However, by way of non-limiting examples of solid carriers, one may cite absorbing gums or polymers or inorganic materials, such as porous polymers, cyclodextrins, wood-based materials, organic or inorganic gels, clays, gypsum talc or zeolites.
[0113] As other non-limiting examples of solid carriers, one may cite encapsulating materials. Examples of such materials may comprise wall-forming and plasticizing materials, such as mono, di- or trisaccharides, natural or modified starches, hydrocolloids, cellulose derivatives, polyvinyl acetates, polyvinylalcohols, proteins or pectins, or yet the materials cited in reference texts such as H. Scherz, Hydrokolloide: Stabilisatoren, Dickungs- und Geliermittel in Lebensmitteln, Band 2 der Schriftenreihe Lebensmittelchemie, Lebensmittelqualitat, Behr's Verlag GmbH & Co., Hamburg, 1996. The encapsulation is a well-known process to a person skilled in the art, and may be performed, for instance, by using techniques such as spray-drying, agglomeration or yet extrusion; or consists of a coating encapsulation, including coacervation and complex coacervation techniques.
[0114] As non-limiting examples of solid carriers, one may cite in particular the core-shell capsules with resins of aminoplast, polyamide, polyester, polyurea or polyurethane type or a mixture thereof (all of said resins are well known to a person skilled in the art) using techniques like phase separation process induced by polymerization, interfacial polymerization, coacervation or altogether (all of said techniques have been described in the prior art), optionally in the presence of a polymeric stabilizer or of a cationic copolymer.
[0115] Resins may be produced by the polycondensation of an aldehyde (e.g., formaldehyde, 2,2- dimethoxyethanal, glyoxal, glyoxylic acid or glycolaldehyde and mixtures thereof) with an amine such as urea, benzoguanamine, glycoluryl, melamine, methylol melamine, methylated methylol melamine, guanazole and the like, as well as mixtures thereof. Alternatively, one may use preformed resins alkylolated polyamines such as those commercially available under the trademark Urac® (origin: Cytec Technology Corp.), Cymel® (origin: Cytec Technology Corp.), Urecoll® or Luracoll® (origin: BASF).
[0116] Others resins one are the ones produced by the polycondensation of an a polyol, like glycerol, and a polyisocyanate, like a trimer of hexamethylene diisocyanate, a trimer of isophorone diisocyanate or xylylene diisocyanate or a Biuret of hexamethylene diisocyanate or a trimer of xylylene diisocyanate with trimethylolpropane (known with the tradename of Takenate®, origin: Mitsui Chemicals), among which a trimer of xylylene diisocyanate with trimethylolpropane and a Biuret of hexamethylene diisocyanate are preferred.
[0117] Some of the seminal literature related to the encapsulation of perfumes by polycondensation of amino resins, namely melamine-based resins with aldehydes includes represented by articles such as those published by K. Dietrich et al. Acta Polymerica, 1989, vol. 40, pages 243, 325 and 683, as well as 1990, vol. 41 , page 91 . Such articles already describe the various parameters affecting the preparation of such core-shell microcapsules following prior art methods that are also further detailed and exemplified in the patent literature. US 4’396'670, to the Wiggins Teape Group Limited is a pertinent early example of the latter. Since then, many other authors have enriched the literature in this field and it would be impossible to cover all published developments here, but the general knowledge in encapsulation technology is significant. More recent publications of pertinency, which disclose suitable uses of such microcapsules, are represented for example by the article of K. Bruyninckx and M. Dusselier, ACS Sustainable Chemistry & Engineering, 2019, vol. 7, pages 8041 -8054H.Y. Lee et al. Journal of Microencapsulation, 2002, vol. 19, pages 559-569, international patent publication WO 01 / 41915 or yet the article of S. Bone et al. Chimia, 2011 , vol. 65, pages 177-181 .
[0118] By “perfumery adjuvant,” it is meant here an ingredient capable of imparting additional added benefit such as a color, a particular light resistance, chemical stability, etc. A detailed description of the nature and type of adjuvant commonly used in perfuming composition cannot be exhaustive, but it has to be mentioned that said ingredients are well known to a person skilled in the art. One may cite as specific non-limiting examples the following: viscosity agents (e.g. surfactants, thickeners, gelling and / or rheology modifiers), stabilizing agents (e.g. preservatives, antioxidant, heat / light and or buffers or chelating agents, such as BHT), coloring agents (e.g. dyes and / or pigments), preservatives (e.g. antibacterial or antimicrobial or antifungal or anti irritant agents), abrasives, skin cooling agents, fixatives, insect repellants, ointments, vitamins and mixtures thereof.
[0119] By “perfumery modulator,” it is understood here an agent having the capacity to affect the manner in which the odor, and in particular the evaporation rate and intensity, of the compositions incorporating said modulator can be perceived by an observer or user thereof, over time, as compared to the same perception in the absence of the modulator. Perfumery modulators are also known as fixative. In particular, the modulator allows prolonging the time during which their fragrance is perceived. Non-limiting examples of suitable modulators may include methyl glucoside polyol; ethyl glucoside polyol; propyl glucoside polyol; isocetyl alcohol; PPG-3 myristyl ether; neopentyl glycol diethylhexanoate; sucrose laurate; sucrose dilaurate, sucrose myristate, sucrose palmitate, sucrose stearate, sucrose distearate, sucrose tristearate, hyaluronic acid disaccharide sodium salt, sodium hyaluronate, propylene glycol propyl ether; dicetyl ether; polyglycerin-4 ethers; isoceteth-5; isoceteth-7, isoceteth-10; isoceteth-12; isoceteth-15; isoceteth-20; isoceteth-25; isoceteth-30; disodium lauroamphodipropionate; hexaethylene glycol monododecyl ether; and their mixtures; neopentyl glycol diisononanoate; cetearyl ethylhexanoate; panthenol ethyl ether, DL- panthenol, N-hexadecyl n-nonanoate, noctadecyl n-nonanoate, a profragrance, cyclodextrin, an encapsulation, and a combination thereof.
[0120] By “flavoring ingredient” it is meant here a compound, which is used in flavoring preparations or compositions to impart a hedonic effect. In other words, such an ingredient, to be considered as being a flavoring one, must be recognized by a person skilled in the art as being able to impart or modify in a positive or pleasant way the taste of a composition, and not just as having a taste. The nature and type of the flavoring ingredients present in the composition do not warrant a more detailed description here, the skilled person being able to select them on the basis of its general knowledge and according to intended use or application and the desired organoleptic effect. In general terms, these flavoring ingredients belong to chemical classes as varied as alcohols, aldehydes, ketones, esters, ethers, acetates, nitriles, terpenoids, nitrogenous or sulphurous heterocyclic compounds and essential oils, and said flavoring ingredients can be of natural or synthetic origin. Many of these ingredients are in any case listed in reference texts such as the book by S. Arctander, Perfume and Flavor Chemicals, 1969, Montclair, New Jersey, USA, or its more recent versions, or in other works of a similar nature, as well as in the abundant patent literature in the field of flavor. It is also understood that said co-ingredients may also be compounds known to release in a controlled manner various types of flavoring compounds, also called proflavor.
[0121] The term “flavor carrier” designates a material which is substantially neutral from a flavor point of view, as far as it does not significantly alter the organoleptic properties of flavoring ingredients. The carrier may be a liquid or a solid.
[0122] Suitable liquid carriers include, for instance, an emulsifying system, i.e., a solvent and a surfactant system, or a solvent commonly used in flavors. A detailed description of the nature and type of solvents commonly used in flavor cannot be exhaustive. Suitable solvents used in flavor include, for instance, propylene glycol, triacetine, caprylic / capric triglyceride (neobee®), triethyl citrate, benzylic alcohol, ethanol, isopropanol, citrus terpenes, vegetable oils such as Linseed oil, sunflower oil or coconut oil, glycerol.
[0123] Suitable solid carriers include, for instance, absorbing gums or polymers, or even encapsulating materials. Examples of such materials may comprise wall-forming and plasticizing materials, such as mono, di- or polysaccharides, natural or modified starches, hydrocolloids, cellulose derivatives, polyvinyl acetates, polyvinylalcohols, xanthan gum, arabic gum, acacia gum or yet the materials cited in reference texts such as H. Scherz, Hydrokolloid : Stabilisatoren, Dickungs- und Geliermittel in Lebensmitteln, Band 2 der Schriftenreihe Lebensmittelchemie, Lebensmittelqualitat, Behr's VerlagGmbH & Co., Hamburg, 1996. Encapsulation is a well-known process to a person skilled in the art, and may be performed, for instance, using techniques such as spray-drying, agglomeration, extrusion, coating, plating, coacervation and the like.
[0124] By “flavor adjuvant”, it is meant here an ingredient capable of imparting additional added benefit such as a color (e.g., caramel), chemical stability, and so on. A detailed description of the nature and type of adjuvant commonly used in flavoring compositions cannot be exhaustive. Nevertheless, such adjuvants are well known to a person skilled in the art who will be able to select them on the basis of its general knowledge and according to intended use or application. One may cite as specific non-limiting examples the following: viscosity agents (e.g., emulsifier, thickeners, gelling and / or rheology modifiers, e.g., pectin or agar gum), stabilizing agents (e.g., antioxidant, heat / light and or buffers agents e.g., citric acid), coloring agents (e.g., natural or synthetic or natural extract imparting color), preservatives (e.g., antibacterial or antimicrobial or antifungal agents, e.g., benzoic acid), vitamins and mixtures thereof.
[0125] By “flavor modulator,” it is meant here an ingredient capable to enhance sweetness, to block bitterness, to enhance umami, to reduce sourness or licorice taste, to enhance saltiness, to enhance a cooling effect, or any combinations of the foregoing. Flavor modulators are also called trigeminal sensates.
[0126] Figure 1 represents, schematically, a particular succession of steps of an embodiment of the method 100 object of the present invention. This computer-implemented method 100 for providing at least one fragrance molecule representation, representing a materializable fragrance molecule for fragrance or flavor product production, which can be grouped into four phases: a database constitution phase, comprising: a step 150 of materially exposing an odorant receptor, corresponding to an exemplar odorant receptor representation, to a fragrance molecule, corresponding to an exemplar fragrance molecule representation, a step 155 of measuring odorant receptor activation values as a result of the material exposition, and a step 145 of recording, in a database, the empirically measured odorant receptor activation values, a machine learning device training phase, comprising: providing 135 an original set of: exemplar fragrance molecule representations, said set of exemplar fragrance molecule representations being representative of materialized fragrance molecules, and at least one associated fragrance molecule feature value, exemplar odorant receptor representation, representative of physically existing odorant receptors, and at least one associated odorant receptor feature value, exemplar odorant receptor activation values for said at least one said exemplar fragrance molecule representation and at least one exemplar odorant receptor representation, representative of empirically measured activation values caused by the material exposition of said exemplar odorant receptor to said exemplar fragrance molecule, and training 140 a machine-learning model using the training set, wherein the machinelearning model is trained to associate at least one feature value with an odorant receptor activation value, a fragrance ingredient representation generation phase, comprising: a step 105 of determining at least one value representative of an odorant receptor activity target caused by the material exposition of said odorant receptor to the fragrance molecule associated with a fragrance molecule representation, a step 110 of generating at least one fragrance molecule representation by a trained machine-learning model using as input value representative of an odorant receptor activity target, and a step 115 of providing at least one of the generated at least one fragrance molecule representation, a fragrance ingredient materialization phase, comprising: a step 175 of sending a digital command representative of an instruction of materialising at least one fragrance ingredient corresponding to at least one fragrance ingredient representation provided and / or a step 180 of materialising at least one fragrance ingredient corresponding to at least one fragrance ingredient representation provided.
[0127] The first, second and fourth phases are optional embodiments of the method subject of the invention.
[0128] The first phase relates to the constitution of a database based on empirical experimentation, which can be used for later training of a machine learning model.
[0129] During this phase, the following steps may be implemented: a step 150 of materially exposing an odorant receptor, corresponding to an exemplar odorant receptor representation, to a fragrance molecule, corresponding to an exemplar fragrance molecule representation, a step 155 of measuring odorant receptor activation values as a result of the material exposition, and a step 145 of recording, in a database, the empirically measured odorant receptor activation values.
[0130] The step 150 of materially exposing an odorant receptor to a fragrance molecule may be performed in a variety of ways which depends on the experimental protocol desired.
[0131] The step 155 of measuring odorant receptor activation values may be performed in a variety of ways which depends on the experimental protocol desired.
[0132] Such an experimental protocol is disclosed below.
[0133] It should be understood that in the context of the present invention, two key relationships may be used: a first relationship, linking fragrance molecule to the activation of ORs and / or a second relationship, linking the activation of ORs to specific tonalities.
[0134] In variants, two key relationships may be used: a first relationship, linking fragrance molecule to the activation of ORs and / or a second relationship, linking the activation of ORs to regions of perceptual space. Such a region of perceptual space can be defined in relative terms between two specific tonalities. For example, a region of perceptual space can correspond to a “somewhere between woody and earthy, but not necessarily either of those”. A region of perceptual space relates to a smell or taste situated in gaps in the capability to express the perception of a smell or taste by survey respondents.
[0135] A region of perceptual space can be modelled by creating a two- or three-dimensional representation of the perceptual space (through a dimensionality reduction algorithm or deep learning), and then using density estimates on that space to find the two-dimensional coordinates between terms that a receptor might be linked to. For example, in 2D space, OR5AU1 has strong projections to COCONUT, LACTONIC, and CREAMY, which are all in a specific region of space, however the true tonality of this receptor is somewhere in the space between these terms which does not have a word in language ontology to describe it.
[0136] Another approach is to create a higher dimensional representation of the space (such as through contrastive learning or variational autoencoders) and then performing a similar projection as above.
[0137] In particular, a given fragrance molecule may interact with several ORs, and one OR may interact with several fragrance molecules. However, the activity of one OR may correspond to one perceived tonality triggered by the activity of one given OR. The perception of a single compound exhibiting several tonalities is resulting from the activation of several ORs encoding said tonalities. This “one-to-one” relationship is not trivial to discover and has major implications in formula and fragrance design, thus emphasizing the importance of the present invention. Indeed, in systems in which the fragrance molecule to tonality is known only, it is impossible to accurately predict or estimate the tonality of a combination of fragrance molecules, given the lack of knowledge of the intermediate layer of interaction between molecule and OR on one side and OR to tonality on another.
[0138] The first relationship may be obtained by putting into contact a sample of fragrance molecules with isolated ORs and by monitoring the activity of each said OR.
[0139] The result of such an experiment may be stored inside a database cross referencing fragrance molecule representation, representative of real fragrance molecules, and the impact of said fragrance molecule on the ORs monitored during the experiment.
[0140] This can correspond to a first embodiment of the step 145 of recording.
[0141] In such methods, as long as the function of an OR is not impaired, it may be used in any form in a method or assay described herein. For example, the OR may be used as follows: tissues or cells which intrinsically express an OR such as olfactory sensory neurons isolated from organism and cultured products thereof; olfactory cell membrane bearing the OR; recombinant cells genetically modified so as to express the OR and cultured products thereof; membrane of the recombinant cells; and artificial lipid bilayer membrane carrying the OR. Indicators for monitoring the activity of olfactory receptors include, for example, a fluorescent calcium indicator dye, a calcium indicator protein (e.g. GCaMP, a genetically encoded calcium indicator), a fluorescent cAMP indicator, a cell mobilization assay, a cellular dynamic mass redistribution assay, a label-free cell based assay, a cAMP response element (CRE) mediated reporter protein, a biochemical cAMP HTRF assay, a beta-arrestin assay, or an electrophysiological recording. In a particular embodiment, a calcium indicator dye is selected that can be used to monitor the activity of olfactory receptors expressed on the membrane of the olfactory neurons (e.g., Fura-2 AM). Molecules may be screened sequentially and the odorant-dependent changes in calcium dye fluorescence are measured using a fluorescent microscope or fluorescent-activated cell sorter (FACS).
[0142] Such methods allow the identification of commonalities in olfactory perception between chemically and organoleptically diverse volatile compounds based on their OR activation profile. Without intending to be limited to any particular theory, OR activation profiles represent better predictors of olfactory tonality for a given volatile compound, compared to physicochemical similarities alone. Without intending to be limited to any particular theory, physiochemical similarity may only partially predict olfactory similarity.
[0143] As an example, olfactory neurons activated by the target agonist, e.g. a molecule with a particular olfactive tonality, are isolated using either a glass microelectrode attached to a micromanipulator or a FACS machine. Mouse olfactory sensory neurons are screened by Ca2+ imaging similar to procedures previously described. Malnic, B., et al. Cell 96, 713-723 (1999); Araneda, R. C. et al. J. Physiol. 555, 743-756 (2004); and WO2014 / 210585. Particularly, a motorized movable microscope stage is used to increase the number of cells that can be screened to at least 1 ,500 per experiment. Since there are approximately 1 ,200 different olfactory receptors in the mouse and each olfactory sensory neuron expresses only 1 of 1 ,200 olfactory receptor genes, this screening capacity will cover virtually the entire mouse odorant receptor repertoire. In other words, the combination of calcium imaging for high-throughput olfactory sensory neuron screening leads to the identification of nearly all of the odorant receptors that respond to a particular profile of odorants. In one aspect, odorant receptors that respond to the target agonist, e.g. with a particular olfactive tonality, can be isolated for receptor identification. For example, at least one neuron is isolated for receptor identification.
[0144] The second relationship may be obtained by associating at least one olfactive tonality with an OR by screening chemically and organoleptically diverse libraries of volatile compounds against the OR to identify the at least one olfactive tonality common to the volatile compounds that are activators of the screened ORs.
[0145] Such a method for associating at least one olfactory tonality to an olfactory receptor can be obtained by providing an OR, contacting the OR and a fragrance molecule having a known at least one tonality, determining whether the compound activates the at least one olfactory receptor, repeating the contacting and determining steps with different compounds having known at least one tonality, categorizing a subset of compounds that activate the OR, identifying an at least one known tonality common to the subset of compounds, and associating the identified at least one known tonality to the OR.
[0146] The common olfactive tonality of compounds activating a given receptor can be associated by comparing the overall description of the compounds and identifying the descriptor that is common among activators. It should be understood that such an olfactive tonality can be semantically described, e.g., by a perfumer, in several ways. Examples of such semantic similarities capturing the same olfactive tonality include for example: marine, watery and ozone; earthy, humus and moss; hay, coumarinic and tonka; celery, fenugreek and maple; muguet and lily-of-the-valley.
[0147] The probability of an ingredient activating a receptor to a given amount can be determined through statistical approaches, such as but not limited to, Bayesian inference and probabilistic distribution models. An example workflow is to gather descriptors for a large chemical library, screen the library against a receptor library, then compute at a series of activity index cutoffs a beta distribution model of the likelihood of descriptors being present in agonists. In such an approach, at least three scenarios can be observed: no meaningful difference from the screening set as the cutoff criteria increases, a convergence onto a tonality (as the cutoff increases the probability approaches 1 ), and an anti-correlation with a tonality (as cutoff increases the probability approaches 0). These models can also be sampled from in an inference step.
[0148] The common olfactive tonality of compounds activating a given receptor can also be determined via explainable Al techniques such as shown in figure 13.
[0149] Figure 13 shows, schematically, a particular embodiment of the method 1300 object of the present invention. This method 1300 for providing at least one odorant receptor representation, representing a physiological odorant receptor, comprises: a step of defining 1305 a set of at least two perceptual descriptor representations and, for each said perceptual descriptor representation, a value representative of a relative weight of said descriptor in relation to at least one other defined descriptor, a step of predicting 1310 at least one activated odorant receptor representation by a trained machine-learning model using as input value the defined set and corresponding relative weight values, and a step of providing 1315 at least one of the predicted at least one activated odorant receptor representation.
[0150] The step of defining 1305 can be performed via any input device or API such as shown in relation to figure 2.
[0151] The step of generating 1310 is performed, for example, by executing instructions representative of a computer software upon a computing device, such as the one shown in figure 2. During this step of generating 1310, the defined set and corresponding values are used as input for a trained machine-learning model which, in return, provides a list of at least one odorant receptor representation. Such odorant receptor representations correspond to odorant receptors activated when the perceptual descriptors are felt.
[0152] In particular variants, the step of generating 1310 is also generates at least one odorant receptor activation value, representative of the strength of the activation of said odorant receptor when a particular perceptual descriptor is felt.
[0153] The step of providing 1315 can be performed via any output device or API such as shown in relation to figure 2.
[0154] In particular embodiments, the method 1300 further comprises at least one of the following steps: providing 1301 an original set of:
[0155] - weighted perceptual descriptors for individual fragrance compounds, odorant receptor activation values, classified into high or low, for said individual fragrance compounds or molecules, rebalancing 1302 the dataset to introduce a bias in favor of either the class of high or low odorant receptor activation values, training 1303 at least one random forest machine learning model to predict high / low odorant receptor activation from the weighted perceptual descriptors of said compounds as a function of the rebalanced original set, downstream of a step of predicting, determining 1311 the importance of perceptual descriptor features, as a function of the SHAP values for the predicted values, and / or assessing 1312 the contribution of a perceptual descriptor to at least one odorant receptor activation as a function of aggregated SHAP values across compounds or molecules for said perceptual descriptor.
[0156] The step of rebalancing 1302 can correspond to oversampling the positive class of high activation fragrance compounds by generating synthetic examples with the SMOTE method or any other method suited to modify the balance of a dataset.
[0157] The step of determining 1311 may further compute the average absolute value of the SHAP values across compounds to determine the importance of each of the descriptor features, where the highest value or highest few values describe the receptor’s tonality. Such a computation may involve methods such as the TreeSHAP method from Lundberg, S.M., Erion, G., Chen, H. et al. From local explanations to global understanding with explainable Al for trees. Nat Mach Intell 2, 56-67 (2020).
[0158] The step of assessing 1312 may further compute, for a given descriptor feature, the average SHAP value across fragrance compounds that are not described with that feature (do not smell like that feature) and calculating the average SHAP value across fragrance compounds that are described with that feature. These reveal when a descriptor contributes to tonality (if having or not having that descriptor leads to increased likelihood of selecting high receptor activation) or when the descriptor is anti-correlated with the receptor’s activation (when having the descriptor leads to higher likelihood of the model selecting low / no activation for the receptor).
[0159] The result of such an experiment may be stored inside a database cross referencing OR representations, representative of real ORs, and information representative of the tonalities associated with the activation of these ORs.
[0160] This can correspond to a second embodiment of the step 145 of recording.
[0161] It should be noted that a pre-processing phase can be included between the first and second phases.
[0162] The number and nature of the preprocessing steps vary based on the particular use-case. Below, several examples are provided.
[0163] For any numerical representation of input / output at least one of the following preprocessing steps may be implemented:
[0164] - applying a mathematical transform (log, square root, etc.),
[0165] - normalizing and / or scaling based on representative statistics of the dataset (e.g. Min-Max Scaling, Normalization, etc.),
[0166] - categorical binning / encoding in which continuous numerical representations are converted to discrete categorical values including one-hot, nominal and ordinal scales, and
[0167] - imputation of unknown values based on statistical or heuristic criteria.
[0168] For fragrance molecule or ingredient input, at least one of the following preprocessing steps may be implemented:
[0169] - gathering fragrance molecule or ingredient representations (name, chemical registry number, SMILES (for “Simplified Molecular Input Line Entry System”)) of said ingredients, and / or
[0170] - converting fragrance molecules or ingredient representation to the appropriate machine representation (Tokenized SMILES, graph, 3d coordinates, fingerprint, physiochemical property).
[0171] Molecular fingerprints refer to unique digital representations of a molecule's chemical structure and properties used to compare and classify molecules.
[0172] For OR input, at least one of the following preprocessing steps may be implemented:
[0173] - gathering OR amino acid sequences data, possibly but not necessarily performing multiple sequence alignment,
[0174] - gathering three-dimensional structures (crystal, cryo-EM, ML-generated) data for said OR,
[0175] - obtaining latent embeddings from ML pretrained models for said OR and / or features of said OR, and / or
[0176] - classifying OR activity into strong (high) vs weak (low) activation.
[0177] A latent embedding refers to a representation of data in a lower-dimensional space that captures some aspects of the original data. The term “latent” suggests that this representation uncovers hidden patterns or structures within the data that might not be immediately apparent in its original form.
[0178] Cryo-EM structures refer to an experimentally determined structure of a biomolecule from cryo-electron microscopy. Commonly used for GPCR transmembrane proteins.
[0179] For tonality input / output, at least one of the following preprocessing steps may be implemented:
[0180] - gathering fragrance molecules or ingredient tonality descriptors and their respective weights, such as available from the GoodScents, FlavorNet and EPA databases, and building a matrix of descriptors as columns and ingredients as rows populated with weights as values / 0 if not that descriptor, and scaled between 0 and 1 ,
[0181] - creating a map based on at least part of the descriptors gathered, to a predefined ontology created, such a map may contain INPUT, OUTPUT, and SOURCE fields, where INPUT refers to the raw descriptors, either fine-grained and specific terms or more broad family level descriptors, from various sources, OUTPUT refers to a predefined ontology of selected terms, and SOURCE defines from where the term in the INPUT field originated - said map may be created manually, automatically through the use of natural language processing and / or large language models and / or character similarity metrics such as the Levenshtein Distance, or a combination of all those previously mentioned,
[0182] - running all of the descriptors through the map to go from freeform description to an organized description, and / or
[0183] - associating organized descriptors with fragrance molecule or ingredient representations (SMILES or IUPAC (for “International Union of Pure and Applied Chemistry”) for example) - When multiple fragrance molecule or ingredient are associated with the same descriptor at different weights (where the weights correspond to how strongly a particular descriptor is associated with a particular fragrance molecule or ingredient), a preference system can be set up, a mathematical operation can be performed (such as calculating the mean of weights for example), or a default value can be assigned.
[0184] For intensity input / output, at least one of the following preprocessing steps may be implemented:
[0185] - gathering raw input values representative of perceived intensity by a panel of users across a range of concentrations or at a single concentration from a database
[0186] - gathering sigmoid-curve approximation fit parameters, including but not limited to ODT, Imax, theta and curve parameter from a database
[0187] The method 100 object of the present invention may further comprise a step 1400 of determining receptor to perceptual descriptor associations, such as shown in figure 14, which comprises the steps of:
[0188] - collecting 1405 a set of at least one perceptual descriptor, - for at least one collected perceptual descriptor, screening 1410 at least one olfactory receptor, represented by an olfactory receptor representation in a computing system, determining 1415 the impact of an ingredient on at least one olfactory receptor, optionally inferring 1420, and providing 1425 an association of at least one collected perceptual descriptor to at least one odorant receptor, said at least one association being used at least during the step of determining 105.
[0189] The step of collecting 1405 may be performed by using any input device such as shown in view of figure 2.
[0190] The step of screening 1410 may be performed by exposing an ingredient to olfactory receptors and measuring the output in terms of perceptual descriptors. Such outputs and associations may be stored in a computer memory.
[0191] The step of determining 1415 may be performed by taking the ingredients screened against a receptor and put said ingredients into different bins depending on how well each ingredient activates the receptor.
[0192] Then for each bin, the Bayesian probability (as a beta probability distribution function) is calculated for "if an ingredient activates a receptor at this level or higher, what is the probability it will have tonality X?". Then, any descriptor (or all of them) can be compared on how that probability distribution changes from the screening set and across each bin. What can be seen is that most tonalities show no real difference across these levels, whereas one or two will show increasing probabilities (and likely converge to a near certainty at the highest activation bin). What can also be seen is the existence of some anti-correlations where the probability decreases to almost no possibility for that descriptor to appear in the highest bin.
[0193] The step of inferring 1420 can be performed by using Bayesian inference to sample from the distribution of each bin to bolster the confidence of the output.
[0194] The step of providing 1425 can be performed, for example, using any output device or API to store values in a computer memory or to show said values on a computer screen for example.
[0195] In particular embodiments, the step 1400 of determining further comprises a step of subsetting data 1411 prior to the step of determining 1415.
[0196] The method 100 object of the present invention may further comprise, prior to the step 140 of providing an original set: a step 160 of generating a conformer for at least one exemplar fragrance molecule representation, a step 165 of computing a three-dimensional point cloud for at least one conformer and of at least one electrostatic potential value for at least one point of the volumetric point cloud. An electrostatic point cloud designates a volumetric surface of a molecule representing the electrostatic potential of a molecule such that each point is represented by four values: three- 1 dimensional coordinates and a value corresponding to the electrostatic potential associated to these three-dimensional coordinates.
[0197] In such embodiments, from a given molecule representation (SMILES, 3d coordinate file, RDKit mol object) conformers are generated. The generation of the conformers can be obtained using any tool to generate them including but not limited to molecular dynamics, software package designed for conformer generation, random coordinate sampling, from databases. In practice, a molecule may have many conformers. Thus, to limit data size, conformers can be clustered using a variety of algorithms to reduce the total number of conformers. Furthermore, conformers can be energy minimized by using traditional force field or quantum mechanics (QM), but this is not necessary.
[0198] Figures 3 to figure 7 further disclose examples of such embodiments.
[0199] Figure 3 represents, schematically, a particular succession of steps of a method using electrostatic point cloud representations to allow for scaffold hopping, fragrance or flavor ingredient comparison or fragrance or flavor ingredient attribute prediction.
[0200] Such a method comprises a step 400 of data gathering, which comprises a step 410 of collecting data, from a database of text, two-dimensional, three-dimensional and / or fourdimensional representations for a set of materialized fragrance or flavor ingredients, to form a training set.
[0201] An example of such data is represented in figure 11 , which represents machine output for a 3D conformer of ethanol molecule (SMILES: CCO). One common example of a 3D output for a molecule consists of x, y, and z coordinates for each individual atom in a molecule. One such file type is XYZ file type shown as Figure 11 . The first line is the total number of atoms in the molecule. The second line is a text-based comment line that can contain any relevant information about this molecule. In this scenario, the text is the SMILES version of the molecule. The subsequent N lines are the individual 3-dimensional coordinates of each atom labeled with the chemical symbol. 3D conformers can also be in other file formats including PDB, SDF, MOL, and many others.
[0202] This method further comprises a step 500 of obtaining an electrostatic surface point cloud for a fragrance or flavor ingredient.
[0203] This step 500 of obtaining comprises:
[0204] - a step 415 of generating a conformer neural network device to associate a three- dimensional representation for at least one part (or point) of a two-dimensional fragrance or flavor ingredient digital representation based on the training set, and
[0205] - a step 420 of computing an electrostatic surface point cloud for a three-dimensional representation obtained from executing the conformer (either on the training set or on other two-dimensional representations of fragrance or flavor ingredient).
[0206] It should be noted that the step 415 of training and the step 420 of computing an electrostatic surface point cloud may not be in immediate succession. In particular embodiments, the step 415 of training is discarded and the step 420 of computing an electrostatic surface point cloud is based on at least one readily available three-dimensional representation of a fragrance or flavor ingredient. This readily available three-dimensional representation may be collected from a database in a dedicated step (not represented herein). Such a database may be populated by associating fragrance or flavor ingredient representations (corresponding to materialized fragrance or flavor ingredients) with values representative of the measurable three-dimensional structure of said fragrance or flavor ingredient.
[0207] The step 415 of generating can be performed, for example, by executing instructions representative of a computer software on a computing device 300, such as the one shown in figure 2. Such a conformer can be obtained from empirical and physics-based models which are readily available.
[0208] It should be noted that a conformer is a conformation of a molecule that lies at any point in the potential energy diagram, such as the minimum. The step 420 of computing an electrostatic surface point cloud can be performed, for example, by executing instructions representative of a computer software on a computing device 300, such as the one shown in figure 2. During this step 420 of computing, a quantum mechanics calculation may be performed.
[0209] This step 420 of computing is agnostic to how the point cloud is generated as long as the output corresponds to a volume representation of a molecule which includes values for parameter representative of electrostatics of the corresponding materialized or materializable fragrance or flavor ingredient.
[0210] In particular embodiments, the step 420 of computing performs quantum mechanics calculations, using deep learning methods that are surrogate to quantum mechanics calculations, such as disclosed in Practical High-Quality Electrostatic Potential Surfaces for Drug Discovery Using a Graph-Convolutional Deep Neural Network, by Prakash Chandra Rathi et al. published in J. Med. Chem. 2020, 63, 16, 8778-8790.
[0211] In such embodiments, the input for the deep learning method can correspond to a 3-d representation similar to what is presented in figure 11 . This input requires x, y, z coordinates of the atoms in a molecule. The output of such a deep learning model can be directly the 4-d electrostatic point cloud or it can be the multipole (monopole, dipole, quadrupole, and / or octuple). These multipoles can then be used to calculate the electrostatic potential surface.
[0212] The calculated point clouds resulting from the step 420 of computing can correspond to a 4-d set of points (array) consisting of the x, y, z coordinates of each point on the point cloud as well as one additional point corresponding to the electrostatic potential at that point. A point cloud can additionally be described by another three points called the surface normal. Furthermore, point clouds can also be represented as meshes and voxels, which have been used for 3D computer vision applications.
[0213] Figure 4 shows the different stages for a particular flavor or fragrance ingredient, in which a two-dimensional digital representation 401 of a flavor or fragrance ingredient is conformed into a three-dimensional digital representation 402 of said flavor or fragrance ingredient, which, after the step 420 of computing, corresponds to an electrostatic surface point cloud representation 403 for said flavor or fragrance ingredient.
[0214] Figure 3 further shows a step 600 of obtaining a neural network model based on the electrostatic point clouds generated. Such a step 600 of obtaining may be performed in order to obtain a trained neural network model which is used during a step 700 of comparing, a step 800 of predicting or a step 900 of scaffold hopping.
[0215] Such a step 600 of obtaining may comprise a step 425 of converting the electrostatic point cloud generated into a graph structure representation. Such a graph structure representation corresponds, for example, to a collection of nodes (or sometimes called vertices) and edges (connecting nodes). One common data structure for graphs consist of at least two arrays:
[0216] - a feature array of size (# nodes, # features) which captures information about the nodes, and
[0217] - an edge index array of size (# of edges, 2) corresponding to the connectivity of nodes.
[0218] Additionally, graph objects can contain edge features arrays, positional information, and other features that may be relevant to the graph.
[0219] It should be noted that graph structure representation corresponds to an input that can be used both for GNNs and Transformer encoding architectures.
[0220] In standard SMILES to graph conversions, each node corresponds to the individual atom (or groups of atoms) in a molecule. Node features would consist of the atomic identity (e.g. C, H, O, N, etc.), aromaticity, charge, chirality or other atom-centric properties. Edges are typically, but not exclusively, the bonds which connect the atoms. Additionally, the edge features may be attributed to the type of bond (e.g. single, double, triple, aromatic, etc.).
[0221] Unlike the standard SMILES to graph conversions, the 4D electrostatic point cloud conversion to graph consists of nodes corresponding to the individual points of the point cloud. Each electrostatic value is one such feature value of the node. Additionally, each node has a corresponding 3D position. Edges are connections between nodes. Since there are no intrinsic “bonds” between electrostatic nodes, the graph can be considered fully connected (in which all nodes are connected to each other), edgeless (in which there are no predefined edges), or edges can be selectively chosen (through down sampling methods as disclosed below).
[0222] The step 425 of converting can be performed, for example, by executing instructions representative of a computer software on a computing device 300, such as the one shown in figure 2.
[0223] For example, during the step 425 of converting, the 4-d electrostatic point clouds is further processed into a graph structure. The graph structure can be a fully connected graph in which all points are connected to one another, an empty graph in which no points are connected, the points can be randomly connected, or additional algorithms can be used to connect points. In practice, using fully connected graphs is computationally expensive. One such algorithm used to reduce computational overhead is a mixed k-Nearest Neighbor / n-Medoids algorithm. In this algorithm, each point is connected to k of its distance-based nearest neighbors (k, for example, can correspond to the number ten). Furthermore, the n-Medoids algorithm can down-sample N clusters whose centers are points on the point cloud. These N cluster centers are then all fully connected to one another. In such variants, the number N can be a value that is equal to 20% of the total number of points in the point cloud. Additional down sampling methods include but are not limited to furthest point sampling, commonly used in computer vision, in which a subsample of nodes is selected that maximize the overall distances between nodes.
[0224] Alternatively, a graph-based structure does not need to have pre-assigned edges when used in specific deep learning architectures, for example transformer architectures. In this scenario, the graph structure can have learnable edges based on an attention mechanism. The attention mechanism determines how each node “attends” to another.
[0225] The step 600 of obtaining may further comprise a step 430 of augmenting the converted graphs by increasing or reducing the number of points in the graph (a process known as “density augmentation”).
[0226] Such a step 430 of augmenting can be performed, for example, by executing instructions representative of a computer software on a computing device 300, such as the one shown in figure 2.
[0227] During this step 430 of augmenting, focuses on the total number of points (which can be called “P”). During this step 430 of augmenting, a “fuzziness” is created in the electrostatic point cloud. A similar comparison is performing pixelation on a 2D image. An image can be taken and either the number of pixels is reduced, which makes the picture fuzzy, or the number of pixels is increased which gives it more definition. In this regard, the density augmentation can be considered a 3D pixelation process.
[0228] During this step 430 of augmenting, additional steps can be taken including random global rotations of the point cloud, masking clusters of points within the point cloud, or adding random noise to either the coordinates or features, or both.
[0229] The step 600 of obtaining may further comprise a step 435 of encoding a graph representation of an electrostatic point cloud representation.
[0230] Such a step 435 of encoding can be performed, for example, by executing instructions representative of a computer software on a computing device 300, such as the one shown in figure 2.
[0231] For example, during this step 435 of encoding, using graphs (augmented or not) as input, an Equivariant Graph Neural Network (“EGNN”) can be used as a machine learning model. These are special GNNs that are equivariant to rotations, reflections, and translations, which are useful for three-dimensional representations of molecules. In particular variants, modifications can be made to break reflection equivariance (for molecular chirality purposes).
[0232] Alternatively, during this step 435 of encoding, transformer architectures consisting of an attention mechanism can be used as a machine learning model using graphs (augmented or not and with or without edge information) as input. The transformer architecture can include one or many transformer layers consisting of alternating attention modules and feedforward modules. The attention module includes an attention weight matrix, which dictates how each node pair attends to one another. Additionally, the attention weight matrix can include an additional bias term (known as attention bias) that can incorporate edge information which can either be spatial or learned edge encoding. Figure 17 is a representative example of the transformer layer.
[0233] The step 600 of obtaining may further comprise a step of aggregating to form a single latent space invariant of the number of nodes in the graph representation.
[0234] Such a step of aggregating can be performed, for example, by executing instructions representative of a computer software on a computing device 300, such as the one shown in figure 2.
[0235] During this step of aggregating, a mathematical operation (sum, max, mean, or attentionbased mechanism) can be employed to aggregate the features of the N nodes to ensure that graphs of different sizes.
[0236] The step 600 of obtaining may further comprise a step 440 of contrastive training of a machine learning device.
[0237] Such a step 440 of contrastive learning can be performed, for example, by executing instructions representative of a computer software on a computing device 300, such as the one shown in figure 2.
[0238] The machine learning paradigm can be based on contrastive learning. Contrastive learning is a self-supervised learning method that creates a learned embedding space such that positive samples are close to one another, and negative samples are farther away in embedding space. There are many methods to perform contrastive learning. This workflow adopts the SimCLR approach, but in principle other methods can be used. The SimCLR approach assumes a data augmentation module in which samples are augmented by some transformation function. Augmented samples that came from the same “parent” sample are considered positive pairs. Augmented samples that are from different “parents” are negative pairs. In the case of point clouds two augmentations were proven useful: conformer augmentation (e.g. conformers from the same molecule are considered negative samples) and density augmentation (the number of points in the point cloud). This density augmentation is a novel augmentation in the contrastive learning paradigm.
[0239] The result from training the model in the contrastive setting results in an embedding space that distinguishes molecules based on their molecular size and electrostatics. Thus, smaller polar molecules are segmented from larger hydrophobic molecules in the embedding space. This embedding space is now extremely valuable for many downstream tasks, especially in relation to ligand binding and mixture modelling.
[0240] As it is understood, the method 100 object of the present invention may further comprise a step 170 of converting at least one three-dimensional point cloud into a graph structure representation, said graph structure representation being used as at least one fragrance molecule feature value.
[0241] Figure 5A further illustrates the conversion from electrostatic point cloud representation to graph representation. In figure 5A, two electrostatic point clouds, 505 and 510, are converted, each, into two separate graph representations (with different numbers of nodes and of nearest neighbors), 506, 507, 511 and 511 , and groups graph representations are associated with either a positive (“+”) or negative pairing.
[0242] Figure 5B further shows the step 435 of encoding and the step 440 of contrastive learning disclosed below.
[0243] The encoding strategy can consist of a set of one or more graph convolution layers in which information between nodes is passed to its neighbors (connected by edges). Alternatively, the encoding strategy can consist of a set of one or more graph-transformer layers in which information between nodes is passed across all nodes or a subset of nodes depending on the spatial encoding and masking.
[0244] Following message-passing steps (graph convolutions) or transformer layers, a graph aggregation step 505 aggregates all the node information using one or more aggregation functions. The aggregation functions may include a sum, average, max, or attention-based aggregation. Finally, a contrastive embedding network can be used to embed the graphs and create an embedding space such that positive samples are closer to each other while negative samples are further apart. In this simple example, the dark gray square has a similar embedding to the light gray square and dark gray diamond. Meanwhile the black triangle and white pentagon are further away in the embedding space. Similar to these simple shapes, molecules have their own shape and color corresponding to the volume and electrostatics of the point cloud.
[0245] The step 700 of comparing may be performed to determine a similarity between fragrance or flavor ingredients.
[0246] The step 800 of predicting may be performed to predict receptor activation, odor or taste tonality, odor or taste intensity, biodegradability and so on.
[0247] The step 900 of scaffold hopping may be performed to generate novel fragrance or flavor ingredients which share a common base structure (the “scaffold”).
[0248] The method object of the present invention may comprise a second, possibly independent, phase relates to the training of a machine learning model which learns to associate odorant receptor - ingredient interactions based on features relative to the odorant receptors and / or relative to the ingredients.
[0249] The second phase comprises the steps of: providing 135 an original set of: exemplar fragrance molecule representations, said set of exemplar fragrance molecule representations being representative of materialized fragrance molecules, and at least one associated fragrance molecule feature value, exemplar odorant receptor representation, representative of physically existing odorant receptors, and at least one associated odorant receptor feature value, exemplar odorant receptor activation values for said at least one said exemplar fragrance molecule representation and at least one exemplar odorant receptor representation, representative of empirically measured activation values caused by the material exposition of said exemplar odorant receptor to said exemplar fragrance molecule, and training 140 a machine-learning model using the training set, wherein the machinelearning model is trained to associate at least one feature value with an odorant receptor activation value.
[0250] While the data used in this second phase are discussed as originating from the first phase, it is obvious that another set of data may be used, based on other experimental protocols.
[0251] It should be noted that there exist significant possibilities to execute this second phase. Below a series of embodiments is presented.
[0252] The step 135 of providing may be performed, for example, via extraction of the original set from at least one database. Such an extraction may be performed via an input device 240 or via a network link 255, connected to the database, such as shown in figure 2. It should be noted that the data may be stored directly on a computer executing instructions representative of the step 140 of training or be accessible via a communication network, such as the Internet or a local area network for example.
[0253] The original set comprises exemplar fragrance molecule representations, which correspond, for example, to the name, chemical registry number, SMILES, IUPAC of said fragrance molecule. The fragrance molecule representations may be all or in part replaced by ingredient representations.
[0254] The original set of exemplar fragrance molecule representations may further comprise values representative of features of said fragrance molecule such as:
[0255] - another machine representation, such as:
[0256] - a tokenized SMILES representation,
[0257] - a two-dimensional or three-dimensional graph representation,
[0258] - a three-dimensional coordinate representation,
[0259] - at least one perceptual descriptor associated with said fragrance molecule,
[0260] - a molecular fingerprint,
[0261] - a measurable or measured physico-chemical value for said fragrance molecule,
[0262] - a machine learning or artificial intelligence based embedding representation, - a three-dimensional point cloud representation associated with at least one electrostatic potential value.
[0263] - a three-dimensional conformer ensemble representation, and / or
[0264] - a mixture representation, which can correspond to either a mixture of different ingredients or to an ingredient comprising different molecules,
[0265] - at least one physico-chemical property value,
[0266] - at least one value representative of an analytical spectra.
[0267] Analytical spectra could be a Mass Spectra (GC-MS or LC-MS), an NMR spectrum, or other analytical chemistry instrumentation that can assign a molecular signature of either a single molecule or complex mixture.
[0268] Physico-chemical properties refer to characteristics of a molecule such as size, shape, polarity, Henry’s constant, and charge that determine its interactions with other molecules and biological systems.
[0269] The original set also comprises exemplar odorant receptor representations, which correspond, for example, to the name of said OR. Such names of ORs can be derived from the GenelD, chromosome locus location.
[0270] The original set of exemplar odorant receptor representations may further comprise values representative of features of said odorant receptor, such as:
[0271] - an amino acid sequence representation,
[0272] - binding site locations for said odorant receptor representations,
[0273] - an amino acid sequence alignment matrix representation,
[0274] - a three-dimensional structure (crystal, cryo-EM, ML-generated) representation,
[0275] - a multi-state three-dimensional structural model representation, which correspond to different activation states of a receptor and are obtained by crystal, cryo-EM, ML-models, and so on - unlike single state, multi-state models distinguish between activation and inhibition of receptor (among other possible states), and / or
[0276] - a latent embedding from machine learning pretrained models for said OR.
[0277] The original set also comprises exemplar odorant receptor activation values associated with at least one exemplar fragrance molecule representation and at least one odorant receptor representation.
[0278] Receptor activity of receptor activation refers to the functional response of a receptor or protein to the binding of a ligand or other stimulus, including activation, inhibition, and enhancement.
[0279] The odorant receptor activation values may correspond to:
[0280] - raw dose response curve values, EC50 and span (affinity / efficacy) or the activity index (mathematical combination thereof), which refer to the amount of activation seen in a dose response assay - essentially, how much activation of the OR was seen at the highest concentration of the ingredient versus how much baseline activation is there of the OR if there is no ingredient present,
[0281] - dose response curve approximation fit parameters, and / or
[0282] - predicted parameters from binding models (agonism and modulation) - which correspond to fit biophysical models describing how groups of ligands compete to bind / activate / inhibit an OR to dose response data for individual ingredients and for simple mixtures and then get parameters per ingredient-OR combination that can be plugged into the biophysical model to predict OR activity to an arbitrary mixture including that ingredient. Binding models can also be derived from molecular docking or other molecular simulation techniques.
[0283] A dose response curve corresponds to a curve for which:
[0284] - an axis represents increasing gas phase concentration of a chemical compound (typically modeled in logarithmic scale),
[0285] - an axis represents increasing perceived psychophysical intensity,
[0286] - for given gas phase concentrations, a sample of perceived intensity by a pool of users, and
[0287] - a sigmoid curve fitting the sample distribution for given gas phase concentrations.
[0288] Dose response curve approximation fit parameters refer to numerical values used to approximate a mathematical curve to experimental data to determine the relationship between biological (or indicator signal) activity and concentration.
[0289] Binding model parameters refer to numerical values used to fit a mathematical model to experimental (measured) data to determine the strength and affinity of a ligand-receptor interaction, likely through the use of simultaneous fitting of multiple ligands at once to determine appropriate parameters.
[0290] It should be noted that input concatenation may be performed for multi-modal models. Such a concatenation may involve information from different types of inputs (chemical and / or biological) and / or information from a single type of input (multiple representations of the same molecule for example). Such a concatenation may use:
[0291] - direct concatenation of inputs,
[0292] - one type of input as a node embedding for a graph structure, and / or
[0293] - cross-attention mechanisms between two inputs.
[0294] Additionally, multi-modal models can arise from creating a latent embedding space in a pretraining step. This is obtained from, for example, SMILES and ORs, which are used to learn representations of both chemicals and receptors (using contrastive learning) in a manner where ORs will have similar latent representations to the fragrance or flavor ingredients that activate said ORs.
[0295] An example of such a multi-modal model 1000 is shown in figure 8, in which 1 to n modalities are used as input, each modality being independently encoded, and a step 1005 of contrastive learning being performed on these independently encoded modalities to generate 1010 a multi-modal embedding. The step 140 of training may be performed by Graph Neural Networks (GNNs), Equivariant GNNs which are used for 3d graphs to ensure equivariance upon rotations and translations, Transformer architectures which are used for string representations, embedding dimensions, a Multi-layer Perceptron (Feed-Forward Neural Network), or a transformer mechanism which handles both conformers and mixture modelling, such as shown in figure 10.
[0296] Figure 9 represents, schematically, a succession of steps of a particular use of the method 1100 of the present invention comprising the use of variational autoencoders and density estimates to represent perceptual space.
[0297] This method 1100 comprises:
[0298] - a step 1105 of retrieving perceptual descriptors and similarities as defined in a descriptor map, such as shown in figure 12, and a rating system which quantifies the relationships between tonalities,
[0299] - a step 1110 of training a variational graph auto encoder, this step of training allowing to treat descriptions as vectors in a regularized latent space that is informed by how perceptually similar tonalities are,
[0300] - a step of 1120 using these latent representations 1115
[0301] - a step 1125 of building density estimates to allow interactions with perceptual neighborhoods 1130 instead of with specific terms (for example, “HAY” and “TONKA” are perceptually related and will have similar representations in the latent space).
[0302] Figure 10 represents an example machine learning method 1200 for multiple instance learning (MIL).
[0303] When dealing with complex mixtures (either difficult to isolate individual components or may contain tens to hundreds of components), the property that is experimentally determined is a complex aggregation of all these ingredients, which is called the “bag” ground truth. In most paradigms (especially in regard to flavor or fragrance ingredients), knowing the individual ground truth (“instance”) is difficult to obtain. This difficulty arises due to the difficulty in isolating individual components of a mixture or requiring many individual experiments for each component. Therefore, the Multiple Instance Learning process is a method to use the “bag” information to potentially access the “instance” information. This learning paradigm will identify patterns in the data to successfully predict both the bag and instance ground truths.
[0304] This can be very useful in identifying the key components in a mixture that lead to the perceived property.
[0305] In this scenario, a mixture of fragrance ingredient representations and conformers are treated as an unordered sequence. This unordered sequence with an additional start token (CLS) is processed by a transformer encoder 1205.
[0306] The now embedded CLS token is passed to a Multi-Layer Perceptron (MLP) 1210 to make a ’’bag” prediction 1215, which corresponds to the end point property including receptor activity, tonality, intensity, etc. This bag prediction 1215 is meant to represent the aggregate of all ingredient representations and conformers. Each individual ingredient representation and / or conformer is passed through an ’’instance” MLP 1220 for property prediction. In most cases, the exact value is unknown for the individual instance. Therefore, an attention-weighted property prediction 1225 can be used to predict the global property.
[0307] Furthermore, during processing, learning strategies may be employed, such as:
[0308] - standard classification or regression learning which includes multi-class and / or multi-label,
[0309] - contrastive learning for pretraining data, or
[0310] - multiple instances learning for conformer ensemble property prediction or mixture modelling.
[0311] Contrastive learning refers to a self-supervised learning paradigm that focuses on extracting meaningful representations by contrasting positive and negative pairs of instances. It assumes that similar instances should be closer in a learned embedding space while dissimilar instances should be farther apart.
[0312] In particular advantageous embodiments, such as shown in figure 6A, an EGNN is used on three-dimensional electrostatics point clouds in a contrastive learning approach to create a novel embedding space. This embedding space can then be used by machine learning algorithms including but not limited to neural network architectures, tree-based models, and linear models to predict receptor activity, tonality, intensity and physiochemical properties for a fragrance ingredient associated with the three-dimensional electrostatics point clouds.
[0313] Figure 6B further shows the results for the architecture shown in figure 6A for classification and regression tasks compared to a 2D-GNN method, popularly used for molecular property prediction. For receptor activity, the accuracy increases using the electrostatic point clouds indicating better performance. In predicting the log P of a set of fragrance molecules, the electrostatic point cloud has a lower Root Mean Squared Error (RMSE) indicating improved performance. In these first two scenarios, neural network architectures were used as the machine learning algorithm. The last graph demonstrates improved performance for predicting the maximum perceivable intensity for fragrance ingredients. The machine learning algorithm used in this last scenario was a Random Forest Classifier to distinguish high and low-intensity fragrance ingredients.
[0314] In particular advantageous embodiments, the above-mentioned embedding space is used in a transformer architecture with multiple instances learning to predict downstream tasks.
[0315] The method object of the present invention may comprise a third, possibly independent, phase which relates to the use of the trained machine learning model to generate ingredient representations.
[0316] This third phase comprises the following steps:
[0317] - a step 105 of determining at least one value representative of an odorant receptor activity target caused by the material exposition of said odorant receptor to the fragrance molecule associated with a fragrance molecule representation, - a step 110 of generating at least one fragrance molecule representation by a trained machine-learning model using as input value representative of an odorant receptor activity target, and
[0318] - a step 115 of providing at least one of the generated at least one fragrance molecule representation.
[0319] The step 105 of determining at least one value representative of an odorant receptor activity target may be performed using an input device 240 of a computing system such as shown in figure 2. Such a determination may result from a choice of a user or by an external command received from another computing system, via an API for example.
[0320] For example, during this step 105 of determining, a user may select ORs individually and associate, for each OR, a positive or negative value corresponding to the odorant receptor activity target for this OR.
[0321] In particular embodiments, the method 100 object of the present invention further comprises a step 120 of defining at least one determined tonality value, the step 110 of determining at least one value representative of an odorant receptor activity target being defined as a function of said at least one defined determined tonality value.
[0322] In such embodiments, a user may select tonalities (corresponding to specific OR activation and deactivation) individually and associate, for each tonality, a positive or negative value which is converted into an odorant receptor activity target for at least one OR associated with the tonality. Such a conversion can be performed via a conversion database obtained via empirical measurement of OR activation values to tonality perception value correspondence. Such empirical measurements are disclosed above.
[0323] The step 110 of generating is performed by using the trained machine learning model with, as an input, the at least one determined value representative of an odorant receptor activity target.
[0324] In particular embodiments, during the step 110 of generating, a value representative of a quantity is generated for at least one fragrance molecule represented by a fragrance molecule representation.
[0325] Such a quantity may be in absolute (weight or volume) or relative (ppm, percentage of total weight) terms.
[0326] Using such a trained machine learning model leads, for example, to obtaining a Matthew’s Correlation Coefficient of 0.59.
[0327] Such embodiments require that, during the step 140 of training, data representative of fragrance molecule quantities is associated with at least some exemplar fragrance molecule representations in the training set.
[0328] In particular embodiments, such as shown in figures 1 and 7A, the step 110 of generating further generates a scaffold representation, representing a materializable molecular scaffold, associated with at least one generated fragrance molecule representation, said method further comprising: a step 125 of selecting a scaffold representation and a step 130 of regenerating at least one fragrance molecule representation by a trained machine-learning model using as input value representative of an odorant receptor activity target and the selected scaffold representation, and
[0329] - the step 115 of providing further providing said regenerated fragrance molecule representation.
[0330] The step 125 of selecting can be performed, for example, by user selection of a scaffold representation upon a GUI.
[0331] The step 130 of regenerating can be performed similarly to the step 110 of generating, with the added constraint of restricting fragrance ingredients which are associated with the selected scaffold representation.
[0332] Figure 7A further shows a particular embodiment of the step 130 of regenerating, which comprises: a step 905 of encoding said fragrance or flavor ingredient representation into an embedding space, a step 910 of decoding said embedding space into a representation for a fragrance or flavor molecule, and a step 915 of training machine learning model to encode-decode.
[0333] Figure 7B shows a particular machine learning training architecture that can be used to obtain an encode-decode machine learning model. This machine learning architecture is based on a transformer architecture that decodes molecular SMILES representations conditioned on the encoder of the electrostatic point clouds. In this example scenario, the electrostatic point clouds of representative molecules are encoded into an embedding space. This embedding space is then used as a condition for the transformer decoder, which is trained to reproduce the SMILES string of the representative molecules. Akin to popular Large Language Models (LLM), the encoded embedding space can be considered a “query” or “question” to the LLM and the ’’response” or “answer” is the SMILES string. Therefore, with this method, one can query molecules of a specific shape and electrostatic signature and use the decoder to suggest molecules that fit said signature.
[0334] Figure 7C depicts example point clouds of two molecules. Scaffold hopping corresponds to the process of going from the benzene ring to the furan ring. These are common bioisosteres used in medicinal chemistry because they both impart similar bioactivity. As one can see, the electrostatic points clouds demonstrate similar features.
[0335] The step 115 of providing is performed, for example, by using an output device 235 of a computing system such as shown in figure 2. Such a step 115 of providing may result in the display of the generated fragrance molecule identifier in a GUI or result in the emission of an external command received to another computing system or to a fragrance molecule assembly device.
[0336] An empirically validated generated fragrance molecule identifier can be obtained, for example, by: - training a machine learning model such as shown in figure 4, for ingredient SMILES representations and amino acid sequences as OR receptor data, in a multi-classification mode between activation classes “no interaction”, “inverse agonism” or “agonism” and modulation classes “no modulatory effects”, “enhancement”, “inhibition”,
[0337] - OR10J5 was discovered to be linked to Lily of the Valley percept or tonality,
[0338] - An ingredient representation associated with OR10J5 activation (based on classification prediction),
[0339] - This ingredient was 2,5,6-TRIMETHYL-2-INDANMETHANOL and the activation of OR10J5 was empirically validated, and the percept of this ingredient was found to contain “lyral” and “muguet” as descriptors.
[0340] Between the second or third phase and the fourth phase, a post-processing phase may be implemented. This post-processing phase may comprise at least one of the following steps:
[0341] - a step of visualization which may comprise:
[0342] - a step of dimensionality reduction of learned embedding spaces, and
[0343] - a step of highlighting attention mechanisms to understand what inputs attends to which output,
[0344] - a step of exploration which may comprise:
[0345] - a step of using embedding spaces as means of similarity or distance scoring between input samples, and
[0346] - a step of optimization which may comprise:
[0347] - using multi-objective optimization algorithms (such as Pareto optimization) to tradeoff multiple tasks / targets for an optimal solution.
[0348] The fourth phase relates to the materialization of a fragrance molecule represented by a fragrance molecule representation generated.
[0349] This fourth phase comprises:
[0350] - a step 175 of sending a digital command representative of an instruction of materialising at least one fragrance ingredient corresponding to at least one fragrance ingredient representation provided and / or
[0351] - a step 180 of materialising at least one fragrance ingredient corresponding to at least one fragrance ingredient representation provided.
[0352] The step 175 of sending a digital command may be performed by any means of communication of a computing system 300 such as shown in figure 2. The recipient of this digital command may be an automated fragrance molecule materialization device.
[0353] The step 180 of materialising may be performed by an automated fragrance molecule materialization device.
[0354] A fragrance molecule materialization device corresponds to any device suited for the production of a fragrance molecule. Such a device can correspond to an automated chemical synthesis robotic system. This device can also correspond to any well-known synthesis devices regularly used in the flavor and fragrance industry.
[0355] Such embodiment may offer several routes to support new ingredient discovery, such as but not limited to previously unused or undiscovered fragrance molecules the field of fragrance design that:
[0356] - are confirmed or predicted to activate / enhance / inhibit desired / undesired ORs,
[0357] - are likely to generate a desired odor quality or perceptual tonality,
[0358] - exhibit a determined intensity for a perceptual tonality,
[0359] - are likely to counteract an undesired odor quality or perceptual tonality, and / or
[0360] - assist in chemists’ approach to new ingredient discovery by offering the ability of successful “scaffold hopping”.
[0361] Figure 2 represents, a block diagram of a device 300 for providing at least one fragrance molecule representation, representing a materializable fragrance molecule for fragrance or flavor product production, which comprises: means 240 of determining at least one value representative of an odorant receptor activity target caused by the material exposition of said odorant receptor to the fragrance molecule associated with a fragrance molecule representation, means 230 of generating at least one fragrance molecule representation by a trained machine-learning model using as input value representative of an odorant receptor activity target, and means 235 of providing at least one of the generated at least one fragrance molecule representation.
[0362] Figure 2 further represents a block diagram that illustrates an example computer system 300 with which an embodiment of the present invention may be implemented. In the example of figure 8, a computer system 205 and instructions for implementing the disclosed technologies in hardware, software, or a combination of hardware and software, are represented schematically, for example as boxes and circles, at the same level of detail that is commonly used by persons of ordinary skill in the art to which this disclosure pertains for communicating about computer architecture and computer systems implementations.
[0363] The computer system 205 includes an input / output (IO) subsystem 220 which may include a bus and / or other communication mechanism(s) for communicating information and / or instructions between the components of the computer system 205 over electronic signal paths. The I / O subsystem 220 may include an I / O controller, a memory controller and at least one I / O port. The electronic signal paths are represented schematically in the drawings, for example as lines, unidirectional arrows, or bidirectional arrows.
[0364] At least one hardware processor 210 is coupled to the I / O subsystem 220 for processing information and instructions. Hardware processor 210 may include, for example, a general-purpose microprocessor or microcontroller and / or a special-purpose microprocessor such as an embedded system or a graphics processing unit (GPU) or a digital signal processor or ARM processor. Processor 210 may comprise an integrated arithmetic logic unit (ALU) or may be coupled to a separate ALU.
[0365] Computer system 205 includes one or more units of memory 225, such as a main memory, which is coupled to I / O subsystem 220 for electronically digitally storing data and instructions to be executed by processor 210. Memory 225 may include volatile memory such as various forms of random-access memory (RAM) or other dynamic storage device. Memory 225 also may be used for storing temporary variables or other intermediate information during execution of instructions to be executed by processor 210. Such instructions, when stored in non-transitory computer-readable storage media accessible to processor 210, can render computer system 205 into a specialpurpose machine that is customized to perform the operations specified in the instructions.
[0366] Computer system 205 further includes non-volatile memory such as read only memory (ROM) 230 or other static storage device coupled to the I / O subsystem 220 for storing information and instructions for processor 210. The ROM 230 may include various forms of programmable ROM (PROM) such as erasable PROM (EPROM) or electrically erasable PROM (EEPROM). A unit of persistent storage 215 may include various forms of non-volatile RAM (NVRAM), such as FLASH memory, or solid-state storage, magnetic disk, or optical disk such as CD-ROM or DVD- ROM and may be coupled to I / O subsystem 220 for storing information and instructions. Storage 215 is an example of a non-transitory computer-readable medium that may be used to store instructions and data which when executed by the processor 210 cause performing computer- implemented methods to execute the techniques herein.
[0367] The instructions in memory 225, ROM 230 or storage 215 may comprise one or more sets of instructions that are organized as modules, methods, objects, functions, routines, or calls. The instructions may be organized as one or more computer programs, operating system services, or application programs including mobile apps. The instructions may comprise an operating system and / or system software; one or more libraries to support multimedia, programming or other functions; data protocol instructions or stacks to implement TCP / IP, HTTP or other communication protocols; file format processing instructions to parse or render files coded using HTML, XML, JPEG, MPEG or PNG; user interface instructions to render or interpret commands for a graphical user interface (GUI), command-line interface or text user interface; application software such as an office suite, Internet access applications, design and manufacturing applications, graphics applications, audio applications, software engineering applications, educational applications, games or miscellaneous applications. The instructions may implement a web server, web application server or web client. The instructions may be organized as a presentation layer, application layer and data storage layer such as a relational database system using structured query language (SQL) or no SQL, an object store, a graph database, a flat file system or other data storage. Computer system 205 may be coupled via I / O subsystem 220 to at least one output device 235. In one embodiment, output device 235 is a digital computer display or Human Machine Interface. Examples of a display that may be used in various embodiments include a touchscreen display or a light-emitting diode (LED) display or a liquid crystal display (LCD) or an e-paper display. Computer system 205 may include other type(s) of output devices 235, alternatively or in addition to a display device. Examples of other output devices 235 include printers, ticket printers, plotters, projectors, sound cards or video cards, speakers, buzzers or piezoelectric devices or other audible devices, lamps or LED or LCD indicators, haptic devices, actuators, or servos.
[0368] At least one input device 240 is coupled to I / O subsystem 220 for communicating signals, data, command selections or gestures to processor 210. Examples of input devices 240 include touchscreens, microphones, still and video digital cameras, alphanumeric and other keys, keypads, keyboards, graphics tablets, image scanners, joysticks, clocks, switches, buttons, dials, slides.
[0369] Another type of input device is a control device 245, which may perform cursor control or other automated control functions such as navigation in a graphical interface on a display screen, alternatively or in addition to input functions. Control device 245 may be a touchpad, a mouse, a trackball, or cursor direction keys for communicating direction information and command selections to processor 210 and for controlling cursor movement on display 235. The input device may have at least two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), that allows the device to specify positions in a plane. Another type of input device is a wired, wireless, or optical control device such as a joystick, wand, console, steering wheel, pedal, gearshift mechanism or other type of control device. An input device 240 may include a combination of multiple different input devices, such as a video camera and a depth sensor.
[0370] In another embodiment, computer system 205 may comprise an Internet of things (loT) device in which one or more of the output device 235, input device 240, and control device 245 are omitted. Or, in such an embodiment, the input device 240 may comprise one or more cameras, motion detectors, thermometers, microphones, seismic detectors, other sensors or detectors, measurement devices or encoders and the output device 235 may comprise a special-purpose display such as a single-line LED or LCD display, one or more indicators, a display panel, a meter, a valve, a solenoid, an actuator or a servo.
[0371] Computer system 205 may implement the techniques described herein using customized hard-wired logic, at least one ASIC or FPGA, firmware and / or program instructions or logic which when loaded and used or executed in combination with the computer system causes or programs the computer system to operate as a special-purpose machine. According to one embodiment, the techniques herein are performed by computer system 205 in response to processor 210 executing at least one sequence of at least one instruction contained in main memory 225. Such instructions may be read into main memory 225 from another storage medium, such as storage 215. Execution of the sequences of instructions contained in main memory 225 causes processor 210 to perform the process steps described herein. In alternative embodiments, hard-wired circuitry may be used in place of or in combination with software instructions.
[0372] The term “storage media” as used herein refers to any non-transitory media that store data and / or instructions that cause a machine to operate in a specific fashion. Such storage media may comprise non-volatile media and / or volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as storage 215. Volatile media includes dynamic memory, such as memory 225. Common forms of storage media include, for example, a hard disk, solid state drive, flash drive, magnetic data storage medium, any optical or physical data storage medium, memory chip, or the like.
[0373] Storage media is distinct from but may be used in conjunction with transmission media. Transmission media participates in transferring information between storage media. For example, transmission media includes coaxial cables, copper wire and fiber optics, including the wires that comprise a bus of I / O subsystem 220. Transmission media can also take the form of acoustic or light waves, such as those generated during radio-wave and infrared data communications.
[0374] Various forms of media may be involved in carrying at least one sequence of at least one instruction to processor 210 for execution. For example, the instructions may initially be carried on a magnetic disk or solid-state drive of a remote computer. The remote computer can load the instructions into its dynamic memory and send the instructions over a communication link such as a fiber optic or coaxial cable or telephone line using a modem. A modem or router local to computer system 205 can receive the data on the communication link and convert the data to a format that can be read by computer system 205. For instance, a receiver such as a radio frequency antenna or an infrared detector can receive the data carried in a wireless or optical signal and appropriate circuitry can provide the data to I / O subsystem 220 such as place the data on a bus. I / O subsystem 220 carries the data to memory 225, from which processor 210 retrieves and executes the instructions. The instructions received by memory 225 may optionally be stored on storage 215 either before or after execution by processor 210.
[0375] Computer system 205 also includes a communication interface 260 coupled to bus 220. Communication interface 260 provides a two-way data communication coupling to network link(s) 265 that are directly or indirectly connected to at least one communication network, such as a network 270 or a public or private cloud on the Internet. For example, communication interface 260 may be an Ethernet networking interface, integrated-services digital network (ISDN) card, cable modem, satellite modem, or a modem to provide a data communication connection to a corresponding type of communications line, for example an Ethernet cable or a metal cable of any kind or a fiber-optic line or a telephone line. Network 270 broadly represents a local area network (LAN), wide-area network (WAN), campus network, internetwork, or any combination thereof. Communication interface 260 may comprise a LAN card to provide a data communication connection to a compatible LAN, or a cellular radiotelephone interface that is wired to send or receive cellular data according to cellular radiotelephone wireless networking standards, or a satellite radio interface that is wired to send or receive digital data according to satellite wireless networking standards. In any such implementation, communication interface 260 sends and receives electrical, electromagnetic, or optical signals over signal paths that carry digital data streams representing various types of information.
[0376] Network link 265 typically provides electrical, electromagnetic, or optical data communication directly or through at least one network to other data devices, using, for example, satellite, cellular, Wi-Fi, or BLUETOOTH technology. For example, network link 265 may provide a connection through a network 270 to a host computer 250.
[0377] Furthermore, network link 265 may provide a connection through network 270 or to other computing devices via internetworking devices and / or computers that are operated by an Internet Service Provider (ISP) 275. ISP 275 provides data communication services through a world-wide packet data communication network represented as Internet 280. A server computer 255 may be coupled to Internet 280. Server 255 broadly represents any computer, data center, virtual machine, or virtual computing instance with or without a hypervisor, or computer executing a containerized program system such as DOCKER or KUBERNETES. Server 255 may represent an electronic digital service that is implemented using more than one computer or instance and that is accessed and used by transmitting web services requests, uniform resource locator (URL) strings with parameters in HTTP payloads, API calls, app services calls, or other service calls. Computer system 205 and server 255 may form elements of a distributed computing system that includes other computers, a processing cluster, server farm or other organization of computers that cooperate to perform tasks or execute applications or services. Server 255 may comprise one or more sets of instructions that are organized as modules, methods, objects, functions, routines, or calls. The instructions may be organized as one or more computer programs, operating system services, or application programs including mobile apps. The instructions may comprise an operating system and / or system software; one or more libraries to support multimedia, programming or other functions; data protocol instructions or stacks to implement TCP / IP, HTTP or other communication protocols; file format processing instructions to parse or render files coded using HTML, XML, JPEG, MPEG or PNG; user interface instructions to render or interpret commands for a graphical user interface (GUI), command-line interface or text user interface; application software such as an office suite, Internet access applications, design and manufacturing applications, graphics applications, audio applications, software engineering applications, educational applications, games or miscellaneous applications. Server 255 may comprise a web application server that hosts a presentation layer, application layer and data storage layer such as a relational database system using structured query language (SQL) or no SQL, an object store, a graph database, a flat file system or other data storage.
[0378] Computer system 205 can send messages and receive data and instructions, including program code, through the network(s), network link 265 and communication interface 260. In the Internet example, a server 255 might transmit a requested code for an application program through Internet 280, ISP 275, local network 270 and communication interface 260. The received code may be executed by processor 210 as it is received, and / or stored in storage 215, or other non-volatile storage for later execution.
[0379] The execution of instructions as described in this section may implement a process in the form of an instance of a computer program that is being executed and consisting of program code and its current activity. Depending on the operating system (OS), a process may be made up of multiple threads of execution that execute instructions concurrently. In this context, a computer program is a passive collection of instructions, while a process may be the actual execution of those instructions. Several processes may be associated with the same program; for example, opening up several instances of the same program often means more than one process is being executed. Multitasking may be implemented to allow multiple processes to share processor 210. While each processor 210 or core of the processor executes a single task at a time, computer system 205 may be programmed to implement multitasking to allow each processor to switch between tasks that are being executed without having to wait for each task to finish. In an embodiment, switches may be performed when tasks perform input / output operations, when a task indicates that it can be switched, or on hardware interrupts. Time-sharing may be implemented to allow fast response for interactive user applications by rapidly performing context switches to provide the appearance of concurrent execution of multiple processes simultaneously. In an embodiment, for security and reliability, an operating system may prevent direct communication between independent processes, providing strictly mediated and controlled inter-process communication functionality.
Claims
1. CLAIMS1. Computer-implemented method for providing at least one fragrance, flavor or aroma molecule representation, representing a materializable fragrance molecule for fragrance, flavor or aroma product production, which comprises: a step of determining at least one value representative of an odorant receptor activity target caused by the material exposition of said odorant receptor to the fragrance molecule associated with a fragrance molecule representation, a step of generating at least one fragrance molecule representation by a trained machinelearning model using as input value representative of an odorant receptor activity target, and a step of providing at least one of the generated at least one fragrance molecule representation.
2. Computer-implemented method according to claim 1 , wherein, during the step of generating, a value representative of a quantity is generated for at least one fragrance molecule represented by a fragrance molecule representation.
3. Computer-implemented method according to claim 1 , which further comprises a step of defining at least one determined tonality value and / or intensity, the step of determining at least one value representative of an odorant receptor activity target being defined as a function of said at least one defined determined tonality value and / or intensity.
4. Computer-implemented method according to claim 3, which further comprises a step of determining receptor to perceptual descriptor associations, which comprises the steps of:- collecting a set of at least one perceptual descriptor,- for at least one collected perceptual descriptor, screening at least one olfactory receptor, represented by an olfactory receptor representation in a computing system,- determining the impact of an ingredient on at least one olfactory receptor, and providing an association of at least one collected perceptual descriptor to at least one odorant receptor, said at least one association being used at least during the step of determining.
5. Computer-implemented method according to claim 1 , wherein the step of generating further generates a scaffold representation, representing a materializable molecular scaffold, associated with at least one generated fragrance molecule representation, said method further comprising: a step of selecting a scaffold representation anda step of regenerating at least one fragrance molecule representation by a trained machinelearning model using as input value representative of an odorant receptor activity target and the selected scaffold representation, and- the step of providing further providing said regenerated fragrance molecule representation.
6. Computer-implemented method according to claim 1 , which comprises the steps of: providing an original set of: exemplar fragrance molecule representations, said set of exemplar fragrance molecule representations being representative of materialized fragrance molecules, and at least one associated fragrance molecule feature value, exemplar odorant receptor representation, representative of physically existing odorant receptors, and at least one associated odorant receptor feature value, exemplar odorant receptor activation values for said at least one said exemplar fragrance molecule representation and at least one exemplar odorant receptor representation, representative of empirically measured activation values caused by the material exposition of said exemplar odorant receptor to said exemplar fragrance molecule, and training a machine-learning model using the training set, wherein the machine-learning model is trained to associate at least one feature value with an odorant receptor activation value, wherein said trained machine-learning model is used during the step of generating.
7. Computer-implemented method according claim 6, which comprises a step of recording, in a database, empirically measured odorant receptor activation values caused by the material exposition of an odorant receptor, corresponding to an exemplar odorant receptor representation, to a fragrance molecule, corresponding to an exemplar fragrance molecule representation, said database being used during the step of providing an original set.
8. Computer-implemented method according claim 7, which comprises: a step of materially exposing an odorant receptor, corresponding to an exemplar odorant receptor representation, to a fragrance molecule, corresponding to an exemplar fragrance molecule representation and a step of measuring odorant receptor activation values as a result of the material exposition.
9. Computer-implemented method according to claim 5, in which at least one fragrance molecule feature value is representative of: at least one perceptual descriptor associated with said fragrance molecule, a molecular fingerprint of said fragrance molecule,a measurable or measured physico-chemical value for said fragrance molecule, a machine-learning representation embedding, a three-dimensional conformer ensemble representation, a mixtures representation, or an analytical spectrum of said fragrance molecule.
10. Computer-implemented method according to claim 6, in which at least one fragrance molecule feature value is representative of a three-dimensional point cloud representation associated with at least one electrostatic potential value.
11. Computer-implemented method according to claim 10, which comprises prior to the step of providing an original set: a step of generating a conformer for at least one exemplar fragrance molecule representation, a step of computing of a three-dimensional point cloud for at least one conformer and of at least one electrostatic potential value for at least one point of the volumetric point cloud.
12. Computer-implemented method according to claim 11 , which comprises a step of converting at least one three-dimensional point cloud into a graph structure representation, said graph structure representation being used as at least one fragrance molecule feature value.
13. Computer-implemented method according to claim 6, in which at least one odorant receptor feature value is representative of: a multi-state three-dimensional model of at least one exemplar odorant receptor, a machine learning latent representation of at least one exemplar odorant receptor, and / or an odorant receptor activation level as a function of fragrance molecule concentration values.
14. Computer-implemented method according to claim 1 , which further comprises a step of sending a digital command representative of an instruction of materialising at least one fragrance ingredient corresponding to at least one fragrance ingredient representation provided.
15. Computer-implemented method according to claim 14, which further comprises a step of materialising at least one fragrance ingredient corresponding to at least one fragrance ingredient representation provided.
16. Computer-implemented method to train a machine learning model to provide at least one fragrance ingredient representation, representing a materializable fragrance ingredient, which comprises the steps of: providing an original set of: exemplar fragrance molecule representations, said set of exemplar fragrance molecule representations being representative of materialized fragrance molecules, and at least one associated fragrance molecule feature value, exemplar odorant receptor representation, representative of physically existing odorant receptors, and at least one associated odorant receptor feature value, exemplar odorant receptor activation values for said at least one said exemplar fragrance molecule representation and at least one exemplar odorant receptor representation, representative of empirically measured activation values caused by the material exposition of said exemplar odorant receptor to said exemplar fragrance molecule, and training a machine-learning model using the training set, wherein the machine-learning model is trained to associate at least one feature value with an odorant receptor activation value.
17. Computer program product which comprises instructions which upon execution by a computer cause the computer to execute the method according to any one of claims 1 to 16.
18. Computer-readable storage medium storing programming instructions which upon execution by a computer cause the computer to execute the method according to any one of claims 1 to 16.
19. Device for providing at least one fragrance molecule representation, representing a materializable fragrance molecule for fragrance or flavor product production, which comprises: means of determining at least one value representative of an odorant receptor activity target caused by the material exposition of said odorant receptor to the fragrance molecule associated with a fragrance molecule representation, means of generating at least one fragrance molecule representation by a trained machinelearning model using as input value representative of an odorant receptor activity target, and means of providing at least one of the generated at least one fragrance molecule representation.
20. Computer-implemented method for providing at least one odorant receptor representation, representing a physiological odorant receptor, which comprises:a step of defining a set of at least two perceptual descriptor representations and, for each said perceptual descriptor representation, a value representative of a relative weight of said descriptor in relation to at least one other defined descriptor, a step of predicting at least one activated odorant receptor representation by a trained machine-learning model using as input value the defined set and corresponding relative weight values, and a step of providing at least one of the predicted at least one activated odorant receptor representation.21 . Method according to claim 20, which further comprises at least one of the steps of: providing an original set of:- weighted perceptual descriptors for individual fragrance compounds, odorant receptor activation values, classified into high or low, for said individual fragrance compounds or molecules, rebalancing the dataset to introduce a bias in favor of either the class of high or low odorant receptor activation values, training at least one random forest machine learning model to predict high / low odorant receptor activation from the weighted perceptual descriptors of said compounds as a function of the rebalanced original set, downstream of a step of predicting, determining the importance of perceptual descriptor features, as a function of the SHAP values for the predicted values, and / or assessing the contribution of a perceptual descriptor to at least one odorant receptor activation as a function of aggregated SHAP values across compounds or molecules for said perceptual descriptor.
Citation Information
Patent Citations
Process for the production of microcapsules
US4396670A
Protein markers associated with bone marrow stem cell differentiation into early progenitor dendritic cells
US60629110P0
Method for producing microcapsules which carry cationic charges
WO2001041915A1
Methods of identifying, isolating and using odorant and aroma receptors
WO2014210585A2
Method for predicting presence or absence of aroma properties or olfactory receptor activation properties in substance
EP4130736A1