Digital smell reformulation
The method addresses the challenge of limited labeled data in digital odor formulation by utilizing weakly-supervised learning and data augmentation to optimize odor inventory, achieving efficient and accurate odor reformulation.
Patent Information
- Application Number
- PCT/IL2025/050472
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-10
- Filing Date
- 2025-06-03
- Publication Date
- 2025-12-11
AI Technical Summary
The effectiveness of applying machine learning to digital odor formulation/reformulation is limited by the lack of labeled data required for training and running the models, making it time-consuming and labor-intensive.
A computerized method using weakly-supervised learning, smell-related data augmentation, and unlabeled data properties to generate similarity and additivity models, optimizing odor inventory to a small set of primary odor items for efficient odor reformulation.
Enables efficient odor reformulation by reducing the need for extensive labeled datasets, optimizing odor item substitutions, and ensuring accurate digital odor reproduction.
Smart Images

Figure IL2025050472_11122025_PF_FP_ABST
Abstract
Description
DIGITAL SMELL REFORMULATIONCROSS-REFERENCES TO RELATED APPLICATIONS
[0001] The present application claims benefit from US Provisional Application No. 63 / 655,123 filed on June 3, 2024, and US Provisional Application No, 63 / 658,018 filed on June 10, 2024, incorporated hereby by reference in their entirety.TECHNICAL FIELD
[0002] The presently disclosed subject matter relates to digital olfactory techniques and, more particularly, to digital olfactory’ techniques involving machine learning models.BACKGROUND
[0003] Digital olfactory is an emerging field that aims to capture, store, recognize and reproduce scents through digital means. Digital olfactory / technologies enable generating, transmitting, and receiving odor-enabled digital media usable for communication, gaming, virtual reality, extended reality, e-commerce, automotive and other applications. Likewise, digital olfactory is usable for a wide range of other applications, for example detection of diseases through breath, air quality surveillance, food and beverage processing, fragrance engineering, etc.
[0004] An odor can be characterized by a detailed recipe (referred to also as “odor formula(s)” or “formula(s)”) informative of a list of ingredients and their ratio in a mixture. Odors’ formulas are fundamental to the development and effectiveness of digital olfactory technologies.
[0005] A process of modifying the composition of odor mixture while aiming to preserve its desired olfactory' character is referred to hereinunder as a reformulation. Odor reformulation enables optimizing / replacing ingredients and proportions thereof in the existing smells, creating new smells, addressing safety and regulatory concerns, inventory management, etc.
[0006] One of the challenges of odor reformulation is to find individual odor items (odors and / or their material sources) that, combined, constitute a mixture that smells like a target odor.
[0007] The problems of odor reformulation have been recognized in conventional art, and various machine learning techniques have been developed to provide a solution. However, the effectiveness of applying machine learning to digital odor formulation / reformulation is limited by the lack of labeled data required for training and running the models.GENERAL DESCRIPTION
[0008] Each given odor item (i.e. an odor and / or its material source(s)) can be characterized by a set of its representative aspects in one or more odor-related spaces. By way of a non-limiting example, in a perceptual-based space, the representative aspects can be informative of descriptors of the odors (e.g. “sweet”, “citrus”, “metallic”, etc.). By way of another non-limiting example, in a receptor-based space the representative aspects can be informative of such aspects of the odors as smell receptors activatable by a given odor; in a chemical-based space the representative aspects can be informative of chemical features of the odors, etc.
[0009] Representative aspects can be associated with representative values. In certain embodiments, the representative values are binary and indicate presence or absence of a given representative aspect in odor’s characteristics. In other embodiments, at least part of the representative aspects can be characterized by non-binary representative values (e.g. values indicative of intensity of a certain scent, receptor’s sensitivity, weight of the aspect in the set of aspects, etc.).
[0010] Two or more odor items can be characterized by their similarity which is indicative of how similar the two or more odors are perceived (by a human observer or by a machine). Alternatively or additionally, two or more odor items can be characterizedby their additivity which is informative of how the perceived smell of a mixture relates to the smells of its individual components.
[0011] For example, in a perceptual-based space the smell of lemon can be represented as a combination of such representative aspects as "fresh" and "citrus” with intensity values of 3 and 10, respectively. The smell of orange can be represented as similar to that of lemon, but with an added "sweet" note. Similarly, the smell of lemon pie can be represented as the smell of lemon combined with a "bakery" note.
[0012] In the above example, additivity refers to the question of what the combined smell of bread and lemon would be. Similarity, on the other hand, can concern whether the combination of bread and lemon resembles the smell of a lemon pie.
[0013] In accordance with certain embodiments of the currently presented subject matter, similarity and additivity can be considered as a regression problem providing a similarity score on a continuous scale. Likewise, similarity and / or additivity can be considered as multi-classification problems predicting for each descriptor one of predefined values (i.e. classes). Similarity and / or additivity of odor items can be defined by applying respective machine learning models trained on the labeled datasets. The labels in such datasets are indicative of similarity and / or additivity of certain odor items and can be obtained using distinction experiments (e.g., distinction tests for similarity, professional evaluations for additivity, etc.). However, such experiments are timeconsuming and labor-intensive, resulting in labeled datasets that can be insufficient for training the corresponding models.
[0014] The inventors have recognized and appreciated that the problem of insufficient labeled datasets for training and running the mathematical models applicable for smells digitization can be solved with the help of weakly-supervised learning adapted in accordance with certain embodiments of the currently presented subject matter. Alternatively or additionally, the problem of insufficient labeled datasets can be solved by smell-related data augmentation and / or using the properties of unlabeled data themselves.
[0015] When searching for substituting an odor item of interest in the formula of the smell, one can consider the entire inventory as a source of potential replacing candidates for the reformulation. However, while inventories can include thousands of items, most formulas have dozens of ingredients and even less. Therefore, it can be desirable to optimize the inventory and reduce it to a small set of primary odor items, which can be used to represent the desired smells.
[0016] In accordance with certain aspects of the currently presented subj ect matter, there is provided a computerized method of substituting an odor item of interest in a formula of a smell. The method comprises, by a computing system: generating for a given inventory of odor items a dataset informative, at least, of similarity of the odor item of interest with other odor items from the inventory; processing data in the generated dataset to reveal one or more odor items meeting one or more similarity criteria, thereby generating one or more candidate odor items; applying to the one or more generated candidate odor items an optimization algorithm to find, among them, a candidate odor item with the best fit to the formula of the smell, thus giving rise to a replacing odor item usable for substituting the odor item of interest in the formula of the smell.
[0017] At least part of the dataset can be generated responsive to a request for substituting the odor item of interest. One or more candidate odor items can be generated responsive to a request for substituting the odor item of interest.
[0018] In accordance with further aspects of the currently presented subj ect matter, for a plurality of odor items of interest in the formula of the smell, the method comprises: separately for each odor item of interest from the plurality of odor items of interest, generating one or more candidate items; applying to candidate odor items obtained for all odor item of interest from the plurality of odor items of interest an optimization algorithm to find a set of candidate odor items with the best overall fit to the formula of the smell, thus giving rise to a replacing set of odor items usable for substituting the plurality of odor items of interest in the formula of the smell.
[0019] In accordance with further aspects of the currently presented subj ect matter, generating the one or more candidate odor items can comprise: for each pair constituted by the odor item of interest and another odor item from the inventory of odor items obtaining a plurality of similarity labels; for each odor item (01) pair, applying a parametric model to obtain a similarity score of the 01 pair, wherein the similarity labels of a given 01 pair are used as the features of the applied parametric model and the similarity score is obtained on a continuous scale; and selecting 01 pairs with similarity score matching, at least, a similarity threshold, thereby giving rise to one or more candidate odor items usable for replacing the odor item of interest.
[0020] In accordance with further aspects of the currently presented subject matter, the similarity labels of a given 01 pair can be obtained by applying to the given 01 pair a plurality of labeling functions, each labelling function resulting from a distinct rule, source and / or type of information, wherein each labeling function is used as a weak classifier to obtain the plurality of similarity labels for the given 01 pair.
[0021] An applied labeling function can be informative of at least one of: perceptual similarity, chemical similarity, similarity of receptors to be involved, similarity of labels provided by domain experts, and proximity of certain numerical features.
[0022] An applied labeling function can be learnt from the dataset or is represented by a fixed rule.
[0023] In accordance with further aspects of the currently presented subject matter, obtaining the plurality of similarity labels can comprise: obtaining multiple evaluations of the 01 pairs in inventory of odor items from the same annotators on different occasions; pairing the obtained multiple evaluations into evaluation pairs; generating one or more labeling functions corresponding to one or more evaluation criteria; and applying the generated labeling function to the evaluation pairs to provide similarity labels for the 01 pairs.
[0024] In accordance with further aspects of the currently presented subject matter, obtaining the plurality of similarity labels is obtained with the help of graph-based modeling. The method can comprise: using a base classifier of similarity of odor items in the inventory to build a graph with nodes representing the odor items and with edges connecting odor items classified by the base classifier as similar; using the graph’s properties to generate a new classifier, wherein: two nodes are considered as similar when a number of similar to both common neighbors exceeds a predefined threshold; a node is considered as un-similar to the odor of interest when a number of its neighbors that are unsimilar to the odor of interest exceeds a predefined threshold; and applying the generated new classifier to the inventory of odor items to provide similarity labels for the 01 pairs.
[0025] In accordance with further aspects of the currently presented subject matter, when the odor pairs are characterized by multiple descriptors, the method can comprise for each given OI pair: obtaining similarity labels for each descriptor of a given OI pair, separately for each descriptor, building a parametric model with a sum of entropies per descriptor, wherein the similarity labels for respective descriptor are used as the features of the parametric model; and processing together the respectively built per-descriptor parametric models to find the weights minimizing the sum of entropies, thereby defining the multi-descriptor similarity score of the pair.
[0026] In accordance with other aspects of the currently presented subject matter, there is provided a computing system configured to perform the operations of the method above.
[0027] In accordance with other aspects of the currently presented subject matter, there is provided a non-transitory computer-readable medium comprising instructions that, when executed by a computing system comprising a memory storing a plurality of program components executable by the computing system, cause the computing system to operate in accordance with the method above.BRIEF DESCRIPTION OF THE DRAWINGSIn order to understand the invention and to see how it can be carried out in practice, embodiments will be described, by way of non-limiting examples, with reference to the accompanying drawings, in which:Figs, la and lb illustrates a generalized flow-chart of a method of reformulating a smell in accordance with certain embodiments of the presently disclosed subject matter;Fig- 2 illustrates a generalized flow-chart of selecting candidate odor items usable for replacing the odor item of interest in accordance with certain embodiments of the presently disclosed subject matter;Fig. 3 illustrates a generalized flow-chart of obtaining similarity labels with the help of data augmentation in accordance with certain embodiments of the presently disclosed subject matter;Fig. 4 illustrates a generalized flow-chart of obtaining similarity labels with the help of graph-based modeling in accordance with certain embodiments of the presently disclosed subject matter;Fig. 5 illustrates a generalized flow-chart of multi-descriptor selecting the candidate odor items in accordance with certain embodiments of the presently disclosed subject matter;Fig. 6 illustrates a method of concurrent validation of additivity and similarity models in accordance with certain embodiments of the presently disclosed subject matter;Fig. 7 illustrates a generalized flow-chart of identifying primary odor items for a target smell. In accordance with certain embodiments of the presently disclosed subject matter; andFig. 8 illustrates a generalized block diagram of a computing system capable of digital smell reformulation in accordance with certain embodiments of the presently disclosed subject matter.DETAILED DESCRIPTION
[0028] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the invention. However, it will be understood by those skilled in the art that the presently disclosed subject matter may be practiced without these specific details. In other instances, well-known methods, procedures, components and circuits have not been described in detail so as not to obscure the presently disclosed subject matter.
[0029] Unless specifically stated otherwise, as apparent from the following discussions, it is appreciated that throughout the specification discussions utilizing terms such as "processing", "applying", "generating", “identifying”, “determining”, or the like, refer to the action(s) and / or process(es) of a computer that manipulate and / or transform data into other data, said data represented as physical, such as electronic, quantities and / or said data representing the physical objects. The term “computer” should be expansively construed to cover any kind of hardware-based electronic device with data processing capabilities including, by way of a non-limiting example, the computing system and processing and memory circuitry therein disclosed in the present application.
[0030] The operations in accordance with the teachings herein may be performed by a computer specially constructed for the desired purposes or by a general-purpose computer specially configured for the desired purpose by a computer program stored in a non-transitory computer-readable storage medium.
[0031] Unless specifically stated otherwise, the terms “mixture of odors in proportion P” “combination of odors in proportion P” or alike cover also a combination of individual material sources of different odors that are taken in proportion P and are perceived simultaneously. It is further noted that, when appropriate, a mixture can be considered as a single item / ingredient / material.
[0032] Unless specifically stated otherwise, the term “odor item” covers the respective odor and material source(s) thereof.
[0033] Bearing this in mind, attention is drawn to Figs, la and lb illustrating a generalized flow-chart of a method of reformulating a smell.
[0034] Creating a new smell or altering an existing formula may require substituting a particular odor item (referred to hereinafter also as an odor item of interest).
[0035] Fig. la illustrates a generalized flow-chart of a substituting a single odor item in the formula. In accordance with certain embodiments of the presently disclosed subject matter, reformulation process comprises generating (101) for a given inventory of odor items a dataset informative, at least, of similarity of the odor item of interest with other odor items from the inventory.
[0036] A computing system generates (102) one or more candidate odor items. Obtaining the candidate odor items comprises processing data in the generated dataset to reveal odor item(s) meeting one or more similarity criteria in relation to the odor item to be substituted.
[0037] Computer-based generation of the dataset informative, at least, of similarity of odor items pairs therein and revealing the candidate odor items is further detailed with reference to Figs. 2 - 5.
[0038] It is noted that at least part of the dataset can be generated in advance and be stored in a database within the computing system or operatively connected to the computer system. At least part of the dataset can be generated responsive to a request for substituting the odor item of interest.
[0039] It is further noted that the pairs with the odor item of interest can be generated for all odor items in inventory or only for odor items meeting initial similarity criteria (e.g. chemical or perceptual). Optionally, the dataset can be informative of similarity of pairs other than pairs with the odor item of interest. Optionally, the dataset can be informative of similarity of all pairs of odor items in the given inventory. Furthermore, the inventory itself can be optimized as further detailed with reference to Fig. 7.
[0040] The computing system further applies (103) to the one or more candidate odor items an optimization algorithm to find, among them, a candidate odor item with the best fit (subject to placed one or more constrains) to the formula of the target smell, thus giving rise to a replacing odor item. The constrains can be related to the odor item of interest and / or to the entire formula.
[0041] Thus, the given odor item can be substituted (104) in the formula by the replacing odor item, thereby enabling reformulation of the target smell.
[0042] Fig. lb illustrates a generalized flow-chart of a substituting multiple odor item in the formula in accordance with certain embodiments of the presently disclosed subject matter. Similar to disclosed with reference to Fig. la, reformulation process comprises generating (111) for a given inventory of odor items a dataset informative, at least, of similarity of each of the odor item of interest with other odor items from the inventory.
[0043] Separately for each odor item of interest, the computing system generates (112) one or more candidate odor items. Obtaining the candidate odor items comprises processing data in the generated dataset to reveal odor item(s) meeting one or more similarity criteria in relation to the respective odor item to be substituted.
[0044] The computing system further applies (113) to all candidate odor items an optimization algorithm to find, among them, a set of candidate odor item enabling the best overall fit to the formula of the target smell, thus giving rise to a replacing set of odor items.
[0045] The optimization algorithm gets as input a list of candidates for each of the target odor items, their similarity scores (further detailed with reference to Fig. 2), and a set of constraints (e.g. constraints on maximum allowed concentration for each replacement and / or a combination of replacements, constraints on availability of candidates, constraints on the total cost after reformulation, etc.). Then the optimization problem is solved to get the best set of replacements that will maximize the total score subject to the constraints.
[0046] Thus, the odor items of interest can be substituted (114) in the formula by the replacing set of odor items, thereby enabling reformulation of the target smell.
[0047] Referring to Fig. 2, there is illustrated a generalized flow-chart of computerized selecting candidate odor items (OIs) usable for replacing the odor item of interest.
[0048] In accordance with certain embodiments of the presently disclosed subject matter, there is provided a plurality of labeling functions that are further applied (201) to each pair constituted by an OI of interest and another OI from the inventory of OIs. Each labelling function in the plurality of labeling functions results from a distinct rule, source and / or type of information.
[0049] A labeling function (LF) is a programmatic rule or model that assigns a label to a data instance based on heuristics, patterns, external knowledge, etc.
[0050] By way of non-limiting example, a labeling function for OIs can be informative of at least one of: perceptual similarity, chemical similarity, similarity of receptors to be involved, similarity of labels provided by domain experts, proximity of certain numerical features, etc.
[0051] Labeling functions can be either learnable from the dataset or be represented by a fixed rule. Such rules can be, for example: two materials being isomers are similar since they are chemically similar, two evaluations of the same material are similar since they represent the same item, etc.
[0052] Some labeling functions can provide Boolean labels, while other labeling functions can also provide a probability of correctly assigning a respective label.
[0053] In accordance with certain embodiments of the presently disclosed subject matter, the computing system further uses the assigned labeling functions as weak classifiers to obtain (202) similarity labels for each of the pairs. A weak classifier is amodel that performs only slightly better than random guessing on a classification task, thus the similarity labels can be only slightly better than the guesses.
[0054] It is noted that the similarity labels of each 01 pair are obtained by applying the same set of labeling functions.
[0055] In certain embodiments, similarity labels can be obtained with the help of data augmentation, i.e. with the help of modifying unlabeled data to generate proxy labels usable in supervised learning tasks. A generalized flow-chart of obtaining similarity labels with the help of data augmentation is illustrated in Fig. 3.
[0056] Obtaining similarity labels comprises obtaining (301) multiple evaluations of OI pairs in the inventory of odor items, the evaluations received from the same annotators (machine or human) on different occasions. The computing system pairs (302) the obtained multiple evaluations into evaluation pairs and generates (304) labeling function(s) corresponding to one or more evaluation criteria. The computing system further applies (305) the generated labeling function to the evaluation pairs to provide similarity labels for the OI pairs.
[0057] The evaluation criteria can be configurable (303). For example, multiple evaluations of the same odor item can be considered as having positive similarity, while evaluations of two randomly selected odor items can be considered as having the negative similarity.
[0058] An evaluation criterion can be configured such that not only A and B smells are different, but also smell A and a mixture of AxBywith concentration of x <75% are different smells. Likewise, an evaluation criterion can be configured to define that multiple evaluations provided by any of different annotators (machine or human) from a certain group are considered as having positive similarity.
[0059] In certain embodiments, the evaluation criteria can be configured to define different members of the same category (e.g., fruits, flowers) as being of “grey similarity”. This means that members of a category do not smell identical but much closertherebetween than to items out of the category. For example, Orange and Mango, Jasmine and Rose, somewhat resemble each other yet represent distinguishable objects. Depending on similarity criteria, the similarity label for the same members of a certain category can be positive or strong negative. For example, depending on the application, it can be defined that Jasmine and Rose are positive (because they belong to the same category) or, alternatively, negative (because they smell different).
[0060] Optionally, the 01 pairs belonging to the same category can be defined as having positive similarity, wherein the fractional positivity level can be defined using distinction tests. Likewise, OI pairs with high chemical match (e.g., A and a mixture of AisBis) can be defined as having positive similarity, wherein the fractional positivity level can be defined using distinction tests. The fractional positivity level is indicative of what is the minimal fraction x of A in the mixture A xB (l-x) before one can distinguish between A and the mixture.
[0061] Thus, the odor items in the semantic monotonic groups (i.e. OI that are similar in category, meaning, concept, interpretation, etc.) can be considered as items with enhanced similarity. Semantic monotonic groups are usable for validating and training similarity models to present the aligned perceptual similarity. In certain embodiments, labeling functions can be generated so to provide the same label to all items in a semantic monotonic group.
[0062] Alternatively or additionally, similarity labels can be obtained using graph-based modeling. A generalized flow-chart of obtaining similarity labels with the help of graph-based modeling is illustrated in Fig. 4.
[0063] Upon obtaining (401) an inventory of odor items and a base classifier of similarity thereof (e.g. a chemical classifier), the computing system builds (402) a graph with nodes representing the odor items and edges that connect items considered as similar in accordance with the base classifier. The computing system further uses the graph’s properties to generate a new classifier (403).
[0064] For example, if two nodes have many common neighbors similar to both, they are likely to be similar. Likewise, a node having many neighbors that are not similar to an odor item of interest is likely to be unsimilar thereto. The new classifier can be built by generating a similarity model with graph-based properties (e.g., shortest distance, min cut, etc.) used as the features.
[0065] The computing system further applies to the inventory of odor items the generated new classifier to provide (404) similarity labels for the OI pairs.
[0066] Referring back to Fig. 2, upon obtaining the similarity labels of the OI pairs, the computing system further applies to the OI pairs a parametric model to obtain (203) a similarity score of each pair, wherein for each given pair its similarity labels are used as the features of the parametric model. By way of non-limiting examples, the parametric model can be a Logistic Regression Model, Naive Bayes Model, Neural Network Model, etc.).
[0067] Using the parametric model with inputs from weak classifiers (similarity labels) for determining OI similarity solves the similarity prediction as a regression task and, thus, enables providing a similarity score on a continuous scale (e.g. from 0 to 1).
[0068] The computing system further selects (204) OI pairs with similarity score matching, at least, a similarity threshold, thereby giving rise to one or more candidate odor items usable for replacing the odor item of interest.
[0069] In certain embodiments, further to similarity score, the candidate odor items can be required to meet one or more additional similarity criteria (e.g. chemical similarity, belonging to the same category of materials, etc.).
[0070] It is noted that in weak supervision, a model can be trained by using a combination of noisy, incomplete, or limited supervision signals, including labeling functions. While a single labeling function might produce noisy or conflicting labels, combining multiple labeling functions allows mitigating the noise and obtaining a more reliable signal.
[0071] The larger the set of labeling functions used, the larger the potential predictive power. However, as the similarity labels can be only somewhat better than a guess, they might disagree with each other.
[0072] Accordingly, in certain embodiments, the computing system can minimize a disagreement between the similarity labels of a given OI pair prior to using them in the parametric model. Alternatively or additionally, the parametric model can be learned with a regularization that favors the classifiers that are in agreement with each other.
[0073] By way of non-limiting example, the disagreement can be minimized by statistical means, e.g. aggregating weighted outputs of the labeling function. The weight of a certain labeling function can be the same for each OI pair or can be different (e.g. resulting from the optimization, chosen a-priory such that a few labeling functions are assigned the same weight, etc.)
[0074] Alternatively or additionally, the disagreements between the similarity labels of a given pair can be minimized by reducing disagreement between the weak classifiers, e.g. by boosting a labeling function selected as a target classifier.
[0075] In certain embodiments, boosting the target classifier can comprise: identifying disagreements between predictions of the target classifier and prediction of the rest of the classifiers, and using the disagreement to improve the target classifier by requiring their agreement for a hit. In case of disagreement between them, training a third classifier on the set of the disagreements to classify the same concept of the rest of the classifier prediction. Then, using the majority of these three classifiers as the prediction (similarity label).
[0076] This boosting process can be further optimized for different goals by defining at least one of the following: the way of aggregating the prediction of the rest of the classifiers; the level of required consensus (e.g. where a reliable label requires full consensus or 75% agreement is enough; what to prefer - precision or recall, etc.).
[0077] Optionally, two or more weak classifiers can be co-trained thereby making stronger each other. The boosting process can be provided for multiple classifiers based on perceptual, biological and chemical datasets, wherein the datasets have a built- in difference between the views, preventing the models from collapsing into one.
[0078] Alternatively or additionally to aggregating the similarity labels, the disagreement of the similarity labels can be minimized by a regressor (e.g. a regression model) that can automatically learn aggregation rules from raw or lightly processed features. Optionally, the regressor can create aggregated weighted sums.
[0079] Alternatively or additionally, the disagreement of similarity labels can be minimized with the help of Maximum Likelihood Estimation (MLE) approach that is configured to learn from noisy labels and combine the outputs of the labeling functions into probabilistic labels. Combining the regressor with MLE enables fitting the parametric model to the inferred labels.
[0080] Thus, in accordance with certain aspects of the presently disclosed subject matter, the technique of minimizing disagreement in a set of weak classifiers can be extended to the regression and the learning problem can be reduced into an optimization problem.
[0081] Referring to Fig. 5, there is illustrated a generalized flow-chart of multidescriptor selecting the candidate odor items.
[0082] As detailed above, a smell can be characterized by different representative aspects (descriptors) in different representing spaces (e.g. by “sweet”, “citrus”, “metallic”, etc. in the perceptual space). In accordance with certain embodiments of the presently disclosed subject matter, the candidates for replacing can be selected in consideration of multiple descriptors.
[0083] For each OI pairs from the inventory of odor items and for each descriptor in a given OI pair, the computing system assigns (501) a plurality of labeling functions and uses the assigned labeling functions as weak classifiers to obtain (502) similaritylabels of the descriptor of the pair. For each given 01 pair, separately for each descriptor, the computing system builds a parametric model with a sum of entropies per descriptor, wherein for each given pair the similarity labels for respective descriptor are used as the features of the parametric model (503).
[0084] Operations (501) - (503) can be provided in a manner detailed with reference to Figs. 2 - 4.
[0085] For each pair, the computing system further processes together the respectively built per-descriptor parametric models to find the weights minimizing the sum of entropies, thereby defining (504) the multi-descriptor similarity score of the pair. The computing system further selects (505) the pairs with similarity score matching one or more similarity criteria, thereby giving rise to one or more candidate odor items usable for replacing the odor item of interest.
[0086] Further to similarity of replacing odor items, creating a new smell or altering an existing formula may require considering additivity of a formula’s components. If linear additivity of odor items (referred to hereinafter also as “smell additivity) holds, the perceived intensity or quality of a mixture would be the linear sum of the individual smells. Digital models estimating the smell additivity attempt to simulate or predict how multiple odor items combine perceptually.
[0087] In a lack of labeled data, a training set for additivity estimations can be obtained with the help of data augmentation detailed below.
[0088] The relationship between a mixture and its ingredients can be represented by smell additivity equations. Let Apibe A diluted to pi percents. Let Api+ BP2 be a mixture in which A appears in pi percents and B in p2. Additivity of A 50 and B50 can be represented by equation A50 +B50 = A50B50, where A 50, B50, A50B50 are perceptual descriptors of the relevant smell and the operator “+” is the additivity model to learn.
[0089] Data augmentation for additivity purposes can include: obtaining a set of odor items; obtaining a set of mixtures of the odor items from the set; obtaining a set ofevaluations of the items and mixtures; and creating additivity equations for a mixture by the items in it, using all permutations of evaluations.
[0090] Further to using descriptors of the ingredients, more equations can be obtained from related odor items. For example, the perceptual sum of odor items leading to a chemical sum with the same ingredients and ratios as the target mixture can be used as an equation (e.g. (^75^25)50 + ( lhs) = A5oBso).
[0091] Alternatively or additionally to the above, the lack of labeled data can be compensated by using unlabeled data to enforce the model to respect such additivity properties as agreement of models, monotonicity, and symmetry.AGREEMENT OF MODELS
[0092] Multiple additivity models can be combined to improve performance. The combination of models could be done with traditional ensemble methods, including boosting, bagging and the likes. In addition to improving performance, careful analysis of the disagreement between models could help identify modeling errors and improve the training of the individual models.MONOTONICITY:
[0093] Chemical similarity can be defined as the ratio of similar ingredients between two formulas (e.g. as disclosed in M. J. Olsson and W. S. Cain. Psychometrics of odor quality discrimination: method for threshold determination. Chemical senses, 25(5):493-499, 2000). For illustration, given A and B are different, the more of A is mixed with B, the result will be less similar to B. Therefore, sim AisB-^B) < sim(A5oB5o,B).
[0094] Therefore, one can automatically generate random cases and identify instances that violate monotonicity. Such violations may result from errors in the additivity model, but can also stem from inaccuracies in the input descriptions. These error-prone cases can serve as valuable inputs for error analysis potentially leading toimprovements in the additivity model. This might result in a different additivity models for a subset of cases, such as additivity when antagonists or trigeminal materials are involved.SYMMETRY:
[0095] The additivity function should be symmetric, meaning A+B=B+A, and associative, such that (A+B)+C=A+(B+C). Preferably, the model shall be constructed as invariant to permutations, including the order of summation. To evaluate these properties, one can generate various combinations and compare the outputs of the summation under different orderings. A significant difference in results indicates a violation of symmetry or associativity. Such discrepancies suggest a problem in the model, as it would produce inconsistent predictions for the same mixture depending solely on the order of its ingredients.
[0096] Thus, in accordance with certain embodiments, using unlabeled data for additivity purposes can include: obtaining a set of items; obtaining a set of evaluations; obtaining additivity model (or models) to validate; choosing a property to validate (e.g., agreement of models, monotonicity, and symmetry); generating random additivity cases that should respect the property; and running the additivity model on the cases identifying violations and their ratios.
[0097] The model trained on such data will respect the desired additivity properties.
[0098] In accordance with further aspects of the currently presented subject matter, there is proposed a method of concurrent validation of additivity and similarity models using synthetic data.
[0099] As illustrated in Fig. 6, the method includes generating (601) a set of mutations Fs,nof a formula F„o.[000100] Mutation operation can include, for example, adding T% of a novel randomly selected ingredient to the recipe, or reducing % of a randomly selected ingredient from the recipe. Given an original formula F, for each G times of applying a mutation operation, one can consider the resulting formula to be g generations, or mutation steps, away from the original formula F. Due to the randomness involved in this process, applying a mutation operation on the same original formula S times while following a different mutation path or strand each time evolves into a different end-result FS,G stemming from the same original formula F, where S represents a respective strand and G represents a respective generation.[000101] Thus, there can be generated the mutation data set comprising the original formula F.o and S *G mutated forms thereof (FS,G).[000102] In certain embodiments, the mutation dataset can be generated as following: defining and validating a set of mutation operations, M; defining an original formula, F.,o; applying a random mutation operation out of M, the result is termed Fs, i; for each strand Fs.n, applying a random mutation operation out of M, the result is termed Fs,n+r, and repeating this G times;[000103] Upon generation of the set of mutations, the computing system defines (602) for each Fs,na perceptual signature using an additivity model to be validated. The perceptual signature is informative of perceptual values of the descriptors characterizing the smell.[000104] Further the computing system uses a similarity model to be validated to define (603) for each Fs,nthe perceptual similarity between the computed perceptual signatures Fs,n andFs,n-i, and to define (604) for each Fs,n (n > 1 ), the perceptual similarity between the computed perceptual signatures Fs,nand F o.[000105] Valid similarity and additivity functions shall enable the following essential properties:The perceptual similarity of an original formula, F„o, to its mutated forms (Fs,x,Fs,z should decrease as the number of mutation steps increases. Therefore, sim(Ffi,Fs,z) < sim(F.o,Fs,x), if x < z.Across all strands, the perceptual similarity of a formula F,xto its mutated form should be proportional to their generational distance. Therefore, for each x,z, simiF^F ) = k * |(x - z)|.The perceptual similarity of an original formula, F.o, to its mutated forms x generations away, should be similar across different strands of evolution (s,r). Therefore, sim(F.fi,Fs,^ = sim(F.fi,Fr^).[000106] The computing system further analyses the results to define compliance (605) with the properties defined as essential for validity, thereby providing concurrent validation of similarity and additivity functions.[000107] If the essential properties are violated, it follows that the similarity model and / or additivity model are wrong. In contrast, if the properties of the generated set of mutations match the essential validation properties, it serves as evidence for two-way validation (cross-validation) of both the similarity and the additivity models. Optionally, the analyses can take into account additional factors such as chemical similarity or the number of ingredients in the original formula, and validating the set of mutation operations as having similar effects, etc.[000108] In accordance with further aspects of the currently presented subj ect matter, it is desirable to optimize the inventory of potential candidates for replacement (referred to hereinunder as “reformulation inventory”) and reduce it to a small set of primary odor items, which can be used to represent the one or more smells requiring reformulation.[000109] A target perceptual smell can be represented by different formulas. The collection of such formulas is referred to hereinafter as a collection covering the perceptualtarget smell, and ingredients in any of the formulas in such collection are referred to hereinafter as "primary ingredients". Thus, a set of primary ingredients is sufficient to enable reformulation(s) of the perceptual target smell.[000110] In accordance with certain embodiments of the currently presented subject matter, the reformulation inventory can be optimized by identifying, for each target smell, a set of primary ingredients and constructing the reformulation inventory exclusively from these primary ingredients.[000111] Identifying the primary ingredients of a perceptual target smell can be provided with the help of supervised learning of “covering relationship” between ingredients in the inventory.[000112] Two odor items can be characterized by covering relationships in a chemical space and in a perceptual space.[000113] For example, considering chemical space, if A is 100% lemon and B is 50% lemon and 50% watermelon, then A can cover B, but B cannot cover A. Likewise, considering perceptual space, lemon can cover orange and lemon can cover the sour in orange.[000114] Covering relationships of odor items in the chemical space can be defined by chemical-related rules:Items can cover themselves;Item can cover its dilution;A significant dilution (e.g., not 99% of pure but 50%) cannot cover the pure. Proper significance level can be obtained from distinction tests (e.g. see a review of the smell distinction methods known in contemporary art in the article of Wise et al. (P. M. Wise, M. J. Olsson, and W. S. Cain. Quantification of odor quality. Chemical senses, 25 (4): 429-443, 2000).;An ingredient of a formula covers the formula;A formula cannot cover its ingredient (given that the rest of the formula is dissimilar enough from the ingredient);Two perceptually dissimilar ingredients cannot cover each other.[000115] By way of non-limiting example, dissimilarity of ingredients can be defined by applying a similarity function (e.g., Euclidian Distance (ED), Jaccard, etc.) to data presented in various perceptual datasets (e.g., TGS, Leffingwell, etc.).[000116] In the perceptual space, covering relationship of ingredients A and B for descriptor d (e.g., sweet, sulphuric, fresh, etc.) can be defined with the help of heuristic perceptual rules. The perceptual rules consider the following handcrafted features characterizing a pair of materials A and B in relation to descriptor dViolations - when A adds new feature(s) to BContributions - when the feature d in B can be reached by diluting A.[000117] In order to reformulate a target smell, one should cover each of its descriptors.[000118] An ingredients can cover ingredient B for descriptor d if its contribution to ingredient B exceeds a contribution threshold and its violation to ingredient B is below a violation threshold.[000119] The heuristic perceptual rules regulating the covering relationships can include the following:[000120] Contribution of an ingredient is evaluated only in relation to its active descriptors. The definition of a descriptor as active (e.g. dominant, salient, characterdefining, etc.) can be provided via domain expertise or via multiple evaluations and reliability tests. A threshold defining an active status of the descriptor can be defined with the help of an auxiliary dataset.[000121] Ingredients can contribute to themselves. For example, given a molecule X which was measured twice. One measurement has three non-zero descriptors, and thesecond measurement also has the same three descriptors. However, the magnitude of these descriptors is not identical. Both these measurements can be labeled as positive examples that can contribute to each other.[000122] Ingredient can contribute to its dilution (ingredient XI 00% can contribute to X50%).[000123] A significant dilution (e.g., not 99% of pure but 50%) cannot contribute any descriptor to pure. X50% cannot contribute to XI 00% because all the descriptors in X50% are lower than XI 00%.[000124] For descriptor d, ingredient A can contribute to ingredient B only when activity level of descriptor d in A is higher enough than activity level of descriptor d in B (e.g. in accordance with data in the auxiliary dataset).[000125] It is noted that the level of “enough” is derived from precision-recall tradeoff of the respective model.[000126] Referring to Fig. 7, there is illustrated a generalized flow-chart of identifying primary odor items for a target smell. In accordance with certain embodiments of the presently disclosed subject matter, the method includes using the chemical-related and the heuristic perception rules to generate (701) a labeled dataset indicative of covering relationships of pairs of odor items; and using the dataset to generate (702) a model predicting the covering relationship in pairs of ingredients (referred to hereinafter as a “coverage model”).[000127] The method further comprises obtaining (703), for each ingredient in a given pair of ingredients, multiple expert evaluations. The multiple evaluations can be provided for values of each descriptor relevant to the ingredients in the given pair of ingredients. The multiple expert evaluations are further combined in evaluation pairs, and the computing system applies the coverage model to each evaluation pair (704) to predict the covering relationship of respective ingredients in each evaluation pair. The computingsystem uses the results of per-evaluation pair prediction to predict (705) the covering relationship of the pair of ingredients.[000128] For example, if we have two ingredients with 3 evaluations each, the coverage model can be applied on all 9 pairs. Intuitively, for good coverage one can expect 9 out of 9 positive results. In unsuitable coverage, one can expect 0 out of 9 positive results.[000129] The prediction for the ingredients can be obtained by statistical analyses of predictions obtained for the respective evaluation pairs (e.g. by majority voting per hit rate, etc.)[000130] Other than using the hit rate, one can use the base classifier predictive performance to estimate the positive rates. By way of non-limiting example, such classifier can be used in a way disclosed in the article by Amit et. al (Amit Idan and Dror G. Feitelson, "Corrective commit probability: a measure of the effort invested in bug fixing." Software Quality Journal 29.4 (2021): 817-861) incorporated hereby by reference.[000131] The computing system uses the predicted covering relationships for the ingredient pairs in inventory to define (706) the primary ingredients.[000132] In accordance with certain embodiments of the presently disclosed subject matter, there is provided an alternative method of identifying the ingredients suitable for a formula of a perceptual target smell.[000133] Suppose that we have numerical perceptual evaluations.[000134] Given a subtraction operatora target description T, and a potential covering item C we can compute the subtraction T-C=S. If S has negative descriptor D, C contributed this descriptor D in amount higher than existing in T.[000135] Since we cannot handle negative mixture in lab, C is not a suitable contributor for T.[000136] If S has no negative descriptors, C is a suitable contributor. C can contribute any of the active descriptors in T.[000137] Note that while we can perform additionby mixing materials and similarity “==” using distinction tests, there is no direct way to represent subtraction.[000138] Instead, we can implement subtraction by a reduction to additivity and similarity and find:[000140] A possible extension is to use a dilution of T and not only pure T.[000141] This can be done by using a direct evaluation of the dilution of T or and estimation of the diluted version evaluation, from arithmetic evaluation or a model that used the pure version evaluation and evaluation level (as done in intensity curves).[000142] Referring to Fig. 8, there is illustrated a generalized block diagram of a computing system capable of digital smell reformulation in accordance with certain embodiments of the presently disclosed subject matter. The illustrated system 800 comprises input / output interface 801 operatively connected processing and memory circuitry (PMC) 802 comprising a processor and a memory (not shown separately within the PMC).[000143] The system is configured to receive data informative of items and / or reformulation requests via input / output interface 801. PMC 802 is configured to execute computer-readable instructions implemented on a non-transitory computer-readable storage medium. The instructions, when executed by PMC 802 cause the computing system to process data informative of the odor items and enable digital smell reformulation as detailed with reference to Figs. 1 - 8. The computing system can maintain one or more datasets detailed with reference to Figs. 1 - 8. Alternatively or additionally, at least part of the datasets and / or data thereof can be maintained in one or more external databases operatively connected to the computing system 800.[000144] The computing system 800 can be a standalone entity, or can be integrated, fully or partly, with other systems.[000145] It is to be understood that the invention is not limited in its application to the details set forth in the description contained herein or illustrated in the drawings. The invention is capable of other embodiments and of being practiced and carried out in various ways. Hence, it is to be understood that the phraseology and terminology employed herein are for the purpose of description and should not be regarded as limiting. As such, those skilled in the art will appreciate that the conception upon which this disclosure is based may readily be utilized as a basis for designing other structures, methods, and systems for carrying out the several purposes of the presently disclosed subject matter.[000146] It will also be understood that the system according to the invention may be, at least partly, implemented on a suitably programmed computer. Likewise, the invention contemplates a computer program being readable by a computer for executing the method of the invention. The invention further contemplates a non-transitory computer- readable memory tangibly embodying a program of instructions executable by the computer for executing the method of the invention.[000147] The references cited above teach background information that may be applicable to the presently disclosed subject matter. Therefore, the full contents of these publications are incorporated by reference herein where appropriate for appropriate teachings of additional or alternative details, features and / or technical background.[000148] Those skilled in the art will readily appreciate that various modifications and changes can be applied to the embodiments of the invention as hereinbefore described without departing from its scope, defined in and by the appended claims.
Claims
CLAIMS1. A computerized method of substituting an odor item of interest in a formula of a smell, the method comprising, by a computing system: generating for a given inventory of odor items a dataset informative, at least, of similarity of the odor item of interest with other odor items from the inventory; processing data in the generated dataset to reveal one or more odor items meeting one or more similarity criteria, thereby generating one or more candidate odor items; and applying to the one or more generated candidate odor items an optimization algorithm to find, among them, a candidate odor item with the best fit to the formula of the smell, thus giving rise to a replacing odor item usable for substituting the odor item of interest in the formula of the smell.
2. The method of Claim 1 , wherein the one or more candidate odor items are generated responsive to a request for substituting the odor item of interest.
3. The method of Claims 1 or 2, wherein there is a plurality of odor items of interest in the formula of the smell, the method comprising: separately for each odor item of interest from the plurality of odor items of interest, generating one or more candidate items; and applying to candidate odor items obtained for all odor item of interest from the plurality of odor items of interest an optimization algorithm to find a set of candidate odor items with the best overall fit to the formula of the smell, thus giving rise to a replacing set of odor items usable for substituting the plurality of odor items of interest in the formula of the smell.
4. The method of any one of Claims 1 - 3, wherein at least part of the dataset is generated responsive to a request for substituting the odor item of interest.
5. The method of any one of Claims 1 - 4, wherein the dataset is informative of similarity of all pairs of odor items in the given inventory, and wherein at least part of the dataset is generated in advance and stored in a database.
6. The method of any one of Claims 1 - 5, wherein generating the one or more candidate odor items comprises: for each pair constituted by the odor item of interest and another odor item from the inventory of odor items obtaining a plurality of similarity labels; for each odor item (01) pair, applying a parametric model to obtain a similarity score of the 01 pair, wherein the similarity labels of a given 01 pair are used as the features of the applied parametric model and the similarity score is obtained on a continuous scale; and selecting 01 pairs with similarity score matching, at least, a similarity threshold, thereby giving rise to one or more candidate odor items usable for replacing the odor item of interest.
7. The method of Claim 6, wherein the similarity labels of a given 01 pair are obtained by applying to the given 01 pair a plurality of labeling functions, each labelling function resulting from a distinct rule, source and / or type of information, wherein each labeling function is used as a weak classifier to obtain the plurality of similarity labels for the given 01 pair.
8. The method of Claim 7, wherein an applied labeling function is informative of at least one of: perceptual similarity, chemical similarity, similarity of receptors to be involved, similarity of labels provided by domain experts, and proximity of certain numerical features.
9. The method of Claims 7 or 8, wherein an applied labeling function is learnt from the dataset or is represented by a fixed rule.
10. The method of Claim 6, wherein obtaining the plurality of similarity labels comprises: obtaining multiple evaluations of the 01 pairs in inventory of odor items from the same annotators on different occasions; pairing the obtained multiple evaluations into evaluation pairs; generating one or more labeling functions corresponding to one or more evaluation criteria; and applying the generated labeling function to the evaluation pairs to provide similarity labels for the 01 pairs.
11. The method of Claim 6, wherein obtaining the plurality of similarity labels is obtained with the help of graph-based modeling.
12. The method of Claim 11, further comprising: using a base classifier of similarity of odor items in the inventory to build a graph with nodes representing the odor items and with edges connecting odor items classified by the base classifier as similar; using the graph’s properties to generate a new classifier, wherein: two nodes are considered as similar when a number of similar to both common neighbors exceeds a predefined threshold; a node is considered as un-similar to the odor of interest when a number of its neighbors that are un-similar to the odor of interest exceeds a predefined threshold; and applying the generated new classifier to the inventory of odor items to provide similarity labels for the 01 pairs.
13. The method of any one of Claims 1 - 12, wherein the odor pairs are characterized by multiple descriptors, the method comprising for each given 01 pair: obtaining similarity labels for each descriptor of a given 01 pair, separately for each descriptor, building a parametric model with a sum of entropies per descriptor, wherein the similarity labels for respective descriptor are used as the features of the parametric model; and processing together the respectively built per-descriptor parametric models to find the weights minimizing the sum of entropies, thereby defining the multi-descriptor similarity score of the pair.
14. A computing system configured to perform the operations of any one of Claims 1- 13.
15. A non-transitory computer-readable medium comprising instructions that, when executed by a computing system comprising a memory storing a plurality of program components executable by the computing system, cause the computing system to operate in accordance with any one of Claims 1-13.
Citation Information
Patent Citations
Predicting human discriminability of odor mixtures
US20200072808A1
Data processing apparatus, learning apparatus, information processing method, and recording medium
US20220343184A1
Application program, smart device, information processing apparatus, information processing system, and information processing method
US9959292B2