Fragrance and Flavor Production

Machine learning models generate perfume and flavor formulations efficiently by predicting ingredient palettes and concentrations, addressing the inefficiencies of manual trial and error, and capturing expert insights.

JP2026508316APending Publication Date: 2026-03-10GIVAUDAN SA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-28
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

The perfume and flavor formulation process is time-consuming and requires multiple iterations, tiring the perfumer's or flavorist's senses, and wastes valuable ingredients due to the reliance on manual trial and error.

Method used

A method using machine learning models, specifically variational autoencoders and generative adversarial networks, to generate ingredient palettes and concentrations, guided by existing knowledge and insights, reducing the need for manual assembly and iteration.

Benefits of technology

Facilitates efficient and accurate generation of fragrance and flavor formulations, capturing expert preferences and traits without the need for extensive manual effort, thus streamlining the creative process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026508316000001_ABST
    Figure 2026508316000001_ABST
Patent Text Reader

Abstract

A computer-implemented method for training a system for generating ingredient blends for fragrances, the method comprising: receiving an input dataset comprising an ingredient palette and ingredient concentrations of known fragrances and / or flavors; training a palette-generating machine learning model using the input dataset, wherein the palette-generating machine learning model, once trained, is configured to generate at least one generated ingredient palette; and training a concentration-generating machine learning model using the input dataset, wherein the concentration-generating machine learning model, once trained, is configured to generate at least one generated formulation, each generated formulation comprising ingredient concentrations of the generated ingredient palette.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD OF THE INVENTION The present invention relates to fragrances and flavors described by a palette of ingredients and associated ingredient concentrations. More specifically, the present invention relates to methods of training and using machine learning (ML) models for use in creating fragrances and / or flavors, as well as data processing devices, computer programs, and computer-readable storage media for carrying out the methods. [Background technology]

[0002] Background of the Invention In the art of perfume design, the generation of formulations (including both fragrances and flavors), including ingredient palettes and ingredient concentrations, is a highly skilled art. For example, state-of-the-art fragrances are prepared using a palette of fragrance ingredients. Central to the formulation generation process is the perfumer's skill in creating combinations inspired by the ingredient palette. However, modern technology is increasingly being used to aid the creative process. For example, ingredient palettes can be stored in computer databases, allowing formulations to be created and displayed on a computer interface.

[0003] The formula may be displayed on a computer interface in the form of a list of the names and amounts or concentrations of the ingredients involved. Based on the olfactory and / or flavor characteristics of each ingredient and the relative proportions in which they are employed, an experienced perfumer or flavorist will be able to form a reasonable mental impression of the odor and / or flavor of the fragrance, which will aid somewhat in the creative process.

[0004] Once the formulation is complete, a digital signal representing the recipe can be sent to an output device to match, mix and dispense (dispense) the fragrances and flavors for further olfactory or flavor evaluation.

[0005] However, even for the most experienced perfumers and flavorists, the creative process involves multiple iterations of formula adjustments, and complex creations require multiple design, distribution, and olfactory or flavor evaluation steps, all of which are time-consuming, tire the perfumer's or flavorist's nose and palate, and waste valuable fragrance and flavor ingredients.

[0006] Therefore, there is a need for technology that allows perfumers, flavorists, evaluators, and any other parties involved in the formulation creation process to streamline the formulation creation process. Summary of the Invention

[0007] SUMMARY OF THE INVENTION The invention is defined in the independent claims, to which reference should be made. Further features are defined in the dependent claims.

[0008] According to one aspect of the present invention, a method for training a system to generate ingredient combinations (or collections of combinations) for a resulting fragrance and / or flavor is provided. The resulting combinations are suitable for subject review by a human or computer, or may be created directly. The method includes receiving an input dataset, the input dataset including a plurality of ingredient palettes of known fragrances and / or flavors, one ingredient palette per known fragrance and / or flavor. The input dataset also includes ingredient concentrations for the ingredients in the palette of each known fragrance and / or flavor.

[0009] The method also includes training a palette-generating ML model (a first ML model) using at least a portion of the input dataset, the palette-generating ML model being configured to generate at least one latent or predictive productive component palette once trained.

[0010] The method also includes training a concentration-generating ML model (second ML model) using at least a portion of the input dataset. Once trained, the concentration-generating ML model is configured to generate at least one potential or predictive product recipe, each product recipe including the component concentrations of one product component palette as generated by the palette-generating ML model. In this manner, the concentration-generating ML model may be considered to "input" the component concentrations of each product component palette.

[0011] Thus, once trained, the system includes a trained palette-generating ML model and a trained density-generating ML model.

[0012] Optionally, the input dataset includes an end-use for each known fragrance and / or flavor. This allows either the palette-generating ML model or the concentration-generating ML model (or both) to be trained with "end-use" condition parameters. The trained model can then be configured to output a palette and concentrations tailored to a specific end-use. Thus, the input dataset may further include an identified end-use, allowing the trained model to generate a palette and / or concentrations for such identified end-use.

[0013] Optionally, the method allows the computer to accept input of a specified end use, where the specified end use can be limited to known fragrance and / or flavor end uses in the input dataset. The method then retrains or fine-tunes the palette-generating ML model. The retraining occurs using the ingredient palette and ingredient concentrations of the known fragrances and / or flavors associated with the specified end use. The retrained palette-generating ML model (or a model with one model for each end use) generates an ingredient palette for the specified end use. Training a concentration-generating ML model then uses the ingredient palette for the specified end use. Because the original task (training the palette-generating ML model) is similar to this new task (retraining the palette-generating ML model), the retraining process enables accurate generation of a palette for the specified end use without the need to obtain a large dataset specifically for the specified end use. Rather, a complete dataset including all end-use fragrances may be suitable for teaching the model learned traits applicable to all end uses, but then a narrower, end-use-specific dataset refines these teachings.

[0014] Optionally, the palette-generating ML model (or indeed the retrained palette-generating ML model(s)) is a variational autoencoder (VAE). Optionally, the VAE is a conditional VAE (CVAE), where some conditions (e.g., known end uses or creators of fragrances / flavors) may be imposed throughout training to enforce the generation of the palette according to the conditions. Alternatively, the palette-generating ML model (or the retrained palette-generating ML model) may be a (first) generative adversarial network (GAN). Again, some conditions may be imposed on this GAN.

[0015] Optionally, the density-generating ML model can be a bucket predictor (BP). Alternatively, the density-generating ML model can be a (second) GAN.

[0016] Optionally, by way of illustration, when the palette generation ML model is a CVAE, once trained, the palette generation ML model may output a score (e.g., a probability) for the presence of each component in each generated component palette. For each generated component palette (and for each component within that palette), the method may determine to include a component in the palette if the component's presence score exceeds some predetermined threshold. The predetermined threshold may be based on the component's percentage presence in the component palettes of known fragrances and / or flavors in the input dataset. In this way, components rarely used in known fragrances and / or flavors appear with a similar score in the generated formulation. Without such modification, a static threshold (e.g., 0.5) may be used, and rarely used components may not appear in the generated formulation.

[0017] Optionally, if the input dataset includes end-uses of known fragrances and / or flavors, and if the palette-generating ML model is a VAE including an encoder-decoder architecture, the training method may involve passing the end-uses of the known fragrances and / or flavors to multiple layers (optionally all layers) in the encoder. Similarly, the method may involve passing the end-uses to multiple layers (optionally all layers) in the decoder. In this way, conditions (end-uses) are linked to the output data of each relevant layer, and the complete model may be trained with the "intent" that the model learns relationships between inputs with the same end-use. That is, conditional links reinforce end-use information during training. Similarly, conditional links may be used in GAN-style palette-generating ML models, where end-uses are reinforced in multiple layers of the underlying network. Conditional links may also be used in trained models to reinforce end-uses when performing inference.

[0018] Optionally, the cardinality-generating ML model may also implement conditional linkages. Illustratively, the cardinality-generating ML model may include an encoder-decoder architecture, and the conditional (end-use) may be enforced as described above. Similarly, the cardinality-generating ML model may be in the form of a GAN, and the end-use may be enforced at multiple layers of the underlying network.

[0019] Optionally, if the palette-generating ML model is a VAE containing multiple layers, the training method may involve normalizing the output of each layer for each training mini-batch of the input dataset and for each layer of the CVAE, ensuring that the input to each layer is stable and thus ensuring that the training process is efficient.

[0020] Optionally, the method may include training (and / or retraining) both the palette-generating ML model and the concentration-generating ML model, followed by running the models to generate at least one generated formulation. Thus, the trained system enables predictive formulation generation without the need to manually assemble ingredient palettes and ingredient concentrations, a task typically only possible for experienced perfumers.

[0021] Optionally, following the generation of at least one formulation, the method involves quantifying each formulation by its expected odor or olfactory characteristics.

[0022] According to another aspect of the present invention, there is provided a method of generating a recipe (or collection of recipes) of ingredients for a resulting fragrance. The method includes running a trained (or retrained) palette-generating ML model to generate at least one potential (or predictive, or candidate) palette of resulting ingredients. The palette-generating ML model may be trained according to other aspects of the present invention.

[0023] The method also includes running the trained concentration-generating ML model to generate at least one potential (generated) formulation, each generated formulation including one component concentration from the generated component palette. The concentration-generating ML model may be trained according to other aspects of the invention.

[0024] Optionally, following the generation of at least one product formulation, the method may include generating instructions for the creation of the formulation according to the product ingredient palette and associated product concentrations. The instructions may be in the form of human user instructions (e.g., amounts of each ingredient to dispense) or in the form of machine instructions suitable for transmission (via wire or wirelessly) to a machine that creates the formulation.

[0025] Optionally, the method may also include having the computer provide a user interface. The user interface may be suitable for displaying at least a subset of the generated recipes. The user interface may be configured to accept user input, allowing the computer to filter or select a subset of the generated recipes for display. In this way, a user can filter a large generated dataset to only recipes that are relevant or interesting to them.

[0026] Optionally, the method allows for user input (to a user interface) to create the formulation. The user interface provides an alternative graphical shortcut, allowing the user to directly set the formulation conditions without the need for manual input of specific ingredients and concentrations, etc.

[0027] According to another aspect of the present invention, there is provided a method of generating a blend, the method including a training method for training an ML model according to another aspect of the present invention, and a blend generation method using an ML model trained according to another aspect of the present invention.

[0028] An apparatus (computer or computer system) or computer program according to preferred aspects may include any combination of method aspects. Methods or computer programs according to further aspects may be described as computer-implemented in that they require processing and memory functionality.

[0029] Apparatus according to preferred aspects is described as being configured to perform a certain function, arranged to perform a certain function, or simply "performing a certain function." This configuration or arrangement can be through the use of hardware or middleware, or any other suitable system. In preferred aspects, the configuration or arrangement is through software.

[0030] Thus, according to one aspect, there is provided a program which, when loaded into at least one computer, configures the computer to become a device according to any of the above device definitions, or any combination thereof.

[0031] Generally, a computer may include the elements listed as configured or arranged to provide the defined functionality, for example, the computer may include memory, processing, and network interfaces, and input devices.

[0032] The invention can be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or in combinations of them. The invention can also be implemented as a computer program or computer program product, i.e., a computer program tangibly embodied in a non-transitory information carrier, for example in a machine-readable storage device or in a propagated signal, for execution by, or to control the operation of, one or more hardware modules.

[0033] A computer program may be in the form of a stand-alone program, a portion of a computer program, or more than one computer program, may be written in any type of programming language, including compiled or interpreted languages, and may be deployed in any form, either as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a data processing environment. A computer program may be deployed to be executed in one module or multiple modules at one site, or distributed across multiple sites and interconnected by a communication network.

[0034] The method steps of the present invention may be performed by one or more programmable processors executing a computer program that performs the functions of the present invention by operating on input data to generate output. Apparatus of the present invention may be implemented as programmed hardware or as special purpose logic circuitry in, for example, an FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit).

[0035] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor accepts instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions connected to one or more memory devices for storing instructions and data.

[0036] The present invention is described in terms of specific embodiments. Other embodiments are within the scope of the following claims. For example, the steps of the invention may be performed in a different order and still achieve desirable results. Multiple test script versions may be compiled and invoked as a unit without using object-oriented programming techniques; for example, elements of a script object may be organized in a structured database or file system, and operations described as being performed by the script object may be performed by a test control program.

[0037] Elements of the invention are described using terms such as "processor," "input device," etc. Those skilled in the art will understand that such functional terms and their equivalents may refer to parts of a system that are spatially separated but combine to perform a defined function. Likewise, the same physical part of a system may provide two or more of the defined functions. For example, separately defined means may be implemented using the same memory and / or processor, as appropriate. [Brief explanation of the drawings]

[0038] Brief description of the drawings The accompanying drawings are referred to by way of example only. [Figure 1] FIG. 1 is a flowchart of a method for training a system for generating a blend of fragrance ingredients, according to an embodiment; [Figure 2] FIG. 2 is a flow chart of a method for producing a fragrance ingredient blend according to an embodiment; [Figure 3] Figure 3 is a schematic diagram of a CVAE for use in generating pallets; [Figure 4] Figure 4 is a schematic diagram of a GAN for use in generating palettes; [Figure 5] Figure 5 is another schematic diagram of a GAN for use in generating palettes; [Figure 6] Figure 6 is a schematic diagram of a BP for use in generating concentrations; [Figure 7] Figure 7 is a schematic diagram of a GAN for use in generating concentrations; [Figure 8] Figure 8 is a schematic diagram of the condition (end use) linkages applied to CVAE for use in the production of pallets; [Figure 9] Figure 9 is a schematic diagram of the training process of a GAN for use – after training – in generating palettes;

[0039] [Figure 10] Figure 10 is a further schematic diagram of the training process of a GAN for use—after training—in generating concentrations; [Figure 11] Figure 11 is a diagram demonstrating the basic principles of the palette evaluation technique using a linear combination / convex embedding approach; [Figure 12] Figure 12 is a boxplot of the number of components in CVAE-generated pallets across distinct end uses (labeled A-Z) with modified frequency thresholds; [Figure 13] Figure 13 is a boxplot of the number of components in CVAE-generated palettes across distinct end uses with frequency threshold modifications, condition linkage modifications, and batch normalization modifications; [Figure 14] Figure 14 is a 2D KDE plot of the two principal components found in the generative palette embedding; [Figure 15] Figure 15 is a 2D KDE plot of the two principal components found for embeddings generated using loss function modified LDA with component ratio preprocessing; [Figure 16] FIG. 16 illustrates the pre-processing of BOM for the determination of olfactory description; [Figure 17] FIG. 17 illustrates the conversion of component concentrations to intensities for the determination of olfactory descriptors; [Figure 18] FIG. 18 is a diagram of an example of a component representation for the determination of an olfactory description; [Figure 19]FIG. 19 illustrates the representation of an example ingredient (Rose Absolute in Color DM) as a path for the determination of its olfactory description;

[0040] [Figure 20] FIG. 20 is a diagram demonstrating the projection of example formulations onto a classification tree for the determination of olfactory description; [Figure 21] Figure 21 illustrates the calculation of blend similarity for each layer of the classification for the determination of olfactory description; [Figure 22] FIG. 22 illustrates the calculation of the weighted sum of all layers for the determination of olfactory description; [Figure 23] Figure 23 illustrates the clustering of formulations for the determination of olfactory descriptors; [Figure 24] FIG. 24 is a pair of diagrams illustrating the olfactory profile of example formulations. [Figure 25] FIG. 25 is an example of a user interface suitable for providing to a user to allow filtering of the resulting formulations; [Figure 26] FIG. 26 is a further example user interface illustrating the results of filtering; [Figure 27] FIG. 27 is a further example user interface illustrating further results of filtering; [Figure 28] FIG. 28 is another example user interface illustrating the results of a separate filtering process; and [Figure 29] FIG. 29 is a diagram of hardware suitable for implementing embodiments.

[0041] Detailed Description The inventors have come to realize that artificial intelligence techniques are well suited to address the shortcomings associated with lengthy trial and error techniques that rely on the capabilities of highly skilled perfumers and flavorists.

[0042] Generative ML algorithms and associated ML models are suitable for the predictive generation of large numbers of fragrances. The inventors have found that decomposing the task of generating a fragrance into the composition task of generating an ingredient palette and ingredient concentrations reliably generates a large number of fragrances that experts would consider suitable for creation.

[0043] Moreover, the generation of predictive formulations using aspects is "guided" by existing knowledge and insights captured within the training dataset. In this way, predictive formulations can capture the traits, habits, and preferences of experienced perfumers and flavorists that users may not necessarily consider.

[0044] FIG. 1 is a flow chart of a method for training a system to generate a blend (or blends) of ingredients for a resulting fragrance (or fragrances) for user review.

[0045] S10: The computer receives an input data set, the input data set including a palette of ingredients of known fragrances and / or flavors. The input data set further includes concentrations (e.g., mass %, volume %, or mole %) of the ingredients in the palette of each known fragrance and / or flavor. Optionally, the input data set further includes an end use (or multiple end uses, if multiple end uses are relevant) for each known fragrance and / or flavor.

[0046] S12 uses the input data set to cause a computer to train a pallet-generating ML model. Once trained, the pallet-generating ML model is configured to generate a plurality of potential or predictive product ingredient pallets. Optionally, these may be generated for a specified end use.

[0047] S14 uses the input data set to cause a computer to train a concentration-generating ML model. Optionally, the training data here can be coupled with one or more condition variables, such as a specified end use. Once trained, the concentration-generating ML model is configured to generate at least one potential or predictive product recipe, each of which includes a component concentration of one of the product component palettes.

[0048] In one example, the palette generation ML model and the concentration ML model are trained on an on-premises development platform. The trained ML model examples are hosted in an S3 bucket on the AWS cloud service. The trained ML model examples are saved as pickle files.

[0049] FIG. 2 is a flow chart of a method for generating a blend or blends of ingredients for a resulting fragrance or fragrances for user review.

[0050] S20 causes a computer to run the trained (or fine-tuned / retrained) palette generation ML model to generate at least one potential product ingredient palette. Optionally, the product palette is for a specified fragrance end use. The trained palette generation ML model is trained using an input dataset including ingredient palettes of known fragrances and / or flavors and ingredient concentrations of the known fragrances and / or flavors (and, optionally, end uses of the known fragrances / flavors).

[0051] S22 causes a computer to execute the trained concentration-generating ML model to generate one or more potential product formulations. The trained concentration-generating ML model is trained using an input data set. Optionally, the data is linked to one or more condition variables, such as a specified end use. Each product formulation includes component concentrations from one product component palette.

[0052] Model Architecture - Palette Generation The skilled reader will appreciate that many known generative models can be used to generate the ingredient palette.

[0053] As one practical example, variational autoencoders (VAEs) have been shown to work successfully. A VAE is an AI algorithm configured to encode and decode information. When encoding information, a VAE maps a large amount of information into a smaller representation. This condensed representation of information is the VAE's latent space, and the original information is hidden in this condensed representation. A decoder maps the latent space back to the original input. Unlike traditional autoencoders, the "variational" nature of a VAE means that it employs variational (Bayesian) inference to learn the probability distribution of the input data. More specifically, the encoder outputs the mean and covariance corresponding to the posterior probability of the given training data, and the decoder obtains sampled latent vectors from the encoder's output to reconstruct the sampled data.

[0054] A modification shown to generate highly suitable predictive fragrances is that of a conditional variational autoencoder (CVAE). A CVAE is a variation of a VAE for cases where several labels or groups can be associated with each data point (e.g., facial features) and conditional generation (by imposing several labels) is the search goal. The architecture of a CVAE is roughly the same (or at least very similar) as that of a VAE, except that label information is shared by both the encoder and decoder.

[0055] In the practical examples below, the end use of the fragrance (e.g., use in bleach products, use in perfume products) is imposed as a "condition." Of course, alternative (or additional) conditions can be imposed, such as known fragrance creators in the BOM. In this case, the "style" of the master creator is learned by the CVAE, and once trained, the CVAE can be used to generate further fragrances in the style of the specific master creator.

[0056] The CVAE employed by the inventors is a neural network made up of dense layers. The learning phase involves encoding the information carried by the training data (represented by the compositional components along with their concentrations) into a vector that, when decoded, returns the initial composition given as input to the model.

[0057] FIG. 3 is a schematic diagram illustrating the generation of a palette using a trained CVAE.

[0058] The encoder outputs (to the decoder) a randomly sampled abstract vector representation of the blend in a latent space (e.g., dimensionality 100) via a multivariate normal distribution, the parameters of which were learned by the CVAE during training. Optionally, the intended end use(s) of the blend formula can be concatenated to this latent space variable. The decoder outputs a reconstructed palette, e.g., in the form of a one-hot encoded vector representation of the generative blend. Each dimension of this vector relates to a particular ingredient in the ingredient catalog, where "0" denotes an ingredient not used in the generative palette and "0" denotes an ingredient used in the generative palette.

[0059] Generative adversarial networks (GANs) have also been shown to be well suited for fragrance generation. GANs are a class of neural network architectures introduced in Goodfellow et al. (2014). GANs are formed by two neural networks: a generator and a discriminator. Given samples from a known low-dimensional distribution (usually multivariate normal), the generator strives to generate samples from a target distribution. Meanwhile, the discriminator strives to distinguish between which samples are real (i.e., part of the training set) and which samples were generated by the generator. Thus, training a GAN essentially involves solving a minimax-type problem in which the generator attempts to fool the discriminator by generating new plausible examples from the problem domain, and the discriminator attempts to classify the examples as real (from the domain) or fake (generated).

[0060] This means that the generator is not trained to minimize the distance to a particular image, but rather to fool the discriminator, which allows the model to learn in an unsupervised manner.

[0061] More precisely, the two networks (generator / discriminator) are obtained by solving the following minimax problem, equation (1):

number

[0062] As mentioned in Arjovsky et al. (2017), there is a caveat when training these minimax-type problems: mode collapse / vanishing gradient. This issue causes learning instability, especially when the discriminator is ahead of the generator. To avoid these issues, the authors propose an alternative way of posing the problem by considering the Wasserstein distance between the original and generated distributions. More specifically, for the so-called Wasserstein GAN (WGAN), the new minimax problem becomes Problem 2:

number

[0063] There is an important technical detail in the previous formula: the discriminator f φ should be within the Lipschitz-1 function family. To meet this condition, various regularization approaches exist, including weight clipping (which has been found to be less effective) and gradient penalty (Gulrajani et al. (2017)). When the target distribution is discrete, palette adaptation is required for better results. In this case, we apply a probabilistic presence transformation, whose main idea is to convert the presence and absence (denoted by 1 and 0, respectively) of each component in the component palette into a continuous random representation. The probabilistic transformation can be implemented through equation (3), which is applied to each component independently.

number

[0064] Figure 4 provides a schematic diagram of a GAN (and equivalently, a WGAN) for palette generation. A generator network (G) accepts noise (z) and generates palettes of ingredients, forming a generative palette G(z). These generative palettes are assigned labels that indicate the generated (y=0) rather than the real compound palette (y=1). A discriminator network (D) attempts to classify the palettes as real or fake (generated).

[0065] Figure 5 is another schematic diagram of an example GAN for palette generation. Noise from a latent space representation of potential palettes (here concatenated with end-use applications) is passed to a generator network (in this example, size 1024, 1024, 1024) to generate blend palettes. A discriminator network (in this example, size 768, 512, 251, 1) distinguishes these generated blend palettes from the real blend palettes (here also concatenated with end-use applications).

[0066] Model Architecture - Cardinality Generation The skilled reader will appreciate that following the generation of the component palette, many known generative models can be used to generate the component concentrations. The inventors have explored the use of bucket predictors and the use of GANs.

[0067] In Zhang, Isola, & Efros (“Colorful Image Colorization”, 2016), the authors address the computer vision problem of image colorization of black-and-white images. They propose transforming the regression problem into a multi-label classification task by constructing buckets of “color usage,” working on the principle that obtaining a rough estimate of color usage intensity is sufficient to obtain realistic-looking coloring. If an object can take on a set of distinct color values, the optimal solution for the Euclidean loss will be the average of that set. In color prediction, this averaging effect favors gray-toned, desaturated results. The inventors came to realize that a similar technique using bucket predictors (BPs) is well suited to generating component intensities (rather than color “intensities”). That is, we can rephrase images as blends, pixels as components, and colors as component proportions. Note that there is also diversity in component usage.

[0068] BP is essentially an encoder / decoder, where the input is a one-hot encoded version of the components (i.e., the palette), and the output is, for each component, an assignment of the bin on a logarithmic scale to which that component's concentration belongs.

[0069] Figure 6 is a schematic diagram of an example of BP for cardinality generation. A compound palette of size n components is passed to a fully connected encoder / decoder architecture (the palette may be concatenated with the end use, in which case the size is instead n components + m labels). The example BP encoder reduces the input to a latent representation of size 100 through a layer of size 300. The example BP decoder then increases the processed data to size 300, and then to a size corresponding to the total number of components. A softmax function is used as the activation function to normalize the output to a score representation (e.g., probability) of the distribution bucket of the components.

[0070] We demonstrate the use of a GAN (e.g., a WGAN) to assign component concentrations to a generated palette. Similar to BP above, this approach is also inspired by image colorization-type problems: in one example, the "noise" input includes a palette extracted from the latent space, or a palette previously generated by a GAN trained for palette generation, coupled with the fragrance end use.

[0071] Figure 7 is a schematic diagram of an example GAN for concentration generation. A blend palette (here linked with the end use) is passed to a generator network (in this example, size 1024, 1024, 1024) to generate blend concentrations. A discriminator network (in this example, size 768, 512, 251, 1) distinguishes between these generated blend concentrations and the true blend concentrations (here also linked with the end use).

[0072] Training Data The skilled reader will understand that the specific structure of the training data may vary depending on the exact architecture and training process used in implementing the embodiment. In this case, the inventors utilized an internal bill of materials (BOM) that provided a dataset of a collection of known formulations. Known formulations as determined by experienced expert perfumers.

[0073] BOM datasets have different levels of detail, with "Level n" providing all the individual raw materials, while "Level 1" provides the ingredients and bases (mixtures of ingredients) available to the perfumer. In the example below, BOM Level 1 is used. In this case, BOM Level 1 contains the following schema: - group_code: A text variable containing the identification code for each recipe. - ingredient_or_group_code: A text variable containing the code for each ingredient. - concentration_over_100: A float8 variable that is the concentration multiplied to a total of 100.

[0074] The BOM is provided in a CSV table format, containing approximately (after pre-processing) 150,000 known fragrances (although, for example, 5,000 or 50,000 fragrances may be used; the same applies to flavors) and the constituents of each fragrance and the concentration of each component.

[0075] The specific steps involved in pre-processing the data before training on practical examples are: - Rename all solvents to one generic name (fragrance solvents are merely "carriers" for the active ingredients; inactive or non-scented ingredients are removed to focus on the ingredients that provide the scent). - Removing functional ingredients (i.e. ingredients that do not provide any fragrance contribution but instead relate, for example, to color or texture and can therefore be ignored for the purposes of fragrance generation). - Eliminate formulas that use other formulas as ingredients. - Eliminate rare ingredients (ingredients used in fewer than 10 recipes).

[0076] More generally, pre-processing may be used to perform any or all of the following steps: - Select a valid formulation (e.g., the following formulation: concentrations total up to 100% by weight; identification code format is expected (e.g., 3 letters followed by the number 3 followed by 3 letters); formulation is non-recursive; formulation has enough ingredients to qualify as a formulation). - Clean up ingredients (e.g. remove rare and odorless ingredients; combine all solvents into one common solvent; combine duplicate ingredients; remove formulations with an insufficient number of ingredients; renormalize when necessary, e.g. by weight %; replace identification group codes with names when possible). - Creating features for training the ML model (e.g., establishing each formula as associated with an array of all ingredients and ingredient concentrations used in the BOM, setting the former to zero if the ingredient is not present in the formula). - Split the data into training and test sets, or training, test and validation sets (for example, in the ratio of 75:25 for the former, respectively, or 80:10:10 for the latter, respectively).

[0077] Of course, if the training data is already in a format suitable for training, some or all of the pre-processing steps may not be necessary.

[0078] In addition to the BOM, various data tables can be collated to quantify the resulting formulation. Illustratively, in this example, a table of odor descriptors for each ingredient can be used. Various systems for quantifying descriptors are available; in this case, a two-level approach was utilized, describing ingredients by their "family" and "characteristic factor" (stored as text variables in this example). The ingredient family and characteristic factor(s) relate to the olfactory characteristics of the raw ingredient. Each raw ingredient has one or two characteristic factors, typically assigned manually by an experienced perfumer.

[0079] Such information at the ingredient level can be used to derive a quantitative olfactory description of the formula (see below for a description of an example implementation). In this example, olfactory features are used during the post-processing stage of the pipeline (after formula generation) to provide additional information. Of course, the skilled reader will understand that quantitative olfactory information (along with the corresponding ingredients) can be used in the training process.

[0080] In addition to the BOM, in which the formulas are expressed as ingredients and concentrations (for each formula, there are as many rows as there are ingredients with their respective concentrations), other tables of data include a table of unscented ingredients (used in one example during BOM pre-processing) and a table of ingredient prices. The latter can be used to quantify the resulting formulas in terms of formula prices. The skilled reader will also appreciate that in other examples, price information (along with the corresponding ingredients, for example) can be used in the training process.

[0081] Model architecture modifications Frequency Threshold For generation using the CVAE architecture / process, each output unit represents a score (e.g., a number between 0 and 1) of the presence of the represented component in the predictive palette. A straightforward approach to including or excluding a component in the predictive blend is to consider a value greater than 0.5 to indicate presence. However, this approach ignores the distribution of component presence within the training data.

[0082] Alternatively, the CVAE approach can be modified so that the presence threshold is specific to each component in the palette; for each component, the threshold can be set equal to the frequency of occurrence of the component in the training set. With this approach, components that occur less frequently in the training dataset should occur with the same frequency in the generated palette. In contrast, with a static threshold of 0.5, these components will rarely occur in any of the generated palettes, despite being present (albeit rare) in the training data.

[0083] conditional linkage In a conditional model, the input typically includes a blend and a label that provides context information (e.g., the end use of the blend). By including this context information, the model can be trained with the intent of learning relationships between inputs with the same label. In many generative models, illustratively in the CVAE and bucket predictors used for palette generation and concentration generation, respectively, context labels may be used at the model's input. To enhance the context information about the input throughout the training process (and thus the inference process), the inventors have come to recognize that the context labels may be reintroduced at various processing stages of the underlying network (e.g., not just as inputs to specific modules of the encoder and / or decoder).

[0084] Figure 8 is a schematic diagram illustrating the conditional connections at each step of the CVAE architecture. The conditional variables (the blend's end-use; here represented as a vector of three values, with the shaded boxes indicating the one-hot encoding of the end-use) are concatenated with the blend component palette as input to the encoder block of the autoencoder architecture. At each layer of the encoder, the conditional variables are concatenated with the output of each layer of the neural network. Similarly, once the latent representations are obtained by the encoder (represented by means and covariances corresponding to the posterior probabilities given the training data), the decoder block of the autoencoder architecture concatenates the conditional variables with the output of each layer of the neural network. As shown in this simple diagram, the reconstructed blend palette is consistent with the (original) blend palette (within the palette, blank boxes indicate the absence of a component, and shaded boxes indicate the presence of a component).

[0085] In summary, the condition labels (e.g., end-use) are one-hot encoded (i.e., converted into a vector filled with zeros except for one coordinate, which is assigned the value “1” and refers to the end-use according to its position) and concatenated into an embedding (abstract compressed representation) of the BOM in the latent space of the CVAE.

[0086] The skilled reader will appreciate that conditional linkage can be applied to GAN-based architectures as well.

[0087] Batch normalization Batch normalization (Ioffe & Szegedy, “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift”, 2015) is a technique that can be used to speed up the neural network training process without sacrificing model performance by normalizing the output of each layer.

[0088] During the training process, for each mini-batch of input, and for each layer, an activation matrix H is obtained and normalized to H', so that each column is normally distributed with mean 0 and variance 1. Before passing this information to the next layer, it may be reparameterized by the learned parameters γ and β, which represent the new variance and mean, respectively:

number

[0089] Conditional GAN Similar to the conditional VAEs discussed above, the inventors recognized the benefits of enforcing conditions (i.e., the fragrance use case in this case) throughout the training and inference stages when using GAN architectures. These techniques consider the goal of enabling conditional generation, where conditions are fed into both the generator and the discriminator. A variation of this (general—not specific to the domain of fragrance generation) was proposed in Mirza & Osindero (“Conditional Generative Adversarial Nets”, 2014).

[0090] Image colorization using GAN In Nazeri et al. ("Image Colorization using Generative Adversarial Networks", 2018), the authors use GANs for image colorization. In this context, the "noise" distribution is the distribution of grayscale images, while the target distribution is the distribution of color images.

[0091] training For each example ML architecture above, conventional training techniques can be implemented (i.e., e.g., VAE / CVAE / GAN for palette generation, BP / GAN for cardinality generation).

[0092] As an example, several architectures and hyperparameters were tested using BOM as training data and readily available GPUs and CPUs. Training a WGAN for ingredient palette generation using the example training dataset takes approximately 1 hour using a GPU. One model that performs well (model output is evaluated using a Random Forest classifier; see the section titled "Evaluation" below) includes the following parameters: - Model type: WGAN with gradient penalization - Latent space dimension: 30 - #Layer:5 - #units: [768,512,256] / [256,512,768] - Normalization: Batch normalization, but not a regularization technique. - Generator output activation: Sigmoid(0,1)

[0093] Of course, other parameter values ​​are suitable. By way of example, as shown in Figure 4, a generator of # units [1024,1024,1024] and a discriminator of # units [768,512,251,1] is suitable and works well.

[0094] Figure 6 is a schematic diagram of the training process of a GAN (and equivalently, a WGAN) for palette generation. The fragrance dataset provides real samples for use by the discriminator (disc / critic). A noise generator inputs noise into the generator network, which generates fake samples that are provided to the discriminator (disc / critic). During training of the discriminator network, the discriminator classifies both real data and fake data from the generator. The discriminator loss function (disc / critic loss) penalizes the loss incurred by the discriminator when it misclassifies real instances as fake or fake instances as real instances. The discriminator updates its weights through backpropagation from the discriminator loss through the discriminator network.

[0095] The generator can be trained by the following steps: sample random noise; create a generator output (i.e., a generated fragrance) sampled from the random noise; obtain a discriminator classification (real or fake) of the generator output; calculate a loss from the discriminator classification; backpropagate through both the discriminator and the generator to obtain a gradient; and use the gradient to change only the generator weights. This is a single iteration of training the generator; the complete GAN training process then alternates between training the discriminator for one or more epochs and training the generator for one or more epochs until the desired convergence is achieved.

[0096] In this example case, data is output at various stages (via the DBoard module) for user observation and assessment. Notably, end-to-end evaluation can be performed using real and fake samples; illustratively, the module "E2EValidator" provides out-of-bag (OOB) error calculations performed at each validation step to keep track of training. See the "Evaluation" section below for a discussion of this end-to-end evaluation.

[0097] Note that the schematic shown for training a GAN is equally valid for the training process for concentration generation if the noise generator is replaced by a real palette (taken, as an example, from a fragrance dataset).

[0098] Figure 10 is a schematic diagram of the training process of a GAN (and equivalently, a WGAN) for concentration generation. Here, a WGAN is used to assign concentrations to a generated palette. The noise input includes the palette previously generated by the trained WGAN, concatenated with fragrance end uses. The palette is written as a 0-1 vector of dimension equivalent to the length of the list of available ingredients (approximately 1600 in this case).

[0099] The generator architecture is a multi-layer perceptron (MLP) with dimensions [1024, 1024, 1024] and activation functions (between layers) [ReLU, ReLU, Sigmoid]. Since some components are not present at the input but have positive cardinality, they are "muted" after the last layer using a mask (input dependent). The generated cardinality (MASK) is based on the input one-hot encoded palette. Palettes with generated cardinality are assigned the label y=0.

[0100] The discriminator input includes blends with concentrations (real blends with a label of y=1, and generated fake blends). The real blends are described by vectors whose coordinates are the concentrations of the blend's components. The discriminator contains three hidden dense layers of size [728, 256, 128] intertwined with a ReLU activation function. The final layer, size [1], returns a score indicating the label (i.e., fake or real) that the discriminator assigned to each blend.

[0101] Regarding the GAN training process of the GAN for palette generation, the optimization of the discriminator and generator weights is performed as described above.

[0102] Model evaluation After successfully generating predictive ingredient palettes, the inventors evaluated the generated palettes in terms of their similarity to actual ingredient palettes assessed by expert perfumers. For the palettes generated by the VAE and CVAE, the inventors evaluated the embeddings (latent representations) obtained by the trained encoders. An embedding in this context is a high-level representation of a palette that highlights the specific details of the processed data.

[0103] The role of these expressions in the developed evaluation techniques can be expressed in three ways. - Be able to explore and / or understand the information encoded by the VAE that generates the palette and visualize the different groups formed by the VAE representation. - Allows visualization of data distributions according to the various generation algorithms previously proposed. - Allows for the identification of any relationships between ingredients and fragrance end uses, as well as relationships between end uses.

[0104] In short, the main questions behind the evaluation of palette generation can be summarized as follows: - Can VAE (or CVAE) representations distinguish end uses? If not, what type of information is encoded? - What part of the pallet distribution is successfully covered by the pallet generation technique?

[0105] More conventional analysis techniques for latent representations using "classical" techniques (principal component analysis, PCT; t-distributed stochastic neighbor embedding, t-Sne; random projection, etc.), in which a cloud of points in the latent representation is plotted, have been found not to provide clear structure or instruction. Illustratively, assigning a color to each point according to some label (e.g., end use) has not been found to be useful in evaluating generative palettes.

[0106] The first evaluation technique for this goal is based on the technique proposed by Liu & Wang (“LatentVis: Investigating and Comparing Variational Auto-Encoders via Their Latent Space”, 2020), which uses latent representations. The authors propose applying a linear transformation to the latent representation obtained by a VAE to capture the semantic direction of the data in the latent representation. The main idea is to obtain this linear transformation by training a linear classifier that predicts several labels (defining “semantic”) from the latent representation of the data. This linear transformation is then used to define the semantic direction in the latent space.

[0107] The second evaluation method uses a linear combination / convex embedding approach instead of operating on the palette directly. The main idea is that the palette is a proportion (p) i For i with end use j, we define the components and end use as plane (v i ) i’ (v e ) e is to represent it as a vector in ∈R2:

number

[0108] Figure 11 demonstrates this second evaluation approach in a schematic manner.

[0109] Of note, there are two related but distinct goals: generation (for example, of an ingredient pallet) and conditional generation (where an end use is imposed on the generating pallet).

[0110] general generation The performance of the generative model can be assessed using Random Forest classification, where a classifier is trained to distinguish between the original and generated palettes. In this regard, two possibilities are considered and explored: - Use the presence or absence of a component as input for the classifier; - Use a bucketized version of the palette (buckets constructed from the proportions of each component) as input to the classifier; this prevents the classifier from relying on the fine details of the values ​​of each concentration. In other words, the classifier cannot rely on the output of the bucket predictor. Therefore, in order to compare real and generated data, we need to ensure that the outputs are all expressed in the same way, i.e., in buckets rather than the actual values ​​of each concentration.

[0111] More precisely, the OOB score of the Random Forest classifier can be determined; the lower the OOB score, the better the generator performs.

[0112] Additionally, for each generated palette, the palette can be evaluated using the Jaccard index (or Jaccard similarity coefficient, distance, or intersection of unions, IoU). For each generated palette, the minimum Jaccard distance of the palette to the training set is calculated as the evaluation criterion. This distance is based solely on the presence or absence of components.

[0113] Condition Evaluation To evaluate how well any conditional generator achieves its goal (i.e., generating pallets for a specified end-use), the inventors trained an end-use classifier on the original pallets. After generation, the accuracy of the end-use prediction of the generated pallets can be measured. That is, if a classifier trained on the original (authentic) data can properly predict the end-use of the generated pallet (e.g., with high accuracy), then the generated pallet can be said to have a "signature of the appropriate end-use."

[0114] Here, the entire pipeline is evaluated in two different ways: the first (quantitative) way as described in the general generation section above (i.e., using the OOB scores of the fake / real classifier based on the bucketed palette); and the second (more qualitative) way, which compares a given distribution from the real palette with that obtained from the generated palette. These are effectively "sanity checks" and involve creating: - Histogram with the number of ingredients per palette; - Boxplot of number of ingredients per end use.

[0115] Generally speaking, the conditional palette generator has been found to be more effective than the non-conditional generator, as quantified in terms of OOB scores and taking into account qualitative comments from experienced perfumers.

[0116] Table 1 below demonstrates the accuracy of the OOB scores and end-use classification of generated palettes generated using CVAE. Various modifications to the underlying CVAE approach have been explored: "T" indicates the use of frequency thresholds; "C" indicates the use of conditional concatenation (at various layers of the CVAE architecture); and "B" indicates the use of batch normalization. Combinations of modifications are indicated by both (or all) of the relevant letters (e.g., "TB" or "TCB").

[0117] Table 1. Evaluation of the palette generation capability of the CVAE model [Table 1]

[0118] In terms of palette generation, all modifications are shown to provide an improvement in the OOB score (i.e., the classifier has a harder time distinguishing genuine palettes from generated palettes using the CVAE with TCB modifications than using palettes generated using the CVAE without modifications). The most significant modification is shown to be the use of a frequency threshold.

[0119] Modification of the TC combination has been shown to provide the most accurate classification potential of the generated palette in terms of end-use. More generally, the use of conditional linkages (C) effectively forces the generator to respect the imposed end-use; this is consistent with the idea that sharing information about end-use between different layers of the CVAE increases the "memory" of this feature.

[0120] Figure 12 is a boxplot demonstrating the distribution of the number of ingredients in the generated palette across distinct end uses (labeled A-Z), where the palette is generated using CVAE with frequency threshold modification (CVAE+T). The OOB score here is 0.57; for reference, CVAE without modification shows an OOB score of 0.66. Note that for this test, the training dataset is distinct from that used for the above test (i.e., a distinct BOM is used). The distribution of ingredients per palette shows, in most cases, smaller palettes (in terms of number of ingredients) than the authentic distribution. However, there are some irregular cases, such as end use E (corresponding to bleach), which show more ingredients than expected. Experienced perfumers indicated that the palette appears authentic, even though the conditional end uses are not always respected.

[0121] Figure 13 is a boxplot demonstrating the distribution of the number of ingredients in the resulting palette across distinct end uses, where the palette is generated using CVAE with frequency threshold modification; conditional linkage modification; and batch normalization modification (CVAE+TCB). The OOB score here is 0.56. Compared to CVAE+T, the distribution of ingredients per palette is broader, resulting in a more realistic pattern. For fine fragrances (end uses M and W), the number of ingredients appears accurate; however, there is a larger, unrealistic pattern for bleach (end use E).

[0122] Pallet Size As a further criterion, we investigated the number of palettes generated to achieve a given component subset. This study is used to assess the completeness of the set of generated palettes using GANs, but the findings are applicable to any generative model. Given a component subset {k1, k2, ..., k m Given}, we may consider how many palettes should be generated to obtain at least one palette with a subset of the ingredients. By considering that each generated palette induces a Bernoulli variable, we can obtain a simple estimate: a palette either contains the subset of ingredients or it does not. This estimate requires estimating the following probabilities:

number

[0123] Using this estimate at hand, we can obtain a simple estimate N(k) of the minimum number of palettes to be generated, during the process (using a probability greater than β) we obtain at least one palette with a subset of the components:

number

[0124] It is worth noting that the aforementioned problem of estimating probabilities is somewhat as difficult as the problem of estimating / generating densities. Furthermore, for each component subset, there exists a different N: the larger the subset size, the larger the estimated N. Finally, the estimation depends heavily on the generating dataset; some component estimates have better estimates than subsets with extremely low probability.

[0125] There are two approaches to estimating N: first, a joint empirical estimate, where one can calculate the ratio between the number of pallets in which a component is present (in the generating dataset) and the total number of generated pallets, J; and second, an empirical estimate for each component, where, for each component k, one estimates the probability that this component is present in the generating pallets using the quotient of the number of pallets in which component k is present divided by J. Then, assuming that component presence is independent, one can bound the target probability by the product of the presence probabilities of each component in the subset. This approach provides a pessimistic bound because it takes into account existing correlations between components.

[0126] In a generated dataset of J = 250,000 pallets (10,000 for each of the 25 possible end uses), for the 4 arbitrary component subsets (benzyl acetate, hedione, florhydral, and peach pure), the joint empirical estimation approach provides estimates for 173 pallets. The empirical estimation for each component approach provides estimates for 22,659,741 pallets.

[0127] Similarly, for the same generative dataset J, for any of the 5 components (Aubepine Paracresol, Yara Yara, Emonil, Evanol, and Galvanone 10), the joint empirical estimation approach provides estimates of 1,496 pallets. The empirical estimation of each component approach provides estimates of 24,071,494 pallets.

[0128] Experienced perfumers note that the first approach (the joint empirical estimation approach) provides numbers that follow generative testing of "traditional" (i.e., not supplemented with AI) fragrances.

[0129] Embedded Evaluation Returning to the first evaluation technique mentioned above (i.e., using a latent representation of the palette; in this case obtained by a VAE or CVAE architecture), the first evaluation technique is used to “guide” the two-dimensional representation in the VAE / CVAE embedding by explicitly using end-use information. That is, the representation is obtained by training a Linear Discriminant Analysis (LDA) classifier that predicts the end-use and retaining two key components:

[0130] In this case, the performance of this classifier is relatively poor (accuracy: 0.51, predicting end-use 14), but this linear transformation of the latent representation can partially separate some end-uses, meaning that some of the information passed through the encoder is retained in the latent representation obtained by the VAE.

[0131] Following the initial implementation of LDA, the inventors recognized that modifications are possible. First, it is possible to pre-process the proportion of use of each component. Each component has a different characteristic usage scale (i.e., each component has a "proper" way of being used across the palette; illustratively, some components are typically used with relatively high concentrations, while other components are usually used with relatively low concentrations). However, what is important is the relative use of the component with respect to its use within the entire dataset. Thus, instead of using the original proportions, LDA can be modified using:

number

[0132] Second, it is possible to apply regularization terms to the representations of the end-uses. The first regularization term (statement (9)) can be used when we are far from the origin and have sufficient distance between the end-uses: this prevents all representations from converging to the origin and ensures that all end-uses are "close" to each other.

number

[0133] A second regularization term (statement (9)) can be used to ensure that the representations do not stray too far from the origin (otherwise the dominant term would be the earlier term and all representations would grow without bound). This is therefore a way of controlling the relative distance from the origin of each end use; it ensures that they are all roughly the same distance from the origin:

number

[0134] Thus, LDA can be modified for training using the following loss function (Equation (10)):

number

[0135] Figure 14 is a 2D KDE plot of the two principal components found for the embedding, where LDA uses the modified loss function described above, but without preprocessing the usage percentages for each ingredient. Representations of six distinct end uses (O, H, D, L, F, Y) are plotted. In this plot, crosses represent palette end uses, and dots represent ingredients within the palette embedding (linear combinations of the palette's ingredients). Experienced perfumers will note that end uses are "fixed" according to ingredient usage (ingredients are generally used more for certain end uses than others).

[0136] Figure 15 shows a 2D KDE plot of the two principal components found for the embeddings, where LDA uses the modified loss function above, with preprocessing of the usage fraction for each component. Qualitatively, these embeddings (and thus their generated palettes) are similar to those shown in Figure 14.

[0137] In summary, the above evaluation techniques demonstrate that palettes generated using various generative ML models are similar to conventionally prepared fragrance palettes. In addition, when an end-use is imposed, the generated palettes match expectations (i.e., for example, a palette generated for end-use A is difficult to distinguish from a known palette for end-use A). That is, palettes generated for a particular end-use are distributed (with respect to each other) in a consistent manner in latent space. In turn, AI-generated palettes are suitable for use in creating fragrances.

[0138] Complete formulation (palette and consistency) evaluation In addition to the above computational evaluation of the product palette, the inventors also investigated the results of the complete formula generation in a qualitative manner: the product formula (including the product palette and associated product concentrations) was compounded into actual fragrances. The product formula selected for testing by the perfumer is sent to the creation software, where the perfumer can modify the formula if necessary, further validate the formula, and generate and deliver instructions for compounding. The compounding is performed robotically, assisted by a manual lab operator if necessary depending on the complexity of the mixing process.

[0139] Table 2 below summarizes the evaluation results of the 10 formulated product formulas. The left column indicates the filter / search applied to the large dataset of product formulas. For example, the top entry shows only product formulas where the dataset was filtered to the end use "female fine fragrance" and where the formula's olfactory signature was calculated as "Family: Fruity" and "Characteristics: Candied Fruit, Raspberry, Blackcurrant" (see below for a discussion of the formula's olfactory signature). As shown in the bottom row, other filters such as the maximum number of ingredients in the palette and the presence of specific ingredients in the palette can also be employed. The middle column indicates the combination of models used to generate the formula. For example, in the top entry, a GAN was used for the palette generation ML model and a bucket predictor was used for the concentration generation ML model. VAER refers to the retrained VAE model. The right column indicates the qualitative evaluation results from an experienced perfumer.

[0140] Table 2. Evaluation of the resulting formulation [Table 2]

[0141] Practical examples The inventors trained at least five models (palette and concentration ML models) using the pandas Python software library.

[0142] The training data (in the format described above) is pre-processed as described above. That is, given the raw BOM of a compound, pre-processing creates tables in the database with cleaned data. Additionally, pre-processing prepares the data for use with Streamlit, an app framework suitable for creating web abbs. Streamlit processing creates olfactory descriptors of the original compound (see below) and the external schema.

[0143] We train a VAE and a GAN palette generation model to generate a palette of fragrances. The inputs for these two models are: - train_set(PandasFormulas): The formulas to be used for training, as specified in catalog.yml under bom_train_formulas - test_set(PandasFormulas): Formulas used for validation, as specified in catalog.yml under bom_test_formulas - parameters(dict): Hyperparameters of the model, present in parameters.yml

[0144] The output of these training processes is: - models: Trained models are saved as pickle files as specified in catalog.yml under vae and gan_presence respectively. - ingredients: Under ingredients, a list of all the ingredients used, stored as pickle files as specified in catalog.yml - labels: A list of all end uses used, stored as pickle files as specified in catalog.yml, under labels

[0145] BP and GAN concentration generation models are trained to generate concentrations for each component given the palette (presence of given components) and end use. Additionally, a Random Forest is trained and its OOB_score is saved and used as an evaluation metric. The inputs for these models are: - train_set(PandasFormulas): The formulas to be used for training, as specified in catalog.yml under bom_train_formulas - test_set(PandasFormulas): Formulas used for validation, as specified in catalog.yml under bom_test_formulas - parameters(dict): Hyperparameters of the model, present in parameters.yml

[0146] The output of these training processes is: - models: The trained models are saved as pickle files as specified in catalog.yml under bp and gan_concentration respectively. - bp_oob: Under bp_oob, a table with OOB_score as specified in catalog.yml - ingredients: Under ingredients, a list of all the ingredients used, stored as pickle files as specified in catalog.yml - labels: A list of all end uses used, stored as pickle files as specified in catalog.yml, under labels

[0147] Additionally, a CVAE retraining model is trained (or fine-tuned). In this case, this is a model that contains multiple VAE models, each retrained for a specific end use. To run the CVAE model, a VAE model must first be trained. Additionally, a Random Forest is trained and its OOB_score is saved and used as an evaluation metric. The inputs for this training process are: - VAE models to be stored as specified in catalog.yml under model:vae - train_set(PandasFormulas): The formulas to be used for training, as specified in catalog.yml under bom_train_formulas - test_set(PandasFormulas): Formulas used for validation, as specified in catalog.yml under bom_test_formulas - parameters(dict): Hyperparameters of the model, present in parameters.yml

[0148] The output for this training process is: - vae_retrained: The final use dictionary of the retrained VAE model, which is saved as a pickle file as specified in catalog.yml under vae_retrained - retrain_score: A pandas dataframe with information about the retraining. It contains the OOB_score before and after retraining for each end use. This is saved as a csv as specified in catalog.yml under retrain_score.

[0149] For blend generation, there are many combinations of models that can be used. For each combination of models, there is a pipeline that can generate a blend (palette and concentrations) by combining the models. In each pair of models, the first is used to generate the palette and the second is used to assign concentrations. For example:

[0150] CVAE-BP and CVAE-GAN - input: ○ parameters(dict): Parameters for creation that exist in parameters.yml ○ vae(CVAE): A trained CVAE palette generation model. ○ bp(FFPropPredictor): A trained concentration assignment model (BP or GAN). ○ train(PandasFormulas): The formulas from training used to calculate the Jaccard score. test (PandasFormulas): The formula from validation used to calculate the Jaccard score. prices(PandasDataFrame): Prices for each component retrieved from the DataBase.

[0151] - output: ○ formulas(PandasDataFrame): In catalog.yml, the path where the generated formulas are saved is specified. For formulas generated by CVAE-BP, it is under formulas. For formulas generated by CVAE-GAN, it is under vae_gan_formulas. ○ formulas_info(PandasDataFrame): Returns the Jaccard score, price, label, Mahalanobis distance for each formula, and `git_commit`, `model_name`, and `date` for each model. In catalog.yml, the path where this information is stored is specified. For CVAEBP, this is stored in the path defined in formulas_info. For CVAE-GAN, this is stored in the path defined in vae_gan_formulas.

[0152] GAN-BP and GAN-GAN - input: ○ parameters(dict): Parameters for creation that exist in parameters.yml ○ palette(GAN): A trained GAN palette generation model. concentration_filler (GAN or BP): a trained concentration assignment model ○ train(PandasFormulas): The formulas from training used to calculate the Jaccard score. test (PandasFormulas): The formula from validation used to calculate the Jaccard score. ○ prices(PandasDataFrame): The prices of the ingredients to calculate the final price of the formula.

[0153] - output: ○ formulas(PandasDataFrame): In catalog.yml, the path where the generated formulas are saved is specified. For formulas generated with GAN-BP, they are stored under gan_bp_formulas. For formulas generated with GAN-GAN, they are stored under gan_gan_formulas. ○ formulas_info (PandasDataFrame): Information about the model (git_commit, model name, date) plus all the information needed for each formulation (Jaccard score, price, labels). In catalog.yml, the path where this information is stored is specified. For formulations generated with GAN-BP, it is under gan_bp_formulas_info. For formulations generated with GAN-GAN, it is under gan_gan_formulas_info.

[0154] CVAER-BP - input: ○ parameters(dict): Parameters for creation that exist in parameters.yml vaes(Dict[str, VAE]): A dictionary of VAE models, the key of which is the end use for which the model was retrained. bp (FFPropPredictor): the model used to determine the concentrations of the ingredients selected by vaestest (PandasFormulas), which is used to determine the fragrance name given to the index.

[0155] - output: cvaer_bp_formulas: A DataFrame containing three columns: "formulas": an integer representing the created formula, "ingredient": the name of the ingredient used in the formula, and "concentration": the concentration. It is saved in the path defined in cvaer_bp_formulas in catalog.yml. cvaer_bp_formulas_info: Returns the Jaccard score, price, label, and Mahalanobis distance for each formula, as well as the 'git_commit', 'model_name', and 'date' for each model. The results are saved in the path defined in cvaer_bp_formulas_info in catalog.yml.

[0156] In the above, the Mahalanobis distance of each blend is the Euclidean distance between the embedding of the generating blend for a given end-use and the center (mean) of the region of the latent space covered by blends for the same end-use.

[0157] After generating the blends, the odor descriptors (for "family_1" and "feature factors") for each blend are calculated.

[0158] After the odor descriptors are calculated, the locally stored fragrances (original and generated) and related information are uploaded to the DB.

[0159] To simplify this process, we implemented the Kedro pipeline structure, a framework for creating modular, reproducible, and maintainable data science code. The entire process of generating a formulation (data preprocessing, model training, palette generation, and concentration addition) can be expressed using the Kedro-viz app.

[0160] The code can be divided into three parts. 1. Data preprocessing & Streamlit data preparation 2. Training the AI ​​model 3. Generate new formulations

[0161] In this working example, the Kedro pipeline is executed in the order given above.

[0162] Olfactory description of the composition After the generation of the formulations (ingredient palette and ingredient concentrations), techniques can be implemented to quantify the olfactory contribution of each formulation.

[0163] The inventors have developed a technique to perform this function. Among other things, this technique provides a means to quantify fragrance families and fragrance feature factors (two ways of describing fragrances) in numerical terms. In summary, this technique involves pre-processing the BOM data; rescaling concentrations to intensities; building component representations; representing components as paths in a classification tree; projecting the formulation composition onto the classification tree; calculating formulation similarity at each layer of the classification; aggregating layers into a global formulation similarity; clustering the formulations using a K-nn label propagator; representing clusters for the user; and naming and querying clusters. In one implementation, this technique is provided in a pandas-based Python framework.

[0164] In broad terms, this technique is based on ingredient concentration statistics for each end use. As an illustration, if the concentration of an ingredient is higher than typical for an end use, then it can be concluded that the perfumer wanted to emphasize this ingredient, and therefore the characterizing factor of that ingredient can be set as one of the characterizing factors of the specific formulation in question.

[0165] Pre-processing Figure 16 is a schematic diagram of the effect of pre-processing on the BOM at level 1. The pre-processing obtains a renormalized BOM at level n. This pre-processing involves the following steps: - Remove non-consumer product formulations and remove "unconventional" end uses (such as fine fragrances) - Full BOM expansion (i.e. converting the sub-formula codes in the BOM into a list of all ingredients contained in the BOM) - Handling formulations that use unclassified ingredients: ○ Removal of components when concentrations are low enough If not, delete the formula - Removal of technical (e.g., odorless) ingredients Renormalize so that the sum is again up to 1 (e.g., renormalize the weight % of the remaining components after removing any potential components).

[0166] Rescale density to intensity Figure 17 is a schematic diagram illustrating the rescaling of concentrations (of components) to intensities. This process involves: - Change the concentration to a logarithmic scale based on 3 - Calculate the mean and standard deviation of these log-scale concentrations - Calculate the "deviation" from the mean by rounding up to 0.

[0167] This deviation can be interpreted as "how many standard deviations above the mean is the concentration?" If the concentration is below the mean, then the deviation is 0. Intensity refers to the additive contribution of each component of the formula to the olfactory profile. The contribution depends on the concentration of each component in the formula, compared on a logarithmic scale to the average concentration of the same component in all formulas. The idea is that a higher-than-normal dose of a component in a formula will impart its olfactory profile to the formula as a whole.

[0168] Specifically, the following equation 11 is used to obtain the intensity from the concentration of each component:

number

[0169] Building component representations An example of an ingredient display is shown in Figure 18. The representation includes a family classification tree and a note classification tree for a particular ingredient.

[0170] Representation of components as paths in a classification tree Figure 19 is a schematic diagram illustrating the representation of an example ingredient (Rose Absolute in Color DM) as a path. For ingredient representation, each ingredient is assigned a family and a note. These families and notes are represented as paths in the family and note classification tree, respectively. A standard (or anchor) ingredient is one whose values ​​are all the same along the note path. In this example, the ingredient is very close to standard rose.

[0171] Projecting mixes onto classification trees Figure 20 demonstrates the projection of an example blend onto a classification tree; the intensity of each component is projected onto each layer of the tree, allowing the blend to be represented at different granularities.

[0172] Calculating blend similarity at each layer of classification Figure 21 illustrates the calculation of blend similarity for each classification layer; in each layer, the cosine distance corresponds to the angle between vectors whose coordinates are the intensities on each possible node. The smaller the angle, the higher the similarity. When the distance is short, the olfactory sense of two blends will be considered close on the layer.

[0173] Collapsing layers into global blend similarities Figure 22 illustrates the calculation of the weighted sum of all layers. By calculating the weighted sum of the distances at each granularity layer, the aggregated pairwise similarity between formulations is calculated. The weights are hyperparameters of the model. In the example ML model for palette and concentration generation above, only layers "family 1" and "feature factors" are employed; the weights of all other layers can be set to 0 to ignore their contributions.

[0174] Clustering blending using K-nn label propagator Figure 23 illustrates the clustering of blends. A K-Nearest Neighbor (K-nn) label propagator algorithm is applied to divide the pool of blends into clusters (note that, as an example, SNB010FSN is the identifier of one generated blend). The number of current neighbors, K, is set to 10; this hyperparameter can be adjusted, if desired, to obtain different cluster distributions. In a practical example, the following numbers are observed: - Retained formula: 165,556 - Number of clusters: 7,083 - Minimum cluster size: 2 - Maximum cluster size: 326 - Average cluster size: 23.37 - Median cluster size: 17

[0175] Representing clusters for users Figure 24 is a pair of diagrams demonstrating the olfactory characteristics of example blends as determined using the above technique. The characteristics of the clusters are represented using the median intensity of the blends within the cluster. The list is truncated to a specific length: for example, 5 for family and 10 for note (characteristic factor). Of course, this can be easily adapted.

[0176] Cluster Naming & Querying The names of the clusters and the filtering criteria applicable to each cluster are extracted as a subset of the representations defined by both the family and the feature factors. All olfactory labels with an intensity higher than a predefined percentage (i.e., threshold) of the maximum intensity will be considered "named" labels. As can be seen in Figure 24, considering a family threshold of 0.5 and a feature factor threshold of 0.38, the example formula can be described with the family "fruit" and the notes (feature factors) "strawberry," "cooked sugar," "butyric acid," and "grass."

[0177] In essence, the described technique for the olfactory description of a formulation involves mapping the formulation's palette and component concentrations to predetermined olfactory values. This allows for quantitative clustering of formulations with respect to their olfactory contribution. In turn, this simplifies the user's formulation assessment, providing knowledge of the expected olfactory characteristics of a large number of formulations without the user having to create the formulation.

[0178] Graphical User Interface After generation of the predicted formulation, optionally a graphical user interface may be provided to the user to allow for quick and efficient observation and analysis of the resulting formulation, followed by the olfactory signature of the resulting formulation.

[0179] FIG. 25 is an example of a user interface suitable for providing to a user to enable filtering of the resulting formulations; in this "Filter" pane, the user can select a subset of formulations that fit their preferences. In the illustrated case, the user filters to show formulations that fall within the "fruity" and "citrus" olfactory families. As shown, additional filters can be applied, including exclusion filters ("Exclude Family," "Exclude Characteristic Factors," "Exclude Ingredients"), as well as inclusion filters ("Include Characteristic Factors," "Include Ingredients"), and filters in terms of the number of ingredients present in the formulation (here, any number between 2 and 90). As shown, drop-down menus can facilitate displaying lists of data (in this case, characteristic factors). The skilled reader will understand that additional filters (e.g., filtering by expected formulation cost) can be implemented.

[0180] FIG. 26 is a further example user interface illustrating the results of filtering; a subset of formulations that match the filter criteria as entered by the user is shown (in this case, there is only one formulation that satisfies the filter criteria). A profile for formulation "KM29EZQ1LM" is displayed, including the formulation's olfactory profile factors and olfactory family. Hovering over a data point provides the user with further information. For example, as shown, hovering over the profile factor "balsam" results in a pop-up indicating that the calculated "balsam" intensity (calculated according to the formulation's olfactory description above) is 1.78.

[0181] FIG. 27 is a further example user interface illustrating further results of filtering. Scrolling provides the user with further information for the formula "KM29EZQ1LM." As shown, this example formula contains five ingredients ("Clonal," "Indole Pure," "Manzanate," "Peach Pure," and "Terpinolene") at the indicated concentrations. In the lower panel, further information is shown, including the expected price of creating the formula; price information can be assessed on a per ingredient basis from a separate table. The model used to generate the formula is also indicated (here, CVAE for palette generation and GAN for concentration generation). Finally, the expected end use of the formula is indicated (here, "Q(Washing)").

[0182] 28 is another example user interface illustrating the results of a separate filtering process. As shown, the filtering parameters implemented here result in a generated combination of 209 pages.

[0183] Also, as shown, the user interface provides functionality that allows a user to easily create a desired formulation. In this case, selecting the "Create" button may generate instructions for creating the formulation "02QJY4N6NT." These instructions may be in the form of instructions for the user to follow (e.g., the amount of each ingredient to mix). Alternatively, the instructions may be in the form of instructions specifically tailored for delivery to a fragrance making machine.

[0184] Hardware 29 is a block diagram of a computing device, such as a personal computer or data storage server, that embodies the present invention and that may be used to implement embodiment methods for generating fragrance ingredient combinations and / or for training a system for generating fragrance ingredient combinations. The computing device comprises a processor 993 and memory 994. Optionally, the computing device also includes a network interface 997 for communicating with other computing devices, for example, a fragrance creation robot or machine.

[0185] For example, an embodiment may consist of a network of such computing devices. Optionally, the computing devices also include one or more input mechanisms, such as a keyboard and mouse 996, and one or more display units, such as a monitor 995. The components may be interconnected via a bus 992.

[0186] The memory 994 may include a computer-readable medium, which term may refer to a single medium or multiple media (e.g., centralized or distributed databases and / or associated caches and servers) configured to carry computer-executable instructions or have data structures stored thereon. Computer-executable instructions may include, for example, instructions and data that can be accessed by a general-purpose computer, a special-purpose computer, or a special-purpose processing device (e.g., one or more processors) and cause it to perform one or more functions or operations. Thus, the term "computer-readable storage medium" may include any medium that can store, encode, or carry a set of instructions for execution by a machine, causing the machine to perform one or more of the methods disclosed herein. Thus, the term "computer-readable storage medium" may be taken to include, but is not limited to, solid-state memory, optical media, and magnetic media. By way of example and not limitation, such computer-readable media may include non-transitory computer-readable storage media including random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory devices (e.g., solid-state memory devices).

[0187] The processor 993 is configured to control the computing device and perform processing operations, for example, executing code stored in memory to implement the various different functions of the training module and inference module described herein and in the claims. The memory 994 stores data that is read and written by the processor 993. The processor referred to herein may include one or more general-purpose processing devices, such as a microprocessor, a central processing unit, or the like. The processor may include a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor implementing other instruction sets or a combination of instruction sets. The processor may also include one or more special-purpose processing devices, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), a network processor, or the like. In one or more aspects, the processor is configured to execute instructions to perform the operations and steps discussed herein.

[0188] The display unit 995 may display representations of data stored by the computing device and may also display cursors and dialog boxes and screens that enable interaction between a user and programs and data stored on the computing device. The input mechanism 996 may allow a user to input data and instructions into the computing device. Illustratively, the display unit 995 and input mechanism 996 may allow a user to interact with a user interface suitable for filtering generated fragrances.

[0189] The network interface (network I / F) 997 may be connected to a network such as the Internet and can be connected to other such computing devices via the network. The network I / F 997 can control the input / output of data from / to other devices via the network. Other peripheral devices such as a microphone, speaker, printer, etc. can be included in the computing device.

[0190] The computing device may include training processing instructions stored in a portion of memory 994, a processor 993 configured to execute the processing instructions, and a portion of memory 994 configured to store ML model training data during execution of the processing instructions. Illustratively, the processor 993 may be configured to access the training data (a BOM of fragrances) stored in memory 994 and implement the forward and backward propagation techniques necessary to arrive at a trained ML model. The trained model (including, by way of example, node topology details and trained weights) may be stored in memory 994 and / or a connected storage unit and delivered / transferred / communicated to a user for execution of the trained ML model.

[0191] The computing device may include production processing instructions stored in a portion of memory 994, a processor 993 configured to execute the processing instructions, and a portion of memory 994 configured to store trained ML model data during execution of the processing instructions. Illustratively, the processor 993 may be configured to access the trained ML model stored in memory 994, execute the trained model, and arrive at a product palette, product ingredient concentrations, and / or product formulation. The product data may be stored in memory 994 and / or a connected storage unit and may be delivered / transferred / communicated to a user for display, for example, or delivered / transferred / communicated to a means for creating a product formulation, for example.

[0192] Methods embodying the present invention may be performed on a computing device such as that illustrated in Figure 29. Such a computing device need not have every component illustrated in Figure 29, but may consist of a subset of those components. Methods embodying the present invention may be performed by a single computing device in communication with one or more data storage servers over a network. The computing device may be the data storage itself that stores the trained fragrance generation ML model or the generated formula.

[0193] Methods embodying the present invention may be performed by a plurality of computing devices operating in conjunction with one another, one or more of which may be data storage servers that store at least a portion of the trained fragrance generation ML model or generated formulation.

Claims

1. 1. A computer-implemented method for training a system for generating blends of fragrance and / or flavor ingredients, the method comprising: receiving an input data set comprising an ingredient palette and ingredient concentrations of known fragrances and / or flavors (S10); training (S12) a palette-generating machine learning model using the input dataset, where the palette-generating machine learning model, once trained, is configured to generate at least one product component palette; and Training a concentration-generating machine learning model using the input dataset (S14), where the concentration-generating machine learning model, once trained, is configured to generate at least one generating recipe, each generating recipe including component concentrations of the generating component palette.

2. The method of claim 1 , wherein the input data set comprises known fragrance end uses.

3. 3. The method of claim 2 further comprising: receiving input of identified end uses from among known fragrance and / or flavor end uses; and training a re-trained palette-generating machine learning model by fine-tuning the trained palette-generating machine learning model using ingredient palettes and ingredient concentrations of known fragrances and / or flavors associated with the identified end-use; and Training of concentration-generating machine learning models using a palette of ingredients for identified end uses.

4. The method of any one of claims 1 to 3, wherein the palette-generating machine learning model is a variational autoencoder (VAE) or a generative adversarial network (GAN).

5. The method according to any one of claims 1 to 4, wherein the cardinality generating machine learning model is a bucket predictor (BP) or a second GAN.

6. The palette generation machine learning model, once trained, outputs a score indicating the presence of each component in each generated component palette; and For each ingredient palette, the palette generation machine learning model is configured to determine that each ingredient is included in the ingredient palette if the ingredient's score or presence exceeds a predetermined threshold, the predetermined threshold being based on the ingredient's presence in known fragrances and / or flavors of the input dataset; The method according to any one of claims 1 to 5.

7. the input data set includes known fragrance and / or flavor end uses; the palette-generating machine learning model is a VAE that includes an encoder-decoder architecture; training the palette-generating machine learning model includes passing known fragrance and / or flavor end uses to an encoder and a decoder of the palette-generating machine learning model; The method according to any one of claims 1 to 6.

8. The palette generation machine learning model is a VAE; and training the palette-generating machine learning model includes, for each training mini-batch of the input dataset, and for each layer, normalizing the output of each layer; The method according to any one of claims 1 to 7.

9. 9. The method of claim 1, further comprising running the trained palette-generating machine learning model and the trained concentration-generating machine learning model to generate at least one resulting formulation.

10. 10. The method of claim 9, further comprising calculating an olfactory descriptor for at least one resultant formulation.

11. 1. A computer-implemented method for generating a blend of ingredients for a fragrance, the method comprising: Executing the trained palette generation machine learning model to generate at least one generated ingredient palette (S20), the trained palette generation machine learning model using an input dataset including ingredient palettes of known fragrances and / or flavors and ingredient concentrations of the known fragrances and / or flavors; and Executing the trained concentration-generating machine learning model to generate at least one resulting formulation (S22), the trained concentration-generating machine learning model using an input dataset, where each resulting formulation includes component concentrations from the resulting component palette.

12. The method of claim 11 , further comprising generating instructions for creating at least one formula according to the ingredients and ingredient concentrations of the corresponding ingredient palette.

13. 13. The method of claim 11 or 12, further comprising: providing a user interface for displaying at least a subset of the at least one product recipe; Accepting user input to the user interface to select at least one subset of the product recipes.

14. 14. The method of claim 13 further comprising: Accepting user input into a user interface to create a recipe.

15. A data processing apparatus comprising a memory and a processor, the memory and the processor being configured to perform a method according to any one of claims 1 to 14.

16. A computer program comprising instructions which, when executed by a computer, cause the computer to carry out the method according to any one of claims 1 to 14.

17. A computer readable medium comprising instructions which, when executed by a computer, cause the computer to carry out the method of any one of claims 1 to 14.