Method for generating a set of manufacturing data making it possible to manufacture a cosmetic composition and associated devices
The use of machine learning models generates manufacturing data for cosmetic compositions, addressing environmental and regulatory challenges by completing data sets with eco-friendly ingredients, optimizing production and performance.
Patent Information
- Application Number
- PCT/EP2025/070199
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-15
- Filing Date
- 2025-07-15
- Publication Date
- 2026-01-22
AI Technical Summary
Existing methods for formulating cosmetic compositions fail to address environmental concerns, regulatory compliance, and optimization of manufacturing processes while ensuring availability and performance of ingredients.
A method utilizing machine learning models, specifically masked predictive transformers, to generate manufacturing data for cosmetic compositions that comply with environmental criteria and performance constraints by completing incomplete data sets with eco-friendly ingredients.
Enables the formulation of sustainable cosmetic compositions that meet regulatory requirements and optimize industrial production, while ensuring ingredient availability and performance.
Smart Images

Figure IMGF000005_0001 
Figure IMGF000005_0002 
Figure IMGF000007_0001
Abstract
Description
[0001] TITLE: Method for generating a set of manufacturing data making it possible to manufacture a cosmetic composition and associated devices
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application claims priority to FR Patent Application No. 24 07748, filed 15 July 2024, which is hereby incorporated by reference.
[0004] TECHNICAL FIELD OF THE INVENTION
[0005] The present invention relates to a method for generating a set of manufacturing data enabling the formulation, preparation and / or manufacture of a cosmetic composition. It also relates to a method for preparing such a composition. It also relates to the devices involved in the above methods.
[0006] TECHNOLOGICAL BACKGROUND OF THE INVENTION
[0007] The formulation of environmentally friendly cosmetic products, i.e. the design and development of which take account of environmental concerns, is becoming a major concern to help meet global challenges.
[0008] It has therefore become essential to propose more sustainable compositions and / or preparation methods and / or raw materials to meet these environmental challenges.
[0009] In particular, in this context, it becomes important to replace certain materials with alternatives with a better environmental footprint, in particular by reducing the use of raw materials from petrochemicals.
[0010] One objective is to propose eco-responsible materials, in particular from green chemistry, in order to limit the environmental impact of the composition (and in addition to its primary and / or secondary packaging) of its manufacture at the end of its life, in particular thanks to materials with a good biodegradability profile and / or from renewable sources.
[0011] The formulation of new cosmetic compositions may also be the result of new regulations.
[0012] For example, as a result of European directives from 2015, producing lipstick or nail polish containing cobalt nitrate was prohibited.
[0013] Reformulation may also be necessary because a raw material may be acceptable in one legislation but not in another.
[0014] More broadly, it may also be desirable to modify a cosmetic composition for the purpose of optimizing its manufacturing cost, optimizing its industrial production, or even depending on the availability of the ingredients making it up.. It may also be desirable to modify it to improve and develop its properties with consumers.
[0015] It may also be useful when it comes to being able to ensure the availability of ingredients or raw materials of certain cosmetic compositions.
[0016] SUMMARY OF THE INVENTION
[0017] There is therefore a need for a method that is easy to implement and that makes it possible to generatea set of manufacturing data of a cosmetic composition responding to a specific context, in particular a reduced environmental footprint.
[0018] For this purpose, the description relates to a method for generating a set of manufacturing data, the set of manufacturing data including manufacturing data of a cosmetic composition able to comply with at least one acceptability criterion, said method being implemented by a computing device and comprising an inference phase, the inference phase comprising a step of:
[0019] - obtaining a set to be completed, the set comprising manufacturing data of a cosmetic composition, and optionally at least one constraint parameter to be met by the chemical composition, and
[0020] - applying a technique to the set to be completed and optionally the at least one constraint parameter to obtain a completed set, the completed set comprising manufacturing data of a cosmetic composition able to comply with said at least one acceptability criterion, the technique including using at least one machine learning model, the at least one machine learning model being capable of completing a set to be completed.
[0021] According to further advantageous aspects of the invention, the generation method comprises one or several of the following features, taken in isolation or in any technically possible combination:
[0022] - the learning model was trained to complete a set to be completed with learning on at least one database containing cosmetic compositions complying with said at least one acceptability criterion and possibly cosmetic compositions not complying with said at least one acceptability criterion.
[0023] - the method further includes a phase of training the model on a training database during which a merit function is used, the merit function favouring models capable of generating completed sets not forming part of the training database and complying with the at least one acceptance criterion.
[0024] - the manufacturing data are a list of raw materials or ingredients, possibly accompanied by their respective proportions in the cosmetic composition. - during the obtaining step, at least one constraint parameter is obtained, a constraint parameter being a completeness parameter defining a number of missing manufacturing data in the set to be completed.
[0025] - during the application phase, the model is used only once, the completeness parameter corresponding to a number of missing data being equal to 1 .
[0026] - a constraint parameter is a global completeness parameter corresponding to a number of missing data greater than 1 , the application phase including several iterations, each iteration comprising the application of the model to the set generated at the previous iteration with a completeness parameter for said iteration corresponding to a number of missing data equal to 1 , the model being applied to the first iteration on the set to be completed.
[0027] - the model includes an encoder and / or decoder.
[0028] - The model is a masked predictive transformer model.
[0029] - the cosmetic composition forms part of a cosmetic product, the manufacturing data being selected from a list consisting of data relating to the cosmetic composition, data relating to a packaging of the cosmetic product and data relating to the manufacturing technique of the cosmetic product used.
[0030] - the method further includes a phase of updating the model by retraining the model using the sets used during the interference phase.
[0031] The description also relates to a method for generating a method for generating at least one set of manufacturing data, the set of manufacturing data including manufacturing data of a cosmetic composition, said method being implemented by a computing device and comprising a generation phase, the generation phase comprising a step of:
[0032] - obtaining a set of manufacturing data and one or more manufacturing data to be modified in the obtained set,
[0033] - calculating a similarity coefficient between each manufacturing data item to be modified and a plurality of candidate manufacturing data,
[0034] - selecting, for each manufacturing data item to be modified, at least one manufacturing data item from the plurality of candidate manufacturing data according to the calculated similarity coefficients, to obtain at least one selected data item,
[0035] - generating at least one set of manufacturing data, to obtain at least one generated set, each generated set being the set obtained in which each manufacturing data item to be modified is replaced by a respective selected manufacturing data item. According to further advantageous aspects of the invention, the generation method comprises one or several of the following features, taken in isolation or in any technically possible combination:
[0036] - each manufacturing data item is associated with several values, each value being a quantification of a respective property of the manufacturing data item, the similarity coefficient of a manufacturing data item dependent on at least one value of the manufacturing data item.
[0037] - a property is a measurable property.
[0038] - a property is a property that can be assessed qualitatively.
[0039] - during the calculation step, the computing device (12) calculates at least one similarity coefficient by applying the following formula:
[0040] Where:
[0041] • C(A, B) denotes the similarity coefficient between the manufacturing data A and the manufacturing data B,
[0042] •xmod,i denotes the ithmanufacturing data item to be modified, i being a non-zero integer,
[0043] •xcand,j denotes the jthcandidate manufacturing data item, j being a non-zero integer,
[0044] • w designates the sum of the weights applied to the p properties
[0045] • wpdesignates the weight applied to the pthproperty (wp= 1 by default)
[0046] • n is the number of properties, o ymodip:value of the pthproperty of xmod i, where p is an integer between
[0047] 1 and n, o yCand,jp : value of the pthproperty ofxcand > o Rp: this value is at least equal to the difference between the 5th and 95th percentile of the distribution of values of the pthproperty.
[0048] - for a specific manufacturing data item to be modified, the selected candidate manufacturing data item is the deletion of the specific manufacturing data item, each set generated being then the set obtained in which the specific manufacturing data item is deleted and each other manufacturing data item possibly to be modified is replaced by a respective selected manufacturing data item.
[0049] - the cosmetic composition forms part of a cosmetic product, the manufacturing data being selected from a list consisting of data relating to the cosmetic composition, data relating to a packaging of the cosmetic product and data relating to the manufacturing technique of the cosmetic product used.
[0050] - the manufacturing data are selected from a list of raw materials, a list of molecular structures and a list of characteristics.
[0051] The description also relates to a method for manufacturing a cosmetic composition, the method being implemented by a manufacturing system, the manufacturing method comprising:
[0052] - a phase of implementing a method for generating at least one set of manufacturing data, the generation method being as previously described, and
[0053] - a phase of manufacturing a cosmetic composition from each set generated during the implementation phase.
[0054] The description also describes a computing device configured to generate at least one set of manufacturing data, the manufacturing data set including manufacturing data of a cosmetic composition and:
[0055] - obtaining a set of manufacturing data and one or more manufacturing data to be modified in the obtained set,
[0056] - calculating a similarity coefficient between each manufacturing data to be modified and a plurality of candidate manufacturing data,
[0057] - selecting, for each manufacturing data to be modified, at least one manufacturing data item from the plurality of candidate manufacturing data according to the calculated similarity coefficients, to obtain at least one selected data item,
[0058] - generating at least one set of manufacturing data, to obtain at least one generated set, each generated set being the set obtained in which each manufacturing data to be modified.
[0059] The description also relates to a computing device configured to generate a set of manufacturing data, the set of manufacturing data comprising manufacturing data of a cosmetic composition complying with at least one acceptability criterion, said calculation device also being configured to :
[0060] - obtain a set to be completed, the set comprising manufacturing data for a cosmetic composition, and optionally at least one constraint parameter to be met by the chemical composition, and - applying a technique to the set to be completed and optionally the at least one constraint parameter to obtain a completed set, the completed set comprising data for manufacturing a cosmetic composition capable of complying with said at least one acceptability criterion, the technique comprising the use of at least one automatic learning model, the at least one automatic learning model being suitable for completing a set to be completed.
[0061] The description also relates to a system for manufacturing a cosmetic composition, the manufacturing system comprising:
[0062] - a computing device as previously described, and
[0063] - a manufacturing device capable of manufacturing a cosmetic composition from the manufacturing data set obtained by the computing device.
[0064] The description also describes a method for predicting a set of characteristics associated with a list of raw materials, said method being implemented by a prediction device and comprising an inference phase, the inference phase comprising a step of:
[0065] - obtaining a list of properties associated with the list of raw materials, and
[0066] - applying a technique to the list of properties to obtain the characteristics of the raw material and / or the list of raw materials, the technique including the use of at least one Bayesian network, said Bayesian network being determined to link characteristics of a raw material to properties.
[0067] According to further advantageous aspects of the invention, the method comprises one or more of the following features, taken in isolation or in any technically possible combination:
[0068] - the method further includes applying a function f linking the value of a characteristic to a distribution of a node and to the list of properties according to the following equation:
[0069] Where: o Xi : value of characteristic i, ->Pi,n)' probability of node i knowing the n parent nodes of i.
[0070] - arcs and distributions of the Bayesian network are determined by an expert group or using known data.
[0071] - the list of characteristics comprises all the characteristics whose associated probability distribution is greater than an associated prediction threshold.
[0072] - the method further includes an update phase during which the user can update the Bayesian network based on expert opinion and known data. - the method includes the application of an unsupervised machine learning algorithm to expert opinions and known data to determine the Bayesian network.
[0073] - a characteristic is a measurable property.
[0074] - a characteristic is a property that can be qualitatively assessed.
[0075] The description also describes a method for selecting one or more lists of candidate raw materials to replace a first list of raw materials and comprising the following steps:
[0076] - determining a list of properties associated with the first list of raw materials,
[0077] - calculating characteristics of the cosmetic composition by implementing a prediction method as previously described,
[0078] - applying a machine learning model to determine a plurality of lists of candidate raw materials whose characteristics are close to the calculated characteristics, and
[0079] - selecting one or more lists of raw materials from the plurality of lists of candidate raw materials according to a selection criterion associated with a threshold assigned to each of the characteristics.
[0080] The description also describes a device for predicting a set of characteristics associated with a list of raw materials, said prediction device being configured to
[0081] - obtain a list of properties associated with the list of raw materials, and
[0082] - apply a technique to the list of properties to obtain the characteristics of the raw material and / or the list of raw materials, the technique including the use of at least one Bayesian network, said Bayesian network being determined to link characteristics of a raw material to properties.
[0083] BRIEF DESCRIPTION OF THE FIGURES
[0084] The invention will become more apparent upon reading the following description, given solely as a non-limiting example, and made with reference to the drawings, wherein: o [Fig. 1] figure 1 is a schematic representation of a manufacturing system comprising a computing device and a manufacturing device, o [Fig. 2] figure 2 is a flow diagram of an example implementation of a method for generating a set of manufacturing data, o [Fig. 3] figure 3 is a block representation of an example of a model used by the generation method of figure 2, o [Fig. 4] figure 4 is a block representation of another example of a model used by the generation method of figure 2, o [Fig. 5] figure 5 is a block representation of yet another example model used by the method for generating figure 2, o [Fig. 6] figure 6 is a flow diagram of an example implementation of another method for generating a set of manufacturing data, o [Fig. 7] figure 7 is a schematic representation of an example of latent space making it possible to better understand the method of figure 6, o [Fig. 8] figure 8 is a flow diagram of an example implementation of a method for manufacturing a cosmetic composition, and o [Fig. 9] figure 9 is a flow diagram of an example implementation of a prediction method.
[0085] DETAILED DESCRIPTION OF PREFERRED EMBODIMENTS
[0086] Description of the manufacturin system
[0087] Figure 1 schematically shows a system 10 for manufacturing a cosmetic product.
[0088] The manufacturing system 10 is suitable for manufacturing the cosmetic product.
[0089] The concept of “cosmetic product” is defined in the section relating to the manufacturing data of a cosmetic product.
[0090] The manufacturing system 10 includes a computing device 12 and a manufacturing device 14.
[0091] The computing device 12 is capable of implementing a method for generating a set of manufacturing data making it possible to manufacture a cosmetic product, or at least a part thereof, i.e. at least one cosmetic composition.
[0092] Several examples of generation methods will be described later. These processes are computer-implemented processes.
[0093] Further information on the manufacturing data included in the set can be found in the corresponding section of the description.
[0094] The computing device 12 is a computer comprising one or more electronic components such as: one or more processors with one or more cores represented collectively by a processor, one or more graphical processing units (GPU), a hard disk, a random access memory (RAM), a display interface or an input / output interface. It will be appreciated that the computer may for example be implemented in the form of a desktop computer, a computer in a vehicle, a tablet or a smartphone.
[0095] In the example of Figure 1 , the computing device 12 comprises a processing unit 16 comprising and a memory 18 associated with the processing unit 16. The processing unit 16 is an electronic circuit design to manipulate and / or transform data represented by electronic or physical quantities in registers of the computer and / or memories into other similar data corresponding to physical data in the memories of registers or of other types of display device, transmission device or storage device.
[0096] As specific examples, the processing unit 16 is made in the form of a programmable logic component, such as an FPGA (standing for Field Programmable Gate Array in English), or an integrated circuit, such as an ASIC (standing for Application Specific Integrated Circuit in English).
[0097] In a variant, when the method is implemented in the form is made in the form of one or more software packages, i.e. in the form of a computer program, also so-called a computer program product, it could also be recorded on a computer-readable medium, not shown. For example, the computer-readable medium is a medium able to record electronic instructions and to be coupled to a bus of a computer system. For example, the readable medium is an optical disk, a magneto-optical disk, a ROM memory, a RAM memory, any type of non-volatile memory (for example FLASH or NVRAM) or a magnetic card. A computer program comprising software instructions is then recorded on the readable medium.
[0098] The manufacturing device 14 is capable of manufacturing the cosmetic product from the manufacturing data set generated by the processing unit 16.
[0099] Description of a generation method based on a set of data to be completed
[0100] The operation of the computing device 12 is now described with reference to Figure 2 which is a flow diagram illustrating an example of implementation of a generation method.
[0101] The generation method is a method aimed at generating a set of manufacturing data of a cosmetic composition capable of complying with at least one acceptability criterion from a set of manufacturing data to be completed.
[0102] The incompleteness of the manufacturing data set to be completed results in a completeness parameter forming a contextual input parameter which, as will be seen below, may be implicit or explicit.
[0103] The completeness parameter reflecting the incompleteness of the manufacturing data set may in particular be expressed in the form of a number of missing raw materials and / or a total number of raw materials to be achieved in the completed composition. It is understood that it may alternatively or additionally be expressed in the form of a constraint parameter, such as a price, performance, environmental compatibility constraint, giving rise to the search for additional raw materials to complete the set. This constraint parameter may be different from the acceptance criterion or correspond to the acceptance criterion.
[0104] In this description, the term “manufacturing data” refers to data making it possible to formulate and / or prepare the cosmetic product and could also be translated as “preparation data”.
[0105] In the remainder of the description, the set that is generated will be called “the completed set” (satisfying the completeness parameter) and the set that is input will be called “set to be completed” (not satisfying the completeness parameter).
[0106] More specifically, in the example to be described, it is assumed that the set to be completed is a list of raw materials constituting a cosmetic composition that, a priori, does not meet the acceptability criterion. The set to be completed is a list that is missing, implicitly or explicitly, one or more raw materials. In particular, the set to be completed is a list whose number of constituent raw materials is less than the number of desired raw materials after application of the completeness parameter.
[0107] For example, when the completeness parameter is the total number of expected raw materials, then the total number of raw materials of the composition to be completed is less than the completeness parameter. When the completeness parameter is the number of raw materials missing or to be added, then said completeness parameter added to the number of raw materials of the composition to be completed corresponds to the number of raw materials of the completed composition sought and to be generated by the model described below. The completeness parameter can also be a range of values giving the model the freedom to add a number of raw materials between a lower and an upper limit. It is also possible to provide only a lower limit value or an upper limit value. If necessary, a default value can also be provided.
[0108] In addition, it is assumed that the completed set is a completed list of raw materials, able to meet at least one acceptance criterion, it being understood that this example can be transposed to any type of manufacturing data as described in the corresponding section, namely the section entitled “Description of the various possible manufacturing data sets”.
[0109] For example, the set to be completed may be an empty set (without raw material).
[0110] The purpose of the method is therefore to generate the data for manufacturing a cosmetic composition comprising, in a physiologically acceptable medium, a plurality of raw materials.
[0111] A raw material is a set of one or more ingredients used in the manufacture of cosmetic products. These raw materials can be of natural or synthetic origin and play various roles in cosmetic formulations, ranging from active agents to texture agents. For example, a raw material may be glycerin, elastin or hyaluronic acid. In the remainder of this description, a raw material is defined by the set of ingredients that make it up, a set of molecular structures of the ingredients that make it up, a set of concentrations of the ingredients that make it up, a set of properties of the ingredients that make it up, or a set of characteristics.
[0112] "Physiologically acceptable" means a medium compatible with keratin materials. "Keratin materials" refers to the skin, mucosa and / or skin appendages. Preferably, the keratin materials are the skin, particularly facial skin, mucosa such as lips, and / or skin appendages such as eyelashes.
[0113] By way of non-limiting example, such a cosmetic composition is a hair treatment or care product (e.g. a shampoo), a skin treatment or care product (e.g. a moisturiser, a sun cream, an anti-aging cream), a make-up product (e.g. lipstick, mascara, foundation, or a gloss.
[0114] In the example described, the generation method comprises a learning phase P20, an inference phase P30 and an update phase P40.
[0115] This is only a simple example, it being understood that in general the P30 inference and P40 update phases are implemented successively to always improve the model that is used.
[0116] It can also be noted that the implementation of the P40 update phase is not mandatory.
[0117] The learning phase P20 may be performed offline, i.e. by a computer different from the computing device 12, and preferably prior to the use of the computing device 12.
[0118] During the learning phase P20, a machine learning model M is trained via a machine learning algorithm A to complete a set to be completed according to a completeness parameter defining, directly or indirectly, a number of manufacturing data missing in the set to be completed.
[0119] A machine learning algorithm A (commonly referred to as the machine learning algorithm) is an algorithm that teaches / trains a model M to find patterns and / or general input-output relationships automatically from training (or learning) data. After learning, the model M could make similar inferences and / or predictions. In the mathematical sense, the model M could be seen as an equation that maps inputs to output. Machine learning algorithms can be categorised according to the types of task to be performed (classification, regression, generation), the learning approach used (supervised, unsupervised, semisupervised, self-supervised and reinforced) and algorithmic theory (neural network, tree model, probabilistic model, polynomial model, Bayesian model, SVM, etc.). The learning algorithm applied in this method relates to self-supervised and reinforced learning for a generation task. It is based on the neural network theory, or more precisely thats of transformer for training and inference / generation.
[0120] The model M at the end of the learning phase is a transformer neural network.
[0121] The neural network includes an ordered succession of neural layers, each of which takes its inputs from the outputs of the preceding layer.
[0122] More specifically, each layer comprises neurons taking their inputs from the outputs of the neurons of the preceding layer, or from the input variables for the first layer.
[0123] Alternatively, more complex neural network structures may be considered with a layer which can be connected to a more distant layer than the immediately previous layer.
[0124] An operation, i.e. a type of processing, to be performed by said neuron in the corresponding processing layer is also associated with each neuron.
[0125] Each layer is connected to the other layers by a plurality of synapses. A synaptic weight is associated with each synapse, and each synapse forms a link between two neurons. It is often a real number, which takes both positive and negative values. In some cases, the synaptic weight is a complex number.
[0126] Each neuron is capable of performing a weighted sum of the value(s) received from the neurons of the preceding layer, each value then being multiplied by the respective synaptic weight of each synapse, or connection, between said neuron and the neurons of the preceding layer, then applying an activation function, typically a non-linear function, to said weighted sum, and outputting from said neuron, in particular to the neurons of the following layer which are connected thereto, the value resulting from the application of the activation function. The activation function allows introducing a non-linearity into the processing performed by each neuron. The sigmoid function, the hyperbolic tangent function, the Heaviside function are examples of activation functions.
[0127] As on optional supplement, each neuron is also capable of also applying a multiplying factor, also so-called bias, to the output of the activation function, and the value output from said neuron is then the product of the bias value and the value obtained from the activation function.
[0128] A convolutional neural network is also sometimes referred to as a convolutional neural network or by the acronym CNN, which refers to the English term “Convolutional Neural Networks”.
[0129] In a convolutional neural network, each neuron in the same layer has exactly the same connection pattern as its neighboring neurons, but at different entry positions. The connection pattern is called a convolution kernel or, more often, “kernel” in reference to the corresponding English name. A fully connected neural layer is a layer in which the neurons of said layer are each connected to all the neurons of the previous layer.
[0130] Such a type of layer is more often referred to by the English term “fully connected”, and sometimes referred to as “dense layer”.
[0131] In a variant, the model M applies linear regression, logistic regression, a decision tree, a random forest, a primary component analysis (PCA), a Naive Bayes algorithm or a k-nearest neighbours algorithm (KNN)
[0132] In a variant, the model M is a support vector machine or a k-means clustering model.
[0133] “Machine learning” here means that the model M is taught to perform a task using a learning algorithm A launched on a computer and applied to training data. The learning of the model M on a task is done through learning algorithms A which are an optimisation process against a well-defined metric. By way of example, the optimisation may be minimising the least squares error between the predicted and actually measured values or maximising the reward in the case of reinforced learning, on the training data.
[0134] The task to be performed here is to generate the completed set of raw materials from a set to be completed or from a subset of the completed set.
[0135] The model M can thus be seen as a generative or predictive model.
[0136] It is also assumed that a training database has been established for this purpose. This learning database contains examples of data that the model M can read repeatedly to understand the implicit patterns or relationships present in the data.
[0137] The generation of such a database varies depending on the type of T ransformer and the learning approach adopted.
[0138] For a Transformer that only contains the encoder and learns through a learning approach called “MLM” (Masking Language Modeling), each learning example consists of a list to be completed with some raw materials masked and another completed list with all raw materials exposed. Each data pair of the base can be generated starting from a real cosmetic composition whose list of raw materials is known and masking one or more raw materials from the list. An actual cosmetic composition is implicitly considered chemically valid and minimally supplemented. When training an encoder type transformer, the list to be completed will be submitted to the input of the model M. The input data information will be propagated in the various layers of the model for processing up to the output layer. The output layer will predict input masked raw materials. The predicted raw materials will be compared with the corresponding raw materials in the completed list by calculating an error. The error will be used to readjust the parameters in the model M via the backpropagation mechanism to reduce this error. This mechanism is performed repetitively in order to gradually reduce the overall error on the prediction of masked raw materials. For a Transformer containing only a decoder, each learning example consists of a completed list of raw materials in a cosmetic composition with a 'START' token added at the beginning of the list and the same list completed with an 'END' token added at the end. Each database data pair can be generated starting from a real cosmetic composition whose list of raw materials is known and adding the 'START' token at the beginning of the list or the 'END' token at the end of the list respectively. An actual cosmetic composition is considered chemically valid and minimally supplemented. When training such a model containing only a decoder, the list containing the 'START' token will be injected into a self- service layer by the 'VALUE', 'KEY' and 'QUERY' entries. The self-service layers encode the information in the list and propagate it to the output layer to predict the list containing the 'END' token. The weights in the self-service layers that are recalculated for each token in the list containing the 'START' token will be masked by a mask called 'CAUSAL'. This 'CAUSAL' masking could mimic a self-regressive generation mechanism. The prediction error is calculated by comparing the raw material sequence generated in such a self- regressive way and the subset list containing the 'END' token. The error will be used to readjust the parameters in the model M via the backpropagation mechanism to reduce this error. This mechanism is performed repetitively in order to gradually reduce the overall error on the prediction of masked raw materials.
[0139] For a Transformer containing an encoder and a decoder, each learning example consists of a list to be completed with a subset of raw materials removed from a completed list of cosmetic composition, and two lists containing a complementary subset of raw materials in the same completed list of cosmetic composition. One of the additional lists is added an additional token 'START' at the beginning of the list and the other is added an additional token 'END' at the END of the list. Each learning data triplet can be generated by splitting a completed list of actual cosmetic composition and adding the token 'START' at the beginning of the additional list or the 'END' at the END of the additional list respectively. An actual cosmetic composition is implicitly considered chemically valid and minimally supplemented. When training a Transformer containing an encoder and a decoder, the list to be completed of a subset of raw materials is injected into an encoder which encodes it into a vector that represents all the essential information for the prediction sequence. This encoded vector will be injected into the decoder by the 'KEY' and 'VALUE' inputs of the second self-service layer. The additional list containing the 'START' token will be injected into the same self-service layer by the 'QUERY' entry, after having been encoded by another previous self-service layer. The information encoded by the second self-service layer will be propagated in the rest of the layers up to the output layer to predict the complementary list containing the 'END' token. The weights in the self-service layers that are recalculated for each token in the list containing the 'START' token will be masked by a mask called 'CAUSAL'. This 'CAUSAL' masking could mimic a self-regressive generation mechanism. The prediction error is calculated by comparing the raw material sequence generated in such a self-regressive way and the subset list containing the 'END' token. The error will be used to readjust the parameters in the model M via the backpropagation mechanism to reduce this error. This mechanism is performed repetitively in order to gradually reduce the overall error on the prediction of masked raw materials.
[0140] Through this simple example, it is clear that it is possible to generate a large database since a single composition can make it possible to generate a plurality of pairs for the training.
[0141] Any training technique can then be used to obtain the model.
[0142] For example, learning is supervised or unsupervised learning or self-supervised learning that has been applied in the process.
[0143] Other techniques are conceivable as reinforcement learning techniques used in the process to encourage the model to propose manufacturing data sets making it possible to manufacture a cosmetic product not part of the learning database and complying with at least one acceptance criterion.
[0144] According to one embodiment, a merit function (also called loss function) is used, the merit function favoring models capable of generating sets of manufacturing data deemed “completed” making it possible to manufacture a cosmetic product not part of the database and complying with at least one acceptance criterion.
[0145] More specifically, learning is accomplished by weight estimation or scores associated with a reward function.
[0146] In this example, the generation of known data is neutral with respect to the reward function.
[0147] On the other hand, generating data that is part of a list of data to be excluded results in a negative weight or score, while generating unknown data results in a positive weight or score.
[0148] By making an average, the model M determines an associated acceptability criterion enabling a future generation to be validated or not.
[0149] For other types of models Msuch as those based on the k-nearest neighbours (KNN) algorithm, training includes storing training data, calculating Euclidean distances between test points and all training points, selecting close neighbours, and predicting by classification or regression (see in particular figure 6).
[0150] During the inference phase P30, the computing device 12 applies the learned model to input data comprising a composition to be completed and a completeness parameter. The completeness parameter can be explicit, for example by directly providing the expected total number of raw materials or the number of missing raw materials to be added. The completeness parameter may also be defined by default or implied and determined from the data of the composition to be completed using a marker or masking token in the data of said composition to be completed. According to a particular embodiment, the completeness parameter is deduced from an additional constraint parameter, such as a price constraint, for example.
[0151] According to the example described, the inference phase P30 including an obtaining step E301 and an application step E302.
[0152] During the obtaining step E301 , the computing device 12 obtains a set to be completed of manufacturing data and the completeness parameter.
[0153] For example, the computing device 12 reads this list and the completeness parameter in the memory 18.
[0154] In addition, in some embodiments, the user may enter a completed list and select at least one raw material to be excluded, with which a MASK token will subsequently be associated. In such a case, the number of MASK tokens determines the completeness parameter (exact, minimum or maximum number of raw materials to be added, for example).
[0155] During the application step E302, the computing device 12 applies a technique to the set to be completed to obtain a completed set. Possibly associated with a probability that said completed set meets the model’s acceptance criterion. Completed sets whose probability of meeting the criterion is greater than a certain threshold (>0.8 for example) will be presented to the user.
[0156] The technique here consists in applying the model M to the set to be completed to obtain a completed set according to the completeness parameter and the acceptance criterion.
[0157] Several topologies are possible for the model M.
[0158] According to a first example, the model M is applied only once.
[0159] In this example, the model M includes at least part of a transformer, more specifically the encoder.
[0160] As a particular example, the model M uses a BERT transformer model (which refers to the corresponding English name of "Bidirectional Encoder Representations from Transformers" meaning literally “bidirectional representations of encoders from transformers”).
[0161] The MPT model is a neural network comprising several layers shown in Figure 3. In general, the model M has several layers, among which two sets of layers play an interesting role.
[0162] The first set is a multi-head attention set. This first set allows the model to estimate the connection between the input elements and to extract relationships between near and far elements with the same efficiency.
[0163] The second set is a direct acting point-to-point neural network.
[0164] Such a neural network is more often referred to by the corresponding English term "Position-Wise Feed-Forward Network".
[0165] The second set serves here to help change the representation and capture more information about the context. A particular example of such an example of a model is now described with reference to figure 3.
[0166] In the example shown, the model M includes 4 encoders. Each encoder includes an “input embedding” layer, a “multi-head attention” layer which comprises two “addition / normalization” sub-layers, a “feed forward” layer and a “classification” layer.
[0167] The “input embedding” or “input vectors” layer is a layer in charge of converting the tokens presented into integer identifiers in a continuous vector format.
[0168] The “multi-head attention” layer continues to encode various tokens previously generated with weightings (also called attention weights). These weights are calculated for each token and represent implicit relational information between the tokens with respect to the final prediction.
[0169] The “addition / normalization” layers consist in concatenating each original token (before passing through the “multi-head attention” layer) at the output of each multi-head attention sub-layer, in order to enrich the number of informative features used for the final prediction. A normalization is then performed to stabilize the learning process.
[0170] The “feed forward” or “direct acting” layer consists of several sublayers of neurons or each of the sublayers is connected to each other. The neurons of each of said sublayers receive weighted inputs from the neurons of the previous layer and transmit their outputs to the following layers. At each activated layer, a scalar product is applied between the weights associated with each neuron and their inputs before application of an activation function.
[0171] All of the training of such a neural network consists in converging the values of said weights to match the theoretical outputs of the training data with the actual outputs.
[0172] The outputs of the “feed forward” layer are “contextualized embedding” or “contextualized vectors”.
[0173] The “classification layer” is used to transform the normalized contextualized vectors concatenated with the tokens entering the feedforward layer into predictions specific to a task. Typically, such predictions are text classification, named entity recognition, or answering questions. In our example, these are raw materials for manufacturing a cosmetic composition.
[0174] According to a second example, the model M is applied several times with an input reinjection of all or part of the completed sets obtained at the output of the model M. The reinjection is possibly accompanied by an update of the completeness parameter (addition of a raw material masked in the data reinjected for an implicit determination of the updated completeness parameter or explicit update of the completeness parameter).
[0175] This means that, at a first calculation iteration, the model M is applied to the set to be completed and outputs a first intermediate set and then, at a second iteration, the model is applied to the first intermediate set to output a second intermediate set and so on until a final iteration during which the application of the model M makes it possible to obtain the completed set.
[0176] According to one embodiment, the number of iterations of the application corresponds to the number of manufacturing data missing in the set to be completed.
[0177] In such a case, at each iteration, the model M adds a manufacturing data item to the set given at the input of the model M at each iteration. According to one embodiment, the overall completeness parameter is greater than 1 missing raw material, for example 3 missing raw materials, each iteration using an intermediate completeness parameter corresponding to 1 missing raw material. Thus, after 3 iterations, the completed composition is obtained.
[0178] In such a case, the model M is a masked predictive transformer model.
[0179] Such a model is often referred to by the abbreviation MPT, which refers to the corresponding English name “Masked Predictive Transformer”.
[0180] The model M is a neural network comprising several layers shown in Figure 4.
[0181] In the example shown, the model M includes an “input embedding” layer, two “multihead attention” layers, three “addition / normalization” layers, a “feed forward” layer, a “linear” layer, a “softmax” layer and an “arg max” layer.
[0182] The linear layer is a layer used in neural networks that applies a linear transformation (or matrix multiplication with a transformation matrix) to input data using weight and bias. The application of this linear layer makes it possible to transform the data received at its input into a dimensionality either lower for dimensional reduction or higher for more complex tasks.
[0183] The softmax layer is a layer used in neural networks that applies a softmax function (or a normalized exponential function) to input data in real numbers, in order to convert them into a format that could represent the distribution of probabilities over a number of choices. It is useful for assigning a probability to each possible response in prediction. The argmax layer always follows the softmax layer. The application of this layer is simply to output the element with the highest value in the probability vector generated by the softmax layer. This element will be taken as the final response for the prediction.
[0184] According to a third example, the model M is applied several times with a reinjection at the input of the second sub-model of the output of the model M.
[0185] This means that, at a first calculation iteration, the model M is applied to the set to be completed and outputs at least one first intermediate set then, at a second iteration, the model is applied to one of the first intermediate sets to output at least one second intermediate set and so on until a final iteration during which the application of the model M makes it possible to obtain the completed set.
[0186] According to one embodiment, the number of iterations of the application corresponds to the number of manufacturing data missing in the set to be completed.
[0187] In such a case, at each iteration, the model M adds a manufacturing data item to the set given at the input of the model M at each iteration.
[0188] In such a case, the model M is a masked predictive transformer model.
[0189] Such a model is often referred to by the abbreviation MPT, which refers to the corresponding English name “Masked Predictive Transformer”.
[0190] The model M is a neural network comprising several layers shown in Figure 5.
[0191] The model M includes a first sub-model and a second sub-model.
[0192] The first submodel includes an “input embedding” layer, a “multi-head attention” layer, two “addition / normalization” layers and a “feed forward” layer.
[0193] The output of the first sub-model is then applied to the input of a second “multi-head attention” layer of the second sub-model including an “input embedding” layer, a first “multihead attention” layer, three “addition / normalization” layers, the second “multi-head attention” layer, a “feed forward” layer, a “linear” layer, a “softmax” layer and an “argmax” layer.
[0194] During the update phase P40, the completed set obtained is, if necessary after checking whether it meets the acceptance criterion or not, integrated with training data and the model M is retrained from the training data obtained.
[0195] For example, the weights associated with each sublayer of the “feedforward” layer are determined by the computing device 12 in order to match the outputs of the model to the training data.
[0196] The generation process thus makes it possible to obtain a complete set.
[0197] Through these examples, it is clear that the machine learning model can also take into account one or more additional input parameters making it possible to determine the acceptability criterion(s). These input parameters are constraint parameters to be complied with by the cosmetic composition (and therefore indirectly by the completed set).
[0198] As an example of additional input data, the number of missing raw materials may be quoted in which case the constraint parameter is a completeness parameter or a quantification of a performance that the cosmetic product should achieve.
[0199] In the case where the constraint parameter is a completeness parameter, it may be envisaged implementing the preceding embodiments as follows.
[0200] During the application phase, the model is used only once, the completeness parameter corresponding to a number of missing data equal to 1 .
[0201] In the iterative embodiment, a constraint parameter is a global completeness parameter corresponding to a number of missing data greater than 1 and the application phase includes several iterations, each iteration comprising applying the model to the set generated at the previous iteration with a completeness parameter for said iteration corresponding to a number of missing data equal to 1 , the model being applied to the first iteration on the set to be completed.
[0202] In the case where the constraint parameter is a quantification of a performance that the cosmetic product should achieve, this quantification makes it possible to define an acceptability criterion.
[0203] More generally, the acceptability criterion is derived from the available information, either only from the constraint parameter(s), in particular when the set to be completed is empty (without raw material), or from the set to be completed, in particular in the absence of constraint parameters or from both, i.e. the raw materials of the set to be completed and the parameters.
[0204] In an extreme case, the acceptability criterion can thus be reduced to a biologically acceptable formula for the part of the body intended to accommodate it.
[0205] If applicable, the term “raw material” may comprise a packaging or member for applying the cosmetic composition, the packaging and the member being then intended to form, together, a cosmetic product.
[0206] Thus, based on the same principle described above, it is possible to provide, to a model, the data of a cosmetic composition considered as incomplete, no longer with respect to a missing chemical raw material, but with respect to a missing application member or packaging to constitute a final cosmetic product, this final cosmetic product constituting the completed cosmetic composition within the meaning of the embodiments described above.
[0207] Of course, it is possible to combine these acceptance criteria to generate a complete set as satisfactory as possible, the model being trained accordingly in accordance with the acceptance criteria to be targeted. Description of a generation method based on a data set to be modified
[0208] The operation of the computing device 12 is now described with reference to Figure 6, which illustrates another example of implementation of another generation method.
[0209] The generation method according to Figure 6 includes a determination phase P50 and a generation phase P60.
[0210] During the determination phase P50, a list of properties and / or molecular structures associated with each of the raw materials is determined.
[0211] During the generation phase P60, the computing device 10 obtains one or more sets to be completed.
[0212] The generation phase P60 includes an obtaining step E601 , a calculation step E602, a selection step E603 and a generation step E604.
[0213] During the obtaining step E601 , the computing device 12 obtains a list of raw materials and one or more raw materials to be modified in the set obtained.
[0214] For example, the user enters the list of raw materials to be modified via the input interface that the computing device 12 reads in the memory 18.
[0215] During the calculation step E602, the computing device 12 calculates a similarity coefficient between each raw material to be modified and a plurality of candidate raw materials.
[0216] According to a particular example, each raw material is associated with several property values.
[0217] The term property is to be understood here in the broad sense as comprising physically measurable properties but also qualitative properties that can be assessed on an appropriate rating scale.
[0218] More specifically, the term property refers to physical, chemical, functional or biological properties.
[0219] For example, a physical property defines the structure of the raw material such as its state, solubility or viscosity.
[0220] For example, a chemical property chemically defines the raw material through its pH, its associated carbon chain length or its oxidation.
[0221] For example, a functional property defines different practical abilities of the raw material such as its hydrating, anti-aging or emollient ability.
[0222] For example, a biological property defines a biological capacity of the raw material as its antibacterial, healing or soothing capacity. Each value is a quantification of a respective property of the raw material, so that the provision of all the values of the properties of a raw material can be seen as the provision of a signature of the raw material.
[0223] The calculated similarity coefficient then depends on at least one value of a property of the raw material.
[0224] Several examples of such a calculation may be given.
[0225] In an embodiment with properties whose value is quantifiable, the calculator 12 applies the following equation:
[0226] Where:
[0227] • C(A, B') means the coefficient of similarity between material A and material B,
[0228] •xmod,i denotes the ithraw material to be modified, i being a non-zero integer,
[0229] •xcand denotes the jthcandidate raw material, j being a non-zero integer,
[0230] • w is the sum of the weights applied to the p properties
[0231] • wpdesignates the weight applied to the pthproperty (wp= 1 by default)
[0232] • n is the number of properties, o Tmodip:value of the pthproperty of xmodii, where p is an integer between 1 and n,
[0233] O ycand p : value of the p,hproperty ofxcand,j ’ o Rp: this value is at least equal to the difference between the 5th and 95th percentile of the distribution of values of the pthproperty.
[0234] • If the value of the property is qualitatively measurable:
[0235] O Sp x-mod,i>xcand,j) 0 If Xmodlj>xcand.j > Otherwise
[0236] According to one example, the number of properties n is equal to the number of different properties of the raw materials of the set given during the obtaining step E601 .
[0237] According to one variant, the calculation step E602 is implemented differently.
[0238] In this variant, the calculation step E602 includes a search step and a calculation step.
[0239] During the search step, the computing device 12 searches for a plurality of reference lists.
[0240] Typically, a reference list is composed of sets that have undergone, or have not undergone, modifications in their associated raw material lists. The reference lists are maintained in a database via a distributed and scalable data structure (SDDS, from the English Shared Domain Data Set) enabling a large number of data to be stored by defining associated characteristics such as the source of said data.
[0241] Typically, such a structure makes it possible to store more than one million reference lists in the database.
[0242] The reference lists relate to cosmetic compositions that may have been modified in the past. Thus each list is associated with a plurality of lists of raw materials.
[0243] In the example described, the reference list includes N iterations and therefore N associated separate raw material lists.
[0244] In addition, the reference list includes the raw material to be modified but all the lists of separate raw materials generated after the Jthiteration no longer contain it (J being smaller than N).
[0245] For example, the search step is performed from an SQL query applied to the database.
[0246] Thus the computing device 12 identifies the N-J lists of associated distinct materials not including the raw material to be modified.
[0247] During the calculation step E602, the computing device 12 calculates a similarity coefficient between the N-J lists of associated separate raw materials and the J lists of raw materials containing the raw material to be modified.
[0248] For example, the computing device 12 applies a function f such as:
[0249] Where: o C A, B) means the similarity coefficient between the set to be completed and the Bthlist of associated separate materials (j < B < N), o B : number of iterations, o X: number of raw materials not shared between the two lists.
[0250] Alternatively, the calculation step E602 is implemented differently.
[0251] In this variant, the calculation step E602 includes a conversion sub-step, a first calculation sub-step, a projection sub-step and a second calculation sub-step.
[0252] In the conversion substep, the molecular structure of the raw material is converted into molecular identifiers.
[0253] For example, such identifiers take the form of SMILES, InchiKey or Molecular fingerprints structures.
[0254] During the first calculation substep, the computing device 12 calculates a plurality of molecular properties from the molecular identifiers. For example, the molecular properties include the molar mass, a partition coefficient, or a number of chemical bonds.
[0255] During the projection sub-step, the computing device 12 identifies the molecular properties related to each other in order to represent the latter in a projection space with reduced dimensions (see Figure 7).
[0256] For example, the computing device 12 minimizes the dimensions of the projection space by identifying the dimensions related to molecular properties being linear compositions of other properties.
[0257] Typically, the computing device 12 can apply a principal component analysis PCA of such a step.
[0258] The principal component analysis is a statistical analysis method that makes it possible to summarize the information contained in a set of functional data, thus making it possible to produce so-called principal component functions of minimum dimension that reproduce a maximum of information on the data studied.
[0259] In other words, functional principal components analysis is a projection of input information into an optimal projection base.
[0260] Finally, during the second calculation step, the computing device 12 calculates the similarity coefficient by applying a clustering algorithm in the projection space of dimension M and calculating the similarity coefficient according to a distance criterion. For example, the similarity coefficient is zero for two raw materials that are not in the same cluster.
[0261] Advantageously, such a feature makes it possible to replace the at least one raw material to be excluded with a new raw material providing one or more priority properties for the operator.
[0262] Alternatively, a method for predicting characteristics of the raw material according to said properties may be used. This device is described in the final section of the present application.
[0263] During the selection step E603, the computing device 12 selects, for each raw material to be modified, at least one raw material from the plurality of candidate raw materials.
[0264] Each raw material thus selected is hereinafter referred to as the selected raw material.
[0265] The selection is made according to the similarity coefficients calculated during the calculation step E602.
[0266] According to a simple example, the selected raw materials are the raw materials of the plurality of candidate raw materials having the highest similarity coefficients. In other examples, the raw materials are selected based on similarity coefficients and a set of criteria associated with each of the candidate raw materials.
[0267] Typically, the set of criteria includes an economic or environmentally friendly indicator such as biodegradability.
[0268] During the generation step E604, the computing device 12 generates all the sets of raw materials in which each raw material to be modified is replaced by a respectively selected raw material.
[0269] The generation method thus makes it possible to obtain one or more completed sets of raw materials.
[0270] In particular, this makes it possible to exclude a raw material by replacing it with one or more others or even to delete it to obtain a new cosmetic composition having the sought properties.
[0271] Such a method then makes it possible to generate manufacturing data useful for the preparation and manufacture of innovative cosmetic compositions adapted to specific constraints.
[0272] Advantageously, the implementation of such a method is easy, in particular relatively fast.
[0273] In addition, this method also allows access to formulation, improvement or replacement levers for anyone manufacturing cosmetic products, from beginners to experienced or even experts.
[0274] Applications of previous generation methods
[0275] The set obtained by either of the previous generation methods is useful for many applications.
[0276] According to the example described, it may be used in a method for preparing and manufacturing a cosmetic product by the manufacturing system 14.
[0277] The manufacturing method comprises an implementation phase P70 and a manufacturing phase P80.
[0278] During the implementation phase P70, the computing device 12 implements either of the previous generation methods.
[0279] For example, the computing device 12 implements the inference phase P40 of the generation method according to the embodiment of Figure 2 or the generation phase P60 of the generation method according to the embodiment of Figure 6.
[0280] The manufacturing phase P80 aims to manufacture the cosmetic product corresponding to the set obtained after implementation during the implementation phase P70. For example, the manufacturing device 14 takes as input the set generated and controls the elements that comprise it so that these elements carry out the manufacture of the cosmetic product corresponding to the set generated by the generation method.
[0281] Other manufacturing methods can be envisaged.
[0282] For example, during a configuration step, the operator can select a packaging and an embodiment to be respected in the form of data and include them in the manufacturing data.
[0283] The manufacturing technique designates a set of criteria to be respected during the manufacture of the cosmetic composition such as an order of incorporation, temperature, mixing elements (speed, time, turbines and / or blades), phase separation or pre-mixing. In particular, the manufacturing technique may be expressed in the form of unitary process operations.
[0284] Description of the various possible man ufactu ring data sets
[0285] The generation methods of figures 2 and 6 have so far been described for a particular example of a list of raw materials for illustrative purposes.
[0286] However, in a general case, manufacturing data can be of any type.
[0287] Depending on the cases, the manufacturing data are selected from a list consisting of data relating to the composition, data relating to the packaging, in particular to a member for applying the cosmetic composition, and data relating to the manufacturing technique used.
[0288] The manufacturing data make it possible, depending on the case, to obtain a cosmetic composition or a cosmetic product.
[0289] A cosmetic composition is composed of a list of raw materials assigning a list of properties to said composition.
[0290] “Cosmetic product” is understood to mean the entirety of a cosmetic composition, a package or packaging, comprising in particular an application member.
[0291] The manufacturing method designates a set of unitary operations to be respected during the manufacture of the cosmetic composition such as an order of incorporation, temperature, humidity or applied force.
[0292] Packaging designates the means for preserving the cosmetic composition. Alternatively, the packaging includes a member or element for applying the associated cosmetic composition such as a lash brush for a mascara or a brush for a foundation
[0293] Data relating to the composition is data representing the chemical composition of the formula. The list of raw materials or ingredients is a particular example of compositional data.
[0294] Another example of composition-related data may be the concentration of the raw material in said composition.
[0295] The representation of the raw material may differ according to embodiments. By way of example, the raw material may be represented by a chemical formula, an identifier, a trade name or a type (emulsifier or fat) and its proportion in the composition may be expressed as a mass or volume percentage, detailed linearly or with preparations
[0296] Instead of a raw material, it is possible to consider the ingredients forming part of the raw materials, the ingredients being similarly represented in various ways (formula, identifier or others).
[0297] The raw material or ingredient content is another example of compositional data.
[0298] Data relating to the manufacturing technique used is data used to characterize the steps of a manufacturing process.
[0299] An example of data relating to the manufacturing technique is a succession of unitary method operations.
[0300] An example of data relating to the manufacturing technique is the order in which the ingredients or raw material are inserted during manufacture, the temperature of placing in the tank, the introduction and mixing times, the speeds and mechanical forces applied"
[0301] Description of a method for predicting characteristics
[0302] As previously explained, previous generation methods use knowledge about raw material lists. It is therefore useful to be able to estimate or predict the characteristics of a raw material and / or a list of raw materials
[0303] For this purpose, a prediction method can be implemented via a prediction device.
[0304] The same remarks as for the computing device apply here for the prediction device and are therefore not repeated.
[0305] An example of implementation of the method for predicting characteristics of a list of raw materials is now described
[0306] The prediction method is a method for estimating a set of characteristics associated with a raw material.
[0307] Alternatively, the method estimates a set of characteristics associated with a plurality of raw materials forming a cosmetic composition. Thus, in the remainder of the description, a raw material also denotes a set of raw materials associated with a list of properties previously determined from the lists of properties of the set of raw materials. The term characteristic here is to be understood in the broad sense as comprising measurable characteristics but also qualitative characteristics that can be assessed appropriately, for example on a rating scale according to defined evaluation criteria.
[0308] Characteristics are information associated with the raw material such as assured functions, technical characteristics, physico-chemical properties or contributions to sensory benefits.
[0309] The characteristics of a raw material are defined by all the technical effects induced by its properties.
[0310] For example, a characteristic of a raw material is its stickiness, its acid resistance or its ability to stabilize the oil phase.
[0311] Such a list of characteristics also includes a value associated with each of the characteristics which may be quantitative or qualitative.
[0312] The prediction method according to Figure 9 includes a determination phase P90 and a prediction phase P100.
[0313] During the determination phase P90, a list of properties associated with the raw material is determined.
[0314] The prediction phase P100 includesan obtaining step E1001 , a calculation step E1002 and a prediction step E1003.
[0315] During the obtaining step E1001 , the computing device 12 obtains a list of properties.
[0316] For example, the user manually enters the list of properties of the raw material via the input interface that the computing device 12 reads in the memory 18. The list of raw material properties can also be retrieved automatically by the system, in particular from a suitable storage database.
[0317] In another example, the computing device 12 receives the list of properties of the raw material by an electronic device via radio communication or wired connection.
[0318] In addition, the user enters values associated with each of said properties. As with the list of properties, the associated values can also be retrieved automatically from a database.
[0319] During the calculation step E1002, the computing device 12 calculates the list of characteristics using the list of properties and a causal Bayesian network.
[0320] A Bayesian network is a model for determining an output value or a distribution of output values according to a plurality of input data influencing (or not) said output value.
[0321] A Bayesian network is a graphical probabilistic model that represents a set of random variables and their conditional dependencies via a directed acyclic graph. It is used to model uncertain knowledge in various fields such as the relationships between characteristics and properties of a cosmetic composition.
[0322] To create such a Bayesian network, it is necessary to determine, before use, a set of variables representing the nodes of the graph, a set of arcs linking the nodes representing conditional dependencies and probabilities for each of the nodes.
[0323] Each node corresponds to a characteristic whose value then depends on the network inputs (values of the properties entered by the user). Each node is associated with a parent influencing its probability distribution.
[0324] The entries (properties) are also represented by nodes without parents.
[0325] The arcs and probabilities associated with each of the nodes enable this dependency. This is because arcs and probabilities are determined by expert appraisal of a group of experts and represent the dependencies between characteristics and properties, elicited probabilities are spoken of.
[0326] Alternatively, arcs and probabilities are determined from known data. Typically, known data is data linking a property to a historically known or recently discovered characteristic.
[0327] The values assigned to each of the probabilities can be changed over time, for example if new knowledge makes it possible to refine the probabilities determined by the expert group. Thus an arc linking a property to a feature means that the property influences said feature.
[0328] An intermediate node is present on each of the arcs making it possible to represent the weight (influence) of the property on said characteristic.
[0329] In addition, intermediate nodes can make it possible to represent a cumulative effect of several properties on a feature linking several weights together.
[0330] For example, the list of characteristics includes all characteristics whose associated probability distribution is greater than an associated prediction threshold.
[0331] Typically, a Bayesian network is determined for each large family of raw materials having common characteristics.
[0332] For example, a large family of raw materials may be water-based thickeners or emollients.
[0333] In the prediction step E1003, the prediction device predicts the list of characteristics and their associated values.
[0334] For example, the prediction device applies a function f linking the distribution to a value of the feature and properties such as:
[0335] Where: o Xi : value of characteristic i, o -, Pi,n)- probability distribution of node i knowing the n parent nodes of i,
[0336] In an embodiment, during an optional update phase, it is possible to update the Bayesian network via the prediction device.
[0337] The term “update” means adding or modifying associated nodes, probabilities, and arcs.
[0338] For example, the user enters new nodes, new arcs, and new associated probabilities via the prediction device and from expert group expertise.
[0339] Alternatively, the Bayesian network, new nodes, new arcs, and new probabilities are determined by applying one or more unsupervised learning algorithms to the known data.
[0340] The unsupervised learning algorithm(s) are part of the group: clustering algorithms, dimensionality reduction algorithm, density modeling algorithm, or neural network algorithm.
[0341] Such unsupervised learning algorithms are presented in the article “Learning Bayesian Networks with the bnlearn R Package” published on July 10, 2010.
[0342] Advantageously, this “hybrid” approach between the expert opinion and the result of one or more learning algorithms makes it possible to increase the accuracy and longevity over time of the results of the causal Bayesian model.
[0343] Advantageously, it is possible to obtain the list of properties of a raw material from these characteristics by going up the graph of the Bayesian network.
[0344] The term “going up” means “determining” and is performed by applying an optimization algorithm or the Bayes theorem.
[0345] Typically, the optimization algorithm is a genetic algorithm.
[0346] The prediction method may be used in a method for selecting one or more lists of raw materials to replace a first list of raw materials used by the computing device.
[0347] Such a selection method includes a determination step, a calculation step, and a selection step.
[0348] During the determination step, a property list associated with the first list of raw materials is determined.
[0349] For example, the user manually enters the list of properties via the input interface that the computing device 12 reads in the memory 18.
[0350] During the calculation step, the computing device 12 calculates the characteristics of the cosmetic composition using a prediction method as presented above. During the selection step, the user selects a list of raw materials from the plurality of candidate lists of raw materials via the computing device 12.
[0351] Alternatively, the computing device 12 selects one or more lists respecting a set of selection criteria. For example, a selection criterion is a threshold assigned to each of the characteristics.
[0352] Advantageously, such a method makes it possible to select one or more lists of raw materials having characteristics close to the input list.
Claims
CLAIMS1. Method for generating a set of manufacturing data, the set of manufacturing data including manufacturing data of a cosmetic composition able to comply with at least one acceptability criterion, said method being implemented by a computing device (12) and comprising an inference phase, the inference phase comprising a step of:- obtaining a set to be completed, the set comprising manufacturing data of a cosmetic composition, and optionally at least one constraint parameter to be met by the chemical composition, and- applying a technique to the set to be completed and optionally the at least one constraint parameter to obtain a completed set, the completed set comprising manufacturing data of a cosmetic composition able to comply with said at least one acceptability criterion, the technique including using at least one machine learning model, the at least one machine learning model being capable of completing a set to be completed.
2. Generation method according to claim 1 , wherein the learning model has been trained to complete a set to be completed with learning on at least one database containing cosmetic compositions complying with said at least one acceptability criterion and optionally cosmetic compositions not complying with said at least one acceptability criterion.
3. Generation method according to claim 1 or 2, wherein the method further includes a phase of training the model on a training database during which a merit function is used, the merit function favoring models capable of generating completed sets not forming part of the training database and complying with the at least one acceptance criterion.
4. Generation method according to any one of claims 1 to 3, wherein the manufacturing data are a list of raw materials or ingredients, optionally accompanied by their respective proportions in the cosmetic composition.
5. Generation method according to any one of claims 1 to 4, wherein, during the obtaining step, at least one constraint parameter is obtained, a constraint parameter beinga completeness parameter defining a number of missing manufacturing data in the set to be completed.
6. Generation method according to claim 5, wherein, during the application phase, the model is used only once, the completeness parameter corresponding to a number of missing data being equal to 1 .
7. Generation method according to claim 5, wherein a constraint parameter is a global completeness parameter corresponding to a number of missing data greater than 1 , the application phase including several iterations, each iteration comprising applying the model to the set generated at the previous iteration with a completeness parameter for said iteration corresponding to a number of missing data equal to 1 , the model being applied to the first iteration on the set to be completed.
8. Generation method according to any one of claims 1 to 6, wherein the model includes an encoder and / or a decoder.
9. Generation method according to any one of claims 1 to 7, wherein the model is a masked predictive transformer model.
10. Generation method according to any one of claims 1 to 8, wherein the cosmetic composition forms part of a cosmetic product, the manufacturing data being selected from a list consisting of data relating to the cosmetic composition, data relating to a packaging of the cosmetic product and data relating to the technique for manufacturing the cosmetic product used.11 . Generation method according to any one of claims 1 to 9, wherein the method includes, among other things, a phase of updating the model by retraining the model using the sets used during the interference phase.
12. Method for manufacturing a cosmetic composition, the method being implemented by a manufacturing system (10), the manufacturing method comprising:- a phase of implementing a method for generating a set of manufacturing data, the generation method being in accordance with any one of claims 1 to 11 , and- a phase of manufacturing a cosmetic composition from the completed set obtained by the implementation phase.
13. Computing device (12) configured to generate a set of manufacturing data, the set of manufacturing data including manufacturing data of a cosmetic composition complying with at least one acceptability criterion, said computing device (12) also being configured to:- obtain a set to be completed, the set comprising manufacturing data of a cosmetic composition, and optionally at least one constraint parameter to be complied with by the chemical composition, and- apply a technique to the set to be completed and optionally the at least one constraint parameter to obtain a completed set, the completed set comprising manufacturing data of a cosmetic composition able to comply with said at least one acceptability criterion, the technique including using at least one machine learning model, the at least one machine learning model being capable of completing a set to be completed.
14. System (10) for manufacturing a cosmetic composition, the manufacturing system (10) comprising:- a computing device (12) according to claim 13, and- a manufacturing device (14) capable of manufacturing a cosmetic composition from the manufacturing data set obtained by the computing device (12).
Citation Information
Patent Citations
Qualitative or quantitative characterization of a coating surface
US20220082508A1
Method, system, and computer program product for creating a recipe for a chemical composition
US20220101453A1