HAIR COLOR SIMULATION USING A GUIDED HAIR COLOR CLASSIFICATION NETWORK

A guided hair simulation model using a hair classifier and semi-supervised learning framework addresses human bias in expert-labeled data, improving hair color classification accuracy and photorealism in virtual hair dye trials.

FR3160260B3Active Publication Date: 2026-04-17LOREAL SA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
FR · FR
Patent Type
Utility models
Current Assignee / Owner
LOREAL SA
Filing Date
2024-03-15
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing hair color classification models face challenges due to human bias in expert-labeled data, which are costly and limited in quantity, leading to inaccurate and synthetic-looking hair simulations.

Method used

A generative hair simulation model is trained using a guided hair classifier model, incorporating annotator confusion matrices and a Mean Teacher semi-supervised learning framework to mitigate label noise and leverage unlabeled data, achieving improved accuracy with only 1,000 labeled data points.

Benefits of technology

The model achieves a 20% higher classification accuracy compared to professional human annotators, enhancing photorealism and color accuracy in hair simulations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000035_0000
    Figure 00000035_0000
  • Figure 00000035_0001
    Figure 00000035_0001
  • Figure 00000036_0000
    Figure 00000036_0000
Patent Text Reader

Abstract

HAIR COLOR SIMULATION USING A GUIDED HAIR COLOR CLASSIFICATION NETWORK Aspects of hair simulation and related networks are proposed, including aspects concerning the training of such networks. A generative hair simulation model is provided and guided during its training by a hair classifier model. The generative model, in one embodiment, is proposed for use in a virtual testing pipeline (VTO), such as for virtually testing hair coloring products. A color mapping network is also provided to process an input image and a target hair color for the generative model in order to define the hair simulation (e.g., as an output image with simulated hair color). Figure for abstract: none
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: SIMULATION OF HAIR COLOR USING A GUIDED HAIR COLOR CLASSIFICATION NETWORK FIELD OF INVENTION

[0001] This disclosure relates to computer image processing and artificial intelligence, including systems and methods for simulating and classifying hair color, and more specifically, a method, device, and system for simulating hair color using a guided hair color classification network. BACKGROUND

[0002] Due to the increasing popularity of online shopping, e-commerce companies are exploring ways to improve the customer experience. Virtual Try-On Technology (VTO) has been developed to help customers search for and virtually explore products, for example, to find the one that best suits them. In some examples of VTO, it can be useful to determine information, particularly from a user-supplied image. For example, in a hair-color-related VTO, a face image (e.g., a portrait) can be processed to determine, by classification, the hair color in the portrait. Figure 1 shows a schematic diagram of a representative processing flow 100 that illustrates a general model in the form of a classifier 102 that receives a color face image 104 (e.g., a portrait) and identifies a hair color 106, in accordance with the prior art.The output, along with the color image of the face, can then be used for a virtual hair dye trial.

[0003] Several challenges arise when developing a hair color classification model. For example, color labels collected from hair experts are subject to human bias, and labeled data are costly to obtain. It is desirable to have an improved hair classification model that can identify hair color.

[0004] A VTO pipeline can simulate hair color, for example, using a generative neural network. Improvements in simulation using generative networks are desired. SUMMARY

[0005] Aspects of hair simulation and corresponding networks are proposed, including aspects concerning the learning of such networks. A model A generative hair simulation model is provided and guided during its training by a hair classifier model. In one embodiment, the generative model is provided for use in a virtual testing pipeline (VTO), for example, to virtually test hair coloring products. A color mapping network is also provided to process an input image and a target hair color for the generative model to define the hair simulation (e.g., as an output image with simulated hair color).

[0006] Aspects of VTO pipelines and corresponding networks are provided, including aspects concerning the training of such networks. Furthermore, according to one embodiment, a classifier model is provided to mitigate the impact of human bias, where the modeling of the distribution of real labels and annotator bias are separated by incorporating annotator confusion matrices into a reference model. To further improve the model's performance by leveraging unlabeled data, the model was trained using a consistency-based semi-supervised learning framework. Using only 1,000 labeled data points, the final classifier model achieved a classification accuracy 20% higher than that of a professional human annotator.The trained model can be used for a wide range of downstream tasks, including as a color classifier to train generative models for hair color translation.

[0007] The following statements present various aspects and features disclosed in the embodiments herein. These aspects and features, as well as others, will be readily understood by those skilled in the art, particularly the aspects of computer program products. It is also understood that aspects / features of computer devices or systems may have corresponding process aspects / features and vice versa.

[0008] Embodiment 1: A computer device comprising a processor coupled to a storage device that stores instructions executable by the processor to cause the computer device to: process an input image and a target hair color using a virtual testing pipeline (VTO) to produce a VTO experiment that simulates the target hair color with the input image to produce an output image; wherein the VTO pipeline comprises a color-refining neural network to generate the output image combining the target hair color and the input image, the color-refining neural network comprising a generative neural network trained under the supervision of a hair classification network comprising a hue classifier that outputs a hair hue value and a reflectance classifier that outputs a value hair reflectance to determine training loss information from output images to train the generative neural network.

[0009] Embodiment 2: Computer device according to embodiment 1, wherein the instructions are executable by the processor to cause the computer device to provide an interface with one or both of: i) a color recommendation engine to recommend a target hair color° and ii) an e-commerce service with which to purchase a product or service or both.

[0010] Embodiment 3: Computer device according to embodiment 1, wherein the VTO pipeline includes a color mapping network trained to produce a color map for input into the generative neural network in order to produce the output image for the VTO experiment, the color map produced from the image features, the image hair color data determined from the image and the target hair color.

[0011] Embodiment 4: Computer device according to embodiment 3, in which the color mapping network includes an encoder to determine features from the input image; a concatenator to combine features, target hair color and image hair color data for processing by a first fully connected block, and a second block to process an intermediate map from the first block with the image hair color data to produce the color map for input into the generative neural network.

[0012] Embodiment 5: Computer device according to embodiment 4, in which the VTO pipeline includes a hair analysis engine to determine image hair color data from the input image.

[0013] Embodiment 6: Computer device according to embodiment 1, wherein the input image and target hair color are defined using RGB values, such that the generative neural network is trained to produce an output image using particular RGB values ​​as guided during training by the hair classification network using hair hue values ​​and reflectance values.

[0014] Embodiment 7: Computer device according to embodiment 6, in which the hair classification network is defined and trained according to an industry standard for hair color classification comprising a respective plurality of classes for hair tint and hair reflectance.

[0015] Embodiment 8: Computer device according to embodiment 1, in which the hair classification network comprises a coding skeleton and the hair shade classifier and reflectance classifier each comprise a respective linear classifier.

[0016] Embodiment 9: Computer device according to embodiment 8, wherein the hair classification network has been defined by training with a cross-entropy loss sum of hue and reflectance determined from i) outputs of hair hue values ​​and hair reflectance values, and ii) target labels for hair hue and hair reflectance, the target labels being prepared for the training images of the hair classification network from respective expert votes by a plurality of experts.

[0017] Embodiment 10: Computer device according to embodiment 9, wherein the target labels include soft labels for at least some of the training images, the soft label being determined from an empirical distribution of expert votes on the respective classes.

[0018] Embodiment 11: A computer device according to embodiment 9, wherein one of the following elements: a hair classification network having been trained with respective annotator confusion matrices comprising a plurality (n) of hue annotator confusion matrices and a plurality (n) of reflectance annotator confusion matrices, wherein n is defined from a total number of experts providing expert votes, and wherein the output of the hue classifier is multiplied by a respective matrix from among the plurality of hue confusion matrices and wherein the output of the reflectance classifier is multiplied by a respective matrix from among the plurality of matrices to predict the vote of a respective expert from among the plurality of experts and wherein the neural network and the respective annotator confusion matrices are trained together or,The hair classification network was trained according to a Mean Teacher framework.

[0019] Embodiment 12: Method of configuring a neural network that classifies hair hue and hair reflectance in an input image, the method comprising: the provision of the neural network, the neural network comprising a coding skeleton coupled to i) a hue classifier that outputs a hair hue value and ii) a reflectance classifier that outputs a hair reflectance value, each classifier comprising a linear classifier; the provision of a plurality of training images associated with respective training labels for each of the hair hue and hair reflectance, the target labels being prepared from respective expert votes by a plurality of experts;and training the neural network using the training images, the training being carried out in accordance with a sum of cross-entropy losses of hue and reflectance determined from i) the classifier outputs of the hair hue values ​​and hair reflectance values ​​and ii) the target labels.

[0020] Embodiment 13: Method according to embodiment 12, in which the training includes training the neural network using a Mean Teacher framework.

[0021] Embodiment 14: Method according to embodiment 12, wherein the training comprises training the neural network using the respective annotator confusion matrices comprising a plurality (n) of hue annotator confusion matrices and a plurality (n) of reflectance annotator confusion matrices, wherein n is defined from a total number of experts providing expert votes, and wherein the output of the hue classifier is multiplied by a respective matrix from among the plurality of hue confusion matrices and wherein the output of the reflectance classifier is multiplied by a respective matrix from among the plurality of reflectance confusion matrices to predict the vote of a respective expert from among the plurality of experts and wherein the neural network and the respective annotator confusion matrices are trained together.

[0022] Embodiment 15: Method comprising: providing a generative neural network that generates an output image combining a target hair color with an input image; providing a hair classification network as a component of a discriminator of the generative neural network, the hair classification network comprising i) a hue classifier that determines a hue of the output hair and ii) a reflectance classifier that determines a reflectance of the output hair; determining the training loss information using the target hair color and the output hair hue, and the output hair reflectance; and training the generative neural network under the direction of the hair classification network using the training loss information.

[0023] Embodiment 16: Method according to embodiment 15, wherein the target hair color, input image and output image are defined using RGB type data and the output hair tint and output hair reflectance are defined in accordance with an industry standard for hair color classification comprising a respective plurality of classes for hair tint and hair reflectance.

[0024] Embodiment 17: Method according to embodiment 15 comprising, prior to training the generative neural network, pre-training a color mapping network configured to process the input image and image hair color data from hair pixels of the input image and the target hair color to provide a color map to the generative neural network to generate the output image from the input image. Brief description of the drawings

[0025] [Fig-1] [Fig. 1] is a schematic diagram showing a classifier model according to prior art.

[0026] [Fig.2] Fig.2 is a schematic diagram showing a classifier model of reference in accordance with an embodiment.

[0027] [Fig.3] Fig.3 is a schematic diagram showing the classifier model of reference of [Fig.2] as adapted with annotator confusion matrices according to one embodiment.

[0028] [Fig.4] [Fig.4] is a diagram of an annotator confusion matrix representative.

[0029] [Fig.5] Fig.5 is a schematic diagram showing the classifier model of reference to [Fig.2] as adapted with a Mean Teacher frame according to one embodiment.

[0030] [Fig.6] The [Fig.6] is a graph, according to one embodiment, showing a comprehensive comparison of model prediction accuracy results between all models and an average human expert.

[0031] [Fig.7A] Fig.7A is a graph showing, according to one embodiment, a comparison of the accuracy of different numbers of training selfie images.

[0032] [Fig.7B] The [Fig.7B] is a graph showing, according to one embodiment, a comparison of the accuracy of different numbers of annotators ([Fig.7B]).

[0033] [Fig.8] Fig.8 is a schematic diagram of a computer device, in accordance with an embodiment.

[0034] [Fig.9A] The [Fig.9A] is a schematic diagram of network systems, according to one embodiment.

[0035] [Fig.9B] The [Fig.9B] is a schematic diagram of network systems, according to one embodiment.

[0036] [Fig.9C] The [Fig.9C] is a schematic diagram of network systems, according to one embodiment. DETAILED DESCRIPTION

[0037] The systems, methods, and techniques presented herein aim to improve existing hair coloring testing technology by leveraging artificial intelligence (AI). Many current hair simulation efforts are primarily based on traditional computer vision engineering techniques, and sometimes the hair texture in the simulated images appears synthetic—for example, obviously simulated and poorly simulated. The current applicant seeks to Using AI to enhance the photorealism of simulated hair images while maintaining color accuracy. A primary objective is color accuracy, and a hair color classification model is described here in one or more embodiments for identifying hair color in selfies.

[0038] According to one embodiment, hair color is defined based on L'Oréal's hair color definition for hair dye products and classifies hair color according to 1) shade and 2) reflectance, namely primary and secondary reflectance. Hair shades are divided into 10 classes, 1 being the lightest and 10 the darkest. In one embodiment, there are 9 reflectance classes for each of primary and secondary reflectance, where primary reflectance indicates the dominant reflectance in the hair, and secondary reflectance indicates the second dominant reflectance. Hair may have only primary reflectance or no reflectance at all. If the absence of reflectance is considered a separate reflectance class, there are 100 combinations of primary and secondary reflectances.Examples: no primary or secondary reflectance, golden primary reflectance without secondary reflectance, ash-colored primary reflectance without secondary reflectance, golden primary reflectance with ash-colored secondary reflectance and golden primary reflectance with mahogany secondary reflectance, etc.

[0039] Labeled datasets from hair coloring experts exist or can be obtained for hair color. However, such datasets, particularly their use, present challenges for building a hair color identification model. For example, hair colors are defined based on the judgment of professional hair experts rather than on scientific color definitions, such as the average RGB color. Consequently, the existing definition of color is subject to human bias. Since the annotations were provided by hair experts, the labeled data are expensive to obtain and may result in a limited number of labeled examples with which to perform training and verification, etc.

[0040] It was desirable to develop a hair color classifier to produce color predictions at least as accurate as those of an average human hair expert. Annotations of 1,000 selfies from 10 hair experts were received to define a dataset to be used to develop the model—the “expert dataset.” The model’s performance was evaluated based on its adequacy with the majority of the hair experts’ annotations or the discrepancy between the predictions and the experts’ annotations.

[0041] The expert dataset was relatively small, especially since the data were obtained gradually as the experts performed their tasks at different times. Furthermore, the combined results illustrated the differences in expert opinion. 1,000 labeled images represent a relatively small dataset given that there were over 100 hair colors. After gradually receiving labels from other hair experts, it was determined that the labels from different hair experts were often inconsistent. In some cases, there was no clear majority vote on the hair color, and all 10 experts thought the hair was different colors. Therefore, one of the objectives of the activities was to mitigate the impact of label noise.

[0042] To address these two challenges, a final model according to one embodiment comprises two improvements on a standard classifier (prior art) (for example, from [Fig. 1]), by adding annotator confusion matrices and using a Mean Teacher semi-supervised learning framework. The addition of annotator confusion matrices aims to mitigate the problem of label noise by modeling the biases of individual experts, while the Mean Teacher framework is a semi-supervised learning technique that mitigates the problem of limited labeled data by using unlabeled data during training. Label noise

[0043] In the real world, data are often labeled by multiple annotators. Since biases in human annotators are inevitable, these labels often contain noise, and annotators may disagree with each other. Studies show that the rate of corrupted labels in real-world datasets ranged from 8% to 39%

[19]

[16] [9][8]. A classic way to deal with label noise is to aggregate annotations using majority voting. However, the majority voting method has some limitations; for example, it only uses the class with the most votes as the aggregated label and ignores the distribution of votes, and there may be scenarios where there is no clear majority vote. Another way to aggregate annotations is to use soft labels in which the soft label values ​​are based on the empirical distribution of votes.

[0044] Model-based approaches have also been explored to handle label noise, including the use of a translation matrix to measure noise and the reweighting of labels according to confidence in them. The transition matrix is ​​widely used to model the relationship between the actual and observed distributions [5]

[13] . However, these methods often fail to account for the fact that different annotators may have different biases. On the other hand, loss reweighting has also been investigated to improve model robustness by assigning more weight to labels or models offering greater confidence

[10] . In particular, Weighted Doctor Net [4] proposed modeling individual annotators separately and then combining the predictions of the individual models using weightings resulting from training. This approach takes into account the overall skill level of annotators across all data classes and ignores the scenario in which annotators may have different skill levels when labeling different classes. Tanno et al.

[17] built upon this approach and combined the use of a transition matrix with new loss weighting, proposing to model annotator labels and real data separately using annotator-specific confusion matrices. Annotator Confusion Matrix

[0045] The annotator confusion model

[17] is a probabilistic model based on the assumptions that the annotators are independent and the label noise is independent of the input images. Based on these two assumptions, the common probability of observing n labels for an image from n annotators can be expressed as follows:

[0046] / -v) = n* jix) Eq 1

[0047] Where ÿ(n) is the annotation of the nth annotator and % is the image, y is the actual color and a(1) y"1,y is a value in the confusion matrix that indicates the probability that the annotator indicates an image as being of class y given that the actual class is y. During training, the annotator-specific confusion matrices are trained jointly with the classifier, in which the classifier learns to predict the actual distribution of labels and each confusion matrix learns the observed distribution of labels as a function of the actual labels for a particular annotator. Semi-supervised learning

[0048] Labeled data are often expensive to obtain, whereas unlabeled data are generally abundant and provide additional information about data distribution. To take advantage of the information contained in unlabeled data, semi-supervised learning (SSL) [3]

[21] is a research area that explores ways to train models using both labeled and unlabeled data. Typical semi-supervised learning techniques include consistency regularization and pseudo-labeling [7].

[0049] Consistency regularization was first introduced by the F model

[14] and suggests that the model should provide similar data points with consistent outputs. To encourage models to do this, the F model applies a loss of consistency between data points with and without noise. Following this, the Time Assembly [6] proposes to improve the model by maintaining an exponential moving average (EMA) prediction of the training data. The Mean Teacher

[18] is an extension of the F model and the Time Assembly, and suggests maintaining an exponential moving average of the model weights through a student-teacher framework. The student model is trained on the basis of supervised classification loss and unsupervised consistency loss with the teacher model. The consistency loss is applied to minimize the differences between the student and teacher model outputs.The teacher model weightings are updated based on the student model's EMA. The Mean Teacher framework has achieved excellent results for classifying semi-supervised learning images.

[0050] Recent state-of-the-art SSL models often combine several SSL techniques, notably the combination of consistency regularization with pseudolabeling [2][1]

[15] . FixMatch

[15] is based on the idea that data augmentation has a significant impact on consistency regularization

[20] , and proposes training the classifier to process consistent output for various augmentations of the same output using pseudolabels. This method produces excellent results in image classification tasks. However, because it relies on large augmentations that change the color of the images, it is not applicable to color classification. Model Architecture Reference Model

[0051] Figure 2 shows a schematic diagram of a processing flow 200 in accordance with to an embodiment in which a reference model 202 receives an image 204 as input and produces output predictions 206 comprising a hair tint 206A and a reflectance 206B. In one embodiment, the reference model 202 uses ResNetl8 202A as its skeleton. The output of the skeleton 202A is fed to two linear classifiers 202B and 202C, which separately predict the tint 206A and the reflectance 206B. In one embodiment, each linear classifier consists of two fully connected linear layers with an activation layer ReLu between the two layers. In one embodiment, the stream 200 represents a training stream at one training time. The components 202 are conserved for use during the inference time.

[0052] In one embodiment, the reference model 202 was trained using a sum of cross-entropy losses of hue and reflectance (e.g., 208) between the model outputs 206A / 206B and the target colors (e.g., soft labels 210) prepared from the expert labels. Hard labels and soft labels

[0053] During the training of the reference model in [Fig. 2], hard labels and soft labels are tested. Hard labels for hue and reflectance are calculated separately based on the class with the most votes. Soft labels are calculated based on the empirical distribution of labels for all annotations. For example, if nine experts label the image as class 1 and one expert labels it as class 2, the soft label will be 0.9 for class 1, 0.1 for class 2, and 0 for all other classes. Regardless of soft or hard labels, the reference models are trained based on cross-entropy loss (e.g., 208).

[0054] Classifier with annotator confusion matrices

[0055] Figure 3 is a schematic diagram of a processing flow 300 showing a classifier 302 according to an embodiment of Figure 2. The classifier 302 is adapted by learning using annotator confusion matrices 304 as described in more detail. The processing flow 300 is useful at a learning stage, and the components 302 are useful at an inference stage.

[0056] In one embodiment, the components 302A to 302C of the classifier 302 are structured in the same way as the components 202A to 202C of the classifier 202 but are adapted, for example, by training with the confusion matrices of annotators 304 as described. The output hue 306A and the reflectance 306B are also similar to the outputs 206A and 206B, but will depend on the trained classifier 302. The outputs of the respective classifiers for the same input image may vary in value.

[0057] According to one embodiment, the annotator confusion matrices 304 comprise a plurality (n) of hue annotator confusion matrices 304A and a plurality (n) of reflectance annotator confusion matrices 304B. Here, "n" refers to the number of annotators.

[0058] Figure 4 illustrates a representative hue annotator confusion matrix 400 according to one embodiment. The rows represent actual classes ("Actual Values") and the columns represent observed annotations ("Expert Annotations").

[0059] In one embodiment, the classifier 302 with annotator confusion matrices 304 is constructed on the basis of the reference model 202, having annotator-specific confusion matrices 304 after the classification layers The respective (202B / 202C) confusion matrices of the reference model 202 are used to predict the observed labels 306A / 206B from the annotators. Assuming there are "n" annotators, n 10*10 confusion matrices are added after the hue classifier and n 100*100 confusion matrices are added after the reflectance classifier. The outputs 306A / 306B of the hue and reflectance classifiers 302B / 302C are multiplied by the respective confusion matrices 304A / 304B to predict the annotations of individual annotators. For example, the outputs of the hue classifier 302B are multiplied by the first hue confusion matrix (e.g., 307) to predict the hue annotation for the first annotator.

[0060] During training, for each training data point, the cross-entropy loss sums of n hues and n reflectances (e.g., 308) are calculated to update the model. Since an annotator has only one hue label and one reflectance label for each training data point, each cross-entropy loss was calculated using hard labels (e.g., 310) instead of soft labels (e.g., 210). The confusion matrices are first randomly assigned values ​​such that diagonal values ​​dominate and the matrices are approximately identical. The matrices are trained with the rest of the model based on the classification loss 308, and a regularization term is added to encourage the trained matrices to converge to the actual confusion matrices (e.g., as in the example in [Fig. 4]). Mean Teacher Framework

[0061] Figure 5 is a schematic diagram of a processing flow 500 showing the reference classifier model 202 of Figure 2 adapted to a Mean Teacher frame for training. The flow 500 comprises a student model 502A and a teacher model 502B, according to one embodiment. Models 502A and 502B are configured as instances of model 202, so details are not shown in Figure 5. The student model 502A produces an output comprising the hue output 506A and the reflectance output 506B, and the teacher model 502B produces an output comprising the hue output 506C and the reflectance output 506D, respectively.

[0062] The student model 502A and the teacher model 502B are first pre-trained using the labeled data. During the semi-supervised learning process, two augmented versions of the same image (e.g., 504A and 504B) are fed separately to each of the 502A / 502B models. The student model 502B is trained using cross-entropy loss (supervised) 508 and consistency loss (unsupervised) 512. Cross-entropy loss (classification loss) is calculated on the basis of the labeled data (e.g., 510), and consistency loss is calculated as corresponding to the root mean square error of the Softmax outputs (506A, 506B, 506C, and 506D) are generated between the student model 502B and the teacher model 502A. The loss of consistency (512) based on the augmented images (504A, 504B) forces the student model 502A to provide consistent outputs at similar data points. After updating the student model 502A, the teacher model 502B is updated based on the exponential moving average of the weights of the student model 502A. According to one embodiment, during inference, only the teacher model 502B is retained and used for evaluation. Experiences Dataset

[0063] The selfie dataset used during the experiments contained 1,000 labeled images and 2,500 unlabeled images. For the labeled data, the hair in the selfie was annotated by 10 professional hair experts, and each label contained a hue, a primary reflectance, and a secondary reflectance. In one embodiment, the range of hue values ​​is from 1 to 10, where 1 corresponds to the darkest hue, and the range of primary and secondary reflectance values ​​is from 0 to 8. Hair color may also have no secondary or primary reflectance.

[0064] The labels of different annotators may not match. However, for most of the labeled images, there were at least two annotations that coincided. Out of the 1,000 images, only 40 did not coincide. When considering the hue, primary, and secondary reflectance values ​​individually, at least two annotators always agreed. To assess the quality of the annotators, Table 1 lists the accuracy and the average absolute hue difference between the labels of individual annotators and the majority vote. Overall agreement on color was about 30%; experts disagreed more on reflectance values ​​(49%) than on hues (64%). In particular, experts agreed on secondary reflectance values ​​only 42% of the time.

[0065] [Table 1]

[0066] [Tables 1] Annotator Avg. 1 2 3 4 5 6 7 8 9 10 Overall Accuracy (%) 30 33 20 31 21 41 39 34 30 28 21 Hue Accuracy (%) 64 81 55 68 53 70 73 66 59 61 55 Reflectance Accuracy (%) 49 45 52 18 18 59 65 67 63 58 49 Primary reflectance accuracy (%) 48 56 47 26 29 59 56 56 52 57 41 Secondary reflectance accuracy (%) 42 33 31 38 32 54 55 49 49 51 31 Experiment configuration

[0067] The reference model 202 and the classifier 302 with the confusion matrices 304 were trained using the stochastic gradient descent (SGD) optimizer. The learning rate was constant for the first 20 iterations and then decreased linearly to zero over the next 65 iterations. Simple data augmentations, including horizontal and vertical flips, rotations, and random translations, were used during training. The batch size was set at 128.

[0068] The Mean Teacher 502B frame teacher model was trained using the SGD optimizer, and the learning rate was brought to zero using simulated annealing based on the cosine function

[12] . For data augmentation, random data augmentations, including horizontal and vertical flipping, rotation, translation, and shear, were used. The batch size was set at 256, with 128 labeled and 128 unlabeled data points. Evaluation metrics

[0069] In one embodiment, the model's performance was evaluated based on the acceptance of the model by the majority of experts and the magnitude of the discrepancy between the predictions and the annotations. Three main metrics were used for the evaluation: accuracy, hue ±1, and mean absolute difference in hue. For hue, reflectance, and overall color, accuracy was measured in terms of the model's prediction of the color that received the most votes from the ten annotators.Since the hues are ordinal data, two additional metrics were used: the hue at ±1 measures the probability that the model's predictions fall within a hue difference from the hue that received the most votes, and the mean absolute difference measures the average distance between the predicted hues and the mean hue annotations for each image. The experimental results were based on a 5-subset cross-validation, and the metrics from the best iteration were reported at each run. Results soft labels / hard labels

[0070] Table 2 compares the prediction performance between an average human expert, the reference model trained with hard labels, and the reference model trained with soft labels, according to one embodiment. In case of disagreement Among the experts, training the model with soft labels allowed for greater insight into vote distribution and improved predictions. The prediction accuracy of the model trained with soft labels was 9% higher than that of the model trained with majority votes, with a 7% increase in hue accuracy and a 6% increase in reflectance accuracy. Another observation was that the model trained with majority votes performed slightly worse than the average human expert, while the model trained with soft labels performed better than the average human expert.

[0071] [Table 2]

[0072] [Tables2] Human Expert Majority Soft labels Overall accuracy (%) 30+7 29+4 38+4 Hue accuracy (%) 64+9 62+6 69+6 Reflectance accuracy (%) 49+18 41+5 47+4 Reflectance accuracy (%) 48+12 60+6 66+6 imaire Classifier with confusion matrices

[0073] In one embodiment, stepwise experiments were conducted to examine the impacts of adding confusion matrices to the model and training the model using the Mean Teacher framework. In short, either model outperformed the average human expert. As shown in Table 3, adding confusion matrices to the model improved overall color accuracy (2%). It should be noted that the reference model trained with soft labels is a special case of the model with confusion matrices where all the confusion matrices were identity matrices. The improvement in color classifications is primarily due to improved reflectance prediction (5%), and the hue prediction accuracy is similar to the reference model. It is understood that this results from more disagreements on reflectance labels among hair experts.

[0074] [Table 3]

[0075] [Tables3] Human Expert Reference Confusion Average Overall Accuracy (%) 30+7 38+4 40+6 50+5 Hue Accuracy (%) 64+9 69+6 67+3 73+4 Reflectance accuracy (%) 49+18 47+4 53+6 58+5 Primary reflectance accuracy (%) 48+12 66+6 68+6 73+2 Secondary reflectance accuracy 42+10 54+3 60+5 63+3

[0076] Although the addition of annotator confusion matrices did not improve hue accuracy, a more detailed analysis showed that its prediction deviated less from the experts' annotations. Table 4 shows a comparison of hue prediction performance between human annotators, the reference model with soft labels, the classifier with annotator confusion matrices, and the Mean Teacher framework. As shown in Table 4, after the addition of the confusion matrices, the percentage of hues within ±1 of the hue receiving the most votes was 1% higher, and the mean absolute differences were 0.04 lower. This revealed that hue prediction was also improved.

[0077] [Table 4]

[0078] [Tables4] Reference Model Confusion Matrix Mean Teacher Hue at +1 (%) 97+0 98+0 98+2 Abs. dif. mean hue 0.44+0.07 0.40+0.07 0.32+0.07 Cadre Mean Teacher

[0079] After integrating the Mean Teacher framework, the overall performance of the model was significantly improved. Tables 3 and 4 show that the framework improved prediction accuracy by 10% compared to the confusion matrix model, and also reduced the difference between the model's prediction and the experts' annotation. The mean absolute difference was reduced by 0.08. Furthermore, when comparing the classifier based on the Mean Teacher framework with an average human expert, there was a 20% improvement in overall color accuracy, with improvements of 9% and 11% in hue and reflectance accuracy, respectively. Figure 6 is a graph, according to one embodiment, showing a comprehensive comparison of the model prediction accuracy results between all models and an average human expert.Figure 600 shows that there is a significant improvement between the accuracy of an average human expert and the Mean Teacher framework model. Ablation study

[0080] To facilitate better data collection planning, an experiment was conducted to investigate the impact of having additional labeled selfies and annotators. In subsequent experiments, the reference model with soft labels was trained to assess the impact of having more images and more labels per image. When examining the impact of additional labeled images, the number of labeled data points varied from 100 to 900. When examining the impact of additional labels per image, the labels of a subset of experts were used to train the model.

[0081] The model's performance improved when it had more selfies or more annotators than expected. However, as shown in Figures 7A and 7B, the model improved even more when it had more labeled images. [Fig. 7A] is a 700 graph showing a comparison of the accuracy of different numbers of training selfies, and [Fig. 7B] is a 710 graph showing a comparison of the accuracy of different numbers of annotators.

[0082] It has been observed that a steady increase in accuracy is achieved when increasing the number of training images from 100 to 900, with overall, hue, and reflectance accuracies increasing in parallel. No downward trend in improvement was observed, even after using 900 images. In contrast, only a minimal improvement in overall accuracy was observed between 5 and 10 annotators; in particular, a smaller improvement in reflectance accuracy was observed after using 8 annotators. This demonstrated that having additional labeled images was more effective in improving model performance. Hair simulation application

[0083] Figure 8 is a functional diagram of a computer device 800 according to one embodiment. In this embodiment, a color classifier according to another embodiment is integrated into a VTO application to provide a hair color simulation. The computer device 800 provides a user computing device, such as a smartphone, tablet, laptop, or other computing device, intended for use by a user such as a consumer of hair dye products or a salesperson assisting such a user. The device may include a component of a larger form factor, such as a kiosk for placement in a retail environment, for example. The device 800 is not exhaustive and is simplified for the sake of brevity.

[0084] The device 800 includes a storage device 802, a processing unit 806, a camera 808, a microphone / speaker 810, a display screen 812 and A communication subsystem 814. In one embodiment, the storage device 802 includes a memory device, for example, one or more types of memory such as RAM, ROM, etc. The storage device 802 may include a long-term storage device such as an integrated circuit disk (SSD) or another type of drive for storing persistent data, etc. The storage device 802 stores computer-readable instructions for execution by the processing unit, such that once executed, the instructions cause the computing device to perform operations such as one or more processes. The processing unit 806 includes one or more central processing units (for example, CPUs), and / or graphics processing units (for example, GPUs) having one or more processors / microprocessors, controllers / microcontrollers, etc. Other types of processors may be used.GPUs can be particularly useful for accelerating graphics processing tasks and / or AI processing tasks (e.g., training and / or inference).

[0085] The camera 808 can be used to take selfies. The microphone and speaker 810 are usually separate devices, but are labeled together here for convenience and represent some of the input (I), output (O), or I / O devices that may be available. Other devices may include a light, a buzzer, a vibrator, a button, a keyboard, a pointing device (e.g., mouse or touchpad, etc.), etc.

[0086] The display screen 812 presents images such as components of a graphical user interface, camera images, etc. In one embodiment, the display screen is a touch screen device, a type of I / O device, configured to receive gesture inputs (e.g., swiping, tapping, etc.) which interact with the screen region(s) and in association with the user interface components (e.g., controls) presented by an application executed by the processing unit 806.

[0087] The communication subsystem 814, in one embodiment, is configured to manage communications between device components and / or between the computing device and external devices such as a remotely located computing device (e.g., web service, cellular network component, printer, etc.). A subcomponent thereof in one embodiment is an antenna for wireless communication. A subcomponent thereof in one embodiment is a wired communication interface (e.g., Ethernet, USB-A, USB-C, Thunderbolt (TM of Intel Corporation), etc.) to be coupled with a suitable cable for wired communication.

[0088] The storage device 804 stores the components of a VTO application and the corresponding data (e.g., 820). Representative components are represented. The VTO application 820 includes the user interface component 822 (e.g., screens, instructions, icons, commands, etc.). The user interface provides outputs to a user and receives inputs such as inputs for the application's workflow, user selections of color choices, etc. A color classifier 824 is provided and includes one of the classifiers 202, 302, and 502 as described previously to determine hair color information (e.g., hue and reflectance). A VTO pipeline 826 is provided to simulate hair color in association with an input image (e.g., on or within the image). A color recommendation engine 828 and an associated data store 830 are provided.In one embodiment, the data store 830 stores hair color data and representative images thereof, such as color sample images or dyed hair, hair reflectance images, etc. The user interface may present one or more choices that a user can select via user input, for example, sorting or filtering examples and presenting them in groups / pages, etc. In one embodiment, the color recommendation engine 828 may include an interface to a chatbot or live agent to discuss a recommendation.

[0089] In one embodiment, recommendations can be made, for example, based on recommendation factors such as a user's personal information, including current hair color and age; hair color trend data; the availability of a local product and / or service for the user (e.g., within a given radius); cost information, etc. In one embodiment, rules or another method of recommendation can be used to determine the recommendation (or recommendations) to be presented to a user.

[0090] In one embodiment, a user can provide an input image with hair (e.g., 832) similar to the input image 204 for classification. The hue and reflectance data 834 are generated by the classifier 824. The hue and reflectance data can be made available to the color recommendation engine 828 for processing (e.g., using color matching rules, color combination rules, etc.) to select one or more colors from the color data store 830 to recommend to a user via the user interface component 822. The recommendation can be displayed on the screen 812. The user can invoke the VTO engine (via an input to a command) for the engine to simulate the color and / or the color and reflectance (e.g., target hair) using image 832 and target hair data as input to produce an output image with simulated hair 836.

[0091] The output image with simulated hair 836 can be presented via the user interface component 822 on the display screen 812. A before-and-after display can be provided for comparison. A comparison between at least two simulated colors (e.g., two different output images) can be displayed for comparison. Products or services, or both, can be purchased via the interface 838. Such an interface 838 can direct the user (e.g., the computer device) to a web-based e-commerce service (e.g., a website (not shown)) to complete the purchase, reservation, or other transaction.

[0092] In one embodiment, the VTO 826 pipeline includes a generative neural network (e.g., a model) configured to generate a simulated hair image using the input image 832 and the target hair data as input. See Figure 9 discussed below.

[0093] Other components stored in the storage device 804 include an operating system 840, a browser 842 (for example, for browsing web pages), an email and / or messaging application 844 (for example: SMS or other type) and a social networking application 846. The output image with simulated hair 836 can be shared (for example, communicated) via the applications 844 and / or 846, for example.

[0094] In another embodiment of the use of the color classifier, not shown in [Fig. 8], the color classifier 824 processes an output image from a generative model configured to generate a simulated hair image using an input image having target hair and color data as input. The classifier determines the color and / or reflectance data for comparison with the target hair data to confirm the accuracy of the generative model.

[0095] The following aspects and features mentioned in the numbered embodiments (class 1 embodiment, class 2 embodiment... class 20 embodiment) will become clear, in particular, from this disclosure.

[0096] Embodiment - Class 1: A computer device comprising a processor coupled to a storage device that stores instructions which, when executed by the processor, cause the computer device to: classify hair color and hair reflectance in an input image using a neural network, the neural network comprising an encoding skeleton coupled to a pair of classifiers comprising i) a color classifier that outputs a hair color value and ii) a reflectance classifier that outputs a hair reflectance value, each of the two classifiers comprising a linear classifier, the neural network having been defined by training with a cross-entropy loss sum of hue and reflectance determined from i) the results of hair hue values ​​and hair reflectance values, and ii) target labels for hair hue and hair reflectance, the target labels prepared for the training images from the respective expert votes by a plurality of experts.

[0097] Embodiment class 2: Computer device according to embodiment class 1, wherein the neural network is configured to classify hair shade and hair reflectance according to a plurality of respective classes to follow an industry standard for hair color.

[0098] Embodiment - Class 3: Computer device according to embodiment - Class 1, wherein the hair reflectance comprises a primary reflectance component and a secondary reflectance component

[0099] Embodiment class 4: Computer device according to embodiment class 1, wherein the target labels include soft labels for at least some of the training images, the soft label being determined from an empirical distribution of expert votes on the respective classes.

[0100] Embodiment class 5: Computer device according to embodiment class 1, wherein the neural network has been trained with respective annotator confusion matrices comprising a plurality (n) of hue annotator confusion matrices and a plurality (n) of reflectance annotator confusion matrices, wherein n is defined from a total number of experts providing expert votes, and wherein the output of the hue classifier is multiplied by a respective hue confusion matrix from among the plurality of hue confusion matrices and wherein the output of the reflectance classifier is multiplied by a respective reflectance confusion matrix from among the plurality of matrices to predict the vote of a respective expert from among the plurality of experts and wherein the neural network and the respective annotator confusion matrices are trained together.

[0101] Embodiment mode - class 6: Computer device according to embodiment mode- class 1, in which the neural network has been trained in accordance with a Mean Teacher framework.

[0102] Embodiment mode -class 7: Computer device according to embodiment mode-Class 6, wherein the neural network includes a teacher network obtained from the Mean Teacher framework.

[0103] Embodiment - Class 8: Method comprising: receiving an input image and classifying hair color and hair reflectance in the input image using a neural network, the neural network comprising an encoding skeleton coupled to a pair of classifiers comprising i) a hue classifier that outputs a hair hue value and ii) a reflectance classifier that outputs a hair reflectance value, each element of the pair of classifiers comprising a linear classifier, the neural network having been defined by learning with a cross-entropy loss sum of hue and reflectance determined from i) outputs of hair hue values ​​and hair reflectance values, and ii) target labels for hair hue and hair reflectance, the target labels being prepared for image learning from respective expert votes by a plurality of experts; the provision of hair hue and hair reflectance.

[0104] Embodiment - class 9: Method according to embodiment- class 8, wherein hair tint and hair reflectance are provided to train a generative model to simulate hair color.

[0105] Embodiment class 10: Method according to embodiment class 9, wherein the neural network is configured to classify hair shade and hair reflectance according to a plurality of respective classes to follow an industry standard for hair color, and wherein hair reflectance comprises a primary reflectance component and a secondary reflectance component.

[0106] Embodiment class 11: Computer device according to embodiment class 9, wherein the neural network has been trained with respective annotator confusion matrices comprising a plurality (n) of hue annotator confusion matrices and a plurality (n) of reflectance annotator confusion matrices, wherein n is defined from a total number of experts providing expert votes, and wherein the output of the hue classifier is multiplied by a respective matrix from among the plurality of hue confusion matrices and wherein the output of the reflectance classifier is multiplied by a respective matrix from among the plurality of reflectance confusion matrices to predict the vote of a respective expert from among the plurality of experts and wherein the neural network and the respective annotator confusion matrices are trained together.

[0107] Implementation mode -class 12: Method according to implementation mode-class 9, in which the neural network has been trained in accordance with a Mean teacher framework.

[0108] Embodiment mode - class 13: Computer device according to embodiment mode- class 12, in which the neural network includes a teacher network obtained from the Mean Teacher framework.

[0109] Embodiment - Class 14: A method according to embodiment - Class 9, wherein the target labels include soft labels for at least some of the images training, the soft label being determined from an empirical distribution of expert votes on the respective classes.

[0110] Embodiment class 15: Method according to embodiment class 14, wherein the input image is associated with at least one of the target labels and wherein step b is carried out to train the neural network using the input image.

[0111] Embodiment class 16: Method according to embodiment class 15, comprising training the neural network with respective annotator confusion matrices comprising a plurality (n) of hue annotator confusion matrices and a plurality (n) of reflectance annotator confusion matrices, wherein n is defined from a total number of experts providing expert votes, and wherein the output of the hue classifier is multiplied by a respective matrix from among the plurality of hue confusion matrices and wherein the output of the reflectance classifier is multiplied by a respective matrix from among the plurality of reflectance confusion matrices to predict the vote of a respective expert from among the plurality of experts and wherein the neural network and the respective annotator confusion matrices are trained together.

[0112] Embodiment - Class 17: Computer software product comprising a non-transient storage device that stores computer-executable instructions which, when executed by a processor of a computer device, train the computer device to: classify hair hue and hair reflectance in an input image using a neural network, the neural network comprising an encoding skeleton coupled to a pair of classifiers comprising i) a hue classifier that outputs a hair hue value and ii) a reflectance classifier that outputs a hair reflectance value, each of the classifiers comprising a linear classifier, the neural network having been defined by training with a cross-entropy loss sum of hue and reflectance determined from i) the outputs of the hair hue values ​​and the hair reflectance values,and ii) target labels for hair color and hair reflectance, the target labels prepared for the training images from the respective expert votes by a plurality of experts.

[0113] Embodiment - class 18: Computer software product according to embodiment- class 17, wherein hair tint and hair reflectance are provided to train a generative model to simulate hair color.

[0114] Embodiment - Class 19: Computer software product according to embodiment - Class 17, wherein the execution of instructions causes the computer device to classify hair color and hair reflectance in accordance with a plurality of respective classes to follow an industry standard for hair color and in which the hair reflectance comprises a primary reflectance component and a secondary reflectance component.

[0115] Embodiment class 20: Computer software product according to embodiment class 17, wherein one or both of the following: the execution of instructions causes the computer device to provide the neural network for classification in which the network has been trained with respective annotator confusion matrices comprising a plurality (n) of hue annotator confusion matrices and a plurality (n) of reflectance annotator confusion matrices, in which n is defined from a total number of experts providing expert votes,and wherein the output of the hue classifier is multiplied by a respective hue confusion matrix from among the plurality of hue confusion matrices, and wherein the output of the reflectance classifier is multiplied by a respective matrix from among the plurality of reflectance matrices to predict the vote of a respective expert from among the plurality of experts; and wherein the neural network and the respective annotator confusion matrices are trained together; or alternatively, the execution of the instructions causes the computing device to provide the neural network for classification in which the network has been trained according to a Mean Teacher framework.

[0116] Classifier-guided training of a color-finement neural network

[0117] In one embodiment, a classifier can guide the training of a color-finement neural network configured to produce rendered images with a target hair color. It is difficult to assess the accuracy of hair rendering based solely on RGB values ​​or simple metrics. In one embodiment, a feedback loop exists within a rendering network system to use a color classifier trained according to this embodiment to guide the rendering network. Figure 9A is a schematic diagram of a 900A network system according to one embodiment providing a training system. Figure 9B is a schematic diagram showing further details of a 900B training network system linked to the 900A system, according to one embodiment.The process, the computer program product, and other aspects can be readily understood by a person skilled in the art from an understanding of the training and inference aspects shown and described.

[0118] The 900A network system displays the input image 104 (for example, image I) and a target hair color 902 supplied to a hair analysis engine 904 to determine the hair information for the hair pixels (by example, from hair segmentation), for example, hair histogram color data for image I. In one embodiment, the color data of image I, target hair 902 (for example, as a hair sample image) and hair histogram include RGB data (for example, 256*3) commonly used for images. In one embodiment, the hair analyzer engine 904 understands or communicates with a hair analyzer such as a hair classifier network to determine the hair pixels in the input image 104. In one embodiment, the image I and color information (e.g., histogram data) for the target hair color 902 and the hairs in image I (e.g., in histogram form (not shown as such in [Fig.9A])) are provided to a color refinement neural network 906 to produce a rendered image 908.In one embodiment, as shown in [Fig. 9B], the color refinement neural network 906 comprises a color mapping network and a generative network configured to modify features of an image it processes, namely hair color. The generative network in one embodiment comprises a model (e.g., generator (G)) which is defined by training guided by the hue and reflectance classifier 910.

[0119] Referring again to [Fig. 9A], the rendered image 908 is provided to the color classifier 910 defined (e.g., trained) according to one embodiment here to classify the hue and reflectance properties in order to produce the color prediction 912 for the image 908. The loss 914 represents a loss determined from the color prediction 912 and the target hair color 902 as used to train the color-finement neural network 906 as described in more detail with [Fig. 9B]. In one embodiment, the system 900A is configured with one or more computing devices (not shown) to provide the computing components (e.g., 904, 906, 910, etc.) and to store data (e.g., 104, 902, 908, 912, 914, etc.).A display device (not shown) may be included to display image 104, 908 and any of the other data, including the output (not shown) of the hair analyzer engine 904, supplied to the color refinement neural network 906.

[0120] In one embodiment, the network system components 900A and 900B can be configured as an inference time system (e.g., following training) to provide a virtual trial experience for simulating a hair color applied to an image. For example, components 904 and 906 are useful for defining a VTO pipeline (e.g., 920) to simulate hair color, for processing an input image to simulate a target hair color and produce a simulated image. In one embodiment, the color refinement neural network 906 comprises a training-defined neural network as guided by the color classifier 910 as described in more detail below.

[0121] According to Figures 9A and 9B, a color mapping network 922 is trained to predict correct rendering parameters based on the classification results of the color classifier (910). As shown in [Fig.9B], the pipeline uses the color mapping network 922 and a generative adversarial network (GAN) 924 with a generator (G) 924A to simulate capillary rendering 908 (i.e. (G})).

[0122] A hair segmentation model (for example, as a component of the 904 hair analyzer engine, and not shown in [Fig. 9B]) is used to extract hair masks, and the 904 engine provides RGB histogram data of hair pixels (Hl) in (i.e., from) the / 104 image. The 922 color mapping network includes an encoder 926 that encodes image features (F) from the / image. The 926 encoder encodes features including, but not limited to, the color and lighting of the hair representation, and in the generator, it is used to regenerate the same hair rendering.

[0123] A concatener 928 of the color mapping network 922 receives inputs: i) RGB histogram of the sample (i.e. target hair color) (HT)

[0124]

[0125]

[0126]

[0127] 902, ii) source RGB hair histogram (Hl) 930 (image I) and iii) image features (F) of encoder 926. The three inputs are concatenated and passed through a fully connected layer block 932 for the output color map (Mp). The output color map Mp and the source RGB hair histogram Hl 930 are concatenated and used as input and passed through another fully connected layer block (FC Block 2 (934)) to obtain the output RGB histogram (H0) to be used by generator (G) 924A to produce its output from image I. Therefore, the rendered output image G1 908 is produced by replacing the hair pixels with the color map. Internal color maps (Mc) and RGB histograms (HG) are used as real data labels to train the 922 color mapping network in a supervised manner. In one embodiment, the following losses are used to train the 922 color mapping network: PcrtC — L Color Mapping P RGB Histogram L2 loss between the actual data color mapping (M c) and the network-predicted color mapping: 2 ^Color-mapping ~ ln \^G ~ Mp)

[0128] L1 loss between the histograms of the real data RGB histogram (H c) and the output RGB histogram (HO):

[0129] j ij i ^RGB histogram ~ Z—y I ■“ G " O i

[0130] with the trained color mapping network 922, the GAN 924 can be trained, as follows, according to one embodiment.

[0131] The output RGB histogram (H0) from the color mapping network 922 is used as a condition with the source image 104 as inputs to the generator (G)924A of the GAN 924. In one embodiment, the GAN 924 is defined as a StarGAN (in accordance with the teaching of Choi, Y., et al, "StarGAN: Unified Generative Adversarial Networks for Multi-Domain translation-to-Image Translation", IEEE Conference on Pattern Recognition and Computer Vision (CVPR), 2018, pp. 8789-8797), such that G1 = G(I, H0).

[0132] It is noted that the previously trained color mapping network 922 is frozen during GAN training. That is, network 922 is not trained or co-trained with the training of GAN 924.

[0133] The discriminating part (discriminator (D) 924B) of the GAN 924 uses two classifiers, one (936) which classifies whether the image is false or real ((G ; ) or I); and the other (910) is the pre-trained classifier which classifies the two parts of the color: the hue and the reflectance.

[0134] The pre-trained classifier guides and enhances the generative capabilities of the GAN network and can be used independently with any GAN architecture for counseling purposes.

[0135] The losses used for the drive of the generator parts 924A and discriminator parts 924B are the same as those defined in the StarGAN document referenced above.

[0136] In one embodiment, the color mapping + GAN is labeled as a color refinement network, because it uses the instructions of the pre-trained hair classification model and the losses indicated above so that the generative model refines its outputs.

[0137] As shown in the simplified block diagram in [Fig. 9C] illustrating a VTO application, in one embodiment, at the inference time (for example, when training is complete and the refinement network is provided to generate output images, in particular from user input), the discriminator component 924B with its two classifiers 910 and Conclusion

[0138] In this work, techniques, etc., are provided for addressing the problem of the limited availability of labeled data and label noise caused by human bias, for classifying the color of attributes in images, through a combination using a semi-supervised learning framework and annotator confusion matrices, the final model achieved 20% greater accuracy than an average human expert.

[0139] The experiments described herein have shown that the confusion matrix is ​​useful for improving color prediction, particularly in scenarios where there is more disagreement among experts. The approach primarily considers label noise caused by individual expert biases and assumes that the biases depend only on the actual class (e.g., without considering other factors such as lighting and background contrast). Due to the noisy nature of subjective annotation, in one embodiment, soft labels are used but not majority votes, as the training labels and confusion matrices are adapted to represent different expert ideas.

[0140] Experiments as described herein have also shown that the Mean Teacher framework helps to significantly improve color prediction. The effectiveness of the current framework depends on the teacher model providing good labels for the student model's learning.

[0141] Unlike normal RGB value prediction, the classification model is designed to classify hair according to a color standard specifically defined for it. In one embodiment, such a standard is three-digit, with the first digit representing the natural hair shades and the second and third digits representing the primary and secondary reflectance of the hair. Such a color prediction adapts to different lighting conditions and therefore accurately predicts dark hair in bright light and light hair in dim light. The color representation is independent of light. The color number prediction can be directly related to how hair products are color-coded (e.g., using the same color standard), which generally cannot be expressed in RGB or simple text information.

[0142] The structure of the model jointly uses one branch for the natural tint (first number) and one branch for the two reflectance determinations in order to obtain better performance.

[0143] A practical implementation may include all or part of the features described herein. These features, characteristics, and various combinations, as well as others, may be expressed in terms of processes, apparatus, systems, means for performing functions, program products, and other ways of combining the features described herein. A number of embodiments have been described. Nevertheless, it is understood that various modifications may be made without departing from the spirit and scope of the processes and techniques described herein. Furthermore, other steps may be provided, or steps may be eliminated, from the described process, and other components may be added to or removed from the described systems.

[0144] Throughout the description of this patent specification, the terms "include," "contain," and variants thereof mean "including but not limited to" and are not intended to exclude (and do not exclude) other components, integers, or steps. Throughout this patent specification, the singular encompasses the plural unless the context requires otherwise. In particular, where the indefinite article is used, this patent specification is to be understood as considering both plurality and singularity unless the context requires otherwise.

[0145] The features, integers, characteristics, or groups described in conjunction with a particular aspect, embodiment, or example of the invention shall be understood as applicable to any other aspect, embodiment, or example, unless inconsistent with them. All features disclosed herein and / or all steps of any method or process so disclosed may be combined in any combination, except combinations in which at least some of these features and / or steps are mutually exclusive. The invention is not limited to the details of the preceding examples or embodiments. The invention extends to any new feature, or any new combination, of the features disclosed in this patent memorandum or to any new feature, or any new combination, of the steps of any method or process disclosed. REFERENCES

[0146] The following documents are cited in the respective title in their respective integrity:

[0147] Berthelot, D., Carlini, N., Cubuk, E.D., Kurakin, A., Sohn, K., Zhang, H., Raffel, C.: Remix-match: Semi-supervised leaming with distribution alignment and augmentation anchoring. arXiv preprint arXiv:1911.09785 (2019)

[0148] Berthelot, D., Carlini, N., Goodfellow, I., Papernot, N., Oliver, A., Raffel, CA: Mixmatch: A holistic approach to semi-supervised leaming. Advances in neural information processing Systems 32 (2019)

[0149] Chapelle, O., Schôlkopf, B., Zien, A.: Introduction to semi-supervised leaming (2006)

[0150] Guan, M., Gulshan, V., Dai, A., Hinton, G.: Who said what: Modeling individual labelers improves classification. In: Proceedings of the AAAI conference on artificial intelligence, vol. 32 (2018)

[0151] Hendrycks, D., Mazeika, M., Wilson, D., Gimpel, K.: Using trusted data to train deep networks on labels corrupted by severe noise. Advances in neural information Processing Systems 31 (2018)

[0152] Laine, S., Aila, T.: Temporal ensembling for semi-supervised learning. arXiv preprint arXiv:1610.02242 (2016)

[0153] Lee, D.H., et al.: Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In: Workshop on challenges in représentation learning, ICML. vol. 3, p. 896 (2013)

[0154] Lee, K.H., He, X., Zhang, L., Yang, L.: Cleannet: Transfer learning for scalable image classifier training with label noise. In: Proceedings of the IEEE conférence on computer vision and pattern récognition, pp. 5447-5456 (2018)

[0155] Li, W., Wang, L., Li, W., Agustsson, E., Van Gool, L.: Base de données Webvision : Visual learning and understanding from web data. arXiv preprint arXiv: 1708.02862 (2017)

[0156] Liu, T., Tao, D.: Classification with noisy labels by importance reweighting. IEEE Transactions on pattern analysis and machine intelligence 38(3), 447-461 (2015)

[0157] Liu, Z., Luo, P., Wang, X., Tang, X.: Deep learning face attributes in the wild. In: Proceedings of International Conférence on Computer Vision (ICCV) (December 2015)

[0158] Loshchilov, I., Hutter, F.: Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv: 1608.03983 (2016)

[0159] Patrini, G., Rozza, A., Krishna Menon, A., Nock, R., Qu, L.: Making deep neural networks robust to label noise: A loss correction approach. In: Proceedings of the IEEE conférence on computer vision and pattern récognition, pp. 1944-1952 (2017)

[0160] Rasmus, A., Berglund, M., Honkala, M., Valpola, H., Raiko, T.: Semi-supervised learning with ladder networks. Advances in neural information processing Systems 28 (2015)

[0161] Sohn, K., Berthelot, D., Carlini, N., Zhang, Z., Zhang, H., Raffel, C.A., Cubuk, E.D., Kurakin, A., Li, C.L.: Fixmatch: Simplifying semi-supervised learning with consistency and confidence. Advances in neural information processing Systems 33, 596-608(2020)

[0162] Song, H., Kim, M., Lee, J.G.: Selfie: Refurbishing unclean samples for robust deep learning. In: International Conférence on Machine Learning. pp. 5907-5915. PMLR (2019)

[0163] Tanno, R., Saeedi, A., Sankaranarayanan, S., Alexander, D.C., Silberman, N.: Learning from noisy labels by regularized estimation of annotator confusion pp. 11244-11253(2019)

[0164] Tarvainen, A., Valpola, H.: Mean teachers are better rôle models: Weight-averaged consistency targets improve semi-supervised deep learning results. Advances in neural information processing Systems 30 (2017)

[0165] Xiao, T., Xia, T., Yang, Y., Huang, C., Wang, X.: Learning from massive noisy labeled data for image classification. In: Proceedings of the IEEE conférence on computer vision and pattern récognition, pp. 2691-2699 (2015)

[0166] Xie, Q., Dai, Z., Hovy, E., Luong, T., Le, Q.: Unsupervised data augmentation for consistency training. Advances in Neural Information Processing Systems 33, 6256-6268 (2020)

[0167] Zhu, X., Goldberg, A.B.: Introduction to semi-supervised learning. Synthesis lectures on artificial intelligence and machine learning 3(1), 1-130 (2009)

Claims

Demands

1. A computing device comprising a processor coupled to a storage device containing instructions which, when executed by the processor, cause the computing device to: (a) process an input image and a target hair color using a virtual testing pipeline (VTO) to produce a VTO experiment which simulates the target hair color with the input image to produce an output image;(b) wherein the VTO pipeline includes a color refinement neural network to generate the output image combining the target hair color and the input image, the color refinement neural network comprising a generative neural network trained under the direction of a hair classification network comprising a hue classifier that outputs a hair hue value and a reflectance classifier that outputs a hair reflectance value to determine training loss information from output images to train the generative neural network.

2. A computer device according to claim 1, wherein instructions are executed by the processor to cause the computer device to offer an interface to either of: i) a color recommendation engine to recommend a target hair color; and ii) an ecommerce service with which to purchase a product or service, or both.

3. A computer device according to claim 1, wherein the VTO pipeline comprises a color mapping network trained to produce a color map for input into the generative neural network in order to produce the output image for the VTO experiment, the color map produced from the image features, image hair color data determined from the image and the target hair color.

4. A computer device according to claim 3, wherein the color mapping network comprises an encoder for determining features from the input image; a concatenator for combining the features, the color of target hair and image hair color data for processing by a first fully connected block, and a second block to process an intermediate map from the first block with the image hair color data to produce the color map for input into the generative neural network.

5. Computer device of claim 4, wherein the VTO pipeline includes a hair analysis engine for determining image hair color data from the input image.

6. A computer device according to claim 1, wherein the input image and target hair color are defined using RGB values, such that the generative neural network is trained to produce an output image using particular RGB values ​​as directed during training by the hair classification network using hue values ​​and hair reflectance values.

7. Computer device according to claim 6, wherein the hair classification network is defined and trained according to an industry standard for hair color classification comprising a respective plurality of classes for hair tint and hair reflectance.

8. Computer device according to claim 1, wherein the hair classification network comprises a coding skeleton and the hair shade classifier and reflectance classifier each comprising a respective linear classifier.

9. Computer device according to claim 8, wherein the hair classification network has been defined by training with a cross-entropy loss sum of hue and reflectance determined from i) outputs of hair hue values ​​and hair reflectance values, and ii) target labels for hair hue and hair reflectance, the target labels being prepared for the training images of the hair classification network from respective expert votes by a plurality of experts.

10. A computer device according to claim 9, wherein: (a) the hair classification network having been trained with respective annotator confusion matrices comprising a plurality (n) of shade annotator confusion matrices and a plurality (n) of reflectance annotator confusion matrices, wherein nest is defined from a total number of experts providing expert votes, and wherein the output of the hue classifier is multiplied by a respective matrix from among the plurality of hue confusion matrices and the output of the reflectance classifier is multiplied by a respective matrix from among the plurality of reflectance matrices to predict the vote of a respective expert from among the plurality of experts, and wherein the neural network and the respective annotator confusion matrices are trained together or (b) the hair classification network having been trained in accordance with a Mean Teacher framework.