HAIR COLOR CLASSIFICATION USING CONSISTENCY REGULATION AND ANNOTATION CONFUSION MATRICES

A neural network-based hair color classification model using annotator confusion matrices and semi-supervised learning addresses human bias and data limitations, achieving improved accuracy and realism in virtual try-on technologies.

FR3160261B3Active Publication Date: 2026-04-17LOREAL SA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
FR · FR
Patent Type
Utility models
Current Assignee / Owner
LOREAL SA
Filing Date
2024-03-15
Publication Date
2026-04-17

Smart Images

  • Figure 00000034_0000
    Figure 00000034_0000
  • Figure 00000034_0001
    Figure 00000034_0001
  • Figure 00000035_0000
    Figure 00000035_0000
Patent Text Reader

Abstract

HAIR COLOR CLASSIFICATION USING CONSISTENCY REGULATION AND ANNOTATION CONFUSION MATRICES. Aspects of hair classification and networks for this purpose are provided, including aspects for training such networks. A classifier model is provided to mitigate the impact of human bias, where the modeling of the actual label distribution and annotator biases are separated by incorporating annotator confusion matrices into a reference model. To further improve the model's performance by leveraging unlabeled data, the model was trained using a consistency-based semi-supervised learning framework. Using only 1000 labeled data points, the final classifier model achieved classification accuracy 20% higher than that of a professional human annotator.The trained model can be used for a wide range of downstream tasks, including as a color classifier to train generative models for hair color translation. Figure for the abstract: none.
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: CLASSIFICATION OF HAIR COLOR USING CONSISTENCY REGULATION AND ANNOTATION CONFUSION MATRICES FIELD OF INVENTION

[0001] This disclosure relates to computer image processing and artificial intelligence, including hair color classification systems and methods, and more specifically, the improvement of hair color classification using consistency regularization matrices and annotation confusion. This disclosure also relates to hair color simulation guided by hair color classification. BACKGROUND

[0002] Due to the increasing popularity of online shopping, e-commerce companies are exploring ways to improve the customer experience. Virtual Try-On Technology (VTO) has been developed to help customers search for and virtually explore products, for example, to find the one that best suits them. In some examples of VTO, it can be useful to determine information, particularly from a user-supplied image. For example, in a hair-color-related VTO, a face image (e.g., a portrait) can be processed to determine, by classification, the hair color in the portrait. Figure 1 shows a schematic diagram of a representative processing flow 100 that illustrates a general model in the form of a classifier 102 that receives a color face image 104 (e.g., a portrait) and identifies a hair color 106, in accordance with the prior art.The output, along with the color image of the face, can then be used for a virtual hair dye trial.

[0003] Several challenges arise when developing a hair color classification model. For example, color labels collected from hair experts are subject to human bias, and labeled data are costly to obtain. It is desirable to have an improved hair classification model that can identify hair color.

[0004] A VTO pipeline can simulate hair color, for example, using a generative neural network. Improvements to the simulation using generative networks are desired. SUMMARY

[0005] Aspects of hair color classification, VTO pipelines, and corresponding networks are provided, including aspects concerning the training of such networks. Furthermore, according to one embodiment, a classifier model is provided to mitigate the impact of human bias, where the modeling of the distribution of real labels and annotator bias are separated by incorporating annotator confusion matrices into a reference model. To further improve the model's performance by leveraging unlabeled data, the model was trained using a consistency-based semi-supervised learning framework. Using only 1000 labeled data points, the final classifier model achieved a classification accuracy 20% higher than that of a professional human annotator.The trained model can be used for a wide range of downstream tasks, including as a color classifier to train generative models for hair color translation.

[0006] The following statements present various aspects and features disclosed in the embodiments hereof. These aspects and features, as well as others, will be readily understood by those skilled in the art, particularly the aspects of computer program products. It is also understood that aspects / features of computer devices or systems may have corresponding process aspects / features and vice versa.

[0007] Declaration 1: A computer device comprising a processor coupled to a storage device that stores instructions which, when executed by the processor, cause the computer device to: classify hair hue and hair reflectance in an input image using a neural network, the neural network comprising a coding skeleton coupled to a pair of classifiers comprising i) a hue classifier that outputs a hair hue value and ii) a reflectance classifier that outputs a hair reflectance value, each element of the pair of classifiers comprising a linear classifier, the neural network having been defined by training with a cross-entropy loss sum of hues and reflectance determined from i) the output hair hue values ​​and the hair reflectance values,and ii) target labels for hair color and hair reflectance, target labels prepared for training images from the respective expert votes by a plurality of experts.

[0008] Declaration 2: Computer device according to Declaration 1, in which the neural network is configured to classify hair shade and hair reflectance according to a plurality of respective classes to follow an industry standard for hair color.

[0009] Declaration 3: Computer device according to Declaration 1, in which the hair reflectance comprises a primary reflectance component and a secondary reflectance component.

[0010] Declaration 4: Computer device according to Declaration 1, wherein the target labels include soft annotations for at least some of the training images, the soft annotation being determined from an empirical allocation of expert votes on the respective classes.

[0011] Declaration 5: Computer device according to Declaration 1, wherein the neural network having been trained with respective annotator confusion matrices comprising a plurality (n) of hue annotator confusion matrices and a plurality (n) of reflectance annotator confusion matrices, wherein n is defined from a total number of experts providing expert votes, and wherein the output of the hue classifier is multiplied by a respective hue confusion matrix from among the plurality of hue confusion matrices and wherein the output of the reflectance classifier is multiplied by a respective reflectance confusion matrix from among the plurality of reflectance confusion matrices to predict the vote of a respective expert from among the plurality of experts and wherein the neural network and the respective annotator confusion matrices are trained together.

[0012] Declaration 6: Computer device according to Declaration 1, in which the neural network has been trained in accordance with a Student-Teacher framework.

[0013] Declaration 7: Computer device according to Declaration 6, in which the neural network includes a teacher network obtained from the Student-Teacher frame.

[0014] Declaration 8: Method comprising: receiving an input image and classifying the hair tint and hair reflectance in the input image using a neural network, the neural network comprising a coding skeleton coupled to a pair of classifiers comprising i) a tint classifier that outputs a hair tint value and ii) a reflectance classifier that outputs a hair reflectance value, each element of the pair of classifiers comprising a linear classifier, the neural network having been defined by learning with a cross-entropy loss sum of tint and reflectance determined from i) outputs of hair tint values ​​and hair reflectance values, and ii) target labels for hair tint and hair reflectance, the target labels being prepared for image learning from respective expert votes by a plurality of experts;providing hair color and hair reflectance.

[0015] Declaration 9: A method according to Declaration 8, wherein hair tint and hair reflectance are provided to train a generative model to simulate hair color.

[0016] Declaration 10: A method according to Declaration 9, wherein the neural network is configured to classify hair hue and hair reflectance according to a plurality of respective classes to follow an industry standard for hair color, and wherein hair reflectance comprises a primary reflectance component and a secondary reflectance component.

[0017] Declaration 11: A method according to Declaration 9, wherein the neural network having been trained with respective annotator confusion matrices comprising a plurality (n) of hue annotator confusion matrices and a plurality (n) of reflectance annotator confusion matrices, wherein n is defined from a total number of experts providing expert votes, and wherein the output of the hue classifier is multiplied by a respective hue confusion matrix from among the plurality of hue confusion matrices and wherein the output of the reflectance classifier is multiplied by a respective reflectance confusion matrix from among the plurality of reflectance confusion matrices to predict the vote of a respective expert from among the plurality of experts and wherein the neural network and the respective annotator confusion matrices are trained together.

[0018] Declaration 12.-Process according to Declaration 9, in which the neural network was trained in accordance with a Student-Teacher framework.

[0019] Declaration 13: Computer device according to Declaration 12, in which the neural network includes a teacher network obtained from the Student-Teacher frame.

[0020] Declaration 14: A method according to Declaration 9, wherein the target labels include soft annotations for at least some of the training images, the soft annotation being determined from an empirical allocation of expert votes on the respective classes.

[0021] Declaration 15: Method according to Declaration 14, wherein the input image is associated with at least one of the target labels and wherein step b is carried out to train the neural network using the input image.

[0022] Declaration 16: A method according to Declaration 15, comprising training the neural network with respective annotator confusion matrices comprising a plurality (n) of hue annotator confusion matrices and a plurality (n) of reflectance annotator confusion matrices, wherein n is defined from a total number of experts providing expert votes, and wherein the output of the hue classifier is multiplied by a confusion matrix of respective hue among the plurality of hue confusion matrices and wherein the output of the reflectance classifier is multiplied by a respective reflectance confusion matrix among the plurality of reflectance confusion matrices to predict the vote of a respective expert from the plurality of experts and wherein the neural network and the respective annotator confusion matrices are trained together.

[0023] Declaration 17: Computer software product comprising a non-transient storage device that stores computer-executable instructions which, when executed by a processor of a computer device, cause the computer device to: classify hair hue and hair reflectance in an input image using a neural network, the neural network comprising a coding skeleton coupled to a pair of classifiers comprising i) a hue classifier that outputs a hair hue value and ii) a reflectance classifier that outputs a hair reflectance value, each pair of classifiers comprising a linear classifier, the neural network having been defined by training with a cross-entropy loss sum of hue and reflectance determined from i) the outputs of the hair hue values ​​and the hair reflectance values,and ii) target labels for hair color and hair reflectance, the target labels being prepared for the training images from the respective expert votes by a plurality of experts.

[0024] Declaration 18: Computer software product according to Declaration 17, wherein the execution of instructions causes the computer to provide the hair shade and hair reflectance to train a generative model to simulate hair color.

[0025] Declaration 19: Computer software product according to Declaration 17, wherein the execution of instructions causes the computer device to classify hair shade and hair reflectance according to a plurality of respective classes to follow an industry standard for hair color and wherein hair reflectance comprises a primary reflectance component and a secondary reflectance component.

[0026] Declaration 20: Computer software product according to Declaration 17, wherein one or both of the following: execution of the instructions causes the computer device to provide the neural network for classification, wherein the network having been trained with respective annotator confusion matrices comprising a plurality (n) of hue annotator confusion matrices and a plurality (n) of reflectance annotator confusion matrices, wherein n is defined from a total number of experts providing the expert votes, and wherein the output of the hue classifier is multiplied by a respective hue confusion matrix from the plurality of hue confusion matrices and wherein the output of the reflectance classifier is multiplied by a respective reflectance confusion matrix from among the plurality of reflectance confusion matrices to predict the vote of a respective expert from among the plurality of experts; and wherein the neural network and the respective annotator confusion matrices are trained together; or the execution of the instructions causes the computing device to provide the neural network for classification in which the network has been trained according to a Student-Teacher framework. Brief description of the drawings

[0027] [Fig. 1] The [Fig. 1] is a schematic diagram showing a model of a classifier according to the prior art.

[0028] [Fig.2] The [Fig.2] is a schematic diagram showing a reference classifier model according to the prior art.

[0029] [Fig.3] The [Fig.3] is a schematic diagram showing the reference classifier model of the [Fig.2] as adapted with annotator confusion matrices according to one embodiment.

[0030] [Fig.4] The [Fig.4] is a diagram of a representative annotator confusion matrix.

[0031] [Fig.5] The [Fig.5] is a schematic diagram showing the reference classifier model of the [Fig.2] as adapted with a Student-Teacher framework in accordance with one embodiment.

[0032] [Fig.6] The [Fig.6] is a graph, according to one embodiment, showing a comprehensive comparison of the model prediction accuracy results between all models and an average human expert.

[0033] [Fig.7A] The [Fig.7A] is a graph showing, according to one embodiment, a comparison of the accuracy of different numbers of training selfie images.

[0034] [Fig.7B] The [Fig.7B] is a graph showing, according to another embodiment, a comparison of the accuracy of different numbers of annotators.

[0035] [Fig.8] The [Fig.8] is a schematic diagram of a computer device according to an embodiment.

[0036] [Fig.9A] The [Fig.9A] is a schematic diagram of a network system according to one embodiment.

[0037] [Fig.9B] The [Fig.9B] is a schematic diagram of a network system according to another embodiment.

[0038] [Fig.9C] The [Fig.9C] is a schematic diagram of a network system according to another embodiment. DETAILED DESCRIPTION

[0039] The systems, methods, and techniques presented herein aim to improve existing hair color testing technology by leveraging artificial intelligence (AI). Many current hair simulation efforts are primarily based on traditional computer vision engineering techniques, and sometimes the hair texture in the simulated images appears synthetic—for example, obviously simulated but poorly simulated. The applicant seeks to use AI to enhance the photorealism of simulated hair images while maintaining color accuracy. A primary objective is color accuracy, and a hair color classification model is described herein in one or more embodiments for identifying hair color in selfies.

[0040] According to one embodiment, hair color is defined based on L'Oréal's hair color definition for hair dye products and classifies hair color according to 1) shade and 2) reflectance, namely primary and secondary reflectance. Hair shades are divided into 10 classes, 1 being the lightest and 10 the darkest. In one embodiment, there are 9 reflectance classes for each of primary and secondary reflectance, where primary reflectance indicates the dominant reflectance in the hair, and secondary reflectance indicates the second dominant reflectance. Hair may have only primary reflectance or no reflectance at all. If the absence of reflectance is considered a separate reflectance class, there are 100 combinations of primary and secondary reflectances.Examples: no primary or secondary reflectance, golden primary reflectance without secondary reflectance, ashy primary reflectance without secondary reflectance, golden primary reflectance with ashy secondary reflectance and golden primary reflectance with mahogany secondary reflectance, etc.

[0041] Labeled datasets from hair coloring experts exist or can be obtained for hair color. However, such datasets, particularly their use, present challenges for building a hair color identification model. For example, hair colors are defined based on the judgment of professional hair experts rather than on scientific color definitions, such as the average RGB color. Consequently, the existing color definition is subject to human bias. Since the annotations were provided by hair experts, the labeled data are expensive to obtain and may result in a limited number of labeled examples for training and verification, etc.

[0042] It was desirable to develop a hair color classifier to produce color predictions at least as accurate as those of an average human hair expert. Annotations of 1000 selfies from 10 hair experts were received to define a dataset to be used to develop the model – the “expert dataset”. The model's performance was evaluated based on its adequacy with the majority of the hair experts' annotations or the discrepancy between the predictions and the experts' annotations.

[0043] The expert dataset was relatively small, especially since the data were obtained gradually as the experts performed their tasks at different times. Furthermore, the combined results illustrated the differences in expert opinion. 1000 labeled images represent a relatively small dataset given that there were over 100 hair colors. After gradually receiving labels from other hair experts, it was determined that the labels from different hair experts were often inconsistent. In some cases, there was no clear majority vote on the hair color, and all 10 experts thought the hair was different colors. Therefore, one of the objectives of the activities was to mitigate the impact of label noise.

[0044] To address these two challenges, a final model according to one embodiment includes two improvements on a standard classifier (prior art) (for example, from [Fig. 1]), by adding annotator confusion matrices and using a Mean Teacher semi-supervised learning framework. The addition of annotator confusion matrices aims to mitigate the problem of label noise by modeling the biases of individual experts, while the Mean Teacher framework is a semi-supervised learning technique that mitigates the problem of limited labeled data by using unlabeled data during training. Label noise

[0045] In the real world, data are often labeled by multiple annotators. Since biases in human annotators are inevitable, these labels often contain noise, and annotators may disagree with each other. Studies show that the rate of corrupted labels in real-world datasets ranged from 8% to 39%

[19]

[16] [9][8]. A classic way to handle label noise is to aggregate annotations using majority voting. However, the majority voting process has some limitations; for example, it only uses the class with the most votes as the aggregated label and ignores the distribution of votes, and there may be scenarios where there is no clear majority vote. Another way to aggregate annotations is to use soft labels in which the values ​​of the soft labels are based on the empirical distribution of votes.

[0046] Model-based approaches have also been explored to handle label noise, including the use of a translation matrix to measure noise and the reweighting of labels based on confidence in them. The transition matrix is ​​widely used to model the relationship between the actual and observed distributions [5]

[13] . However, these methods often fail to account for the fact that different annotators may have different biases. On the other hand, loss reweighting has also been investigated to improve model robustness by assigning more weight to labels or models offering greater confidence

[10] . In particular, Weighted Doctor Net [4] proposed modeling individual annotators separately and then combining the predictions of the individual models using weights resulting from training.This method takes into account the overall skill level of annotators across all data classes and ignores the scenario where annotators may have different skill levels when labeling different classes. Tanno et al.

[17] were inspired by this and combined the use of a transition matrix with a new weighting of losses and proposed to model annotator labels and real data separately using annotator-specific confusion matrices. Annotator confusion matrix

[0047] The annotator confusion model

[17] is a probabilistic model based on the assumptions that the annotators are independent and the noise in the labels is independent of the input images. Based on these two assumptions, the common probability of observing n labels for an image from n annotators can be expressed as follows: [00481

[0049] where y(n) is the annotation of the nth annotator and x is the image, y is the actual color and a(1) y"',ycst a value in the confusion matrix that indicates the probability that the annotator indicates an image as being of class y given that the actual class is y. During training, the annotator-specific confusion matrices are trained jointly with the classifier, in which the classifier learns to predict the actual distribution of labels and each confusion matrix learns the observed distribution of labels as a function of the actual labels for a particular annotator. Semi-supervised learning

[0050] Labeled data are often expensive to obtain, whereas unlabeled data are generally abundant and provide additional information about data distribution. To take advantage of the information contained in unlabeled data, semi-supervised learning (SSL) [3]

[21] is a research area that explores ways to train models using both labeled and unlabeled data. Typical semi-supervised learning techniques include consistency regularization and pseudo-labeling [7].

[0051] Consistency regularization was first introduced by the F model

[14] and suggests that the model should provide similar data points with consistent outputs. To encourage models to do this, the F model applies a loss of consistency between data points with and without noise. Following this, the Time Assembly [6] proposes to improve the model by maintaining an exponential moving average (EMA) prediction of the training data. The Student-Teacher model

[18] is an extension of the F model and the Time Assembly, and suggests maintaining an exponential moving average of the model weights through a student-teacher framework. The student model is trained on the basis of supervised classification loss and unsupervised consistency loss with the teacher model. The consistency loss is applied to minimize the differences between the student and teacher model outputs.The teacher model weightings are updated based on the student model's EMA. The Student-Teacher framework has achieved excellent results for semi-supervised learning image classification benchmarks.

[0052] Recent state-of-the-art SSL models often combine several SSL techniques, notably the combination of consistency regularization with pseudolabeling [2][1]

[15] . FixMatch

[15] is based on the idea that data augmentation has a significant impact on consistency regularization

[20] , and proposes training the classifier so that it processes consistent output for various augmentations of the same output using pseudolabels. This method produces excellent results in image classification tasks. However, because it relies on strong augmentations that change the color of the images, it is not applicable for color classification. Model architecture Reference model

[0053] Figure 2 shows a schematic diagram of a processing flow 200 according to an embodiment in which a reference model 202 receives an image 204 as input and produces output predictions 206 comprising a hair color 206A and a reflectance 206B. In one embodiment, the model of Reference 202 uses ResNetl8 202A as its skeleton. The output of skeleton 202A is fed to two linear classifiers 202B and 202C, which separately predict hue 206A and reflectance 206B. In one embodiment, each linear classifier consists of two fully connected linear layers with a ReLu activation layer between them. In one embodiment, the flow 200 represents a training flow at one training time. The components 202 are retained for use during the inference time.

[0054] In one embodiment, the reference model 202 was trained using a sum of cross-entropy losses of hue and reflectance (e.g., 208) between the model outputs 206A / 206B and the target colors (e.g., soft annotations 210) prepared from the expert labels.

[0055] "Soft" and "hard" labels

[0056] During the training of the reference model of [Fig.2], the "soft" annotations Both "hard" and "soft" annotations are tested. Hard annotations for hue and reflectance are calculated separately based on the class with the most votes. Soft annotations are calculated based on the empirical distribution of labels for all annotations. For example, if nine experts label the image as class 1 and one expert labels it as class 2, the soft annotation will be 0.9 for class 1, 0.1 for class 2, and 0 for all other classes. Regardless of whether the annotations are soft or hard, the reference models are trained based on cross-entropy loss (e.g., 208).

[0057] Classifier with annotator confusion matrices

[0058] Figure 3 is a schematic diagram of a processing flow 300 showing a classifier 302 according to an embodiment of Figure 2. The classifier 302 is adapted by learning using annotator confusion matrices 304 as described in more detail. The processing flow 300 is useful at a learning stage, and the components 302 are useful at an inference stage.

[0059] In one embodiment, the components 302A to 302C of the classifier 302 are structured similarly to the components 202A to 202C of the classifier 202 but are adapted, for example, by training with the confusion matrices of annotators 304 as described. The output hue 306A and the reflectance 306B are also similar to the outputs 206A and 206B, but will depend on the trained classifier 302. The outputs of the respective classifiers for the same input image may vary in value.

[0060] According to one embodiment, the annotator confusion matrices 304 comprise a plurality (n) of hue annotator confusion matrices 304A and a plurality (n) of reflectance annotator confusion matrices 304B. Here, "n" refers to the number of annotators.

[0061] Figure 4 is an illustration of a representative hue annotator confusion matrix 400, according to one embodiment. The rows represent actual classes ("Actual Values") and the columns represent observed annotations ("Expert Annotations").

[0062] In one embodiment, the classifier 302 with annotator confusion matrices 304 is built on top of the reference model 202, having annotator-specific confusion matrices 304 after the respective classification layers (202B / 202C) of the reference model 202 to predict the observed labels 306A / 206B from the annotators. Assuming there are "n" annotators, n 10*10 confusion matrices are added after the hue classifier and n 100*100 confusion matrices are added after the reflectance classifier. The outputs 306A / 306B of the hue and reflectance classifiers 302B / 302C are multiplied by the respective confusion matrices 304A / 304B to predict the individual annotator annotations. For example, the outputs of the hue classifier 302B are multiplied by the first hue confusion matrix (e.g., 307) to predict the hue annotation for the first annotator.

[0063] During training, for each training data point, the cross-entropy loss sums of n hues and n reflectances (e.g., 308) are calculated to update the model. Since an annotator has only one hue label and one reflectance label for each training data point, each cross-entropy loss was calculated using hard annotations (e.g., 310) instead of soft annotations (e.g., 210). The confusion matrices are first randomly assigned values ​​such that diagonal values ​​dominate and the matrices are approximately identical. The matrices are trained with the rest of the model based on the classification loss 308, and a regularization term is added to encourage the trained matrices to converge to the actual confusion matrices (e.g., as in the example in [Fig. 4]). Student-Teacher Framework

[0064] Figure 5 is a schematic diagram of a processing flow 500 showing the reference classifier model 202 of Figure 2 adapted to a student-teacher framework for training. The flow 500 comprises a student model 502A and a teacher model 502B, according to one embodiment. Models 502A and 502B are configured as instances of model 202, so details are not shown in Figure 5. The student model 502A produces an output comprising the hue output 506A and the reflectance output 506B, and the teacher model 502B produces an output comprising the hue output 506C and the reflectance output 506D, respectively.

[0065] The student model 502A and the teacher model 502B are first pre-trained using the labeled data. During the semi-supervised learning process, two augmented versions of the same image (e.g., 504A and 504B) are fed separately to each of the 502A / 502B models. The student model 502B is trained using the cross-entropy loss (supervised) 508 and the consistency loss (unsupervised) 512. The cross-entropy loss (classification loss) is calculated on the basis of the labeled data (e.g., 510), and the consistency loss is calculated as corresponding to the root mean square error of the softmax outputs (506A, 506B and 506C, 506D) between the student model 502B and the teacher model 502A. The loss of consistency 512 based on the augmented images 504A, 504B forces the pupil model 502A to provide consistent outputs at similar data points.After the update of the student model 502A, the teacher model 502B is updated based on the exponential moving average of the weights of the student model 502A. According to one embodiment, during inference, only the teacher model 502B is retained and used for evaluation. Experiences Dataset

[0066] The selfie dataset used during the experiments contained 1000 labeled images and 2500 unlabeled images. For the labeled data, the hair in the selfie was annotated by 10 professional hair experts, and each label contained a hue, a primary reflectance, and a secondary reflectance. In one embodiment, the hue values ​​range from 1 to 10, where 1 corresponded to the darkest hue, and the primary and secondary reflectance values ​​range from 0 to 8. The hair color may also have no secondary or primary reflectance.

[0067] The labels of different annotators may not match. However, for most of the labeled images, there were at least two annotations that coincided. Out of the 1000 images, only 40 had no concordance. When considering the hue, primary, and secondary reflectance values ​​individually, at least two annotators always agreed. To assess the quality of the annotators, Table 1 lists the accuracy and the average absolute hue difference between the labels of individual annotators and the majority vote. Overall concordance on color was about 30%; experts disagreed more on reflectance values ​​(49%) than on hues (64%). In particular, experts agreed on secondary reflectance values ​​only 42% of the time.

[0068] [Table 1]

[0069] Table 1

[0070] Configuring Experiences Annotator Avg. 1 2 3 4 5 6 7 8 9 10 Overall Accuracy (%) 30 33 20 31 21 41 39 34 30 28 21 Hue Accuracy (%) 64 81 55 68 53 70 73 66 59 61 55 Reflectance Accuracy (%) 49 45 52 18 18 59 65 67 63 58 49 Primary Reflectance Accuracy (%) 48 56 47 26 29 59 56 56 52 57 41 Secondary Reflectance Accuracy (%) 42 33 31 38 32 54 55 49 49 51 31

[0071]

[0072]

[0073]

[0074] The reference model 202 and the classifier 302 with the confusion matrices 304 were trained using the stochastic gradient descent (SGD) optimizer. The learning rate was constant for the first 20 periods and then decreased linearly to zero over the following 65 periods. Simple data augmentations, including horizontal and vertical flips, rotations, and random translations, were used during training. The batch size was set at 128. The teacher model of the 502B Student-Teacher framework was trained using the SGD optimizer, and the learning rate was reset to zero using cosine cancellation

[12] . For data augmentation, random data augmentations, including horizontal and vertical flipping, rotation, translation, and shear, were used. The batch size was set at 256, with 128 labeled and 128 unlabeled data points. Evaluation metrics In one embodiment, model performance was evaluated based on the acceptance of the model by the majority of experts and the magnitude of the discrepancy between predictions and annotations. Three main metrics were used for evaluation: accuracy, hue ±1, and mean absolute difference in hue. For hue, reflectance, and overall color, accuracy was measured in terms of the model prediction corresponding to the color that received the most votes from the ten annotators.Since the hues are ordinal data, two additional measures were used: the hue at ±1 measures the probability that the model's predictions fall within a hue difference from the hue that received the most votes, and the mean absolute difference measures the average distance between the predicted hues and the mean hue annotations for each image. The results. of the experiment were based on 5-time cross-validation and, at each run, the metrics of the best period were reported. Results Soft annotations / hard annotations

[0075] Table 2 compares the prediction performance between an average human expert, the reference model trained with hard annotations, and the reference model trained with soft annotations, according to one embodiment. In cases of disagreement between the experts, training the model with soft annotations provided more insight into the distribution of votes and led to improved predictions. The prediction accuracies of the model trained with soft annotations were 9% higher than those of the model trained with majority votes, with a 7% increase in hue accuracy and a 6% increase in reflectance accuracy. Another observation is that the model trained with majority votes performed slightly worse than an average human expert, while the model trained with soft annotations performed better than the average human expert.

[0076] [Table 2]

[0077] Table 2 Classifier with confusion matrices Human Expert Majority Annotations (Dou) Overall Accuracy (%) 30±7 29+4 38+4 Hue Accuracy (%) 64±9 62+6 69+6 Reflectance Accuracy (%) 49±18 41+5 47+4 Primary Reflectance Accuracy 48±12 60+6 66+6

[0079] In one embodiment, stepwise experiments were conducted to examine the impacts of adding confusion matrices to the model and training the model using the Student-Teacher framework. In short, either model outperformed the average human expert. As shown in Table 3, adding confusion matrices to the model improved overall color accuracy (2%). It is worth noting that the reference model trained with soft annotations is a special case of the model with confusion matrices where all the confusion matrices were identity matrices. The improvement in color classifications is primarily due to improved reflectance prediction (5%) and prediction accuracy. The shades are similar to the reference model. It is understood that this results from more disagreements among hair experts regarding reflectance labels.

[0080] [Table 3]

[0081] [Tables3] Human Expert Reference Confusion Average Overall Accuracy (%) 30+7 38+4 40+6 50+5 Hue Accuracy (%) 64+9 69+6 67+3 73+4 Reflectance Accuracy (%) 49+18 47+4 53+6 58+5 Primary Reflectance Accuracy (%) 48+12 66+6 68+6 73+2 Secondary Reflectance Accuracy 42+10 54+3 60+5 63+3

[0082] Although the addition of annotator confusion matrices did not improve hue accuracy, a more detailed analysis showed that its prediction deviated less from the experts' annotations. Table 4 shows a comparison of hue prediction performance between human annotators, the reference model with soft annotations, the classifier with annotator confusion matrices, and the Student-Teacher framework. As shown in Table 4, after the addition of the confusion matrices, the percentage of hues within ±1 of the hue receiving the most votes was 1% higher, and the mean absolute differences were 0.04 lower. This revealed that hue prediction was also improved.

[0083] Reference Model Student-Teacher Confusion Matrix

[0084] [Table 4]

[0085] Table 4 Student-Teacher Framework Reference Model Confusion Matrix Student-Teacher Tint at +1 (%) 97+0 98+0 98+2 Abs. Diff. Mean Tint 0.44+0.07 0.40+0.07 0.32+0.07

[0087] After the integration of the Student-Teacher framework, the overall performance of the model was significantly improved. Tables 3 and 4 show that the framework improved the prediction accuracy by 10% compared to the confusion matrix model, and that it also reduced the difference between the model's predictions. and expert annotation. The mean absolute difference was reduced by 0.08. Furthermore, when comparing the classifier based on the Student-Teacher framework with an average human expert, there was a 20% improvement in overall color accuracy, with improvements of 9% and 11% in hue and reflectance accuracy, respectively. Figure 6 is a 600-dot graph, according to one embodiment, showing a comprehensive comparison of model prediction accuracy results between all models and an average human expert. The 600-dot graph demonstrates a significant improvement in accuracy between that of an average human expert and the Student-Teacher framework model. Ablation study

[0088] To facilitate better data collection planning, an experiment was conducted to investigate the impact of having additional labeled selfies and annotators. In subsequent experiments, the reference model with soft annotations was trained to assess the impact of having more images and more labels per image. When examining the impact of additional labeled images, the number of labeled data points varied from 100 to 900. When examining the impact of additional labels per image, the labels of a subset of experts were used to train the model.

[0089] The model's performance improved when it had more selfies or more annotators than expected. However, as shown in Figures 7A and 7B, the model improved even more when it had more labeled images. Figure 7A is a 700 graph showing a comparison of the accuracy of different numbers of training selfies, and Figure 7B is a 710 graph showing a comparison of the accuracy of different numbers of annotators.

[0090] It has been observed that a steady increase in accuracy is achieved when increasing the number of training images from 100 to 900, with overall, hue, and reflectance accuracies increasing in parallel. No downward trend in improvement was observed, even after using 900 images. In contrast, minimal improvement in overall accuracy was observed between 5 and 10 annotators; in particular, minimal improvement in reflectance accuracy was observed after using 8 annotators. This demonstrated that having additional labeled images was more effective in improving model performance. Hair simulation application

[0091] Figure 8 is a schematic diagram of a computer device 800 according to one embodiment. In the embodiment, a color classifier according to one embodiment of the present text is integrated into an application of VTO ​​to provide a hair color simulation. The computing device 800 proposes a user computing device such as a smartphone, tablet, laptop, or other computing device for use by a user such as a consumer of hair dye products or a salesperson assisting such a user. The device may include a component of a larger form factor such as a kiosk for placement in a retail environment, for example. Device 800 is not exhaustive and is simplified for brevity.

[0092] The device 800 includes a storage device 802, a processing unit 806, a camera 808, a microphone / speaker 810, a display screen 812, and a communication subsystem 814. In one embodiment, the storage device 802 includes a memory device, for example, one or more types of memory such as RAM, ROM, etc. The storage device 802 may include a long-term storage device such as an integrated circuit disk (SSD) or another type of drive for storing persistent data, etc. The storage device 802 stores computer-readable instructions for execution by the processing unit, such that once executed, the instructions cause the computer device to perform operations such as one or more processes.The 806 processing unit comprises one or more central processing units (e.g., CPUs) and / or graphics processing units (e.g., GPUs) having one or more processors / microprocessors, controllers / microcontrollers, etc. Other types of processors may be used. GPUs can be particularly useful for accelerating graphics processing tasks and / or AI processing tasks (e.g., training and / or inference).

[0093] The camera 808 can be used to take selfies. The microphone and speaker 810 are usually separate devices, but are labeled together here for convenience and represent some of the input (I), output (O), or I / O devices that may be available. Other devices may include a light, a buzzer, a vibrator, a button, a keyboard, a pointing device (e.g., mouse or touchpad, etc.), etc.

[0094] The display screen 812 presents images such as components of a graphical user interface, camera images, etc. In one embodiment, the display screen is a touch screen device, a type of I / O device, configured to receive gesture inputs (e.g. swipe, tap, etc.) which interact with the screen region(s) and in association with the user interface components (e.g., controls) presented by an application executed by the processing unit 806.

[0095] The communication subsystem 814, in one embodiment, is configured to manage communications between device components and / or between the computing device and external devices such as a remotely located computing device (e.g., web service, cellular network component, printer, etc.). A subcomponent thereof in one embodiment is an antenna for wireless communication. A subcomponent thereof in one embodiment is a wired communication interface (e.g., Ethernet, USB-A, USB-C, Thunderbolt (TM of Intel Corporation), etc.) to be coupled with a suitable cable for wired communication.

[0096] The storage device 804 stores the components of a VTO application and the corresponding data (e.g., 820). Representative components are shown. The VTO application 820 includes the user interface component 822 (e.g., screens, instructions, icons, commands, etc.). The user interface provides outputs to a user and receives inputs such as inputs for the application workflow, user selections of color choices, etc. A color classifier 824 is provided and includes one of the classifiers 202, 302, and 502 as described previously for determining hair color information (e.g., hue and reflectance). A VTO pipeline 826 is provided for simulating hair color in association with an input image (e.g., on or within the latter). A color recommendation engine 828 and an associated data store 830 are provided.In one embodiment, the data store 830 stores hair color data and representative images thereof, such as colored strand or colored hair images, hair reflectance images, etc. The user interface may present one or more choices that a user can select via user input, for example, sorting or filtering examples and presenting them in groups / pages, etc. In one embodiment, the color recommendation engine 828 may include an interface to a chatbot or live agent to discuss a recommendation.

[0097] In one embodiment, recommendations can be made, for example, based on recommendation factors such as a user's personal information, including current hair color and age; hair color trend data; the availability of a product and / or service locally for the user (e.g., within a certain radius); cost information, etc. In one embodiment, rules or another method of recommendation can be used to determine the recommendation (or recommendations) to be presented to a user.

[0098] In one embodiment, a user can provide an input image with hair (e.g., 832) similar to the input image 204 for classification. The hue and reflectance data 834 are generated by the classifier 824. The hue and reflectance data can be made available to the color recommendation engine 828 for processing (e.g., using color matching rules, color complementation rules, etc.) to select one or more colors from the color data store 830 to recommend to a user via the user interface component 822. The recommendation can be displayed on the screen 812.The user can invoke the VTO engine (via an input to a command) to have the engine simulate color and / or color and reflectance (e.g., target hair) using the 832 image and target hair data as input to produce an output image with simulated hair 836.

[0099] The output image with simulated hair 836 can be presented via the user interface component 822 on the display screen 812. A before-and-after display can be provided for comparison. A comparison between two or more simulated colors (for example, two different output images) can be displayed for comparison. Products or services, or both, can be purchased via the interface 838.

[0100] Such an interface 838 can direct the user (e.g., the computer device) to a web-based e-commerce service (e.g., a website (not shown)) to carry out the purchase, reservation or other.

[0101] In one embodiment, the VTO 826 pipeline includes a generative neural network (e.g., a model) configured to generate a simulated hair image using the input image 832 and the target hair data as input. See Figure 9 discussed below.

[0102] Other components stored in the storage device 804 include an operating system 840, a browser 842 (for example, for browsing web pages), an email and / or messaging application 844 (for example: SMS or other type) and a social networking application 846. The output image with simulated hair 836 can be shared (for example, communicated) via the applications 844 and / or 846, for example.

[0103] In another embodiment of using the color classifier not shown in [Fig. 8], the color classifier 824 processes an output image from a generative model configured to generate a simulated hair image using an input image having target hair and color data as input. The classifier determines the color and / or reflectance data. such as for comparison with target hair data to confirm the accuracy of the generative model.

[0104] Classifier-guided training of a color-finement neural network

[0105] In one embodiment, a classifier can guide the training of a color-finement neural network that is configured to produce rendered images with a target hair color. It is difficult to assess the accuracy of hair rendering based solely on RGB values ​​or simple metrics. In one embodiment, there is a feedback loop in a rendering network system to use a color classifier trained according to an embodiment of the present to guide the rendering network. Figure 9A is a schematic diagram of a 900A network system according to an embodiment providing a training system. Figure 9B is a schematic diagram showing further details of a 900B training network system linked to the 900A system, according to an embodiment.The process, the computer program product, and other aspects can be readily understood by a person skilled in the art from an understanding of the training and inference aspects shown and described.

[0106] The 900A network system displays the input image 104 (e.g., image 1) and a target hair color 902 provided to a hair analysis engine 904 to determine hair information for hair pixels (e.g., from hair segmentation), for example, hair histogram color data for image 1. In one embodiment, the image color data 1, target hair 902 (e.g., as a hair strand image), and hair histogram include RGB data (e.g., 256*3) commonly used for images. In one embodiment, the hair analyzer engine 904 understands or communicates with a hair analyzer such as a hair classifier network to determine the hair pixels in the input image 104.In one embodiment, the image I and color information (e.g., histogram data) for the target hair color 902 and the hair in image I (e.g., in histogram form (not shown as such in [Fig. 9A])) are provided to a color-finement neural network 906 to produce a rendered image 908. In one embodiment, as shown in [Fig. 9B], the color-finement neural network 906 includes a color-mapping network and a generative network configured to modify features of an image it processes, namely hair color. The generative network in one embodiment includes a model (e.g., generator (G)) that is defined by training guided by the hue and reflectance classifier 910.

[0107] Referring again to [Fig. 9A], the rendered image 908 is provided to the color classifier 910 defined (e.g., trained) according to one embodiment of the present to classify the hue and reflectance properties in order to produce the color prediction 912 for the image 908. The loss 914 represents a loss determined from the color prediction 912 and the target hair color 902 as used to train the color-fine neural network 906 as described in more detail with [Fig. 9B]. In one embodiment, the system 900A is configured with one or more computing devices (not shown) to provide the computing components (e.g., 904, 906, 910, etc.) and to store data (e.g., 104, 902, 908, 912, 914, etc.).A display device (not shown) may be included to display image 104, 908 and any of the other data, including the output (not shown) of the hair analyzer engine 904 supplied to the color refinement neural network 906.

[0108] In one embodiment, the network system components 900A and 900B can be configured as an inference time system (e.g., following training) to provide a virtual trial experience for simulating a hair color applied to an image. For example, components 904 and 906 are useful for defining a VTO pipeline (e.g., 920) to simulate hair color, processing an input image to simulate a target hair color and producing a simulated image. In one embodiment, the color refinement neural network 906 comprises a trained-defined neural network as guided by the color classifier 910 as described in more detail below.

[0109] According to Figures 9A and 9B, a 922 color mapping network is trained to predict correct rendering parameters based on the classification results of the color classifier (910). As shown in [Fig.9B], the pipeline uses the color mapping network 922 and a generative adversarial network (GAN) 924 with a generator (G)924A to simulate capillary rendering 908 (i.e. (G})).

[0110] A hair segmentation model (for example, as a component of the hair analyzer engine 904, and not shown in [Fig. 9B]) is used to extract hair masks and the engine 904 provides RGB histogram data of hair pixels (Hl) in (i.e. from) image 1 104. The color mapping network 922 includes an encoder 926 which encodes image features (F) from image 1. The encoder 926 encodes features including, but not limited to, the color and lighting of the hair representation and, in the generator, it is used to regenerate the same hair rendering. [YES]

[0112]

[0113]

[0114]

[0115]

[0116]

[0117]

[0118]

[0119]

[0120]

[0121] A concatenator 928 of the color mapping network 922 receives inputs: i) RGB histogram of the sample (i.e. target hair color) (H r) 902, ii) source (image 7) RGB histogram of the hair (777 ) 930 and iii) image features (F) of the encoder 926. The three inputs are concatenated and passed through a block of fully connected layers 932 for output color mapping (Mp). The output color map Mp and the source RGB hair histogram H1 930 are concatenated and used as input and passed through another block of fully connected layers (FC Block 2 (934)) to obtain the output RGB histogram (770) to be used by the generator (G)924A to produce its output from image 7. Thus, the rendered output image G i 908 is produced by replacing the hair pixels with the color map. Internal color maps (Mc) and RGB histograms (77G) are used as real data labels to train the 922 color mapping network in a supervised manner. In one embodiment, the following losses are used to train the 922 color mapping network: PcrtC — L Color Mapping + P RGB Histogram L2 loss between the actual data color mapping (M c) and the network-predicted color mapping: Color lithography — G " ^p) L1 loss between the histograms of the real data RGB histogram (77c) and the output RGB histogram (770): ^RGB Histogram ~ E,»c- HJ with the trained color mapping network 922, the GAN 924 can be trained, in particular as follows, according to one embodiment. The RGB output histogram (Ho) from the 922 color mapping network is used as a condition with the source image 104 as inputs to the generator (G) 924A of the 924 GAN. In one embodiment, the 924 GAN is defined as a StarGAN (in accordance with the teaching of Choi, Y., et al, "StarGAN: Unified Generative Adversarial Networks for Multi-Domain translation-to-Image Translation", IEEE Conference on Pattern Recognition and Computer Vision (CVPR), 2018, pp. 8789-8797), so that G7 = G(I, Ho). It is noted that the previously trained 922 color mapping network is frozen during GAN training. That is, the 922 network is not trained or co-trained with the 924 GAN training. The discriminating part (discriminator (D) 924B) of GAN 924 uses two classifiers, one (936) that classifies whether the image is false or real ((G7) or 7); and the other (910) is the pre-trained classifier which classifies the two parts of the color: hue and reflectance.

[0122] The pre-trained classifier guides and enhances the generative capabilities of the GAN network and can be used independently with any GAN architecture for guidance.

[0123] The losses used for the drive of the generator parts 924A and discriminator parts 924B are the same as those defined in the StarGAN document referenced above.

[0124] In one embodiment, the color mapping + GAN is labeled as a color refinement network, because it uses the instructions of the pre-trained hair classification model and the losses indicated above so that the generative model refines its outputs.

[0125] As shown in the simplified block diagram in [Fig. 9C] illustrating a VTO application, in one embodiment, at the inference time (for example, when training is complete and the refinement network is provided to generate output images, in particular from user input), the discriminator component 924B with its two classifiers 910 and Conclusion

[0126] In this work, techniques, etc., are provided to address the problem of limited availability of labeled data and label noise caused by human bias, for classifying the color of attributes in images, through a combination of a semi-supervised learning framework and annotator confusion matrices. The final model achieved an accuracy 20% higher than that of an average human expert.

[0127] The experiments described herein have shown that the confusion matrix is ​​useful for improving color prediction, particularly in scenarios where there is more disagreement among experts. The approach primarily considers label noise caused by individual expert biases and assumes that the biases depend only on the actual class (e.g., without considering other factors such as lighting and background contrast). Due to the noisy nature of subjective annotation, in one embodiment, soft annotations are used but not majority votes, as the training labels and confusion matrices are adapted to represent different expert ideas.

[0128] The experiments described herein have also shown that the Student-Teacher framework helps to significantly improve color prediction. The effectiveness of the current framework depends on the teacher model providing good labels for the student model to learn.

[0129] Unlike normal RGB value prediction, the classification model is designed to classify hair according to a color standard specifically defined for it. In one embodiment, such a standard consists of three digits, the first digit representing the natural hair shades and the second and third digits representing the primary and secondary reflectance of the hair. Such a color prediction adapts to different lighting conditions and therefore accurately predicts dark hair in bright light and light hair in dim light. The color representation is independent of light. The color digit prediction can be directly related to how hair products are color-coded (e.g., using the same color standard), which cannot generally be expressed in RGB or simple text information.

[0130] The model structure uses one branch for the natural tint (first number) and one branch for both reflectance determinations together in order to obtain better performance.

[0131] The following aspects and characteristics cited in the numbered statements (Statement-Sim 1, Statement-Sim 2... Statement-Sim 17) will become clear from the disclosure herein, in particular:

[0132] Declaration-Sim 1: Computing device comprising a processor coupled to a storage device that stores instructions executable by the processor to cause the computing device to: process an input image and a target hair color using a virtual testing pipeline (VTO) to produce a VTO experiment that simulates the target hair color with the input image to produce an output image;in which the VTO pipeline includes a color refinement neural network to generate the output image combining the target hair color and the input image, the color refinement neural network comprising a generative neural network trained under the supervision of a hair classification network comprising a hue classifier that outputs a hair hue value and a reflectance classifier that outputs a hair reflectance value to determine training loss information from output images to train the generative neural network.

[0133] Declaration-Sim 2: Computing device according to Declaration-Sim 1, wherein instructions are executable by the processor to cause the computing device to provide an interface with one or both of: i) a color recommendation engine to recommend a target hair color and ii) an e-commerce service with which to purchase a product or service or both.

[0134] Declaration-Sim 3: Computing device according to Declaration-Sim 1, wherein the VTO pipeline includes a color mapping network trained to produce a color map for input into the generative neural network in order to produce the output image for the VTO experiment, the color map produced from the image features, image hair color data determined from the image and the target hair color.

[0135] Declaration-Sim 4: Computing device according to Declaration-Sim 3, wherein the color mapping network includes an encoder for determining features from the input image; a concatenator for combining features, target hair color and image hair color data for processing by a first fully connected block, and a second block for processing an intermediate map from the first block with the image hair color data to produce the color map for input into the generative neural network.

[0136] Declaration-Sim 5: Computer device according to Declaration-Sim 4, wherein the VTO pipeline includes a hair analysis engine for determining image hair color data from the input image.

[0137] Declaration-Sim 6. 'Computer device according to Declaration-Sim 1, in which The input image and target hair color are defined using RGB values, such that the generative neural network is trained to produce an output image using particular RGB values ​​as guided during training by the hair classification network using hue values ​​and hair reflectance values.

[0138] Declaration-Sim 7: Computer device according to Declaration-Sim 6, in which the hair classification network is defined and trained according to an industry standard for hair color classification comprising a respective plurality of classes for hair tint and hair reflectance.

[0139] Declaration-Sim ^.'Computer device according to Declaration-Sim 1, wherein the hair classification network includes a coding skeleton and each of the hair hue classifier and reflectance classifier includes a respective linear classifier.

[0140] Declaration-Sim 9: Computer device according to Declaration-Sim 8, wherein the hair classification network has been defined by training with a cross-entropy loss sum of hue and reflectance determined from i) outputs of hair hue values ​​and hair reflectance values, and ii) target labels for hair hue and hair reflectance, the target labels being prepared for the training images of the hair classification network from respective expert votes by a plurality of experts.

[0141] Declaration-Sim 70,-Computer device according to Declaration-Sim 9, wherein the target labels include soft annotations for at least some of the training images, the soft annotation being determined from an empirical allocation of expert votes on the respective classes.

[0142] Declaration-Sim 11: Computer device according to Declaration-Sim 9, wherein one of: the hair classification network having been trained with respective annotator confusion matrices comprising a plurality (n) of hue annotator confusion matrices and a plurality (n) of reflectance annotator confusion matrices, wherein (n) is defined from a total number of experts providing expert votes,and wherein the output of the hue classifier is multiplied by a respective hue confusion matrix from among the plurality of hue confusion matrices, and wherein the output of the reflectance classifier is multiplied by a respective reflectance confusion matrix from among the plurality of reflectance confusion matrices to predict the vote of a respective expert from among the plurality of experts, and wherein the neural network and the respective annotator confusion matrices are trained together; or the hair classification network was trained according to a Student-Teacher framework.

[0143] Declaration-Sim 12: Method of setting up a neural network that classifies hair hue and hair reflectance on an input image, the method comprising: providing the neural network, the neural network comprising a coding skeleton coupled to i) a hue classifier that outputs a hair hue value and ii) a reflectance classifier that outputs a hair reflectance value, each classifier comprising a linear classifier; providing a plurality of training images associated with respective training labels for each of the hair hue and hair reflectance, the target labels being prepared from respective expert votes by a plurality of experts;and training the neural network using the training images, the training being carried out in accordance with a sum of cross-entropy losses of hue and reflectance determined from i) the classifier outputs of the hair hue values ​​and the hair reflectance values, and ii) the target labels.

[0144] Declaration-Sim 13: The Method according to Declaration-Sim 12, wherein the training includes training the neural network using a Student-Teacher framework.

[0145] Declaration-Sim 14: A method according to Declaration-Sim 12, wherein the training comprises training the neural network using respective annotator confusion matrices comprising a plurality (n) of hue annotator confusion matrices and a plurality (n) of confusion matrices of reflectance annotators, wherein n is defined from a total number of experts providing expert votes, and wherein the output of the hue classifier is multiplied by a respective hue confusion matrix from among the plurality of hue confusion matrices and in which the output of the reflectance classifier is multiplied by a respective reflectance confusion matrix from among the plurality of reflectance confusion matrices to predict the vote of a respective expert from among the plurality of experts; and wherein the neural network and the respective annotator confusion matrices are trained together.

[0146] Declaration-Sim 15: Method comprising: providing a generative neural network that generates an output image combining a target hair color with an input image; providing a hair classification network as a component of a discriminator of the generative neural network, the hair classification network comprising i) a hue classifier that determines a hue of the output hair and ii) a reflectance classifier that determines a reflectance of the output hair; determining training loss information using the target hair color and the output hair hue, and the output hair reflectance; and training the generative neural network under the direction of the hair classification network using the training loss information.

[0147] Declaration-Sim 16: Method according to Declaration-Sim 15, wherein the target hair color, input image and output image are defined using RGB type data and the output hair tint and output hair reflectance are defined in accordance with an industry standard for hair color classification comprising a respective plurality of classes for hair tint and hair reflectance.

[0148] Declaration-Sim 7 7.-Procedure according to Declaration-Sim 15 comprising, before Generative neural network training, pre-training of a color mapping network configured to process the input image and image hair color data from hair pixels of the input image and the target hair color to provide a color map to the generative neural network to generate the output image from the input image.

[0149] A practical implementation may include all or part of the features described herein. These aspects, features, and various combinations, as well as others, may be expressed in the form of processes, devices, systems, means of performing functions, program products, and other ways, combining the features described herein. A number of embodiments have been described. Nevertheless, it is understood that various modifications may be made without departing from the spirit and scope of the processes and techniques described herein. Furthermore, other Steps may be provided, or steps may be eliminated, from the described process, and other components may be added to or removed from the described systems. Consequently, other embodiments fall within the scope of the following claims.

[0150] Throughout the description and claims of this document, the terms “include,” “contain,” and their variations mean “including but not limited to” and are not intended to exclude (and do not exclude) other components, whole numbers, or steps. Throughout this document, the singular encompasses the plural unless the context requires otherwise. In particular, where the indefinite article is used, this document is to be understood as considering plurality as well as singularity, unless the context requires otherwise.

[0151] The features, integers, characteristics, or groups described in conjunction with a particular aspect, embodiment, or example of the invention shall be understood as applicable to any other aspect, embodiment, or example, unless they are inconsistent with them. All features disclosed herein (including the claims, abstract, and accompanying drawings), and / or all steps of any method or process so disclosed, may be combined in any combination, except combinations where at least some of these features and / or steps are mutually exclusive. The invention is not limited to the details of the preceding examples or embodiments.The invention extends to any new feature, or any new combination, of the features disclosed in this document (including any claim, abstract and any attached drawing) or to any new feature, or any new combination, of the steps of any disclosed method or process. REFERENCES.

[0152] Berthelot, D., Carlini, N., Cubuk, ED, Kurakin, A., Sohn, K., Zhang, H., Raffel, C.: Remix match: Semi-supervised learning with distribution alignment and augmentation anchoring. arXiv preprint arXiv:1911.09785 (2019)

[0153] Berthelot, D., Carlini, N., Goodfellow, I., Papernot, N., Oliver, A., Raffel, CA: Mixmatch: A holistic approach to semi-supervised learning. Advances in neural information processing Systems 32 (2019)

[0154] Chapelle, O., Schôlkopf, B., Zien, A.: Introduction to semi-supervised learning (2006)

[0155] Guan, M., Gulshan, V., Dai, A., Hinton, G.: Who said what: Modeling individual labelers improves classification. Dans : Actes du congrès de l’AAAI sur l’intelligence artificielle, vol. 32 (2018)

[0156] Hendrycks, D, Mazeika, M, Wilson, D, Gimpel, K : Using trusted data to train deep networks on labels corrupted by severe noise. Advances in neural information Processing Systems 31 (2018)

[0157] Laine, S., Aila, T.: Temporal ensembling for semi-supervised learning. arXiv preprint arXiv:1610.02242 (2016)

[0158] Lee, D.H. et al.: Pseudo-label : The simple and efficient semi-supervised learning method for deep neural networks. Dans : Workshop on challenges in représentation learning, ICML. vol. 3, p. 896 (2013)

[0159] Lee, K.H., He, X., Zhang, L., Yang, L. : Cleannet : Transfer learning for scalable image classifier training with label noise. Dans : Proceedings of the IEEE conférence on computer vision and pattern récognition, pp. 5447-5456 (2018)

[0160] Li, W., Wang, L., Li, W., Agustsson, E., Van Gool, L : Base de données Webvision : Visual learning and understanding from web data. arXiv preprint arXiv: 1708.02862 (2017)

[0161] Liu, T., Tao, D. : Classification with noisy labels by importance reweighting. IEEE Transactions on pattern analysis and machine intelligence 38(3), 447-461 (2015)

[0162] Liu, Z., Luo, P., Wang, X., Tang, X.: Deep learning face attributes in the wild. Dans : Proceedings of International Conférence on Computer Vision (ICCV) (Décembre 2015)

[0163] Loshchilov, I., Hutter, F. : Sgdr : Stochastic gradient descent with warm restarts. arXiv preprint arXiv: 1608.03983 (2016)

[0164] Patrini, G., Rozza, A., Krishna Menon, A., Nock, R., Qu, L. : Making deep neural networks robust to label noise: A loss correction approach. Dans : Proceedings of the IEEE conférence on computer vision and pattern récognition, pp. 1944-1952 (2017)

[0165] Rasmus, A., Berglund, M., Honkala, M., Valpola, H., Raiko, T. : Semi-supervised learning with ladder networks. Advances in neural information processing Systems 28 (2015)

[0166] Fils, K., Berthelot, D., Carlini, N., Zhang, Z., Zhang, H., Raffel, C.A., Cubuk, E.D., Kurakin, A., Li, C.L.: Fixmatch : Simplifying semi-supervised learning with consistency and confidence. Advances in neural information processing Systems 33, 596-608(2020)

[0167] Song, H., Kim, M., Lee, J.G.: Selfie: Refurbishing unclean samples for robust deep learning. Dans : International Conférence on Machine Learning. pp. 5907-5915. PMLR (2019)

[0168] Tanno, R., Saeedi, A., Sankaranarayanan, S., Alexander, D.C., Silberman, N.: Learning from noisy labels by regularized estimation of annotator confusion pp. 11244-11253(2019)

[0169] Tarvainen, A., Valpola, H.: Mean teachers are better rôle models: Weight-averaged consistency targets improve semi-supervised deep learning results. Advances in neural information processing Systems 30 (2017)

[0170] Xiao, T., Xia, T., Yang, Y., Huang, C., Wang, X.: Learning from massive noisy labeled data for image classification. Dans : Proceedings of the IEEE conférence on computer vision and pattern récognition, pp. 2691-2699 (2015)

[0171] Xie, Q., Dai, Z., Hovy, E., Luong, T., Le, Q.: Unsupervised data augmentation for consistency training. Advances in Neural Information Processing Systems 33, 6256-6268 (2020)

[0172] Zhu, X., Goldberg, A.B.: Introduction to semi-supervised learning. Synthesis lectures on artificial intelligence and machine learning 3(1), 1-130 (2009)

Claims

Demands

1. A computing device comprising a processor coupled to a storage device containing instructions which, when executed by the processor, cause the computing device to: classify hair color and hair reflectance on an input image using a neural network, the neural network comprising a coding skeleton coupled to a pair of classifiers comprising i) a color classifier that outputs a hair color value and ii) a reflectance classifier that outputs a hair reflectance value, each element of the classifier pair comprising a linear classifier, the neural network having been defined by learning with a cross-entropy loss sum of color and reflectance determined from i) the outputs of the hair color values ​​and the hair reflectance values, and ii) the target labels for the hair color and the hair reflectance,The target labels are prepared for image training based on votes from a plurality of experts.

2. A computer device according to claim 1, wherein the neural network is configured to classify hair hue and hair reflectance according to a plurality of respective classes to follow an industry standard for hair color.

3. A computer device according to claim 1, wherein the hair reflectance comprises a primary reflectance component and a secondary reflectance component

4. A computer device of claim 1, wherein the target labels include soft annotations for at least some of the training images, the soft annotation being determined from an empirical allocation of expert votes on the respective classes.

5. A computer device of claim 1, wherein the neural network has been trained with respective annotator confusion matrices comprising a plurality (n) of hue annotator confusion matrices and a plurality (n) of reflectance annotator confusion matrices, where n is defined from a total number of experts providing expert votes, and wherein the output of the hue classifier is multiplied by a respective hue confusion matrix from among the plurality of hue confusion matrices and wherein the output of the reflectance classifier is multiplied by a respective reflectance confusion matrix from among the plurality of reflectance confusion matrices to predict the vote of a respective expert from among the plurality of experts and wherein the neural network and the respective annotator confusion matrices are trained together.

6. Computer device according to claim 1, wherein the neural network has been trained in accordance with a Student-Teacher framework.

7. Computer device according to claim 6, wherein the neural network comprises a teacher network obtained from the Student-Teacher frame.