An image classification method for cross-domain self-taken face acne grading with self-adaptive skin color

By fusing cross-domain adaptive and image classification in deep neural networks, combining adversarial generation networks and adaptive skin tone gating classification models, the problem of cross-domain data domain offset of acne image data set is solved, and a more accurate selfie face acne grading effect is achieved.

CN115035068BActive Publication Date: 2025-05-27HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210680710.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-15
Publication Date
2025-05-27
Estimated Expiration
2042-06-15

AI Technical Summary

Technical Problem

The existing technology has misdiagnosis in the acne severity rating, and the acne image data set is highly professional in labeling and scarce data resources, making it difficult to effectively solve the problem of cross-domain data domain offset.

Method used

A deep neural network that integrates cross-domain adaptive and image classification is adopted to realize style conversion between source domain and target domain samples through adversarial generation network, reduce domain offset, and build an adaptive skin color gating classification model, and use the multi-core maximizing mean difference method for feature space alignment.

Benefits of technology

In the self-portraited face acne grading task in different data domains, the accuracy and robustness of the model can be improved, and the image classification of different skin colors can be better adapted to the rate of misdiagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115035068B_ABST
    Figure CN115035068B_ABST
Patent Text Reader

Abstract

The present invention proposes an image classification method for cross-domain selfie face acne grading with adaptive skin color. The steps of the present invention are as follows: 1. Between the source domain and the target domain, use the adversarial generative network model for cross-domain data augmentation to reduce the domain shift. 2. Construct two gating networks to adaptively learn the optimal sample weights. Among them, construct an expert gating network to adaptively learn the optimal feature weights and a skin color gating network to adaptively learn the optimal skin color weights. 3. Between the source domain and the target domain, use the multi-kernel maximum mean discrepancy method to align the sample features, aiming to reduce the domain bias between the source domain and the target domain. 4. Establish a multi-task end-to-end deep learning model according to the above steps, train the entire network on a specific dataset, and test the performance of the final model on the test set. The present invention can adaptively learn the most appropriate sample weight allocation for a specific dataset, and has strong practicality and universality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image classification, and in particular to an image classification method for grading acne on a self-portrait face with self-adaptive skin color. Background Art

[0002] With the progress of the times and the rapid development of medical cosmetology, people are beginning to pay more and more attention to the health of their skin, and skin care has gradually become a hot topic. Acne, commonly known as blackheads, is the most common skin disease. About 80% of adolescents will suffer from acne, of which about 3% of males and 12% of females have acne symptoms that persist into adulthood. Due to the impact of the epidemic, people have to wear masks for protection. This makes acne-related diseases more serious and occurs at more different ages. However, acne can also leave scars and pigmentation, which often leads to low self-esteem and depression. Therefore, a large number of acne patients are in urgent need of more precise treatment. The severity of acne is crucial for dermatologists to make accurate and standardized treatment plans.

[0003] However, when dermatologists rate the severity of patients' skin acne, they may misdiagnose due to individual subjectivity, lack of experience and other factors, which indirectly causes the patient's condition to continue to deteriorate. With the development of deep learning technology, various intelligent auxiliary diagnosis algorithms based on deep learning have made continuous breakthroughs and innovations, and have achieved relatively impressive results in various traditional diagnostic tasks and emerging tasks in medical image analysis, such as medical image classification, detection, segmentation, registration, and image-based guided treatment and intervention, showing extraordinary potential.

[0004] Unlike the natural image field, which already has a series of public, complete, large-scale labeled datasets, such as MNIST, CIFAR, ImageNet, etc., the annotation process of acne image datasets is highly professional, and labeled datasets are relatively scarce. Due to the existence of facial privacy issues, large-scale, complete and available data resources are even rarer. At the same time, changes in image background, lighting, and patient skin color will bring about changes in the data domain. In this case, using domain adaptation methods to mine discriminative information from small sample data or a small amount of labeled data is an effective way to meet the characteristics of current acne image neighborhood resources.

[0005] Domain adaptation aims to fit the distribution differences between different data domains. Generally, it is assumed that the tasks in different domains are the same. In machine learning problems, it is usually assumed that the test data and the training data have the same distribution. However, if this assumption is not verified, the performance of the model on the test set will decrease significantly. In computer vision applications, such distribution differences (domain shifts) are common in real life and may be the result of changes in conditions (such as background, light intensity, etc.). In classification tasks, it uses the knowledge learned from the labeled data in one or more related source domains to guide the classifier for the unlabeled data in the target domain. In real life, most of the data in the target domain is unlabeled or the amount of data to be labeled is too large, involving a large amount of manual work. Therefore, in this case, domain adaptation techniques can be selected to build the target model. Summary of the Invention

[0006] The present invention provides an image classification method for cross-domain self-taken face acne grading with adaptive skin color. Traditional acne grading methods first detect the existing acne, then calculate the number of detected acne, and grade the image according to this number. However, different from traditional acne grading methods, this method integrates cross-domain adaptation and image classification in a unified deep neural network to complete an end-to-end deep learning model, which can directly complete the acne grading classification task for self-taken images and simultaneously complete the cross-domain adaptation task between the source domain and the target domain. In terms of the cross-domain adaptation task between the source domain and the target domain, an adversarial generative network model is used to achieve the style conversion between the source domain samples and the target domain samples, thereby reducing the domain shift between the source domain samples and the target domain samples and realizing the cross-domain adaptation between the source domain and the target domain; in terms of the image classification of self-taken face acne grading, a cross-domain adaptation module and an adaptive skin color sensitive module are used to make the model achieve better results on relevant data sets.

[0007] An image classification method for cross-domain self-taken face acne grading with adaptive skin color is as follows:

[0008] Step (1): Image data preprocessing

[0009] Since the image sizes in the data set are different, in order to adapt to the deep learning framework, it is necessary to perform size transformation on the images before the model starts training to unify the image sizes in the data set.

[0010] Step (2): Cross-domain data augmentation

[0011] The present invention introduces a publicly available dataset as the source domain to assist the learning of the classifier in the target domain. Due to differences in sample illumination and dissimilar sample backgrounds between the source domain data and the target domain data, there is a domain shift between the source domain data and the target domain data. Therefore, in the source domain and the target domain, an adversarial generative network model is used for cross-domain data augmentation to reduce the domain shift.

[0012] Step (3): Construct an adaptive skin tone gating classification model

[0013] The present invention believes that skin tone variation has a certain impact on acne grading. For example, in the case of fair skin, when acne, allergies, or similar phenomena occur, the classifier can more easily identify acne to give a more accurate skin quality evaluation result; on the contrary, in the case of dark skin, acne is not easily recognizable and it is more difficult for the classifier to make a correct evaluation. Therefore, we propose to construct an adaptive skin tone gating classification model. The skin tone gating classification model can adaptively learn the weight vector regarding the skin tone of the sample, making the performance of the classification model better. In the source domain dataset and the target domain dataset, the skin tone is divided into five categories: white, neutral, tan, brown, and black according to the individual type angle (ITA°) index.

[0014] Step (4): Feature space alignment

[0015] Due to the domain shift between the source domain data and the target domain data, there will also be a domain shift between the real images and the generated images in the source domain and the target domain respectively. Therefore, a feature-based domain alignment loss is designed to bridge the domain gap. In this method, the multi-kernel maximum mean discrepancy method is used.

[0016] Step (5): Model training and testing

[0017] A multi-task end-to-end deep learning model is established according to the above steps. On a specific dataset, the network parameters are trained through the backpropagation algorithm until the entire network model converges, and then the performance of the final model is tested on the test set.

[0018] The image data preprocessing described in step (1) is specifically as follows:

[0019] Since the sizes of each image in the dataset are different, first, all images are uniformly adjusted to a certain fixed size through the bilinear interpolation method. Secondly, the resized images are randomly cropped to obtain image data with a size of 256*256. Finally, the images are normalized.

[0020] The cross-domain data augmentation described in step (2) has the following specific process:

[0021] 2-1. The source domain is a set containing n images, and each image has a corresponding skin quality category label, denoted as Is = {(x i , y i ) | 1 ≤ i ≤ n}; The target domain is a set containing m images, and each image has a corresponding skin quality category label, denoted as where y i ∈ Y and are the class labels of the images x i and respectively. Y = {1, 2, …, N c} represents the class label space, and N c is the total number of classes.

[0022] is an image generator that converts samples in the source domain into images with the style of samples in the target domain. The set of generated images in the target domain is denoted as is an image generator that converts samples in the target domain into images with the style of samples in the source domain. The set of generated images in the source domain is denoted as

[0023] The specific formula is defined as follows:

[0024]

[0025] where x i ∈ I s , denotes the loss function; |||| 1 represents taking the 1-norm.

[0026] 2 - 2. Construct a pair of image discriminators in the image space and the feature space, denoted as and and a pair of feature discriminators, denoted as and used to discriminate whether the samples passing through the image discriminator network in the source domain are real images or generated images; used to discriminate whether the samples passing through the image discriminator network in the target domain are real images or generated images. used to determine whether the features extracted from the source domain through the classification network come from real image samples or from generated image samples; used to determine whether the features extracted from the target domain through the classification network come from real image samples or from generated image samples.

[0027] where the discriminant losses in the image space and the feature space are specifically as shown in Equation 3:

[0028]

[0029] where Denote the classification network, s represents the source domain, t represents the target domain, and d takes values of s and t.

[0030]

[0031] Among them, l i represents the true label of the i-th sample image. When used to calculate the discriminant loss in the image space When When used to calculate the discriminant loss in the feature space When is an intermediate parameter variable, see Formula 2.

[0032] The construction of the adaptive gating classification model described in step (3) is as follows:

[0033] The present invention designs an expert gating network and a skin color gating network to adaptively learn the optimal sample weights of the expert network and the sub-network respectively to solve the skin color label noise. This way also ensures the reasonable contribution of each sample to all sub-networks. In particular, the number of sub-networks is set different from the number of skin color categories to break the one-to-one mapping between the number of sub-networks and the number of skin color categories.

[0034] First, the sample passes through multiple parallel expert networks to obtain corresponding feature vectors. At the same time, the sample passes through the expert gating network to adaptively learn a gating weight vector. The gating weight vector is multiplied by the feature vector to obtain the final weight-combined feature vector. Then the weight-combined feature vector passes through multiple sub-networks at the same time, and the features extracted by the multiple sub-networks are correspondingly output; the features extracted by the sub-networks are multiplied by the skin color proportion weight vector adaptively learned by the sample passing through the skin color gating network to obtain the final combined feature.

[0035] In this method, there is a classification network in both the source domain and the target domain. Therefore, the adaptive gating network has three structures, namely the classification network only in the source domain, the classification network only in the target domain, and the classification network jointly in the source domain and the target domain. The specific definitions are as follows:

[0036]

[0037] Among them, N′ t represents the number of sub-networks, N e represents the number of expert networks, represents the fully connected layer of the i-th sub-network, M represents the aggregation layer, E j represents the j-th expert network, G j represents the j-th element of the N e weight vector, W i represents the i-th element of the N′ t weight vector.

[0038]

[0039] Among them, τ(x) represents the skin color class label of sample x, and e τ (x) represents N′ t The τ(x)-th element of the rights protection weight vector is 1; represents the skin color class label of the real image corresponding to the generated image sample, represents N′ t The -th element of the rights protection weight vector is 1; γ and γ′ respectively represent the weight parameters of the real image and the generated image.

[0040] The feature space alignment described in step (4) is specifically as follows:

[0041] In this method, the maximum mean discrepancy loss of multiple kernels is used for the real image samples and the generated image samples in their respective domains in the source domain and the target domain. Therefore, the cross-domain generated images from another domain are more compatible with the real images in a specific domain, ensuring the rationality of sharing the same classification model for real images and generated images. The maximum mean discrepancy loss of multiple kernels is specifically defined as follows:

[0042]

[0043] Among them is the weight parameter controlling the source domain, is the weight parameter controlling the target domain; φ s is the mapping function in the source domain, and φ t is the mapping function in the target domain; represents taking the 2-norm; E represents taking the expectation. The maximum mean discrepancy loss of multiple kernels can be applied to a single domain or two domains with shared or non-shared mapping functions.

[0044] The construction of the multi-task deep learning model described in step (5) specifically means that after establishing an end-to-end framework according to steps (2), (3), and (4), on a specific data set, the classification loss, the image generation loss, and the domain alignment loss are optimized simultaneously. During the training process, the parameters of the adversarial generation network model and the classification model are updated simultaneously to obtain the final model and test the training effect on the test set.

[0045] Advantages of the present invention:

[0046] Based on the idea of cross - domain image generation and adaptive learning of gated networks, a classification model for grading the severity of acne on self - taken faces is proposed. By introducing a publicly available source - domain dataset as an auxiliary dataset, a framework for grading the severity of acne based on deep transfer learning is proposed. An expert gated network and a skin - tone gated network model are designed to adaptively learn the correlation between labels and the skin - tone in the feature space. Brief Description of the Drawings

[0047] Figure 1 It is a schematic diagram of the specific process of the method of the present invention.

[0048] Figure 2 It is a schematic diagram of the network framework constructed in the method of the present invention.

[0049] Figure 3 Experiment on the combination of the number of expert networks and the number of sub - networks.

[0050] Specific Implementation Details

[0051] The following further specifically describes the present invention in conjunction with Figure 1 , 2. The present invention proposes a cross - domain self - taken face acne grading method with adaptive skin - tone. The steps of the present invention are as follows: 1. Between the source domain and the target domain, use the adversarial generative network model for cross - domain data augmentation to reduce the domain shift. 2. Construct two gated networks to adaptively learn the optimal sample weights. Among them, construct an expert gated network to adaptively learn the optimal feature weights, and a skin - tone gated network to adaptively learn the optimal skin - tone weights. 3. Between the source domain and the target domain, use the multi - kernel maximum mean discrepancy method to align the sample features, aiming to reduce the domain bias between the source domain and the target domain. 4. Establish a multi - task end - to - end deep learning model according to the above steps, train the entire network on a specific dataset, and test the performance of the final model on the test set. The present invention can adaptively learn the most suitable sample weight distribution for a specific dataset, and has strong practicality and universality.

[0052] The specific implementation steps of the present invention are as follows:

[0053] The first step:

[0054] We use the Acne04 dataset as the source domain, and use the Acnehdu, AcnehduP, and AcnePGP datasets as the target domains respectively to verify our classification model for grading the severity of acne on self - taken faces. When training the model, first adjust the size of the images in the four datasets to 286 * 286 through bilinear interpolation, then randomly crop each image to 256 * 256, and finally normalize the pixel values of the images. When testing the model, the data - processing process is the same as that during training.

[0055] Step 2:

[0056] In this method, the Acne04 dataset is the source domain dataset we introduced, and the Acnehdu, AcnehduP, and AcnePGP datasets are used as target domain datasets respectively. Among them, the following process is illustrated by taking the Acne04 and Acnehdu datasets as examples.

[0057] First, the sample image in the source domain dataset Acne04 is denoted as x s , and through the generator network the generated image corresponding to x s is obtained, denoted as x s2t ; Second, the sample image x t in the target domain dataset Acnehdu passes through the generator network to obtain the generated image corresponding to x t , denoted as x t2s . In the source domain, both x s and x t2s participate in the training of the classification model in the source domain; in the target domain, both x t and x s2t will participate in the training of the classification model in the target domain.

[0058] After that, in the source domain, x s and x t2s pass through the discriminator D s , and in the target domain, x t and x s2t pass through the discriminator D t . D s and D t will discriminate whether the sample is a real image or a generated image. At the same time, there will also be a pair of discriminators in the feature space extracted by the classification network, with the same function.

[0059] Step 3:

[0060] We first tested different structures in the three datasets, and the results are shown in Table 1 below. We found that due to the existence of domain shift, the optimal gated network structure is different in different target domain datasets. At the same time, we evaluated two important parameters in the gated classification network: the number of expert networks (N e ) and the number of sub-networks (N′ t ), and the results are shown in Table 2. We found that 1) the optimal N′ t may be different from the number of skin color categories and the larger the number of sub-networks (N′ t ), the easier it is to obtain better performance; 2) the optimal number of expert networks (N e ) tends to be less than the optimal number of sub-networks (N′ t), which means sharing the correlation and effectiveness of the underlying feature space across multiple skin tones. In our method, a gating network is utilized to adaptively learn the optimal number of subnetworks and expert networks for a specific dataset.

[0061] Table 1 Ablation experiments of gating classification network modules with different structures

[0062]

[0063]

[0064] Step 4:

[0065] We apply the maximum mean discrepancy (MK-MMD) loss with multiple kernels on the generated and real image datasets to achieve cross-domain alignment. On AcnePGP, we first test different numbers of kernels K and alignment structures using Gaussian kernels under a model with cross-domain data augmentation. The results are shown in Table 2:

[0066] Table 2 Experiments on the multiple-kernel maximum mean difference loss module

[0067] Type Only in the source domain Only in the target domain Joint sharing Joint non-sharing K=3 0.872 0.849 0.847 0.857 K=5 0.857 0.766 0.857 0.852 K=8 0.864 0.766 0.766 0.849

[0068] Specifically, we construct the MK-MMD loss on the source domain only, the target domain only, the two domains with shared mapping, and the two domains without shared mapping, respectively. We observe that the alignment structure has a significant impact on the performance results. To obtain better performance, we apply the best alignment structure for each dataset under three Gaussian kernels.

Claims

1. An image classification method for cross - domain self - portrait human face acne grading with adaptive skin color, characterized in that it includes the following steps: Step (1): Image data pre - processing, unifying the image sizes in the dataset; Step (2): Cross - domain data augmentation, introducing a publicly available dataset as the source domain to assist the learning of the classifier in the target domain; Step (3): Constructing an adaptive skin - color gated classification model; Step (4): Feature space alignment, designing a feature - based domain alignment loss to bridge the domain gap; Step (5): Model training and testing; establishing a multi - task end - to - end deep learning model, and on a specific dataset, training the network parameters through the back - propagation algorithm until the entire network model converges; The cross - domain data augmentation described in Step (2) has the following specific process: 2-1. The source domain is a set containing n images, and each image has a corresponding skin quality category label, denoted as I s = {(x i , y i ) | 1 ≤ i ≤ n}; The target domain is a set containing m images, and each image has a corresponding skin quality category label, denoted as where y i ∈ Y and are the category labels of images x i and respectively, Y = {1, 2, …, N c} represents the category label space, and N c is the total number of categories; is an image generator that converts samples in the source domain into images with the style of samples in the target domain. The set of generated images in the target domain is denoted as is an image generator that converts samples in the target domain into images with the style of samples in the source domain. The set of generated images in the source domain is denoted as The specific formula is defined as follows: where x i ∈ I s , denotes the loss function; || || 1 represents taking the 1-norm; 2-2. Construct a pair of image discriminators in the image space and the feature space, denoted as and and a pair of feature discriminators, denoted as and which are used to determine whether the samples passing through the image discriminator network in the source domain are real images or generated images; which are used to determine whether the samples passing through the image discriminator network in the target domain are real images or generated images; which are used to judge whether the features extracted from the source domain through the classification network come from real image samples or generated image samples; which are used to judge whether the features extracted from the target domain through the classification network come from real image samples or generated image samples; Among them, the discriminant losses in the image space and the feature space are specifically shown in the following formulas: Among them represents the classification network, s represents the source domain, t represents the target domain, and d takes values of s and t; where, l i represents the true label of the i-th sample image; when used to calculate the discriminative loss in the image space when when used to calculate the discriminative loss in the feature space when ψ(x) is an intermediate parameter variable; The specific process of constructing the adaptive skin - color gated classification model described in Step (3) is as follows: Designing an expert gating network and a skin - color gating network to adaptively learn the optimal sample weights of the expert network and the sub - network respectively to solve the skin - color label noise; and setting the number of sub - networks different from the number of skin - color categories, thus breaking the one - to - one mapping between the number of sub - networks and the number of skin - color categories; First, the sample passes through multiple parallel expert networks to obtain corresponding feature vectors. At the same time, the sample passes through the expert gating network to adaptively learn a gating weight vector, and the gating weight vector is multiplied by the feature vector to obtain the finally weighted combined feature vector; Then, the weighted combined feature vector passes through multiple sub - networks simultaneously, and the features extracted by the multiple sub - networks are correspondingly output; the features extracted by the sub - networks are multiplied by the skin - color proportion weight vector adaptively learned by the sample passing through the skin - color gating network to obtain the final combined feature.

2. An image classification method for cross - domain self - portrait human face acne grading with adaptive skin color according to claim 1, characterized in that it has the following additional technical features: There is a classification network in both the source domain and the target domain. Therefore, the adaptive skin - color gated network has three structures, namely the classification network only in the source domain, the classification network only in the target domain, and the classification network jointly in the source domain and the target domain. The specific definitions are as follows: Among them, N' t represents the number of sub-networks, N e represents the number of expert networks, represents the fully connected layer of the i-th sub-network, M represents the aggregation layer, E j represents the j-th expert network, G j represents N e the j-th element of the weight vector for rights protection, W i represents N' t the i-th element of the weight vector for rights protection; Among them, τ(x) represents the skin color category label of sample x, and e τ(x) represents the N′ t the τ(x)-th element of the rights protection weight vector is 1; represents the skin color category label of the real image corresponding to the generated image sample, represents the N′ t the -th element of the rights protection weight vector is 1; γ and γ′ respectively represent the weight parameters of the real image and the generated image.

3. An image classification method for cross - domain self - portrait human face acne grading with adaptive skin color according to claim 1 or 2, characterized in that the specific process of the feature space alignment described in Step (4) is as follows: Using the multi - kernel maximum mean discrepancy loss for the real image samples and the generated image samples in their respective domains in the source domain and the target domain; therefore, the cross - domain generated images from another domain are more compatible with the real images in a specific domain, ensuring the rationality of sharing the same classification model for real images and generated images; the multi - kernel maximum mean discrepancy loss is specifically defined as follows: where is the weight parameter for controlling the source domain, is the weight parameter for controlling the target domain; φ s is the mapping function in the source domain, φ t is the mapping function in the target domain; denotes taking the 2-norm; E denotes taking the expectation; the maximum mean discrepancy loss for multiple kernels is applied to a single domain or two domains with shared or non-shared mapping functions.