Soft-Label Based Facial Acne Classification Method, System, Device and Medium
Through the soft tag-based facial acne classification method, the combination of teacher network and student network is used to dynamically adjust the characteristics and emphasize weight, solving the problem of low classification accuracy of acne images caused by coupling in the existing technology, and achieving higher classification accuracy.
Patent Information
- Application Number
- CN202411609086.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-12
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-11-12
AI Technical Summary
In the prior art, due to the coupling of acne images between different levels, the model has low accuracy in classification of acne images.
The facial acne classification method based on soft labels is adopted. By building a teacher network and student network, using soft labels to train students networks, dynamically adjust features and emphasize weights, and improve the accuracy of the model classification of acne images.
It achieves higher feature accuracy, improves low-level learning efficiency, enhances the classification accuracy of acne images, and solves the problem of low classification accuracy caused by coupling.
Smart Images

Figure CN119600331B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, relates to the classification of facial acne, and particularly relates to a method, system, device and medium for classifying facial acne based on soft labels. Background Art
[0002] Acne is a very common skin disease. Acne is a multifactorial disease of the pilosebaceous unit, which can be clinically manifested as mild comedonal acne or fulminant acne with systemic symptoms. According to relevant research, about 80% of teenagers have suffered from acne, and 3% of boys and 12% of girls still have untreated acne even in adulthood. The scars and pigmentation caused by acne will undoubtedly cause great damage to the psychological and mental health of patients. As a key step in the diagnosis and treatment of acne, accurate grading of the severity of acne is very important for obtaining a treatment plan suitable for patients.
[0003] When classifying acne, most often doctors will classify it based on the overall facial condition of the patient and their experience. When classifying, it usually combines standard lesion counting and experience-based whole-image assessment. Common grading methods such as the Hayashi grading method or the Pillsbury grading method divide the severity level of acne into four grades, including: "mild" (slight), "moderate" (medium), "severe" (severe), and "very severe" (extremely severe). Manual classification by doctors not only greatly increases the workload of doctors, but also has low annotation efficiency.
[0004] With the development of artificial intelligence and deep neural network technologies in recent years, it is very necessary to automatically grade the severity of acne with the help of a computer for patients to understand their own conditions and assist doctors in formulating treatment plans.
[0005] The invention patent application with the application number 202310795157.7 discloses an acne grading method, device, equipment and medium. The method includes: obtaining a facial image of a target user, wherein the facial image includes an image corresponding to acne; detecting the facial image through a pre-trained detection model to obtain at least one acne type and the number of acne corresponding to the acne type; grading the acne type and the number of acne corresponding to the acne type through a pre-trained grading model to obtain the acne grade of the target user. This solution realizes the analysis of the facial image of the target user to obtain the acne grade of the target user, making the acne grade independent of the doctor's experience and making the division of the acne grade more objective and accurate.
[0006] Similar to the above-mentioned invention patent application, when training a model in the prior art, most of the labels used are hard labels, that is, labels with a value of either 0 or 1. The problem with such labels is that the information is too concentrated. In natural images, even if hard labels are used to train the model, due to the large differences between classes, good results can be obtained. However, in the classification of acne images, since there are overlapping parts between acne images of different grades, if hard labels are still used to train the model, the classification accuracy of the model for acne images is low, and it is urgent to improve the classification accuracy of the model for acne images. Summary of the Invention
[0007] The purpose of the present invention is to provide a facial acne classification method, system, device and medium based on soft labels in order to solve the problem in the prior art that the classification accuracy of the model for acne images is low due to the coupling between acne images of different grades.
[0008] The present invention specifically adopts the following technical solutions to achieve the above purpose:
[0009] A facial acne classification method based on soft labels, comprising the following steps:
[0010] Step S1, obtaining facial acne sample data;
[0011] Obtaining facial acne sample images, the facial acne sample images including facial acne sample images with hard labels and facial acne sample images without labels;
[0012] Step S2, building a facial acne classification model;
[0013] Building a facial acne classification model, the facial acne classification model including a teacher network and a student network;
[0014] Step S3, training the facial acne classification model;
[0015] Using the facial acne sample images with hard labels in Step S1 to pre-train the teacher network in Step S2, inputting the facial acne sample images without labels in Step S1 into the pre-trained teacher network, and the teacher network outputs soft labels; using the facial acne sample images without labels in Step S1 and the soft labels output by the teacher network to train the student network in Step S2;
[0016] When training the facial acne classification model, the loss function is:
[0017]
[0018]
[0019]
[0020] Among them, represents the loss function of the target category, represents the loss function of non-target categories, represents the predicted probability of the teacher network for the target category, represents the predicted probability of the student network for the target category, represents the predicted probability of the teacher network for non-target categories, represents the predicted probability of the student network for non-target categories, represents the dynamic weight vector used to emphasize features, represents the difference between the teacher network and the student network in non-target categories, represents the th category, represents the total number of categories, represents the target category, represents the ith component of the dynamic weight vector, represents the ith component of the vector, represents the ith component of the vector, represents the independent target category probability corresponding to the teacher network, represents the independent target category probability corresponding to the student network;
[0021] Step S4, facial acne classification;
[0022] Obtain a facial acne image to be classified and input it into the student network, and the student network outputs the classification result.
[0023] Furthermore, in step S1, the facial acne sample image is preprocessed, and the size of the facial acne sample image is adjusted to 224*224*3.
[0024] Furthermore, the structures of the teacher network and the student network are the same, and both include a horizontal first convolutional layer, a horizontal second convolutional layer, a horizontal third convolutional layer, and a horizontal fully connected layer arranged in sequence. The output of the horizontal first convolutional layer serves as the input of the horizontal second convolutional layer and the vertical first convolutional layer, and the output of the vertical first convolutional layer serves as the input of the vertical first fully connected layer; the output of the horizontal second convolutional layer serves as the input of the horizontal third convolutional layer and the vertical second convolutional layer, and the output of the vertical second convolutional layer serves as the input of the vertical second fully connected layer; the output of the horizontal third convolutional layer serves as the input of the horizontal fully connected layer and the vertical third convolutional layer, and the output of the vertical third convolutional layer serves as the input of the vertical third fully connected layer.
[0025] Further, in step S3, when training the facial acne classification model, the loss function is respectively used between the output of each fully connected layer and the output of each previous fully connected layer to calculate the loss.
[0026] Further, in step S3, each component of the weight is expressed as:
[0027]
[0028] where represents the th category, represents the value of the output of the teacher network on the target class, represents the value of the output of the teacher network on the th non-target class, represents a scaling factor used to scale the dynamic weight to a suitable range, represents the initial value of the dynamic weight to improve the stability of the method.
[0029] Further, in step S3, during training, the learning rate is set to 0.001, is set to 0.0001, the batch size is set to 32, and the parameters and are respectively set to 1.3 and 8, and the entire training process is optimized using the SGD optimizer.
[0030] A facial acne classification system based on soft labels includes:
[0031] A facial acne sample data acquisition module for acquiring facial acne sample images, where the facial acne sample images include facial acne sample images with hard labels and facial acne sample images without labels;
[0032] A facial acne classification model construction module for constructing a facial acne classification model, where the facial acne classification model includes a teacher network and a student network;
[0033] A facial acne classification model training module for pre-training the teacher network in the facial acne classification model construction module using the facial acne sample images with hard labels in the facial acne sample data acquisition module, inputting the facial acne sample images without labels in the facial acne sample data acquisition module into the pre-trained teacher network, and the teacher network outputs soft labels; training the student network in the facial acne classification model construction module using the facial acne sample images without labels in the facial acne sample data acquisition module and the soft labels output by the teacher network;
[0034] When training a facial acne classification model, the loss function is as follows:
[0035]
[0036]
[0037]
[0038] Among them, represents the loss function of the target category, represents the loss function of the non-target category, represents the prediction probability of the teacher network for the target category, represents the prediction probability of the student network for the target category, represents the prediction probability of the teacher network for the non-target category, represents the prediction probability of the student network for the non-target category, represents the dynamic weight vector used to emphasize features, represents the difference between the teacher network and the student network in the non-target category, represents the th category, represents the total number of categories, represents the target category, represents the dynamic weight vector 's i-th component, represents the i-th component of the vector 's i-th component, represents the i-th component of the vector 's i-th component, represents the independent target category probability corresponding to the teacher network, represents the independent target category probability corresponding to the student network;
[0039] The facial acne classification module is used to obtain the facial acne image to be classified and input it into the student network, and the student network outputs the classification result.
[0040] A computer device includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor is caused to execute the steps of the above method.
[0041] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the processor is caused to execute the steps of the above method.
[0042] The beneficial effects of the present invention are as follows:
[0043] 1. The features of the present invention are accurate to the component level. In previous methods, the connections and different degrees of differences between acne samples and other categories were ignored. Therefore, each class was roughly given the same attention in the model, resulting in more difficulty in distinguishing samples that are difficult to distinguish and more confusion in easily confused samples. The present invention dynamically emphasizes the features of each class according to the model's cognitive ability for samples, enabling the model to learn more knowledge about anti-confusion from a sample, effectively solving the problem that the classification accuracy of acne images is low due to the coupling between acne images of different grades. This advantage is proposed by this method. It is achieved.
[0044] 2. The present invention can achieve more intensive feature transfer. When existing methods perform feature transfer, most of them only consider the features between the highest layer and the label or between adjacent two layers. The present invention enables the lower layer to learn features of different granularities by transferring features from the high layer to each layer, thereby enhancing the learning efficiency of the lower layer and improving the classification accuracy of acne images.
[0045] 3. The present invention considers the similarity of acne features. Most current methods ignore the similarity between acne images and simply process them according to natural image processing. The method of the present invention considers the degree of difference between each sample and each class. This advantage is achieved through the re-distillation module proposed by the present invention, improving the classification accuracy of acne images.
[0046] 4. The distribution characteristics of acne samples in the present invention. Most previous models used hard labels for model training, and hard labels with overly concentrated information are often not applicable to acne images with similarities between images. Moreover, most existing methods for generating soft labels are class-based, that is, each class corresponds to a soft label, which ignores the intra-class differences between acne images. The method of this application dynamically generates soft labels unique to each sample according to the different situations of each sample, which can better reflect the real situation of the sample. This advantage is achieved through the point-to-point architecture of knowledge distillation and is also the first method to apply the knowledge distillation architecture to acne classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is a schematic flowchart of the present invention;
[0048] Figure 2 is a schematic structural diagram of the teacher network and the student network in the present invention;
[0049] Figure 3 is a schematic diagram during the training of the teacher network and the student network in the present invention;
[0050] Figure 4 is a visual comparison diagram of various knowledge distillation methods in the present invention. Detailed implementation manners
[0051] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention.
[0052] Therefore, based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0053] Embodiment 1
[0054] This embodiment provides a facial acne classification method based on soft labels, which is used to classify facial acne images to obtain classification results. As Figure 1 shown, the classification method specifically includes the following steps:
[0055] Step S1, obtaining facial acne sample data;
[0056] Obtaining facial acne sample images, where the facial acne sample images include facial acne sample images with hard labels and facial acne sample images without labels.
[0057] The sample data in this embodiment mainly includes two parts. One is the publicly available facial acne dataset ACNE04, and the other is the dataset ACNE-HX collected from West China Hospital of Sichuan University. Among them, the dataset ACNE04 contains 1457 facial acne images, and the acne data in all datasets are divided into four acne severity levels. The dataset ACNE-HX includes 1189 facial acne images, and the acne data included therein are divided into eight acne severity levels. In order to make the acne severity levels in the two datasets the same, that is, both are set to four levels. For example, the first and second levels in the dataset ACNE-HX are divided into a new level, and the third and fourth levels are divided into another new level, and so on, so that the levels of the two datasets are both divided into four acne severity levels.
[0058] After the sample data is collected, the data needs to be preprocessed accordingly to meet the requirements of the model. In this embodiment, the data is in a three-dimensional format (picture length * picture width * number of channels), and the size of the facial acne sample images is adjusted to 224 * 224 * 3.
[0059] After the preprocessing of the sample data is completed, the sample data is divided into a dataset. In this embodiment, the sample dataset is divided into two parts, a training set and a test set. The training set is used to train the model and adjust the model weights; the test set is used to test the performance of the finally obtained model. For the dataset used, the training set and the test set are divided in a ratio of 4:1.
[0060] Step S2, construct a facial acne classification model;
[0061] Construct a facial acne classification model, which includes a teacher network and a student network. Among them, the teacher model will be pre-trained first, and then the teacher model will guide the training of the student model. The upper layer of the student model will also be re-distilled to guide the training of the lower layer. Finally, the student model is used to make corresponding predictions.
[0062] The structures of the teacher network and the student network are the same. The specific structure is as Figure 2 shown. They both include a horizontal first convolutional layer, a horizontal second convolutional layer, a horizontal third convolutional layer, and a horizontal fully connected layer arranged in sequence. The output of the horizontal first convolutional layer is used as the input of the horizontal second convolutional layer and the vertical first convolutional layer. The output of the vertical first convolutional layer is used as the input of the vertical first fully connected layer; the output of the horizontal second convolutional layer is used as the input of the horizontal third convolutional layer and the vertical second convolutional layer. The output of the vertical second convolutional layer is used as the input of the vertical second fully connected layer; the output of the horizontal third convolutional layer is used as the input of the horizontal fully connected layer and the vertical third convolutional layer. The output of the vertical third convolutional layer is used as the input of the vertical third fully connected layer.
[0063] Step S3, train the facial acne classification model;
[0064] Use the facial acne sample images with hard labels in step S1 to pre-train the teacher network in step S2. Input the unlabeled facial acne sample images in step S1 into the pre-trained teacher network, and the teacher network outputs soft labels. Use the unlabeled facial acne sample images in step S1 and the soft labels output by the teacher network to train the student network in step S2.
[0065] As Figure 3 shown, the sample images are respectively input into the teacher network and the student network. The teacher network guides the student network to learn and updates the parameters of the student network; the loss is calculated through the hard labels (predicted classification, binary classification step, and non-target step) of the teacher model and the prediction results (binary classification step, non-target class step, and prediction step) output by the student model, and the parameters of the student network are updated. When training the facial acne classification model, the output of each fully connected layer and the output of each previous fully connected layer are respectively used with the loss function to calculate the loss. Among them, the loss function is expressed as:
[0066]
[0067]
[0068]
[0069] Among them, represents the loss function of the target class, represents the loss function of non-target classes, represents the prediction probability of the teacher network for the target class, represents the prediction probability of the student network for the target class, represents the prediction probability of the teacher network for non-target classes, represents the prediction probability of the student network for non-target classes, represents the dynamic weight vector used to emphasize features, represents the difference between the teacher network and the student network in non-target classes, represents the th class, represents the total number of classes, represents the target class, represents the dynamic weight vector of the i-th component, represents the vector of the i-th component, represents the vector of the i-th component, represents the independent target class probability corresponding to the teacher network, represents the independent target class probability corresponding to the student network.
[0070] In this embodiment, since the teacher network often has a large fitting ability, its prediction on the target class will be very confident, which will lead to tending to 0. When it is used as the coefficient, it will cause the loss contributed by to tend to 0. And is associated with non-target classes, so it will damage the model's cognitive ability for non-target classes, resulting in the model being confused between target classes and non-target classes. Therefore, in this embodiment, it is necessary to replace the weight of with a more reasonable value to improve the distillation efficiency between the teacher network and the student network, which is also one of the core innovations of this embodiment.
[0071] Acne images have sequentiality, that is, adjacent acne images with severe levels will show more similarities. Each component represents the next optimization strength of the model for the corresponding class. Therefore, the weight of each component also needs to be determined dynamically. For those classes close to the target, since their features are similar to the current sample, a smaller weight needs to be assigned to the component in to prevent overemphasis of the excessive weight on it from causing the model to be confused between it and the target class; for those classes far from the target, since their similarity to the target class is low, a larger weight can be assigned to them to emphasize their features. At the same time, it should be noted that the weight assignment is also dynamic because the situations between each sample and each sample are different.
[0072] Let the above weight be , then regarding it can be expressed as:
[0073]
[0074] where, is generated by the output of the fully connected layer of the teacher network . The output of the teacher network hides the knowledge it wants to transfer to the student network, and performing transformation processing on it can extract more refined knowledge.
[0075] The weight each component of can be calculated as:
[0076]
[0077] where, represents the th category, , represents the value of the output of the teacher network on the target class, represents the value of the output of the teacher network on the th non-target class, represents the adjustment factor used to scale the dynamic weight to a suitable range, represents the initial value of the dynamic weight to improve the stability of the method. Since and the difference between may fluctuate greatly, so is used to scale it to a suitable range in this embodiment; at the same time, in order to prevent the situation that the difference is too small from hindering the model learning, a parameter also needs to be added to provide the initial value.
[0078] The error backpropagation process is used to optimize model parameters. After calculating the corresponding loss, the gradient corresponding to each model parameter can be calculated using this loss, and the gradient descent algorithm is used to optimize the parameters. In actual execution, some parameters are added to control the optimization process. In this embodiment, the learning rate is set to 0.001, is set to 0.0001, the batch size during training is set to 32, and the parameters 、 are set to 1.3 and 8 respectively, and the entire training process is optimized using the SGD optimizer.
[0079] As Figure 2 shown, the loss is calculated between the final probability output of the student network and the soft labels generated by the teacher model. In addition, for each horizontal convolutional layer in the student network, the predicted probability of this layer is obtained by adding an additional classifier (vertical convolutional layer and vertical fully connected layer), and the loss is calculated between the predicted probability of this layer and the predicted probabilities of each previous horizontal convolutional layer. Since the information contained in the later horizontal convolutional layers is more abundant than that in the previous horizontal convolutional layers, the information of the later horizontal convolutional layers is passed to the previous horizontal convolutional layers in this way, realizing the distillation of the student network by itself. And by passing information through the fully connected layer, the transmission efficiency of high-level information can be improved, that is, re-distillation is realized.
[0080] After the student network is trained, the student network is tested to verify the actual performance of the model.
[0081] The dataset used for testing in this embodiment includes the publicly available acne image dataset ACNE04 and the ACNE-HX dataset collected by West China Hospital of Sichuan University. The test results on the ACNE04 dataset are as follows:
[0082]
[0083] Among them, the meanings of the above five indicators are as follows:
[0084] 1) Accuracy: Measures the classification prediction accuracy of the model, the higher the better;
[0085] 2) Precision: Among the detected positive examples, the proportion of the number of truly positive examples in the detected positive examples, the higher the better;
[0086] 3) Sensitivity: The proportion of detected positive examples in all truly positive examples, the higher the better;
[0087] 4) Specificity: The proportion of detected negative examples in all truly negative examples, the higher the better;
[0088] 5) Youden index: Measures the comprehensive ability of the model, the higher the better.
[0089] As can be seen from the above table, compared with the other methods, the method proposed in this application has achieved optimal performance in all indicators. In addition, the test results on the ACNE-HX dataset are as follows:
[0090]
[0091] It can be seen that the method of this application also shows excellent performance on ACNE-HX. In addition, the visualization results of the method of this application are compared with those of another similar method, and the results are as Figure 4 shown.
[0092] Figure 4 In, each column represents a case, each row represents the results of the corresponding model, and the red area represents the area of attention of the model for the sample when making predictions. From Figure 4 it can be seen that the method of this application is more sensitive to the area where acne is located. Therefore, for samples that are originally similar to the characteristics of adjacent grades and are prone to misjudgment, it can better classify them correctly.
[0093] Step S4, facial acne classification;
[0094] Obtain the facial acne image to be classified and input it into the student network, and the student network outputs the classification result.
[0095] Embodiment 2
[0096] This embodiment provides a facial acne classification system based on soft labels, which is used to classify facial acne images to obtain classification results. The classification system specifically includes:
[0097] A facial acne sample data acquisition module, which is used to acquire facial acne sample images, and the facial acne sample images include facial acne sample images with hard labels and facial acne sample images without labels.
[0098] The sample data of this embodiment mainly includes two parts. One is the publicly available facial acne dataset ACNE04, and the other is the dataset ACNE-HX collected from West China Hospital of Sichuan University. Among them, the dataset ACNE04 contains 1457 facial acne images, and the acne data in all datasets are divided into four acne severity levels. The dataset ACNE-HX includes 1189 facial acne images, and the acne data contained therein are divided into eight acne severity levels. In order to make the acne severity levels in the two datasets the same, that is, both are set to four levels. For example, the first and second levels in the dataset ACNE-HX are divided into a new level, and the third and fourth levels are divided into another new level, and so on. The levels of the two datasets are both divided into four acne severity levels.
[0099] After the sample data is collected, the data needs to be preprocessed accordingly to meet the requirements of the model. In this embodiment, the data is in a three-dimensional format (image length * image width * number of channels), and the size of the facial acne sample image is adjusted to 224 * 224 * 3.
[0100] After the sample data preprocessing is completed, the sample data is divided into datasets. In this embodiment, the sample dataset is divided into two parts: a training set and a test set. The training set is used to train the model and adjust the model weights; the test set is used to test the performance of the finally obtained model. For the dataset used, the training set and the test set are divided in a ratio of 4:1.
[0101] The facial acne classification model building module is used to build a facial acne classification model, and the facial acne classification model includes a teacher network and a student network. Among them, the teacher model will be pre-trained first, and then the teacher model will guide the training of the student model. The high-level of the student model will also be re-distilled to guide the training of the low-level. Finally, the student model is used to make corresponding predictions.
[0102] The structures of the teacher network and the student network are the same. The specific structure is as Figure 2 shown. They both include a horizontal first convolutional layer, a horizontal second convolutional layer, a horizontal third convolutional layer, and a horizontal fully connected layer arranged in sequence. The output of the horizontal first convolutional layer serves as the input of the horizontal second convolutional layer and the vertical first convolutional layer. The output of the vertical first convolutional layer serves as the input of the vertical first fully connected layer; the output of the horizontal second convolutional layer serves as the input of the horizontal third convolutional layer and the vertical second convolutional layer. The output of the vertical second convolutional layer serves as the input of the vertical second fully connected layer; the output of the horizontal third convolutional layer serves as the input of the horizontal fully connected layer and the vertical third convolutional layer. The output of the vertical third convolutional layer serves as the input of the vertical third fully connected layer.
[0103] The facial acne classification model training module is used to pre-train the teacher network in the facial acne classification model building module with the facial acne sample images with hard labels in the facial acne sample data acquisition module, input the unlabeled facial acne sample images in the facial acne sample data acquisition module into the pre-trained teacher network, and the teacher network outputs soft labels; use the unlabeled facial acne sample images in the facial acne sample data acquisition module and the soft labels output by the teacher network to train the student network in the facial acne classification model building module;
[0104] When training the facial acne classification model, the output of each fully connected layer and the output of each previous fully connected layer are respectively used with the loss function to calculate the loss. Among them, the loss function is expressed as:
[0105]
[0106]
[0107]
[0108] Among them, represents the loss function of the target class, represents the loss function of non-target classes, represents the predicted probability of the teacher network for the target class, represents the predicted probability of the student network for the target class, represents the predicted probability of the teacher network for non-target classes, represents the predicted probability of the student network for non-target classes, represents the dynamic weight vector used to emphasize features, represents the difference between the teacher network and the student network in non-target classes, represents the th class, represents the total number of classes, represents the target class, represents the dynamic weight vector 's i-th component, represents the 's i-th component, represents the 's i-th component, represents the independent target class probability corresponding to the teacher network, represents the independent target class probability corresponding to the student network.
[0109] In this embodiment, since the teacher network often has a large fitting ability, its prediction on the target class will be very confident, which will lead to tending to 0. When it is used as the coefficient, it will cause the loss contributed by to tend to 0. And is associated with non-target classes, so it will damage the model's cognitive ability for non-target classes, resulting in the model being confused between target classes and non-target classes. Therefore, this embodiment needs to replace the weight of with a more reasonable value to improve the distillation efficiency between the teacher network and the student network, which is also one of the core innovations of this embodiment.
[0110] Acne images have sequentiality, that is, adjacent acne images with severe grades will show more similarities. Each component represents the next optimization strength of the model for the corresponding class. Therefore, the weight of each component also needs to be determined dynamically. For those classes close to the target, since their features are similar to the current sample, a smaller weight needs to be assigned to the component in to prevent overemphasis by too large a weight from causing the model to be confused between it and the target class; for those classes far from the target, since their similarity to the target class is low, a larger weight can be assigned to them to emphasize their features. At the same time, it should be noted that the weight assignment is also dynamic because the situations between each sample and each sample are different.
[0111] Let the above weight , then regarding can be expressed as:
[0112]
[0113] where, is generated by the output of the fully connected layer of the teacher network. The output of the teacher network hides the knowledge it wants to pass to the student network, and performing transformation processing on it can extract more refined knowledge.
[0114] The weight each component of can be calculated as:
[0115]
[0116] where, represents the th category, , represents the value of the output of the teacher network on the target class, represents the value of the output of the teacher network on the th non-target class, represents a scaling factor used to scale the dynamic weight to a suitable range, represents the initial value of the dynamic weight, used to improve the stability of the method. Since and the difference between may fluctuate greatly, so is used to scale it to a suitable range in this embodiment; at the same time, in order to prevent the situation that the difference is too small from hindering the model learning, a parameter also needs to be added, used to provide the initial value.
[0117] The error backpropagation process is used to optimize the model parameters. After calculating the corresponding loss, the gradient corresponding to each model parameter can be calculated using this loss, and the gradient descent algorithm is used to optimize the parameters. In actual execution, some parameters are added to control the optimization process. In this embodiment, the learning rate is set to 0.001, set to 0.0001, the batch size during training is set to 32, and the parameters 、 are set to 1.3 and 8 respectively. The entire training process is optimized using the SGD optimizer.
[0118] As Figure 2 shown, the loss is calculated between the final probability output of the student network and the soft label generated by the teacher model. In addition, for each horizontal convolutional layer in the student network, the predicted probability of this layer is obtained by adding an additional classifier (vertical convolutional layer and vertical fully connected layer), and the loss is calculated between the predicted probability of this layer and the predicted probability of each previous horizontal convolutional layer. Since the information contained in the later horizontal convolutional layer is more abundant than that in the previous horizontal convolutional layer, the information of the later horizontal convolutional layer is transmitted to the previous horizontal convolutional layer in this way, realizing the distillation of the student network by itself. And by transmitting information through the fully connected layer, the transmission efficiency of high-level information can be improved, that is, re-distillation is achieved.
[0119] The facial acne classification module is used to obtain the facial acne image to be classified and input it into the student network, and the student network outputs the classification result.
[0120] Embodiment 3
[0121] A computer device includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor is caused to execute the steps of the facial acne classification method based on soft labels.
[0122] Among them, the computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can perform human-computer interaction with the user through means such as a keyboard, a mouse, a remote control, a touchpad, or a voice control device.
[0123] The memory at least includes one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or D-interface display memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory may also be an external storage device of the computer device, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device. Of course, the memory may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the memory is commonly used to store the operating system and various application software installed on the computer device, such as the program code of the soft-label-based facial acne classification method. In addition, the memory can also be used to temporarily store various data that have been output or will be output.
[0124] In some embodiments, the processor may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor is generally used to control the overall operation of the computer device. In this embodiment, the processor is used to run the program code stored in the memory or process data, such as running the program code of the soft-label-based facial acne classification method.
[0125] Embodiment 4
[0126] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to execute the steps of the soft-label-based facial acne classification method.
[0127] Wherein, the computer-readable storage medium stores an interface display program, and the interface display program can be executed by at least one processor to cause the at least one processor to execute the steps of the soft-label-based facial acne classification method as described above.
[0128] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server or network device, etc.) to execute the soft-label-based facial acne classification method described in the embodiments of the present application.
Claims
1. A facial acne classification method based on soft labels, characterized in that: The following steps are involved: Step S1, obtaining facial acne sample data; Acquire facial acne sample images, where the facial acne sample images include facial acne sample images with hard labels and facial acne sample images without labels; Step S2, building a facial acne classification model; Build a facial acne classification model, which includes a teacher network and a student network; Step S3, training a facial acne classification model; The teacher network in step S2 is pre-trained using the facial acne sample images with hard labels in step S1, and the unlabeled facial acne sample images in step S1 are input into the pre-trained teacher network, and the teacher network outputs soft labels; the student network in step S2 is trained using the unlabeled facial acne sample images in step S1 and the soft labels output by the teacher network; When training a facial acne classification model, the loss function is: in, represents the target category loss function, represents the non-target category loss function, represents the predicted probability of the target category by the teacher network, represents the predicted probability of the target category by the student network, represents the predicted probability of the teacher network for non-target categories, represents the predicted probability of the student network for non-target categories, represents the dynamic weight vector used to emphasize the feature, represents the difference between the teacher network and the student network on non-target categories, Indicates categories, represents the total number of categories, represents the target category, Represents a dynamic weight vector The i-th component of Representation vector The i-th component of Representation vector The i-th component of represents the independent target category probability corresponding to the teacher network, represents the probability of independent target categories corresponding to the student network; Step S4, facial acne classification; Obtain the facial acne image to be classified and input it into the student network, and the student network outputs the classification result; In step S3, the weight Each component of is expressed as: in, Indicates categories, represents the output of the teacher network The value on the target class, represents the output of the teacher network In the values on non-target classes, Represents the adjustment factor, which is used to scale the dynamic weight to an appropriate range. Represents the initial value of the dynamic weight to improve the stability of the method.
2. A facial acne classification method based on soft labels as claimed in claim 1, characterized in that: In step S1, the facial acne sample image is preprocessed, and the size of the facial acne sample image is adjusted to 224*224*3.
3. A facial acne classification method based on soft labels as claimed in claim 1, characterized in that: The structures of the teacher network and the student network are the same, both of which include a first horizontal convolutional layer, a second horizontal convolutional layer, a third horizontal convolutional layer and a horizontal fully connected layer arranged in sequence. The output of the first horizontal convolutional layer is used as the input of the second horizontal convolutional layer and the first vertical convolutional layer, and the output of the first vertical convolutional layer is used as the input of the first vertical fully connected layer; the output of the second horizontal convolutional layer is used as the input of the third horizontal convolutional layer and the second vertical convolutional layer, and the output of the second vertical convolutional layer is used as the input of the second vertical fully connected layer; the output of the third horizontal convolutional layer is used as the input of the horizontal fully connected layer and the third vertical convolutional layer, and the output of the third vertical convolutional layer is used as the input of the third vertical fully connected layer.
4. A facial acne classification method based on soft labels as claimed in claim 3, characterized in that: In step S3, when training the facial acne classification model, the output of each fully connected layer is compared with the output of each previous fully connected layer using the loss function Calculate the loss.
5. The facial acne classification method based on soft labels according to claim 1, characterized in that: In step S3, during training, the learning rate is set to 0.
001. Set to 0.0001, batch size to 32, parameters , They are set to 1.3 and 8 respectively, and the entire training process is optimized using the SGD optimizer.
6. A facial acne classification system based on soft labels, characterized in that: include: A facial acne sample data acquisition module is used to acquire facial acne sample images, wherein the facial acne sample images include facial acne sample images with hard labels and facial acne sample images without labels; Facial acne classification model building module, used to build a facial acne classification model, which includes a teacher network and a student network; A facial acne classification model training module is used to pre-train the teacher network in the facial acne classification model building module using the facial acne sample images with hard labels in the facial acne sample data acquisition module, and input the pre-trained teacher network using the facial acne sample images without labels in the facial acne sample data acquisition module, and the teacher network outputs soft labels; The student network in the facial acne classification model building module is trained using the unlabeled facial acne sample images in the facial acne sample data acquisition module and the soft labels output by the teacher network; When training a facial acne classification model, the loss function is: in, represents the target category loss function, represents the non-target category loss function, represents the predicted probability of the target category by the teacher network, represents the predicted probability of the target category by the student network, represents the predicted probability of the teacher network for non-target categories, represents the predicted probability of the student network for non-target categories, represents the dynamic weight vector used to emphasize the feature, represents the difference between the teacher network and the student network on non-target categories, Indicates categories, represents the total number of categories, represents the target category, Represents a dynamic weight vector The i-th component of Representation vector The i-th component of Representation vector The i-th component of represents the independent target category probability corresponding to the teacher network, represents the probability of independent target categories corresponding to the student network; The facial acne classification module is used to obtain facial acne images to be classified and input them into the student network, and the student network outputs the classification results; In the facial acne classification model training module, the weight Each component of is expressed as: in, Indicates categories, represents the output of the teacher network The value on the target class, represents the output of the teacher network In the values on non-target classes, Represents the adjustment factor, which is used to scale the dynamic weight to an appropriate range. Represents the initial value of the dynamic weight to improve the stability of the method.
7. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Acne grading method, device, equipment and medium
CN116863522A
Model training method and device, equipment and storage medium
CN117392484A
Lung sound classification method and system based on knowledge distillation, terminal, and storage medium
WO2022073285A1