Prediction Method, Device, Electronic Device and Storage Medium for Acne Categories
By training multiple teacher models and distilling knowledge, multi-layer loss function is constructed to train student models, the problem of low accuracy in acne category recognition in the existing technology is solved, and higher prediction accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202111609463.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-24
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-12-24
AI Technical Summary
In the prior art, the accuracy of identifying acne category is low, and it is difficult to effectively predict the category of acne.
By obtaining image data sets of multiple categories of acne, multiple teacher models are trained, and the student model is distilled through multiple teacher models, multi-layer loss function is constructed to train the student model, and finally predict the target image based on the trained student model.
It improves the accuracy of acne category prediction, enhances the robustness and accuracy of student models, and can more effectively identify and predict acne categories.
Smart Images

Figure CN114266897B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of image processing technology, and in particular, to a method, device, electronic device, and storage medium for predicting the type of acne. Background Art
[0002] With the rapid development of mobile communication technology and the improvement of people's living standards, various intelligent terminals have been widely used in people's daily work and life, making people more and more accustomed to using software such as APPs. As a result, the demand for APPs with functions such as beauty selfies and skin detection by taking pictures has become more and more. Therefore, many users hope that such APPs can automatically analyze the acne situation on the face and propose targeted skin improvement plans according to the type of acne.
[0003] Currently, classification algorithms often adopt ensemble classification algorithms, which are a collection of neural networks, and their outputs are combined through weighted averaging or voting. However, the accuracy of recognition by ensemble classification algorithms is relatively low. Summary of the Invention
[0004] The main technical problem to be solved by the embodiments of the present application is to provide a method, device, electronic device, and storage medium for predicting the type of acne, so as to improve the accuracy of predicting the type of acne.
[0005] In a first aspect, a method for predicting the type of acne provided in the embodiments of the present application includes:
[0006] Obtain an image data set, where the image data set includes images of various types of acne;
[0007] Based on the image data set, train a plurality of preset teacher models, where the teacher models include various different network structures;
[0008] Perform knowledge distillation on a preset student model through a plurality of teacher models to train the student model and obtain a trained student model;
[0009] According to the trained student model, predict a target image containing acne to obtain the predicted acne type of the target image.
[0010] In some embodiments, training the student model includes:
[0011] Construct a multi-layer loss function and train the student model based on the multi-layer loss function.
[0012] In some embodiments, the multi-layer loss function includes at least one of a similarity loss function, a class loss function, and a cross-entropy loss function.
[0013] In some embodiments, the multi-layer loss function is:
[0014]
[0015] Among them, Loss is the multi-layer loss function, and L l1-sim is the similarity loss function, and L KD is the class loss function, and L s is the cross-entropy loss function, i is the class of acne, c is the size of the feature map, is the feature map of the teacher model, is the feature map of the student model, and n is the number of acne classes, is the probability value of the i-th class of acne predicted by the teacher model, is the probability value of the i-th class of acne predicted by the student model, and y i is the true acne class.
[0016] In some embodiments, knowledge distillation is performed on a preset student model by multiple teacher models, including:
[0017] According to the multiple trained teacher models, feature extraction is performed on the images in the image dataset to determine multiple first feature maps, where each teacher model corresponds to a first feature map;
[0018] In each iteration, a second feature map is determined, and a teacher model is randomly selected to perform knowledge distillation on the student model, where the second feature map has the same size as the first feature map.
[0019] In some embodiments, training the student model based on the multi-layer loss function includes:
[0020] Performing iterative training on the student model based on the multi-layer loss function;
[0021] If the number of iterations is greater than the first number threshold, or the loss of the student model is less than the first loss threshold, stop the iterative training.
[0022] In some embodiments, the acne classes include at least one of acne, post-acne erythema, inflammatory papules, pustules, nodules, and cysts.
[0023] In a second aspect, an acne class prediction device provided by an embodiment of the present application includes:
[0024] A dataset acquisition unit for acquiring an image dataset, where the image dataset includes images of various types of acne;
[0025] A teacher model training unit for training a preset multiple teacher models based on the image dataset, where the teacher models include various different network structures;
[0026] A student model training unit, configured to perform knowledge distillation on a preset student model through multiple teacher models to train the student model and obtain a trained student model;
[0027] A pimple category prediction unit, configured to predict a target image containing pimples according to the trained student model to obtain the pimple category of the predicted target image.
[0028] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0029] A memory and one or more processors, where the one or more processors are configured to execute one or more computer programs stored in the memory, and when the one or more processors execute the one or more computer programs, the electronic device implements the method as in the first aspect.
[0030] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the method as in the first aspect.
[0031] The beneficial effects of the embodiments of the present application: Different from the prior art, a method, device, electronic device and storage medium for predicting pimple categories provided by the embodiments of the present application, the method includes: obtaining an image data set, where the image data set includes images of various types of pimples; based on the image data set, training a preset multiple teacher models, where the teacher models include various different network structures; performing knowledge distillation on a preset student model through multiple teacher models to train the student model and obtain a trained student model; predicting a target image containing pimples according to the trained student model to obtain the pimple category of the predicted target image.
[0032] On the one hand, training multiple teacher models through a data set including images of various types of pimples enables the multiple teacher models to learn the characteristics of various types of pimples; on the other hand, performing knowledge distillation on a preset student model through multiple teacher models to train the student model and obtain a trained student model enables the student model to better extract the knowledge learned in the teacher models, thereby improving the robustness and accuracy of the student model, and further enabling the present application to improve the accuracy of predicting pimple categories. Description of the Drawings
[0033] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated, and the drawings in the drawings do not constitute a proportional limitation.
[0034] Figure 1 It is a schematic diagram of the application environment of a method for predicting acne types provided by an embodiment of the present application;
[0035] Figure 2 It is a schematic flowchart of a method for predicting acne types provided by an embodiment of the present application;
[0036] Figure 3 It is a schematic diagram of a teacher model training a student model provided by an embodiment of the present application;
[0037] Figure 4 It is a schematic flowchart of the iterative training of a student model provided by an embodiment of the present application;
[0038] Figure 5 It is a schematic structural diagram of a device for predicting acne types provided by an embodiment of the present application;
[0039] Figure 6 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0040] The present application will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any form. It should be noted that those of ordinary skill in the art can make several modifications and improvements without departing from the concept of the present application. These all belong to the protection scope of the present application.
[0041] In order to make the purpose, technical solution and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0042] It should be noted that if there is no conflict, the various features in the embodiments of the present application can be combined with each other, and all are within the protection scope of the present application. In addition, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the flowchart. In addition, the terms "first", "second", "third", etc. used herein do not limit the data and execution order, but only distinguish the same items or similar items with basically the same function and role.
[0043] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used in this specification in the description of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The term "and / or" used in this specification includes any and all combinations of one or more of the related listed items.
[0044] In addition, the technical features involved in the various embodiments of this application described below can be combined with each other as long as they do not conflict with each other.
[0045] Before elaborating on this application in detail, the nouns and terms involved in the embodiments of this application are described. The nouns and terms involved in the embodiments of this application are applicable to the following explanations:
[0046] (1) A neural network, also simply referred to as neural networks (NNs) or called a connection model (Connection Model), is an algorithmic mathematical model that mimics the behavioral characteristics of animal neural networks and performs distributed parallel information processing. A neural network relies on the complexity of the system and adjusts the relationships between a large number of internal nodes to achieve the purpose of processing information. Specifically, a neural network can be composed of neural units, which can be specifically understood as a neural network with an input layer, hidden layers, and an output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the intermediate layers are all hidden layers. Among them, a neural network with many hidden layers is called a deep neural network (deep neural network, DNN). The operation of each layer in a neural network can be described by the mathematical expression y = a(W·x + b). Physically, the operation of each layer in a neural network can be understood as completing the transformation from the input space (the set of input vectors) to the output space (i.e., the row space to the column space of the matrix) through five operations on the input space. These five operations include: 1. Dimensionality increase / dimensionality reduction; 2. Magnification / reduction; 3. Rotation; 4. Translation; 5. "Bending". Among them, the operations of 1, 2, and 3 are completed by "W·x", the operation of 4 is completed by "+b", and the operation of 5 is achieved by "a()". The reason for using the word "space" here is that the objects to be classified are not individual things but a class of things. Space refers to the set of all individuals of this class of things. Among them, W is the weight matrix of each layer of the neural network, and each value in this matrix represents the weight value of a neuron in this layer. This matrix W determines the space transformation from the input space to the output space described above, that is, W of each layer of the neural network controls how to transform the space. The purpose of training a neural network, that is, ultimately obtaining the weight matrices of all layers of the trained neural network. Therefore, the training process of a neural network is essentially the process of learning the way to control space transformation, and more specifically, learning the weight matrix.
[0047] It should be noted that in the embodiments of the present application, based on the models adopted for machine learning tasks, they are essentially neural networks. Common components in neural networks include convolutional layers, pooling layers, normalization layers, transposed convolutional layers, etc. By assembling these common components in the neural network, a model is designed. When the model parameters (weight matrices of each layer) are determined such that the model error meets the preset conditions or the number of adjusted model parameters reaches the preset threshold, the model converges.
[0048] Among them, the convolutional layer is configured with multiple convolutional kernels, and each convolutional kernel is set with a corresponding stride to perform convolutional operations on the image. The purpose of convolutional operations is to extract different features of the input image. The first convolutional layer may only be able to extract some low-level features such as edges, lines, and corners, etc. Deeper convolutional layers can iteratively extract more complex features from the low-level features.
[0049] The transposed convolutional layer is used to map a low-dimensional space to a high-dimensional space while maintaining their connection relationships / patterns (here the connection relationship refers to the connection relationship during convolution). The transposed convolutional layer is configured with multiple convolutional kernels, and each convolutional kernel is set with a corresponding stride to perform transposed convolutional operations on the image. Generally, in the framework library used to design neural networks (such as the PyTorch library), there is a built-in upsample() function. By calling this upsample() function, the low-dimensional to high-dimensional space mapping can be achieved.
[0050] The pooling layer (pooling) mimics the human visual system and can reduce the dimension of data or represent the image with higher-level features. Common operations of the pooling layer include max pooling, average pooling, stochastic pooling, median pooling, and combined pooling, etc. Generally speaking, pooling layers are periodically inserted between the convolutional layers of a neural network to achieve dimension reduction.
[0051] The normalization layer is used to perform normalization operations on all neurons in the intermediate layer to prevent gradient explosion and gradient disappearance.
[0052] (2) Loss function refers to a function that maps the values of a random event or its associated random variables to non - negative real numbers to represent the "risk" or "loss" of the random event. A loss function is a non - negative real - valued function used to quantify the difference between the predicted label and the true label in model prediction. In applications, the loss function is usually associated with the learning criterion and the optimization problem, that is, the model is solved and evaluated by minimizing the loss function. For example, it is used in parametric estimation in statistics and machine learning. During the process of training a neural network, since we hope that the output of the neural network is as close as possible to the value we really want to predict, we can compare the predicted value of the current network with the real target value, and then update the weight matrix of each layer of the neural network according to the difference between them. (Of course, there is usually an initialization process before the first update, that is, parameters are pre - configured for each layer in the neural network). For example, if the predicted value of the network is too high, we adjust the weight matrix to make it predict lower, and keep adjusting until the neural network can predict the real target value. Therefore, it is necessary to pre - define "how to compare the difference between the predicted value and the target value", which is the loss function or the objective function. They are important equations for measuring the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then the training of the neural network becomes a process of minimizing this loss as much as possible.
[0053] (3) Knowledge Distillation, i.e., (KD), refers to a model compression method and a training method based on the "teacher - student network concept". It extracts the knowledge contained in a trained model into another model. By introducing a soft - target related to the teacher network (a complex but superior - performing inference network) as part of the total loss, the training of the student network (a simple and low - complexity network) is induced to achieve knowledge transfer. That is, first train another more complex (usually an ensemble of multiple networks) teacher network (Teacher Network), and use the output of the large network as the soft - target to train the student network (Student Network).
[0054] The technical solution of the present application will be specifically described below with reference to the accompanying drawings of the specification.
[0055] Please refer to Figure 1 , Figure 1It is a schematic diagram of the application environment of a method for predicting acne categories provided by an embodiment of the present application;
[0056] As Figure 1 shown, the application environment 100 includes: an electronic device 101 and a server 102, and the electronic device 101 and the server 102 communicate through a wired or wireless communication method.
[0057] Among them, the electronic device 101 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. A client can be provided in the electronic device 101, and the client can be a video client, a browser client, an online shopping client, an instant messaging client, etc., and the present application does not limit the type of the client.
[0058] The electronic device 101 and the server 102 can be directly or indirectly connected through a wired or wireless communication method, and the present application does not limit this. The electronic device 101 can obtain a target image, predict the acne category of the target image, and display the target image and its acne category on a visualization interface. Among them, the target image can be an image stored in the memory of the electronic device 101 or an image received from other devices.
[0059] Alternatively, the electronic device 101 can receive the acne category of the target image sent by the server 102, and display the target image and its acne category on a visualization interface. The user can browse the face images stored in the electronic device, and trigger a prediction instruction for the acne category of the face image by triggering the acne category prediction button corresponding to any face image. The electronic device can respond to the acne category prediction instruction, obtain a face image through an image acquisition device, and use the face image as the target image. Among them, the image acquisition device can be built into the electronic device 101 or externally connected to the electronic device 101, and the present application does not limit this.
[0060] The electronic device 101 can send the target image to the server 102, and receive the acne category prediction value of the target image returned by the server 102, and then display the target image and its acne category prediction value on a visualization interface so that the user can understand the prediction result of the target image.
[0061] It can be understood that the electronic device 101 can generally refer to one of multiple electronic devices, and only the electronic device 101 is used as an example in the embodiment of the present application. Those skilled in the art can know that the number of the above-mentioned electronic devices can be more or less. For example, the above-mentioned electronic device can be only one, or the above-mentioned electronic devices can be dozens or hundreds, or more in number, and the embodiment of the present application does not limit the number and type of the electronic devices.
[0062] Among them, the server 102 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0063] The server 102 and the electronic device 101 can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here. The server 102 can maintain a face image database for storing multiple face images. The server 102 can receive the acne category prediction instruction and the target image sent by the electronic device 101, and based on the acne category prediction instruction and the target image, predict the acne category of the target image to obtain the acne category prediction value of the target image, and then send the acne category prediction value of the target image to the electronic device 101.
[0064] It can be understood that the number of the above-mentioned servers 102 can be more or less, and the embodiments of this application do not limit this. Of course, the server 102 can also include other functional servers to provide more comprehensive and diverse services.
[0065] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a method for predicting acne categories provided by an embodiment of this application;
[0066] Among them, the method for predicting acne categories is applied to an electronic device. Specifically, the execution subject of the method for predicting acne categories is one or more processors of the electronic device.
[0067] As Figure 2 shown, the method for predicting acne categories includes:
[0068] Step S201: Obtain an image data set, where the image data set includes images of various types of acne;
[0069] Specifically, the image dataset consists of face images of various types of acne, that is, the image dataset includes an acne category dataset. For example, each image in the image dataset includes a face, and each image is a three-channel color image. By collecting multiple acne images, where each acne image is labeled with a category label, and the category label is used to characterize the acne category of the acne image. In some embodiments, the acne categories include comedones, post-acne erythema, inflammatory papules, pustules, nodules, and cysts, a total of six categories. The acne images are labeled with category labels through the one-hot category label algorithm. For example, [0, 1, 0, 0, 0, 0] indicates that the acne category is post-acne erythema.
[0070] Further, in order to provide additional intra-class and inter-class relationships, the embodiment of the present application also processes the category labels of the acne categories to obtain soft labels. For example, [0, 1, 0, 0, 0, 0] indicates that the acne category is post-acne erythema. After processing, the obtained soft label is [0, 1, 0.8, 0.02, 0.04, 0.03, 0]. It can be understood that the position corresponding to each acne category in the soft label represents the probability value of that category.
[0071] In the embodiment of the present application, since the acne images are small, it is necessary to perform a normalization operation on the acne images. Specifically, the size of the acne images is adjusted to a preset resolution. For example, the resolution of the acne images is adjusted to 40 * 40. The size of the acne images in the embodiment of the present application can also be other resolutions, which are not limited herein.
[0072] It can be understood that the image dataset can be a color ID photo or a color selfie collected by an image acquisition device, etc. It can also be understood that the image dataset can also be data in an existing open-source face library. Among them, the open-source face library can be the FERET face database, the CMU Multi-PIE face database, or the YALE face database, etc. Here, the source of the image samples is not limited, as long as the image is a color image including a face and acne, such as an RGB format face image.
[0073] Step S202: Based on the image dataset, train a plurality of preset teacher models, where the teacher models include various different network structures;
[0074] Specifically, multiple teacher models are preset. Among them, the teacher models include a variety of different network structures, such as at least one of Densenet network structure, Googlenet network structure, Resnet network structure, and VGG network structure. It can be understood that each network structure corresponds to a network model, that is, the Densenet network structure corresponds to the Densenet model, the Googlenet network structure corresponds to the Googlenet model, the Resnet network structure corresponds to the Resnet model, and the VGG network structure corresponds to the VGG model.
[0075] In the embodiment of the present application, in order to enable the student model to learn the features of various types of acne learned by multiple different teacher models, further, each teacher model in the multiple teacher models is set to include at least one of the Densenet model, Googlenet model, Resnet model, and VGG model, and the network structures of each teacher model are different to avoid duplicate teacher models, thereby accelerating the training of the student model and improving the training efficiency.
[0076] In the embodiment of the present application, training multiple teacher models includes:
[0077] Construct a multi-class cross-entropy loss function for each teacher model, and perform model training on each teacher model to enable each teacher model to predict the probability values of each acne category. Among them, the multi-class cross-entropy loss function is used to multiply the true label category by the predicted label category to obtain the corresponding loss. For example, the true label y = [0, 1, 0, 0, 0, 0], the predicted category p = [0.1, 0.8, 0.1, 0.0, 0.0], and the finally calculated loss loss = -y * logp.
[0078] In the embodiment of the present application, the training of the teacher model is optimized for parameters through the Adam algorithm.
[0079] Step S203: Perform knowledge distillation on the preset student model through multiple teacher models to train the student model and obtain the trained student model;
[0080] Specifically, performing knowledge distillation on the preset student model through multiple teacher models includes:
[0081] According to the trained multiple teacher models, perform feature extraction on the images in the image dataset to determine multiple first feature maps, where each teacher model corresponds to a first feature map;
[0082] In each iteration, determine the second feature map, and randomly select a teacher model to perform knowledge distillation on the student model, where the second feature map has the same size as the first feature map.
[0083] For example, assume that the input is a pimple image of size 40*40. In each iteration, a teacher model is randomly selected. The teacher model is at least one of the Densenet model, the Googlenet model, the Resnet model, and the VGG model. The teacher model performs feature extraction to obtain a first feature map. For example, the size of the first feature map is 20*20, 10*10, or 5*5. The student model also needs to obtain a second feature map of the same size as the first feature map. At this time, whether it is the teacher model or the student model, the same number of convolutional kernels are used on the feature maps of the corresponding size, so both the teacher model and the student model output feature maps of the same size and corresponding scale. For example, the feature maps output by the teacher model and the student model are all of sizes 20*20*32, 10*10*64, and 5*5*128.
[0084] Specifically, training the student model includes:
[0085] Constructing a multi-layer loss function and training the student model based on the multi-layer loss function.
[0086] Among them, the multi-layer loss function includes at least one of a similarity loss function, a class loss function, and a cross-entropy loss function.
[0087] Specifically, the multi-layer loss function is:
[0088]
[0089] Among them, Loss is the multi-layer loss function, L l1-sim is the similarity loss function, L KD is the class loss function, L s is the cross-entropy loss function, i is the class of the pimple, c is the size of the feature map, is the feature map of the teacher model, is the feature map of the student model, n is the number of pimple classes, is the probability value of the teacher model predicting the pimple of the i-th class, is the probability value of the student model predicting the pimple of the i-th class, y i is the true pimple class.
[0090] For example: the size c of the feature map = {20, 10, 5}, n is the number of pimple classes, that is, the number of classes, and respectively represent the feature maps of the same size in the teacher model and the student model, and respectively represent the probability values of the teacher model and the student model predicting the pimple of the i-th class, y iIt is a real acne category, which is represented by the one-hot category label algorithm.
[0091] Specifically, please refer to Figure 3 , Figure 3 which is a schematic diagram of a teacher model training a student model provided by an embodiment of the present application;
[0092] As Figure 3 shown, by inputting acne images into multiple teacher models and a student model. For example, input the same acne image into teacher model A (Teacher_A) - teacher model N (Teacher_N) respectively, and input another acne image into the student model (Student), where the sizes of the two acne images are the same.
[0093] By calculating the feature maps output by each teacher model and the student model, a similarity loss L is obtained l1-sim , where the similarity loss is used to subtract the position and size of the corresponding feature maps of the teacher model and the student model, so that the features learned by the feature maps of the corresponding size of the student model are similar to the features learned by the feature maps of the corresponding size of the teacher model.
[0094] At the same time, in order to better utilize the soft labels output by the teacher model, the KD loss function is used for the category loss function of the student model training to calculate the KD loss L KD , so that the model is optimized in the direction of minimizing the KL loss value during training to better learn the similar features between acne categories, where the KD loss
[0095] Furthermore, due to the misclassification of the soft labels output by the teacher model, therefore, a cross-entropy loss function is added to the multi-layer loss function to calculate the cross-entropy loss L s , where the cross-entropy loss is the loss between the result predicted by the student model and the true label.
[0096] It can be understood that the model structure of the student model in the embodiments of the present application can be any network model, such as at least one of the Densenet model, Googlenet model, Resnet model, VGG model, and Mobilenet model. The ultimate goal of different network models is to predict the acne category. However, due to different learned features, in order to meet the model size, detection speed, and classification accuracy shown in the public image classification data ImageNet, in the embodiments of the present application, the model structure of the student model is preferably the Mobilenet model, so that the model size of the student model is appropriate, the detection speed is fast, and the classification accuracy is high, which is beneficial to better predicting the acne category.
[0097] Step S204: Predict the target image containing acne according to the trained student model to obtain the acne category of the predicted target image.
[0098] Specifically, after the student model is trained, the trained student model is called to predict the target image containing acne to obtain the acne category of the predicted target image.
[0099] In the embodiments of the present application, multiple teacher models are used to perform knowledge distillation on the student model to train the student model and obtain the trained student model, so that the student model can better extract the knowledge learned in the teacher model, thereby improving the robustness and accuracy of the student model, and further enabling the present application to improve the accuracy of predicting the acne category.
[0100] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of the iterative training of a student model provided by the embodiments of the present application;
[0101] As Figure 4 shown, the process of the iterative training of the student model includes:
[0102] Step S401: Construct multiple teacher models and a student model;
[0103] Specifically, the network structure of each teacher model among the multiple teacher models is different. For example, each teacher model is one of the Densenet model, Googlenet model, Resnet model, and VGG model. The network structure of the student model is the Mobilenet model.
[0104] Step S402: Construct a multi-layer loss function;
[0105] Specifically, the multi-layer loss function includes a similarity loss function, a category loss function, and a cross-entropy loss function. For example, the multi-layer loss function is:
[0106]
[0107] Among them, Loss is the multi-layer loss function, and L l1-sim is the similarity loss function, and L KD is the class loss function, and L s is the cross-entropy loss function. i is the class of acne, c is the size of the feature map, is the feature map of the teacher model, is the feature map of the student model, and n is the number of acne classes, is the probability value of the i-th class of acne predicted by the teacher model, is the probability value of the i-th class of acne predicted by the student model, and y i is the true acne class.
[0108] Step S403: Iteratively train the student model based on the multi-layer loss function;
[0109] Specifically, in each training, a teacher model is randomly selected to guide the training of the student model, so that the student model can learn the feature extraction method and feature fusion method of the teacher model. For example: the teacher model is the teacher network, and the student model is the student network. The loss or intermediate features of the teacher network are used to constrain the loss and intermediate features of the student network, so as to learn the feature extraction method and feature fusion method of the teacher network. It can be understood that both the teacher model and the student model are generative networks. For example: deep learning image segmentation networks, such as: UNet network, which are used to generate prediction results of acne classes for different input target images through loss functions or constraint conditions.
[0110] In the embodiments of the present application, by using the differences in the neural network structures of different teacher models and using the combination of different outputs as supervision to guide the training of the student model, it is possible to better extract the knowledge learned in the teacher model under the premise of ensuring the model size and prediction speed, thereby improving the robustness and accuracy of the student model.
[0111] Specifically, knowledge distillation is performed on the student model through the teacher model, including:
[0112] Calculate the distance between the feature layers in the teacher model and the feature layers in the student model, and use the calculated distance as the loss function to train the student model, so that the feature layers in the student model are close to the feature layers in the teacher model.
[0113] Knowledge distillation is performed on the student model through the teacher model, enabling the student model to learn the feature extraction method and feature fusion method of the teacher model, which can improve the network expression ability of the student model and is conducive to improving the accuracy of predicting acne categories.
[0114] Step S404: Whether the number of iterations is greater than the first number threshold;
[0115] Specifically, in the embodiment of the present application, the Adam algorithm (Adaptive Moment Estimation Algorithm) is used to optimize the model parameters. For example: the number of iterations is set to 500 times, the initial learning rate is set to 0.001, the weight decay is set to 0.0005, and every 50 iterations, the learning rate decays to 1 / 10 of the original.
[0116] It can be understood that the Adam algorithm (Adaptive Moment Estimation Algorithm) can be regarded as the combination of the momentum method and the RMSprop algorithm. It not only uses momentum as the parameter update direction but also can adaptively adjust the learning rate.
[0117] Specifically, it is judged whether the number of iterations is greater than the first number threshold, which is preset, for example: set to 500 times. If the number of iterations is greater than the first number threshold, then go to step S406: Training is completed; if the number of iterations is not greater than the first number threshold, then go to step S405: Whether the loss of the student model is less than the first loss threshold.
[0118] It can be understood that the first number threshold is specifically set according to specific needs and is not limited here.
[0119] Step S405: Whether the loss of the student model is less than the first loss threshold;
[0120] Specifically, it is judged whether the loss of the student model is less than the first loss threshold. If so, then go to step S406: Training is completed; if not, then return to step S403: Iteratively train the student model based on the multi-layer loss function.
[0121] In the embodiment of the present application, judging whether the loss of the student model is less than the first loss threshold, that is, judging whether the loss calculated by the multi-layer loss function is less than the first loss threshold, so as to determine whether to end the iterative process in advance when the number of iterations is less than the first number threshold, that is, stop iterative training to quickly obtain the trained student model. Among them, the multi-layer loss function is:
[0122]
[0123] In the embodiments of the present application, the first loss threshold can be set to 0.0005, 0.001. It can be understood that the first loss threshold is specifically set according to specific needs and is not limited herein.
[0124] Step S406: Training is completed;
[0125] Specifically, after the training is completed, a trained student model is obtained. At this time, the trained student model can be called to predict the acne category of the target image.
[0126] In the embodiments of the present application, by providing a method for predicting acne categories, the method includes: obtaining an image data set, where the image data set includes images of various categories of acne; based on the image data set, training a plurality of preset teacher models, where the teacher models include various different network structures; performing knowledge distillation on a preset student model through the plurality of teacher models to train the student model and obtain a trained student model; and predicting a target image containing acne according to the trained student model to obtain the predicted acne category of the target image.
[0127] On the one hand, training a plurality of teacher models through a data set including images of various categories of acne enables the plurality of teacher models to learn the characteristics of various categories of acne; on the other hand, performing knowledge distillation on a preset student model through the plurality of teacher models to train the student model and obtain a trained student model enables the student model to better extract the knowledge learned in the teacher models, thereby improving the robustness and accuracy of the student model, and further enabling the present application to increase the probability of predicting acne categories.
[0128] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a device for predicting acne categories provided by the embodiments of the present application;
[0129] Among them, the device for predicting acne categories is applied to an electronic device. Specifically, the device for predicting acne categories is applied to one or more processors of the electronic device.
[0130] As Figure 5 shown, the device 50 for predicting acne categories includes:
[0131] A data set acquisition unit 501, configured to obtain an image data set, where the image data set includes images of various categories of acne;
[0132] A teacher model training unit 502, configured to train a plurality of preset teacher models based on the image data set, where the teacher models include various different network structures;
[0133] A student model training unit 503, configured to perform knowledge distillation on a preset student model through multiple teacher models to train the student model and obtain a trained student model;
[0134] A pimple category prediction unit 504, configured to predict a target image containing pimples according to the trained student model to obtain the pimple category of the predicted target image.
[0135] In the embodiments of the present application, the pimple category prediction device may also be built by hardware devices. For example, the pimple category prediction device may be built by one or more than two chips, and each chip may work in coordination to complete the pimple category prediction method described in each of the above embodiments. For another example, the pimple category prediction device may also be built by various logic devices, such as a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a single-chip microcomputer, an ARM processor (Advanced RISC Machines, ARM), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of these components.
[0136] The pimple category prediction device in the embodiments of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.
[0137] The pimple category prediction device in the embodiments of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.
[0138] The acne category prediction device provided by the embodiments of this application can achieve Figure 2 For the various processes achieved, to avoid repetition, they will not be elaborated here.
[0139] It should be noted that the above acne category prediction device can execute the acne category prediction method provided by the embodiments of this application, and has the corresponding functional modules and beneficial effects for executing the method. For the technical details not described in detail in the embodiments of the acne category prediction device, reference can be made to the acne category prediction method provided by the embodiments of this application.
[0140] In the embodiments of this application, by providing an acne category prediction device, including: a dataset acquisition unit for acquiring an image dataset, where the image dataset includes images of various categories of acne; a teacher model training unit for training a plurality of preset teacher models based on the image dataset, where the teacher models include various different network structures; a student model training unit for performing knowledge distillation on a preset student model through the plurality of teacher models to train the student model and obtain a trained student model; an acne category prediction unit for predicting the acne category of a target image containing acne according to the trained student model to obtain the predicted acne category of the target image.
[0141] On the one hand, by training a plurality of teacher models with a dataset including images of various categories of acne, the plurality of teacher models can learn the characteristics of various categories of acne; on the other hand, by performing knowledge distillation on the student model through the plurality of teacher models to train the student model and obtain a trained student model, the student model can better extract the knowledge learned in the teacher models, thereby improving the robustness and accuracy of the student model, and further enabling this application to increase the probability of predicting the acne category.
[0142] The embodiments of this application also provide an electronic device. Please refer to Figure 6 , Figure 6 which is a schematic hardware structure diagram of an electronic device provided by the embodiments of this application;
[0143] As Figure 6 shown, the electronic device 60 includes at least one processor 601 and a memory 602 connected communicatively ( Figure 6 taking bus connection and one processor as an example).
[0144] Among them, the processor 601 is used to provide computing and control capabilities to control the electronic device 60 to perform corresponding tasks. For example, it controls the electronic device 60 to execute the acne category prediction method in any of the above method embodiments, including: obtaining an image data set, where the image data set includes images of various categories of acne; based on the image data set, training a plurality of preset teacher models, where the teacher models include various different network structures; performing knowledge distillation on a preset student model through the plurality of teacher models to train the student model and obtain a trained student model; and predicting a target image containing acne according to the trained student model to obtain the acne category of the predicted target image.
[0145] On the one hand, training a plurality of teacher models through a data set including images of various categories of acne enables the plurality of teacher models to learn the characteristics of various categories of acne; on the other hand, performing knowledge distillation on the student model through the plurality of teacher models to train the student model and obtain a trained student model enables the student model to better extract the knowledge learned in the teacher models, thereby improving the robustness and accuracy of the student model, and further enabling this application to increase the probability of predicting the acne category.
[0146] The processor 601 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0147] The memory 602, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the acne category prediction method in the embodiments of the present application. By running the non-transitory software programs, instructions, and modules stored in the memory 602, the processor 601 can implement the acne category prediction method in any of the following method embodiments. Specifically, the memory 602 may include volatile memory (VM), such as random access memory (RAM); the memory 602 may also include non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD), or other non-transitory solid-state storage devices; the memory 502 may further include a combination of the above types of memories.
[0148] In the embodiments of the present application, the memory 602 may further include memories remotely provided with respect to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.
[0149] In the embodiments of the present application, the electronic device 60 may further have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. The electronic device 60 may further include other components for implementing the functions of the device, which will not be elaborated herein.
[0150] The embodiments of the present application further provide a computer-readable storage medium, such as a memory including program code, and the above program code can be executed by a processor to complete the acne category prediction method in the above embodiments. For example, the computer-readable storage medium may be a read-only memory (ROM), random access memory (RAM), compact disc read-only memory (CDROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0151] The embodiments of the present application also provide a computer program product, which includes one or more pieces of program code stored in a computer-readable storage medium. The processor of the electronic device reads the program code from the computer-readable storage medium, and the processor executes the program code to complete the method steps of the acne category prediction method provided in the above embodiments.
[0152] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by hardware related to program code. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disk, etc.
[0153] Through the description of the above embodiments, those of ordinary skill in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; under the idea of the present application, the technical features in the above embodiments or different embodiments can also be combined, and the steps can be implemented in any order, and there are many other changes in different aspects of the present application as described above. For the sake of brevity, they are not provided in detail; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for predicting the type of acne, characterized in that, including: Obtain an image dataset, where the image dataset includes images of various types of acne; Based on the image dataset, train a plurality of preset teacher models, where the teacher models include various different network structures; Perform knowledge distillation on a preset student model through the plurality of teacher models to train the student model and obtain a trained student model; According to the trained student model, predict a target image containing acne to obtain the acne category of the predicted target image; Training the student model includes: Construct a multi-layer loss function and train the student model based on the multi-layer loss function; The performing knowledge distillation on a preset student model through the plurality of teacher models includes: According to the trained plurality of teacher models, extract features from the images in the image dataset to determine a plurality of first feature maps, where each teacher model corresponds to one first feature map; In each iteration, determine a second feature map with the same size as the first feature map, and randomly select a teacher model to perform knowledge distillation on the student model, where the second feature map is obtained by the student model.
2. The method according to claim 1, characterized in that, The multi-layer loss function includes at least one of a similarity loss function, a class loss function, and a cross-entropy loss function.
3. The method according to claim 2, characterized in that, The multi-layer loss function is: Among them, Loss is the multi-layer loss function, and L l1-sim is the similarity loss function, and L KD is the class loss function, and L s is the cross-entropy loss function, i is the class of acne, c is the size of the feature map, is the feature map of the teacher model, is the feature map of the student model, n is the number of acne classes, is the probability value of the i-th class of acne predicted by the teacher model, is the probability value of the i-th class of acne predicted by the student model, and y i is the true acne class.
4. The method according to claim 1, characterized in that, The training the student model based on the multi-layer loss function includes: Iteratively train the student model based on the multi-layer loss function; If the number of iterations is greater than a first number threshold, or the loss of the student model is less than a first loss threshold, stop the iterative training.
5. The method according to any one of claims 1 - 4, characterized in that, The acne category includes at least one of comedones, post-acne erythema, inflammatory papules, pustules, nodules, and cysts.
6. A device for predicting the type of acne, characterized in that, including: A dataset acquisition unit for obtaining an image dataset, where the image dataset includes images of various types of acne; A teacher model training unit for training a plurality of preset teacher models based on the image dataset, where the teacher models include various different network structures; A student model training unit for performing knowledge distillation on a preset student model through the plurality of teacher models to train the student model and obtain a trained student model; training the student model includes: constructing a multi-layer loss function and training the student model based on the multi-layer loss function; the performing knowledge distillation on a preset student model through the plurality of teacher models includes: according to the trained plurality of teacher models, extracting features from the images in the image dataset to determine a plurality of first feature maps, where each teacher model corresponds to one first feature map; in each iteration, determining a second feature map with the same size as the first feature map, and randomly selecting a teacher model to perform knowledge distillation on the student model, where the second feature map is obtained by the student model; An acne category prediction unit for predicting a target image containing acne according to the trained student model to obtain the acne category of the predicted target image.
7. An electronic device, characterized in that, including: A memory and one or more processors, the one or more processors being configured to execute one or more computer programs stored in the memory, and when the one or more processors execute the one or more computer programs, causing the electronic device to implement the method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program including program instructions that, when executed by a processor, cause the processor to execute the method according to any one of claims 1-5.
Citation Information
Patent Citations
Decentralized artificial intelligence (AI) / machine learning training system
US20220344049A1