Method, apparatus, device and storage medium for image multi-label classification
A multi-label classification model with multiple activation layers and optional multi-head attention addresses the challenge of classifying abstract concepts by simplifying network structure and improving accuracy in multi-label image classification.
Patent Information
- Application Number
- CN202111574142.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-21
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-12-21
AI Technical Summary
It is difficult for the prior art to effectively classify images in multi-labels, especially for abstract concept images, detection and semantic segmentation models are difficult to define detection target boxes and semantic segmentation masks, and the existing model network structure is large and resource consumption is high.
A multi-label classification model with multiple activation layers is adopted, and a multi-head attention mechanism is added to the model to build a lightweight network structure, and the model is trained through the training set to improve classification accuracy.
It realizes rapid convergence and efficient training of lightweight network structures, improves the accuracy and service performance of multi-label classification, and can complete multi-label classification of images faster.
Smart Images

Figure CN114443877B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of image processing, and particularly relates to a method, device, equipment and storage medium for multi-label classification of images. Background Art
[0002] Currently, there are a vast number of images on the network. Performing multi-label classification on images helps to perform structured analysis and processing on the images.
[0003] In related technologies, detection and semantic segmentation models are usually used to perform multi-label classification on images. However, this method is more suitable for the case where the object is a specific thing. For abstract concepts, it is difficult to define the detection target box and the mask of semantic segmentation. For example, it is difficult to classify the situation of whether there is light. Summary of the Invention
[0004] This application proposes a method, device, equipment and storage medium for multi-label classification of images, and uses a multi-label classification model with multiple activation layers for multi-label classification. The structure of the multi-label classification model is simple, with a small amount of computation, has a lighter network structure, can converge faster during training, improves the model training efficiency, and has higher service performance under the same resources.
[0005] The first aspect embodiment of this application proposes a method for multi-label classification of images, including:
[0006] Obtain a training set, where the training set includes sample images labeled with multiple classification labels;
[0007] Construct a network model structure for multi-label classification, where the network model structure includes multiple activation layers, and the number of activation layers is equal to the number of classification labels;
[0008] Train the network model structure according to the training set to obtain a trained multi-label classification model.
[0009] In some embodiments of this application, the constructing a network model structure for multi-label classification includes:
[0010] Based on a preset classification model, construct a backbone classifier;
[0011] Connect the backbone classifier with multiple activation layers.
[0012] In some embodiments of this application, the constructing a network model structure for multi-label classification includes:
[0013] Based on a preset classification model, construct a backbone classifier;
[0014] Connect the backbone classifier with a multi-head attention layer;
[0015] Connect the multi-head attention layer to multiple activation layers.
[0016] In some embodiments of the present application, the preset classification model includes an EfficientNet network;
[0017] Remove the normalization layer of the EfficientNet network to obtain the backbone classifier.
[0018] In some embodiments of the present application, training the network model structure according to the training set to obtain a trained multi-label classification model includes:
[0019] Obtain sample images from the training set;
[0020] Input the sample images into the backbone classifier to output feature vectors corresponding to each classification label;
[0021] Input each of the feature vectors into the multiple activation layers respectively to obtain prediction probabilities corresponding to each classification label;
[0022] Calculate the loss value of the current training cycle through a preset loss function according to the prediction probabilities corresponding to each classification label.
[0023] In some embodiments of the present application, training the network model structure according to the training set to obtain a trained multi-label classification model includes:
[0024] Obtain sample images from the training set;
[0025] Input the sample images into the backbone classifier to output feature vectors corresponding to each classification label;
[0026] Input the feature vectors corresponding to each classification label into the multi-head attention layer to output a multi-head attention matrix corresponding to each classification label;
[0027] Input each of the multi-head attention matrices into the multiple activation layers respectively to obtain prediction probabilities corresponding to each classification label;
[0028] Calculate the loss value of the current training cycle through a preset loss function according to the prediction probabilities corresponding to each classification label.
[0029] In some embodiments of the present application, the method further includes:
[0030] Obtain an image to be classified;
[0031] Classify the image to be classified through the trained multi-label classification model.
[0032] In some embodiments of the present application, classifying the image to be classified by the trained multi-label classification model includes:
[0033] Inputting the image to be classified into the trained multi-label classification model to obtain the prediction probability corresponding to each classification label;
[0034] Determining the classification labels with prediction probabilities greater than a preset threshold as the classification labels to which the image to be classified belongs.
[0035] An embodiment of the second aspect of the present application provides an apparatus for multi-label classification of images, including:
[0036] An acquisition module, configured to acquire a training set, where the training set includes sample images labeled with multiple classification labels;
[0037] A model construction module, configured to construct a network model structure for multi-label classification, where the network model structure includes multiple activation layers, and the number of activation layers is equal to the number of classification labels;
[0038] A model training module, configured to train the network model structure according to the training set to obtain a trained multi-label classification model.
[0039] An embodiment of the third aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor runs the computer program to implement the method described in the first aspect above.
[0040] An embodiment of the fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement the method described in the first aspect above.
[0041] The technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0042] In the embodiments of the present application, a multi-label classification model with multiple activation layers is used to perform multi-label classification on images. The structure of the multi-label classification model is simple and the amount of computation is small. Further, a multi-head attention mechanism is added to the multi-label classification model, so that the multi-label classification model can learn the correlation between different classification labels and improve the accuracy of multi-label classification. Both multi-label classification models provided in the present application have a lighter network structure, can converge faster during the training process, improve the model training efficiency, and the multi-label classification model trained in the present application has higher service performance under the same resources.
[0043] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. Description of the Drawings
[0044] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components.
[0045] In the drawings:
[0046] Figure 1 A flowchart of a method for image multi-label classification provided by an embodiment of the present application is shown;
[0047] Figure 2 A schematic structural diagram of a network model for multi-label classification provided by an embodiment of the present application is shown;
[0048] Figure 3 A schematic structural diagram of MBConv provided by an embodiment of the present application is shown;
[0049] Figure 4 A schematic structural diagram of the EfficientNet network provided by an embodiment of the present application is shown;
[0050] Figure 5 A schematic structural diagram of a multi-label classification model with an EfficientNet network as the backbone classifier provided by an embodiment of the present application is shown;
[0051] Figure 6 A schematic structural diagram of another network model for multi-label classification provided by an embodiment of the present application is shown;
[0052] Figure 7 A schematic structural diagram of the attention mechanism provided by an embodiment of the present application is shown;
[0053] Figure 8 A schematic structural diagram of the multi-head attention mechanism provided by an embodiment of the present application is shown;
[0054] Figure 9 A schematic structural diagram of another multi-head attention mechanism provided by an embodiment of the present application is shown;
[0055] Figure 10 A schematic structural diagram of another multi-label classification model with an EfficientNet network as the backbone classifier provided by an embodiment of the present application is shown;
[0056] Figure 11 The figure shows a schematic structural diagram of an apparatus for multi-label classification of images provided by an embodiment of the present application;
[0057] Figure 12 The figure shows a schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0058] Figure 13 The figure shows a schematic diagram of a storage medium provided by an embodiment of the present application. Detailed implementation manners
[0059] Hereinafter, exemplary embodiments of the present application will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that a more thorough understanding of the present application can be obtained and the scope of the present application can be fully conveyed to those skilled in the art.
[0060] It should be noted that unless otherwise specified, the technical terms or scientific terms used in the present application should have the ordinary meanings understood by those skilled in the art to which the present application belongs.
[0061] Hereinafter, a method, an apparatus, a device, and a storage medium for multi-label classification of images according to embodiments of the present application will be described with reference to the accompanying drawings.
[0062] Currently, there are a large number of images in the network, and it is of great significance to perform structured analysis and processing on these images. The existing image analysis and processing methods mainly focus on classification, detection, semantic segmentation, etc. In daily life, pictures are essentially multi-labeled. Image classification, as a single-label analysis method, is subject to certain limitations. Detection and semantic segmentation can solve the multi-label problem, but there is a huge problem of annotation workload, and it is more suitable for solving problems in specific scenarios. In addition, detection and segmentation are more suitable for the case where the object is a specific thing. For abstract concepts, it is difficult to define the detection target box and the semantic segmentation mask, such as whether there is light. Multi-label classification is the most applicable method for picture structured analysis and processing in general scenarios.
[0063] In the related art, the TresnetASL model is used to perform multi-label classification on images, but the network structure of this model is relatively large. In some scenarios, a lighter-weight network is used without reducing the network performance.
[0064] Based on this, the embodiments of the present application provide a method for multi-label classification of images. This method uses a multi-label classification model with multiple activation layers to perform multi-label classification on images. The structure of the multi-label classification model is simple and the amount of computation is small. Further, a multi-head attention mechanism is added to the multi-label classification model, enabling the multi-label classification model to learn the correlations between different classification labels and improving the accuracy of the multi-label classification model in performing multi-label classification.
[0065] See Figure 1 , and the method specifically includes the following steps:
[0066] Step 101: Obtain a training set, which includes sample images labeled with multiple classification labels.
[0067] Obtain a large number of sample images and label multiple classification labels in each sample image. The classification labels can include any labels to be classified, such as whether it contains a portrait, whether it is a medical staff member, whether it is a surgical procedure diagram, whether it contains a surgical site, etc. Each classification label is a binary classification problem, and specific categories are represented by different values. For example, for the classification label of whether it surrounds a medical staff member, a value of 1 indicates that the person in the picture is a medical staff member; while a value of 0 for the classification label indicates that the person in the picture is not a medical staff member.
[0068] The classification labels are determined by the specific business requirements of the multi-label classification of images, and the embodiments of the present application do not limit the specific content of the classification labels.
[0069] It should be noted that a sample image can belong to multiple classifications at the same time. Assuming that 1 represents that the sample image belongs to a certain classification, multiple classification labels in a sample image can have values of 1 at the same time.
[0070] Step 102: Construct a network model structure for multi-label classification. This network model structure includes multiple activation layers, and the number of activation layers is equal to the number of classification labels.
[0071] In some embodiments of the present application, based on a preset classification model, a backbone classifier is constructed. The backbone classifier is connected to multiple activation layers to obtain a network model structure for multi-label classification. Figure 2 Shows a schematic diagram of this network model structure. In the figure, activation layer 1, activation layer 2,..., activation layer N are schematically drawn. In actual applications, the number of activation layers is equal to the number of classification labels.
[0072] The preset classification model can be any one of the 9 efficient classification networks from EfficientNet B0 to EfficientNet B8, or any other classification network such as ResNet, vittransformer, etc. Here, the case where the preset classification model is the EfficientNet B0 network is taken as an example for illustration. Table 1 shows the main network structure of the EfficientNet B0 network. The structure between Conv3×3 and FC in the EfficientNet B0 network is used as the backbone classifier.
[0073] Table 1
[0074] Operation layer Resolution Number of channels Number of layers Conv3×3 224×224 32 1 MBConv1, k3×3 122×122 16 1 MBConv6, k3×3 122×122 24 2 MBConv6, k5×5 56×56 40 2 MBConv6, k3×3 28×28 80 3 MBConv6, k5×5 14×14 112 3 MBConv6, k5×5 14×14 192 4 MBConv6, k3×3 7×7 320 1 Conv1×1&Pooling&FC 7×7 1280 1
[0075] Among them, in the EfficientNet B0 network, MBConv comes from the InvertedResidualBlock in the MobileNetV3 network. An SE (Squeeze-and-Excitation) module is added to MBConv, and the structure of MBConv is as Figure 3 shown. The Swish activation function is used as the activation function in the EfficientNet network, and the structure of the EfficientNet network is as Figure 4 shown. The feature vector output by the fully connected layer FC is input into the softmax activation layer, and finally the prediction probability is output. From Figure 4 it can be seen that the EfficientNet network can only output the prediction probability of one classification through the softmax activation layer and cannot achieve multi-label classification.
[0076] In the embodiment of the present application, the softmax normalization layer in the EfficientNet network is removed to obtain the backbone classifier. The fully connected layer FC of the backbone classifier is respectively connected to multiple activation layers to obtain a network model structure for multi-label classification. Among them, the number of activation layers is equal to the number of classification labels, and each activation layer uses the sigmoid function as the activation function. This network model structure is as Figure 5 shown, Figure 5 where the multiple Sigmoid layers are multiple activation layers, and each activation layer outputs the prediction probability of one classification label respectively. Figure 5 In
[0077] In some other embodiments of the present application, a backbone classifier is constructed based on a preset classification model. Then, the backbone classifier is first connected to the multi-head attention layer, and then the multi-head attention layer is connected to multiple activation layers to obtain a network model structure for multi-label classification. Among them, the process of constructing the backbone classifier is the same as the construction process described above and will not be elaborated here. Figure 6 FIG. shows a schematic diagram of a network model structure including a multi-head attention layer. In the figure, activation layer 1, activation layer 2,..., activation layer N are schematically drawn. In practical applications, the number of activation layers is equal to the number of classification labels.
[0078] The attention mechanism Attention is used to calculate the "degree of relevance", which is expressed as mapping the query (Q) and key-value pairs to the output. The mapping formula is shown below. Among them, the query, each key, and each value are all vectors, and the output is the weighted sum of all values in V, where the weights are calculated from the query and each key. The structure of the attention mechanism Attention is as Figure 7 shown.
[0079] First, the similarity between Q and K is calculated through the above formula, and the obtained similarity is subjected to the Softmax operation for normalization. Finally, for the calculated weights, a weighted sum calculation is performed on all values in V to obtain the Attention vector.
[0080] The multi-head attention mechanism is based on the above attention mechanism. Q, K, and V are calculated in groups, and finally the results are concatenated. The structure of the multi-head attention mechanism is as Figure 8 and 9 shown. The formula representation of the multi-head attention mechanism is:
[0081] MultiHead(Q, K, V) = Concat(head1,..., head h )W O
[0082] Among them, Q, K, and V respectively represent the vectors of the query, key, and value, W represents the weight, and W Q , W K , W V , W O respectively represent the weight matrices of the query, key, value, and output Out, and head i represents the i-th head.
[0083] In the embodiments of the present application, a multi-head attention layer is added between the backbone classifier and multiple activation layers. Taking the EfficientNet network as an example for the backbone classifier, the constructed network model structure for multi-label classification is as Figure 10 shown. Figure 10 The multiple Sigmoid layers in Figure 10 are multiple activation layers, and each activation layer outputs the prediction probability of a classification label respectively.
[0084] By learning the correlation between different categories through the multi-head attention layer, the multi-head attention layer outputs the feature vector corresponding to each classification label, and then through multiple activation layers, the sigmoid function is used to calculate the prediction probability corresponding to each classification label according to the feature vector corresponding to each classification label respectively, realizing multi-label classification based on images and improving the accuracy of the model's multi-label classification.
[0085] Step 103: Train the network model structure according to the training set to obtain a trained multi-label classification model.
[0086] In some embodiments of the present application, the network model structure as Figure 2 shown is constructed in step 102. When training this network model structure, first obtain sample images from the training set, and the number of obtained sample images can be Figure 2 the number of sample images corresponding to the batch size of the network model structure shown. Input the obtained sample images into the backbone classifier to output the feature vector corresponding to each classification label. Input each feature vector into multiple activation layers respectively. The number of feature vectors output by the backbone classifier is equal to the number of classification labels and the number of activation layers. Input each feature vector into different activation layers respectively, and each activation layer outputs the prediction probability corresponding to the feature vector it receives, that is, the prediction probability corresponding to each classification label is obtained. According to the prediction probability corresponding to each classification label, calculate the loss value of the current training cycle through a preset loss function.
[0087] The preset loss function can be ASL (Auto Seg-Loss, automatic loss function), or any other binary classification loss function. The formula of the ASL loss function is as follows:
[0088]
[0089] where, L + is the positive sample loss value, L - is the negative sample loss value, p is the prediction probability output by the activation layer, and γ is the focusing parameter. p m= max(p - m, 0), where m is a hyperparameter used to adjust the amplitude of p adjustment.
[0090] During the above model training process, the AdamW optimizer is used. This optimizer is easy to tune parameters and can train a model with the same performance as SGD (Stochastic Gradient Descent) + Moment. The learning rate scheduler can adopt the CosineAnnealingWarmRestarts learning rate scheduling formula as shown below. The cosine annealing learning rate can enable the model to jump out of local optimal solutions, thus training a better model.
[0091]
[0092] Among them, η min is the minimum learning rate, η max is the initial learning rate, T cur is the number of epochs (training rounds) after the last learning rate reset, T i represents after how many epochs (training rounds) the learning rate is reset. When T cur = T i , set η t = η min , when the learning rate is reset and T cur = 0, set η t = η max .
[0093] The embodiments of this application do not limit which specific loss function, optimizer, and learning rate scheduler to use. The above only gives some loss functions, optimizers, and learning rate schedulers by way of example. In actual applications, appropriate loss functions, optimizers, and learning rate schedulers can be selected according to requirements.
[0094] After calculating the loss value of the current training cycle in the above manner, it is judged whether the number of currently trained cycles has reached a preset number. If so, stop training, and according to the model parameters of the training cycle with the smallest loss value in the trained cycles and Figure 2 the network model structure shown, obtain the trained multi-label classification model. If the number of currently trained cycles has not reached the preset number, continue training until the number of training times reaches the preset number, and then obtain the finally trained multi-label classification model in the above manner.
[0095] In some embodiments of this application, step 102 constructs a network model structure as Figure 6 shown. Training this network model structure, first obtain sample images from the training set. The number of obtained sample images can be Figure 6Batch size sample images corresponding to the shown network model structure. Input the obtained sample images into the backbone classifier, and output the feature vectors corresponding to each classification label. The number of feature vectors output by the backbone classifier is equal to the number of classification labels. Input the feature vectors corresponding to each classification label into the multi-head attention layer, and output the multi-head attention matrix corresponding to each classification label. The number of multi-head attention matrices output by the multi-head attention layer is equal to both the number of classification labels and the number of activation layers. Input each multi-head attention matrix into each activation layer respectively, and each activation layer outputs the prediction probability corresponding to the multi-head attention matrix it receives, that is, the prediction probability corresponding to each classification label is obtained. According to the prediction probability corresponding to each classification label, calculate the loss value of the current training cycle through a preset loss function.
[0096] The process of calculating the loss value of the current training cycle and the convergence process of model training are both the same as the above training Figure 2 The process of the model network structure shown, and train it in the above manner to obtain Figure 6 A multi-label classification model with the shown structure.
[0097] After training a multi-label classification model in the above manner, the multi-label classification model can be deployed on devices that need to provide multi-label classification services. After deploying this service, the multi-label classification model can be used to perform multi-label classification on images.
[0098] Specifically, obtain the image to be classified; classify the image to be classified through the trained multi-label classification model. Input the image to be classified into the trained multi-label classification model to obtain the prediction probability corresponding to each classification label. Determine the classification label to which the image to be classified belongs as the classification label with a prediction probability greater than the preset threshold.
[0099] If the structure of the deployed multi-label classification model is as Figure 2 shown, input the image to be classified into the backbone classifier of the multi-label classification model, and output the feature vectors corresponding to each classification label. Input each feature vector into different activation layers respectively, and each activation layer calculates the prediction probability of the classification label corresponding to its received feature vector using the sigmoid algorithm. Determine the classification label to which the image to be classified belongs as the classification label with a prediction probability greater than the preset threshold, and return the determined classification label to which the image to be classified belongs to the client that calls this service.
[0100] Figure 2 The multi-label classification model with the shown structure has a more lightweight network structure, can converge faster during training, and improves the training efficiency of the model. Moreover, using this multi-label classification model can accurately achieve multi-label classification of images.
[0101] If the structure of the deployed multi-label classification model is as Figure 6 shown, the image to be classified is input into the backbone classifier of the multi-label classification model, and the feature vectors corresponding to each classification label are output. Each feature vector is input into the multi-head attention layer, and the multi-head attention matrix corresponding to each classification label is output. Each multi-head attention matrix is respectively input into different activation layers, and each activation layer calculates the prediction probability of the classification label corresponding to its received multi-head attention matrix using the sigmoid algorithm. The classification labels with prediction probabilities greater than the preset threshold are determined as the classification labels to which the image to be classified belongs, and the determined classification labels to which the image to be classified belongs are returned to the client that calls this service.
[0102] Figure 6 The multi-label classification model with the structure shown has a smaller scale and can converge faster during training, improving the training efficiency of the model. And this multi-label classification model includes a multi-head attention layer, and during training, the multi-head attention layer learns the correlation between different classification labels. When performing multi-label classification on the image to be classified through this multi-label classification model, the multi-head attention layer applies the correlation it has learned between different classification labels to the distinction between each classification label, making the accuracy of the final multi-label classification higher.
[0103] In the embodiments of the present application, a multi-label classification model with multiple activation layers is used to perform multi-label classification on images. The structure of the multi-label classification model is simple and the amount of computation is small. Further, a multi-head attention mechanism is added to the multi-label classification model, enabling the multi-label classification model to learn the correlation between different classification labels and improving the accuracy of multi-label classification. Both of the multi-label classification models provided in the present application have a lighter network structure, can converge faster during training, improve the model training efficiency, and the multi-label classification models trained in the present application have higher service performance under the same resources.
[0104] The embodiments of the present application also provide an apparatus for image multi-label classification, which is used to execute the method for image multi-label classification provided in any of the above embodiments. As Figure 11 shown, the apparatus includes:
[0105] An acquisition module 201, configured to acquire a training set, where the training set includes sample images labeled with multiple classification labels;
[0106] A model construction module 202, configured to construct a network model structure for performing multi-label classification, where the network model structure includes multiple activation layers, and the number of activation layers is equal to the number of classification labels;
[0107] A model training module 203, configured to train the network model structure according to the training set to obtain a trained multi-label classification model.
[0108] The model construction module 202 is configured to construct a backbone classifier based on a preset classification model; connect the backbone classifier with multiple activation layers.
[0109] The model construction module 202 is configured to construct a backbone classifier based on a preset classification model; connect the backbone classifier with a multi-head attention layer; connect the multi-head attention layer with multiple activation layers.
[0110] The preset classification model includes an EfficientNet network; the model construction module 202 is configured to remove the normalization layer of the EfficientNet network to obtain a backbone classifier.
[0111] The model training module 203 is configured to obtain sample images from a training set; input the sample images into the backbone classifier to output feature vectors corresponding to each classification label; input each feature vector into multiple activation layers respectively to obtain prediction probabilities corresponding to each classification label; calculate the loss value of the current training cycle through a preset loss function according to the prediction probabilities corresponding to each classification label.
[0112] The model training module 203 is configured to obtain sample images from a training set; input the sample images into the backbone classifier to output feature vectors corresponding to each classification label; input the feature vectors corresponding to each classification label into the multi-head attention layer to output a multi-head attention matrix corresponding to each classification label; input each multi-head attention matrix into multiple activation layers respectively to obtain prediction probabilities corresponding to each classification label; calculate the loss value of the current training cycle through a preset loss function according to the prediction probabilities corresponding to each classification label.
[0113] The apparatus further includes: a classification module, configured to obtain an image to be classified; classify the image to be classified through a trained multi-label classification model.
[0114] The classification module is configured to input the image to be classified into the trained multi-label classification model to obtain prediction probabilities corresponding to each classification label; determine the classification labels whose prediction probabilities are greater than a preset threshold as the classification labels to which the image to be classified belongs.
[0115] In the embodiments of the present application, a multi-label classification model with multiple activation layers is used to perform multi-label classification on images. The structure of the multi-label classification model is simple and the amount of computation is small. Further, a multi-head attention mechanism is added to the multi-label classification model, so that the multi-label classification model can learn the correlation between different classification labels and improve the accuracy of multi-label classification. Both multi-label classification models provided in the present application have a lighter network structure, can converge faster during the training process, improve the model training efficiency, and the multi-label classification model trained in the present application has higher service performance under the same resources.
[0116] Embodiments of the present application further provide an electronic device to execute the method for multi-label classification of images described above. Please refer to Figure 12 FIG. shows a schematic diagram of an electronic device provided by some embodiments of the present application. As Figure 12 shown, the electronic device 8 includes: a processor 800, a memory 801, a bus 802, and a communication interface 803. The processor 800, the communication interface 803, and the memory 801 are connected through the bus 802. A computer program that can run on the processor 800 is stored in the memory 801. When the processor 800 runs the computer program, it executes the method for multi-label classification of images provided by any one of the foregoing embodiments of the present application.
[0117] Among them, the memory 801 may include a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 803 (which can be wired or wireless), a communication connection is established between the device network element and at least one other network element. The Internet, wide area network, local area network, metropolitan area network, etc. can be used.
[0118] The bus 802 may be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 801 is used to store a program. After receiving an execution instruction, the processor 800 executes the program. The method for multi-label classification of images disclosed in any one of the foregoing embodiments of the present application can be applied to the processor 800 or implemented by the processor 800.
[0119] The processor 800 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 800 or the instructions in the form of software. The above-mentioned processor 800 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 801, and the processor 800 reads the information in the memory 801 and combines its hardware to complete the steps of the above method.
[0120] The electronic device provided by the embodiments of the present application and the method for multi-label classification of images provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by it.
[0121] The embodiments of the present application also provide a computer-readable storage medium corresponding to the method for multi-label classification of images provided in the foregoing embodiments. Please refer to Figure 13 , which shows that the computer-readable storage medium is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will execute the method for multi-label classification of images provided in any of the foregoing embodiments.
[0122] It should be noted that the examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated here one by one.
[0123] The computer-readable storage medium provided by the above embodiments of the present application and the method for image multi-label classification provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.
[0124] It should be noted that:
[0125] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present application may be practiced without these specific details. In some instances, well-known structures and techniques have not been shown in detail in order not to obscure the understanding of this specification.
[0126] Similarly, it should be understood that, in order to streamline the present application and assist in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together in a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting the following schematic: that the claimed present application requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, where each claim stands on its own as a separate embodiment of the present application.
[0127] In addition, those skilled in the art will understand that although some of the embodiments described herein include certain features included in other embodiments but not others, the combination of features of different embodiments is within the scope of the present application and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.
[0128] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for multi-label classification of images, characterized in that, Including: Obtain a training set, where the training set includes sample images labeled with multiple classification labels; Construct a network model structure for multi-label classification, where the network model structure includes multiple activation layers, and the number of activation layers is equal to the number of classification labels, and each activation layer outputs the prediction probability of one classification label; The constructing the network model structure for multi-label classification includes: constructing a backbone classifier based on a preset classification model; connecting the backbone classifier with multiple activation layers; Train the network model structure according to the training set to obtain a trained multi-label classification model; the training the network model structure according to the training set to obtain a trained multi-label classification model includes: obtaining a sample image from the training set; inputting the sample image into the backbone classifier to output a feature vector corresponding to each classification label; inputting each feature vector into different activation layers among the multiple activation layers, and each activation layer outputs the prediction probability corresponding to the received feature vector, that is, obtaining the prediction probability corresponding to each classification label; calculating the loss value of the current training cycle through a preset loss function according to the prediction probability corresponding to each classification label.
2. The method according to claim 1, wherein The constructing the network model structure for multi-label classification includes: Construct a backbone classifier based on a preset classification model; Connect the backbone classifier with a multi-head attention layer; Connect the multi-head attention layer with multiple activation layers.
3. The method according to claim 2, wherein The preset classification model includes an EfficientNet network; Remove the normalization layer of the EfficientNet network to obtain the backbone classifier.
4. The method according to claim 2, wherein The training the network model structure according to the training set to obtain a trained multi-label classification model includes: Obtaining a sample image from the training set; Inputting the sample image into the backbone classifier to output a feature vector corresponding to each classification label; Inputting the feature vector corresponding to each classification label into the multi-head attention layer to output a multi-head attention matrix corresponding to each classification label; Inputting each multi-head attention matrix into the multiple activation layers respectively to obtain the prediction probability corresponding to each classification label; Calculating the loss value of the current training cycle through a preset loss function according to the prediction probability corresponding to each classification label.
5. The method according to any one of claims 1-2 and 4, characterized in that, The method further includes: Obtaining an image to be classified; Classifying the image to be classified through the trained multi-label classification model.
6. The method according to claim 5, wherein The classifying the image to be classified through the trained multi-label classification model includes: Inputting the image to be classified into the trained multi-label classification model to obtain the prediction probability corresponding to each classification label; Determining the classification label to which the image to be classified belongs as the classification label with a prediction probability greater than a preset threshold.
7. An apparatus for multi-label classification of images, characterized in that, Including: An acquisition module for acquiring a training set, where the training set includes sample images labeled with multiple classification labels; The model construction module is used to construct a network model structure for multi-label classification. The network model structure includes multiple activation layers, and the number of the activation layers is equal to the number of the classification labels. Each activation layer outputs the prediction probability of a classification label respectively. The model construction module is further used to: construct a backbone classifier based on a preset classification model; connect the backbone classifier with the multiple activation layers; The model training module is used to train the network model structure according to the training set to obtain a trained multi-label classification model. The model training module is further used to: obtain a sample image from the training set; input the sample image into the backbone classifier to output a feature vector corresponding to each classification label; input each of the feature vectors into different activation layers among the multiple activation layers respectively, and each activation layer outputs the prediction probability corresponding to the received feature vector, that is, the prediction probability corresponding to each classification label is obtained; calculate the loss value of the current training cycle through a preset loss function according to the prediction probability corresponding to each classification label.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor runs the computer program to implement the method according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the method according to any one of claims 1-6.
Citation Information
Patent Citations
Image processing method and device, medium and electronic device
CN109543773A