Image Processing Method, Apparatus and Related Devices Based on a Multi-Task Model
By designing feature extraction and embedded representation structures in parallel in multi-task model, and adjusting parameters through phased training loss, the problems of poor similarity embedding characterization and classification overfitting are solved, and the multi-task learning effect of image processing is improved.
Patent Information
- Application Number
- CN202110827411.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-21
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-07-21
AI Technical Summary
In the image re-retrieval task of existing multi-task models, when similarity embedded representation and classification are learned in joint learning, it is easy to lead to poor similarity embedding representation, affecting the classification effect, and the classification model is easily overfitted, making it difficult to effectively retain the relative relationship of similarity embedding in categories.
Using the image processing method based on multi-task model, the feature extraction structure and embedded representation structure are designed in parallel, and the first target loss is generated using predicted embedded representation, the feature extraction and embedded representation structure parameters are adjusted, and the second target loss is generated based on the category prediction results and labels are used to train in stages to avoid overfitting the embedded representation by the classification structure.
It improves the training effect of multi-task model, ensures the relative relationship of similarity embedding in categories, prevents overfitting of classification structures, and improves the accuracy of image classification and embedded representation.
Smart Images

Figure CN113822324B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer vision technology. Specifically, it relates to an image processing method, apparatus, electronic device, and computer-readable medium based on a multi-task model. Background Art
[0002] In the image deduplication retrieval task, the quality of the similarity embedding used to represent image similarity is very important. The purpose of similarity embedding is to make the distance between the same images very small and the distance between different images very large. Similarity embedding is based on images as the granularity, which is different from the conventional classification embedding (based on categories as the granularity). Classification embedding requires that the distance between images of the same category is small and the distance between images of different categories is large. Conventional similarity learning does not need to consider image category information. However, in general applications, in addition to using similarity embedding to deduplicate images, it is also necessary to classify or assign multiple labels to images. A direct approach is to add a fully connected layer for classification after the similarity embedding layer of the model. However, similarity embedding often relies on extracting significant foreground objects in the image. When an image has no significant foreground (such as a lake, grassland, blue sky), it is easy to cause poor representation of similarity embedding and thus poor classification effect.
[0003] Therefore, there is a need for a new image processing method, apparatus, electronic device, and computer-readable medium based on a multi-task model.
[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure. Summary of the Invention
[0005] Embodiments of the present disclosure provide an image processing method, apparatus, electronic device, and computer-readable medium based on a multi-task model, thereby at least to a certain extent avoiding the influence of the embedding representation on the classification effect and improving the training effect of the multi-task model.
[0006] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be partially learned through the practice of the present disclosure.
[0007] An embodiment of the present disclosure provides an image processing method based on a multi-task model, including: obtaining a sample image and the class label of the sample image; processing the sample image through a feature extraction structure in the multi-task model to obtain image features; processing the image features through an embedded representation structure in the multi-task model to obtain a predicted embedded representation of the sample image; determining a first target loss according to the predicted embedded representation, so as to adjust the parameters of the feature extraction structure and the embedded representation structure in the multi-task model according to the first target loss, and obtaining a multi-task model in the first training stage; processing the image features through a classification structure parallel to the embedded representation structure in the multi-task model to obtain a class prediction result of the sample image; determining a second target loss according to the class prediction result, the class label and the predicted embedded representation; adjusting the parameters of the feature extraction structure, the embedded representation structure and the classification structure in the multi-task model in the first training stage according to the second target loss, and obtaining the trained multi-task model, so as to perform predictions of image classification and embedded representation according to the trained multi-task model.
[0008] An embodiment of the present disclosure provides an image processing apparatus based on a multi-task model, including: a sample acquisition module configured to obtain a sample image and the class label of the sample image; a feature extraction module configured to process the sample image through a feature extraction structure in the multi-task model to obtain image features; an embedded representation module configured to process the image features through an embedded representation structure in the multi-task model to obtain a predicted embedded representation of the sample image; a first training module configured to determine a first target loss according to the predicted embedded representation, so as to adjust the parameters of the feature extraction structure and the embedded representation structure in the multi-task model according to the first target loss, and obtaining a multi-task model in the first training stage; a class prediction module configured to process the image features through a classification structure parallel to the embedded representation structure in the multi-task model to obtain a class prediction result of the sample image; a second training module configured to determine a second target loss according to the class prediction result, the class label and the predicted embedded representation, adjust the parameters of the feature extraction structure, the embedded representation structure and the classification structure in the multi-task model in the first training stage according to the second target loss, and obtaining the trained multi-task model, so as to perform predictions of image classification and embedded representation according to the trained multi-task model.
[0009] In an exemplary embodiment of the present disclosure, when the first training module "determines the first target loss according to the predicted embedded representation", it includes: a sample pair generation sub-module configured to generate a sample pair according to the sample image, where the two sample images included in the sample pair are a first image and a second image, and the distance between the actual embedded representations of the first image and the second image is less than a distance threshold; a global triplet sub-module configured to form a global triplet sample by combining a sample image with a different class label from the first image in the sample pair and the sample pair; a local triplet sub-module configured to form a local triplet sample by combining a sample image with the same class label as the first image in the sample pair and the sample pair; a first loss sub-module configured to generate the first target loss according to the predicted embedded representations of each sample image in the global triplet sample and the local triplet sample.
[0010] In an exemplary embodiment of the present disclosure, the global triplet sub-module includes: a sample pair image generation unit configured to, for each sample pair, randomly select a sample image from the remaining sample pairs as the sample pair image of each remaining sample pair; a first sample pair image unit configured to determine a first target sample pair image whose class label is different from the class label of the first image in the sample pair; a first distance calculation unit configured to calculate the distance between each first target sample pair image and the first image in the sample pair to obtain the first distance of each first target sample pair image; a first distance sorting unit configured to sort the first target sample pair images in ascending order of the first distance; a global triplet unit configured to form a global triplet sample by combining the first a first target sample pair images in the sorting result and the sample pair respectively, where a is an integer greater than 0, and the global triplet sample includes the first image, the second image of the sample pair, and the first target sample pair image as the third image.
[0011] In an exemplary embodiment of the present disclosure, the local triple sub-module includes: a sample pair image generation unit configured to, for each sample pair, randomly select a sample image from the remaining sample pairs as the sample pair image for each of the remaining sample pairs; a second sample pair image unit configured to determine a second target sample pair image whose class label is the same as the class label of the first image in the sample pair; a second distance calculation unit configured to calculate the distance between each second target sample pair image and the first image in the sample pair to obtain the second distance of each second target sample pair image; a second distance sorting unit configured to sort the second target sample pair images in ascending order of the second distance; and a local triple unit configured to form b local triple samples by combining the first b second target sample pair images in the sorting result with the sample pair respectively, where b is an integer greater than 0, and the local triple sample includes the first image, the second image of the sample pair, and the second target sample pair image as the third image.
[0012] In an exemplary embodiment of the present disclosure, the first loss sub-module includes: a global embedded representation loss calculation unit configured to determine a global embedded representation loss based on the predicted embedded representations of the first image, the second image, and the third image in the global triple sample; a local embedded representation loss calculation unit configured to determine a local embedded representation loss based on the predicted embedded representations of the first image, the second image, and the third image in the local triple sample; and a first loss calculation unit configured to determine the first target loss based on the weighted calculation result of the global embedded representation loss and the local embedded representation loss.
[0013] In an exemplary embodiment of the present disclosure, the global embedded representation loss calculation unit includes: a first positive sample pair distance calculation sub-unit configured to calculate the distance between the predicted embedded representations of the first image and the second image in the global triple sample to obtain the positive sample pair distance of the global triple sample; a first negative sample pair distance calculation sub-unit configured to calculate the distance between the predicted embedded representations of the first image and the third image in the global triple sample to obtain the negative sample pair distance of the global triple sample; and a global embedded representation loss calculation sub-unit configured to determine the global embedded representation loss based on the positive sample pair distance and the negative sample pair distance of the global triple sample.
[0014] In an exemplary embodiment of the present disclosure, the local embedded representation loss calculation unit includes: a second positive sample pair distance calculation sub-unit, configured to calculate the distance between the predicted embedded representations of the first image and the second image in the local triple sample, and obtain the positive sample pair distance of the local triple sample; a second negative sample pair distance calculation sub-unit, configured to calculate the distance between the predicted embedded representations of the first image and the third image in the local triple sample, and obtain the negative sample pair distance of the local triple sample; a local embedded representation loss calculation sub-unit, configured to determine the local embedded representation loss according to the positive sample pair distance and the negative sample pair distance of the local triple sample.
[0015] In an exemplary embodiment of the present disclosure, the second training module includes: a classification loss calculation sub-module, configured to determine a classification loss according to the class prediction result of the sample image and the class label; a second loss calculation sub-module, configured to determine the second target loss according to the weighted calculation result of the classification loss and the first target loss.
[0016] In an exemplary embodiment of the present disclosure, the classification loss calculation sub-module includes: a class prediction result partitioning unit, configured to determine a first prediction result of the class prediction result of the sample image under the class label of the sample image and Nc - 1 second prediction results under the other Nc - 1 classes outside the class label of the sample image, where Nc is the total number of classes and Nc is an integer greater than 1; a labeled class loss calculation unit, configured to determine a labeled class loss according to the first prediction result, the class label of the sample image, and a first weight; a single unlabeled class loss calculation unit, configured to determine Nc - 1 single unlabeled class losses according to the Nc - 1 second prediction results, the class label of the sample image, and a second weight, where the first weight and the second weight are preset values and the sum of the first weight and the second weight is a fixed value; an unlabeled class loss calculation unit, configured to determine the average value of the Nc - 1 single unlabeled class losses as the unlabeled class loss; a classification loss calculation unit, configured to determine the classification loss according to the labeled class loss and the unlabeled class loss.
[0017] In an exemplary embodiment of the present disclosure, the classification structure includes a convolutional unit and a fully connected unit connected in sequence; wherein, the class prediction module includes: a convolutional operation sub-module, configured to perform a convolutional operation on the image feature according to the convolutional unit to obtain a convolutional output; a classification prediction sub-module, configured to process the convolutional output through the fully connected unit to obtain the class prediction result of the sample image.
[0018] In an exemplary embodiment of the present disclosure, when the second training module "performs prediction of image classification and embedded representation according to the trained multi-task model", it includes: an image acquisition sub-module configured to obtain an image to be predicted; an image prediction sub-module configured to process the image to be predicted through the trained multi-task model to obtain the embedded representation and classification result of the image to be predicted.
[0019] An embodiment of the present disclosure provides an electronic device, including: at least one processor; a storage device for storing at least one program, which, when executed by the at least one processor, causes the at least one processor to implement the image processing method based on a multi-task model as described in the above embodiment.
[0020] An embodiment of the present disclosure provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, it implements the image processing method based on a multi-task model as described in the above embodiment.
[0021] In the technical solution provided by some embodiments of the present disclosure, when training a multi-task model using a sample image and the class label of the sample image, first, the sample image is processed by the feature extraction structure in the multi-task model to obtain image features; then the image features are simultaneously input into the classification structure and the embedded representation structure, so that the classification structure and the embedded structure share the underlying feature extraction structure, which can save the inference time of feature extraction. And the parallel design of the classification structure and the embedded structure can reduce the influence of classification on the learning of embedded representation. In addition, first, a first target loss is generated according to the predicted embedded representation to adjust the parameters of the feature extraction structure and the embedded representation structure to obtain a multi-task model in the first training stage, and then a second target loss is generated according to the predicted embedded representation, the class prediction result, and the class label to adjust the parameters of the feature extraction structure, the embedded representation structure, and the classification structure in the multi-task model in the first training stage to obtain the trained multi-task model. This phased training and learning method can take into account the characteristic that the embedded representation task is more difficult to converge than the classification task, and improve the effect of multi-task learning through a two-stage learning method of first pre-training the embedded representation structure and then fine-tuning the network in combination with the classification structure, which can effectively prevent overfitting of the classification structure and ensure the embedded representation.
[0022] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the accompanying drawings:
[0024] Figure 1 Schematically shows a structural diagram of a combined learning model for similarity embedded representation and classification according to the related art.
[0025] Figure 2 Shows a schematic diagram of an exemplary system architecture of an image processing method or apparatus based on a multi-task model to which embodiments of the present disclosure can be applied.
[0026] Figure 3 Schematically shows a flowchart of an image processing method based on a multi-task model according to an embodiment of the present disclosure.
[0027] Figure 4 Schematically shows a flowchart of an image processing method based on a multi-task model according to another embodiment of the present disclosure.
[0028] Figure 5 Schematically shows a flowchart of an image processing method based on a multi-task model according to still another embodiment of the present disclosure.
[0029] Figure 6 Schematically shows a flowchart of an image processing method based on a multi-task model according to yet another embodiment of the present disclosure.
[0030] Figure 7 Schematically shows a flowchart of an image processing method based on a multi-task model according to yet another embodiment of the present disclosure.
[0031] Figure 8 Schematically shows a flowchart of an image processing method based on a multi-task model according to yet another embodiment of the present disclosure.
[0032] Figure 9 Schematically shows a flowchart of an image processing method based on a multi-task model according to yet another embodiment of the present disclosure.
[0033] Figure 10 Schematically shows a flowchart of an image processing method based on a multi-task model according to yet another embodiment of the present disclosure.
[0034] Figure 11 Schematically shows a flowchart of an image processing method based on a multi-task model according to yet another embodiment of the present disclosure.
[0035] Figure 12 Schematically shows the structural diagram of a multi - task model according to an embodiment of the present disclosure.
[0036] Figure 13 Schematically shows the structural diagram of a multi - task model according to another embodiment of the present disclosure.
[0037] Figure 14 Schematically shows the schematic diagram of a triple sample according to an embodiment of the present disclosure.
[0038] Figure 15 Schematically shows the block diagram of an image processing device based on a multi - task model according to an embodiment of the present disclosure.
[0039] Figure 16 Shows the structural schematic diagram of an electronic device suitable for implementing the embodiments of the present disclosure. Detailed implementation manners
[0040] Now, example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art.
[0041] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well - known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present disclosure.
[0042] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in at least one hardware module or integrated circuit, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0043] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the contents and operations / steps, nor do they necessarily have to be executed in the described order. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0044] Computer Vision (CV) technology is a science that studies how to enable machines to "see". Further speaking, it refers to using cameras and computers to replace human eyes for tasks such as object recognition, tracking, and measurement in machine vision, and further performing graphic processing to make the images processed by the computer more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision researches related theories and technologies, and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0045] In computer vision technology, image recognition refers to recognition at the category level, without considering specific instances of the object, but only considering the category of the object (such as people, dogs, cats, birds, etc.) for recognition and giving the category to which the object belongs. A typical example is the recognition task in the large-scale general object recognition open-source dataset (such as imagenet), which recognizes which of the 1000 categories a certain object belongs to.
[0046] A joint learning model of similarity embedded representation and classification in related technologies is as Figure 1 shown. After the model embedded representation unit 101, a fully connected layer 102 for classification is added, and the losses of the two models (classification loss 110 and embedded representation loss 120) are learned simultaneously during learning.
[0047] However, Figure 1 The joint learning model of similarity embedded representation and classification shown has the following defects: The model embedded representation unit 101 is overfitted to classification, and the fully connected layer 102 is prone to overfitting in learning the one-hot targets of classification. And this overfitting directly affects the representation of the global similarity embedding of the image through gradient update. As a result, the similarity embedding cannot meet the requirements for image representation, and further, images with similar foregrounds but different backgrounds (or partially different backgrounds) cannot be distinguished at the similarity embedding level.
[0048] Current multi-task models have defects that can easily lead to poor similarity embedded representations and thus affect classification results. On the other hand, the purpose of similarity embedding is to represent images. In principle, there is a potential relative relationship: the differences in similarity embedding at the feature level should be significantly different according to whether two images are in the same category (the similarity embedding features of two different images but in the same category are close in a distinguishable case, and far apart for different categories). When this relationship is not satisfied, unreasonable sorting in the recall results is likely to occur. Finally, since the classification model has to learn one-hot outputs, the final output branch is prone to overfitting to the classification task, which causes the heatmap of this branch to focus on the parts related to classification in the image, and this will affect the representation of the full-image feature embedding. How to effectively maintain the relative relationship of similarity embedding in terms of categories and reasonably combine the learning of classification tasks is a problem.
[0049] Therefore, a new image processing method, apparatus, electronic device, and computer-readable medium based on a multi-task model are needed.
[0050] Figure 2 The schematic diagram shows an exemplary system architecture to which the image processing method or apparatus based on a multi-task model according to the embodiments of the present disclosure can be applied.
[0051] As Figure 2 shown, the system architecture 200 may include one or more of the terminal devices 201, 202, 203, the network 204, and the server 205. The network 204 is used to provide a medium for communication links between the terminal devices 201, 202, 203 and the server 205. The network 204 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0052] It should be understood that Figure 2 the numbers of terminal devices, networks, and servers in
[0053] Users can use terminal devices 201, 202, and 203 to interact with server 205 via network 204 to receive or send messages, etc. Terminal devices 201, 202, and 203 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, portable computers, desktop computers, wearable devices, virtual reality devices, smart homes, smart TVs, smart in-vehicle devices, and so on. A client can be installed on terminal devices 201, 202, and 203, such as a video client, an information stream client, a browser client, etc., but the present disclosure is not limited thereto.
[0054] Server 205 can be a server that provides various services. For example, terminal device 203 (which can also be terminal device 201 or 202) uploads a sample image and the class label of the sample image to server 205. Server 205 can obtain the sample image and the class label of the sample image; process the sample image through the feature extraction structure in the multi-task model to obtain image features; process the image features through the embedded representation structure in the multi-task model to obtain the predicted embedded representation of the sample image; determine a first target loss according to the predicted embedded representation, so as to adjust the parameters of the feature extraction structure and the embedded representation structure in the multi-task model according to the first target loss to obtain the multi-task model in the first training stage; process the image features through the classification structure parallel to the embedded representation structure in the multi-task model to obtain the class prediction result of the sample image; determine a second target loss according to the class prediction result, the class label, and the predicted embedded representation; adjust the parameters of the feature extraction structure, the embedded representation structure, and the classification structure in the multi-task model in the first training stage according to the second target loss to obtain the trained multi-task model. And feedback the trained multi-task model to terminal device 203. Further, terminal device 203 can obtain an image to be predicted; process the image to be predicted through the trained multi-task model to obtain the embedded representation and classification result of the image to be predicted.
[0055] Figure 3 Schematically shows a flowchart of an image processing method based on a multi-task model according to an embodiment of the present disclosure. The method provided by the embodiments of the present disclosure can be processed by any electronic device with computing and processing capabilities, such as server 205 and / or terminal devices in the above Figure 2 embodiments. In the following embodiments, server 205 is taken as an example of the execution subject for illustration, but the present disclosure is not limited thereto.
[0056] As Figure 3As shown, the image processing method based on a multi-task model provided by an embodiment of the present disclosure may include the following steps.
[0057] In step S310, a sample image and a class label of the sample image are obtained.
[0058] In an embodiment of the present disclosure, the class label of each sample image is used to represent the classification category of the sample image. For example, reference may be made to the 1000 classes (such as dogs, cats, etc.) that often appear in ImageNet.
[0059] In step S320, the sample image is processed by a feature extraction structure in the multi-task model to obtain image features.
[0060] In an embodiment of the present disclosure, the multi-task model may be a deep learning model, which can be used for processing tasks of similarity embedded representation (hereinafter referred to as embedded representation) and classification of images. The feature extraction structure may, for example, adopt a 101-layer residual network (ResNet-101). In the training initialization stage, the parameters in ResNet-101 may adopt the parameters pre-trained on the ImageNet dataset.
[0061] The structure of ResNet-101 may be as shown in Table 1.
[0062] Table 1 Structure Table of ResNet-101 Feature Modules
[0063]
[0064]
[0065] In step S330, the image features are processed by an embedded representation structure in the multi-task model to obtain a predicted embedded representation of the sample image.
[0066] In an embodiment of the present disclosure, the embedded representation structure may include a pooling layer, a fully connected layer, and a feature regularization layer connected in sequence. The embedded representation structure may be as shown in Table 2.
[0067] Table 2 Structure Table of Embedded Representation Structure
[0068]
[0069]
[0070] In Table 2, N is the dimension of the embedded representation, such as 1*123, and the feature regularization layer may, for example, adopt an Euclidean norm regularization layer (L2 normalization).
[0071] In step S340, a first target loss is determined according to the predicted embedded representation, so as to adjust the parameters of the feature extraction structure and the embedded representation structure in the multi-task model according to the first target loss, and obtain the multi-task model in the first training stage.
[0072] In the embodiments of the present disclosure, for example, the parameters of the structures shown in Table 1 and Table 2 can be adjusted according to the first target loss, and the adjusted multi-task model obtained is the multi-task model in the first training stage. Among them, the parameters of the feature extraction structure and the embedded representation structure can be adjusted with a learning rate of 0.05.
[0073] In step S350, the image features are processed by the classification structure parallel to the embedded representation structure in the multi-task model to obtain the class prediction result of the sample image.
[0074] In the embodiments of the present disclosure, the structure diagram of the multi-task model can be seen Figure 12 . As Figure 12 shown, the multi-task model may include a feature extraction structure 1210, an embedded representation structure 1220, and a classification structure 1230 parallel to the embedded representation structure 1220.
[0075] In an exemplary embodiment, as Figure 12 shown, the classification structure 1230 may include a convolution unit 1231 and a fully connected unit 1232 connected in sequence. In this step, the convolution operation can be performed on the image features according to the convolution unit 1231 to obtain a convolution output; the convolution output is processed by the fully connected unit 1232 to obtain the class prediction result of the sample image. In Figure 12 , since the result output by the feature extraction structure 1210 directly supports the similarity embedded representation learning of the embedded representation structure 1220, in order to avoid the last layer of the feature extraction structure 1210 (such as conv5_x in Table 1) being affected by classification prematurely, in this embodiment, an additional convolution module is added to the classification structure 1230, which can avoid the gradient backpropagation of classification reaching conv5_x too quickly and thus affecting the effect of the embedded representation structure 1220 (this is because the gradient backpropagation calculates the gradient layer by layer from the output end of the network until the model output layer, and the modules closer to the model output layer are more easily affected by the output layer loss). At the same time, it can enable the classification structure 1230 to further extract key part features related to classification on the global features of conv5_x, ensuring the classification effect.
[0076] The classification structure 1230 can be as shown in Table 3.
[0077] Table 3 Structure Table of Classification Structure
[0078]
[0079] Among them, the classification structure includes a convolutional unit Conv6_x, a pooling layer Pool, and a fully connected layer Fc connected in sequence, and the probabilities of Nc classifications are predicted.
[0080] In step S360, a second target loss is determined according to the class prediction result, the class label, and the predicted embedded representation.
[0081] In this functional embodiment, the classification loss can be determined according to the class prediction result and the class label of the sample image; the second target loss is determined according to the weighted calculation result of the classification loss and the first target loss.
[0082] In step S370, the parameters of the feature extraction structure, the embedded representation structure, and the classification structure in the multi-task model in the first training stage are adjusted according to the second target loss, and a trained multi-task model is obtained to perform image classification and prediction of the embedded representation according to the trained multi-task model.
[0083] In the embodiment of the present disclosure, for the classification structure, a learning rate of 0.05 can be adopted, and for the feature extraction structure and the embedded representation structure, a learning rate of 0.005 can be adopted to avoid the effect of the classification structure being too fast on the feature extraction structure. The trained multi-task model can be regarded as the multi-task model in the second training stage. For the sample set of M sample images, the two-stage learning method adopted in the embodiment of the present disclosure can be specifically as follows.
[0084] 1) The first stage
[0085] First, train the branches of Table 1 and Table 2. The learning rate of all network structures is 0.05. Calculate the first target loss in each round of iteration, calculate the gradient, and update the network parameters. After M / bs rounds (completing one full-data learning), 1 epoch of learning is completed, and continue to iterate for the next epoch of learning. The learning rate drops to 0.1 times the previous value every 10 epochs. Terminate when K epochs of training are reached (or when the average first target loss for 10 consecutive epochs does not decrease).
[0086] 2) The second stage
[0087] Train the branches of Table 1, Table 2, and Table 3. Except for the classification structure of Table 3 (learning rate is 0.05), the learning rate of other network structures is 0.005. Calculate the classification loss and the first target loss to calculate the second target loss and update the network parameters. The learning rate drops to 0.1 times the previous value every 10 epochs. Terminate when K epochs of training are reached (or when the average second target loss for 10 consecutive epochs does not decrease).
[0088] The image processing method based on a multi-task model provided by the embodiments of the present disclosure, when training the multi-task model using a sample image and the class label of the sample image, first processes the sample image using the feature extraction structure in the multi-task model to obtain image features; then inputs the image features into the classification structure and the embedded representation structure simultaneously, enabling the classification structure and the embedded structure to share the underlying feature extraction structure, which can save the inference time of feature extraction. Moreover, the parallel design of the classification structure and the embedded structure can reduce the impact of classification on the embedded representation learning. In addition, first generate a first target loss based on the predicted embedded representation to adjust the parameters of the feature extraction structure and the embedded representation structure, obtaining the multi-task model in the first training stage, and then generate a second target loss based on the predicted embedded representation, the class prediction result, and the class label to adjust the parameters of the feature extraction structure, the embedded representation structure, and the classification structure in the multi-task model of the first training stage, obtaining the trained multi-task model. This phased training and learning method can take into account the characteristic that the embedded representation task is more difficult to converge than the classification task, and improve the effect of multi-task learning through a two-stage learning of first pre-training the embedded representation structure and then fine-tuning the network in combination with the classification structure, which can effectively prevent overfitting of the classification structure and ensure the embedded representation.
[0089] In an exemplary embodiment, the image processing method based on a multi-task model of the present disclosure may further include the following steps 1) and 2).
[0090] In step 1), obtain an image to be predicted.
[0091] In the embodiments of the present disclosure, the image to be predicted may be an object that is currently received and requires multi-task recognition. The multi-task recognition includes an image embedded representation task and an image classification task.
[0092] In step 2), process the image to be predicted through the trained multi-task model to obtain the embedded representation and the classification result of the image to be predicted.
[0093] In the embodiments of the present disclosure, the structure of the trained multi-task model may be as Figure 13 shown. After training, the image to be predicted can be used as the input of the multi-task model. First, process the image to be predicted through the feature extraction structure 1210 to obtain an image feature output; process the image feature output through the embedded representation structure 1220 to obtain the embedded representation 1221 of the image to be predicted; process the image feature output through the classification structure 1230 to obtain the classification result 1233 of the image to be predicted, where the classification structure 1230 may include a convolutional unit 1231 and a fully connected unit 1232 connected in sequence.
[0094] The embedded representation of the image to be predicted obtained in the embodiments of the present disclosure can be used for image deduplication (removing identical and similar images) and other downstream image characterization tasks.
[0095] Figure 4 FIG. schematically shows a flowchart of an image processing method based on a multi-task model according to another embodiment of the present disclosure.
[0096] As Figure 4 shown, in the above Figure 3 When determining the first target loss according to the predicted embedded representation in step S340 of the embodiment, the following steps may be further included.
[0097] In step S410, a sample pair is generated according to the sample image. The two sample images included in the sample pair are the first image and the second image, and the distance between the actual embedded representations of the first image and the second image is less than the distance threshold.
[0098] In the embodiments of the present disclosure, the actual embedded representation of the first image refers to the true embedded representation of the first image, and the actual embedded representation of the second image refers to the true embedded representation of the second image. The distance threshold is used to measure similar images. For the first image and the second image, if the distance between their actual embedded representations is less than the distance threshold, it is considered that the first image and the second image are similar images; if the distance between their actual embedded representations is greater than or equal to the distance threshold, it is considered that the first image and the second image are not similar images.
[0099] In step S420, a sample image different from the class label of the first image in the sample pair and the sample pair are combined to form a global triplet sample.
[0100] In the embodiments of the present disclosure, for each sample pair, a sample image can be selected from the sample pairs with different sample pair labels from the sample pair, and the selected sample image and the sample pair are combined to form a global triplet sample. Among them, the sample pair label is to randomly select a sample image from the sample pair (that is, the first image or the second image), and the class label of the randomly selected sample image is determined as the sample pair label of the sample pair.
[0101] In step S430, a sample image with the same class label as the first image in the sample pair and the sample pair are combined to form a local triplet sample.
[0102] In the embodiments of the present disclosure, for each sample pair, a sample image can be selected from the sample pairs with the same sample pair label as the sample pair, and the selected sample image and the sample pair are combined to form a global triplet sample.
[0103] When generating global triple samples and local triple samples, for the existing sample pair images, randomly select one sample image from them to label the class label of the image. The labeled class labels can refer to the 1000 labels that often appear in ImageNet (such as dogs, cats, etc.).
[0104] Taking the sample pairs as input, the following mining is performed on each batch (e.g., bs sample pairs) of sample pairs to obtain triples: For a certain sample pair x: From the samples of the remaining bs - 1 sample pairs (randomly select one image from each sample pair as the sample pair image of that sample pair): (1) Find the sample pair images that are different in class from the sample pair x, calculate their distances from the sample pair x, sort them in ascending order of distance, and take the top 10 sample pair images as negative samples, and respectively form global triple samples with the sample pair x. Therefore, each sample pair generates 10 global triple samples; (2) Find the sample pair images that are the same in class as the sample pair x, calculate their distances from the sample pair x, sort them in ascending order of distance, take the top 10 sample pair images as negative samples, and respectively form local triple samples with the sample pair x. Therefore, each sample pair generates 10 local triple samples; The entire batch obtains 20 * bs triple samples (including global triple samples and local triple samples).
[0105] In step S440, a first target loss is generated according to the predicted embedded representations of each sample image in the global triple samples and the local triple samples.
[0106] In the embodiments of the present disclosure, a global embedded representation loss can be generated according to the global triple samples, a local embedded representation loss can be generated according to the local triple samples, and the first target loss is determined according to the weighted calculation result of the global embedded representation loss and the local embedded representation loss.
[0107] Figure 5 Schematically shows a flowchart of an image processing method based on a multi-task model according to another embodiment of the present disclosure.
[0108] As Figure 5 shown, step S420 in the above Figure 4 embodiment can further include the following steps.
[0109] In step S510, for each sample pair, randomly select one sample image from the remaining sample pairs as the sample pair image of each of the remaining sample pairs.
[0110] In the embodiments of the present disclosure, the sample pair image of each sample pair can be the first image or the second image of that sample pair.
[0111] In step S520, a sample pair image whose class label is different from that of the first image in the sample pair is determined as a first target sample pair image.
[0112] In the embodiments of the present disclosure, when comparing the sample pair images of the remaining sample pairs with the class label of the first image in the current sample pair for the current sample pair, the first image in the current sample pair is regarded as the sample pair image of the current sample pair. However, this is only for convenience of description. In the present application, the first image of the current sample pair is a hypothetical sample image randomly selected from the current sample pair. In the embodiments of the present disclosure, the second image in the current sample pair can also be regarded as the sample pair image of the current sample pair to compare with the class labels of the sample pair images of the remaining sample pairs.
[0113] In step S530, the distance between each first target sample pair image and the first image in the sample pair is calculated to obtain the first distance of each first target sample pair image.
[0114] In the embodiments of the present disclosure, the distance between two images can be the Euclidean distance obtained according to the actual embedded representations of the two images.
[0115] In step S540, the first target sample pair images are sorted in ascending order of the first distance.
[0116] In the embodiments of the present disclosure, the sorting can be performed in ascending order of the first distance from small to large.
[0117] In step S550, the first a first target sample pair images in the sorting result and the sample pair are respectively combined into a global triplet samples, where a is an integer greater than 0. The global triplet samples include the first image, the second image of the sample pair, and the first target sample pair image as the third image.
[0118] In the embodiments of the present disclosure, the value of a can be, for example, 10. However, this is only an example, and the present disclosure does not make special limitations on the specific value of a. Among them, for each sample pair, it can form a global triplet samples.
[0119] Figure 6 Schematically shows a flowchart of an image processing method based on a multi-task model according to still another embodiment of the present disclosure.
[0120] As Figure 6 shown, step S430 in the above Figure 4 embodiments can further include the following steps.
[0121] In step S610, for each sample pair, a sample image is randomly selected from the remaining sample pairs as the sample pair image of each of the remaining sample pairs.
[0122] Step S610 in the embodiments of the present disclosure may adopt steps similar to those of step S510, which will not be elaborated here.
[0123] In step S620, the sample pair images with the same class label as the class label of the first image in the sample pair are determined as the second target sample pair images.
[0124] For the current sample pair, when comparing the sample pair images of the remaining sample pairs with the class label of the first image in the current sample pair, the first image in the current sample pair is regarded as the sample pair image of the current sample pair. However, this is only for convenience of description. In the present application, for the first image of the current sample pair, it is an assumed sample image randomly selected from the current sample pair. In the embodiments of the present disclosure, the second image in the current sample pair may also be regarded as the sample pair image of the current sample pair to compare with the class labels of the sample pair images of the remaining sample pairs.
[0125] In step S630, the distances between each second target sample pair image and the first image in the sample pair are calculated to obtain the second distances of each second target sample pair image.
[0126] In the embodiments of the present disclosure, the distance between two images may be the Euclidean distance obtained according to the actual embedded representations of the two images.
[0127] In step S640, the second target sample pair images are sorted in ascending order of the second distances.
[0128] In the embodiments of the present disclosure, the sorting may be performed from the smallest to the largest second distance.
[0129] In step S650, the first b second target sample pair images in the sorting result are respectively combined with the sample pair to form b local triple sample pairs. b is an integer greater than 0. The local triple sample pair includes the first image, the second image of the sample pair, and the second target sample pair image as the third image.
[0130] In the embodiments of the present disclosure, the value of b may be, for example, 10. However, this is only an example, and the present disclosure does not make special limitations on the specific value of b. Among them, for each sample pair, it can form b local triple sample pairs.
[0131] Figure 7 The flowchart of an image processing method based on a multi-task model according to still another embodiment of the present disclosure is schematically shown.
[0132] As Figure 7 shown, step S440 in the above Figure 4 embodiments may further include the following steps.
[0133] In step S710, a global embedding representation loss is determined based on the predicted embedding representations of the first image, the second image, and the third image in the global triple samples.
[0134] In the embodiments of the present disclosure, the global embedding representation loss L can be determined according to the following formula (1). tr1 .
[0135]
[0136] Where a 1 is the first image in the global triple sample, p 1 is the second image in the global triple sample, and n 1 is the third image in the global triple sample. represents the Euclidean distance between the predicted embedding representations of the first image and the second image in the global triple sample. The Euclidean distance between the predicted embedding representations of the first image and the third image in the global triple sample. α 1 is the first margin value, which can take a value of 1.2, for example.
[0137] The purpose of this global embedding representation loss is to make the distance between the first image and the second image in the global triple sample greater than the distance to the third image by the first margin value.
[0138] In step S720, a local embedding representation loss is determined based on the predicted embedding representations of the first image, the second image, and the third image in the local triple samples.
[0139] In the embodiments of the present disclosure, the local embedding representation loss L can be determined according to the following formula (2). tr2 .
[0140]
[0141] Where a 2 is the first image in the local triple sample, p 2 is the second image in the local triple sample, and n 2 is the third image in the local triple sample. represents the Euclidean distance between the predicted embedding representations of the first image and the second image in the local triple sample. is the Euclidean distance between the predicted embedding representations of the first image and the third image in the local triple sample. α 2 is the second margin value, which can take a value of 0.6, for example.
[0142] The purpose of this local embedding representation loss is to make the distance between the first image and the second image in the local triple sample greater than the distance to the third image by the first margin value.
[0143] In step S730, a first target loss is determined according to the weighted calculation result of the global embedded representation loss and the local embedded representation loss.
[0144] In the embodiments of the present disclosure, the weights of the global embedded representation loss and the local embedded representation loss can be taken as 1, or adjusted according to empirical values. The embodiments of the present disclosure do not make special limitations on this comparison. After obtaining the first target loss, for example Figure 12 in which the weight of the first target loss is 1, and the weight of the classification loss is set to 0 to obtain the weighted loss.
[0145] In this embodiment, for the global triplet samples and the local triplet samples, where the first image and the second image form a positive sample pair, and the first image and the second image form a negative sample pair. The positive sample pair has class labels with the same semantics and are the same or extremely similar images. The negative sample pair in the global triplet samples do not have the same semantics because the class labels are different, and the negative sample pair in the local triplet samples have the same semantics because the class labels are the same. The first target loss in this embodiment can make the distance between negative sample pairs with the same semantics (i.e., the negative sample pairs in the local triplet samples) closer than that between negative sample pairs with different semantics (i.e., the negative sample pairs in the global triplet samples), so that the closer the embedded representations are, the more similar the semantics are. Furthermore, the distance between classes is increased and the distance within classes is decreased, which is helpful for both the embedded representation learning and the classification learning, and can effectively maintain the relative relationship of the similarity embeddings in terms of classes and reasonably combine the classification task learning. Figure 14 Schematically shows a schematic diagram of triplet samples (including global triplet samples and local triplet samples) according to an embodiment of the present disclosure. As Figure 14 shown, C1, C2, C3, C4, C5, C6 are six categories respectively, and a, p, n1, n2 are sample images. For the global triplet sample (a, p, n2), a is the first image with a class label of C1; p is the second image with a class label of C1; n2 is the third image with a class label of C2. For the local triplet sample (a, p, n1), a is the first image with a class label of C1; p is the second image with a class label of C1; n1 is the third image with a class label of C1. The first target loss determined in the above embodiment can make the distance of the negative sample pair (a, n2) that does not belong to the same class greater than the distance of the negative sample pair (a, n1) that belongs to the same class.
[0146] Figure 8 Schematically shows a flowchart of an image processing method based on a multi-task model according to another embodiment of the present disclosure.
[0147] AsFigure 8 As shown above Figure 7 Step S710 in the above embodiment may further include the following steps.
[0148] In step S810, calculate the distance between the predicted embedded representations of the first image and the second image in the global triplet sample to obtain the positive sample pair distance of the global triplet sample.
[0149] In the embodiment of the present disclosure, the positive sample pair distance of the global triplet sample may be, for example, the
[0150] In step S820, calculate the distance between the predicted embedded representations of the first image and the third image in the global triplet sample to obtain the negative sample pair distance of the global triplet sample.
[0151] In the embodiment of the present disclosure, the negative sample pair distance of the global triplet sample may be, for example, the
[0152] In step S830, determine the global embedded representation loss according to the positive sample pair distance and the negative sample pair distance of the global triplet sample.
[0153] In the embodiment of the present disclosure, the difference between the positive sample pair distance and the negative sample pair distance of the global triplet sample may be calculated, and the sum value of the difference and the first margin value is compared with 0, and the larger one of them is determined as the global embedded representation loss.
[0154] Figure 9 Schematically shows a flowchart of an image processing method based on a multi-task model according to still another embodiment of the present disclosure.
[0155] As Figure 9 shown above Figure 7 Step S720 in the above embodiment may further include the following steps.
[0156] In step S910, calculate the distance between the predicted embedded representations of the first image and the second image in the local triplet sample to obtain the positive sample pair distance of the local triplet sample.
[0157] In the embodiment of the present disclosure, the positive sample pair distance of the local triplet sample may be, for example, the
[0158] In step S920, calculate the distance between the predicted embedded representations of the first image and the third image in the local triplet sample to obtain the negative sample pair distance of the local triplet sample.
[0159] In the embodiment of the present disclosure, the negative sample pair distance of the local triplet sample may be, for example, the
[0160] In step S930, a local embedding representation loss is determined according to the positive sample pair distance and the negative sample pair distance of the local triple samples.
[0161] In the embodiments of the present disclosure, the difference between the positive sample pair distance and the negative sample pair distance of the local triple samples can be calculated, and the sum value of the difference and the second margin value is compared with 0, and the larger one of them is determined as the local embedding representation loss.
[0162] Figure 10 A flowchart of an image processing method based on a multi-task model according to still another embodiment of the present disclosure is schematically shown.
[0163] As Figure 10 shown, step S360 in the above Figure 3 shown embodiment may further include the following steps.
[0164] In step S1010, a classification loss is determined according to the class prediction result and the class label of the sample image.
[0165] In the embodiments of the present disclosure, a cross-entropy loss can be generated based on the class prediction result and the class label as the classification loss. The classification loss L class can be expressed as the following formula (3).
[0166]
[0167] where, p ic represents the predicted probability that the sample image i belongs to the c classification, y ic represents whether the class label of the sample image i is c, if it is c, then y ic = 1, otherwise it is 0. N is the number of sample images, N is an integer greater than 0, Nc is the total number of classes, and Nc is an integer greater than 1.
[0168] In step S1020, a second target loss is determined according to the weighted calculation result of the classification loss and the first target loss.
[0169] In the embodiments of the present disclosure, the second target loss L total can be obtained, for example, by the following formula (4).
[0170] L total = w1L class + w2L tr1 + w3L tr2 (4)
[0171] where, L class is the classification loss, L tr1 is the global embedding representation loss, L tr2is the local embedded representation loss. w1, w2, and w3 are the weights of each item, which can be 1 or adjusted according to empirical values. Among them, w2L tr1 +w3L tr2 is the first target loss. The second target loss can be expressed, for example, as the weighted loss in Figure 12 . Among them, w2L tr1 +w3L tr2 is the first target loss.
[0172] Figure 11 Schematically shows a flowchart of an image processing method based on a multi-task model according to another embodiment of the present disclosure.
[0173] As Figure 11 shown, Figure 10 the step S1010 in the shown embodiment may further include the following steps.
[0174] In step S1110, determine the first prediction result of the class prediction result of the sample image under the class label of the sample image and the Nc - 1 second prediction results under the other Nc - 1 classes outside the class label of the sample image, where Nc is the total number of classes and Nc is an integer greater than 1.
[0175] In the embodiments of the present disclosure, the class prediction result may include the prediction result of the sample image under each class. For a certain sample image, the form of its class prediction result can be expressed, for example, as (0.1, 0.15, 0.2, 0.8, 0.3), where 0.1 is the probability that the sample image belongs to the first class, and 0.15 is the probability that the sample image belongs to the second class. Assume that there are 5 classes in total (i.e., Nc = 5) and the class label is the 4th class. Then the first prediction result is 0.8, 0.1 is the second prediction result under the 1st class, and 0.15 is the second prediction result under the 2nd class.
[0176] In step S1120, determine the labeled class loss according to the first prediction result, the class label of the sample image, and the first weight.
[0177] In the embodiments of the present disclosure, for each sample image, the first cross-entropy loss can be calculated according to the first prediction result and weighted by 0.7 to obtain the labeled class loss. ε = 0.7 is the first weight. The labeled class loss can be expressed as (1 - ε)*Loss1, if(i = y), where ε is set to 0.3 and Loss1 is the cross-entropy loss calculated according to the first prediction result. i = y represents the prediction probability of the labeled class.
[0178] In step S1130, Nc - 1 single unlabeled class losses are determined according to Nc - 1 second prediction results, the class label of the sample image, and the second weight. The first weight and the second weight are preset values, and the sum of the first weight and the second weight is a fixed value.
[0179] In the embodiments of the present disclosure, for each second prediction result (a total of Nc - 1), its cross - entropy loss can be calculated as its single unlabeled class loss. The single unlabeled class loss can be expressed as ε*Loss2,if(i≠y). ε is set to 0.3, and Loss2 is the cross - entropy loss calculated according to each second prediction result. Among them, the first weight is ε, the second weight is 1 - ε, and the sum of the first weight and the second weight is 1, but the present disclosure is not limited thereto.
[0180] In step S1140, the average value of the Nc - 1 single unlabeled class losses is determined as the unlabeled class loss.
[0181] In step S1150, the classification loss is determined according to the labeled class loss and the unlabeled class loss.
[0182] In the embodiments of the present disclosure, for the i - th sample, the classification of its classification loss can be expressed as shown in formula (5).
[0183]
[0184] Furthermore, substituting formula (5) into formula (3) to obtain the classification loss.
[0185] The following introduces the device embodiments of the present disclosure, which can be used to execute the above - mentioned image processing method based on the multi - task model of the present disclosure. For the details not disclosed in the device embodiments of the present disclosure, please refer to the embodiments described in the above - mentioned image processing method based on the multi - task model of the present disclosure.
[0186] Figure 15 A block diagram of an image processing device based on a multi - task model according to an embodiment of the present disclosure is schematically shown.
[0187] Refer to Figure 15 As shown, an image processing device 1500 based on a multi - task model according to an embodiment of the present disclosure may include: a sample acquisition module 1510, a feature extraction module 1520, an embedded representation module 1530, a first training module 1540, a class prediction module 1550, and a second training module 1560.
[0188] The sample acquisition module 1510 can be configured to acquire a sample image and the class label of the sample image.
[0189] The feature extraction module 1520 can be configured to process a sample image through the feature extraction structure in the multi-task model to obtain image features.
[0190] The embedded representation module 1530 can be configured to process the image features through the embedded representation structure in the multi-task model to obtain the predicted embedded representation of the sample image.
[0191] The first training module 1540 can be configured to determine a first target loss according to the predicted embedded representation, so as to adjust the parameters of the feature extraction structure and the embedded representation structure in the multi-task model to obtain the multi-task model in the first training stage.
[0192] The category prediction module 1550 can be configured to process the image features through the classification structure parallel to the embedded representation structure in the multi-task model to obtain the category prediction result of the sample image.
[0193] The second training module 1560 can be configured to determine a second target loss according to the category prediction result, the category label and the predicted embedded representation, and adjust the parameters of the feature extraction structure, the embedded representation structure and the classification structure in the multi-task model in the first training stage to obtain the trained multi-task model, so as to perform predictions on image classification and embedded representation according to the trained multi-task model.
[0194] When training the multi-task model by using the sample image and the category label of the sample image, the image processing device based on the multi-task model provided by the embodiments of the present disclosure first processes the sample image through the feature extraction structure in the multi-task model to obtain image features; then inputs the image features into the classification structure and the embedded representation structure at the same time, so that the classification structure and the embedded structure share the underlying feature extraction structure, which can save the inference time of feature extraction. And the parallel design of the classification structure and the embedded structure can reduce the influence of classification on the learning of embedded representation. In addition, a first target loss is first generated according to the predicted embedded representation to adjust the parameters of the feature extraction structure and the embedded representation structure to obtain the multi-task model in the first training stage, and then a second target loss is generated according to the predicted embedded representation, the category prediction result and the category label to adjust the parameters of the feature extraction structure, the embedded representation structure and the classification structure in the multi-task model in the first training stage to obtain the trained multi-task model. This phased training and learning method can take into account the characteristic that the embedded representation task is more difficult to converge than the classification task, and improve the effect of multi-task learning through two-stage learning of pre-training the embedded representation structure first and then fine-tuning the network by combining the classification structure, which can effectively prevent overfitting of the classification structure and ensure the embedded representation.
[0195] In an exemplary embodiment, when the first training module 1540 "determines the first target loss according to the predicted embedded representation", it may include: a sample pair generation sub-module, configured to generate a sample pair according to a sample image, the sample pair including two sample images, i.e., a first image and a second image, and the distance between the actual embedded representations of the first image and the second image being less than a distance threshold; a global triplet sub-module, configured to form a global triplet sample by combining a sample image with a different class label from the first image in the sample pair and the sample pair; a local triplet sub-module, configured to form a local triplet sample by combining a sample image with the same class label as the first image in the sample pair and the sample pair; a first loss sub-module, configured to generate a first target loss according to the predicted embedded representation of each sample image in the global triplet sample and the local triplet sample.
[0196] In an exemplary embodiment, the global triplet sub-module may include: a sample pair image generation unit, configured to, for each sample pair, randomly select a sample image from the remaining sample pairs as the sample pair image of each of the remaining sample pairs; a first sample pair image unit, configured to determine a sample pair image with a class label different from that of the first image in the sample pair as the first target sample pair image; a first distance calculation unit, configured to calculate the distance between each first target sample pair image and the first image in the sample pair to obtain the first distance of each first target sample pair image; a first distance sorting unit, configured to sort the first target sample pair images in ascending order of the first distance; a global triplet unit, configured to form a global triplet sample by combining the first a first target sample pair images in the sorting result and the sample pair respectively, where a is an integer greater than 0, and the global triplet sample includes the first image, the second image, and the first target sample pair image as the third image of the sample pair.
[0197] In an exemplary embodiment, the local triplet sub-module may include: a sample pair image generation unit, configured to, for each sample pair, randomly select a sample image from the remaining sample pairs as the sample pair image of each of the remaining sample pairs; a second sample pair image unit, configured to determine a sample pair image with a class label the same as that of the first image in the sample pair as the second target sample pair image; a second distance calculation unit, configured to calculate the distance between each second target sample pair image and the first image in the sample pair to obtain the second distance of each second target sample pair image; a second distance sorting unit, configured to sort the second target sample pair images in ascending order of the second distance; a local triplet unit, configured to form a local triplet sample by combining the first b second target sample pair images in the sorting result and the sample pair respectively, where b is an integer greater than 0, and the local triplet sample includes the first image, the second image, and the second target sample pair image as the third image of the sample pair.
[0198] In an exemplary embodiment, the first loss sub-module may include: a global embedded representation loss calculation unit, configured to determine a global embedded representation loss according to the predicted embedded representations of the first image, the second image, and the third image in the global triplet sample; a local embedded representation loss calculation unit, configured to determine a local embedded representation loss according to the predicted embedded representations of the first image, the second image, and the third image in the local triplet sample; a first loss calculation unit, configured to determine a first target loss according to the weighted calculation result of the global embedded representation loss and the local embedded representation loss.
[0199] In an exemplary embodiment, the global embedded representation loss calculation unit may include: a first positive sample pair distance calculation sub-unit, configured to calculate the distance between the predicted embedded representations of the first image and the second image in the global triplet sample to obtain the positive sample pair distance of the global triplet sample; a first negative sample pair distance calculation sub-unit, configured to calculate the distance between the predicted embedded representations of the first image and the third image in the global triplet sample to obtain the negative sample pair distance of the global triplet sample; a global embedded representation loss calculation sub-unit, configured to determine a global embedded representation loss according to the positive sample pair distance and the negative sample pair distance of the global triplet sample.
[0200] In an exemplary embodiment, the local embedded representation loss calculation unit may include: a second positive sample pair distance calculation sub-unit, configured to calculate the distance between the predicted embedded representations of the first image and the second image in the local triplet sample to obtain the positive sample pair distance of the local triplet sample; a second negative sample pair distance calculation sub-unit, configured to calculate the distance between the predicted embedded representations of the first image and the third image in the local triplet sample to obtain the negative sample pair distance of the local triplet sample; a local embedded representation loss calculation sub-unit, configured to determine a local embedded representation loss according to the positive sample pair distance and the negative sample pair distance of the local triplet sample.
[0201] In an exemplary embodiment, the second training module 1560 may include: a classification loss calculation sub-module, configured to determine a classification loss according to the class prediction result and the class label of the sample image; a second loss calculation sub-module, configured to determine a second target loss according to the weighted calculation result of the classification loss and the first target loss.
[0202] In an exemplary embodiment, the classification loss calculation sub-module may include: a class prediction result division unit, configured to determine a first prediction result of the class prediction result of the sample image under the class label of the sample image and Nc - 1 second prediction results under the other Nc - 1 classes outside the class label of the sample image, where Nc is the total number of classes and Nc is an integer greater than 1; a labeled class loss calculation unit, configured to determine a labeled class loss according to the first prediction result, the class label of the sample image, and a first weight; a single unlabeled class loss calculation unit, configured to determine Nc - 1 single unlabeled class losses according to the Nc - 1 second prediction results, the class label of the sample image, and a second weight, where the first weight and the second weight are preset values and the sum of the first weight and the second weight is a fixed value; an unlabeled class loss calculation unit, configured to determine the average value of the Nc - 1 single unlabeled class losses as the unlabeled class loss; a classification loss calculation unit, configured to determine a classification loss according to the labeled class loss and the unlabeled class loss.
[0203] In an exemplary embodiment, the classification structure may include a convolutional unit and a fully connected unit connected in sequence; wherein, the class prediction module 1550 may include: a convolutional operation sub-module, configured to perform a convolutional operation on the image features according to the convolutional unit to obtain a convolutional output; a classification prediction sub-module, configured to process the convolutional output through the fully connected unit to obtain the class prediction result of the sample image.
[0204] In an exemplary embodiment, when the second training module 1560 performs prediction of image classification and embedded representation according to the trained multi-task model, it may include: an image acquisition sub-module, configured to acquire an image to be predicted; an image prediction sub-module, configured to process the image to be predicted through the trained multi-task model to obtain the embedded representation and classification result of the image to be predicted.
[0205] Figure 16 The structural schematic diagram of an electronic device suitable for implementing the embodiments of the present disclosure is shown. It should be noted that Figure 16 The shown electronic device 1600 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0206] Such as Figure 16As shown, the electronic device 1600 includes a central processing unit (CPU) 1601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1602 or a program loaded from a storage section 1608 into a random access memory (RAM) 1603. In the RAM 1603, various programs and data required for system operation are also stored. The CPU 1601, ROM 1602, and RAM 1603 are connected to each other via a bus 1604. An input / output (I / O) interface 1605 is also connected to the bus 1604.
[0207] In some embodiments, the following components may be connected to the I / O interface 1605: an input section 1606 including a keyboard, a mouse, etc.; an output section 1607 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1608 including a hard disk, etc.; and a communication section 1609 including a network interface card such as a LAN card, a modem, etc. The communication section 1609 performs communication processing via a network such as the Internet. A drive 1610 is also connected to the I / O interface 1605 as needed. A removable medium 1611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1610 as needed so that a computer program read from it can be installed into the storage section 1608 as needed.
[0208] In particular, according to an embodiment of the present disclosure, the processes described below with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1609, and / or installed from the removable medium 1611. When the computer program is executed by a central processing unit (CPU) 1601, various functions defined in the system of the present application are executed.
[0209] It should be noted that the computer-readable medium shown in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection having at least one wire, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0210] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a portion of code that contains at least one executable instruction for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0211] The modules and / or units and / or subunits involved in the embodiments of the present disclosure can be implemented in software or in hardware, and the described modules and / or units and / or subunits can also be provided in a processor. Among them, the names of these modules and / or units and / or subunits do not constitute a limitation to the modules and / or units and / or subunits themselves in certain cases.
[0212] As another aspect, the present application also provides a computer-readable medium, which can be included in the electronic device described in the above embodiments; or can exist alone without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the methods described in the following embodiments. For example, the electronic device can implement each step as shown in Figure 3 or Figure 4 or Figure 5 or Figure 6 or Figure 7 or Figure 8 or Figure 9 or Figure 10 or Figure 11 shown.
[0213] It should be noted that although several modules or units or subunits of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units or subunits described above can be embodied in one module or unit or subunit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units or subunits.
[0214] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable an electronic device to execute the methods according to the embodiments of the present disclosure.
[0215] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0216] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. An image processing method based on a multi-task model, characterized in that Including: Obtaining a sample image and the class label of the sample image; Processing the sample image through a feature extraction structure in a multi-task model to obtain image features; Processing the image features through an embedded representation structure in the multi-task model to obtain a predicted embedded representation of the sample image; Generating a sample pair according to the sample image, where the two sample images included in the sample pair are a first image and a second image, and the distance between the actual embedded representations of the first image and the second image is less than a distance threshold; Combining a sample image with a different class label from the first image in the sample pair and the sample pair to form a global triplet sample; Combining a sample image with the same class label as the first image in the sample pair and the sample pair to form a local triplet sample; Generating a first target loss according to the predicted embedded representations of each sample image in the global triplet sample and the local triplet sample, so as to adjust the parameters of the feature extraction structure and the embedded representation structure in the multi-task model according to the first target loss, and obtaining a multi-task model in the first training stage; Processing the image features through a classification structure parallel to the embedded representation structure in the multi-task model to obtain a class prediction result of the sample image; Determining a classification loss according to the class prediction result of the sample image and the class label; Determining a second target loss according to a weighted calculation result of the classification loss and the first target loss; Adjusting the parameters of the feature extraction structure, the embedded representation structure, and the classification structure in the multi-task model in the first training stage according to the second target loss to obtain the trained multi-task model, so as to perform image classification and prediction of embedded representation according to the trained multi-task model.
2. The method according to claim 1, characterized in that, Combining a sample image with a different class label from the first image in the sample pair and the sample pair to form a global triplet sample includes: For each sample pair, randomly selecting a sample image from the remaining sample pairs as the sample pair image of each remaining sample pair; Determining a first target sample pair image whose class label is different from the class label of the first image in the sample pair; Calculating the distance between each first target sample pair image and the first image in the sample pair to obtain the first distance of each first target sample pair image; Sorting the first target sample pair images in ascending order of the first distance; Combining the first a first target sample pair images in the sorting result and the sample pair to form a global triplet samples respectively, where a is an integer greater than 0, and the global triplet sample includes the first image, the second image of the sample pair, and the first target sample pair image as the third image.
3. The method according to claim 2, characterized in that Combining a sample image with the same class label as the first image in the sample pair and the sample pair to form a local triplet sample includes: For each sample pair, randomly selecting a sample image from the remaining sample pairs as the sample pair image of each remaining sample pair; Determining a second target sample pair image whose class label is the same as the class label of the first image in the sample pair; Calculate the distance between each second target sample pair image and the first image in the sample pair to obtain the second distance of each second target sample pair image; Sort the second target sample pair images in ascending order of the second distance; Combine the first b second target sample pair images in the sorting result with the sample pair to form b local triple sample pairs respectively, where b is an integer greater than 0, and the local triple sample pair includes the first image, the second image of the sample pair, and the second target sample pair image as the third image.
4. The method according to claim 3, wherein, Generating the first target loss according to the predicted embedded representations of each sample image in the global triple sample pair and the local triple sample pair includes: Determine the global embedded representation loss according to the predicted embedded representations of the first image, the second image, and the third image in the global triple sample pair; Determine the local embedded representation loss according to the predicted embedded representations of the first image, the second image, and the third image in the local triple sample pair; Determine the first target loss according to the weighted calculation result of the global embedded representation loss and the local embedded representation loss.
5. The method according to claim 4, wherein Determining the global embedded representation loss according to the predicted embedded representations of the first image, the second image, and the third image in the global triple sample pair includes: Calculate the distance between the predicted embedded representations of the first image and the second image in the global triple sample pair to obtain the positive sample pair distance of the global triple sample pair; Calculate the distance between the predicted embedded representations of the first image and the third image in the global triple sample pair to obtain the negative sample pair distance of the global triple sample pair; Determine the global embedded representation loss according to the positive sample pair distance and the negative sample pair distance of the global triple sample pair.
6. The method according to claim 4, wherein Determining the local embedded representation loss according to the predicted embedded representations of the first image, the second image, and the third image in the local triple sample pair includes: Calculate the distance between the predicted embedded representations of the first image and the second image in the local triple sample pair to obtain the positive sample pair distance of the local triple sample pair; Calculate the distance between the predicted embedded representations of the first image and the third image in the local triple sample pair to obtain the negative sample pair distance of the local triple sample pair; Determine the local embedded representation loss according to the positive sample pair distance and the negative sample pair distance of the local triple sample pair.
7. The method according to claim 1, wherein Determining the classification loss according to the class prediction result of the sample image and the class label includes: Determine the first prediction result of the class prediction result of the sample image under the class label of the sample image and the Nc - 1 second prediction results under the other Nc - 1 classes outside the class label of the sample image, where Nc is the total number of classes and Nc is an integer greater than 1; Determine the labeled class loss according to the first prediction result, the class label of the sample image, and the first weight; Determine Nc - 1 single unlabeled class losses according to the Nc - 1 second prediction results, the class label of the sample image, and the second weight, where the first weight and the second weight are preset values, and the sum of the first weight and the second weight is a fixed value; Determine the average value of the Nc-1 single unlabeled class losses as the unlabeled class loss; Determine the classification loss according to the labeled class loss and the unlabeled class loss.
8. The method according to claim 1, wherein The classification structure includes a convolutional unit and a fully connected unit connected in sequence; wherein, processing the image features through the classification structure parallel to the embedded representation structure in the multi-task model to obtain the class prediction result of the sample image includes: Perform a convolution operation on the image features according to the convolutional unit to obtain a convolution output; Process the convolution output through the fully connected unit to obtain the class prediction result of the sample image.
9. The method according to claim 1, wherein Performing predictions of image classification and embedded representation according to the trained multi-task model includes: Obtain the image to be predicted; Process the image to be predicted through the trained multi-task model to obtain the embedded representation and classification result of the image to be predicted.
10. An image processing apparatus based on a multi-task model, characterized in that, Including: A sample acquisition module configured to acquire a sample image and the class label of the sample image; A feature extraction module configured to process the sample image through the feature extraction structure in the multi-task model to obtain image features; An embedded representation module configured to process the image features through the embedded representation structure in the multi-task model to obtain the predicted embedded representation of the sample image; A first training module configured to generate sample pairs according to the sample image, the two sample images included in the sample pair being a first image and a second image, the distance between the actual embedded representations of the first image and the second image being less than a distance threshold; forming a global triplet sample by combining a sample image with a different class label from the class label of the first image in the sample pair and the sample pair; forming a local triplet sample by combining a sample image with the same class label as the first image in the sample pair and the sample pair; generating a first target loss according to the predicted embedded representations of each sample image in the global triplet sample and the local triplet sample, so as to adjust the parameters of the feature extraction structure and the embedded representation structure in the multi-task model according to the first target loss, and obtain the multi-task model in the first training stage; A class prediction module configured to process the image features through the classification structure parallel to the embedded representation structure in the multi-task model to obtain the class prediction result of the sample image; A second training module configured to determine a classification loss according to the class prediction result and the class label of the sample image; determine a second target loss according to the weighted calculation result of the classification loss and the first target loss, and adjust the parameters of the feature extraction structure, the embedded representation structure and the classification structure in the multi-task model in the first training stage according to the second target loss to obtain the trained multi-task model, so as to perform predictions of image classification and embedded representation according to the trained multi-task model.
11. An electronic device, characterized in that, Including: At least one processor; A storage device for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1-9.
12. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the method according to any one of claims 1-9 is implemented.
Citation Information
Patent Citations
Project name and category identification method and device
CN112085012A
Text-to-Visual Machine Learning Embedding Techniques
US20200380298A1