Image annotation model training and image annotation method and device
By training an image annotation model that includes feature extraction, prototype extraction, correction and classification modules, and using prototype vectors to cross-compare image features, the inaccurate coverage problem caused by image-level label training in the prior art is solved, and efficient pixel-level annotation and more stable segmentation results are achieved.
Patent Information
- Application Number
- CN202110976261.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-08-24
AI Technical Summary
The prior art weakly supervised image segmentation trained with image-level tags in image processing usually causes inaccurate coverage of the target area, resulting in high labor consumption and low efficiency.
By training a image annotation model including feature extraction module, prototype extraction module, correction module and classification module, the image feature cross-comparison is used for image features, and pixel-level annotation is automatically performed to reduce the need for manual annotation.
More accurate target and background separation is achieved, manpower consumption is reduced, image annotation efficiency is improved, and model parameters are adjusted by minimizing loss, improving the stability of segmentation results.
Smart Images

Figure CN113673607B_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present specification relate to the field of computer technology, and in particular, to training of image annotation models and methods and devices for image annotation. Background Art
[0002] Image processing has a wide range of applications in daily production or life. For example: laser anti-counterfeiting detection based on laser and other technologies, document region segmentation, panoramic segmentation, target recognition, etc. In these applications, weakly supervised image segmentation trained with image-level labels often causes inaccurate coverage of the target area in the process of generating pseudo-true values. This is because the object activation map is trained with categorical targets and lacks the ability to generalize. In order to more accurately separate the target from the background, pixel-level annotation is usually involved, that is, pixel-by-pixel annotation of categories. This may bring great labor consumption to the annotation work. Summary of the invention
[0003] One or more embodiments of this specification describe a method and device for training an image annotation model, and a method and device for object annotation using the trained image annotation model, in order to solve one or more problems mentioned in the background technology.
[0004] According to a first aspect, a training method for an image annotation model is provided, wherein the image annotation model is used to perform pixel-level annotation on images with classification labels, the image annotation model includes a feature extraction module, a prototype extraction module, a correction module, and a classification module, and the method includes: obtaining a first image and a second image from a sample set, wherein the first image and the second image both have an image-level first category label; processing the first image and the second image respectively through a pre-trained feature extraction module to obtain a corresponding first feature map and a second feature map; and extracting a plurality of prototype vectors from the first feature map and the second feature map respectively using the prototype extraction module, wherein a single prototype vector corresponds to a single feature point on the corresponding feature map. , and corresponding to the corresponding activation value that satisfies the activation condition; through the correction module, for each prototype vector extracted from the first feature map and the second feature map, a pairwise similarity comparison is performed, and according to the maximum similarity between a single prototype vector and other prototype vectors, the first feature map and the second feature map are respectively corrected to obtain a first corrected feature map and a second corrected feature map; according to the first corrected feature map and the second corrected feature map, the first image and the second image are respectively classified by the classification module to obtain respective corresponding classification results, and the classification results include pixel-level annotation results; based on the classification results, the model loss of the image annotation model is determined, so as to adjust the undetermined parameters of the image annotation model with the goal of minimizing the model loss.
[0005] In one embodiment, the feature extraction module includes a first convolution block composed of multiple convolution layers, the convolution results of each convolution layer in the first convolution block have the same number of channels, the first feature map includes each convolution result of the multiple convolution layers in the first convolution block performing convolution operations on the first image, and the second feature map includes each convolution result of the multiple convolution layers in the first convolution block performing convolution operations on the second image.
[0006] In one embodiment, the use of the prototype extraction module to extract multiple prototype vectors from the first feature map and the second feature map respectively includes extracting multiple prototype vectors from the first feature map in the following manner: detecting the activation values corresponding to the feature points in the first feature map, wherein the single activation value of a single feature point is positively correlated with the absolute value of the feature value of the single feature point in each channel; selecting multiple feature points from the feature points that meet the activation conditions as candidate feature points; and for a single candidate feature point, constructing a corresponding single prototype vector according to its feature value in each channel.
[0007] In one embodiment, the activation condition is that the activation value is greater than a predetermined activation threshold; the selecting of multiple feature points from the feature points that meet the activation condition as candidate feature points includes at least one of the following: taking all feature points whose activation values are greater than the predetermined activation threshold as candidate feature points; randomly selecting a predetermined number of feature points from the feature points whose activation values are greater than the predetermined activation threshold as candidate feature points; and selecting a predetermined number of feature points from the feature points whose activation values are greater than the predetermined activation threshold as candidate feature points in descending order of activation values.
[0008] In one embodiment, the first feature map and the second feature map are corrected according to the maximum similarity between the single prototype vector and other prototype vectors to obtain the first corrected feature map and the second corrected feature map, respectively, including: for a single prototype vector, using the maximum similarity between the single prototype vector and other prototype vectors as the confidence of the eigenvalue of the corresponding single feature point on the first feature map / the second feature map; correcting each eigenvalue of the single feature point in the first corrected feature map / the second corrected feature map according to the product of the confidence and the corresponding eigenvalue, thereby correcting the first feature map and the second feature map to the corresponding first corrected feature map and the second corrected feature map, respectively.
[0009] In one embodiment, the model loss includes a first loss for a first image and a second loss for the second image. For the first image, the classification result includes a first labeling result at the pixel level and a first classification result at the image level. The first loss includes: a first classification loss determined by comparing the first classification result with the first category label; and a first correction loss determined by comparing the first labeling result with a second labeling result determined using the first feature map.
[0010] In one embodiment, the first classification loss is determined by a cross entropy between the first classification result and the first category label.
[0011] In one embodiment, the first correction loss is determined by: processing the first feature map through the classification module to obtain a second labeling result at the pixel level; comparing the labeling difference between the first labeling result and the second labeling result pixel by pixel; and determining the first correction loss by using the sum of the labeling differences corresponding to each pixel.
[0012] In one embodiment, the first annotation result is a result obtained by classifying the first image by using a classification module through boundary thinning.
[0013] In one embodiment, the undetermined parameters of the image annotation model include the undetermined parameters in the prototype extraction module, the correction module, and the classification module.
[0014] In one embodiment, the method further includes: based on the current batch of training samples including the first image and the second image, detecting each model loss corresponding to the current batch of training samples and multiple consecutive batches of training samples; when the change in the sliding average of each model loss is less than a predetermined loss value, determining that the image annotation model training is completed.
[0015] According to a second aspect, a method for image annotation is provided, which is used to perform pixel-level annotation on images with classification labels in a sample set through a pre-trained image annotation model, wherein the image annotation model includes a feature extraction module, a prototype extraction module, a correction module, and a classification module, and the method includes: obtaining a first image and a second image from the sample set, wherein the first image and the second image both have an image-level first category label; processing the first image and the second image through the pre-trained feature extraction module to obtain a first feature map and a second feature map, respectively; using the prototype extraction module, extracting multiple prototype vectors from the first feature map and the second feature map, respectively, a single prototype vector corresponds to a single feature point on the corresponding feature map, and has a corresponding activation value that satisfies an activation condition; through the correction module, performing pairwise similarity comparisons on each prototype vector extracted from the first image and the second image, and correcting the first feature map and the second feature map according to the maximum similarity between the single prototype vector and other prototype vectors to obtain a first corrected feature map and a second corrected feature map; using the classification module to classify the first image and the second image according to the first corrected feature map and the second corrected feature map, respectively, to obtain the corresponding pixel-level annotation results.
[0016] According to a third aspect, a training method for an image annotation model is provided, wherein the image annotation model is used to annotate an image with a classification label at the pixel level, the image annotation model includes a feature extraction module, a prototype extraction module, a correction module, and a classification module, and the method includes: obtaining a first image from a sample set, wherein the first image corresponds to a first category label; processing the first image by a pre-trained feature extraction module to obtain a first feature map; extracting multiple prototype vectors from the first feature map by using the prototype extraction module, wherein a single prototype vector corresponds to a single feature point on the corresponding feature map and has a corresponding activation value that satisfies an activation condition; performing similarity comparisons with respective reference vectors in a reference vector set for each prototype vector extracted from the first image by using the correction module, and correcting the first feature map to obtain a first corrected feature map according to the maximum similarity between the single prototype vector and each reference vector, wherein the reference vector is extracted from the image corresponding to the first category label; classifying the first image by using the classification module according to the first corrected feature map to obtain a first classification result, wherein the first classification result includes a first annotation result at the pixel level; determining a model loss of the image annotation model based on the first classification result, thereby adjusting the undetermined parameters of the image annotation model with the goal of minimizing the model loss.
[0017] In one embodiment, the benchmark vectors in the benchmark vector set are determined in the following manner: using a pre-trained feature extraction module to extract corresponding feature maps for each image in the sample set; in each feature map, selecting each candidate feature point whose activation value is greater than a first activation threshold, and the first activation threshold is compared with the activation condition to filter out feature points with higher activation values; for a single candidate feature point, according to its feature value in each channel, constructing a corresponding single benchmark vector, and adding it to the benchmark vector set.
[0018] In one embodiment, the method further includes: for the first prototype vector extracted from the first feature map, detecting whether its first maximum similarity with each benchmark vector is greater than a predetermined similarity threshold; when the first maximum similarity is greater than the predetermined similarity threshold, adding the first prototype vector as a benchmark vector to the benchmark vector set.
[0019] In one embodiment, the model loss includes a first loss for a first image, and for the first image, the classification result includes a first labeling result at the pixel level and a first classification result at the image level, and the first loss includes: a first classification loss determined by comparing the first classification result with the first category label; and a first correction loss determined by comparing the first labeling result with a second labeling result determined using the first feature map.
[0020] In one embodiment, the method further includes: based on the current batch of training samples including the first image, detecting each model loss corresponding to the current batch of training samples and multiple consecutive batches of training samples; when the change in the sliding average of each model loss is less than a predetermined loss value, determining that the image annotation model training is completed.
[0021] According to a fourth aspect, a method for image annotation is provided, for annotating images with classification labels in a sample set at the pixel level through a pre-trained image annotation model, the image annotation model comprising a feature extraction module, a prototype extraction module, a correction module, and a classification module, the method comprising: obtaining a first image from the sample set, wherein the first image corresponds to a first category label; processing the first image through the pre-trained feature extraction module to obtain a first feature map; extracting multiple prototype vectors from the first feature map using the prototype extraction module, wherein a single prototype vector corresponds to a single feature point on the corresponding feature map and has a corresponding activation value that satisfies an activation condition; performing a similarity comparison with each benchmark vector in a benchmark vector set for each prototype vector extracted from the first image through the correction module, and correcting the first feature map to obtain a first corrected feature map according to the maximum similarity between the single prototype vector and each benchmark vector; and classifying the first image using the classification module according to the first corrected feature map to obtain a first pixel-level annotation result for the first image.
[0022] According to a fifth aspect, a training device for an image annotation model is provided, wherein the image annotation model is used to annotate an image with a classification label at the pixel level, wherein the image annotation model comprises a feature extraction module, a prototype extraction module, a correction module, and a classification module, and the device comprises:
[0023] An acquisition unit, configured to acquire a first image and a second image from a sample set, wherein both the first image and the second image have a first category label at the image level;
[0024] A feature extraction unit is configured to process the first image and the second image through a pre-trained feature extraction module to obtain a first feature map and a second feature map respectively;
[0025] A prototype extraction unit is configured to use the prototype extraction module to extract multiple prototype vectors from the first feature map and the second feature map respectively, where a single prototype vector corresponds to a single feature point on the corresponding feature map and has a corresponding activation value that satisfies the activation condition;
[0026] A correction unit is configured to perform a pairwise similarity comparison on each prototype vector extracted from the first image and the second image through the correction module, and correct the first feature map and the second feature map respectively according to the maximum similarity between a single prototype vector and other prototype vectors to obtain a first corrected feature map and a second corrected feature map;
[0027] a classification unit configured to classify the first image and the second image respectively using a classification module according to the first corrected feature map and the second corrected feature map to obtain respective corresponding classification results, wherein the classification results include pixel-level annotation results;
[0028] An adjustment unit is configured to determine a model loss of the image annotation model based on the classification result, thereby adjusting the undetermined parameters of the image annotation model with the goal of minimizing the model loss.
[0029] According to a sixth aspect, a device for image annotation is provided, which is used to perform pixel-level annotation on images with classification labels in a sample set by using a pre-trained image annotation model, wherein the image annotation model includes a feature extraction module, a prototype extraction module, a correction module, and a classification module, and the device includes:
[0030] An acquisition unit, configured to acquire a first image and a second image from a sample set, wherein both the first image and the second image have a first category label at the image level;
[0031] A feature extraction unit is configured to process the first image and the second image through a pre-trained feature extraction module to obtain a first feature map and a second feature map respectively;
[0032] A prototype extraction unit is configured to use the prototype extraction module to extract multiple prototype vectors from the first feature map and the second feature map respectively, where a single prototype vector corresponds to a single feature point on the corresponding feature map and has a corresponding activation value that satisfies the activation condition;
[0033] A correction unit is configured to perform a pairwise similarity comparison on each prototype vector extracted from the first image and the second image through the correction module, and correct the first feature map and the second feature map respectively according to the maximum similarity between a single prototype vector and other prototype vectors to obtain a first corrected feature map and a second corrected feature map;
[0034] The labeling unit is configured to classify the first image and the second image respectively according to the first corrected feature map and the second corrected feature map by using a classification module to obtain respective corresponding pixel-level labeling results.
[0035] According to a seventh aspect, a training device for an image annotation model is provided, wherein the image annotation model is used to annotate an image with a classification label at the pixel level, wherein the image annotation model comprises a feature extraction module, a prototype extraction module, a correction module, and a classification module, and the device comprises:
[0036] An acquisition unit is configured to acquire a first image from a sample set, wherein the first image corresponds to a first category label;
[0037] a feature extraction unit, configured to process the first image through a pre-trained feature extraction module to obtain a first feature map;
[0038] A prototype extraction unit is configured to extract a plurality of prototype vectors from the first feature map using the prototype extraction module, wherein a single prototype vector corresponds to a single feature point on the corresponding feature map and has a corresponding activation value that satisfies an activation condition;
[0039] A correction unit is configured to compare the similarity of each prototype vector extracted from the first image with each reference vector in the reference vector set through the correction module, and correct the first feature map according to the maximum similarity between the single prototype vector and each reference vector to obtain a first corrected feature map;
[0040] a classification unit, configured to classify the first image using a classification module according to the first corrected feature map to obtain a first classification result, wherein the first classification result includes a first pixel-level annotation result;
[0041] An adjustment unit is configured to determine a model loss of the image annotation model based on the first classification result, thereby adjusting the undetermined parameters of the image annotation model with the goal of minimizing the model loss.
[0042] According to an eighth aspect, a device for image annotation is provided, which is used to perform pixel-level annotation on images with classification labels in a sample set by using a pre-trained image annotation model, wherein the image annotation model includes a feature extraction module, a prototype extraction module, a correction module, and a classification module, and the device includes:
[0043] An acquisition unit is configured to acquire a first image from a sample set, wherein the first image corresponds to a first category label;
[0044] a feature extraction unit, configured to process the first image through a pre-trained feature extraction module to obtain a first feature map;
[0045] A prototype extraction unit is configured to extract a plurality of prototype vectors from the first feature map using the prototype extraction module, wherein a single prototype vector corresponds to a single feature point on the corresponding feature map and has a corresponding activation value that satisfies an activation condition;
[0046] A correction unit is configured to compare the similarity of each prototype vector extracted from the first image with each reference vector in the reference vector set through the correction module, and correct the first feature map according to the maximum similarity between the single prototype vector and each reference vector to obtain a first corrected feature map;
[0047] The labeling unit is configured to classify the first image using a classification module according to the first corrected feature map to obtain a first labeling result at the pixel level for the first image.
[0048] According to a ninth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method described in any one of the first to fourth aspects.
[0049] According to the tenth aspect, a computing device is provided, comprising a memory and a processor, wherein executable code is stored in the memory, and when the processor executes the executable code, the method described in any one of the first to fourth aspects is implemented.
[0050] Through the device and method provided in the embodiments of this specification, for the images in the training set, pixel-level annotation can be performed through image-level category labels. In the specific image annotation model training and image annotation process, the features between different images are cross-compared through prototype vectors, so as to further explore the target area in the image, and non-target areas can also be screened out to achieve weakly supervised segmentation tasks. In the loss determination process, not only the classification loss is considered, but also the similarity between the corrected segmentation result and the original segmentation result, so that the segmentation result is more stable. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0052] Figure 1 It is a schematic diagram of an application scenario under the technical conception of this specification;
[0053] Figure 2 A schematic diagram of an implementation architecture according to the technical concept of this specification is shown;
[0054] Figure 3 A schematic diagram showing the training process of an image annotation model according to an embodiment of the present specification;
[0055] Figure 4 A schematic diagram of a specific network architecture under the concept of cross-comparison of multiple images is shown;
[0056] Figure 5 A schematic diagram showing a prototype vector extraction of a specific example;
[0057] Figure 6 A schematic diagram of an image annotation process according to an embodiment of the present specification is shown;
[0058] Figure 7 A schematic diagram showing the training process of an image annotation model according to an embodiment of the present specification;
[0059] Figure 8 A schematic diagram of a specific network architecture under the concept of processing a single image at a time is shown;
[0060] Fig. 9 A schematic diagram of an image annotation process according to another embodiment of the present specification is shown;
[0061] Fig.10 A schematic diagram showing the image annotation effect under the technical concept of this specification;
[0062] Fig.11 A schematic block diagram showing a training device for an image annotation model according to an embodiment of the present specification;
[0063] Fig.12 A schematic block diagram of an image annotation device according to an embodiment of the present specification is shown. DETAILED DESCRIPTION
[0064] The solution provided in this specification is aimed at the business scenario of target recognition of images. The solution provided in this specification is described below in conjunction with the accompanying drawings. It is worth noting that the technical solution of this application involves image processing, and the accompanying drawings involve some images or computer screen shots. In order to illustrate more clearly, the color blocks of these images are not eliminated, and the clarity after conversion into grayscale images does not affect the expression of the essence of the solution.
[0065] Figure 1 An example of an application scenario of the technical architecture of this specification is shown. The application scenario is a laser pattern area recognition scenario of a certificate. Figure 1 As shown in the figure, the left side is a photo of a certificate. In order to protect data privacy, the important information of the certificate in the figure is covered by color blocks. The purpose of this scene in the application is to mark the recognition target area on the right side from the certificate image on the left. The recognition target varies depending on the anti-counterfeiting settings of the certificate. Figure 1 In the example of , the anti-counterfeiting identification target is divided into two areas. The correspondence between the two areas and the document image area is represented by two-way arrow lines 101 and 102, respectively, and the arrows at both ends of the same line point to corresponding areas.
[0066] Understandably, except Figure 1 The document anti-counterfeiting mark recognition scenario shown in the figure, the technical architecture of this specification can also be applied to document segmentation (such as identifying text areas, table areas, image areas, etc. in segmented documents), panoramic segmentation and other scenarios, without limitation here.
[0067] refer to Figure 1As shown, in order to identify the target in the image, the corresponding image recognition model can be trained. In conventional technology, the weakly supervised image recognition model (or image segmentation model) is directly trained with image-level labels (such as classification labels for the entire image), which usually produces inaccurate coverage of the target area in the process of generating the prediction result. This is because the classification label lacks the ability to generalize. In some schemes, training samples are also annotated with annotation boxes. However, many target recognition scenarios may require more accurate target area recognition. It can be understood that there may be overlaps between the recognition targets and other entities in many images. In similar situations, for the accuracy of the image recognition model, pixel-by-pixel annotated training samples are more needed to train the image recognition model. When annotating an image at the pixel level, if it is labeled pixel by pixel by manpower, although a more accurate annotation result can be obtained, the progress is slow and it takes a lot of effort. In this way, the ability requirements and workload of the annotation personnel are greatly improved, thereby increasing the labor cost and reducing the efficiency of model training.
[0068] If target category annotation is performed on the image, such as whether an image contains a real laser area of a certificate, the annotation speed will be greatly accelerated. Accordingly, this specification aims to train an image annotation model to further perform pixel-by-pixel annotation on images in a sample set that are annotated with image-level category labels (such as real certificates containing anti-counterfeiting groups, non-real certificates that do not contain anti-counterfeiting patterns, and other annotation results). In other words, the image annotation model can be used to annotate targets in an image, such as annotating the anti-counterfeiting code area, text / picture / table area, or other target areas pixel by pixel. Sample images annotated pixel by pixel can be used to train image recognition models.
[0069] It can be understood that, usually, the diversity of the target area is much smaller than the diversity of the background. For example, in the recognition scenario where the target is a person, the area of the person in the image can include the head, torso, arms, etc., and the contours of the target area have similarities, while the background can be the sky, roads, various buildings, woods, grass, etc. Therefore, if the target area is extracted as a prototype or reference, with the help of the diversity of training images, the information of the prototype can be well propagated between different images during training, and similar areas can be activated from other samples. On the contrary, if the prototype corresponds to the background area that is mistakenly activated by the classified target, since the diversity of the background is obviously greater, these prototypes will find it difficult to find other prototypes that they are similar to, and they will be filtered out, and their information will propagate much slower than that of the target area, or even not propagated.
[0070] like Figure 2In the given conceptual example, after image B is processed by the convolutional neural network CNN, only the head region 201 of the bird is activated, and the body region 202 is not activated, while in image A, both the head and body regions of the bird are activated. If the body features in image A (such as the corresponding body regions 203, 204, etc.) are used to guide image B to find its inactivated body regions and highlight them, a fully activated annotated image of the bird region as shown in image C can be obtained. Based on this concept, a regional prototype extraction network can be constructed to extract prototypes corresponding to the feature regions of the target image and cross-compare these prototypes. In other words, it is proposed to capture prototypes of unique feature regions related to the target in sample images of the same target category, so that prototypes extracted from different samples can be cross-utilized to enhance feature regions or filter non-feature regions in their feature maps, so that the target features of the context in the training set can be fully utilized.
[0071] The technical concept provided in this specification proposes a solution of exploring the cross-image target transferability between training samples using a regional prototype network. Specifically, in the model training process, training samples with only image-level classification labels are used, for example, the classification label corresponding to a single training sample is a target category (such as sheep, bird, correct anti-counterfeiting mark, etc.). The target features are extracted through such training samples to construct a prototype vector set about the target. Then, the similarity between other feature vectors of the training samples and each prototype vector is detected. On the one hand, by utilizing the objective fact that the diversity of the target area is much smaller than the diversity of the background, a higher confidence is transferred between similar prototype vectors of different sample images, so that more target-related areas are activated. On the other hand, by utilizing the characteristic that the diversity of the background is much greater than the diversity of the target area, a lower confidence is transferred in areas without similar prototype vectors in different sample images, so that non-target areas are filtered. Therefore, the target can be automatically annotated at the pixel level on the basis of the image-level label through the cross-transfer of features between images.
[0072] This technical concept can identify similar object parts (i.e. target-related areas) in different images by comparing regional features, so that the identified target areas are propagated between images with high confidence to discover new target areas, and non-target areas with low confidence. This method simplifies the annotation process of training samples based on the commonality of target areas and saves labor costs.
[0073] The technical details of the design concept of this specification are described below in conjunction with specific embodiments.
[0074] Please refer to Figure 3 As shown, Figure 3A schematic diagram of the training process of an image annotation model according to an embodiment is given. The execution subject of the process can be any computer, device or server with certain computing power. It can be understood that the training process of the image annotation model can include training of multiple batches of training samples. In conventional technology, at least one sample image is used in a batch, and in a single batch, image samples are often processed one by one, and the model loss is determined for a batch of images. According to the implementation architecture of this specification, in Figure 3 In the illustrated embodiment, at least two sample images are used in a batch, and in a single batch, image samples are cross-processed in pairs to determine the model loss for a batch of images.
[0075] In general, the image annotation model can include feature extraction module, prototype extraction module, correction module, and classification module. Figure 3 As shown, the training process of the image annotation model may include: step 301, obtaining a first image and a second image from a sample set, wherein the first image and the second image both have a first category label; step 302, processing the first image and the second image respectively through a pre-trained feature extraction module to obtain a first feature map and a second feature map corresponding to each other; step 303, using a prototype extraction module, extracting a plurality of prototype vectors from the first feature map and the second feature map respectively, wherein a single prototype vector corresponds to a single feature point on the corresponding feature map and corresponds to a corresponding activation value that satisfies the activation condition; step 304, through a correction module, extracting a plurality of prototype vectors from the first feature map and the second feature map respectively, wherein a single prototype vector corresponds to a single feature point on the corresponding feature map and corresponds to a corresponding activation value that satisfies the activation condition; step 305, and each prototype vector extracted from the second feature map is compared for similarity in pairs, and the first feature map and the second feature map are respectively corrected according to the maximum similarity between a single prototype vector and other prototype vectors to obtain a first corrected feature map and a second corrected feature map; step 305, according to the first corrected feature map and the second corrected feature map, the first image and the second image are respectively classified by a classification module to obtain respective corresponding classification results, and the classification results include pixel-level annotation results; step 306, based on the classification results, the model loss of the image annotation model is determined, so as to adjust the pending parameters of the image annotation model with the goal of minimizing the model loss.
[0076] In order to more clearly describe the training process of the image annotation model, the following Figure 3 In order to describe the relevant steps more intuitively, Figure 4 A process architecture diagram is shown. Figure 4 This is a process architecture diagram of a specific example in the training process of the image annotation model, which does not represent the only architecture.
[0077] First, in step 301, a first image and a second image are obtained from a sample set, wherein both the first image and the second image have a first category label. It can be understood that in order to perform target recognition on an image, the training sample used can be pre-labeled with a category label. The category label can be manually labeled or labeled by a pre-trained category labeling model. For the accuracy of the labeling, manual labeling is usually performed. The category label is an image-level labeling result, that is, for the overall labeling result of the entire image, the category label is, for example: target (such as sheep) image, non-target (such as non-sheep) image, or target 1 (such as sheep), target 2 (such as cow), target 3 (such as horse), non-target (such as non-cow, sheep, horse, etc.), or digital form of 1, 0, vector form such as (1, 0, 0), and so on.
[0078] Under the technical framework of this specification, in order to make full use of the common features of the targets in different sample images to accurately mark the relevant areas, images of the same category can be selected for cross-processing. It is assumed here that the images selected for cross-processing include a first image and a second image with a first category label. The first category label here can be a label corresponding to any one of the recognition targets. The first image and the second image can be any two images from the training samples of the first category label, which can be randomly selected images or two images obtained according to a predetermined rule (such as arrangement order), which is not limited here. Figure 4 In the figure, there are images 401 and 402 of the target “sheep”.
[0079] It is worth noting that, under the implementation framework of this specification, the number of images to be cross-processed at a time is at least 2, and in practice it can be two or more than two. Figure 3 , Figure 4 In the illustrated embodiment, for the convenience of description, the cross processing of two images is taken as an example for description.
[0080] Next, through step 302, the first image and the second image are processed respectively by a pre-trained feature extraction module to obtain a first feature map corresponding to the first image and a second feature map corresponding to the second image. The purpose of the feature extraction module is to extract features related to the target, such as head features and ear features of a sheep. It can be understood that the target recognition process is usually the result of a joint judgment of multiple features, and the features extracted by the feature extraction module are usually related to the recognition of the target. In practice, the features extracted by the feature extraction module may not be divided into regions based on the head, ears, etc. visible to the human eye, but other features customized by the model through deep learning.
[0081] In order to extract features from the first image and the second image, a feature extraction module can be pre-trained. In this specification, the feature extraction module can be the front part of the classification model for classifying the target. It can be understood that the classification model can be used to extract various target-related features from the image, and fuse the extracted features to obtain an output result consistent with the corresponding category label as the basis for classification. In the field of image processing, the classification model can often be implemented by a convolutional neural network, and the features on the image are extracted via the convolution kernel. After the classification model is trained, it can be considered that the target features on the image can be extracted using the corresponding convolution kernel. These extracted features can be visible to the naked eye and can be distinguished by concrete features, such as the head features and leg features of the target "sheep", etc., or they can be abstract and cannot be recognized by the naked eye. Other features are not limited here.
[0082] The trained classification model can be considered that the first half is used for feature extraction, and the second half is used for feature fusion and classification processing. Here, the first half and the second half can be pre-divided parts, for example, called feature extraction module and classification module respectively. Those skilled in the art can understand that for specific neural networks, especially deep neural networks, there is often no clear division between the first half and the second half. In this way, several layers arranged in the front (such as 10 layers) can be taken as the first half, that is, the feature extraction part, as the feature extraction module.
[0083] The feature extraction module extracts relevant features from the image and can generate a corresponding feature map to represent it. Generally, the feature map can be an array with more channels than the sample image and the resolution of a single channel is less than the resolution of the sample image. For example, the sample image is 1960×1024 pixels, with three channels of R, G, and B, which can be recorded as 1960×1024×3-dimensional data. The feature map has 512 channels, and the number of feature points of a single channel is 64×64, which is recorded as 64×64×512-dimensional feature data. In this way, the data corresponding to a feature point can represent the characteristics of an area composed of multiple pixels. For example, a feature point describes the regional characteristics of 30×16 pixels. In this step 302, the feature data obtained for the first image and the second image are respectively referred to as the first feature map and the second feature map.
[0084] In practice, in some embodiments, the feature map may also be consistent with the pixels of the sample image, or the number of feature points may be greater than the pixels (in this case, the feature of a pixel may be represented by multiple feature points), which is not limited here.
[0085] In the feature extraction module, different convolutional layers extract different features. Usually, the convolutional layers arranged in the front can extract more detailed features, and the convolutional layers arranged in the back have a larger receptive field. Therefore, according to a possible design, in order to extract feature maps with different meanings, feature data can be collected separately in multiple convolutional layers and used together as feature maps of the corresponding image. Figure 4 As shown, n different feature maps 403 can be obtained through n different convolutional layers, denoted as f1 to f n . f1 to f in feature map 403 n They can be collectively referred to as n first feature maps corresponding to the first image 401, or collectively referred to as first feature maps. As can be seen from the foregoing, the first feature map can extract the feature map corresponding to the target in the first image. For example, in the feature map 403, each small circle can represent each feature point.
[0086] In one embodiment, the number of channels of the first feature maps taken from different convolutional layers is consistent, for example, 512 channels, which makes it easier for features between different images to activate each other. In a specific example, the convolutional neural network can be divided into "blocks", and the number of channels of the feature maps output by each convolutional layer in a single "block" is consistent. In this way, the corresponding n feature maps can be determined based on the outputs of different convolutional layers in the same "block". Where n is a positive integer. Figure 4 In FIG. 4 , the feature map 403 taken for the first image 401 may be a feature map output by each convolution layer in a convolution block.
[0087] It can be understood that for different images, the size of the pixel area corresponding to the target in the image frame is also different. Therefore, by combining multiple feature maps to describe an image, on the one hand, a feature map that can extract more detailed features can be selected, and on the other hand, a feature map with a larger receptive field can be selected. n When the resolutions of the feature maps are inconsistent, the resolutions can be kept consistent by interpolation, upsampling, downsampling, etc. In this way, the areas corresponding to the single feature points of each feature map are consistent.
[0088] Similarly, the second image 402 may also correspond to n feature maps as second feature maps, which will not be described in detail here.
[0089] Furthermore, through step 303, the prototype extraction module is used to extract multiple prototype vectors from the first feature map and the second feature map respectively. Among them, prototype vectors can be understood as vectors used to represent unit areas on an image. Since for a certain target (such as a sheep), the area on the image that describes a certain part of the target (such as the head) has similarity, the purpose of extracting prototype vectors is to use prototype vectors to represent feature areas, so as to use the similarity of vectors to find feature areas that are not extracted in one image through feature areas of another image, or to filter out non-feature areas that are mistakenly extracted as feature areas.
[0090] In order to represent the feature region with a prototype vector, the vector of the corresponding feature point can be determined by the eigenvalues of each channel of the corresponding feature point on the feature map. In the case where the probability of the feature point being mapped to the target region in the image is high, its vector can be extracted as a feature vector.
[0091] It is understandable that not all feature points on the feature map correspond to the feature area of the target. Figure 2 In the case where the target is a pigeon, the area outside the pigeon, including the pebbles and beach area, does not belong to the characteristic area of the target pigeon. Figure 4 In the figure, the areas outside the goats in image 401 and image 402 do not belong to the target area. That is to say, in the feature extraction process, whether each feature point is mapped to the feature area of the target needs to be represented by a probability value or a credibility value. This probability value or credibility value also indicates the importance of the current feature point to the corresponding target. In this specification, the value used to represent the probability or credibility of whether the feature point is mapped to the feature area of the target can be recorded as an activation value. The larger the activation value, the higher the probability that the corresponding feature point corresponds to the feature area of the target in the image. Those skilled in the art can understand that since the feature extraction module is pre-trained and the activation value is determined by the feature map extracted by the feature extraction module, how to make the size of the activation value able to express the probability that the corresponding feature point corresponds to the feature area of the target in the image can be controlled by the loss function and the overall network architecture during the pre-training process of the feature extraction model, which will not be repeated here. Depending on the network settings, the method of determining the activation value is also different.
[0092] In one embodiment, during training, the feature extraction module sets one of the channels as an activation channel, and the value on the activation channel represents the size of the activation value of the corresponding feature point. For example, in the example of the 64×64×512-dimensional feature map mentioned above, one of the channels represents the activation channel, and for a single feature point among the 64×64 feature points, the feature values on the other 511 channels can represent the corresponding area, and each value on the activation channel represents the activation value of each feature point.
[0093] In another embodiment, in the feature map extracted by the feature extraction module, the value on each channel itself represents the importance of the corresponding feature point. The feature values of the same feature point on each channel can jointly determine the importance of the feature point. At this time, the activation value of a single feature point can be positively correlated with the absolute value of the feature value corresponding to each channel. In a specific example, the activation value of a single feature point is the square root of the feature value corresponding to each channel. Specifically, assuming that the number of channels is s, the feature point x i The corresponding eigenvalues on each channel are x i1 、x i2 ……x is , then its activation value is .
[0094] In more embodiments, the activation value of a single feature point may also be determined by other reasonable methods, which will not be described in detail here.
[0095] In order to extract the prototype vector, we can first select candidate feature points that can express the target area on the image as much as possible. According to the activation value, we can filter out candidate feature points from each region vector according to the predetermined activation conditions. The region corresponding to the candidate feature point usually has a greater probability of being the target feature region. The activated feature point is the point that is initially screened to express the target region, for example Figure 4 Each feature point filled with gray on the feature map. The activation conditions are, for example: the activation value is greater than a predetermined threshold (such as 0.7); arranged in a predetermined number in descending order of activation values; and so on. Among them, the activation value greater than the predetermined threshold is more universal, and when selecting candidate feature points in a predetermined number in descending order of activation values, due to the inconsistent target sizes in each image, the corresponding number of feature points is inconsistent. In order to ensure that each image selects feature points that correspond to the target area as much as possible, it is necessary to reasonably control the "predetermined number", such as using a smaller predetermined number (but this may cause images with larger target areas to miss many valid feature points).
[0096] When selecting candidate feature points, all feature points that meet the activation conditions can be selected as candidate feature points, or a part of the feature points can be selected as candidate feature points. Specifically, in one embodiment, all feature points whose activation values are greater than a predetermined activation threshold can be selected as candidate feature points; in another embodiment, a predetermined number of feature points can be randomly selected from feature points whose activation values are greater than a predetermined activation threshold as candidate feature points; in another embodiment, a predetermined number of feature points can be selected from feature points whose activation values are greater than a predetermined activation threshold in descending order of activation values as candidate feature points. In more embodiments, the prototype vector can be determined in more ways, which will not be described in detail here.
[0097] For a single candidate feature point, a corresponding single prototype vector is constructed based on its feature values in each channel. For example, the feature values of each channel can be used as the values of each dimension of the prototype vector to construct the prototype vector, or the normalized values of the feature values of each channel can be used as the values of each dimension of the prototype vector to construct the prototype vector, and so on. For a more intuitive description of the prototype vector extraction process, please refer to Figure 5 shown. Figure 5 The feature maps shown can be f1 to f n Any feature map f in i The number of channels of this feature map is 4. In practice, the number of channels can also be other numbers.
[0098] It can be understood that after feature extraction, an area in the original image is mapped to a feature point on the feature map. On each channel of the feature map, a single feature area is mapped to the same feature point, for example, the feature point at the 10th row and 20th column. In other words, the corresponding feature points on each channel of the feature map are expressing the same area. Therefore, for a certain candidate feature point, the feature values on each channel can be arranged in sequence to form a multi-dimensional vector to represent the corresponding area features. As shown in the example above, a 64×64×512-dimensional feature map can be represented by 64×64 512-dimensional vectors representing a 64×64 area. Figure 5 In , the points 501, 502, 503, and 504 on the four feature channels all correspond to the same row and column values, and therefore represent the same feature point. This is equivalent to cutting out a cuboid on each channel from top to bottom to describe the feature point and the corresponding area, and finally obtaining the corresponding prototype vector. For example, in Figure 5 It can be written as (x 501 , x 502 , x 503 , x 504 ).
[0099] It is worth noting that the above process of determining the prototype vector focuses more on the selection principle of the prototype vector, so the process of determining the candidate feature points is also described. In fact, in the process of determining the prototype vector, determining the candidate feature points is not a necessary step. For example, in a specific example, a predetermined number of feature points can be directly selected in descending order of activation values to extract the prototype vector. In another specific example, a predetermined number of feature points with activation values greater than a predetermined threshold can be directly randomly selected to extract the prototype vector. In addition, in the case where a single image corresponds to multiple corresponding feature maps, the prototype vector can be selected from each feature map. For example, for n feature maps f1 to f n Construct a total number of N prototype vectors.
[0100] In addition, since the feature extraction module is pre-trained according to the classification labels of the image set through samples, it can usually extract the target features. Generally, the larger the activation value, the more likely it is to correspond to the target area, and the smaller the activation value, the more likely it is to correspond to the background area. Therefore, the candidate feature points determined by the activation value, or the feature points corresponding to the prototype vector, usually filter out most of the background areas that are irrelevant to the target, such as the sky, sea, etc. in the image with the bird as the target.
[0101] Then, through step 304, the correction module is used to perform pairwise similarity comparisons on the prototype vectors extracted from the first feature map and the second feature map, and the first feature map and the second feature map are corrected according to the maximum similarity between a single prototype vector and other prototype vectors to obtain a first corrected feature map and a second corrected feature map, respectively.
[0102] It can be understood that for two different images, the backgrounds are often quite different, and the backgrounds and targets are also quite different, while the features of the targets are consistent. Therefore, if a prototype vector can find a prototype vector with a high degree of similarity in other images or itself, the prototype vector is more likely to correspond to the feature area of the target. However, due to the differences in the color and brightness of the original images, the image values of each channel of a single image are different. For example, the RGB value of a pixel in one image is (120, 60, 180), and the RGB value of a pixel in another image is (60, 30, 90), and the colors displayed by these two pixels may only be different in brightness. However, their processed results may have a certain correlation. A region is composed of multiple pixels, so the processing results of two similar regions are also correlated. In addition, if a prototype vector has no similarity with the prototype vectors in itself and other images, the prototype vector is more likely to correspond to the background area. Therefore, by comparing the similarity between prototype vectors, similar target areas can be mined and non-target areas can be filtered out.
[0103] After comparing the similarities between the prototype vectors pairwise, the highest similarity obtained by comparing the single prototype vectors can be used as the confidence (i.e., degree of confidence) that the prototype vector corresponds to the feature area of the target. Usually, the confidence can be defined as an interval by two endpoint values. The closer to one of the endpoints, the smaller the confidence, and the closer to the other endpoint, the greater the confidence. For example, when the confidence is a 0-1 interval defined by endpoints 0 and 1, the closer the confidence is to 0, the smaller the possibility that the feature point is mapped to the feature area of the target, and the closer the confidence is to 1, the greater the possibility that the feature point is mapped to the feature area of the target. In practice, the confidence can be determined in different ways. In this specification, the confidence is positively correlated with the maximum correlation between the prototype vector and other prototype vectors.
[0104] In some optional implementations, a single prototype vector in the first image can be compared with multiple prototype vectors in the second image one by one. Similarly, a single prototype vector in the second image can be compared with multiple prototype vectors in the first image one by one. Assuming that the number of prototype vectors in the first image and the second image is N, at least N prototype vectors are compared. 2 The purpose of this is to mine similar feature areas in the first image and the second image.
[0105] For the convenience of description, any prototype vector in the first image can be called the first prototype vector. The following takes the first prototype vector in the first image as an example to illustrate its confidence determination process. Assuming that there are N prototype vectors corresponding to the second image, the first prototype vector is compared with multiple prototype vectors in the second image one by one for similarity to obtain N similarities. Then, the largest similarity (such as 0.78, etc.) is selected from the N similarities as the confidence of the first prototype vector. Among them, the similarity comparison can be determined by various vector similarity comparison methods such as cosine similarity, standardized Euclidean distance, correlation coefficient, information entropy, etc. For example, X and Y represent the two prototype vectors currently being compared for similarity, and the cosine similarity can be: cos (X, Y).
[0106] In this way, the prototype vector corresponding to the target area (such as the foreground area) can usually find a prototype vector with a higher similarity among the prototype vectors corresponding to the other image, so it has a higher confidence, while the prototype vector of the non-target area (such as the background area) usually has a lower similarity with the prototype vector corresponding to the other image, so it has a lower confidence.
[0107] In a possible design, a single image itself may also have some activation regions that can be reference mined, e.g. Figure 3In the second image 302, there are multiple sheep, and some feature regions between the multiple sheep can also be used as references to complement and activate each other. In this way, in some other optional embodiments, for the first prototype vector, its similarity with 2N-1 prototype vectors outside itself can be detected. For example, when the first image and the second image are currently involved in the cross comparison, p n is any prototype vector in the first image or the second image, and s n (x, y) represents the similarity between prototype vectors, N is the number of prototype vectors corresponding to a single image, and f N When (x, y) represents other prototype vectors in the feature map, it is recorded as: Among them, the prototype vector p n The confidence level is the corresponding 2N s n The maximum value in. n The confidence of is expressed as FM(x, y) .
[0108] It is worth noting that the similarity between the prototype vector and itself is usually the greatest. If the confidence is determined by the maximum similarity, when the prototype vector is compared with itself, the confidence of each prototype vector is the maximum confidence, which loses the meaning of confidence. Therefore, in the process of determining the confidence, the prototype vector is often not compared with itself. The confidence that the area corresponding to the prototype vector is the target area
[0109] It can be understood that when there are multiple images to be cross-compared, a single prototype vector can be compared with the prototype vectors corresponding to each image, which will not be described in detail here.
[0110] In this way, each confidence level corresponding to each feature point (such as a candidate feature point) corresponding to each prototype vector can be determined according to each prototype vector. According to each confidence level, each feature map can also be modified to further determine the target area and filter out non-target areas that meet the aforementioned activation conditions.
[0111] In order to use confidence to correct the first feature map, the second feature map and other feature maps extracted by the feature extraction module, various reasonable methods can be used to increase the feature value of feature points with higher confidence, and conversely, reduce the feature value of feature points with lower confidence.
[0112] In a possible design, the first feature map, the second feature map, etc. can be corrected by using the product of the confidence determined according to the prototype vector and each eigenvalue of the corresponding feature point. For example, in one embodiment, for a single channel of a single feature point in a single feature map, the eigenvalue can be replaced by the product of the corresponding confidence and the eigenvalue on the channel, thereby forming a corrected eigenvalue. For another example, in another embodiment, for a single channel of a single feature point in a single feature map, the eigenvalue can also be replaced by the sum of the product of the corresponding eigenvalue and the corresponding confidence and the eigenvalue on the channel to form a corrected value. For example, if the eigenvalue of a channel is 150 and the confidence is 0.7, the corrected value can be 150×(1+0.7)=255. In an optional implementation, the corrected value can also be set with a maximum value, such as 255. When the calculated corrected value is greater than the maximum value, the maximum value can be uniformly taken, or normalization can be performed according to the actual maximum value as a normalization coefficient. For example, the eigenvalue is 160, and the correction value is 160×(1+0.7)=271 according to the confidence level of 0.7, which is greater than the maximum value of 255. In this case, the correction value can be taken as the maximum value of 255, or normalized according to the maximum correction value (for example, 360), such as modifying the maximum correction value 360 to the maximum value of 255 as the maximum correction value, and the correction value 271 is corrected by the ratio of 255 to 360, such as correcting to 271×(255 / 360)≈192. In addition, for feature points that do not correspond to the prototype vector, their eigenvalues can be left unchanged, or other reasonable processing can be performed, which is not limited here.
[0113] In one embodiment, a confidence distribution map may be constructed based on the confidence, and used to perform operations on the corresponding feature map by element-by-element multiplication to obtain a modified feature map. Figure 4 As shown, the confidence distribution map constructed for the first feature map 403 is, for example, an array 405. It is worth noting that the array 405 is a general identifier. In practice, confidence maps can also be constructed for each feature map. In the confidence map, the confidence corresponding to the candidate feature point can be used as the corresponding element, and the position that is not selected as the candidate feature point can be supplemented by a predetermined value. For example, the corresponding confidence of the feature point that meets the activation condition but is not selected as the candidate feature point can be set to a first predetermined value, such as 1, or a maximum confidence value, such as 0.9, and the corresponding confidence of the feature point that does not meet the activation condition can be set to a second predetermined value, such as 0, or a minimum confidence value, such as 0.1. Then, the confidence distribution map is multiplied point by point with the corresponding feature map to obtain a modified feature map. For example, a modified feature map 406 is obtained for the first feature map. Similarly, a second modified feature map can be obtained for the second feature map. In the case where the images participating in this round of cross-matching also include other images, other modified feature maps can also be obtained in a similar manner.
[0114] In more embodiments, the modified feature map can also be obtained by other methods, which will not be described in detail here. It can be understood that compared with the feature map extracted by the feature extraction module, the modified feature map may have a small increase, unchanged or greatly reduced feature value corresponding to the feature point of the non-target area (such as corresponding to a smaller confidence), and the feature value of the feature point of the target area (such as corresponding to a larger confidence) may increase significantly, remain unchanged or slightly decrease, thereby better widening the gap to better identify the target and mark it.
[0115] Furthermore, in step 305, the first image and the second image are classified respectively by using a classification module according to the first corrected feature map and the second corrected feature map to obtain respective corresponding classification results.
[0116] Usually, in order to ensure the accuracy of target recognition, it is necessary to perform pixel-level segmentation on the image. For example, when performing human target recognition on a person holding an apple, the pixels contained in the part of the human body covered by the apple need to be excluded. Therefore, the image can be segmented at the pixel level by mapping the activation value of the feature map or the feature value on the feature map to each pixel of the initial image. That is, the recognition results are annotated pixel by pixel, for example, each pixel of the human body corresponding to the target person is annotated as 1, and other pixels are annotated as 0, or Figure 2 In the method, different colors are selected according to the size of the feature value to represent each pixel (for example, each pixel in the activated area is represented by color, the pixel with the largest activation value is represented by red, followed by orange, yellow, etc.). The classification result here can at least include the pixel-level annotation result for the corresponding image.
[0117] like Figure 4 As shown, for the first image, the second image and other images to be cross-compared, the corrected feature map can be processed by the classification module to obtain a classification result. The classification process of each image is the same, and here, the first image is taken as an example for explanation.
[0118] refer to Figure 4 As shown, by processing the corrected feature map 405 by the classification module, a classification result image 408 can be obtained. The classification result image 408 includes at least pixel-level annotation results, which can annotate the target area and non-target area for each pixel. In the case of multiple targets, pixels belonging to various targets can also be annotated. According to an optional implementation method, in order to utilize the labels of the image set, the classification result image 408 can also include image-level classification results. For example, the image is classified as a target image or a non-target image, etc. The image classification result can be represented by a numerical value, for example, each category corresponds to a numerical value. The image classification result can also be represented in the form of a vector, and each dimension of the vector corresponds to the probability of classification into each category.
[0119] Next, through step 306, the model loss of the image annotation model is determined based on the above classification results, so as to adjust the undetermined parameters of the image annotation model with the goal of minimizing the model loss. It can be understood that in the supervised learning field, the model loss can usually be determined by comparing the model output results (such as the above classification results) with the target results. Under the framework of this specification, the pixel-level annotation results are weakly supervised by the image-level classification labels to determine the model loss.
[0120] It is understandable that in some implementations, it is hoped that the image-level classification results and the corresponding labels are consistent to ensure the basic classification and recognition capabilities of the image annotation model. Therefore, the model loss can include classification loss. The so-called classification loss is the difference between the classification result obtained by the image annotation model and the corresponding category label (here is the first category label). When the classification result is represented by a vector, the classification loss can be measured by one or more methods such as cross entropy, mean square error, DL distance, etc. Taking cross entropy as an example, the classification loss L c For example, it can be written as:
[0121]
[0122] Among them, v defines the probability of classification into the current category i, and u(i) defines the true category label of the current image. For the first image and the second image, it is the first category label. u(i) and v(i) can be represented by numerical values or vectors.
[0123] According to the technical concept of this specification, if the corrected feature map deviates greatly from the feature map extracted by the feature extraction module, the corrected feature map may be distorted. Therefore, according to a possible design, a self-supervised correction loss (e.g., L self ), to narrow this gap. The correction loss can be described by the difference between the corrected feature map and the feature map extracted from the corresponding image. This difference can be represented by a norm, for example. Taking the first image as an example, the norm of the image value of the first feature map and the first corrected feature map can be determined feature point by feature point, and summed for each feature point. For example, it can be recorded as:
[0124]
[0125] Where HW is the resolution of the feature map, f N and f N ' respectively represent the first modified feature map and the first feature map. The meaning of this formula is to sum the square of the difference between the H×W feature points in the first modified feature map and the first feature map pixel by pixel. Figure 4As shown, the classification module processes the feature map 403 extracted by the feature extraction module to map the activation area of the feature map 403 to each pixel of the first image to obtain a second annotation result 407. The classification result image 408 includes the first annotation result obtained by the classification module mapping the corrected feature map 406 to each pixel of the first image. The first annotation result and the second annotation result 407 in the classification result image 408 can be respectively used as f in the above formula. N and f N ', determine the model loss corresponding to the first image.
[0126] In an optional implementation, the model loss determined for one image is the sum of its classification loss and correction loss. The model loss corresponding to multiple images is the sum of the losses corresponding to each image. In an optional embodiment, the model loss can be determined for each image of the current batch, and the sum can be calculated for each image to obtain the model loss of the current batch. Then, according to the gradient of the undetermined parameters in the image annotation model, the direction of reducing the model loss is determined, so that the undetermined parameters are adjusted using methods such as gradient descent and Newton's method, thereby training the image annotation model.
[0127] Here, the pending parameters in the image annotation model may at least include the pending parameters in the prototype extraction module and the correction module. Among them, in the case where the feature extraction module and the classification module are two parts of a pre-trained classification model, the pending parameters may not be included. In some embodiments, the pending parameters in the image annotation model may also include pending parameters in at least one of the feature extraction module and the classification module. For example, in the case where the classification module is unrelated to the pre-trained classification model, the pending parameters are included. For another example, in order to make the feature extraction module better adapt to the pixel-level image annotation task, the pending parameters in the pre-trained feature extraction module can still be further adjusted as the pending parameters in the image annotation model.
[0128] Usually, a model training stop condition can also be preset. When the stop condition is met, the training of the image annotation model is stopped. The stop condition can be, for example, the image annotation results corresponding to each image (such as Figure 4 f N ') convergence, model loss convergence, gradient convergence of undetermined parameters, etc. The corresponding parameter convergence can be achieved by changing the amount less than the predetermined value (such as 1 / 10 3 ), the mean tends to be stable, etc. In one embodiment, the image annotation results (such as Figure 4 f N ') Take a sliding average in multiple rounds, such as taking a sliding average in 5 rounds, and when the average value is less than a predetermined value, it is determined that the image annotation result has converged, thereby ending the training.
[0129] Understandably, Figure 3 The illustrated embodiment processes multiple images together through an image annotation model, thereby mining target areas and filtering out non-target areas through cross-comparison of features between images. In this embodiment, a batch of training samples may include one or more groups of sample images, and multiple images in a group of sample images are referenced to each other. Optionally, a single image may also provide a reference for itself.
[0130] For the trained image annotation model, you can Figure 6 The process shown in FIG. Figure 6 As shown, the image annotation process may include:
[0131] Step 601, obtaining a first image and a second image from a sample set, wherein both the first image and the second image have a first category label at the image level;
[0132] Step 602, processing the first image and the second image by a pre-trained feature extraction module to obtain a first feature map and a second feature map respectively;
[0133] Step 603, using a prototype extraction module, extracting multiple prototype vectors from the first feature map and the second feature map respectively, where a single prototype vector corresponds to a single feature point on the corresponding feature map and has a corresponding activation value that satisfies the activation condition;
[0134] Step 604: Perform a pairwise similarity comparison on each prototype vector extracted from the first image and the second image through a correction module, and correct the first feature map and the second feature map according to the maximum similarity between a single prototype vector and other prototype vectors to obtain a first corrected feature map and a second corrected feature map respectively.
[0135] Step 605: Classify the first image and the second image respectively using a classification module according to the first corrected feature map and the second corrected feature map to obtain corresponding pixel-level annotation results.
[0136] It is worth mentioning that Figure 6 The image annotation process shown generally has a good annotation effect only on the sample set used to train the image annotation model, and therefore, steps 601 to 605 are substantially consistent with steps 301 to 305. Figure 6 The process shown is Figure 3 The difference is that there is one less step 305 for determining the model loss to adjust the pending parameters. At the same time, in step 605, it is only necessary to obtain the pixel-level image annotation results without paying attention to whether there are classification results for the image set.
[0137] Based on the above technical concept, a solution with a simpler processing structure can also be envisioned. For example, the prototype vectors of multiple images are collected in a prototype vector set, and the image annotation model can directly process one image and mine the target area for the current image through the prototype vector set, filtering out non-target areas.
[0138] According to this assumption, Figure 7 Another embodiment is shown. A scheme for collecting prototype vectors in a prototype vector set is also provided. When processing a single image, its prototype vector can be compared with the vectors in the prototype vector set to determine the confidence level, and the single image can be corrected to complete the image annotation.
[0139] like Figure 7 As shown in Figure 1, the training process of an implemented image annotation model is shown. Figure 7 As shown, the process includes the following steps: step 701, obtaining a first image from a sample set, wherein the first image corresponds to a first category label; step 702, processing the first image through a pre-trained feature extraction module to obtain a first feature map; step 703, extracting multiple prototype vectors from the first feature map using a prototype extraction module, wherein a single prototype vector corresponds to a single feature point on the corresponding feature map and has a corresponding activation value that satisfies the activation condition; step 704, using a correction module, comparing the similarity of each prototype vector extracted from the first image with each reference vector in a reference vector set, and correcting the first feature map to obtain a first corrected feature map according to the maximum similarity between the single prototype vector and each reference vector, wherein the reference vector is extracted from an image corresponding to the first category label; step 705, classifying the first image using a classification module according to the first corrected feature map to obtain a first classification result, wherein the first classification result includes a first annotation result at the pixel level; step 706, determining a model loss of an image annotation model based on the first classification result, thereby adjusting the undetermined parameters of the image annotation model with the goal of minimizing the model loss.
[0140] Combine the following Figure 8 The architecture diagram is shown in Figure 1 and the relevant steps are described in detail.
[0141] First, in step 701, a first image is obtained from a sample set. The first image may be any image in the sample set that has a corresponding first category label. For example Figure 8 The first category label is an image-level classification label pre-annotated for the first image, such as a target image, a non-target image, or a first target image, a second target image, a third target image, etc. The first category label can also be represented by a number, such as 0, 1, etc.
[0142] Next, through step 702, the first image is processed by a pre-trained feature extraction module to obtain a first feature map. The purpose of the feature extraction module is to extract features related to the target, such as the head features and ear features of a sheep. It can be understood that the target recognition process is usually the result of a joint judgment of multiple features, and the features extracted by the feature extraction module are usually related to the recognition of the target. In practice, the features extracted by the feature extraction module may not be divided into regions based on the head, ears, etc. visible to the human eye, but other features customized by the model through deep learning.
[0143] The feature extraction module can be pre-trained. For example, the feature extraction module can be the front part of the classification model for classifying the target. In the field of image processing, the classification model can often be implemented through a convolutional neural network, which extracts features on the image through the convolution kernel. After the classification model is trained, it can be considered that the target features on the image can be extracted using the corresponding convolution kernel. These extracted features can be visible to the naked eye and can be distinguished by concrete features, such as the head features and leg features of the target "sheep", or they can be abstract and cannot be recognized by the naked eye. Other features are not limited here. The trained classification model can be considered that the first half is used for feature extraction, and the second half is used for feature fusion and classification processing. The second half can be called a classification module, for example.
[0144] exist Figure 8 In the example, the first image is extracted into a feature map 802. The feature map 802 may include n feature maps f1 to f n . Where n is a positive integer greater than or equal to 1. A single feature map can contain multiple channels. For consistency, n feature maps f n Can have a consistent number of channels. f1 to f n For example, the corresponding n feature maps can be determined according to the outputs of different convolutional layers in the same "block".
[0145] Then, in step 703, a plurality of prototype vectors are extracted from the first feature map using a prototype extraction module. A single prototype vector corresponds to a single feature point on the corresponding feature map and has a corresponding activation value that satisfies the activation condition. In order to represent the feature region with a prototype vector, the vector of the corresponding feature point can be determined by the feature value of each channel of the corresponding feature point on the feature map. In the case where the probability of a feature point being mapped to the target region in the image is high, its vector can be extracted as a feature vector.
[0146] It is understandable that not all feature points on the feature map correspond to the feature area of the target. Therefore, feature points that are obviously not in the target area can be filtered out through activation conditions. Feature points that do not meet the activation conditions usually correspond to non-target areas. In one embodiment, during training, the feature extraction module sets one of the channels as an activation channel, and the value on the activation channel represents the size of the activation value of the corresponding feature point. In another embodiment, in the feature map extracted by the feature extraction module, the value on each channel itself represents the importance of the corresponding feature point. For example, the activation value of a single feature point can be positively correlated with the absolute value of the feature value corresponding to it on each channel. In more embodiments, the activation value of a single feature point can also be determined by other reasonable methods.
[0147] In order to extract the prototype vector, you can first select candidate feature points that can express the target area on the image as much as possible. Candidate feature points can be selected from each area vector according to the activation value and the predetermined activation condition. The activation condition is, for example: the activation value is greater than a predetermined threshold (such as 0.7); arranged in a predetermined number in the order of activation value from large to small; and so on. When selecting candidate feature points, all feature points that meet the activation conditions can be selected as candidate feature points, or a part of the feature points can be selected as candidate feature points. For example: all feature points with activation values greater than the predetermined activation threshold can be selected as candidate feature points; a predetermined number of feature points can be randomly selected from feature points with activation values greater than the predetermined activation threshold as candidate feature points; a predetermined number of feature points can also be selected from feature points with activation values greater than the predetermined activation threshold as candidate feature points in the order of activation values from large to small, and so on. For candidate feature points, the prototype vector can be determined according to their corresponding eigenvalues.
[0148] Next, through step 704, the correction module compares the similarity of each prototype vector extracted from the first image with each reference vector in the reference vector set, and corrects the first feature map according to the maximum similarity between the single prototype vector and each reference vector to obtain a first corrected feature map.
[0149] The reference vector may be a reference vector for comparing the target area. The reference vector set may correspond to the category label one by one. For a first image with a first category label, a reference vector in the reference vector set corresponding to the first category label may be used as a reference.
[0150] The reference vectors in the reference vector set corresponding to the first category label may be extracted from the image corresponding to the first category label.
[0151] In one embodiment, a pre-trained feature extraction module can be used to extract corresponding feature maps from each image corresponding to the first category label in the sample set, and then, in each feature map, each candidate feature point with an activation value greater than a first activation threshold is selected. The first activation threshold is compared with the activation condition to screen out feature points with higher activation values (when the activation condition is greater than the second activation threshold, the first activation threshold is greater than the second activation threshold). Then, for a single candidate feature point, a corresponding single reference vector is constructed according to its feature value in each channel and added to the reference vector set. Afterwards, the reference vector in the reference vector set is used as a reference for each image of the first category label.
[0152] In another embodiment, the reference vector set may be empty or contain very few reference vectors at the initial stage, and during the iteration of each cycle, a prototype vector (such as a prototype vector having an activation value greater than a first activation threshold) may be selected according to the current image to be added to the reference vector set. When the reference vector set is empty at the initial stage, the prototype vectors of the image itself may be compared for similarity.
[0153] According to one embodiment, in the current cycle, when the current image is the first image, for the first prototype vector extracted from the first feature map, it is detected whether its first maximum similarity with each benchmark vector is greater than a predetermined similarity threshold. When the first maximum similarity is greater than the predetermined similarity threshold, the first prototype vector is added to the benchmark vector set as a benchmark vector.
[0154] In one embodiment, the single prototype vector in the first image is compared with the reference vectors in the reference vector set corresponding to the first category label one by one to obtain each similarity. The similarity comparison can be determined by various vector similarity comparison methods such as cosine similarity, standardized Euclidean distance, correlation coefficient, information entropy, etc. Among them, for a single prototype vector, the confidence can be determined according to the highest similarity among the similarities. For example, the confidence is the highest similarity among the similarities, or other values positively correlated with the highest similarity among the similarities.
[0155] In this way, the confidences corresponding to the feature points (such as candidate feature points) corresponding to the prototype vectors can be determined according to the prototype vectors. According to the confidences, the first feature map can be modified to further determine the target area and filter out the non-target area that meets the above activation conditions.
[0156] In one possible design, the first feature map can be corrected by using the product of the confidence determined according to the prototype vector and each eigenvalue of the corresponding feature point to obtain a first corrected feature map. For example, in one embodiment, for a single channel of a single feature point in a single feature map, the eigenvalue can be replaced by the product of the corresponding confidence and the eigenvalue on the channel to form a corrected eigenvalue. For another example, in another embodiment, for a single channel of a single feature point in a single feature map, the eigenvalue can also be replaced by the sum of the product of the corresponding eigenvalue and the corresponding confidence and the eigenvalue on the channel to form a corrected value. In more embodiments, the first feature map can also be corrected in more ways, which will not be repeated here.
[0157] Further, in step 705, the first image is classified using a classification module according to the first corrected feature map to obtain a first classification result. The first classification result includes a first pixel-level annotation result. The first annotation result is a recognition result annotated pixel by pixel, for example, each pixel of the human body corresponding to the target person is annotated as 1, and other pixels are annotated as 0, and so on.
[0158] Then, through step 706, the model loss of the image annotation model is determined based on the first classification result, so as to adjust the undetermined parameters of the image annotation model with the goal of minimizing the model loss.
[0159] The model loss may include a first loss for the first image, and for the first image, the classification result also includes a first classification result at the image level. The first loss may specifically include: a first classification loss determined by comparing the first classification result with the first category label; and a first correction loss determined by comparing the first annotation result with the second annotation result determined using the first feature map.
[0160] Further, based on the current batch of training samples including the first image, the model losses corresponding to the current batch of training samples and the previous batches of training samples are detected. When the change of the sliding average of the losses of the models is less than a predetermined value, it is determined that the image annotation model training is completed.
[0161] It is worth mentioning that Figure 7 The process shown is Figure 3 The process shown is similar to the process for a single image, except that Figure 7 In the process shown, one image is processed at a time, and the prototype vector is compared with the reference vector in the reference vector set. Figure 3 In the illustrated process, multiple images are processed at once and cross-compared based on the selected prototype vectors. Figure 7 The architecture used by the process (such as Figure 8 shown), and Figure 3 The architecture used by the process shown (such as Figure 4 ) is more concise, but in comparison, Figure 7 The process needs to maintain a reference vector set for each category label, and when the number of images with the corresponding category label is large, the number of reference vectors in the reference vector set is large. In an optional implementation, the reference vectors in the reference vector set can also be screened according to similarity. For example, when the similarity between two reference vectors (or between a reference vector and a candidate reference vector) is greater than a predetermined screening threshold, one of them can be screened out. When the image standard model training is completed, the reference vector set is fixed.
[0162] according to Figure 7 The image standard model trained by the process shown can also be used for pixel-level annotation of images in the sample set. Fig. 9 The process of annotating an image to be annotated by using a trained image annotation model is shown. The process includes: step 901, obtaining a first image from a sample set, wherein the first image corresponds to a first category label; step 902, processing the first image by a pre-trained feature extraction module to obtain a first feature map; step 903, extracting multiple prototype vectors from the first feature map by using a prototype extraction module, wherein a single prototype vector corresponds to a single feature point on the corresponding feature map and has a corresponding activation value that satisfies an activation condition; step 904, comparing the similarity of each prototype vector extracted from the first image with each reference vector in a reference vector set by using a correction module, and correcting the first feature map to obtain a first corrected feature map according to the maximum similarity between the single prototype vector and each reference vector; step 905, classifying the first image by using a classification module according to the first corrected feature map to obtain a first pixel-level annotation result for the first image.
[0163] Fig. 9 The process shown is Figure 7 The process shown is similar, except that the steps of determining model loss and adjusting model parameters are omitted.
[0164] In addition, Figure 3 , Figure 6 , Figure 7 , Fig. 9 In the illustrated process, an edge refinement step can be added, and edge refinement is performed after pixel-level annotation is completed, thereby improving the segmentation effect. The edge refinement can be implemented using conventional technologies such as CONTA and RPNet, which will not be described in detail here.
[0165] Looking back at the above process, for the images in the training set, pixel-level annotation can be performed through image-level category labels. In the specific image annotation model training and image annotation process, the prototype vector is used to cross-compare the features between different images, so as to further explore the target area in the image, and to filter out non-target areas to achieve weakly supervised segmentation tasks. In the loss determination process, not only the classification loss is considered, but also the similarity between the corrected segmentation result and the original segmentation result, so that the segmentation result is more stable.
[0166] refer to Fig.10 As shown, a schematic diagram of the effect of evaluating the annotation performance of the image annotation model trained under the architecture of this specification in the PASCAL VOC 2012 and MS COCO training sets is shown. Fig.10 In the figure, the brighter area represents the segmentation results of various machine learning models for the target. Each row represents a segmentation target. The first column is the original image, and each of the other columns represents a segmentation method. Among them, the column corresponding to "Ours" shows the segmentation effect achieved by the solution under the technical concept of this specification, and the "GT" (Ground Tures) column represents the effect of manual annotation. Fig.10 It can be seen that the segmentation scheme of this specification is closer to the "GT" effect than other segmentation schemes.
[0167] According to another aspect, an embodiment of the present specification also provides a training device for an image annotation model. The image annotation model can be used to annotate images with classification labels at the pixel level. The image annotation model includes a feature extraction module, a prototype extraction module, a correction module, and a classification module. Fig.11 FIG. 1 shows a training apparatus 1100 for an image annotation model according to an embodiment. Fig.11 In the embodiment, the device 1100 includes: an acquisition unit 1101, a feature extraction unit 1102, a prototype extraction unit 1103, a correction unit 1104, a classification unit 1105, and an adjustment unit 1106.
[0168] In the case where at least the first image and the second image intersect the mining target area, in a single execution cycle of the apparatus 1100:
[0169] An acquisition unit 1101 is configured to acquire a first image and a second image from a sample set, wherein both the first image and the second image have a first category label at the image level;
[0170] The feature extraction unit 1102 is configured to process the first image and the second image through a pre-trained feature extraction module to obtain a first feature map and a second feature map respectively;
[0171] The prototype extraction unit 1103 is configured to use the prototype extraction module to extract multiple prototype vectors from the first feature map and the second feature map respectively, where a single prototype vector corresponds to a single feature point on the corresponding feature map and has a corresponding activation value that satisfies the activation condition;
[0172] The correction unit 1104 is configured to perform a pairwise similarity comparison on each prototype vector extracted from the first image and the second image through the correction module, and correct the first feature map and the second feature map according to the maximum similarity between a single prototype vector and other prototype vectors to obtain a first corrected feature map and a second corrected feature map respectively;
[0173] The classification unit 1105 is configured to classify the first image and the second image respectively using a classification module according to the first corrected feature map and the second corrected feature map to obtain respective corresponding classification results, wherein the classification results include pixel-level annotation results;
[0174] The adjustment unit 1106 is configured to determine the model loss of the image annotation model based on the classification result, so as to adjust the undetermined parameters of the image annotation model with the goal of minimizing the model loss.
[0175] In the case of cross-mining the target region using a single image and a reference vector set, in a single execution cycle of the apparatus 1100:
[0176] An acquisition unit 1101 is configured to acquire a first image from a sample set, wherein the first image corresponds to a first category label;
[0177] A feature extraction unit 1102 is configured to process the first image through a pre-trained feature extraction module to obtain a first feature map;
[0178] The prototype extraction unit 1103 is configured to extract a plurality of prototype vectors from the first feature map using a prototype extraction module, wherein a single prototype vector corresponds to a single feature point on the corresponding feature map and has a corresponding activation value satisfying an activation condition;
[0179] The correction unit 1104 is configured to compare the similarity of each prototype vector extracted from the first image with each reference vector in the reference vector set through the correction module, and correct the first feature map according to the maximum similarity between the single prototype vector and each reference vector to obtain a first corrected feature map;
[0180] The classification unit 1105 is configured to classify the first image using a classification module according to the first corrected feature map to obtain a first classification result, where the first classification result includes a first pixel-level annotation result;
[0181] The adjustment unit 1106 is configured to determine the model loss of the image annotation model based on the first classification result, so as to adjust the undetermined parameters of the image annotation model with the goal of minimizing the model loss.
[0182] According to one aspect, a corresponding image annotation device is also provided, such as Fig.12 As shown, the image annotation device 1200 may include an acquisition unit 1201 , a feature extraction unit 1202 , a prototype extraction unit 1203 , a correction unit 1204 , and an annotation unit 1205 .
[0183] In the case of cross-labeling of multiple selected images, during the process of labeling the multiple images, the apparatus 1200:
[0184] An acquisition unit 1201 is configured to acquire a first image and a second image from a sample set, wherein both the first image and the second image have a first category label at the image level;
[0185] The feature extraction unit 1202 is configured to process the first image and the second image through a pre-trained feature extraction module to obtain a first feature map and a second feature map respectively;
[0186] The prototype extraction unit 1203 is configured to use the prototype extraction module to extract multiple prototype vectors from the first feature map and the second feature map respectively, where a single prototype vector corresponds to a single feature point on the corresponding feature map and has a corresponding activation value that satisfies the activation condition;
[0187] The correction unit 1204 is configured to perform a pairwise similarity comparison on each prototype vector extracted from the first image and the second image through the correction module, and correct the first feature map and the second feature map according to the maximum similarity between a single prototype vector and other prototype vectors to obtain a first corrected feature map and a second corrected feature map respectively;
[0188] The labeling unit 1205 is configured to classify the first image and the second image respectively according to the first corrected feature map and the second corrected feature map by using a classification module to obtain corresponding pixel-level labeling results.
[0189] During the process of the device 1200 annotating a single image:
[0190] An acquisition unit 1201 is configured to acquire a first image from a sample set, wherein the first image corresponds to a first category label;
[0191] The feature extraction unit 1202 is configured to process the first image through a pre-trained feature extraction module to obtain a first feature map;
[0192] The prototype extraction unit 1203 is configured to extract a plurality of prototype vectors from the first feature map using a prototype extraction module, wherein a single prototype vector corresponds to a single feature point on the corresponding feature map and has a corresponding activation value that satisfies an activation condition;
[0193] The correction unit 1204 is configured to compare the similarity of each prototype vector extracted from the first image with each reference vector in the reference vector set through the correction module, and correct the first feature map according to the maximum similarity between the single prototype vector and each reference vector to obtain a first corrected feature map;
[0194] The labeling unit 1205 is configured to classify the first image using a classification module according to the first corrected feature map to obtain a first labeling result at the pixel level for the first image.
[0195] It is worth mentioning that Fig.11 The device embodiment shown is Figure 3 , Figure 7 Corresponding to the method embodiment shown, Fig.12 The device embodiment shown is Figure 6 , Fig. 9 The method embodiment shown corresponds to the embodiment shown in FIG. Figure 3 , Figure 7 , Figure 6 , Fig. 9 The corresponding descriptions are applicable to Fig.11 , Fig.12 The embodiments in the corresponding scenarios will not be described in detail here.
[0196] According to another embodiment, there is also provided a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute Figure 3 , Figure 6 , Figure 7 or Fig. 9 Any of the methods described.
[0197] According to another embodiment of the present invention, a computing device is provided, including a memory and a processor. The memory stores executable code. When the processor executes the executable code, the above-described Figure 3 , Figure 6 , Figure 7 or Fig. 9 Any of the methods described.
[0198] Those skilled in the art should be aware that in one or more of the above examples, the functions described in the embodiments of this specification can be implemented by hardware, software, firmware, or any combination thereof. When implemented by software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0199] The above specific implementation methods further explain in detail the purpose, technical solutions and beneficial effects of the technical concept of this specification. It should be understood that the above are only specific implementation methods of the technical concept of this specification, and are not used to limit the protection scope of the technical concept of this specification. Any modifications, equivalent substitutions, improvements, etc. made on the basis of the technical solutions of the embodiments of this specification should be included in the protection scope of the technical concept of this specification.
Claims
1. A training method for an image annotation model, wherein the image annotation model is used to annotate an image with a classification label at the pixel level, wherein the image annotation model comprises a feature extraction module, a prototype extraction module, a correction module, and a classification module, wherein the method comprises: Acquire a first image and a second image from a sample set, wherein both the first image and the second image have a first category label at the image level; Processing the first image and the second image respectively through a pre-trained feature extraction module to obtain a corresponding first feature map and a second feature map; Utilizing the prototype extraction module, extracting a plurality of prototype vectors from the first feature map and the second feature map respectively, wherein a single prototype vector corresponds to a single feature point on the corresponding feature map and has a corresponding activation value that satisfies an activation condition, and the magnitude of the activation value is used to characterize the probability that the corresponding feature point corresponds to a feature region of the target in the image; Through the correction module, for each prototype vector extracted from the first feature map and the second feature map, a pairwise similarity comparison is performed, and according to the maximum similarity between a single prototype vector and other prototype vectors, the first feature map and the second feature map are corrected to obtain a first corrected feature map and a second corrected feature map, respectively, and the corrected feature value of a single feature point is determined based on the product of the corresponding maximum similarity and its own feature value; According to the first corrected feature map and the second corrected feature map, the first image and the second image are classified by a classification module to obtain respective corresponding classification results, wherein the classification results include pixel-level annotation results; The model loss of the image annotation model is determined based on the classification result, so as to adjust the undetermined parameters of the image annotation model with the goal of minimizing the model loss.
2. The method according to claim 1, wherein: The feature extraction module includes a first convolution block composed of multiple convolution layers, the convolution results of each convolution layer in the first convolution block have the same number of channels, the first feature map includes each convolution result of the multiple convolution layers in the first convolution block performing convolution operations on the first image, and the second feature map includes each convolution result of the multiple convolution layers in the first convolution block performing convolution operations on the second image.
3. The method according to claim 1, wherein: Using the prototype extraction module, extracting multiple prototype vectors from the first feature map and the second feature map respectively includes extracting multiple prototype vectors from the first feature map in the following manner: Detect activation values corresponding to feature points in the first feature map; Selecting multiple feature points from the feature points that meet the activation conditions as candidate feature points; For a single candidate feature point, a corresponding single prototype vector is constructed according to its feature value in each channel.
4. The method according to claim 3, wherein: The activation condition is that the activation value is greater than a predetermined activation threshold; the selecting of multiple feature points from the feature points satisfying the activation condition as candidate feature points includes at least one of the following: All feature points whose activation values are greater than a predetermined activation threshold are taken as candidate feature points; Randomly select a predetermined number of feature points from feature points whose activation values are greater than a predetermined activation threshold as candidate feature points; Among the feature points whose activation values are greater than a predetermined activation threshold, a predetermined number of feature point parts are selected as candidate feature points in descending order of activation values.
5. The method according to claim 1, wherein: The maximum similarity between a single prototype vector and other prototype vectors is used as the confidence of the feature value of the corresponding single feature point on the first feature map / the second feature map.
6. The method according to claim 1, wherein: The model loss includes a first loss for the first image and a second loss for the second image. For the first image, the classification result includes a first pixel-level annotation result and a first image-level classification result. The first loss includes: a first classification loss determined by comparing the first classification result with the first category label; and A first modified loss is determined by comparing the first labeling result with a second labeling result determined using the first feature map.
7. The method according to claim 6, wherein: The first classification loss is determined by a cross entropy between the first classification result and the first category label.
8. The method according to claim 6, wherein: The first modified loss is determined by: Processing the first feature map via the classification module to obtain a second annotation result at a pixel level; Compare the first labeling result and the second labeling result pixel by pixel; The first correction loss is determined by using the sum of the labeled difference values corresponding to each pixel.
9. The method according to claim 6, wherein: The first labeling result is a result obtained by classifying the first image by using a classification module and then thinning the boundary.
10. The method according to claim 1, wherein: The undetermined parameters of the image annotation model include the undetermined parameters in the prototype extraction module, the correction module, and the classification module.
11. The method according to claim 1, wherein: The method further comprises: According to the current batch of training samples including the first image and the second image, detecting each model loss corresponding to the current batch of training samples and a plurality of consecutive batches of training samples; When the change in the sliding average of the losses of each model is less than the predetermined loss value, it is determined that the training of the image annotation model is completed.
12. A method for image annotation, for annotating images with classification labels in a sample set at the pixel level using a pre-trained image annotation model, wherein the image annotation model comprises a feature extraction module, a prototype extraction module, a correction module, and a classification module, and the method comprises: Acquire a first image and a second image from a sample set, wherein both the first image and the second image have a first category label at the image level; Processing the first image and the second image through a pre-trained feature extraction module to obtain a first feature map and a second feature map respectively; Using a prototype extraction module, extract multiple prototype vectors from the first feature map and the second feature map respectively, wherein a single prototype vector corresponds to a single feature point on the corresponding feature map and has a corresponding activation value that satisfies an activation condition, and the magnitude of the activation value is used to characterize the probability that the corresponding feature point corresponds to a feature region of the target in the image; Through the correction module, for each prototype vector extracted from the first image and the second image, a pairwise similarity comparison is performed, and according to the maximum similarity between a single prototype vector and other prototype vectors, the first feature map and the second feature map are corrected to obtain a first corrected feature map and a second corrected feature map, respectively, and the corrected feature value of a single feature point is determined based on the product of the corresponding maximum similarity and its own feature value; According to the first corrected feature map and the second corrected feature map, the first image and the second image are respectively classified using a classification module to obtain corresponding pixel-level annotation results.
13. A training method for an image annotation model, wherein the image annotation model is used to annotate an image with a classification label at the pixel level, the image annotation model comprises a feature extraction module, a prototype extraction module, a correction module, and a classification module, the method comprising: Acquire a first image from the sample set, wherein the first image corresponds to a first category label; Processing the first image through a pre-trained feature extraction module to obtain a first feature map; Extracting multiple prototype vectors from the first feature map using a prototype extraction module, where a single prototype vector corresponds to a single feature point on the corresponding feature map and has a corresponding activation value that satisfies an activation condition; By means of a correction module, for each prototype vector extracted from the first image, a similarity comparison is performed with each reference vector in the reference vector set, and the first feature map is corrected to obtain a first corrected feature map according to the maximum similarity between a single prototype vector and each reference vector, wherein the reference vector is extracted from the image corresponding to the first category label; According to the first corrected feature map, classify the first image using a classification module to obtain a first classification result, wherein the first classification result includes a first pixel-level annotation result; The model loss of the image annotation model is determined based on the first classification result, so as to adjust the undetermined parameters of the image annotation model with the goal of minimizing the model loss.
14. The method according to claim 13, wherein: The reference vectors in the reference vector set are determined in the following manner: Use the pre-trained feature extraction module to extract the corresponding feature map from each image in the sample set; In each feature graph, selecting each candidate feature point having an activation value greater than a first activation threshold, wherein the first activation threshold screens out feature points with higher activation values than the activation condition; For a single candidate feature point, a corresponding single reference vector is constructed according to its feature value in each channel, and is added to the reference vector set.
15. The method according to claim 13, wherein: The method further comprises: For the first prototype vector extracted from the first feature map, detecting whether the first maximum similarity between the first prototype vector and each reference vector is greater than a predetermined similarity threshold; When the first maximum similarity is greater than a predetermined similarity threshold, the first prototype vector is added as a reference vector to a reference vector set.
16. The method according to claim 13, wherein: The model loss includes a first loss for a first image. For the first image, the classification result includes a first pixel-level annotation result and a first image-level classification result. The first loss includes: a first classification loss determined by comparing the first classification result with the first category label; and A first modified loss is determined by comparing the first labeling result with a second labeling result determined using the first feature map.
17. The method according to claim 13, wherein: The method further comprises: According to the current batch of training samples including the first image, detecting the model losses corresponding to the current batch of training samples and the previous batches of training samples; When the change in the sliding average of the losses of each model is less than the predetermined loss value, it is determined that the training of the image annotation model is completed.
18. A method for image annotation, for annotating images with classification labels in a sample set at the pixel level using a pre-trained image annotation model, wherein the image annotation model comprises a feature extraction module, a prototype extraction module, a correction module, and a classification module, and the method comprises: Acquire a first image from the sample set, wherein the first image corresponds to a first category label; Processing the first image through a pre-trained feature extraction module to obtain a first feature map; Extracting multiple prototype vectors from the first feature map using a prototype extraction module, where a single prototype vector corresponds to a single feature point on the corresponding feature map and has a corresponding activation value that satisfies an activation condition; By means of the correction module, for each prototype vector extracted from the first image, the similarity is compared with each reference vector in the reference vector set, and the first feature map is corrected according to the maximum similarity between the single prototype vector and each reference vector to obtain a first corrected feature map; According to the first corrected feature map, the first image is classified using a classification module to obtain a first pixel-level annotation result for the first image.
19. A training device for an image annotation model, the image annotation model is used to annotate images with classification labels at the pixel level, the image annotation model includes a feature extraction module, a prototype extraction module, a correction module, and a classification module, the device includes: An acquisition unit, configured to acquire a first image and a second image from a sample set, wherein both the first image and the second image have a first category label at the image level; A feature extraction unit is configured to process the first image and the second image through a pre-trained feature extraction module to obtain a first feature map and a second feature map respectively; A prototype extraction unit is configured to use the prototype extraction module to extract multiple prototype vectors from the first feature map and the second feature map respectively, wherein a single prototype vector corresponds to a single feature point on the corresponding feature map and has a corresponding activation value that satisfies the activation condition, and the magnitude of the activation value is used to represent the probability that the corresponding feature point corresponds to a feature area about the target in the image; The correction unit is configured to perform a pairwise similarity comparison on each prototype vector extracted from the first image and the second image through the correction module, and correct the first feature map and the second feature map respectively according to the maximum similarity between a single prototype vector and other prototype vectors to obtain a first corrected feature map and a second corrected feature map, wherein the corrected feature value of a single feature point is determined based on the product of the corresponding maximum similarity and its own feature value; a classification unit configured to classify the first image and the second image respectively using a classification module according to the first corrected feature map and the second corrected feature map to obtain respective corresponding classification results, wherein the classification results include pixel-level annotation results; An adjustment unit is configured to determine a model loss of the image annotation model based on the classification result, thereby adjusting the undetermined parameters of the image annotation model with the goal of minimizing the model loss.
20. An image annotation device, used for annotating images with classification labels in a sample set at the pixel level using a pre-trained image annotation model, wherein the image annotation model comprises a feature extraction module, a prototype extraction module, a correction module, and a classification module, and the device comprises: An acquisition unit, configured to acquire a first image and a second image from a sample set, wherein both the first image and the second image have a first category label at the image level; A feature extraction unit is configured to process the first image and the second image through a pre-trained feature extraction module to obtain a first feature map and a second feature map respectively; A prototype extraction unit is configured to use the prototype extraction module to extract multiple prototype vectors from the first feature map and the second feature map respectively, wherein a single prototype vector corresponds to a single feature point on the corresponding feature map and has a corresponding activation value that satisfies the activation condition, and the magnitude of the activation value is used to represent the probability that the corresponding feature point corresponds to a feature area about the target in the image; The correction unit is configured to perform a pairwise similarity comparison on each prototype vector extracted from the first image and the second image through the correction module, and correct the first feature map and the second feature map respectively according to the maximum similarity between a single prototype vector and other prototype vectors to obtain a first corrected feature map and a second corrected feature map, wherein the corrected feature value of a single feature point is determined based on the product of the corresponding maximum similarity and its own feature value; The labeling unit is configured to classify the first image and the second image respectively according to the first corrected feature map and the second corrected feature map by using a classification module to obtain respective corresponding pixel-level labeling results.
21. A training device for an image annotation model, the image annotation model is used to annotate images with classification labels at the pixel level, the image annotation model includes a feature extraction module, a prototype extraction module, a correction module, and a classification module, the device includes: An acquisition unit is configured to acquire a first image from a sample set, wherein the first image corresponds to a first category label; a feature extraction unit, configured to process the first image through a pre-trained feature extraction module to obtain a first feature map; A prototype extraction unit is configured to extract a plurality of prototype vectors from the first feature map using the prototype extraction module, wherein a single prototype vector corresponds to a single feature point on the corresponding feature map and has a corresponding activation value that satisfies an activation condition; A correction unit is configured to compare the similarity of each prototype vector extracted from the first image with each reference vector in the reference vector set through the correction module, and correct the first feature map according to the maximum similarity between the single prototype vector and each reference vector to obtain a first corrected feature map; a classification unit, configured to classify the first image using a classification module according to the first corrected feature map to obtain a first classification result, wherein the first classification result includes a first pixel-level annotation result; An adjustment unit is configured to determine a model loss of the image annotation model based on the first classification result, thereby adjusting the undetermined parameters of the image annotation model with the goal of minimizing the model loss.
22. An image annotation device, used for annotating images with classification labels in a sample set at the pixel level using a pre-trained image annotation model, wherein the image annotation model comprises a feature extraction module, a prototype extraction module, a correction module, and a classification module, and the device comprises: An acquisition unit is configured to acquire a first image from a sample set, wherein the first image corresponds to a first category label; a feature extraction unit, configured to process the first image through a pre-trained feature extraction module to obtain a first feature map; A prototype extraction unit is configured to extract a plurality of prototype vectors from the first feature map using the prototype extraction module, wherein a single prototype vector corresponds to a single feature point on the corresponding feature map and has a corresponding activation value that satisfies an activation condition; A correction unit is configured to compare the similarity of each prototype vector extracted from the first image with each reference vector in the reference vector set through the correction module, and correct the first feature map according to the maximum similarity between the single prototype vector and each reference vector to obtain a first corrected feature map; The labeling unit is configured to classify the first image using a classification module according to the first corrected feature map to obtain a first labeling result at the pixel level for the first image.
23. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 18.
24. A computing device comprising a memory and a processor, characterized in that: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 18 is implemented.
Citation Information
Patent Citations
Weak supervision image semantic segmentation method, system and device based on cross-image association
CN111723814A