Labeling method and labeling device

By using a neural network to obtain the pixel point feature values ​​in the video image and propagate the annotation results based on the similarity of the feature values, the problem of low accuracy of the video semantic segmentation annotation in the prior art is solved, and the segmentation accuracy of the video semantic segmentation neural network is improved.

CN120092258APending Publication Date: 2025-06-03YINWANG INTELLIGENT TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280101337.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-11-30
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The accuracy of the annotation method used for video semantic segmentation in the prior art is low, resulting in the segmentation accuracy of the training video semantic segmentation neural network.

Method used

The first neural network obtains the characteristic values ​​of the pixel point set in the image to be marked and the marked image, and determines the annotation of the pixel point set in the image to be marked based on the similarity of the characteristic values ​​and the annotation result in the marked image.

Benefits of technology

It improves the accuracy of the annotation results, enhances the segmentation accuracy of the video semantic segmentation neural network, and reduces the dependence on the manual annotation of the annotator.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120092258A_ABST
    Figure CN120092258A_ABST
Patent Text Reader

Abstract

The invention provides a labeling method and a labeling device. According to the labeling method provided by the invention, when a labeling device labels a first pixel point set in an unlabeled first image, a first feature value of the first pixel point set in the first image and a second feature value of a second pixel point set in a second image are obtained through a first neural network, the first neural network is used for acquiring feature values of pixel points in the image; and determining the mark of the first pixel point set according to the similarity between the first feature value and the second feature value and the mark of the second pixel point set so as to improve the accuracy of automatic marking.
Need to check novelty before this filing date? Find Prior Art

Description

Annotation Method and Annotation Device

[0001] This application relates to the field of artificial intelligence technology, and in particular, to an annotation method and an annotation device.

[0002] Semantic segmentation refers to classifying each pixel point in an image. Semantic segmentation can be achieved through a deep learning neural network. A deep learning neural network for video semantic segmentation can be called a video semantic segmentation neural network. The video semantic segmentation neural network can be trained with training data, which includes videos and annotations of pixel points in each frame of the video, and the annotation of a pixel point is used to indicate the classification result of that pixel point.

[0003] Currently, a method for obtaining the annotation of pixel points in an image in a video is as follows: 1) Obtain a video, and an annotator annotates the pixel points of some images in the video; 2) Establish a matching relationship between the pixel points in any two temporally consecutive images based on gray-scale similarity; 3) According to this matching relationship, propagate the annotation result of the pixel points in the previous image to the corresponding pixel points in the subsequent image.

[0004] However, the accuracy of the annotation results obtained by the above method is relatively low.

[0005]

[0006] This application provides an annotation method and device, which can improve the accuracy of annotation.

[0007] In a first aspect, this application provides an annotation method, including: obtaining, by a first neural network, a first eigenvalue of a first pixel point set in a first image, where the first neural network is used to obtain the eigenvalue of a pixel point in an image, and the first pixel point set includes at least one pixel point; obtaining, by the first neural network, a second eigenvalue of a second pixel point set in a second image, where the second pixel point set includes at least one pixel point; determining the annotation of the first pixel point set according to the similarity between the first eigenvalue and the second eigenvalue and the annotation of the second pixel point set.

[0008] Among them, the first image can be considered as an unannotated or to-be-annotated image, and the second image can be considered as an annotated image.

[0009] In this embodiment, when the annotation device annotates the first pixel point set in the first image to be annotated based on the second annotated image, it determines the annotation of the first pixel point set based on the similarity between the eigenvalue (also referred to as the first eigenvalue) obtained after inputting the first pixel point set in the first image into the first neural network and the eigenvalue (also referred to as the second eigenvalue) obtained after inputting the second pixel point set in the second image into the first neural network, as well as the annotation of the second pixel point set. In this application, the eigenvalue obtained through the neural network is also referred to as the semantic eigenvalue.

[0010] It should be understood that since the features extracted by the neural network can be considered as abstract semantic features related to tasks such as classification, compared with features such as pixel grayscale and optical flow, they are less sensitive to interferences such as illumination, deformation, occlusion, and rapid displacement. In addition, the semantic features have less noise compared with the features extracted by self-supervised learning. Therefore, in this embodiment, when establishing the matching relationship between the pixel point sets of the first image and the second image based on the similarity between the semantic eigenvalues, the second pixel point set determined from the second image that matches the first pixel point set can be made more accurate. Further, the accuracy of automatic annotation can be improved.

[0011] Combined with the first aspect, in a possible implementation manner, obtaining the first eigenvalue of the first pixel point set in the first image through the first neural network includes: obtaining the feature map output by the first neural network in the first feature extraction layer after the first image is input; obtaining the first target semantic feature map according to the feature map output by the first feature extraction layer; and determining the first eigenvalue of the first pixel point set according to the first target semantic feature map.

[0012] Exemplarily, the annotation device may first scale the feature map output by the first feature extraction layer to the first resolution to obtain the first target semantic feature map; then, obtain the first eigenvalue of the first pixel point set according to the first target semantic feature map with the first resolution.

[0013] Exemplarily, the annotation device may directly determine the feature map output by the first feature extraction layer as the first target semantic feature map without scaling, and then, based on the size relationship between the resolution of the feature map output by the first feature extraction layer and the resolution of the first image, determine the pixel value of the pixel point set corresponding to the first pixel point set in the first target semantic feature map as the first eigenvalue of the first pixel point set.

[0014] In this implementation manner, since the first eigenvalue corresponding to the first pixel point set is determined only according to the first target semantic feature map obtained from one feature extraction layer of the first neural network, the first eigenvalue can be calculated relatively quickly.

[0015] In combination with the first aspect, in a possible implementation, the method further includes: obtaining a feature map output by a first neural network at a second feature extraction layer after a first image is input into the first neural network; wherein, obtaining a first target semantic feature map according to the feature map output by the first feature extraction layer includes: obtaining the first target semantic feature map according to the feature map output by the first feature extraction layer and the feature map output by the second feature extraction layer.

[0016] In specific implementation, when obtaining the first target semantic feature map according to the feature map output by the first feature extraction layer and the feature map output by the second feature extraction layer, the feature map output by the first feature extraction layer and the feature map output by the second feature extraction layer can be simultaneously scaled to the same resolution and fused, and then the fused feature map is determined as the first target semantic feature map.

[0017] In this implementation, since the determined first target semantic feature map is based not only on the feature map output by the first feature extraction layer but also on the feature map output by the second feature extraction layer in the first neural network, it should be understood that different feature extraction layers extract different types of features. Therefore, compared with obtaining the first eigenvalue of the first pixel point set in the first image by using the feature map output by a single feature extraction layer alone, the obtained first eigenvalue is more accurate.

[0018] In combination with the first aspect, in a possible implementation, the first target semantic feature map is consistent with the resolution of the first image.

[0019] In this implementation, when the first target semantic feature map is consistent with the resolution of the first image, the first eigenvalue of the first pixel point set in the first image can be obtained relatively quickly from the first target semantic feature map, that is, the time spent on determining the first eigenvalue of the first pixel point set can be relatively short.

[0020] In combination with the first aspect, in a possible implementation, the method further includes: obtaining a second target semantic feature map corresponding to a second image through a first neural network; reducing the first target semantic feature map to a third target semantic feature map with a second resolution, where the second resolution is smaller than the first resolution; reducing the second target semantic feature map to a fourth target semantic feature map with the second resolution; determining a third pixel point set corresponding to the first pixel point set in the third target semantic feature map according to the first resolution and the second resolution; determining a fourth pixel point set with the highest similarity of pixel values to the third pixel point set from the fourth target semantic feature map; determining a set of second pixel point sets corresponding to the fourth pixel point set from the second image according to the first resolution and the second resolution; where determining the annotation of the first pixel point set according to the similarity between the first eigenvalue and the second eigenvalue and the annotation of the second pixel point set includes: determining the annotation of the second pixel point set with the highest similarity between the eigenvalue in a set of second pixel point sets and the first eigenvalue as the annotation of the first pixel point set.

[0021] In this implementation, when determining the second pixel point set matching the first pixel point set from the second image, since the first target semantic feature map and the second target semantic feature map are first scaled to a smaller resolution, then a pixel matching relationship is established at this smaller resolution, and finally, based on the size relationship between the smaller resolution and the resolution of the first target semantic feature map, the second pixel point set matching the first pixel point set is determined from the second image. It should be understood that this method has higher computational efficiency compared to directly establishing a pixel matching relationship from the first target semantic feature map and the second target semantic feature map.

[0022] In combination with the first aspect, in a possible implementation, the similarity includes any one of the following measurement methods: the Euclidean distance between eigenvalues, the cosine distance between eigenvalues, and the inner product similarity between eigenvalues.

[0023] Optionally, when the similarity is measured by the Euclidean distance between eigenvalues or the cosine distance between eigenvalues, the nearest neighbor search algorithm is used to determine the second pixel point set that best matches the first pixel point set; or, when the similarity is measured by the inner product similarity between eigenvalues, the maximum inner product search method is used to determine the second pixel point set that best matches the first pixel point set.

[0024] Second aspect, the present application provides a pixel matching method, including: obtaining, by a first neural network, first eigenvalue of a first pixel point set in a first image, where the first neural network is used to obtain eigenvalue of pixel points in an image, and the first pixel point set includes at least one pixel point; obtaining, by the first neural network, second eigenvalue of a second pixel point set in a second image, where the second pixel point set includes at least one pixel point; establishing a matching relationship between the pixel point sets of the first image and the second image according to the similarity between the first eigenvalue and the second eigenvalue.

[0025] In combination with the second aspect, in a possible implementation, obtaining, by the first neural network, first eigenvalue of a first pixel point set in a first image includes: obtaining a feature map output by the first neural network at a first feature extraction layer after the first image is input into the first neural network; obtaining a first target semantic feature map according to the feature map output by the first feature extraction layer; determining the first eigenvalue of the first pixel point set according to the first target semantic feature map.

[0026] In combination with the second aspect, in a possible implementation, the method further includes: obtaining a feature map output by the first neural network at a second feature extraction layer after the first image is input into the first neural network; where obtaining the first target semantic feature map according to the feature map output by the first feature extraction layer includes: obtaining the first target semantic feature map according to the feature map output by the first feature extraction layer and the feature map output by the second feature extraction layer.

[0027] In combination with the second aspect, in a possible implementation, the resolution of the first target semantic feature map is consistent with a first resolution, and the first resolution is the resolution of the first image.

[0028] In combination with the second aspect, in a possible implementation, the method further includes:

[0029] obtaining, by the first neural network, a second target semantic feature map corresponding to the second image; shrinking the first target semantic feature map to a third target semantic feature map with a second resolution, where the second resolution is less than the first resolution; shrinking the second target semantic feature map to a fourth target semantic feature map with the second resolution;

[0030] Determine a third set of pixel points corresponding to the first set of pixel points in the third target semantic feature map according to the first resolution and the second resolution; determine a fourth set of pixel points with the highest similarity in pixel values between the fourth target semantic feature map and the third set of pixel points; determine a set of second pixel points corresponding to the fourth set of pixel points in the second image according to the first resolution and the second resolution; determine the second pixel point with the highest similarity between the feature value in the set of second pixel points and the first feature value as the second pixel point matching the first set of pixel points.

[0031] Combined with the second aspect, in a possible implementation, the similarity includes any of the following measurement methods: the Euclidean distance between feature values, the cosine distance between feature values, and the inner product similarity between feature values.

[0032] Optionally, the method further includes: when the similarity is measured by the Euclidean distance between feature values or the cosine distance between feature values, use the nearest neighbor search algorithm to determine the most matching second pixel point set; or, when the similarity is measured by the inner product similarity between feature values, use the maximum inner product search method to determine the most matching second pixel point set.

[0033] In a third aspect, the present application provides a labeling device, including: an acquisition module, configured to obtain a first feature value corresponding to a first set of pixel points in a first image through a first neural network, where the first neural network is used to obtain feature values of pixel points in the image, and the first set of pixel points includes at least one pixel point; the acquisition module is further configured to: obtain a second feature value corresponding to a second set of pixel points in a second image through the first neural network, where the second set of pixel points includes at least one pixel point; a processing module, configured to determine the label of the first set of pixel points according to the similarity between the first feature value and the second feature value and the label of the second set of pixel points.

[0034] Combined with the third aspect, in a possible implementation, after the first image is input into the first neural network, the feature map output by the first neural network at the first feature extraction layer; the processing module is specifically configured to: obtain a first target semantic feature map according to the feature map output by the first feature extraction layer; determine the first feature value of the first set of pixel points according to the first target semantic feature map.

[0035] Combined with the third aspect, in a possible implementation, after the first image is input into the first neural network, the feature map output by the first neural network at the second feature extraction layer; the processing module is further configured to: obtain the first target semantic feature map according to the feature map output by the first feature extraction layer and the feature map output by the second feature extraction layer.

[0036] In combination with the third aspect, in a possible implementation manner, the resolution of the first target semantic feature map is the same as the first resolution, and the first resolution is the resolution of the first image.

[0037] In combination with the third aspect, in a possible implementation manner, the obtaining module is further configured to: obtain a second target semantic feature map corresponding to the second image through the first neural network; the processing module is further configured to: reduce the first target semantic feature map to a third target semantic feature map with a second resolution, where the second resolution is less than the first resolution; reduce the second target semantic feature map to a fourth target semantic feature map with the second resolution; determine a third pixel point set corresponding to the first pixel point set in the third target semantic feature map according to the first resolution and the second resolution; determine a fourth pixel point set with the largest similarity of pixel values between the fourth pixel point set and the third pixel point set from the fourth target semantic feature map; determine a set of second pixel point sets corresponding to the fourth pixel point set from the second image according to the first resolution and the second resolution; and determine the annotation of the second pixel point set with the largest similarity between the eigenvalue in the set of second pixel point sets and the first eigenvalue as the annotation of the first pixel point set.

[0038] In combination with the third aspect, in a possible implementation manner, the similarity includes any one of the following measurement methods: the Euclidean distance between eigenvalues, the cosine distance between eigenvalues, and the inner product similarity between eigenvalues.

[0039] Optionally, when the similarity is measured by the Euclidean distance between eigenvalues or the cosine distance between eigenvalues, the nearest neighbor search algorithm is used to determine the second pixel point set that best matches the first pixel point set; or, when the similarity is measured by the inner product similarity between eigenvalues, the maximum inner product search method is used to determine the second pixel point set that best matches the first pixel point set.

[0040] Fourth aspect, the present application provides a pixel matching device, including: an obtaining module, configured to obtain a first eigenvalue corresponding to a first pixel point set in a first image through a first neural network, where the first neural network is used to obtain the eigenvalues of pixel points in the image, and the first pixel point set includes at least one pixel point; the obtaining module is further configured to: obtain a second eigenvalue corresponding to a second pixel point set in a second image through the first neural network, where the second pixel point set includes at least one pixel point; and a processing module, configured to establish a matching relationship between the pixel point sets of the first image and the second image according to the similarity between the first eigenvalue and the second eigenvalue.

[0041] In combination with the fourth aspect, in a possible implementation manner, the obtaining module is specifically configured to: obtain the feature map output by the first neural network at the first feature extraction layer after the first image is input into the first neural network; the processing module is specifically configured to: obtain a first target semantic feature map according to the feature map output by the first feature extraction layer; and determine a first feature value of the first pixel point set according to the first target semantic feature map.

[0042] In combination with the fourth aspect, in a possible implementation manner, the obtaining module is further configured to: obtain the feature map output by the first neural network at the second feature extraction layer after the first image is input into the first neural network; the processing module is further configured to: obtain the first target semantic feature map according to the feature map output by the first feature extraction layer and the feature map output by the second feature extraction layer.

[0043] In combination with the fourth aspect, in a possible implementation manner, the resolution of the first target semantic feature map is the same as the first resolution, and the first resolution is the resolution of the first image.

[0044] In combination with the fourth aspect, in a possible implementation manner, the obtaining module is further configured to: obtain a second target semantic feature map corresponding to a second image through the first neural network; the processing module is further configured to: reduce the first target semantic feature map to a third target semantic feature map with a second resolution, where the second resolution is less than the first resolution; reduce the second target semantic feature map to a fourth target semantic feature map with the second resolution; determine a third pixel point set corresponding to the first pixel point set in the third target semantic feature map according to the first resolution and the second resolution; determine a fourth pixel point set with the largest similarity of pixel values between the fourth target semantic feature map and the third pixel point set; determine a set of second pixel point sets corresponding to the fourth pixel point set in the second image according to the first resolution and the second resolution; and determine the second pixel point set with the largest similarity between the feature values in the set of second pixel point sets and the first feature value as the second pixel point set matching the first pixel point set.

[0045] In combination with the fourth aspect, in a possible implementation manner, the similarity includes any one of the following measurement methods: the Euclidean distance between feature values, the cosine distance between feature values, and the inner product similarity between feature values.

[0046] In a fifth aspect, the present application provides a labeling device, including: a memory and a processor; the memory is used to store program instructions; the processor is used to call the program instructions in the memory to execute the method described in any one of the first aspect or the second aspect.

[0047] Sixth aspect, the present application provides a vehicle, including the annotation device described in any one of the third aspect and / or the matching device described in any one of the fourth aspect.

[0048] Seventh aspect, the present application provides a chip, including at least one processor and a communication interface, the communication interface and the at least one processor are interconnected by a line, and the at least one processor is configured to run a computer program or instruction to execute the method described in any one of the first aspect or the second aspect.

[0049] Eighth aspect, the present application provides a computer-readable medium, which stores program code for computer execution, and the program code includes instructions for executing the method described in any one of the first aspect or the second aspect.

[0050] Ninth aspect, the present application provides a computer program product, which includes computer program code, and when the computer program code runs on a computer, it enables the computer to implement the method described in any one of the first aspect or the second aspect.

[0051] FIG. 1 is a schematic structural diagram of an application scenario provided by the present application;

[0052] FIG. 2 is a schematic flow diagram of an annotation method provided by an embodiment of the present application;

[0053] FIG. 3 is a schematic structural diagram of obtaining a first target semantic feature map through a neural network provided by an embodiment of the present application;

[0054] FIG. 4 is a schematic structural diagram of obtaining a first target semantic feature map through a neural network provided by another embodiment of the present application;

[0055] FIG. 5 is a schematic structural diagram of obtaining a first target semantic feature map through a neural network provided by yet another embodiment of the present application;

[0056] FIG. 6 is a schematic diagram of the process of establishing a matching relationship provided by an embodiment of the present application;

[0057] FIG. 7 is a schematic flow diagram of a method for annotating each pixel point of each frame of image in a video provided by an embodiment of the present application;

[0058] FIG. 8 is a schematic diagram of a constructed feature pyramid provided by an embodiment of the present application;

[0059] FIG. 9 is a schematic diagram of the process of establishing a matching relationship provided by another embodiment of the present application;

[0060] FIG. 10 is a schematic diagram of matching different pixel points in the same frame of image to the same pixel point in another frame of image provided by an embodiment of the present application;

[0061] FIG. 11 is a structural schematic diagram of an annotation device provided by an embodiment of the present application;

[0062] FIG. 12 is a structural schematic diagram of an annotation device provided by another embodiment of the present application.

[0063] In order to make the purpose and technical solutions of the present application clearer and more intuitive, the annotation method provided by the present application will be described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0064] With the rapid development of artificial intelligence technology, video semantic segmentation technology has become a research hotspot and has become one of the most important technologies in autonomous driving perception, which is of great significance for vehicle decision-making and scene understanding.

[0065] Specifically, video semantic segmentation refers to classifying each pixel point in each frame of an image in a video, which can be implemented through a deep learning neural network. Usually, the deep learning neural network used for video semantic segmentation is also called a video semantic segmentation neural network.

[0066] Exemplarily, FIG. 1 is a schematic diagram of an application scenario of the present application. As shown in FIG. 1, in order to obtain a video semantic segmentation neural network, training data 101 including a video and the annotation of each pixel point in each frame of the video (the annotation of the pixel point is used to indicate the classification result of the pixel point) is required to train the neural network. After that, for a new video data, the classification result of each pixel point in each frame of the new video data can be obtained by using the trained video semantic segmentation neural network.

[0067] It can be seen from the application scenario shown in FIG. 1 that before training the video semantic segmentation model, a large number of videos need to be obtained first, and each pixel point in each frame of the video needs to be annotated to obtain training data. Since the image data used for training is usually very large, it is undoubtedly very time-consuming if only annotators are used to annotate each pixel point in each frame of the video. Therefore, how to efficiently complete the annotation of the training data of the semantic segmentation neural network becomes a very critical issue.

[0068] Currently, a method for annotating pixel points in an image in a video is as follows: 1) Obtain a video, and an annotator annotates the pixel points of some images in the video; 2) Establish a matching relationship between the pixel points in any two temporally consecutive images based on gray-scale similarity; 3) According to this matching relationship, propagate the annotation result of the pixel points in the previous image to the corresponding pixel points in the subsequent image.

[0069] However, the accuracy of the annotation results obtained by the above method is relatively low, which in turn leads to a relatively low segmentation accuracy of the video semantic segmentation neural network obtained by training.

[0070] In view of this, an embodiment of the present application provides an annotation method to improve the accuracy of the annotation results.

[0071] FIG. 2 is a schematic flowchart of an annotation method provided by an embodiment of the present application. As shown in FIG. 2, the annotation method of the present application includes: S201, S202, S203, and S204. The method of the present application can be executed by an annotation device.

[0072] It should be noted here that the specific deployment of the annotation device in this embodiment is not limited. For example, it can be deployed in a server, or it can also be deployed on the vehicle side.

[0073] S201, obtaining a first feature value of a first pixel point set in a first image through a first neural network, where the first neural network is used to obtain the feature value of a pixel point in the image, and the first pixel point set includes at least one pixel point.

[0074] In this embodiment, the first image refers to an image that includes a first pixel point set and needs to be annotated but has not been annotated yet. Among them, the annotation of a pixel point is used to indicate the classification result of the pixel point. For ease of description, the first image in the present application is also referred to as the first image to be annotated.

[0075] It should be noted here that the present application embodiment does not limit how the first image is specifically obtained. Exemplarily, the first image can be a single captured image, or it can be a frame image extracted from a video.

[0076] Specifically, in this embodiment, the number of pixel points included in the first pixel point set is not limited. For example, the first pixel point set can include only one pixel point, or the first pixel point set can include multiple pixel points.

[0077] In this embodiment, after the annotation device obtains the first image, the first image will be input into the first neural network. Among them, the first neural network in this embodiment is used to obtain the feature value of each pixel point in the image to obtain the first feature value of the first pixel point set. In the present application, the feature value corresponding to the pixel point set obtained based on the first neural network is also referred to as a semantic feature value.

[0078] For example, when the first pixel point set only includes one pixel point (assuming this pixel point is called the first pixel point), at this time, obtaining the first eigenvalue corresponding to the first pixel point set in the first image through the first neural network can be understood as obtaining the first eigenvalue corresponding to the first pixel point in the first image. That is, in this scenario, the obtained eigenvalue is at the pixel point level. In addition, it should be understood that in this scenario, after taking the first image as the input of the first neural network, the semantic eigenvalue corresponding to each pixel point can be obtained.

[0079] For another example, the first pixel point set can also include multiple pixel points (for example, N pixel points, where N is a positive integer greater than 1). Among them, these multiple pixel points can be obtained after some algorithms, such as clustering using optical information, light intensity and other information. At this time, when obtaining the first eigenvalue corresponding to the first pixel point set in the first image through the first neural network, the first image can be input into the first neural network first to obtain the semantic eigenvalue of each of the N pixel points, and then based on the semantic eigenvalue of each of the N pixel points, the semantic eigenvalue corresponding to the first pixel point set can be obtained.

[0080] Optionally, the first neural network in this embodiment can be a pre-trained deep neural network. It should be noted that a neural network usually includes multiple layers. Except for the first layer (i.e., the input layer) used to receive the first image, any other layer can be considered as a feature extraction layer for feature extraction of the first image.

[0081] It should be noted here that the specific structure of the first neural network in this embodiment is not limited. For example, the first neural network can be a fully connected neural network, or the first neural network is a convolutional neural network.

[0082] In this embodiment, in specific implementation, one implementation manner of obtaining the first eigenvalue of the first pixel point set in the first image through the first neural network is: obtaining the feature map output by the first neural network in a feature extraction layer (referred to as the first feature extraction layer) after the first image is input; then obtaining the first target semantic feature map according to the feature map output by the first feature extraction layer; and finally determining the first eigenvalue of the first pixel point set according to the first target semantic feature map.

[0083] Among them, the first target semantic feature map obtained according to the feature map output by the first feature extraction layer refers to the feature map that can obtain the first eigenvalue of the first pixel point set.

[0084] It should be understood that the resolution of the feature map output by the first feature extraction layer may be different from that of the first image, but there is a certain proportional relationship between the resolutions.

[0085] Therefore, in one embodiment, as shown in FIG. 3, the feature map output by the first feature extraction layer can be directly used as the first target semantic feature map, and then based on the proportional relationship between the resolutions, the pixel values of the pixel set corresponding to the first pixel point set in the first image in the feature map output by the first feature extraction layer are determined as the first feature values of the first pixel point set in the first image.

[0086] Alternatively, in another embodiment, as shown in FIG. 4, after obtaining the feature map output by the first feature extraction layer, the feature map output by a certain feature extraction layer can be scaled to a feature map with the first resolution (the resolution of the first image is referred to as the first resolution) through interpolation, that is, scaled to a feature map with the same resolution as the first image; then the feature map with the first resolution is determined as the first target semantic feature map, and the first feature value (i.e., the semantic feature value) corresponding to the first pixel point set is determined based on the feature map with the first resolution.

[0087] It should be noted here that in the example shown in FIG. 3, taking the second-to-last feature extraction layer as the first feature extraction layer is only an example and does not constitute a limitation of this application. In actual applications, the user can define which layer is used as the first feature extraction layer by themselves.

[0088] Optionally, when obtaining the first target semantic feature map based on the feature map output by the first feature extraction layer, feature maps output by other feature extraction layers (such as a second feature extraction layer) can also be obtained, and then the first target semantic feature is obtained based on the feature map output by the first feature extraction layer and the feature map output by the second feature extraction layer.

[0089] That is, in this embodiment, when obtaining the first target semantic feature map, it is not only based on the feature map output by a single feature extraction layer, but the feature maps output by multiple feature extraction layers of the first neural network are used (which specific layers to use can be predefined by the user).

[0090] Exemplarily, as shown in FIG. 5, the feature maps output by the last three feature extraction layers of the first neural network after the first image is input into the first neural network can be obtained first, and then the feature map output by each feature extraction layer among the three feature extraction layers is scaled to a feature map with the first resolution and fused to obtain the first target semantic feature map. It can be understood that the first feature value of the first pixel point set can be determined through the first target semantic feature map.

[0091] It can also be understood that since the feature maps output by different feature extraction layers may reflect different types of features of the first image, using the feature maps output by multiple feature extraction layers of the first neural network to obtain the first target semantic feature map results in higher accuracy of the first eigenvalue of the first pixel point set compared to obtaining the first target semantic feature map by using the feature map output by a single feature extraction layer.

[0092] S202. Obtain the second eigenvalue of the second pixel point set in the second image through the first neural network, where the second pixel point set includes at least one pixel point.

[0093] In this embodiment, the second image can be regarded as an image that has been annotated. For ease of description, this application also refers to the second image as the annotated second image.

[0094] In this embodiment, after the annotation device obtains the first image, it inputs the second image into the first neural network to obtain the second eigenvalue of the second pixel point set.

[0095] In other words, in this embodiment, the second image is also input into the first neural network, and based on the first neural network, the pixel values of each pixel point in the second image are subjected to feature extraction to obtain the second eigenvalue of the second pixel point set.

[0096] It should be noted here that the second pixel point set in this embodiment is a set including at least one pixel point.

[0097] Specifically, when obtaining the second eigenvalue of the second pixel point set in the second image through the first neural network, it is similar to obtaining the first eigenvalue of the first pixel point set in the first image through the first neural network. The specific implementation can refer to the detailed description of obtaining the first eigenvalue of the first pixel point set in the first image in S201, which will not be elaborated here.

[0098] S203. Determine the second pixel point set in the second image that matches the first pixel point set according to the similarity between the first eigenvalue and the second eigenvalue.

[0099] Specifically, in this embodiment, after obtaining the first eigenvalue of the first pixel point set in the first image and the second eigenvalue of the second pixel point set in the second image, that is, after obtaining the semantic eigenvalue of the first pixel point set in the first image and the semantic eigenvalue of the second pixel point set in the second image, when the annotation device annotates the first pixel point set in the first image, it first determines the second pixel point set that matches the first pixel point set based on the similarity between the first eigenvalue and the second eigenvalue. Then, it determines the annotation of the matching second pixel point set as the annotation of the first pixel point set.

[0100] Exemplarily, when the annotation device determines a second pixel point set matching the first pixel point set from the second image based on the similarity between the first eigenvalue and the second eigenvalue, one implementation is as follows: The second pixel point set in the second image with an eigenvalue similarity greater than a preset value between the second eigenvalue and the first eigenvalue is used as the matching second pixel point set.

[0101] Furthermore, the second pixel point set in the second image with the greatest eigenvalue similarity between the second eigenvalue and the first eigenvalue can be used as the matching second pixel point set.

[0102] Specifically, when determining the second pixel point set in the second image with the greatest eigenvalue similarity between the second eigenvalue and the first eigenvalue, one implementation is as follows: The first target semantic feature map is reduced to a third target semantic feature map with a second resolution, where the second resolution is less than the first resolution; the second target semantic feature map is reduced to a fourth target semantic feature map with the second resolution; according to the first resolution and the second resolution, a third pixel point set corresponding to the first pixel point set in the third target semantic feature map is determined; a fourth pixel point set with the greatest similarity in pixel values between the third pixel point set is determined from the fourth target semantic feature map; according to the first resolution and the second resolution, a set of second pixel point sets corresponding to the fourth pixel point set is determined from the second image; finally, the second pixel point set with the greatest similarity between the eigenvalues in the set of second pixel point sets and the first eigenvalue is determined as the second pixel point set matching the first pixel point set.

[0103] Among them, the semantic feature map with the second resolution is the semantic feature map obtained by reducing the resolution of the first target semantic feature map and the second target semantic feature map.

[0104] In this implementation method, a matching relationship is first established at the second resolution based on the similarity between eigenvalues, and then a set of second pixel point sets in the second image is obtained based on the size relationship between the second resolution and the first resolution, and the second pixel point set matching the first pixel point set is determined from the set of second pixel point sets based on the greatest eigenvalue similarity.

[0105] Optionally, the first target semantic feature map can also be reduced to other preset resolutions to obtain the feature pyramids of the first image and the second image respectively. Among them, the other preset resolution is less than the resolution of the original image and greater than the second resolution (i.e., the second resolution is the minimum resolution).

[0106] In this case, when a matching relationship is established at the second resolution based on the similarity between the eigenvalue, when obtaining a set of second pixel point sets in the second image based on the size relationship between the second resolution and the first resolution, the following method can be adopted: As shown in Figure 6, according to the scaling relationship, obtain a set of pixel point sets corresponding to the unmatched lowest resolution semantic feature map of the pixel points matched in the semantic feature map at the lowest resolution corresponding to the second image; then, based on the maximum eigenvalue similarity, find the pixel point set matched by the first pixel point set at the unmatched lowest resolution from this set of pixel point sets. At this time, the unmatched lowest resolution becomes the matched resolution; finally, repeat the above process until the second pixel point set matched by the first pixel point in the second image is determined.

[0107] In this implementation method, when determining the second pixel point set matched with the first pixel point set in the second image, since the first target semantic feature map and the second target semantic feature map are first scaled to a smaller resolution, then a pixel matching relationship is established at the smaller resolution, and finally, based on the size relationship between the smaller resolution and the resolution of the original image, the second pixel point set matched with the first pixel point set is determined from the second image. It should be understood that this method has higher computational efficiency compared to directly establishing a pixel matching relationship from the first target semantic feature map and the second target semantic feature map.

[0108] Specifically, when determining the second pixel point set with the maximum eigenvalue similarity between the second eigenvalue and the first eigenvalue in the second image, another implementation method is: for the first pixel point set in the first image, use a specific search method to search for a set of second pixel point sets in the second image, and then use the second pixel point set with the maximum eigenvalue similarity between the first eigenvalue as the second pixel point set matched with the first pixel point set. Specifically, if the Euclidean distance or cosine similarity method is used to measure the similarity, the nearest neighbor search method is used to search for a set of second pixel point sets in the second image and then match based on the similarity between the eigenvalues. If the inner product distance method is used to measure the similarity, search methods such as the maximum inner product search are used to search for a set of second pixel point sets in the second image and then match based on the similarity between the eigenvalues.

[0109] It should be noted here that this embodiment does not limit how to define the similarity between the first eigenvalue and the second eigenvalue. For example, the similarity between the first eigenvalue and the second eigenvalue can be the Euclidean distance between the first eigenvalue and the second eigenvalue, the cosine distance between the first eigenvalue and the second eigenvalue, the inner product similarity between the first eigenvalue and the second eigenvalue, etc.

[0110] S204. Determine the annotation of the second pixel point set that matches the first pixel point set as the annotation of the first pixel point set.

[0111] In this embodiment, when it is determined that the second pixel point set in the second image that has been annotated is a pixel point set that matches the first pixel point set in the first image to be annotated, the annotation device determines the annotation of the first pixel point set as the annotation of the matching second pixel point set, that is, the annotations of the first pixel point set and the second pixel point set are the same.

[0112] In this embodiment, when the annotation device annotates the first pixel point set in the first image based on the second image that has been annotated, it is based on the similarity between the eigenvalue (also referred to as the first eigenvalue) obtained after inputting the first pixel point set in the first image into the first neural network and the eigenvalue (also referred to as the second eigenvalue) obtained after inputting the second pixel point set in the second image into the first neural network, as well as the annotation of the second pixel point set to determine the annotation of the first pixel point set. In this application, the eigenvalue obtained through the neural network is also referred to as the semantic eigenvalue. It should be understood that since the features extracted by the neural network can be considered as abstract semantic features related to tasks such as classification, compared with features such as pixel grayscale and optical flow, they are less sensitive to interferences such as illumination, deformation, occlusion, and rapid displacement. In addition, the semantic features have less noise compared to the features extracted by self-supervised learning. Therefore, in this embodiment, when establishing the matching relationship between the pixels of the first image and the second image based on the similarity between the semantic eigenvalues, the second pixel point set determined from the second image that matches the first pixel point set can be made more accurate, and further, the accuracy of automatic annotation can be improved.

[0113] The above describes the method of annotating the first pixel point set in the unannotated first image based on the second image that has been annotated. It should be understood that in many cases, many images need to be annotated simultaneously. For example, each pixel point in each frame of the video needs to be annotated. Next, a specific embodiment will be combined to illustrate a method of annotating each pixel point in each frame of the video.

[0114] As shown in FIG. 7, the method of annotating each pixel point in each frame of the video includes: S701, S702, S703, and S704.

[0115] S701. Obtain the video to be annotated, where the video to be annotated includes N frames of images.

[0116] Specifically, the video to be annotated in this embodiment can be considered as a video composed of N frames of images. More specifically, among these N frames of images, each pixel point in some of the images has been annotated by an annotator, and each pixel point in some other images is to be annotated.

[0117] S702. Use the first neural network to obtain the target semantic feature maps corresponding to each of the N frames of images, and obtain the feature values of each pixel point in each frame of image based on the target semantic feature maps corresponding to each frame of image.

[0118] Among them, the description of the first neural network can refer to the description in the embodiment shown in FIG. 2, and will not be elaborated here.

[0119] In the first implementation scheme, for each frame of the N frames of images in S701, use this frame of image as the input image of the first neural network, and obtain the feature map output by a certain feature extraction layer (which can be defined by the user) of the first neural network after this frame of image is input into the first neural network, and then scale the feature map output by this certain feature extraction layer to the first resolution to obtain the target semantic feature map, where the first resolution is the resolution of this frame of image. It should be understood that the pixel value at a certain position on the target semantic feature map is the feature value of the pixel point at the same position in the corresponding frame of image.

[0120] In the second implementation scheme, for each frame of the N frames of images in S701, use this frame of image as the input image of the first neural network, and obtain L (L>1) feature maps output by L feature extraction layers (which specific feature extraction layers these L feature extraction layers are can be defined by the user) of the first neural network after this frame of image is input into the first neural network, and then scale the L feature maps to L feature maps with the first resolution, where the first resolution is the resolution of the input image. Finally, determine the feature map obtained after fusing the L feature maps with the first resolution as the target semantic feature map.

[0121] It should be understood that since the feature maps output by different feature extraction layers may reflect different types of features of the image, the feature values of each pixel point determined by using the target feature map obtained from the feature maps output by the L feature extraction layers of the first neural network are more accurate than the feature values of each pixel point determined by using the feature map output by a single feature extraction layer to obtain the target feature map.

[0122] In this application, since the first target feature map or the second target feature map is an image used to obtain the feature value of each pixel point, in this application, the first target feature map or the second target feature map is also referred to as a pixel-level semantic feature map.

[0123] Among them, for a more detailed description of using the first neural network to obtain the pixel-level semantic feature map corresponding to each frame of image, it can be analogized to the description of using the first neural network to obtain the feature value of each pixel point in the first image in the embodiment shown in FIG. 2, and will not be elaborated here.

[0124] S703. Scale the target semantic feature maps corresponding to each frame of image to several preset resolutions, construct a feature pyramid for each frame of image, and establish a pixel point matching relationship for two consecutive frames of images using the feature pyramids corresponding to the two consecutive frames of images.

[0125] Specifically, in implementation, it includes:

[0126] Step 1: As shown in Figure 8, first, for the target semantic feature maps of each frame of the N frames of images obtained in S702, successively scale them down to several resolutions preset by the user from large to small, and then, together with the target semantic feature maps obtained in S702, construct a feature pyramid for each frame of image. Among them, different layers of the feature pyramid correspond to different resolutions.

[0127] In this embodiment, the image obtained by scaling the target semantic feature map to a certain resolution is called the semantic feature map at that certain resolution.

[0128] Step 2: For two consecutive frames of images, obtain the semantic feature maps with the lowest resolution in their respective corresponding feature pyramids. And aiming at the maximum similarity between semantic feature values, establish a matching relationship between the pixel points of the semantic feature maps with the lowest resolution corresponding to the two consecutive frames of images.

[0129] Here, the maximum similarity between semantic feature values can be measured by the shortest Euclidean distance between semantic features, the largest cosine similarity, the largest inner product similarity, etc., which does not constitute a limitation of this application.

[0130] Step 3: For two consecutive frames of images, establish a matching relationship between the pixel points of the two consecutive frames of images according to the matching relationship and scaling relationship between the pixel points of the semantic feature maps with the lowest resolution corresponding to the two consecutive frames of images.

[0131] In specific implementation, for two consecutive frames of images, assuming one frame of image is called the first frame image and the other frame of image is called the second frame image. At this time, when determining the matching pixel point of the first pixel point (which can be any pixel point in the first frame image) in the first frame image in the second frame image, as shown in Figure 9, first, through the scaling relationship, determine the positions of the first pixel point in the first frame image in the semantic feature maps of different resolutions; then, according to the matching relationship between the pixel points in the semantic feature map with the lowest resolution that has been matched corresponding to the first frame image and the second frame image, determine the pixel point in the semantic feature map with the lowest resolution corresponding to the second image that matches the first pixel point; after that, according to the scaling relationship, obtain a set of pixel points corresponding to the semantic feature map with the lowest resolution that has not been matched for the pixel point in the semantic feature map with the lowest resolution corresponding to the second image that matches; then, calculate the similarity between the semantic feature value of the pixel point corresponding to the first pixel point in the semantic feature map with the lowest resolution that has not been matched and the feature values of each pixel point in this set of pixel points, and determine the pixel point in the semantic feature map with the lowest resolution that has not been matched that matches the first pixel point. At this time, the semantic feature map with the lowest resolution that has not been matched becomes the semantic feature map with the resolution that has been matched; finally, repeat the above process until the pixel point in the pixel-level semantic feature map (with the highest resolution) corresponding to the second image that matches the first pixel point is determined.

[0132] It should be understood that there may be a situation where different pixel points in the same frame of image are matched to the same pixel point in another frame of image. Exemplarily, as shown in Figure 10, t and t + 1 represent two consecutive frames of images at time t and time t + 1, and the arrows point to the matching pixels. Two different pixel points in the image corresponding to time t + 1 are both matched to the same pixel point in the image corresponding to time t. And for the method provided in this embodiment, there is no requirement for a one-to-one correspondence between the pixels of two consecutive frames. A newly emerged pixel in a certain frame can also establish a many-to-one matching relationship with an existing pixel in another frame. Therefore, in fast-changing scenarios with a large number of newly added pixels and complex occlusion change scenarios, this application has stronger generalization performance and accuracy compared with methods such as optical flow.

[0133] S704, according to the partial images (also called sparse annotation frames) already annotated by the annotator and the pixel point matching relationship between two consecutive frames of images, use the label propagation method to gradually spread the labels from the sparse annotation frames to the unannotated frames to complete the annotation of all frames.

[0134] Specifically, in implementation, it includes:

[0135] Step 1: For any unlabeled frame, set it as the target frame. If neither of its two adjacent frames is labeled, the target frame is not processed temporarily. If only one of the adjacent frames is labeled, set the labeled frame as the reference frame. For any target pixel in the target frame, according to the matching relationship obtained in S703, find its matching pixel in the reference frame, and use the label of the matching pixel as the label of the target pixel. If both adjacent frames of the target frame are labeled, for any target pixel in the target frame, take the label of the matching pixel with a shorter distance among the adjacent frame matching pixels as the target pixel label. If the distances are the same, take the label of any matching pixel as the target label.

[0136] Step 2: Repeat Step 1 until all frames are labeled.

[0137] In this embodiment, by propagating the label of the annotator from one pixel to another, this label does not need to be consistent with the label used during the training of the pre-trained deep learning model. Therefore, this application can cover more labeling tasks. At the same time, using pixel semantic similarity for label propagation has higher accuracy compared to directly performing pre-labeling using a pre-trained deep learning model or performing pre-labeling after fine-tuning the pre-trained deep learning model with a small number of labeled frames.

[0138] In the above, in combination with FIGS. 2 to 10, the labeling method of the embodiment of the present application has been described in detail. Next, the labeling device according to the embodiment of the present application will be described in detail in combination with FIGS. 11 and 12.

[0139] FIG. 11 is a structural schematic diagram of a labeling device provided by an embodiment of the present application. Specifically, as shown in FIG. 11, the labeling device includes: an acquisition module 1101 and a processing module 1102.

[0140] In the first embodiment, the acquisition module 1101 is used to obtain the first eigenvalue of the first pixel point set in the first image through the first neural network. The first neural network is used to obtain the eigenvalue of the pixel points in the image, and the first pixel point set includes at least one pixel point; the acquisition module 1101 is further used to obtain the second eigenvalue of the second pixel point set in the second image through the first neural network, and the second pixel point set includes at least one pixel point; the processing module 1102 is used to determine the label of the first pixel point set according to the similarity between the first eigenvalue and the second eigenvalue, and the label of the second pixel point set.

[0141] In a possible implementation manner, the acquisition module 1101 is specifically used to: obtain the feature map output by the first neural network in the first feature extraction layer after the first image is input into the first neural network; the processing module 1102 is specifically used to: obtain the first target semantic feature map according to the feature map output by the first feature extraction layer; determine the first eigenvalue of the first pixel point set according to the first target semantic feature map.

[0142] In a possible implementation, the obtaining module 1101 is further configured to: obtain the feature map output by the first neural network at the second feature extraction layer after the first image is input into the first neural network; the processing module 1102 is further configured to: obtain a first target semantic feature map according to the feature map output by the first feature extraction layer and the feature map output by the second feature extraction layer.

[0143] In a possible implementation, the resolution of the first target semantic feature map is the same as the first resolution, and the first resolution is the resolution of the first image.

[0144] In a possible implementation, the obtaining module 1101 is further configured to: obtain a second target semantic feature map corresponding to a second image through the first neural network; the processing module 1102 is further configured to: reduce the first target semantic feature map to a third target semantic feature map with a second resolution, where the second resolution is less than the first resolution; reduce the second target semantic feature map to a fourth target semantic feature map with the second resolution; determine a third pixel point set corresponding to the first pixel point set in the third target semantic feature map according to the first resolution and the second resolution; determine a fourth pixel point set with the largest similarity of pixel values to the third pixel point set from the fourth target semantic feature map; determine a set of second pixel point sets corresponding to the fourth pixel point set from the second image according to the first resolution and the second resolution; determine the annotation of the second pixel point set with the largest similarity between the feature value in the set of second pixel point sets and the first feature value as the annotation of the first pixel point set.

[0145] In a possible implementation, the similarity includes any one of the following measurement methods: the Euclidean distance between feature values, the cosine distance between feature values, and the inner product similarity between feature values.

[0146] FIG. 12 is a structural schematic diagram of an annotation device provided by another embodiment of the present application. The device shown in FIG. 12 can be used to execute the method described in any of the foregoing embodiments.

[0147] As shown in FIG. 12, the device 1200 of this embodiment includes: a memory 1201 and a processor 1202. Optionally, the device 1200 further includes a communication interface 1203 and a bus 1204. Among them, the memory 1201, the processor 1202, and the communication interface 1203 are communicatively connected to each other through the bus 1204.

[0148] The memory 1201 may be a read only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1201 may store a program, and when the program stored in the memory 1201 is executed by the processor 1202, the processor 1202 is configured to execute each step of the method shown in FIGS. 2 to 10.

[0149] The processor 1202 may be a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, configured to execute relevant programs to implement the method shown in FIGS. 2 to 10 of the present application.

[0150] The processor 1202 may also be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the method of FIGS. 2 to 10 of the embodiments of the present application may be completed by the integrated logic circuit in the hardware of the processor 1202 or by instructions in software form.

[0151] The above-mentioned processor 1202 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0152] The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read only memory, a programmable read only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 1201, and the processor 1202 reads the information in the memory 1201 and combines its hardware to complete the functions required to be executed by the units included in the device of the present application. For example, it may execute each step / function of the embodiments shown in FIGS. 2 to 10.

[0153] The communication interface 1203 may use, but is not limited to, a transceiver device such as a transceiver to implement communication between the device 1200 and other devices or communication networks.

[0154] The bus 1204 may include a path for transmitting information between various components of the device 1200 (e.g., the memory 1201, the processor 1202, the communication interface 1203).

[0155] It should be understood that the device 1200 shown in the embodiments of the present application may be an electronic device, or alternatively, may be a chip configured in an electronic device. The device 1200 may be deployed at the vehicle end, or may also be deployed at the server end.

[0156] The above embodiments may be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more collections of available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. The semiconductor media may be a solid-state drive.

[0157] It should be understood that the term “and / or” herein is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B may be singular or plural. In addition, the character “ / ” herein generally represents an “or” relationship between the associated objects before and after, but may also represent an “and / or” relationship, which can be specifically understood with reference to the context.

[0158] In this application, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or a similar expression means any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0159] It should be understood that in various embodiments of this application, the magnitudes of the sequence numbers of the above - mentioned processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.

[0160] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0161] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0162] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0163] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0164] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0165] If the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs.

[0166] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

A labeling method characterized in that it includes obtaining a first eigenvalue of a first pixel point set in a first image through a first neural network, where the first neural network is used to obtain the eigenvalue of a pixel point in an image, and the first pixel point set includes at least one pixel point; obtaining a second eigenvalue of a second pixel point set in a second image through the first neural network, where the second pixel point set includes at least one pixel point; and determining the label of the first pixel point set according to the similarity between the first eigenvalue and the second eigenvalue and the label of the second pixel point set. The method according to claim 1 characterized in that the obtaining a first eigenvalue of a first pixel point set in a first image through a first neural network includes: obtaining a feature map output by the first neural network in a first feature extraction layer after the first image is input into the first neural network; obtaining a first target semantic feature map according to the feature map output by the first feature extraction layer; and determining the first eigenvalue of the first pixel point set according to the first target semantic feature map. The method according to claim 2 characterized in that the method further includes: obtaining a feature map output by the first neural network in a second feature extraction layer after the first image is input into the first neural network; where the obtaining a first target semantic feature map according to the feature map output by the first feature extraction layer includes: obtaining the first target semantic feature map according to the feature map output by the first feature extraction layer and the feature map output by the second feature extraction layer. The method according to claim 2 or 3 characterized in that the resolution of the first target semantic feature map is the same as a first resolution, and the first resolution is the resolution of the first image. The method according to claim 4 characterized in that the method further includes: obtaining a second target semantic feature map corresponding to the second image through the first neural network; reducing the first target semantic feature map to a third target semantic feature map with a second resolution, where the second resolution is less than the first resolution; reducing the second target semantic feature map to a fourth target semantic feature map with the second resolution; determining a third pixel point set corresponding to the first pixel point set in the third target semantic feature map according to the first resolution and the second resolution; determining a fourth pixel point set with the largest similarity of pixel values between the fourth pixel point set and the third pixel point set from the fourth target semantic feature map; determining a set of second pixel point sets corresponding to the fourth pixel point set from the second image according to the first resolution and the second resolution; where the determining the label of the first pixel point set according to the similarity between the first eigenvalue and the second eigenvalue and the label of the second pixel point set includes: determining the label of the second pixel point set with the largest similarity between the eigenvalue in the set of second pixel point sets and the first eigenvalue as the label of the first pixel point set. The method according to any one of claims 1 to 5 characterized in that The similarity includes any one of the following measurement methods: the Euclidean distance between eigenvalues, the cosine distance between eigenvalues, and the inner product similarity between eigenvalues. A labeling device characterized in that it includes: an acquisition module, configured to obtain a first eigenvalue of a first pixel point set in a first image through a first neural network, where the first neural network is used to obtain eigenvalues of pixel points in an image, and the first pixel point set includes at least one pixel point; the acquisition module is further configured to: obtain a second eigenvalue of a second pixel point set in a second image through the first neural network, where the second pixel point set includes at least one pixel point; a processing module, configured to determine a label of the first pixel point set according to the similarity between the first eigenvalue and the second eigenvalue, and the label of the second pixel point set. The device according to claim 7 characterized in that the acquisition module is specifically configured to: obtain a feature map output by the first neural network at a first feature extraction layer after the first image is input into the first neural network; the processing module is specifically configured to: obtain a first target semantic feature map according to the feature map output by the first feature extraction layer; determine a first eigenvalue of the first pixel point set according to the first target semantic feature map. The device according to claim 8 characterized in that the acquisition module is further configured to: obtain a feature map output by the first neural network at a second feature extraction layer after the first image is input into the first neural network; the processing module is further configured to: obtain the first target semantic feature map according to the feature map output by the first feature extraction layer and the feature map output by the second feature extraction layer. The device according to claim 8 or 9 characterized in that the resolution of the first target semantic feature map is the same as a first resolution, and the first resolution is the resolution of the first image. The device according to claim 10 characterized in that the acquisition module is further configured to: obtain a second target semantic feature map corresponding to the second image through the first neural network; the processing module is further configured to: reduce the first target semantic feature map to a third target semantic feature map with a second resolution, where the second resolution is less than the first resolution; reduce the second target semantic feature map to a fourth target semantic feature map with the second resolution; determine a third pixel point set corresponding to the first pixel point set in the third target semantic feature map according to the first resolution and the second resolution; determine a fourth pixel point set with the largest similarity of pixel values between the fourth target semantic feature map and the third pixel point set; determine a set of second pixel point sets corresponding to the fourth pixel point set in the second image according to the first resolution and the second resolution; determine the label of the second pixel point set with the largest similarity between the eigenvalues in the set of second pixel point sets and the first eigenvalue as the label of the first pixel point set. The device according to any one of claims 7 to 11 characterized in that The similarity includes any one of the following measurement methods: the Euclidean distance between eigenvalues, the cosine distance between eigenvalues, and the inner product similarity between eigenvalues. An annotation device for training data of a semantic segmentation neural network, characterized in that it includes: a memory and a processor; the memory is used to store program instructions; the processor is used to call the program instructions in the memory to execute the image processing method according to any one of claims 1 to 6. A chip, characterized in that it includes at least one processor and a communication interface, the communication interface and the at least one processor are interconnected by a line, and the at least one processor is used to run a computer program or instruction to execute the method according to any one of claims 1 to 6. A computer-readable medium, characterized in that the computer-readable medium stores program code for computer execution, and the program code includes instructions for executing the method according to any one of claims 1 to 6. A computer program product, which includes computer program code, characterized in that when the computer program code runs on a computer, the computer is caused to implement the method according to any one of claims 1 to 6.