Image annotation method, device, electronic device and storage medium

By tracking and detecting multi-frame images in the video stream and labeling attribute information, the problem of high cost and long period of manual labeling in image recognition is solved, and efficient and low-cost image labeling is achieved.

CN114419493BActive Publication Date: 2025-05-27BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111642573.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-05-27
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

In the prior art, image recognition requires a large amount of manual labeling data, resulting in high labeling costs and long cycles, which are difficult to effectively reduce.

Method used

By obtaining multiple frame images from the video stream to be marked, tracking detection is performed to determine the detection frame information, determine the target frame image and object attribute information, and then label the detection frame attribute information.

Benefits of technology

It greatly reduces the annotation cost and cycle, improves the annotation efficiency, and shortens the annotation time by uniformly labeling the attribute information of the object.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114419493B_ABST
    Figure CN114419493B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, apparatus, electronic device, and storage medium for image annotation, relating to the field of image processing technologies, and particularly to artificial intelligence fields such as computer vision, deep learning, and cloud services. The specific solution is as follows: Obtain multiple frames of images from a video stream to be annotated; perform tracking detection on the multiple frames of images to determine the detection box information included in each frame of image; based on the detection box information included in each frame of image, determine the target frame image corresponding to the identifier of each detection box; based on the target frame image corresponding to the identifier of each detection box, determine the attribute information of the object corresponding to the identifier of each detection box; and perform attribute information annotation on the detection boxes included in each frame of image according to the attribute information of the object corresponding to the identifier of each detection box and the detection box information included in each frame of image. By performing unified attribute information annotation on the object in each frame of image according to the attribute information of the object corresponding to the identifier of the detection box, the annotation cost and annotation cycle are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technologies, and particularly to artificial intelligence fields such as computer vision, deep learning, and cloud services. Specifically, it relates to an image annotation method, apparatus, electronic device, and storage medium. Background Art

[0002] When performing image recognition, a large amount of labeled data is required. In the related art, data annotation is usually carried out manually, with a large amount of labeled data, high costs, and a long annotation cycle.

[0003] Therefore, how to reduce the annotation cost and annotation cycle is an urgent problem to be solved. Summary of the Invention

[0004] The present disclosure provides an image annotation method, apparatus, electronic device, and storage medium.

[0005] According to one aspect of the present disclosure, there is provided an image annotation method, including:

[0006] Obtaining multiple frames of images from a video stream to be annotated;

[0007] Performing tracking detection on the multiple frames of images to determine the detection box information included in each frame of the images, where the detection box information includes the identifier, position, and / or size of the detection box;

[0008] Determining, according to the detection box information included in each frame of the images, the target frame image corresponding to the identifier of each detection box;

[0009] Determining, based on the target frame image corresponding to the identifier of each detection box, the attribute information of the object corresponding to the identifier of each detection box;

[0010] Performing attribute information annotation on the detection boxes included in each frame of the images according to the attribute information of the object corresponding to the identifier of each detection box and the detection box information included in each frame of the images.

[0011] According to another aspect of the present disclosure, there is provided an image annotation apparatus, including:

[0012] An obtaining module, configured to obtain multiple frames of images from a video stream to be annotated;

[0013] A detection module, configured to perform tracking detection on the multiple frames of images to determine the detection box information included in each frame of the images, where the detection box information includes the identifier, position, and / or size of the detection box;

[0014] A first determination module, configured to determine a target frame image corresponding to the identifier of each detection box according to the detection box information included in each frame of the image;

[0015] A second determination module, configured to determine the attribute information of the object corresponding to the identifier of each detection box based on the target frame image corresponding to the identifier of each detection box;

[0016] An annotation module, configured to perform attribute information annotation on the detection boxes included in each frame of the image according to the attribute information of the object corresponding to the identifier of each detection box and the detection box information included in each frame of the image.

[0017] According to another aspect of the present disclosure, there is provided an electronic device, including:

[0018] At least one processor; and

[0019] A memory communicatively connected to the at least one processor; wherein,

[0020] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in the above embodiments.

[0021] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the above embodiments.

[0022] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, and the computer program implements the steps of the method described in the above embodiments when executed by a processor.

[0023] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0024] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0025] Figure 1 is a schematic flowchart of an image annotation method provided by an embodiment of the present disclosure;

[0026] Figure 2 is a schematic flowchart of an image annotation method provided by another embodiment of the present disclosure;

[0027] Figure 3Schematic flowchart of an image annotation method provided by another embodiment of the present disclosure;

[0028] Figure 4 Schematic diagram of an image annotation process provided by another embodiment of the present disclosure;

[0029] Figure 5 Schematic structural diagram of an image annotation device provided by an embodiment of the present disclosure;

[0030] Figure 6 Block diagram of an electronic device for implementing the image annotation method of the embodiments of the present disclosure. Detailed implementation manners

[0031] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0032] The following describes an image annotation method, device, electronic device, and storage medium of the embodiments of the present disclosure with reference to the accompanying drawings.

[0033] Artificial intelligence is a discipline that studies the use of computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), and it has technical fields at both the hardware and software levels. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, as well as deep learning, big data processing technology, and knowledge graph technology.

[0034] Computer vision is a science that studies how to enable machines to "see", which refers to using cameras and computers to replace the human eye to perform machine vision such as target recognition, tracking, and measurement on targets, and further perform graphic processing to make the computer-processed images more suitable for human eye observation or transmission to instruments for detection.

[0035] Deep learning is a new research direction in the field of machine learning. Deep learning is to learn the internal laws and representation levels of sample data, and the information obtained during these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Its ultimate goal is to enable machines to have the ability of analysis and learning like humans, and be able to recognize data such as text, images, and sounds.

[0036] Figure 1 Schematic flowchart of an image annotation method provided by an embodiment of the present disclosure.

[0037] As Figure 1 shown, the annotation method for the image includes:

[0038] Step 101, obtain multiple frames of images from the video stream to be annotated.

[0039] Since there are a large number of images in the video stream, in order to improve the annotation efficiency, in the present disclosure, multiple frames of images can be extracted from the video stream to be annotated according to a preset frame extraction frequency or a preset frame extraction interval.

[0040] For example, if the video stream to be annotated is a 10-minute video stream with a frame rate of 25 frames per second, one frame of image can be extracted from the video stream every 5 frames, so that 3000 frames of images can be obtained.

[0041] It should be noted that the frame extraction mode or the extraction time interval can be determined according to actual needs, and the present disclosure does not limit this.

[0042] Step 102, perform tracking detection on the multiple frames of images to determine the detection box information included in each frame of image.

[0043] Among them, the detection box information may include but is not limited to the identifier, position, size, etc. of the detection box.

[0044] In the present disclosure, the multiple frames of images obtained can be sequentially input into an object detection model trained in advance, and the detection model is used to perform tracking detection on the objects in the multiple frames of images to determine the detection box information included in each frame of image. Among them, the object detection model can be trained by means of deep learning.

[0045] In addition, the object for tracking detection can be a vehicle or a person, etc., which can be determined according to actual needs, and the present disclosure does not limit this.

[0046] For example, tracking detection can be performed on multiple frames of images obtained from a video stream captured by a camera at a certain intersection to determine the detection box information of vehicles in each frame of image.

[0047] It can be understood that performing tracking detection on multiple frames of images can make the identifiers of the detection boxes of the same object in different frames of images the same. For example, if a vehicle appears in the first frame of image and the second frame of image, then the identifier of the detection box of the vehicle in the two frames of images is the same, that is, the vehicle has the same identifier in the two frames of images.

[0048] Step 103, determine the target frame image corresponding to the identifier of each detection box according to the detection box information included in each frame of image.

[0049] In the present disclosure, cropping can be performed according to the coordinates of the detection boxes included in each frame of image to obtain sub-images corresponding to each detection box, and according to the identifier of each detection box, sub-images corresponding to the identifier of the same detection box can be determined, so that sub-images corresponding to the identifier of each detection box can be determined. Then, according to the sizes of the sub-images corresponding to the identifier of each detection box, target sub-images can be determined from the sub-images, and the frame image where the target sub-image is located can be determined as the target frame image corresponding to the identifier of each detection box. Among them, the target sub-image can be the sub-image with the largest size, or the sub-image with the relatively largest object and the least truncation and occlusion, etc.

[0050] Step 104: Based on the target frame image corresponding to the identifier of each detection box, determine the attribute information of the object corresponding to the identifier of each detection box.

[0051] In the present disclosure, a pre-trained recognition model can be used to recognize the object corresponding to the identifier of the detection box in the target frame image corresponding to the identifier of each detection box to determine the attribute information of the object.

[0052] If the target frame image is determined based on the target sub-image corresponding to the identifier of the detection box, then the target sub-image can be recognized to determine the attribute information of the object corresponding to the identifier of the detection box.

[0053] For example, if the object is a vehicle, the determined attribute information of the vehicle may include the type of the vehicle (such as a truck, a lorry, a muck truck, a sedan, a tricycle, a motorcycle, a bicycle, etc.), the color of the vehicle (such as red, white, black, orange, silver, pink, cyan, purple, green, etc.), the orientation of the vehicle (such as forward, backward, left, right, etc.), etc.

[0054] Step 105: According to the attribute information of the object corresponding to the identifier of each detection box and the detection box information included in each frame of image, perform attribute information annotation on the detection boxes included in each frame of image.

[0055] In the present disclosure, based on the attribute information of the object corresponding to the identifier of the detection box determined from the target frame image, attribute information annotation can be performed on the detection boxes corresponding to the identifier of the detection box in other frame images.

[0056] For example, in multiple frame images of a video stream, vehicle A and vehicle B appear. Among them, the attribute information of vehicle A is determined according to the third frame image, and the attribute information of vehicle B is determined based on the fourth frame image. Then, according to the attribute information of vehicle A determined from the third frame image, attribute information annotation can be performed on the detection box of vehicle A in other frame images, and according to the attribute information of vehicle B determined from the fourth frame image, attribute information annotation can be performed on the detection box of vehicle B in other frame images.

[0057] During implementation, the attribute information of the object corresponding to the identifier of each detection box in each frame of image and the detection box information included in each frame of image can be used to label the attribute information of each detection box, so as to obtain the labeling result corresponding to each frame of image. Among them, the labeling result includes each detection box information, the attribute information of the object corresponding to the identifier of each detection box, etc.

[0058] The image labeling method according to the embodiments of the present disclosure can uniformly label the attribute information of the detection box in each frame of image based on the attribute information of the object corresponding to the identifier of each detection box, thereby greatly reducing the labeling cost.

[0059] For example, when it is necessary to label a 10-minute video stream with a frame rate of 25 frames per second, and there are a total of 100 vehicles in the video stream, with an average of 10 vehicles per frame. If one frame is selected for labeling every 5 frames, 3000 frames of images need to be labeled. In the related art, if 10 vehicle detection boxes are labeled per frame, and if the vehicle has a total of 40 attributes and each attribute item has an average of 5 attribute values, then a total of 3000 * 10 * 40 * 5 = 6 million attribute labels need to be labeled.

[0060] However, by using the image labeling method of the present disclosure, the number of labels to be labeled is directly proportional to the number of vehicles appearing in the video, 100 * 40 * 5 = 20,000 attribute labels, reducing the cost by 300 times.

[0061] The image labeling method of the present disclosure can be applied to the evaluation of video analysis applications. For example, a certain video can be first analyzed by the video analysis application to be evaluated, and the position information of the object, the attribute information of the object, etc. are output, and the video is processed by using the image labeling method of the present disclosure to obtain the labeling result of each frame of image in the video stream. Then, according to the comparison between the labeling result and the output result of the video analysis application, the video analysis application is evaluated according to the comparison result.

[0062] In addition, the video stream labeled by the image labeling method of the present disclosure can also be used as training data to train a model, so as to obtain a model capable of analyzing the video stream and obtaining the detection box information, attribute information, etc. of the objects in the video stream.

[0063] The image annotation method according to the embodiments of the present disclosure determines the information of the detection boxes included in each frame of image by performing tracking annotation on multiple frames of images obtained from the video stream to be annotated, and determines the target frame image corresponding to the identifier of each detection box based on the detection box information included in each frame of image. Based on the target frame image corresponding to the identifier of each detection box, the attribute information of the object corresponding to the identifier of each detection box is determined, and the detection boxes included in each frame of image are annotated with attribute information according to the attribute information of the object corresponding to the identifier of each detection box and the detection box information included in each frame of image. Thus, by using the attribute information of the object corresponding to the identifier of the detection box determined according to the target frame image to uniformly annotate the attribute information of the object in each frame of image, the annotation cost is greatly reduced and the annotation cycle is shortened.

[0064] Figure 2 It is a schematic flowchart of the image annotation method provided by another embodiment of the present disclosure.

[0065] As Figure 2 shown, the image annotation method includes:

[0066] Step 201, obtain multiple frames of images from the video stream to be annotated.

[0067] Step 202, perform tracking detection on the multiple frames of images to determine the detection box information included in each frame of image.

[0068] In the present disclosure, steps 201 - 202 are similar to the content described in the above embodiments, so details are not described herein again.

[0069] Step 203, determine the detection box information corresponding to the same detection box identifier according to the detection box identifier in the detection box information included in each frame of image.

[0070] Since the same object may appear in multiple consecutive frames of images, that is, the detection boxes of the same object may be included in multiple frames of images, in the present disclosure, the detection box information corresponding to the same detection box identifier in multiple frames of images can be determined according to the detection box identifier in the detection box information included in each frame of image.

[0071] Step 204, determine the target detection box corresponding to the identifier of each detection box according to the size and / or position of the detection box in the detection box information corresponding to the same detection box identifier.

[0072] Since the object may be moving, its position in multiple frames of images may change, resulting in different sizes and positions of the detection boxes of the object in different frames of images. Therefore, the accuracy of the attribute information of the object determined based on different frames of images is also different. Based on this, in the present disclosure, the target detection box of the object can be determined from the detection boxes of the same object, that is, the target detection box corresponding to the detection box identifier can be determined from the detection boxes corresponding to the identifier of the same detection box.

[0073] Since the larger the size of the detection box, the larger and clearer the object is, the sizes of the detection boxes can be compared to determine the detection box with the largest size among the detection boxes as the target detection box.

[0074] Due to different shooting angles, the position of the object in the image may be different. Therefore, the detection boxes within the preset area of the image among the detection boxes can also be determined as the target detection boxes. Among them, the preset area can be the middle area, the left area, the right area, etc., which can be determined according to needs, and the present disclosure does not limit this.

[0075] For example, due to the position of the camera, the larger the proportion of the object in the left area of the captured image in the image, the left area can be used as the preset area.

[0076] Alternatively, the detection boxes with sizes greater than the threshold and positions within the preset area of the image among the detection boxes can also be determined as the target detection boxes. Among them, the threshold can be determined according to actual needs, and the present disclosure does not limit this.

[0077] In the present disclosure, the target detection box can be determined according to the sizes and / or positions of the detection boxes, enriching the determination method of the target detection box and meeting diverse requirements.

[0078] Step 205: Determine the frame image where the target detection box corresponding to the identifier of each detection box is located as the target frame image corresponding to the identifier of each detection box.

[0079] Since the attribute information of the object corresponding to the identifier of the detection box is relatively easy to identify in the target detection box corresponding to the identifier of the detection box, the frame image where the target detection box corresponding to the identifier of each detection box is located can be determined as the target frame image corresponding to the identifier of each detection box.

[0080] For example, if the target frame image corresponding to the identifier of a certain detection box is the fourth frame image among the captured multiple frame images, the fourth frame image can be used as the target frame image corresponding to the identifier of this detection box.

[0081] Step 206: Determine the attribute information of the object corresponding to the identifier of each detection box based on the target frame image corresponding to the identifier of each detection box.

[0082] Step 207: Label the attribute information for each detection box included in each frame of image according to the attribute information of the object corresponding to the identifier of each detection box and the detection box information included in each frame of image.

[0083] In this disclosure, Steps 206 - 207 are similar to the content described in the above embodiments, so they will not be elaborated here.

[0084] In the embodiment of this disclosure, when determining the target frame image corresponding to the identifier of each detection box according to the detection box information included in each frame of image, by determining the detection box information corresponding to the identifier of the same detection box according to the identifier of each detection box in each frame of image, determining the target detection box based on the size and / or position of the detection box in each detection information, and determining the target frame image based on the frame image where the target detection box is located, the accuracy of the target frame image is improved, and the accuracy of the annotation is improved.

[0085] Figure 3 It is a schematic flowchart of an image annotation method provided by another embodiment of this disclosure.

[0086] As Figure 3 shown, the image annotation method includes:

[0087] Step 301: Determine the scene to which the video stream belongs.

[0088] In practical applications, the objects to be annotated are different in different scenes, and the moving speeds of the objects may also be different, so the durations that appear in the video stream will also be different.

[0089] For example, for the video stream captured by a road intersection camera device, the object to be annotated is a vehicle, while for the video stream captured by a camera device at the entrance of a certain park, the object to be annotated is a pedestrian. It can be seen that the objects to be annotated are different in the two scenes, and the moving speeds of the objects are also different.

[0090] In this disclosure, the scene to which the video stream belongs can be determined according to the location information of the shooting device of the video stream.

[0091] Step 302: Determine the frame extraction mode corresponding to the video stream according to the belonging scene.

[0092] Among them, the frame extraction mode can refer to the frame extraction frequency, frame extraction interval, etc.

[0093] Since the scenes to which the video streams belong are different, the durations of the objects in the video streams to be annotated in the video streams may be different. Therefore, in this disclosure, the frame extraction mode corresponding to the video stream can be determined according to the scene to which the video stream belongs and the corresponding relationship between the scene and the frame extraction mode.

[0094] For example, the moving speed of vehicles in the video stream of a certain intersection is faster than the moving speed of pedestrians in the video stream of the park entrance. If the frame extraction interval corresponding to the video stream of the intersection is relatively large, it is possible that some vehicles are relatively small in the extracted images or do not appear in the multiple frames of extracted images, which will affect the accuracy of annotation. Therefore, the frame extraction interval corresponding to the video stream of the intersection can be smaller than that of the video stream of the park entrance.

[0095] Step 303: Obtain multiple frames of images from the video stream based on the frame extraction mode.

[0096] Step 304: Perform tracking detection on the multiple frames of images to determine the detection box information included in each frame of image.

[0097] Step 305: Determine the target frame image corresponding to the identifier of each detection box according to the detection box information included in each frame of image.

[0098] Step 306: Determine the attribute information of the object corresponding to the identifier of each detection box based on the target frame image corresponding to the identifier of each detection box.

[0099] Step 307: Perform attribute information annotation on the detection boxes included in each frame of image according to the attribute information of the object corresponding to the identifier of each detection box and the detection box information included in each frame of image.

[0100] In this disclosure, steps 303 - 307 are similar to the content described in the above embodiments, so they will not be elaborated here.

[0101] In the embodiments of this disclosure, when obtaining multiple frames of images from the video stream to be annotated, the scene to which the video stream belongs is determined; according to the belonging scene, the frame extraction mode corresponding to the video stream is determined; and multiple frames of images are obtained from the video stream based on the frame extraction mode. Thus, by determining the frame extraction mode corresponding to the video stream according to the scene to which the video stream to be annotated belongs, multiple frames of images are obtained from the video stream using the frame extraction mode corresponding to the scene, improving the annotation accuracy.

[0102] In order to further reduce the annotation cost, in an embodiment of this disclosure, attribute information annotation can be performed on the detection boxes included in the target area of each frame of image.

[0103] In practical applications, for different scenes, the areas that users are concerned about may be different. Therefore, in this disclosure, users can input the position information of the area of concern in the scene to which the video stream belongs, and thus the position information of the area of concern can be obtained. Then, according to the position information of the area of concern, the target area in each frame of image can be determined, and tracking detection is performed on the target area in each frame of image to determine the detection box information included in each frame of image.

[0104] For example, in the video stream of a certain intersection, if the user is more concerned about the vehicles on a certain side of the road, then the position information of that side of the road can be used to determine the area where that side of the road is located in each frame of the image, that is, the target area, so as to perform tracking detection on the target area to determine the detection box information included in the target area in each frame of the image.

[0105] For another example, in the video stream of a certain intersection, if the user is more concerned about the vehicles within a preset range from the intersection, the position information corresponding to the preset range from the intersection can be used to determine the target area in each frame of the image, so as to perform tracking detection on the vehicles in the target area, which can greatly reduce the number of vehicles to be labeled and reduce the labeling cost.

[0106] In the embodiments of the present disclosure, when performing tracking detection on multiple frames of images to determine the detection box information included in each frame of the image, the position information of the area of interest in the scene to which the video stream belongs is obtained; according to the position information of the area of interest, the target area in each frame of the image is determined; and tracking detection is performed on the target area in each frame of the image to determine the detection box information included in each frame of the image. Thus, by determining the target area in each frame of the image according to the position information of the area of interest in the scene to which the video stream belongs and performing tracking detection on the target area in each frame of the image, the personalized labeling requirements are met, the number of objects to be labeled is reduced, and the labeling cost is reduced.

[0107] To further illustrate the above embodiments, the following is combined with Figure 4 for illustration. Figure 4 FIG. is a schematic diagram of the labeling process of an image provided by another embodiment of the present disclosure.

[0108] As Figure 4 shown, the video stream to be labeled is obtained, and the video stream to be labeled is frame-decomposed for the entire frame. After that, the frame extraction interval can be determined according to the scene to which the video stream to be labeled belongs, and multiple frames of images can be obtained from the video stream according to the frame extraction interval.

[0109] After obtaining multiple frames of images, a detection model can be used for tracking detection to determine the detection box information included in each frame of the image. Among them, the detection box information includes the identification, size, position, coordinates, etc. of the detection box. Alternatively, the model can also be used to first determine the size, position, etc. of the detection box included in each frame of the image, and then determine the identification of the detection box corresponding to the same object according to the position of the detection box in each frame of the image.

[0110] To improve the accuracy of labeling, in the present disclosure, the detection box can also be corrected manually, such as correcting the incorrect detection box, determining the detection box information for the missed object, correcting the detection box, etc., so as to improve the accuracy of the detection box information.

[0111] After that, cropping can be performed according to the positions of the detection boxes in the detection box information included in each frame of image to obtain sub-images corresponding to each detection box, and based on the identifiers of each detection box in each frame of image, the sub-images corresponding to the identifier of the same detection box are determined, the target sub-image is determined from the sub-images, and the attribute information of the object corresponding to the identifier of the detection box in each frame of image is annotated based on the target sub-image. After obtaining the annotation result, the annotation is completed.

[0112] Alternatively, it is also possible to determine the target frame image corresponding to the identifier of each detection box according to the detection box information included in each frame of image, determine the attribute information of the object corresponding to the identifier of each detection box based on the target frame image corresponding to the identifier of each detection box, and perform attribute information annotation on the detection boxes included in each frame of image according to the detection box information included in each frame of image and the attribute information of the object corresponding to the identifier of each detection box to obtain the annotation result.

[0113] To implement the above embodiments, the present disclosure also proposes an image annotation device. Figure 5 FIG. is a schematic structural diagram of an image annotation device provided by an embodiment of the present disclosure.

[0114] As Figure 5 shown, the image annotation device 500 includes:

[0115] An acquisition module 510, configured to acquire multiple frames of images from a video stream to be annotated;

[0116] A detection module 520, configured to perform tracking detection on the multiple frames of images to determine the detection box information included in each frame of the image, where the detection box information includes the identifier, position, and / or size of the detection box;

[0117] A first determination module 530, configured to determine the target frame image corresponding to the identifier of each detection box according to the detection box information included in each frame of the image;

[0118] A second determination module 540, configured to determine the attribute information of the object corresponding to the identifier of each detection box based on the target frame image corresponding to the identifier of each detection box;

[0119] An annotation module 550, configured to perform attribute information annotation on the detection boxes included in each frame of the image according to the attribute information of the object corresponding to the identifier of each detection box and the detection box information included in each frame of the image.

[0120] In a possible implementation manner of the embodiments of the present disclosure, the first determination module 530 includes:

[0121] The first determination unit is configured to determine, according to the identifier of the detection box in the detection box information included in each frame of the image, the detection box information corresponding to the identifier of the same detection box;

[0122] The second determination unit is configured to determine, according to the size and / or position of the detection box in the detection box information corresponding to the identifier of the same detection box, the target detection box corresponding to the identifier of each detection box;

[0123] The third determination unit is configured to determine the frame image where the target detection box corresponding to the identifier of each detection box is located as the target frame image corresponding to the identifier of each detection box.

[0124] In a possible implementation manner of the embodiment of the present disclosure, the second determination unit is configured to:

[0125] Compare the sizes of the detection boxes to determine the detection box with the largest size among the detection boxes as the target detection box;

[0126] Alternatively, determine the detection box whose position is within a preset area of the image among the detection boxes as the target detection box;

[0127] Alternatively, determine the detection box whose size is greater than a threshold and whose position is within a preset area of the image among the detection boxes as the target detection box.

[0128] In a possible implementation manner of the embodiment of the present disclosure, the obtaining module 510 is configured to:

[0129] Determine the scene to which the video stream belongs;

[0130] According to the scene to which it belongs, determine the frame extraction mode corresponding to the video stream;

[0131] Based on the frame extraction mode, obtain the multiple frames of images from the video stream.

[0132] In a possible implementation manner of the embodiment of the present disclosure, the detection module 520 is configured to:

[0133] Obtain the position information of the area of interest in the scene to which the video stream belongs;

[0134] According to the position information of the area of interest, determine the target area in each frame of the image;

[0135] Perform tracking detection on the target area in each frame of the image to determine the detection box information included in each frame of the image.

[0136] It should be noted that the explanations of the foregoing embodiments of the image annotation method also apply to the image annotation device of this embodiment, so details are not described herein again.

[0137] The annotation device for an image according to an embodiment of the present disclosure determines information on detection frames included in each frame of image by performing tracking annotation on multiple frames of images obtained from a video stream to be annotated, and determines a target frame image corresponding to the identifier of each detection frame based on the detection frame information included in each frame of image. Based on the target frame image corresponding to the identifier of each detection frame, attribute information of the object corresponding to the identifier of each detection frame is determined, and the detection frames included in each frame of image are annotated with attribute information according to the attribute information of the object corresponding to the identifier of each detection frame and the detection frame information included in each frame of image. Thus, by annotating the object in each frame of image with the attribute information of the object corresponding to the identifier of the detection frame determined according to the target frame image, the annotation cost is greatly reduced and the annotation cycle is shortened.

[0138] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0139] Figure 6 FIG. shows a schematic block diagram of an exemplary electronic device 600 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0140] As Figure 6 shown, the device 600 includes a computing unit 601 that can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 602 or a computer program loaded from a storage unit 608 into a RAM (Random Access Memory) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An I / O (Input / Output) interface 605 is also connected to the bus 604.

[0141] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as a keyboard, mouse, etc.; output unit 607, such as various types of displays, speakers, etc.; storage unit 608, such as a disk, optical disc, etc.; and communication unit 609, such as a network card, modem, wireless communication transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0142] Computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 601 include but are not limited to CPU (Central Processing Unit), GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. Computing unit 601 executes the various methods and processes described above, such as the method for annotating images. For example, in some embodiments, the method for annotating images can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by computing unit 601, one or more steps of the method for annotating images described above can be executed. Alternatively, in other embodiments, computing unit 601 can be configured to execute the method for annotating images by any other suitable means (e.g., by means of firmware).

[0143] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SoCs (System-On-Chip systems), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0144] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0145] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only Memory), or a flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0146] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0147] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a blockchain network.

[0148] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of high management difficulty and weak business scalability existing in traditional physical hosts and VPS services (Virtual Private Server). The server may also be a server of a distributed system or a server combined with a blockchain.

[0149] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, which, when executed by an instruction processor in the computer program product, executes the image annotation method proposed in the above embodiments of the present disclosure.

[0150] It should be understood that various forms of the processes shown above may be used, with steps reordered, added, or deleted. For example, the steps described in the present disclosure may be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitation is imposed herein.

[0151] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A method for annotating an image, comprising: obtaining multiple frames of images from a video stream to be annotated; performing tracking detection on the multiple frames of images to determine the detection box information included in each frame of the images, wherein the detection box information includes the identifier, position, and / or size of the detection box, and the identifiers of the detection boxes of the same object in different frames of images are the same; determining, according to the detection box information included in each frame of the images, the target frame image corresponding to the identifier of each detection box; determining, based on the target frame image corresponding to the identifier of each detection box, the attribute information of the object corresponding to the identifier of each detection box; annotating the attribute information of the detection boxes included in each frame of the images according to the attribute information of the object corresponding to the identifier of each detection box and the detection box information included in each frame of the images.

2. The method according to claim 1, wherein, the determining, according to the detection box information included in each frame of the images, the target frame image corresponding to the identifier of each detection box includes: determining, according to the identifier of the detection box in the detection box information included in each frame of the images, the detection box information corresponding to the same detection box identifier; determining, according to the size and / or position of the detection box in the detection box information corresponding to the same detection box identifier, the target detection box corresponding to the identifier of each detection box; determining the frame image where the target detection box corresponding to the identifier of each detection box is located as the target frame image corresponding to the identifier of each detection box.

3. The method according to claim 2, wherein, the determining, according to the size and / or position of the detection box in the detection box information corresponding to the same detection box identifier, the target detection box corresponding to the identifier of each detection box includes: comparing the sizes of the detection boxes to determine the detection box with the largest size among the detection boxes as the target detection box; alternatively, determining the detection box whose position is within a preset area of the image as the target detection box; alternatively, determining the detection box whose size is greater than a threshold and whose position is within a preset area of the image as the target detection box.

4. The method according to claim 1, wherein, the obtaining multiple frames of images from a video stream to be annotated includes: determining the scene to which the video stream belongs; determining the frame extraction mode corresponding to the video stream according to the belonging scene; obtaining the multiple frames of images from the video stream based on the frame extraction mode.

5. The method according to claim 1, wherein, the performing tracking detection on the multiple frames of images to determine the detection box information included in each frame of the images includes: obtaining the position information of the area of interest of the scene to which the video stream belongs; determining the target area in each frame of the images according to the position information of the area of interest; performing tracking detection on the target area in each frame of the images to determine the detection box information included in each frame of the images.

6. An image annotation device, comprising: an obtaining module, configured to obtain multiple frames of images from a video stream to be annotated; A detection module, configured to perform tracking detection on the multi-frame images to determine the detection box information included in each frame of the images, where the detection box information includes the identifier, position, and / or size of the detection box, and the identifiers of the detection boxes of the same object in different frame images are the same; A first determination module, configured to determine the target frame image corresponding to the identifier of each detection box according to the detection box information included in each frame of the images; A second determination module, configured to determine the attribute information of the object corresponding to the identifier of each detection box based on the target frame image corresponding to the identifier of each detection box; A labeling module, configured to perform attribute information labeling on the detection boxes included in each frame of the images according to the attribute information of the object corresponding to the identifier of each detection box and the detection box information included in each frame of the images.

7. The apparatus according to claim 6, wherein, the first determination module includes: A first determination unit, configured to determine the detection box information corresponding to the same detection box identifier according to the detection box identifier in the detection box information included in each frame of the images; A second determination unit, configured to determine the target detection box corresponding to the identifier of each detection box according to the size and / or position of the detection boxes in the detection box information corresponding to the same detection box identifier; A third determination unit, configured to determine the frame image where the target detection box corresponding to the identifier of each detection box is located as the target frame image corresponding to the identifier of each detection box.

8. The apparatus according to claim 7, wherein, the second determination unit is configured to: Compare the sizes of the detection boxes to determine the detection box with the largest size among the detection boxes as the target detection box; Or, determine the detection box whose position is within a preset area of the image as the target detection box; Or, determine the detection box whose size is greater than a threshold and whose position is within a preset area of the image as the target detection box.

9. The apparatus according to claim 6, wherein, the acquisition module is configured to: Determine the scene to which the video stream belongs; Determine the frame extraction mode corresponding to the video stream according to the belonging scene; Acquire the multi-frame images from the video stream based on the frame extraction mode.

10. The apparatus according to claim 6, wherein, the detection module is configured to: Acquire the position information of the area of interest in the scene to which the video stream belongs; Determine the target area in each frame of the images according to the position information of the area of interest; Perform tracking detection on the target area in each frame of the images to determine the detection box information included in each frame of the images.

11. An electronic device, including: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-5.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are for causing the computer to execute the method according to any one of claims 1-5.

13. A computer program product comprising a computer program which, when executed by a processor, implements the steps of the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Video frame information labeling method and device, equipment and storage medium

    CN110503074A

  • Video semi-automatic target labeling method integrating target detection and tracking

    CN110929560A