Image annotation method, device and electronic equipment
By extracting the foreground image in the depth image and determining its minimum external rectangle, combined with the visualization effect of the pseudo-color image, the problem of low depth image labeling efficiency is solved, efficient and accurate data labeling is achieved, and the development of depth image research and 3D technology is promoted.
Patent Information
- Application Number
- CN202111401632.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-24
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-11-24
AI Technical Summary
In the prior art, manual data labeling is inefficient, time-consuming and labor-consuming, and most methods are suitable for color images, making it difficult to efficiently label depth images.
By obtaining the background image and the depth image, extracting the foreground image and determining its minimum external rectangle, assigning label information and generating an annotation file, converting the depth image into a pseudo-color image for re-checking and correcting the annotation data.
The demand for manual labeling of candidate boxes is reduced, the labeling efficiency and accuracy of the data set is greatly improved, and the depth image data set with higher confidence is obtained, which supports the rapid development of deep image research and the development of 3D-related technologies.
Smart Images

Figure CN114119695B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image detection technology, and in particular to an image annotation method, device and electronic equipment. Background Art
[0002] In recent years, deep learning technology has been increasingly used in the field of image object detection due to its powerful feature learning capabilities. Preparing training data is one of the prerequisites for deep learning.
[0003] At present, the preparation of training data mostly relies on manual data annotation. Annotators need to perform a lot of repetitive judgments and operations to complete the data annotation of images. Data annotation is a very tedious and time-consuming task, which requires a lot of manpower and time costs. Therefore, a more efficient annotation solution is urgently needed. Summary of the invention
[0004] In view of this, the embodiments of the present application provide an image annotation method, device, and electronic device, which can solve one or more technical problems in the related art.
[0005] In a first aspect, an embodiment of the present application provides an image annotation method, comprising:
[0006] Get the background image and the depth image containing the foreground;
[0007] Acquire a foreground image using the depth image and the background image, and extract a foreground contour of the foreground image to determine a minimum circumscribed rectangle of the foreground contour;
[0008] Assigning label information to the foreground image, and generating a label file corresponding to the depth image according to the label information and the coordinate information of the minimum bounding rectangle, wherein the label information refers to category information of the foreground;
[0009] The depth image is converted into a pseudo-color image, and the annotation file is corrected based on the pseudo-color image to obtain a corrected annotation data set.
[0010] In this embodiment, on the one hand, the foreground image is obtained by using the depth image and the background image, and then the minimum enclosing rectangle of the foreground and its coordinate information are determined, thereby reducing the manual annotation of candidate boxes and greatly improving the annotation efficiency of the data set. On the other hand, converting the depth image into a pseudo-color image facilitates users to recheck the annotation file, improves the recheck efficiency, and also improves the accuracy of the annotation, and obtains a data set with higher confidence. On the other hand, the annotation file can be applied to the depth image, realizing a method for quickly annotating the depth image, facilitating annotation on the depth image and then training and learning, and can quickly carry out research on the depth image and promote the development of 3D related technologies.
[0011] As an implementation of the first aspect, converting the depth image into a pseudo-color image includes:
[0012] Acquire a chromaticity diagram, wherein the chromaticity diagram includes a mapping relationship between color values and pixel values;
[0013] Normalizing the depth image to obtain a normalized image corresponding to the depth image;
[0014] Each of the normalized images is mapped into a pseudo-color image according to the chromaticity diagram.
[0015] As an implementation of the first aspect, extracting a foreground contour of the foreground image to determine a minimum circumscribed rectangle of the foreground contour includes:
[0016] Performing morphological operations and binarization processing on the foreground image to obtain a binarized image;
[0017] A foreground contour in the binary image is extracted, and a minimum bounding rectangle of the foreground contour is determined.
[0018] As an implementation manner of the first aspect, the annotation file also includes image information of the depth image, and the image information includes the length, width, channel, path and image name of the image.
[0019] As an implementation of the first aspect, generating a label file corresponding to the depth image according to the label information of the foreground and the coordinate information of the minimum bounding rectangle includes:
[0020] The label information corresponding to the foreground of the depth image, the image information of the depth image, and the coordinate information of the minimum bounding rectangle are written into a labeling file in a preset format to obtain a labeling file corresponding to the depth image.
[0021] As an implementation of the first aspect, obtaining the corrected labeled data set includes:
[0022] The label information and the coordinate information included in the annotation file are checked one by one using the pseudo-color image to obtain a corrected data set.
[0023] As an implementation of the first aspect, using the pseudo-color image to perform one-to-one verification on the label information and the coordinate information included in the annotation file includes:
[0024] The annotation file is corrected using an annotation tool to obtain a corrected annotation data set; wherein the annotation tool corrects the annotation file by determining whether the label information and coordinate information included in the annotation file displayed by the pseudo-color image are correct.
[0025] In a second aspect, an embodiment of the present application provides an image annotation device, including:
[0026] An acquisition module, used for acquiring a background image and a depth image including a foreground;
[0027] An extraction module, configured to obtain a foreground image using the depth image and the background image, and extract a foreground contour of the foreground image to determine a minimum circumscribed rectangle of the foreground contour;
[0028] A file generation module, used to give label information to the foreground image, and generate a label file corresponding to the depth image according to the label information and the coordinate information of the minimum bounding rectangle, wherein the label information is the category information of the foreground;
[0029] A conversion module, used for converting the depth image into a pseudo-color image;
[0030] The annotation module is used to correct the annotation file based on the pseudo-color image to obtain a corrected annotation data set.
[0031] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the image annotation method as described in the first aspect or any implementation manner of the first aspect are implemented.
[0032] In a fourth aspect, an embodiment of the present application provides a computer storage medium, wherein the computer storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the image annotation method as described in any implementation method of the first aspect or the first direction are implemented.
[0033] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device can implement the steps of the image annotation method described in the first aspect or any implementation of the first aspect.
[0034] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0036] Figure 1 It is a structural schematic diagram of an electronic device provided by an embodiment of the present application;
[0037] Figure 2 This is a schematic diagram of an implementation flow of an image annotation method provided by an embodiment of the present application;
[0038] Figure 3 This is a schematic diagram of a specific implementation process of step S160 in an image annotation method provided in an embodiment of the present application;
[0039] Figure 4 is a structural schematic diagram of an image annotation device provided by an embodiment of the present application;
[0040] Figure 5 It is a structural schematic diagram of another image annotation device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0041] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0042] The term "and / or" as used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0043] "One embodiment" or "some embodiments" described in the specification of this application means that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the sentences "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0044] Furthermore, in the description of the present application, “plurality” means two or more.
[0045] It should also be understood that, unless otherwise clearly specified or limited, the term "connection" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a direct connection, or an indirect connection through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0046] The current manual data annotation method is time-consuming, labor-intensive, costly, and inefficient. In addition, most data annotation methods on the market are suitable for color images, and rarely consider data annotation for depth images.
[0047] Therefore, an embodiment of the present application provides an image annotation method, which can realize rapid annotation of images, and further realize rapid annotation of depth images to obtain depth images with annotation information, thereby facilitating research on depth images and promoting the development of three-dimensional (3D) related technologies.
[0048] In order to illustrate the technical solution described in this application, a specific embodiment is provided below for illustration.
[0049] Figure 1 The schematic diagram of the structure of an electronic device provided in one embodiment of the present application. The electronic device includes but is not limited to computers, tablets, laptops, netbooks, servers and other electronic devices. The present embodiment of the application does not impose any restrictions on the specific type of electronic device.
[0050] In some embodiments of the present application, the electronic device may include one or more processors 10 ( Figure 1 Only one is shown in the figure), a memory 11 and a computer program 12 stored in the memory 11 and executable on one or more processors 10, for example, a program for image annotation. When one or more processors 10 execute the computer program 12, the various steps in the image annotation method embodiment described later can be implemented. Alternatively, when one or more processors 10 execute the computer program 12, the functions of various modules / units in various image annotation device embodiments described later can be implemented.
[0051] Exemplarily, the computer program 12 may be divided into one or more modules / units, one or more modules / units are stored in the memory 11 and executed by the processor 10 to complete the present application. One or more modules / units may be a series of computer program instruction segments that can complete specific functions, and the instruction segments are used to describe the execution process of the computer program 12 in the processing unit. For example, the computer program 12 may be divided into the following modules. The specific functions of each module are as follows:
[0052] An acquisition module, used for acquiring a background image and a depth image including a foreground;
[0053] An extraction module, configured to obtain a foreground image using the depth image and the background image, and extract a foreground contour of the foreground image to determine a minimum circumscribed rectangle of the foreground contour;
[0054] A file generation module, used to give label information to the foreground image, and generate a label file corresponding to the depth image according to the label information and the coordinate information of the minimum bounding rectangle, wherein the label information is the category information of the foreground;
[0055] A conversion module, used for converting the depth image into a pseudo-color image;
[0056] The annotation module is used to correct the annotation file based on the pseudo-color image to obtain a corrected annotation data set.
[0057] Those skilled in the art will understand that Figure 1 The electronic device is merely an example of an electronic device and does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.
[0058] The processor 10 may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0059] The memory 11 may be an internal storage unit of the processing unit, such as a hard disk or memory of the processing unit. The memory 11 may also be an external storage device of the processing unit, such as a plug-in hard disk equipped on the processing unit, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 11 may also include both an internal storage unit of the processing unit and an external storage device. The memory 11 is used to store computer programs and other programs and data required by the processing unit. The memory 11 may also be used to temporarily store data that has been output or is to be output.
[0060] An embodiment of the present application further provides another preferred embodiment of an electronic device. In this embodiment, the electronic device includes one or more processors, and the one or more processors are used to execute the following program modules stored in the memory:
[0061] An acquisition module, used for acquiring a background image and a depth image including a foreground;
[0062] An extraction module, configured to obtain a foreground image using the depth image and the background image, and extract a foreground contour of the foreground image to determine a minimum circumscribed rectangle of the foreground contour;
[0063] A file generation module, used to give label information to the foreground image, and generate a label file corresponding to the depth image according to the label information and the coordinate information of the minimum bounding rectangle, wherein the label information is the category information of the foreground;
[0064] A conversion module, used for converting the depth image into a pseudo-color image;
[0065] The annotation module is used to correct the annotation file based on the pseudo-color image to obtain a corrected annotation data set.
[0066] Figure 2 1 is a schematic diagram of an implementation flow of an image annotation method provided in an embodiment of the present application. The image annotation method in this embodiment is applicable to situations where depth images need to be annotated. The image annotation method in this embodiment can be executed by an electronic device. As an example and not a limitation, the image annotation method can be applied to Figure 1 Electronic equipment shown.
[0067] like Figure 2 As shown, the image annotation method may include: step S110 to step S160.
[0068] S110, acquiring a background image and a depth image including a foreground.
[0069] Specifically, the application scenario usually includes a foreground target (or foreground, or target) and a background.
[0070] For the same angle in the same scene, the acquisition module can be used to capture at least one depth image and at least one background image. At least one depth image contains the foreground object and the background, that is, a depth panoramic image, and at least one background image only contains the background RGB image or the depth image. By changing the captured scene and / or camera angle, a large number of images can be obtained.
[0071] As one implementation method, first in a certain scene, multiple frames of background images and multiple frames of depth images are respectively shot at different acquisition module angles, and then change the shooting scene, and shoot multiple frames of background images and multiple frames of depth images at different acquisition module angles. Multiple scenes can be changed.
[0072] As another implementation method, first shoot multiple frames of background images and multiple frames of depth images at a certain scene and a certain acquisition module angle, then change the shooting scene, and shoot multiple frames of background images and multiple frames of depth images at the same acquisition module angle. Multiple scenes can be changed.
[0073] In some embodiments, multiple frames of background images and multiple frames of depth images captured at a certain acquisition module angle in a certain scene can form a background image sequence frame and a depth image sequence frame. The background image can be placed before the depth image to facilitate subsequent background modeling and improve efficiency.
[0074] S120: Acquire a foreground image corresponding to the depth image based on the background image and the depth image.
[0075] More specifically, step S120 performs target foreground extraction, preferably extracting the target foreground in the depth image based on a background modeling method.
[0076] In some embodiments, the background is first modeled based on the depth image and the background image, and then the target foreground in the sequence frames is detected using the background subtraction method. The present application embodiment does not specifically limit the method of background modeling.
[0077] As a non-limiting example, extracting target foreground based on background modeling may include the following steps:
[0078] 1) Background modeling: The process of background modeling is the process of learning a background image sequence frame. In the training stage, by learning a background image sequence frame to extract the background features in this sequence frame, a mathematical model is established to describe the background and form a background model.
[0079] 2) Detection stage: Subtract the detection image (i.e., depth image) from the background model to obtain the foreground target. Specifically, the image to be detected, i.e., the depth panoramic image, is processed using the background model. The background subtraction method is generally used to extract pixels with different properties from the detection image and the background model. The image composed of these pixels is the foreground target, i.e., the foreground image.
[0080] In some embodiments, if the acquired foreground image is a foreground grayscale image, the subsequent step S130 can be directly entered, or the grayscale image can be converted into a binary image before entering the subsequent step S130. In some other embodiments, if the acquired foreground image is a foreground binary image, the subsequent step S130 can be directly entered.
[0081] It should be understood that other methods may be used to extract the foreground image based on the depth image and the background image, and the present application embodiment does not specifically limit this. For example, the foreground image of any depth image may be extracted by subtracting the background image.
[0082] S130, performing morphological operations on the foreground image to obtain a binary image.
[0083] In some embodiments, morphological operations can be used to remove noise and reduce background interference to improve the accuracy of subsequent annotation results, thereby providing training data with higher confidence. Among them, morphological operations include but are not limited to dilation, erosion, opening or closing operations, etc., and this application does not limit the method of morphological operations.
[0084] As a non-limiting example, based on the foreground soft segmentation image obtained in step S120, that is, based on the foreground binary image, a 3*3 structure matrix is selected, and the elements in the structure matrix are all 1. With a step size of 1, each pixel in the foreground soft segmentation image is scanned, and a logical AND operation is performed between the structure matrix and the foreground soft segmentation image. If the values of the structure matrix and the foreground soft segmentation image are both 1, the pixel at that point in the output image is 1, and in other cases the pixel of the output image is 0. This process is called an erosion process, which can reduce the foreground soft segmentation image by one circle.
[0085] As another non-limiting example, based on the foreground soft segmentation image obtained in step S61, a 3*3 structure matrix is selected, and the elements in the structure matrix are all 1. The structure matrix and the foreground soft segmentation image are used to perform a logical AND operation. If the values of the structure matrix and the foreground soft segmentation image are both 0, the pixel at that point in the output image is 0. In other cases, the pixel of the output image is 1. This process is called an expansion process, which can expand the foreground soft segmentation image by one circle.
[0086] It should be understood that the above-mentioned process of first corrosion and then expansion is called an opening operation, which can be used to eliminate noise points, separate objects at fine points, and smooth the boundaries of larger objects without significantly changing their areas.
[0087] S140, extracting a foreground contour in the binary image, and determining a minimum circumscribed rectangle of the foreground contour.
[0088] In some embodiments, edge detection is performed on the binary image to extract the foreground contour. In other embodiments, the binary image is first smoothed and filtered to eliminate some noise, and then edge detection is performed to extract the foreground contour with higher accuracy.
[0089] In some implementations, the image smoothing filtering method includes but is not limited to: interpolation method, linear smoothing method or convolution method, etc. Different smoothing filtering methods can be selected according to the actual image noise. For example, for salt and pepper noise, a linear smoothing method is used. It should be noted that the embodiment of the present application does not limit the smoothing filtering method.
[0090] In some implementations, edge detection may be performed based on an edge detection operator, including but not limited to a Sobel operator, a Roberts operator, a Prewitt operator, a Canny operator, a Lagrange operator, etc. It should be noted that the embodiment of the present application does not limit the edge detection method.
[0091] After extracting the foreground contour, the minimum bounding rectangle of the foreground contour is determined. The minimum bounding rectangle may include a minimum area bounding rectangle and a minimum perimeter bounding rectangle.
[0092] In some embodiments, the minimum bounding rectangle of the foreground contour may be determined by a direct calculation method, an equal-interval rotation search method, or an improved method thereof. The present application does not specifically limit the method for determining the minimum bounding rectangle.
[0093] Furthermore, all the contours found can be conditionally judged first, interfering contours that do not meet the conditions can be removed, and then the minimum circumscribed rectangle of the contour can be found. The condition is a pre-set judgment condition based on one or more parameters such as the perimeter of the contour, the area of the contour, the centroid of the contour, the aspect ratio of the contour, etc. If the parameters meet the pre-set judgment conditions, they are retained, and if they do not meet the pre-set judgment conditions, they are removed.
[0094] It should be understood that in some other embodiments, the circumscribed rectangle may also be other shapes. This is only an exemplary description and cannot be interpreted as a specific limitation of the present application.
[0095] S150 , assigning label information to the foreground image, and generating a label file corresponding to each depth image based on the label information and the coordinate information of the minimum bounding rectangle.
[0096] The label information refers to the category information of the foreground or target, and can be input by the annotator.
[0097] The coordinate information of the minimum bounding rectangle refers to the coordinate information of the candidate box of the foreground or target. The pixel coordinates corresponding to any pixel point on the minimum bounding rectangle in the foreground image can be selected as the coordinate information. For example, any pixel point on the minimum bounding rectangle in the foreground image can select the pixel point of the upper left corner, upper right corner, lower left corner or lower right corner of the rectangle. This is only an exemplary description and cannot be interpreted as a specific limitation of the present application.
[0098] In order to facilitate the description of the scheme of the embodiment of the present application, in the embodiment of the present application, the human body is used as an example of a foreground target, and the preparation of training data for human body detection is used as an example for description. It should be understood that the exemplary description cannot be interpreted as a specific limitation of the present application. In other embodiments, the foreground target may also include objects and / or animals, etc.
[0099] As a non-limiting example, training data for human detection is prepared, the human detection model is a two-classification model, and the label information may include human (or person) or non-human. It should be understood that if there is a human target in a certain depth image, the label information corresponding to the human target is human, and the coordinate information of the minimum circumscribed rectangle corresponding to the human target is specific coordinate information; if there is no human target in a certain depth image but there is a non-human target, the label information and coordinate information in the depth image are empty or do not exist.
[0100] As another non-limiting example, training data for human detection is prepared, and the human detection model is a model with more than three categories, and the label information may include human body, animal, object, etc. It should be understood that if there is a human target in a certain depth image, the label information corresponding to the human target is human body, and the coordinate information of the minimum circumscribed rectangle corresponding to the human target is specific coordinate information; if there is no human target in a certain depth image but there is an animal or object, etc., the label information and coordinate information in the depth image are empty or do not exist. It should be noted that in models with more than three categories, human body or non-human body can be further subdivided, and the embodiments of the present application are not limited to this.
[0101] In some embodiments, the label information corresponding to the depth image and the coordinate information of the minimum circumscribed rectangle in the foreground image corresponding to the depth image are written into a preset format annotation file to obtain an annotation file. It should be noted that the preset format annotation file may include an xml annotation file of VOC, a txt annotation file of yolo, or a json annotation file of coco. The embodiment of the present application does not specifically limit the format of the annotation file.
[0102] In some other embodiments, the label information and image information corresponding to the depth image and the coordinate information of the minimum bounding rectangle in the foreground image corresponding to the depth image are written into a preset format annotation file to obtain the annotation file.
[0103] More specifically, when acquiring the depth image captured by the camera, each depth image may have its own image information. The image information includes, but is not limited to, one or more combinations of the length, width, channel, path, and image name of the image. Therefore, in some embodiments, the image information is also written into the annotation file.
[0104] S160, converting the depth image into a pseudo-color image.
[0105] For depth images, it is difficult for the human eye to identify the target and perceive the change in depth. Therefore, each depth image can be converted into a pseudo-color image to achieve a better visualization effect and facilitate subsequent manual re-inspection.
[0106] In some embodiments, Figure 3 As shown, converting the depth image into a pseudo-color image includes steps S161 to S163.
[0107] S161, obtaining a chromaticity map (colormap), the chromaticity map including a mapping relationship between color values and pixel values.
[0108] In the embodiment of the present application, there are many kinds of chromaticity diagrams, and a chromaticity diagram suitable for the application scenario can be selected. In the embodiment of the present application, the chromaticity diagram is selected to be a chromaticity diagram that can largely distinguish the human body from the background.
[0109] In some implementations, a chromaticity map suitable for a human body detection scene is selected, and the chromaticity map may be selected as colormap_jet.
[0110] In some implementations, the color value may be an RGB value.
[0111] S162, normalizing the depth image to obtain a normalized image corresponding to the depth image;
[0112] Normalize each pixel value in the depth image to a range of 0 to 255, which may be a closed interval [0, 8000]. That is, the original range of each pixel value in the depth image is 0 to 8000, and through the normalization operation, these pixel value ranges are normalized to a range of 0 to 255.
[0113] As a non-limiting example, if the maximum pixel value in a depth image is a, and the depth value of a pixel point j in the depth image is z, then the pixel point j is normalized to between [0,225], then:
[0114]
[0115] Among them, G(j) represents the normalized value of pixel j.
[0116] S163, mapping the normalized image into a pseudo-color image according to the chromaticity diagram.
[0117] Since the chromaticity diagram includes the correspondence between color values and pixel values, the color value corresponding to the pixel value of each pixel in each normalized image (i.e., the value after normalizing the depth value) can be determined according to the chromaticity diagram, thereby mapping the normalized image into a pseudo-color image.
[0118] S170, using the pseudo-color image of the depth image, correct the annotation file corresponding to each depth image to obtain a corrected annotation data set.
[0119] In some embodiments, the annotation file corresponding to the depth image is first matched one-to-one with the pseudo-color image of the depth image to obtain the aligned annotation file and the pseudo-color image, and then the label information and the coordinate information of the minimum circumscribed matrix included in the annotation file corresponding to each depth image are corrected based on the pseudo-color image of the depth image to obtain a corrected annotation data set.
[0120] In some embodiments, step S170 also includes: using a labeling tool to correct the labeling data set to obtain a corrected labeling data set; wherein the labeling tool can be used to simultaneously display the labeling image and the labeling file, so as to further correct the labeling file by determining whether the label information and coordinate information included in the labeling file displayed by the pseudo-color image are correct.
[0121] After the automatic annotation of the depth image is completed by finding the minimum enclosing rectangle of the foreground contour and assigning label information, there may be situations where the position of the candidate box (i.e. the coordinate information of the target or foreground) is not accurate, the label is missed, or the label is wrong. Therefore, manual re-inspection is required to correct the annotation box to improve the accuracy of the annotation box. Since the visualization effect of the depth image is poor, manual re-inspection using pseudo-color images corresponding to each depth image with good visualization effect can allow users to perceive the target quickly and friendly, facilitate users to perform manual re-inspection, improve the efficiency and accuracy of re-inspection, and obtain a depth image dataset with high annotation accuracy.
[0122] As a non-limiting example, the annotation tool displays each pair of one-to-one corresponding annotation files and annotated pseudo-color images through a display screen, that is, the label information and the minimum circumscribed rectangle in the annotation file are displayed on the pseudo-color image, which is convenient for the user (i.e., the annotation personnel) to conduct a re-examination. The user can check the label information and target coordinate information by comparing the pseudo-color image and the annotation file. When an error is found, the user can input manual re-examination data through an external device such as a microphone, a mouse, a keyboard, etc. The annotation tool corrects the annotation file according to the manual re-examination data input by the user, thereby obtaining a corrected annotation file, and obtaining a corrected annotation data set, that is, a labeled depth image set. Further, after manual re-examination, an accurate annotated data set will be obtained, and the annotated depth image can be converted into an annotated pseudo-color image to obtain an annotation file corresponding to both the depth image and the pseudo-color image, forming an annotated depth image data set. For example, the annotated depth image data set is used as training data to obtain a human body detection model based on a depth image.
[0123] It should be noted that if the user manually rechecks and finds that the information in the annotation file does not need to be modified, the original annotation file information will be retained.
[0124] In some embodiments, one or more pairs of one-to-one corresponding annotation files and pseudo-color images may be displayed on the same screen.
[0125] As a non-limiting example, the labeling tool is labelimg software, which performs batch labeling and correction on pseudo-color images in which each frame corresponds to the labeling file.
[0126] Specifically, after the user opens the annotation tool labelimg software, what needs to be checked mainly are the label information in the annotation file and the coordinate information of the minimum enclosing rectangle. For image information, labelimg software will automatically read the original image information. If there is a discrepancy with the image information in the annotation file, the image information will be automatically modified and written into the annotation file.
[0127] In the embodiments of the present application, on the one hand, the foreground image is extracted based on the depth image and the background image, and then the minimum enclosing rectangle of the foreground and its coordinate information are determined, thereby reducing the number of manually labeled candidate boxes and greatly improving the efficiency of data set annotation. On the other hand, converting the depth image into a pseudo-color image facilitates users to recheck the annotation file, improves the efficiency of rechecking, and also improves the accuracy of annotation, and obtains a data set with higher confidence. On the other hand, the annotation file can be applied to the depth image, realizing a method for quickly annotating the depth image, and then training and learning can be carried out, so that research on depth images can be quickly carried out and the development of 3D related technologies can be promoted.
[0128] It should be noted that Figure 2The step numbers of the illustrated embodiment should not be interpreted as limiting the time sequence of the steps. It should be understood that in some other embodiments, the order of the steps can be swapped according to the logical relationship between the steps without affecting the implementation of the present solution. As a non-limiting example, step S160 is performed at any time after step S110 and before step S170.
[0129] Corresponding to the above image annotation method, an embodiment of the present application further provides an image annotation device. For details not described in detail in the image annotation device, please refer to the description of the above method.
[0130] Figure 4 1 is a schematic diagram of the structure of an image annotation device provided in an embodiment of the present application. The image annotation device comprises: an acquisition module 51 , an extraction module 52 , a file generation module 53 , a conversion module 54 and an annotation module 55 .
[0131] The acquisition module 51 is used to acquire a background image and a depth image including a foreground;
[0132] An extraction module 52, configured to obtain a foreground image using the depth image and the background image, and extract a foreground contour of the foreground image to determine a minimum circumscribed rectangle of the foreground contour;
[0133] A file generating module 53, configured to assign label information to the foreground image, and generate a labeling file corresponding to each of the depth images according to the label information and the coordinate information of the minimum bounding rectangle, wherein the label information is the category information of the foreground;
[0134] A conversion module 54, configured to convert each of the depth images into a pseudo-color image;
[0135] The annotation module 55 is used to correct the annotation file based on the pseudo-color image to obtain a corrected annotation data set.
[0136] In one embodiment, the acquisition module 51 is used to acquire images of multiple scenes, the images including a depth image including a foreground and a background image not including a foreground, wherein the background image may be a depth image or a color image. It should be noted that the acquisition module 51 includes a depth camera, which may be, but is not limited to, a camera based on a light time-of-flight method (time-of-flight, TOF) such as indirect time-of-flight (iToF) or direct time-of-flight (dToF), based on binocular vision, or based on structured light; in another embodiment, the acquisition module 51 may also be a color camera for acquiring a color image (i.e., a background image) containing only the background, which is not limited here.
[0137] In order to collect a sufficient amount of image data for deep learning, the shooting scene can be changed multiple times, and for each scene, multiple frames of depth images and background images can be collected, so that training data in multiple different scenes can be obtained.
[0138] As a non-limiting example, after the acquisition module 51 is started, the acquisition module 51 is used to capture images, that is, to capture depth images and background images. When capturing images, for any camera angle in each scene, a small number of background images, such as 50 frames of background images, can be captured first, and then multiple frames of depth images can be captured. It should be noted that the background image refers to an image that does not contain a human body but contains a background in the depth image, and the depth image refers to an image that contains a human body and a background; when capturing images, multiple frames of depth images and background images can be saved as sequence frames respectively for subsequent calculation and processing, which is not limited here.
[0139] In some embodiments, the conversion module 54 is specifically configured to:
[0140] Obtain a chromaticity diagram, the chromaticity diagram including a mapping relationship between color values and pixel values;
[0141] Normalizing each of the depth images to obtain a normalized image corresponding to each of the depth images;
[0142] Each of the normalized images is mapped into a pseudo-color image according to a chromaticity diagram.
[0143] In one implementation, acquiring a chromaticity diagram includes: acquiring a chromaticity diagram of a human body detection scene.
[0144] In some embodiments, the extraction module 52 is specifically used to:
[0145] Performing morphological operations on the foreground image to obtain a binary image;
[0146] A foreground contour in the binary image is extracted, and a minimum bounding rectangle of the foreground contour is determined.
[0147] In some embodiments, Figure 4 Based on the embodiment shown, Figure 5 FIG. 1 is an image annotation device provided by an embodiment of the present application. Figure 5 As shown, the image annotation device further includes: a corresponding module 56. It should be understood that other details not described in detail can be found in Figure 4 The embodiment shown.
[0148] The corresponding module 56 is used to make a one-to-one correspondence between the annotation file corresponding to the depth image and the pseudo-color image of the depth image, and obtain the aligned annotation file and the pseudo-color image.
[0149] In some embodiments, the file generation module 53 is specifically used to:
[0150] The label information corresponding to the foreground of each of the depth images and the coordinate information of the minimum bounding rectangle are written into a labeling file in a preset format to obtain a labeling file corresponding to the depth image.
[0151] In some embodiments, the file generation module 53 is specifically used to:
[0152] The label information corresponding to the foreground of each of the depth images, the image information of the depth image, and the coordinate information of the minimum bounding rectangle are written into a labeling file in a preset format to obtain a labeling file corresponding to each of the depth images.
[0153] In some embodiments, further acquiring the annotated data set includes: converting the annotated depth image into an annotated pseudo-color image to obtain an annotated depth image data set.
[0154] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0155] An embodiment of the present application also provides an electronic device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor implements the steps in any of the above-mentioned image annotation method embodiments when executing the computer program.
[0156] The embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various image annotation method embodiments can be implemented.
[0157] An embodiment of the present application provides a computer program product. When the computer program product is executed on an electronic device, the electronic device can implement the steps in each of the above-mentioned image annotation method embodiments.
[0158] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0159] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0160] In the embodiments provided in the present application, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0161] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0162] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0163] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electric carrier signals and telecommunication signals.
[0164] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. An image annotation method, characterized in that: include: Get the background image and the depth image containing the foreground; Acquire a foreground image using the depth image and the background image, and extract a foreground contour of the foreground image to determine a minimum circumscribed rectangle of the foreground contour; Assigning label information to the foreground image, and generating a label file corresponding to the depth image according to the label information and the coordinate information of the minimum bounding rectangle, wherein the label information refers to category information of the foreground; Converting the depth image into a pseudo-color image, and correcting the annotation file based on the pseudo-color image to obtain a corrected annotation data set, including: using the pseudo-color image to check the label information and the coordinate information included in the annotation file one by one to obtain a corrected data set; The method of using the pseudo-color image to proofread the label information and the coordinate information included in the annotation file one by one includes: using an annotation tool to correct the annotation file to obtain a corrected annotation data set; wherein the annotation tool corrects the annotation file by determining whether the label information and coordinate information included in the annotation file displayed by the pseudo-color image are correct.
2. The image annotation method according to claim 1, characterized in that: The converting the depth image into a pseudo-color image comprises: Acquire a chromaticity diagram, wherein the chromaticity diagram includes a mapping relationship between color values and pixel values; Normalizing the depth image to obtain a normalized image corresponding to the depth image; Each of the normalized images is mapped into a pseudo-color image according to the chromaticity diagram.
3. The image annotation method according to claim 1, characterized in that: The step of extracting a foreground contour of the foreground image to determine a minimum circumscribed rectangle of the foreground contour comprises: Performing morphological operations and binarization processing on the foreground image to obtain a binarized image; A foreground contour in the binary image is extracted, and a minimum bounding rectangle of the foreground contour is determined.
4. The image annotation method according to any one of claims 1 to 3, characterized in that: The generating a labeling file corresponding to the depth image according to the label information and the coordinate information of the minimum bounding rectangle includes: The label information corresponding to the foreground of the depth image, the image information of the depth image, and the coordinate information of the minimum bounding rectangle are written into a labeling file in a preset format to obtain a labeling file corresponding to the depth image.
5. The image annotation method according to claim 4, characterized in that: The annotation file also includes image information of the depth image, and the image information includes the length, width, channel, path and image name of the image.
6. An image annotation device, characterized in that: include: An acquisition module, used for acquiring a background image and a depth image including a foreground; An extraction module, configured to obtain a foreground image using the depth image and the background image, and extract a foreground contour of the foreground image to determine a minimum circumscribed rectangle of the foreground contour; A file generation module, used to give label information to the foreground image, and generate a label file corresponding to the depth image according to the label information and the coordinate information of the minimum bounding rectangle, wherein the label information is the category information of the foreground; A conversion module, used for converting the depth image into a pseudo-color image; A labeling module, used to correct the labeling file based on the pseudo-color image to obtain a corrected labeling data set, including: using the pseudo-color image to check the label information and the coordinate information included in the labeling file one by one to obtain the corrected data set; The method of using the pseudo-color image to proofread the label information and the coordinate information included in the annotation file one by one includes: using an annotation tool to correct the annotation file to obtain a corrected annotation data set; wherein the annotation tool corrects the annotation file by determining whether the label information and coordinate information included in the annotation file displayed by the pseudo-color image are correct.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the image annotation method according to any one of claims 1 to 5 is implemented.
8. A computer storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the image annotation method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Method for automatically generating image training data, image processing equipment and storage device
CN112861899A
An image data target detection inclination frame automatic labeling method and system
CN113313751A