Automatic object annotation methods, apparatus, electronic devices and storage media

By identifying landmarks in images and using their correspondence with 3D models for contour matching, objects are automatically labeled, solving the problem of high-cost labeling in deep learning model training and achieving efficient image labeling and 3D model acquisition.

CN116266402BActive Publication Date: 2026-06-02HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
Filing Date
2021-12-10
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In existing technologies, the cost of labeling massive amounts of sample images during the training of deep learning models is high. How to reduce the cost of manual labeling of sample images has become an urgent technical problem to be solved.

Method used

By acquiring images containing markers and objects to be labeled, marker recognition technology is used to determine the target marker information. Based on the pre-set correspondence between the marker information and the object's 3D model, contour matching is performed to determine the location information of the object to be labeled, thereby automatically labeling the object.

Benefits of technology

It enables automatic object annotation, reduces image annotation costs, improves annotation efficiency, and can automatically obtain the 3D model of the object to be annotated, reducing manual workload.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116266402B_ABST
    Figure CN116266402B_ABST
Patent Text Reader

Abstract

This application provides an automatic object annotation method, apparatus, electronic device, and storage medium. The method involves acquiring a first image to be annotated, containing markers and objects to be annotated; identifying markers in the first image to obtain target marker information; determining the target 3D model corresponding to the target marker information according to a pre-set correspondence between the marker information and the object's 3D model; performing contour matching on the first image to be annotated based on the target 3D model to determine the position information of the objects to be annotated in the first image; and annotating the objects to be annotated in the first image according to the position information of the objects to be annotated. This achieves automatic object annotation, reducing image annotation costs and increasing annotation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to methods, apparatus, electronic devices and storage media for automatic object annotation. Background Technology

[0002] With the development of artificial intelligence technology, computer vision technology, especially deep learning-based computer vision technology, has developed rapidly. In computer vision technology, deep learning models need to be trained on a large amount of labeled data. For example, when using computer vision technology to identify vehicles, a large number of sample images labeled with vehicles are selected to train the deep learning model.

[0003] In the process of training deep learning models, the cost of labeling massive amounts of sample images has become the biggest cost in the training process. How to reduce the cost of manual labeling of sample images has become a technical problem that urgently needs to be solved. Summary of the Invention

[0004] The purpose of this application is to provide an automatic object annotation method, apparatus, electronic device, and storage medium to achieve automatic object annotation in images, thereby reducing image annotation costs. The specific technical solution is as follows:

[0005] Firstly, this application provides an automatic object annotation method, the method comprising:

[0006] Obtain the first image to be labeled, which contains the markers and the objects to be labeled;

[0007] The first image to be labeled is subjected to marker recognition to obtain target marker information;

[0008] Based on the pre-set correspondence between marker information and the object's three-dimensional model, determine the target three-dimensional model corresponding to the target marker information;

[0009] Based on the target 3D model, contour matching is performed on the first image to be labeled to determine the position information of the object to be labeled in the first image to be labeled.

[0010] According to the position information of the object to be labeled in the first image to be labeled, the object to be labeled is labeled in the first image to be labeled.

[0011] In one possible implementation, the marker is a QR code, the target marker information is target QR code information, and the correspondence is the correspondence between the QR code information and the three-dimensional model of the object.

[0012] The step of identifying target markers in the first image to be labeled to obtain target marker information includes:

[0013] The first image to be labeled is identified using QR code recognition technology to obtain the target QR code information in the first image to be labeled.

[0014] In one possible implementation, the method further includes:

[0015] Acquire multiple sample images containing the marker and the object to be labeled, acquired by an image acquisition device, wherein the marker is set at multiple key points of the object to be labeled in the multiple sample images;

[0016] Determine the position of the marker in each of the sample images;

[0017] Obtain the pose information of the image acquisition device when acquiring each of the sample images;

[0018] For each sample image, based on the pose information of the image acquisition device when acquiring the sample image and the position of the marker in the sample image, the position of the marker in the world coordinate system corresponding to the sample image is determined;

[0019] Based on the position of the marker corresponding to each sample image in the world coordinate system, a three-dimensional model of the object to be labeled is established.

[0020] Obtain the marker information of the marker and establish the correspondence between the marker information of the marker and the three-dimensional model of the object to be labeled.

[0021] In one possible implementation, acquiring the pose information of the image acquisition device when acquiring each of the sample images includes:

[0022] Based on each of the sample images, the pose information of the image acquisition device when acquiring each of the sample images is determined using the Simultaneous Localization and Mapping (SLAM) algorithm.

[0023] In one possible implementation, determining the position of the marker in the world coordinate system for each sample image, based on the pose information of the image acquisition device when acquiring the sample image and the position of the marker in the sample image, includes:

[0024] For each sample image, based on the pose information of the image acquisition device when acquiring the sample image and the position of the marker in the sample image, the SLAM algorithm is used to determine the position of the marker in the world coordinate system corresponding to the sample image.

[0025] In one possible implementation, the method further includes:

[0026] The marker is placed at the key points of the object to be labeled, and the image acquisition device is used to acquire a sample image containing the marker and the object to be labeled;

[0027] Adjust the pose of the image acquisition device and / or the position of the marker at the object to be labeled, and use the image acquisition device to acquire a sample image containing the marker and the object to be labeled;

[0028] Repeat the above steps: adjust the position of the image acquisition device and / or the position of the marker at the object to be labeled, and use the image acquisition device to acquire sample images containing the marker and the object to be labeled until the acquisition termination condition is met.

[0029] In one possible implementation, the method further includes:

[0030] Based on the obtained position of the marker in the world coordinate system and the current pose information of the image acquisition device, the position of the key points of the object to be labeled in the image coordinate system of the image acquisition device is determined to obtain the key point image position;

[0031] Based on the obtained key point image positions, a rectangular bounding box is fitted;

[0032] The key point image positions and the rectangles are displayed on the screen corresponding to the image acquisition device.

[0033] In one possible implementation, the method further includes:

[0034] Obtain a second image to be labeled that contains the object to be labeled but does not contain any markers;

[0035] Based on the target 3D model, contour matching is performed on the second image to be labeled to determine the position information of the object to be labeled in the second image to be labeled.

[0036] According to the position information of the object to be labeled in the second image to be labeled, the object to be labeled is labeled in the second image to be labeled.

[0037] Secondly, embodiments of this application provide an automatic object labeling device, the device comprising:

[0038] The image acquisition module is used to acquire a first image containing markers and objects to be labeled.

[0039] The marker information recognition module is used to recognize markers in the first image to be labeled and obtain target marker information.

[0040] The 3D model determination module is used to determine the target 3D model corresponding to the target marker information according to the pre-set correspondence between marker information and the 3D model of the object;

[0041] The location information determination module is used to perform contour matching on the first image to be labeled based on the target 3D model to determine the location information of the object to be labeled in the first image to be labeled.

[0042] The object annotation module is used to annotate the object in the first image according to the position information of the object in the first image to be annotated.

[0043] In one possible implementation, the marker is a QR code, the target marker information is target QR code information, and the correspondence is the correspondence between the QR code information and the three-dimensional model of the object.

[0044] The marker information recognition module is specifically used to: use QR code recognition technology to perform QR code recognition on the first image to be labeled, and obtain the target QR code information in the first image to be labeled.

[0045] In one possible implementation, the device further includes:

[0046] The sample image acquisition module is used to acquire multiple sample images containing the marker and the object to be labeled, which are acquired by the image acquisition device, wherein the marker is set at multiple key points of the object to be labeled in the multiple sample images;

[0047] A marker location determination module is used to determine the location of the marker in each of the sample images;

[0048] The pose information acquisition module is used to acquire pose information when the image acquisition device acquires each of the sample images;

[0049] The world coordinate determination module is used to determine the position of the marker in the world coordinate system for each sample image based on the pose information of the image acquisition device when the sample image is acquired and the position of the marker in the sample image.

[0050] A 3D model building module is used to build a 3D model of the object to be labeled based on the position of the marker corresponding to each sample image in the world coordinate system.

[0051] The correspondence establishment module is used to obtain the marker information of the marker and establish the correspondence between the marker information of the marker and the three-dimensional model of the object to be labeled.

[0052] In one possible implementation, the pose information acquisition module is specifically used to: determine the pose information of the image acquisition device when acquiring each of the sample images using a simultaneous localization and mapping (SLAM) algorithm, based on each of the sample images.

[0053] In one possible implementation, the world coordinate determination module is specifically used to: for each sample image, based on the pose information of the image acquisition device when acquiring the sample image and the position of the marker in the sample image, use the SLAM algorithm to determine the position of the marker corresponding to the sample image in the world coordinate system.

[0054] In one possible implementation, the device further includes:

[0055] The marker setting module is used to set the marker at the key points of the object to be labeled, and to use the image acquisition device to acquire a sample image containing the marker and the object to be labeled;

[0056] The sample image acquisition module is used to adjust the pose of the image acquisition device and / or the position of the marker at the object to be labeled, and to acquire a sample image containing the marker and the object to be labeled using the image acquisition device;

[0057] The acquisition completion judgment module is used to call the sample image acquisition module to repeatedly acquire sample images until the acquisition termination condition is met.

[0058] In one possible implementation, the device further includes:

[0059] The rectangular frame display module is used to determine the position of the key points of the object to be labeled in the image coordinate system of the image acquisition device based on the obtained position of the marker in the world coordinate system and the current pose information of the image acquisition device, thereby obtaining the key point image position; based on the obtained key point image position, a rectangular frame is fitted; and the key point image position and the rectangular frame are displayed on the display screen corresponding to the image acquisition device.

[0060] In one possible implementation, the image acquisition module is further configured to: acquire a second image to be labeled that contains the object to be labeled but does not contain any markers;

[0061] The location information determination module is further configured to perform contour matching on the second image to be labeled based on the target 3D model to determine the location information of the object to be labeled in the second image to be labeled;

[0062] The object annotation module is further configured to annotate the object in the second image according to the position information of the object in the second image.

[0063] Thirdly, embodiments of this application provide an electronic device, including a processor and a memory;

[0064] The memory is used to store computer programs;

[0065] When the processor executes the program stored in the memory, it implements any of the automatic object annotation methods described in this application.

[0066] Fourthly, embodiments of this application provide a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, which, when executed by a processor, implements any of the object automatic annotation methods described in this application.

[0067] Fifthly, embodiments of this application provide a computer program product containing instructions, characterized in that, when the computer program product is run on a computer, it causes the computer to execute any of the object automatic annotation methods described in this application.

[0068] Beneficial effects of the embodiments in this application:

[0069] The automatic object annotation method, apparatus, electronic device, and storage medium provided in this application acquire a first image to be annotated, including markers and objects to be annotated; perform marker recognition on the first image to be annotated to obtain target marker information; determine the target 3D model corresponding to the target marker information according to a pre-set correspondence between the marker information and the object's 3D model; perform contour matching on the first image to be annotated based on the target 3D model to determine the position information of the objects to be annotated in the first image to be annotated; and annotate the objects to be annotated in the first image to be annotated according to the position information of the objects to be annotated in the first image to be annotated. This achieves automatic object annotation, reducing image annotation costs and increasing annotation efficiency. Furthermore, using markers to obtain the 3D model corresponding to the object to be annotated allows for automatic acquisition of the object's 3D model without manual setting, reducing manual workload and increasing image annotation efficiency. Of course, implementing any product or method of this application does not necessarily require achieving all of the above advantages simultaneously. Attached Figure Description

[0070] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0071] Figure 1 This is a first schematic diagram of the automatic object annotation method according to an embodiment of this application;

[0072] Figure 2 This is a second schematic diagram of the automatic object annotation method according to an embodiment of this application;

[0073] Figure 3 This is a third schematic diagram of the automatic object annotation method according to an embodiment of this application;

[0074] Figure 4 This is a fourth schematic diagram of the automatic object annotation method according to an embodiment of this application;

[0075] Figure 5 This is a first schematic diagram of the object to be labeled in an embodiment of this application;

[0076] Figure 6 This is a second schematic diagram of the object to be labeled in an embodiment of this application;

[0077] Figure 7 This is a schematic diagram of a three-dimensional sparse point cloud model according to an embodiment of this application;

[0078] Figure 8 This is a schematic diagram of an automatic object labeling device according to an embodiment of this application;

[0079] Figure 9 This is a schematic diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0080] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of this application.

[0081] This application provides an automatic object annotation method, see [link to relevant documentation]. Figure 1 The method includes:

[0082] S101, Obtain the first image to be labeled, which includes the markers and the objects to be labeled.

[0083] The automatic object labeling method of this application embodiment can be implemented by an electronic device with image processing capabilities. In one example, the electronic device can be a handheld electronic device, such as a smart camera, a hard disk recorder, or a smartphone. In another example, the electronic device can also be a personal computer or a server.

[0084] The first image to be labeled includes markers and objects to be labeled. Markers need to have distinctive appearance features to facilitate accurate identification using computer vision technology. The specific type of marker can be customized based on the actual situation; for example, a marker can be a QR code, a black and white checkerboard pattern, or other specific images. Objects to be labeled can be any object that needs to be labeled, such as vehicles, buildings, industrial parts, animals, or plants. The objects to be labeled here can be related to… Figure 3 The objects to be labeled in the illustrated embodiments can be the same object, or they can be different objects of the same type, such as two cars of the same model.

[0085] S102, perform marker recognition on the first image to be labeled to obtain target marker information.

[0086] Computer vision technology is used to identify markers in a first image to be labeled, obtaining the marker information of the markers contained therein, which is called target marker information. In one example, the marker information can be the marker's identifier, etc., and an identifier can be pre-set for each marker as its marker information. In another example, markers with the same visual features have the same marker information, while markers with different visual features have different marker information.

[0087] In one possible implementation, the marker is a QR code, the target marker information is target QR code information, and the step of identifying the marker in the first image to be labeled to obtain the target marker information includes: using QR code recognition technology to identify the QR code in the first image to be labeled to obtain the target QR code information in the first image to be labeled. In one example, the QR code information can be character information, which uniquely corresponds to a three-dimensional model of a type of object; this correspondence is a pre-set correspondence between marker information and the three-dimensional model of the object. Alternatively, the QR code information can be address information or index information, which uniquely points to a three-dimensional model of a type of object; this pointing relationship is a pre-set correspondence between marker information and the three-dimensional model of the object.

[0088] S103, determine the target three-dimensional model corresponding to the target marker information according to the pre-set correspondence between the marker information and the object's three-dimensional model.

[0089] In one example, the correspondence is between QR code information and the 3D model of an object. For instance, QR code information A corresponds to the 3D model of a vehicle, QR code information B corresponds to the 3D model of a traffic light, and so on. The 3D model corresponding to the target marker information is obtained by querying the correspondence according to the target marker information; this is called the target 3D model.

[0090] S104, perform contour matching on the first image to be labeled based on the target 3D model to determine the position information of the object to be labeled in the first image to be labeled.

[0091] The method utilizes the target 3D model to perform contour matching on the first image to be labeled, thereby determining the location information of the object to be labeled in the first image. In one example, based on the target 3D model, 2D models of the object to be labeled can be obtained from multiple angles. Each 2D model is then compared with an image region in the first image to be labeled, thus obtaining the location information of the object to be labeled in the first image. In another example, using brute-force matching, the target 3D model is adjusted by a preset unit step size at each step, and the adjusted target 3D model is projected onto a 2D plane to obtain a 2D model representing the object's 2D contour. The current 2D model is then matched with an image region in the first image to be labeled. If the match is successful, the location information of the object to be labeled in the first image is obtained; if the match fails, the angle of the target 3D model is adjusted by a unit step size, and contour matching is performed again until a match is successful or the target 3D model at all angles is matched. In one example, the location information of the object to be labeled can be the pixel region corresponding to the object.

[0092] S105, according to the position information of the object to be labeled in the first image to be labeled, label the object to be labeled in the first image to be labeled.

[0093] After obtaining the location information of the object to be labeled in the first image to be labeled, the object can be labeled in the first image according to that location information. In one example, the object to be labeled can be labeled in the first image using a rectangular label box.

[0094] In this embodiment, automatic object annotation is implemented, which can reduce the cost of image annotation and increase the efficiency of image annotation. In addition, by using markers to obtain the 3D model corresponding to the object to be annotated, the 3D model of the object to be annotated can be obtained automatically without manual setting, reducing the amount of manual work and increasing the efficiency of image annotation.

[0095] In one possible implementation, see Figure 2 The method further includes:

[0096] S201, Obtain a second image to be labeled that contains the object to be labeled but does not contain any markers.

[0097] After using markers to determine the target 3D model corresponding to the object to be labeled, an image acquisition device can be used to acquire a second image to be labeled that contains the object to be labeled but does not contain markers.

[0098] S202, based on the target 3D model, perform contour matching on the second image to be labeled to determine the position information of the object to be labeled in the second image to be labeled.

[0099] The target 3D model is used to perform contour matching on the second image to be labeled, thereby determining the location information of the object to be labeled in the second image.

[0100] S203, according to the position information of the object to be labeled in the second image to be labeled, label the object to be labeled in the second image to be labeled.

[0101] For example, in the process of annotating a vehicle, the target 3D model of the vehicle is first called using the corresponding marker. Then, in the subsequent annotation process of the vehicle, since the target 3D model of the vehicle has already been called, it can be used again without the need to call the target 3D model using the marker. In this case, the target 3D model can be directly used to perform contour matching on the second image to be annotated that does not contain the marker, thereby realizing the annotation of the object.

[0102] In this embodiment, contour matching is performed on the second image to be labeled that does not contain markers using the target 3D model. This allows for the labeling of objects in the second image to be labeled that contains markers, resulting in an image labeled without markers. This reduces the impact of markers on the training results during subsequent model training.

[0103] The 3D model of an object can be established through various modeling methods, including manual modeling, 3D laser scanning modeling, 2D image modeling with depth information, or SLAM (Simultaneous Localization and Mapping) algorithm modeling. In one possible implementation, see... Figure 3 The method further includes:

[0104] S301, acquire multiple sample images containing the marker and the object to be labeled, acquired by an image acquisition device, wherein the marker is set at multiple key points of the object to be labeled in the multiple sample images.

[0105] Image acquisition devices can be monocular cameras, binocular cameras, or smartphones with camera capabilities. Each sample image includes at least one marker. The marker's position on the object to be labeled can be the same or different in different sample images. However, for all sample images, the markers in these sample images need to represent the positions of multiple key points on the object to be labeled. In one example, to prevent duplicate acquisition, the positions of the marker and the object to be labeled are not all the same in different sample images. In another example, the marker's position on the object to be labeled may differ in different sample images, and / or the angle at which the object to be labeled is captured may differ in different sample images. The key points of the object to be labeled can be customized according to the actual situation. Key points are used to represent the outline of the object to be labeled and can be points on the outline line of the object to be labeled. In one example, significant corner positions on the object to be labeled can be selected as key points. It is understood that there will be some error in the setting of markers in actual scenarios. The marker may be placed exactly on the key point or at a small distance from the key point, as long as it can represent the outline of the object to be labeled.

[0106] S302, determine the position of the marker in each of the sample images.

[0107] Computer vision technology is used to determine the location of markers in sample images. In one example, the marker is a QR code. QR code recognition technology can be used to identify the QR codes in each sample image to obtain the location of the QR codes in each sample image.

[0108] S303, Obtain the pose information of each sample image acquired by the image acquisition device.

[0109] The pose information of an image acquisition device can include its position information (e.g., its position in a world coordinate system) and its attitude information (e.g., the shooting angle when acquiring sample images). In one example, the image acquisition device may be equipped with one or more of a gyroscope, a geomagnetic sensor, and an accelerometer to obtain its pose information.

[0110] In one possible implementation, obtaining the pose information of the image acquisition device when acquiring each of the sample images includes: determining the pose information of the image acquisition device when acquiring each of the sample images using a SLAM algorithm based on each of the sample images.

[0111] SLAM, also known as CML (Concurrent Mapping and Localization) algorithm, refers to a algorithm that places a robot in an unknown location within an unknown environment, allowing the robot to gradually create a complete map of that environment as it moves. Specifically, SLAM can use 2D images acquired by an image acquisition device to model the unknown environment, obtain the position and orientation of the image acquisition device in that environment, and obtain the positions of various objects in that environment. Using SLAM requires that the objects to be labeled be stationary, meaning they will not move or deform in the world coordinate system. The specific calculation process of the SLAM algorithm can be found in related technologies, and is not specifically limited in this application.

[0112] S304, For each sample image, based on the pose information of the image acquisition device when acquiring the sample image and the position of the marker in the sample image, determine the position of the marker corresponding to the sample image in the world coordinate system.

[0113] In this embodiment, the world coordinate system refers to the coordinate system of the real world where the sample to be labeled is located. It can be a coordinate system of latitude, longitude and altitude, or a three-dimensional coordinate system that is custom-established for the scene where the sample to be labeled is located.

[0114] In one example, the extrinsic parameters of the image acquisition device can be obtained. Based on the position of the marker in the sample image, the posture information of the image acquisition device when acquiring the sample image, and the extrinsic parameters of the image acquisition device, the position of the marker in the three-dimensional coordinate system of the image acquisition device can be obtained. Then, based on the position information of the image acquisition device in the world coordinate system and the position of the marker in the three-dimensional coordinate system of the image acquisition device, the position of the marker in the world coordinate system can be obtained.

[0115] In one example, the SLAM algorithm can be used to obtain the position of the marker in the world coordinate system. In one possible implementation, determining the position of the marker in the world coordinate system corresponding to each sample image, based on the pose information of the image acquisition device when acquiring the sample image and the position of the marker in the sample image, includes: for each sample image, using the SLAM algorithm, determining the position of the marker in the world coordinate system corresponding to the sample image, based on the pose information of the image acquisition device when acquiring the sample image and the position of the marker in the sample image.

[0116] S305, Based on the position of the marker corresponding to each sample image in the world coordinate system, establish a three-dimensional model of the object to be labeled.

[0117] Markers are placed at the keypoints of the object to be labeled. Markers in multiple sample images can be placed at different keypoints of the object. The position of the object in the world coordinate system is also the position of its keypoints in the world coordinate system. Therefore, the positions of the keypoints of the object in the world coordinate system can be used to construct a 3D model of the object in the world coordinate system. In one example, the 3D model here is a 3D sparse point cloud model, for example, a 3D model composed of keypoints represented by markers.

[0118] S306, Obtain the marker information of the marker and establish the correspondence between the marker information of the marker and the three-dimensional model of the object to be labeled.

[0119] Establish a correspondence between the marker information of the marker and the 3D model of the object to be labeled, so that the 3D model of the object to be labeled can be retrieved directly based on the marker, thereby achieving rapid and automatic labeling of the object.

[0120] In this embodiment, by combining QR codes and SLAM algorithms, the positions of key points of the object to be labeled in the three-dimensional world coordinate system can be obtained using two-dimensional images. By fusing SLAM algorithms with QR code recognition and interpretation, the problem of difficult interaction between two-dimensional images and three-dimensional scenes is solved. Compared with automatic image tracking and labeling, the method of using QR codes combined with SLAM algorithms significantly reduces engineering and manual costs; obtaining key points of the labeled object from QR codes uses an accurate contour description method, resulting in high accuracy of the labeling results.

[0121] The process of acquiring sample images is described below. In one possible implementation, see [link to relevant documentation]. Figure 4 The method further includes:

[0122] S401, the marker is placed at the key point of the object to be labeled, and the image acquisition device is used to acquire a sample image containing the marker and the object to be labeled.

[0123] In one example, see Figure 5 Taking a natural gas pipeline interface as an example, a pre-defined QR code is placed as a marker at key points on the object to be labeled, for example... Figure 6 As shown.

[0124] S402, adjust the pose of the image acquisition device and / or the position of the marker at the object to be labeled, and use the image acquisition device to acquire a sample image containing the marker and the object to be labeled.

[0125] Adjust the angle and position of the image acquisition device to capture the object to be labeled, thereby obtaining the object to be labeled containing the marker in different poses; place the marker on different key points of the object to be labeled, thereby obtaining the positions of different key points of the object to be labeled.

[0126] S403, repeat the above steps: S402 adjust the position of the image acquisition device and / or the position of the marker at the object to be labeled, and use the image acquisition device to acquire a sample image containing the marker and the object to be labeled until the acquisition termination condition is met.

[0127] Repeat step S402 until the acquisition termination condition is met. The acquisition termination condition can be customized according to actual conditions. For example, the acquisition termination condition can be that a preset number of sample images have been acquired, where the preset number can be customized according to actual conditions, but it must ensure that the preset number of sample images is sufficient to build a 3D model of the object to be labeled; for example, the acquisition termination condition can be a user-triggered command to stop acquisition, etc. In one example, the 3D sparse point cloud model of eight key points of a natural gas pipeline interface can be obtained as follows: Figure 7 As shown.

[0128] In this embodiment, a three-dimensional model is obtained by sampling two-dimensional images, which enables the automatic generation of the three-dimensional model. It has good adaptability to the scene and has significant advantages in terms of lighting conditions, camera imaging effects, and indoor and outdoor environments.

[0129] In one possible implementation, such as Figure 4 The automatic object annotation method shown allows for real-time sampling of the images to be annotated (including the first and second images) and real-time annotation of the objects within those images. After acquiring the images, the objects can be automatically annotated in real time. During the automatic annotation process, QR codes can be used for on-site annotation; almost no preparation is required, only the QR codes need to be pre-made. The annotation is completed immediately upon sampling, eliminating the need for subsequent annotation.

[0130] To make it easier for users to perceive the effect of the 3D model creation, in one possible implementation, the method further includes:

[0131] Step 1: Based on the obtained position of the marker in the world coordinate system and the current pose information of the image acquisition device, determine the position of the key points of the object to be labeled in the image coordinate system of the image acquisition device to obtain the key point image position.

[0132] The markers represent the key points of the object to be labeled. The position of the markers in the world coordinate system is the position of the key points of the object to be labeled in the actual coordinate system. Based on the real-time pose information of the image acquisition device, the transformation relationship between the image coordinate system of the image acquisition device and the world coordinate system can be obtained. Therefore, the position of the key points in the image coordinate system, that is, the image position of the key points, can be obtained.

[0133] Step 2: Based on the obtained key point image positions, fit a rectangular box.

[0134] In one example, when there is only one keypoint image location, no bounding box fitting is performed. When there are at least two keypoint image locations, a bounding box can be obtained by fitting a rectangle to each keypoint image location. For methods of fitting rectangles using multiple points, please refer to the rectangle fitting methods in related technologies. In one example, the keypoint image locations can be used as the corner points of the bounding box to fit the largest possible bounding box, ensuring that each keypoint image location falls within and on the bounding box.

[0135] Step 3: Display the key point image position and the rectangle on the display screen corresponding to the image acquisition device.

[0136] The display screen corresponding to the image acquisition device can be a built-in display screen or an external display screen. Displaying the key point image positions and rectangles on the display screen of the image acquisition device allows users to intuitively perceive the effect of the 3D model creation and the rectangle annotation results, making it convenient for users to adjust the position of markers in real time to obtain a 3D model with better annotation effects.

[0137] This application also provides an automatic object labeling device, see [link to relevant documentation]. Figure 8 The device includes:

[0138] The image acquisition module 801 is used to acquire a first image to be labeled, which includes markers and objects to be labeled.

[0139] The marker information recognition module 802 is used to perform marker recognition on the first image to be labeled to obtain target marker information;

[0140] The 3D model determination module 803 is used to determine the target 3D model corresponding to the target marker information according to the pre-set correspondence between marker information and the 3D model of the object;

[0141] The location information determination module 804 is used to perform contour matching on the first image to be labeled based on the target 3D model to determine the location information of the object to be labeled in the first image to be labeled.

[0142] The object annotation module 805 is used to annotate the object in the first image according to the position information of the object in the first image to be annotated.

[0143] In one possible implementation, the marker is a QR code, the target marker information is target QR code information, and the correspondence is the correspondence between the QR code information and the three-dimensional model of the object.

[0144] The marker information recognition module is specifically used to: use QR code recognition technology to perform QR code recognition on the first image to be labeled, and obtain the target QR code information in the first image to be labeled.

[0145] In one possible implementation, the device further includes:

[0146] The sample image acquisition module is used to acquire multiple sample images containing the marker and the object to be labeled, which are acquired by the image acquisition device, wherein the marker is set at multiple key points of the object to be labeled in the multiple sample images;

[0147] A marker location determination module is used to determine the location of the marker in each of the sample images;

[0148] The pose information acquisition module is used to acquire pose information when the image acquisition device acquires each of the sample images;

[0149] The world coordinate determination module is used to determine the position of the marker in the world coordinate system for each sample image based on the pose information of the image acquisition device when the sample image is acquired and the position of the marker in the sample image.

[0150] A 3D model building module is used to build a 3D model of the object to be labeled based on the position of the marker corresponding to each sample image in the world coordinate system.

[0151] The correspondence establishment module is used to obtain the marker information of the marker and establish the correspondence between the marker information of the marker and the three-dimensional model of the object to be labeled.

[0152] In one possible implementation, the pose information acquisition module is specifically used to: determine the pose information of the image acquisition device when acquiring each of the sample images using a simultaneous localization and mapping (SLAM) algorithm, based on each of the sample images.

[0153] In one possible implementation, the world coordinate determination module is specifically used to: for each sample image, based on the pose information of the image acquisition device when acquiring the sample image and the position of the marker in the sample image, use the SLAM algorithm to determine the position of the marker corresponding to the sample image in the world coordinate system.

[0154] In one possible implementation, the device further includes:

[0155] The marker setting module is used to set the marker at the key points of the object to be labeled, and to use the image acquisition device to acquire a sample image containing the marker and the object to be labeled;

[0156] The sample image acquisition module is used to adjust the pose of the image acquisition device and / or the position of the marker at the object to be labeled, and to acquire a sample image containing the marker and the object to be labeled using the image acquisition device;

[0157] The acquisition completion judgment module is used to call the sample image acquisition module to repeatedly acquire sample images until the acquisition termination condition is met.

[0158] In one possible implementation, the device further includes:

[0159] The rectangular frame display module is used to determine the position of the key points of the object to be labeled in the image coordinate system of the image acquisition device based on the obtained position of the marker in the world coordinate system and the current pose information of the image acquisition device, thereby obtaining the key point image position; based on the obtained key point image position, a rectangular frame is fitted; and the key point image position and the rectangular frame are displayed on the display screen corresponding to the image acquisition device.

[0160] In one possible implementation, the image acquisition module is further configured to: acquire a second image to be labeled that contains the object to be labeled but does not contain any markers;

[0161] The location information determination module is further configured to perform contour matching on the second image to be labeled based on the target 3D model to determine the location information of the object to be labeled in the second image to be labeled;

[0162] The object annotation module is further configured to annotate the object in the second image according to the position information of the object in the second image.

[0163] This application also provides an electronic device, including: a processor and a memory;

[0164] The aforementioned memory is used to store computer programs;

[0165] When the processor executes the computer program stored in the memory, it implements any of the automatic object annotation methods described in this application.

[0166] Optional, see Figure 9 The electronic device in this application embodiment also includes a communication interface 902 and a communication bus 904, wherein the processor 901, the communication interface 902, and the memory 903 communicate with each other through the communication bus 904.

[0167] The communication bus mentioned in the above electronic devices can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0168] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0169] The memory may include RAM (Random Access Memory) or NVM (Non-Volatile Memory), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0170] The processors mentioned above can be general-purpose processors, including CPUs (Central Processing Units), NPs (Network Processors), etc.; they can also be DSPs (Digital Signal Processors), ASICs (Application Specific Integrated Circuits), FPGAs (Field-Programmable Gate Arrays), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0171] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the object automatic annotation methods described in this application.

[0172] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the object automatic annotation methods described in the above embodiments.

[0173] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0174] It should be noted that, in this document, the technical features of the various alternative solutions can be combined to form solutions as long as they are not contradictory, and these solutions are all within the scope of this application. Relational terms such as "first" and "second" are merely used to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0175] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, embodiments of devices, electronic devices, and storage media are basically similar to the method embodiments, and therefore the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0176] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.

Claims

1. An automatic object annotation method, characterized in that, The method includes: Obtain the first image to be labeled, which contains the markers and the objects to be labeled; The first image to be labeled is subjected to marker recognition to obtain target marker information; According to the pre-set correspondence between marker information and the object's three-dimensional model, the target three-dimensional model corresponding to the target marker information is determined, wherein the three-dimensional model is a three-dimensional sparse point cloud model; Based on the target 3D model, contour matching is performed on the first image to be labeled to determine the position information of the object to be labeled in the first image to be labeled. According to the position information of the object to be labeled in the first image to be labeled, the object to be labeled is labeled in the first image to be labeled.

2. The method according to claim 1, characterized in that, The marker is a QR code, the target marker information is the target QR code information, and the correspondence is the correspondence between the QR code information and the three-dimensional model of the object; The step of identifying target markers in the first image to be labeled to obtain target marker information includes: The first image to be labeled is identified using QR code recognition technology to obtain the target QR code information in the first image to be labeled.

3. The method according to claim 1, characterized in that, The method further includes: Acquire multiple sample images containing the marker and the object to be labeled, acquired by an image acquisition device, wherein the marker is set at multiple key points of the object to be labeled in the multiple sample images; Determine the position of the marker in each of the sample images; Obtain the pose information of the image acquisition device when acquiring each of the sample images; For each sample image, based on the pose information of the image acquisition device when acquiring the sample image and the position of the marker in the sample image, the position of the marker in the world coordinate system corresponding to the sample image is determined; Based on the position of the marker corresponding to each sample image in the world coordinate system, a three-dimensional model of the object to be labeled is established. Obtain the marker information of the marker and establish the correspondence between the marker information of the marker and the three-dimensional model of the object to be labeled.

4. The method according to claim 3, characterized in that, The step of obtaining the pose information of each sample image acquired by the image acquisition device includes: Based on each of the sample images, the pose information of the image acquisition device when acquiring each of the sample images is determined using the Simultaneous Localization and Mapping (SLAM) algorithm.

5. The method according to claim 4, characterized in that, For each sample image, determining the position of the marker in the world coordinate system based on the pose information of the image acquisition device when acquiring the sample image and the position of the marker in the sample image includes: For each sample image, based on the pose information of the image acquisition device when acquiring the sample image and the position of the marker in the sample image, the SLAM algorithm is used to determine the position of the marker in the world coordinate system corresponding to the sample image.

6. The method according to claim 3, characterized in that, The method further includes: The marker is placed at the key points of the object to be labeled, and the image acquisition device is used to acquire a sample image containing the marker and the object to be labeled; Adjust the pose of the image acquisition device and / or the position of the marker at the object to be labeled, and use the image acquisition device to acquire a sample image containing the marker and the object to be labeled; Repeat the above steps: adjust the position of the image acquisition device and / or the position of the marker at the object to be labeled, and use the image acquisition device to acquire sample images containing the marker and the object to be labeled until the acquisition termination condition is met.

7. The method according to any one of claims 3-6, characterized in that, The method further includes: Based on the obtained position of the marker in the world coordinate system and the current pose information of the image acquisition device, the position of the key points of the object to be labeled in the image coordinate system of the image acquisition device is determined to obtain the key point image position; Based on the obtained key point image positions, a rectangular bounding box is fitted; The key point image positions and the rectangles are displayed on the screen corresponding to the image acquisition device.

8. The method according to claim 1, characterized in that, The method further includes: Obtain a second image to be labeled that contains the object to be labeled but does not contain any markers; Based on the target 3D model, contour matching is performed on the second image to be labeled to determine the position information of the object to be labeled in the second image to be labeled. According to the position information of the object to be labeled in the second image to be labeled, the object to be labeled is labeled in the second image to be labeled.

9. An automatic object labeling device, characterized in that, The device includes: The image acquisition module is used to acquire a first image containing markers and objects to be labeled. The marker information recognition module is used to recognize markers in the first image to be labeled and obtain target marker information. The 3D model determination module is used to determine the target 3D model corresponding to the target marker information according to the pre-set correspondence between the marker information and the 3D model of the object, wherein the 3D model is a 3D sparse point cloud model; The location information determination module is used to perform contour matching on the first image to be labeled based on the target 3D model to determine the location information of the object to be labeled in the first image to be labeled. The object annotation module is used to annotate the object in the first image according to the position information of the object in the first image to be annotated.

10. The apparatus according to claim 9, characterized in that, The marker is a QR code, the target marker information is the target QR code information, and the correspondence is the correspondence between the QR code information and the three-dimensional model of the object; The marker information recognition module is specifically used to: use QR code recognition technology to perform QR code recognition on the first image to be labeled, and obtain the target QR code information in the first image to be labeled. The device further includes: The sample image acquisition module is used to acquire multiple sample images containing the marker and the object to be labeled, which are acquired by the image acquisition device, wherein the marker is set at multiple key points of the object to be labeled in the multiple sample images; A marker location determination module is used to determine the location of the marker in each of the sample images; The pose information acquisition module is used to acquire pose information when the image acquisition device acquires each of the sample images; The world coordinate determination module is used to determine the position of the marker in the world coordinate system for each sample image based on the pose information of the image acquisition device when the sample image is acquired and the position of the marker in the sample image. A 3D model building module is used to build a 3D model of the object to be labeled based on the position of the marker corresponding to each sample image in the world coordinate system. The correspondence establishment module is used to obtain the marker information of the marker and establish the correspondence between the marker information of the marker and the three-dimensional model of the object to be labeled; The pose information acquisition module is specifically used to: determine the pose information of the image acquisition device when acquiring each of the sample images using the Simultaneous Localization and Mapping (SLAM) algorithm, based on each of the sample images; The world coordinate determination module is specifically used to: for each sample image, based on the pose information of the image acquisition device when acquiring the sample image and the position of the marker in the sample image, use the SLAM algorithm to determine the position of the marker corresponding to the sample image in the world coordinate system; The device further includes: The marker setting module is used to set the marker at the key points of the object to be labeled, and to use the image acquisition device to acquire a sample image containing the marker and the object to be labeled; The sample image acquisition module is used to adjust the pose of the image acquisition device and / or the position of the marker at the object to be labeled, and to acquire a sample image containing the marker and the object to be labeled using the image acquisition device; The acquisition completion judgment module is used to call the sample image acquisition module to repeatedly acquire sample images until the acquisition termination condition is met. The device further includes: The rectangular frame display module is used to determine the position of the key points of the object to be labeled in the image coordinate system of the image acquisition device based on the obtained position of the marker in the world coordinate system and the current pose information of the image acquisition device, thereby obtaining the key point image position; based on the obtained key point image position, a rectangular frame is fitted; and the key point image position and the rectangular frame are displayed on the display screen corresponding to the image acquisition device. The image acquisition module is further configured to: acquire a second image to be labeled that contains the object to be labeled but does not contain any markers; The location information determination module is further configured to perform contour matching on the second image to be labeled based on the target 3D model to determine the location information of the object to be labeled in the second image to be labeled; The object annotation module is further configured to annotate the object in the second image according to the position information of the object in the second image.

11. An electronic device, characterized in that, Including processor and memory; The memory is used to store computer programs; When the processor executes the program stored in the memory, it implements the automatic object annotation method according to any one of claims 1-8.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the automatic object annotation method according to any one of claims 1-8.

13. A computer program product containing instructions, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the automatic object annotation method according to any one of claims 1-8.