Label information generation method, device, system, electronic equipment and storage medium

By utilizing the motion parameters and pose of image acquisition devices during deep learning model training, the pose of objects to be labeled is automatically annotated, solving the problems of high labeling costs and the impact on model accuracy, and realizing an efficient and low-cost labeling method.

CN116266388BActive Publication Date: 2025-11-04GUANGDONG GAOHANG INTELLECTUAL PROPERTY OPERATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111506095.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-10
Publication Date
2025-11-04
Estimated Expiration
2041-12-10

AI Technical Summary

Technical Problem

In the process of training deep learning models, the annotation of massive sample images is costly and manual annotation is inefficient. At the same time, using images containing markers to train the model will affect the model accuracy.

Method used

By acquiring images with and without markers captured by the image acquisition device during its movement, and utilizing the movement parameters and orientation of the image acquisition device, the pose of the object to be labeled is automatically annotated, generating annotation information without markers.

Benefits of technology

It achieves reduced annotation costs and improved annotation efficiency without affecting model accuracy, and automatically annotates images that do not contain landmarks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116266388B_ABST
    Figure CN116266388B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a labeling information generation method, device, system, electronic equipment and storage medium. The first type of image and the second type of image collected by an image collection device during movement are obtained, the movement parameters of the image collection device, the posture of the image collection device when collecting the first type of image and the second type of image, and the second type of image does not include the first marker. The pose of the first marker in the first type of image is determined. According to the pre-determined relative pose relationship between the first marker and the object to be labeled, and the pose of the first marker in the first type of image, the pose of the object to be labeled in the first type of image is determined. According to the movement parameters, the posture of the image collection device, and the pose of the object to be labeled in the first type of image, the pose of the object to be labeled in the second type of image is determined. Based on the pose of the object to be labeled in the second type of image, the labeling information of the object to be labeled in the second type of image is generated. The image without the marker in the automatic labeling is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to methods, apparatus, systems, electronic devices and storage media for generating annotation information. Background Technology

[0002] With the development of artificial intelligence technology, computer vision technology, especially deep learning-based computer vision technology, has developed rapidly. In computer vision technology, deep learning models need to be trained on a large amount of labeled data. For example, when using computer vision technology to identify vehicles, a large number of sample images labeled with vehicles are selected to train the deep learning model.

[0003] In the training process of deep learning models, the cost of labeling massive sample images has become the biggest cost in the training process. In order to reduce the problems of low efficiency and high cost of manual labeling, related technologies set up markers around the object to be labeled, obtain the pose relationship between the markers and the object to be labeled, keep the pose relationship unchanged, and take an image containing the markers and the object to be labeled. Thus, the object to be labeled can be automatically labeled in the image based on the pose of the markers in the image.

[0004] To facilitate the acquisition of the pose of the marker, the appearance features of the marker must be obvious and easily distinguishable. However, although the objects to be labeled are automatically marked in the images obtained by the above method, the appearance features of the marker are too obvious. Using images containing the marker to train deep learning models will seriously affect the accuracy of deep learning models. Summary of the Invention

[0005] The purpose of this application is to provide a method, apparatus, system, electronic device, and storage medium for generating annotation information, so as to ensure that automatically annotated images do not contain markers. The specific technical solution is as follows:

[0006] In a first aspect, embodiments of this application provide a method for generating annotation information, including:

[0007] Acquire a first type of image and a second type of image acquired by an image acquisition device during movement, the movement parameters of the image acquisition device, and the posture of the image acquisition device when acquiring the first type of image and the second type of image, wherein the first type of image includes an object to be labeled and a first marker, the second type of image includes an object to be labeled but does not include the first marker, and the relative pose of the object to be labeled and the first marker in the world coordinate system remains unchanged;

[0008] For each image of the first type, determine the pose of the first marker in that image;

[0009] For each first type of image, the pose of the object to be labeled in the first type of image is determined based on the predetermined relative pose relationship between the first marker and the object to be labeled, and the pose of the first marker in the first type of image.

[0010] For each second type of image, the pose of the object to be labeled in the second type of image is determined based on the movement parameters of the image acquisition device, the posture of the image acquisition device when acquiring the second type of image, and the pose of the object to be labeled in the first type of image.

[0011] Based on the pose of the object to be labeled in each of the second type of images, annotation information of the object to be labeled in each of the second type of images is generated.

[0012] In one possible implementation, the method further includes:

[0013] Acquire a third type of image captured by an image acquisition device, wherein the third type of image includes an object to be labeled with a second marker and a first marker, the second marker being set on a key point of the object to be labeled;

[0014] For each third-class image, determine the relative pose relationship between the second marker and the first marker in that third-class image;

[0015] For each third-class image, based on the relative pose relationship between the second marker and the first marker in the third-class image, the relative pose relationship between the first marker and the object to be labeled is determined.

[0016] In one possible implementation, determining the pose of the object to be labeled in the second type of image for each second type of image, based on the movement parameters of the image acquisition device, the posture of the image acquisition device when acquiring the second type of image, and the pose of the object to be labeled in the first type of image, includes:

[0017] For each second type of image, based on the movement parameters of the image acquisition device and the posture of the image acquisition device when acquiring the second type of image, the pose transformation parameters of the image acquisition device when acquiring the specified first type of image and when acquiring the second type of image are determined, and the first pose transformation parameters of the second type of image are obtained.

[0018] Based on the first pose transformation parameter of the second type of image, determine the image coordinate transformation relationship between the specified first type of image and the second type of image;

[0019] Based on the image coordinate transformation relationship of the second type of image and the pose of the object to be labeled in the specified first type of image, the pose of the object to be labeled in the second type of image is determined.

[0020] In one possible implementation, determining the pose of the object to be labeled in the second type of image for each second type of image, based on the movement parameters of the image acquisition device, the posture of the image acquisition device when acquiring the second type of image, and the pose of the object to be labeled in the first type of image, includes:

[0021] For each second type of image, the first world coordinates when the image acquisition device acquires the specified first type of image and the second world coordinates when the image acquisition device acquires the second type of image are determined according to the movement parameters of the image acquisition device;

[0022] The third world coordinates of the object to be labeled are determined based on the pose of the object to be labeled in the specified first type of image, the pose of the image acquisition device when acquiring the specified first type of image, and the first world coordinates.

[0023] Based on the pose of the object to be labeled in the second type of image when the image acquisition device acquires the second type of image, the second world coordinates, and the third world coordinates, the pose of the object to be labeled in the second type of image is determined.

[0024] In one possible implementation, the method further includes:

[0025] Based on the movement parameters of the image acquisition device and the posture, determine the pose transformation parameters of the image acquisition device when acquiring at least two first-type images, and obtain the second pose transformation parameters;

[0026] Based on the pose of the object to be labeled in one of the at least two first-class images and the second pose transformation parameters, the predicted pose of the object to be labeled in the other first-class images in the at least two first-class images is determined;

[0027] Based on the pose of the object to be labeled in other first-class images among the at least two first-class images and the predicted pose, determine the pose correction parameters;

[0028] The pose correction parameters are used to correct the pose of the object to be labeled in the second type of image.

[0029] Secondly, embodiments of this application provide a labeling information generation system, including:

[0030] Image acquisition equipment, first marker generation equipment, computing equipment;

[0031] The first marker generating device is used to generate a first marker at a specified location;

[0032] The image acquisition device is used to acquire a first type of image and a second type of image of the object to be labeled during the movement, and to record the movement parameters and the posture when acquiring the first type of image and the second type of image. The first type of image includes a first marker, the second type of image does not include the first marker, and the relative pose of the object to be labeled and the first marker in the world coordinate system remains unchanged.

[0033] The computing device is configured to: determine the pose of the first marker in each first-type image; determine the pose of the object to be labeled in each first-type image based on a predetermined relative pose relationship between the first marker and the object to be labeled, and the pose of the first marker in the first-type image; determine the pose of the object to be labeled in each second-type image based on the movement parameters of the image acquisition device, the posture of the image acquisition device when acquiring the second-type image, and the pose of the object to be labeled in the first-type image; and generate labeling information for the object to be labeled in each second-type image based on the pose of the object to be labeled in each second-type image.

[0034] In one possible implementation, the system further includes:

[0035] The second marker is placed on the key point of the object to be marked;

[0036] The image acquisition device is also used to acquire a third type of image of the object to be labeled, wherein the third type of image includes the first marker;

[0037] The computing device is further configured to determine the relative pose relationship between the second marker and the first marker in the third type of image; and based on the relative pose relationship between the second marker and the first marker in the third type of image, to determine the relative pose relationship between the first marker and the object to be labeled.

[0038] In one possible implementation, the computing device is specifically configured to: for each second type of image, determine the pose transformation parameters of the image acquisition device when acquiring a specified first type of image and when acquiring the second type of image, based on the movement parameters of the image acquisition device and the pose of the image acquisition device when acquiring the second type of image, to obtain the first pose transformation parameter of the second type of image; determine the image coordinate transformation relationship between the specified first type of image and the second type of image based on the first pose transformation parameter of the second type of image; and determine the pose of the object to be labeled in the second type of image based on the image coordinate transformation relationship of the second type of image and the pose of the object to be labeled in the specified first type of image.

[0039] In one possible implementation, the computing device is specifically configured to: for each second type of image, determine, based on the movement parameters of the image acquisition device, the first world coordinates when the image acquisition device acquires a specified first type of image and the second world coordinates when acquiring the second type of image; determine, based on the pose of the object to be labeled in the specified first type of image, the pose of the image acquisition device when acquiring the specified first type of image, and the first world coordinates, determine the third world coordinates of the object to be labeled in the second type of image; and determine the pose of the object to be labeled in the second type of image based on the pose of the image acquisition device when acquiring the second type of image, the second world coordinates, and the third world coordinates.

[0040] In one possible implementation, the computing device is further configured to: determine pose transformation parameters for the image acquisition device when acquiring at least two first-type images based on the movement parameters of the image acquisition device and the pose, thereby obtaining second pose transformation parameters; determine the predicted pose of the object to be labeled in other first-type images among the at least two first-type images based on the pose of the object to be labeled in one of the at least two first-type images and the second pose transformation parameters; determine pose correction parameters based on the pose of the object to be labeled in the other first-type images among the at least two first-type images and the predicted pose; and correct the pose of the object to be labeled in the second-type image using the pose correction parameters.

[0041] Thirdly, embodiments of this application provide a labeling information generation apparatus, including:

[0042] The first parameter acquisition module is used to acquire a first type of image and a second type of image acquired by the image acquisition device during movement, the movement parameters of the image acquisition device, and the posture of the image acquisition device when acquiring the first type of image and the second type of image. The first type of image includes an object to be labeled and a first marker, the second type of image includes an object to be labeled but does not include the first marker, and the relative pose of the object to be labeled and the first marker in the world coordinate system remains unchanged.

[0043] The first pose determination module is used to determine the pose of the first marker in each first type of image.

[0044] The second pose determination module is used to determine the pose of the object to be labeled in the first type of image for each first type of image based on the pre-determined relative pose relationship between the first marker and the object to be labeled, and the pose of the first marker in the first type of image.

[0045] The third pose determination module is used to determine the pose of the object to be labeled in the second type of image for each second type of image based on the movement parameters of the image acquisition device, the pose of the image acquisition device when acquiring the second type of image, and the pose of the object to be labeled in the first type of image.

[0046] The annotation information generation module is used to generate annotation information for the objects to be annotated in each of the second type of images based on their poses.

[0047] In one possible implementation, the device further includes:

[0048] The second parameter acquisition module is used to acquire a third type of image acquired by the image acquisition device. The third type of image includes an object to be labeled with a second marker and a first marker. The second marker is set on a key point of the object to be labeled.

[0049] The first pose relationship determination module is used to determine the relative pose relationship between the second marker and the first marker in each third type of image;

[0050] The second pose relationship determination module is used to determine the relative pose relationship between the first marker and the object to be labeled for each third type of image, based on the relative pose relationship between the second marker and the first marker in the third type of image.

[0051] In one possible implementation, the third pose determination module is specifically used to: for each second type of image, determine the pose transformation parameters of the image acquisition device when acquiring a specified first type of image and when acquiring the second type of image based on the movement parameters of the image acquisition device and the pose of the image acquisition device when acquiring the second type of image, and obtain the first pose transformation parameter of the second type of image;

[0052] Based on the first pose transformation parameter of the second type of image, determine the image coordinate transformation relationship between the specified first type of image and the second type of image;

[0053] Based on the image coordinate transformation relationship of the second type of image and the pose of the object to be labeled in the specified first type of image, the pose of the object to be labeled in the second type of image is determined.

[0054] In one possible implementation, the third pose determination module is specifically used to: for each second type of image, determine the first world coordinates when the image acquisition device acquires the specified first type of image and the second world coordinates when acquiring the second type of image, based on the movement parameters of the image acquisition device;

[0055] The third world coordinates of the object to be labeled are determined based on the pose of the object to be labeled in the specified first type of image, the pose of the image acquisition device when acquiring the specified first type of image, and the first world coordinates.

[0056] Based on the pose of the object to be labeled in the second type of image when the image acquisition device acquires the second type of image, the second world coordinates, and the third world coordinates, the pose of the object to be labeled in the second type of image is determined.

[0057] In one possible implementation, the device further includes:

[0058] The pose correction module is used to determine pose transformation parameters when the image acquisition device acquires at least two first-type images based on the movement parameters and the posture of the image acquisition device, thereby obtaining second pose transformation parameters; determine the predicted pose of the object to be labeled in the other first-type images among the at least two first-type images based on the pose of the object to be labeled in one of the at least two first-type images and the second pose transformation parameters; determine pose correction parameters based on the pose of the object to be labeled in the other first-type images among the at least two first-type images and the predicted pose; and correct the pose of the object to be labeled in the second-type images using the pose correction parameters.

[0059] Fourthly, embodiments of this application provide an electronic device, including a processor and a memory;

[0060] The memory is used to store computer programs;

[0061] When the processor executes the program stored in the memory, it implements any of the annotation information generation methods described in this application.

[0062] Fifthly, embodiments of this application provide a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, which, when executed by a processor, implements any of the annotation information generation methods described in this application.

[0063] Sixthly, embodiments of this application also provide a computer program product containing instructions that, when run on a computer, cause the computer to execute any of the annotation information generation methods described above.

[0064] Beneficial effects of the embodiments in this application:

[0065] The annotation information generation method, apparatus, system, electronic device, and storage medium provided in this application embodiment acquire a first type of image and a second type of image acquired by an image acquisition device during movement, the movement parameters of the image acquisition device, and the posture of the image acquisition device when acquiring the first type of image and the second type of image. The first type of image includes an object to be annotated and a first marker, while the second type of image includes the object to be annotated but does not include the first marker. The method involves determining the pose of the first marker in the first type of image; determining the pose of the object to be annotated in the first type of image based on the predetermined relative pose relationship between the first marker and the object to be annotated, and the pose of the first marker in the first type of image; determining the pose of the object to be annotated in the second type of image based on the movement parameters and posture of the image acquisition device and the pose of the object to be annotated in the first type of image; and generating annotation information for the object to be annotated in the second type of image based on the pose of the object to be annotated in the second type of image.

[0066] The second type of image is an image that does not contain markers; images that have been automatically labeled do not contain markers. Of course, implementing any product or method of this application does not necessarily require achieving all of the advantages described above simultaneously. Attached Figure Description

[0067] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0068] Figure 1 This is a first schematic diagram of the annotation information generation system according to an embodiment of this application;

[0069] Figure 2 This is a second schematic diagram of the annotation information generation system according to an embodiment of this application;

[0070] Figure 3 This is a schematic diagram of a method for generating annotation information according to an embodiment of this application;

[0071] Figure 4 This is a schematic diagram of a possible implementation of step S304 in an embodiment of this application;

[0072] Figure 5 This is a schematic diagram illustrating another possible implementation of step S304 in an embodiment of this application;

[0073] Figure 6 This is a schematic diagram of a labeling information generation device according to an embodiment of this application;

[0074] Figure 7This is a schematic diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0075] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of this application.

[0076] This application provides a system for generating annotation information; see [link to documentation]. Figure 1 ,include:

[0077] Image acquisition device 11, first marker generation device 12, computing device 13;

[0078] The first marker generating device 12 is used to generate a first marker at a designated location;

[0079] The image acquisition device 11 is used to acquire a first type of image and a second type of image of the object to be labeled during the movement, and to record the movement parameters and the posture when acquiring the first type of image and the second type of image. The first type of image includes a first marker, the second type of image does not include the first marker, and the relative posture of the object to be labeled and the first marker in the world coordinate system remains unchanged.

[0080] The computing device 13 is configured to: determine the pose of the first marker in each first-type image; determine the pose of the object to be labeled in each first-type image based on a predetermined relative pose relationship between the first marker and the object to be labeled, and the pose of the first marker in the first-type image; determine the pose of the object to be labeled in each second-type image based on the movement parameters of the image acquisition device, the posture of the image acquisition device when acquiring the second-type image, and the pose of the object to be labeled in the first-type image; and generate labeling information for the object to be labeled in each second-type image based on the pose of the object to be labeled in each second-type image.

[0081] The pose in this application includes position and orientation. Position can be understood as the coordinates of an object, and orientation can be understood as the angle of an object. The first marker generating device can be a projector or laser, etc., used to generate the first marker at a specified location. The first marker needs to have obvious appearance features so that it can be accurately identified from the image using computer vision technology. The specific form of the first marker can be customized according to the actual situation. In one example, the first marker can be a QR code, a black and white checkerboard image, etc.

[0082] Image acquisition devices can acquire images of objects to be labeled. These objects can be any object requiring labeling, specifically configured according to the training needs of the deep learning model. For example, if the deep learning model is used for vehicle recognition, the objects to be labeled can be vehicles. Image acquisition devices can include monocular or binocular cameras. Besides image acquisition, they also need to have positioning and attitude acquisition capabilities. For example, an image acquisition device may include one or more of a gyroscope, a geomagnetic sensor, and an accelerometer to obtain the device's movement parameters or attitude. Another example is a pan-tilt unit (PTZ) to obtain the device's attitude. The movement parameters of the image acquisition device can be its velocity, acceleration, or world coordinates; the attitude specifically refers to the attitude of the camera within the device.

[0083] In addition, image acquisition devices may also include self-propelled modules and wireless modules. The self-propelled module is used to enable the movement of the image acquisition device. In one example, the self-propelled module can be an AGV (Automated Guided Vehicle). The wireless module is used for data interaction between the image acquisition device and the computing device. In one example, the wireless module can also be used to obtain the movement parameters of the image acquisition device through wireless positioning technology.

[0084] The computing device is used to generate annotation information for the objects to be labeled in the second type of image. In one example, the computing device can be a smartphone, tablet, dedicated embedded device, laptop, or personal computer. In another example, the computing device can also be used to control the pose of the image acquisition device. In yet another example, the computing device can also be used to control the first marker generation device to generate or stop generating the first marker.

[0085] The first marker generation device generates a first marker, maintaining the relative pose between the first marker and the object to be labeled. An image acquisition device then acquires images of the object to be labeled during the movement. During image acquisition, the first marker generation device intermittently generates the first marker (while maintaining the relative pose between the first marker and the object to be labeled). Therefore, the images acquired by the image acquisition device will include two types of images: a first type containing both the first marker and the object to be labeled, and a second type containing the object to be labeled but excluding the first marker. Simultaneously with image acquisition, the image acquisition device also needs to record its own movement parameters and orientation.

[0086] In one example, during the process of image acquisition by the image acquisition device, after the first marker generation device generates a first marker for a first preset duration, the first marker generation device can be turned off for a second preset duration. That is, after generating the first marker for the first preset duration, there is a second preset duration during which the first marker is not generated. This process is repeated to obtain the first type of image and the second type of image.

[0087] The relative pose relationship between the first marker and the object to be labeled can be predetermined. The first marker has obvious visual effects, and the computing device can accurately acquire the pose of the first marker in the first type of image. Then, combined with the predetermined relative pose relationship between the first marker and the object to be labeled, the pose of the object to be labeled in the first type of image can be obtained. Based on the movement parameters and attitude of the image acquisition device, using a preset 3D tracking algorithm, the object to be labeled can be tracked, thereby obtaining the pose of the object to be labeled in the second type of image. The pose of the object to be labeled in the world coordinate system is fixed, but the pose of the image acquisition device in the world coordinate system is changing. Determining the pixel region of the object to be labeled in the image data acquired by the image acquisition device is to achieve the tracking of the object to be labeled. In one example, based on the movement parameters and attitude of the image acquisition device, the pose of the image acquisition device in the world coordinate system can be obtained. Since the pose of the object to be labeled in the world coordinate system is fixed, the position of the object to be labeled in the image coordinate system of the image acquisition device can be obtained based on the pose of the image acquisition device in the world coordinate system and the pose of the object to be labeled in the world coordinate system, thus achieving the tracking of the object to be labeled.

[0088] The preset 3D tracking algorithm can be found in the 3D tracking algorithm in related technologies. In one example, the coordinate transformation relationship between the first type of image and the second type of image can be obtained based on the movement parameters and posture of the image acquisition device. Based on the coordinate transformation relationship between the first type of image and the second type of image and the pose of the object to be labeled in the first type of image, the pose of the object to be labeled in the second type of image can be obtained. Based on the pose of the object to be labeled in the second type of image, the labeling of the object to be labeled in the second type of image is completed, and the labeling information is obtained, thereby realizing that the automatically labeled image does not contain markers.

[0089] The following describes how to determine the relative pose relationship between the first marker and the object to be labeled. In one possible implementation, see... Figure 2 The system further includes: a second marker, which is disposed on a key point of the object to be marked;

[0090] The image acquisition device is also used to acquire a third type of image of the object to be labeled, wherein the third type of image includes the first marker;

[0091] The computing device is further configured to, for each third type of image, determine the relative pose relationship between the second marker and the first marker in the third type of image; and, for each third type of image, determine the relative pose relationship between the first marker and the object to be labeled based on the relative pose relationship between the second marker and the first marker in the third type of image.

[0092] Before determining the relative pose between the first marker and the object to be labeled, the positions of each part must first be set. The first marker generating device (such as a projection device) can be mounted on a tripod and set to continuously generate the first marker. The position of the generated first marker is adjusted so that the distance between the first marker and the object to be labeled is within a preset range. This preset range needs to ensure that the image acquisition device can simultaneously capture the first marker and the object to be labeled at the guided shooting distance.

[0093] The second marker can be a metal or plastic sheet printed with a special pattern (such as a QR code or a black and white checkerboard pattern); the second marker can also be a marker generated by the first marker generation device that is different from the first marker. A 3D model of the object to be labeled is acquired using a computing device, and key points on the 3D model are determined. In one example, easily identifiable and placement locations of the second marker on the object to be labeled can be selected as key points; for example, key points can be distinctive corner locations. In another example, the second marker can be placed or projected onto the key points of the object to be labeled. A corresponding program is started on the computing device, which acquires a third type of image captured by an image acquisition device and the pose at the time of acquisition. The relative pose relationship between the first and second markers in the third type of image is detected. Combining the position of the second marker (key point) on the 3D model of the object to be labeled and the pose at the time of acquisition of the third type of image, the relative pose relationship between the first marker and the object to be labeled in the world coordinate system can be obtained, and this positional relationship data is stored.

[0094] The second marker needs to have distinct visual features, and these features must differ from those of the first marker to allow for differentiation. The second marker is placed on key points of the object to be labeled, thus allowing the relative pose relationship between the second marker and the object to be labeled to be obtained. In one example, the relative pose relationship between the second marker and the object to be labeled specifically refers to the relative pose relationship between the second marker and the 3D model of the object to be labeled. Images of the object to be labeled, with the second marker attached, are acquired using an image acquisition device; these are called third-class images. The third-class images also include the image of the first marker. The first and second markers are identified from the third-class images, thus obtaining their relative pose relationship. Then, based on the relative pose relationships of the first and second markers, and the relative pose relationship between the second marker and the object to be labeled, the relative pose relationship between the first marker and the object to be labeled can be obtained. In one example, the relative pose relationship between the first marker and the object to be labeled can be the relative pose relationship between the two in the camera coordinate system; in another example, the relative pose relationship between the first marker and the object to be labeled can be the relative pose relationship between the two in the world coordinate system.

[0095] It is understood that in the annotation information generation method of this application embodiment, it is necessary to keep the relative pose of the object to be annotated and the first marker unchanged in the world coordinate system. If the relative pose of the two changes, the relative pose relationship between the first marker and the object to be annotated needs to be recalibrated.

[0096] In one example, the coordinate transformation relationship between the image coordinate system of the first type of image and the image coordinate system of the second type of image can be directly calculated. In one possible implementation, the computing device is specifically used to: for each second type of image, determine the pose transformation parameters of the image acquisition device when acquiring the specified first type of image and when acquiring the second type of image, based on the movement parameters of the image acquisition device and the pose of the image acquisition device when acquiring the second type of image, to obtain the first pose transformation parameter of the second type of image; determine the image coordinate transformation relationship between the specified first type of image and the second type of image based on the first pose transformation parameter of the second type of image; and determine the pose of the object to be labeled in the second type of image based on the image coordinate transformation relationship of the second type of image and the pose of the object to be labeled in the specified first type of image. In one example, the first pose transformation parameter includes a first position change parameter (e.g., which can be represented by a translation vector) and a first pose change parameter (e.g., which can be represented by a rotation matrix).

[0097] In one example, a transformation relationship between the image coordinate system and the world coordinate system of a first type of image can be established, as well as a transformation relationship between the image coordinate system and the world coordinate system of a second type of image, thereby indirectly obtaining the image coordinate transformation relationship between the first type of image and the second type of image. In one possible implementation, the computing device is specifically used for: for each second type of image, determining the first world coordinates when the image acquisition device acquires a specified first type of image and the second world coordinates when acquiring the second type of image, based on the movement parameters of the image acquisition device; determining the third world coordinates of the object to be labeled based on the pose of the object to be labeled in the specified first type of image, the pose of the image acquisition device when acquiring the specified first type of image, and the first world coordinates; and determining the pose of the object to be labeled in the second type of image based on the pose of the image acquisition device when acquiring the second type of image, the second world coordinates, and the third world coordinates.

[0098] To further improve the accuracy of the annotation information of the object to be labeled, pose correction parameters can be calculated and used to correct the pose of the object to be labeled. In one possible implementation, the computing device is further configured to: determine pose transformation parameters when the image acquisition device acquires at least two first-type images based on the movement parameters of the image acquisition device and the pose, and obtain second pose transformation parameters; determine the predicted pose of the object to be labeled in the other first-type images among the at least two first-type images based on the pose of the object to be labeled in one of the at least two first-type images and the second pose transformation parameters; determine the pose correction parameters based on the pose of the object to be labeled in the other first-type images among the at least two first-type images and the predicted pose; and use the pose correction parameters to correct the pose of the object to be labeled in the second-type image.

[0099] In one example, the position correction parameters include correction parameters per unit distance in the X, Y, and Z directions of the world coordinate system, which can be represented as (x, y, z). Based on the second attitude transformation parameters and the attitude error, the attitude correction parameters can be obtained. In another example, the attitude correction parameters include correction parameters per unit angle in the azimuth, pitch, and roll angles, which can be represented as (h, t, r). The pose correction parameters are represented by the position correction parameters and the attitude correction parameters.

[0100] For example, if the image acquisition device moves from the position of acquiring the first type of image A to the position of acquiring the first type of image B, and the vector of translation of the image acquisition device in the position is (a, b, c), then the second position transformation parameter in the second pose transformation parameter is (a, b, c). The angle of rotation of the image acquisition device in the attitude is composed of three angles: azimuth angle, roll angle, and pitch angle, which are represented as (d, e, f) respectively. Then the second attitude transformation parameter is (d, e, f).

[0101] Based on the predicted pose and the true pose of the object to be labeled in image B (assuming the pose determined by the markers is the true pose), the translation vector (X, Y, Z) from the predicted pose to the true pose can be calculated. The rotation angles from the predicted pose to the true pose are azimuth, roll, and pitch, denoted as (H, T, R). Then, the position correction parameters (x, y, z) = (X / a, Y / b, Z / c) and the attitude correction parameters (h, t, r) = (H / d, T / e, R / f) when the image acquisition device moves a unit distance can be calculated. When there are multiple other images of type I, the mean values ​​of the position correction parameters and the attitude correction parameters can be calculated, thereby improving the accuracy of the pose correction parameters.

[0102] After obtaining the pose of the object to be labeled in the second type of image based on the first type of image data, the pose of the object to be labeled in the second type of image can be corrected by combining the pose correction parameters and the pose transformation parameters of the image acquisition device when moving from the position of acquiring the first type of image data to the position of acquiring the second type of image data. For example, if the pose transformation parameters are a translation vector of (g, m, i) and rotation angles: azimuth, roll, and pitch angles are expressed as (j, k, l), then it is necessary to perform (g*x, m*y, i*z) correction on the position and (j*h, k*t, l*r) correction on the attitude.

[0103] In one possible implementation, the annotation information generation system further includes:

[0104] The image display module is used to display images acquired by the image acquisition device in real time, and to display the bounding boxes of the objects to be labeled according to their poses in the images. The images include a first type of image and a second type of image. In one example, when the bounding box of the object to be labeled in the image display module coincides with the object to be labeled, the corresponding second type of image and the corresponding bounding box in the second type of image are output as sample data for subsequent model training.

[0105] This application also provides a method for generating annotation information, see [link to relevant documentation]. Figure 3 ,include:

[0106] S301, acquire a first type of image and a second type of image acquired by the image acquisition device during movement, the movement parameters of the image acquisition device, and the posture of the image acquisition device when acquiring the first type of image and the second type of image, wherein the first type of image includes an object to be labeled and a first marker, the second type of image includes the object to be labeled but does not include the first marker, and the relative pose of the object to be labeled and the first marker in the world coordinate system remains unchanged.

[0107] The annotation information generation method in this application embodiment can be implemented by a computing device. In one example, the computing device can be a smartphone, tablet, dedicated embedded device, laptop or personal computer, etc.

[0108] S302, for each first type of image, determine the pose of the first marker in the first type of image.

[0109] The first marker has obvious visual features, and the pose of the first marker is detected in the first type of image using computer vision technology.

[0110] S303, for each first type of image, determine the pose of the object to be labeled in the first type of image based on the predetermined relative pose relationship between the first marker and the object to be labeled, and the pose of the first marker in the first type of image.

[0111] The first type of image includes a first marker. Therefore, the pose of the object to be labeled in the first type of image can be directly obtained from the pose of the first marker in the first type of image and the relative pose relationship between the first marker and the object to be labeled. The method of obtaining the pose of the object to be labeled in the first type of image using the pose of the first marker in the first type of image can be found in related technologies, and is not specifically limited in this application. In one example, the predetermined relative pose relationship between the first marker and the object to be labeled can be their relative pose relationship in the world coordinate system. Based on the pose of the first marker in the first type of image and its pose in the world coordinate system, the transformation relationship between the image coordinate system and the world coordinate system of the first type of image can be obtained. Thus, the relative pose relationship between the first marker and the object to be labeled in the world coordinate system can be converted into their relative pose relationship in the image coordinate system of the first type of image. Based on the pose of the first marker in the first type of image and the relative pose relationship between the first marker and the object to be labeled in the image coordinate system of the first type of image, the pose of the object to be labeled in the first type of image can be obtained.

[0112] S304, for each second type of image, determine the pose of the object to be labeled in the second type of image based on the movement parameters of the image acquisition device, the posture of the image acquisition device when acquiring the second type of image, and the pose of the object to be labeled in the first type of image.

[0113] A preset 3D tracking algorithm can be used to track the object to be labeled based on the movement parameters and posture of the image acquisition device and the pose of the object to be labeled in the first type of image, thereby obtaining the pose of the object to be labeled in the second type of image.

[0114] S305, Based on the pose of the object to be labeled in each of the second type of images, generate annotation information for the object to be labeled in each of the second type of images.

[0115] The annotation information of the object to be annotated can be the coordinates of the object to be annotated in the second type of image. In one example, the object to be annotated can be annotated by a rectangle, and the annotation information can be represented by the coordinates of the four vertices of the rectangle, or by the coordinates of the lower left corner vertex of the rectangle, the length and width of the rectangle.

[0116] In one possible implementation, the method further includes:

[0117] Step 1: Acquire a third type of image captured by an image acquisition device. The third type of image includes an object to be labeled with a second marker and a first marker. The second marker is set on a key point of the object to be labeled.

[0118] Step 2: For each third type of image, determine the relative pose relationship between the second marker and the first marker in that third type of image.

[0119] Step 3: For each third type of image, based on the relative pose relationship between the second marker and the first marker in the third type of image, determine the relative pose relationship between the first marker and the object to be labeled.

[0120] The second marker is placed on the key point of the object to be labeled. The third type of image and the pose of the third type of image are acquired by the image acquisition device. The relative pose relationship between the first marker and the second marker in the third type of image is detected. By combining the position of the second marker (key point) on the 3D model of the object to be labeled and the pose of the third type of image, the relative pose relationship between the first marker and the object to be labeled in the world coordinate system can be obtained and stored.

[0121] In one example, the coordinate transformation relationship between the image coordinate systems of the first type of image and the second type of image can be directly calculated; in one possible implementation, see [link to implementation details]. Figure 4 For each second type of image, determining the pose of the object to be labeled in the second type of image based on the movement parameters of the image acquisition device, the pose of the image acquisition device when acquiring the second type of image, and the pose of the object to be labeled in the first type of image includes:

[0122] S3041, for each second type of image, based on the movement parameters of the image acquisition device and the posture of the image acquisition device when acquiring the second type of image, determine the pose transformation parameters of the image acquisition device when acquiring the specified first type of image and when acquiring the second type of image, and obtain the first pose transformation parameter of the second type of image.

[0123] The pose transformation parameters include the pose change parameters and attitude change parameters of the world coordinate system acquired by the image acquisition device. Based on the movement parameters of the image acquisition device, the attitude when acquiring the first type of image, and the attitude when acquiring the second type of image, the pose transformation parameters of the image acquisition device from acquiring the first type of image to acquiring the second type of image can be obtained, which are called the first pose transformation parameters.

[0124] S3042, Based on the first pose transformation parameter of the second type of image, determine the image coordinate transformation relationship between the specified first type of image and the second type of image.

[0125] Based on the first pose transformation parameter and the pre-labeled camera extrinsic parameters, the transformation relationship between the image coordinates of the first type of image and the image coordinates of the second type of image can be obtained, which is called the image coordinate transformation relationship between the first type of image and the second type of image.

[0126] S3043, Based on the image coordinate transformation relationship of the second type of image and the pose of the object to be labeled in the specified first type of image, determine the pose of the object to be labeled in the second type of image.

[0127] Based on the pose of the object to be labeled in the first type of image, the pose of the object to be labeled in the second type of image is determined by using the image coordinate transformation relationship between the first type of image and the second type of image.

[0128] In one example, a transformation relationship between the image coordinate system and the world coordinate system of a first type of image can be established, as well as a transformation relationship between the image coordinate system and the world coordinate system of a second type of image, thereby indirectly obtaining the image coordinate transformation relationship between the first type of image and the second type of image; in one possible implementation, see... Figure 5 For each second type of image, determining the pose of the object to be labeled in the second type of image based on the movement parameters of the image acquisition device, the pose of the image acquisition device when acquiring the second type of image, and the pose of the object to be labeled in the first type of image includes:

[0129] S304a, for each second type of image, the first world coordinates when the image acquisition device acquires the specified first type of image and the second world coordinates when the second type of image is acquired are determined according to the movement parameters of the image acquisition device.

[0130] S304b, determine the third world coordinates of the object to be labeled based on the pose of the object to be labeled in the specified first type of image, the pose of the image acquisition device when acquiring the specified first type of image, and the first world coordinates.

[0131] S304c, Based on the pose of the object to be labeled in the second type of image when the image acquisition device acquires the second type of image, the second world coordinates, and the third world coordinates, determine the pose of the object to be labeled in the second type of image.

[0132] To further improve the accuracy of the annotation information of the object to be annotated, pose correction parameters can be calculated and used to correct the pose of the object to be annotated. In one possible implementation, the method further includes:

[0133] Step 1: Based on the movement parameters of the image acquisition device and the posture, determine the pose transformation parameters when the image acquisition device acquires at least two first-type images, and obtain the second pose transformation parameters.

[0134] Step 2: Based on the pose of the object to be labeled in one of the at least two first-class images and the second pose transformation parameters, determine the predicted pose of the object to be labeled in the other first-class images among the at least two first-class images.

[0135] Step 3: Determine the pose correction parameters based on the pose of the object to be labeled in the other first-class images among the at least two first-class images and the predicted pose.

[0136] Step four: Correct the pose of the object to be labeled in the second type of image using the pose correction parameters.

[0137] In one example, the position correction parameters include correction parameters per unit distance in the X, Y, and Z directions of the world coordinate system, which can be represented as (x, y, z). Based on the second attitude transformation parameters and the attitude error, the attitude correction parameters can be obtained. In another example, the attitude correction parameters include correction parameters per unit angle in the azimuth, pitch, and roll angles, which can be represented as (h, t, r). The pose correction parameters are represented by the position correction parameters and the attitude correction parameters.

[0138] For example, if the image acquisition device moves from the position of acquiring the first type of image A to the position of acquiring the first type of image B, and the vector of translation of the image acquisition device in the position is (a, b, c), then the second position transformation parameter in the second pose transformation parameter is (a, b, c). The angle of rotation of the image acquisition device in the attitude is composed of three angles: azimuth angle, roll angle, and pitch angle, which are represented as (d, e, f) respectively. Then the second attitude transformation parameter is (d, e, f).

[0139] Based on the predicted pose and the true pose of the object to be labeled in image B (assuming the pose determined by the markers is the true pose), the translation vector (X, Y, Z) from the predicted pose to the true pose can be calculated. The rotation angles from the predicted pose to the true pose are azimuth, roll, and pitch, denoted as (H, T, R). Then, the position correction parameters (x, y, z) = (X / a, Y / b, Z / c) and the attitude correction parameters (h, t, r) = (H / d, T / e, R / f) when the image acquisition device moves a unit distance can be calculated. When there are multiple other images of type I, the mean values ​​of the position correction parameters and the attitude correction parameters can be calculated, thereby improving the accuracy of the pose correction parameters.

[0140] After obtaining the pose of the object to be labeled in the second type of image based on the first type of image data, the pose of the object to be labeled in the second type of image can be corrected by combining the pose correction parameters and the pose transformation parameters of the image acquisition device when moving from the position of acquiring the first type of image data to the position of acquiring the second type of image data. For example, if the pose transformation parameters are a translation vector of (g, m, i) and rotation angles: azimuth, roll, and pitch angles are expressed as (j, k, l), then it is necessary to perform (g*x, m*y, i*z) correction on the position and (j*h, k*t, l*r) correction on the attitude.

[0141] This application also provides a labeling information generation device, see [link to relevant documentation]. Figure 6 ,include:

[0142] The first parameter acquisition module 601 is used to acquire a first type of image and a second type of image acquired by the image acquisition device during movement, the movement parameters of the image acquisition device, and the posture of the image acquisition device when acquiring the first type of image and the second type of image. The first type of image includes an object to be labeled and a first marker, the second type of image includes an object to be labeled but does not include the first marker, and the relative pose of the object to be labeled and the first marker in the world coordinate system remains unchanged.

[0143] The first pose determination module 602 is used to determine the pose of the first marker in each first type of image.

[0144] The second pose determination module 603 is used to determine the pose of the object to be labeled in the first type of image for each first type of image based on the pre-determined relative pose relationship between the first marker and the object to be labeled, and the pose of the first marker in the first type of image.

[0145] The third pose determination module 604 is used to determine the pose of the object to be labeled in the second type of image for each second type of image based on the movement parameters of the image acquisition device, the pose of the image acquisition device when acquiring the second type of image, and the pose of the object to be labeled in the first type of image.

[0146] The annotation information generation module 605 is used to generate annotation information for the objects to be annotated in each of the second type of images based on the pose of the objects to be annotated in each of the second type of images.

[0147] In one possible implementation, the device further includes:

[0148] The second parameter acquisition module is used to acquire a third type of image acquired by the image acquisition device. The third type of image includes an object to be labeled with a second marker and a first marker. The second marker is set on a key point of the object to be labeled.

[0149] The first pose relationship determination module is used to determine the relative pose relationship between the second marker and the first marker in each third type of image;

[0150] The second pose relationship determination module is used to determine the relative pose relationship between the first marker and the object to be labeled for each third type of image, based on the relative pose relationship between the second marker and the first marker in the third type of image.

[0151] In one possible implementation, the third pose determination module is specifically used to: for each second type of image, determine the pose transformation parameters of the image acquisition device when acquiring a specified first type of image and when acquiring the second type of image based on the movement parameters of the image acquisition device and the pose of the image acquisition device when acquiring the second type of image, and obtain the first pose transformation parameter of the second type of image;

[0152] Based on the first pose transformation parameter of the second type of image, determine the image coordinate transformation relationship between the specified first type of image and the second type of image;

[0153] Based on the image coordinate transformation relationship of the second type of image and the pose of the object to be labeled in the specified first type of image, the pose of the object to be labeled in the second type of image is determined.

[0154] In one possible implementation, the third pose determination module is specifically used to: for each second type of image, determine the first world coordinates when the image acquisition device acquires the specified first type of image and the second world coordinates when acquiring the second type of image, based on the movement parameters of the image acquisition device;

[0155] The third world coordinates of the object to be labeled are determined based on the pose of the object to be labeled in the specified first type of image, the pose of the image acquisition device when acquiring the specified first type of image, and the first world coordinates.

[0156] Based on the pose of the object to be labeled in the second type of image when the image acquisition device acquires the second type of image, the second world coordinates, and the third world coordinates, the pose of the object to be labeled in the second type of image is determined.

[0157] In one possible implementation, the device further includes:

[0158] The pose correction module is used to determine pose transformation parameters when the image acquisition device acquires at least two first-type images based on the movement parameters and the posture of the image acquisition device, thereby obtaining second pose transformation parameters; determine the predicted pose of the object to be labeled in the other first-type images among the at least two first-type images based on the pose of the object to be labeled in one of the at least two first-type images and the second pose transformation parameters; determine pose correction parameters based on the pose of the object to be labeled in the other first-type images among the at least two first-type images and the predicted pose; and correct the pose of the object to be labeled in the second-type images using the pose correction parameters.

[0159] This application also provides an electronic device, including: a processor and a memory;

[0160] The aforementioned memory is used to store computer programs;

[0161] When the processor executes the computer program stored in the memory, it implements any of the annotation information generation methods described in this application.

[0162] Optional, see Figure 7 The electronic device in this application embodiment also includes a communication interface 702 and a communication bus 704, wherein the processor 701, the communication interface 702, and the memory 703 communicate with each other through the communication bus 704.

[0163] The communication bus mentioned in the above electronic devices can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0164] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0165] The memory may include RAM (Random Access Memory) or NVM (Non-Volatile Memory), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0166] The processors mentioned above can be general-purpose processors, including CPUs (Central Processing Units), NPs (Network Processors), etc.; they can also be DSPs (Digital Signal Processors), ASICs (Application Specific Integrated Circuits), FPGAs (Field-Programmable Gate Arrays), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0167] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the annotation information generation methods described in this application.

[0168] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the annotation information generation methods described in the above embodiments.

[0169] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0170] It should be noted that, in this document, the technical features of the various alternative solutions can be combined to form solutions as long as they are not contradictory, and these solutions are all within the scope of this application. Relational terms such as "first" and "second" are merely used to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0171] The various embodiments in this specification are described in a related manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other.

[0172] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.

Claims

1. A method for generating annotation information, characterized in that, include: Acquire a first type of image and a second type of image acquired by an image acquisition device during movement, the movement parameters of the image acquisition device, and the posture of the image acquisition device when acquiring the first type of image and the second type of image, wherein the first type of image includes an object to be labeled and a first marker, the second type of image includes an object to be labeled but does not include the first marker, and the relative pose of the object to be labeled and the first marker in the world coordinate system remains unchanged; For each image of the first type, determine the pose of the first marker in that image; For each first type of image, the pose of the object to be labeled in the first type of image is determined based on the predetermined relative pose relationship between the first marker and the object to be labeled, and the pose of the first marker in the first type of image. For each second type of image, the pose of the object to be labeled in the second type of image is determined based on the movement parameters of the image acquisition device, the posture of the image acquisition device when acquiring the second type of image, and the pose of the object to be labeled in the first type of image. Based on the pose of the object to be labeled in each of the second type of images, annotation information of the object to be labeled in each of the second type of images is generated.

2. The method according to claim 1, characterized in that, The method further includes: Acquire a third type of image captured by an image acquisition device, wherein the third type of image includes an object to be labeled with a second marker and a first marker, the second marker being set on a key point of the object to be labeled; For each third-class image, determine the relative pose relationship between the second marker and the first marker in that third-class image; For each third-class image, based on the relative pose relationship between the second marker and the first marker in the third-class image, the relative pose relationship between the first marker and the object to be labeled is determined.

3. The method according to claim 1, characterized in that, For each second type of image, determining the pose of the object to be labeled in the second type of image based on the movement parameters of the image acquisition device, the posture of the image acquisition device when acquiring the second type of image, and the pose of the object to be labeled in the first type of image includes: For each second type of image, based on the movement parameters of the image acquisition device and the posture of the image acquisition device when acquiring the second type of image, the pose transformation parameters of the image acquisition device when acquiring the specified first type of image and when acquiring the second type of image are determined, and the first pose transformation parameters of the second type of image are obtained. Based on the first pose transformation parameter of the second type of image, determine the image coordinate transformation relationship between the specified first type of image and the second type of image; Based on the image coordinate transformation relationship of the second type of image and the pose of the object to be labeled in the specified first type of image, the pose of the object to be labeled in the second type of image is determined.

4. The method according to claim 1, characterized in that, For each second type of image, determining the pose of the object to be labeled in the second type of image based on the movement parameters of the image acquisition device, the posture of the image acquisition device when acquiring the second type of image, and the pose of the object to be labeled in the first type of image includes: For each second type of image, the first world coordinates when the image acquisition device acquires the specified first type of image and the second world coordinates when the image acquisition device acquires the second type of image are determined according to the movement parameters of the image acquisition device; The third world coordinates of the object to be labeled are determined based on the pose of the object to be labeled in the specified first type of image, the pose of the image acquisition device when acquiring the specified first type of image, and the first world coordinates. Based on the pose of the object to be labeled in the second type of image when the image acquisition device acquires the second type of image, the second world coordinates, and the third world coordinates, the pose of the object to be labeled in the second type of image is determined.

5. The method according to any one of claims 1-4, characterized in that, The method further includes: Based on the movement parameters of the image acquisition device and the posture, determine the pose transformation parameters of the image acquisition device when acquiring at least two first-type images, and obtain the second pose transformation parameters; Based on the pose of the object to be labeled in one of the at least two first-class images and the second pose transformation parameters, the predicted pose of the object to be labeled in the other first-class images in the at least two first-class images is determined. Based on the pose of the object to be labeled in other first-class images among the at least two first-class images and the predicted pose, determine the pose correction parameters; The pose correction parameters are used to correct the pose of the object to be labeled in the second type of image.

6. A system for generating annotation information, characterized in that, include: Image acquisition equipment, first marker generation equipment, computing equipment; The first marker generating device is used to generate a first marker at a specified location; The image acquisition device is used to acquire a first type of image and a second type of image of the object to be labeled during the movement, and to record the movement parameters and the posture when acquiring the first type of image and the second type of image. The first type of image includes a first marker, the second type of image does not include the first marker, and the relative pose of the object to be labeled and the first marker in the world coordinate system remains unchanged. The computing device is configured to: determine the pose of the first marker in each first-type image; determine the pose of the object to be labeled in each first-type image based on a predetermined relative pose relationship between the first marker and the object to be labeled, and the pose of the first marker in the first-type image; determine the pose of the object to be labeled in each second-type image based on the movement parameters of the image acquisition device, the posture of the image acquisition device when acquiring the second-type image, and the pose of the object to be labeled in the first-type image; and generate labeling information for the object to be labeled in each second-type image based on the pose of the object to be labeled in each second-type image.

7. The system according to claim 6, characterized in that, The system also includes: The second marker is placed on the key point of the object to be marked; The image acquisition device is also used to acquire a third type of image of the object to be labeled, wherein the third type of image includes the first marker; The computing device is further configured to determine the relative pose relationship between the second marker and the first marker in the third type of image; and based on the relative pose relationship between the second marker and the first marker in the third type of image, to determine the relative pose relationship between the first marker and the object to be labeled.

8. The system according to claim 6, characterized in that, The computing device is specifically configured to: for each second type of image, determine the pose transformation parameters of the image acquisition device when acquiring a specified first type of image and when acquiring the second type of image, based on the movement parameters of the image acquisition device and the pose of the image acquisition device when acquiring the second type of image, and obtain the first pose transformation parameter of the second type of image; determine the image coordinate transformation relationship between the specified first type of image and the second type of image based on the first pose transformation parameter of the second type of image; and determine the pose of the object to be labeled in the second type of image based on the image coordinate transformation relationship of the second type of image and the pose of the object to be labeled in the specified first type of image.

9. The system according to claim 6, characterized in that, The computing device is specifically configured to: for each second type of image, determine the first world coordinates when the image acquisition device acquires the specified first type of image and the second world coordinates when acquiring the second type of image, based on the movement parameters of the image acquisition device; determine the third world coordinates of the object to be labeled based on the pose of the object to be labeled in the specified first type of image, the pose of the image acquisition device when acquiring the specified first type of image, and the first world coordinates; and determine the pose of the object to be labeled in the second type of image based on the pose of the image acquisition device when acquiring the second type of image, the second world coordinates, and the third world coordinates.

10. The system according to any one of claims 6-9, characterized in that, The computing device is further configured to: determine pose transformation parameters for the image acquisition device when acquiring at least two first-type images based on the movement parameters of the image acquisition device and the pose, thereby obtaining second pose transformation parameters; determine the predicted pose of the object to be labeled in other first-type images among the at least two first-type images based on the pose of the object to be labeled in one of the at least two first-type images and the second pose transformation parameters; determine pose correction parameters based on the pose of the object to be labeled in other first-type images among the at least two first-type images and the predicted pose; and correct the pose of the object to be labeled in the second-type image using the pose correction parameters.

11. A device for generating annotation information, characterized in that, include: The first parameter acquisition module is used to acquire a first type of image and a second type of image acquired by the image acquisition device during movement, the movement parameters of the image acquisition device, and the posture of the image acquisition device when acquiring the first type of image and the second type of image. The first type of image includes an object to be labeled and a first marker, the second type of image includes an object to be labeled but does not include the first marker, and the relative pose of the object to be labeled and the first marker in the world coordinate system remains unchanged. The first pose determination module is used to determine the pose of the first marker in each first type of image. The second pose determination module is used to determine the pose of the object to be labeled in the first type of image for each first type of image based on the pre-determined relative pose relationship between the first marker and the object to be labeled, and the pose of the first marker in the first type of image. The third pose determination module is used to determine the pose of the object to be labeled in the second type of image for each second type of image based on the movement parameters of the image acquisition device, the pose of the image acquisition device when acquiring the second type of image, and the pose of the object to be labeled in the first type of image. The annotation information generation module is used to generate annotation information for the objects to be annotated in each of the second type of images based on their poses.

12. The apparatus according to claim 11, characterized in that, The device further includes: The second parameter acquisition module is used to acquire a third type of image acquired by the image acquisition device. The third type of image includes an object to be labeled with a second marker and a first marker. The second marker is set on a key point of the object to be labeled. The first pose relationship determination module is used to determine the relative pose relationship between the second marker and the first marker in each third type of image; The second pose relationship determination module is used to determine the relative pose relationship between the first marker and the object to be labeled for each third type of image, based on the relative pose relationship between the second marker and the first marker in the third type of image. The pose correction module is used to determine pose transformation parameters when the image acquisition device acquires at least two first-type images based on the movement parameters and the posture of the image acquisition device, thereby obtaining second pose transformation parameters; determine the predicted pose of the object to be labeled in the other first-type images among the at least two first-type images based on the pose of the object to be labeled in one of the at least two first-type images and the second pose transformation parameters; determine pose correction parameters based on the pose of the object to be labeled in the other first-type images among the at least two first-type images and the predicted pose; and correct the pose of the object to be labeled in the second-type image using the pose correction parameters. The third pose determination module is specifically used, for each second type of image, to determine, based on the movement parameters of the image acquisition device and the posture of the image acquisition device when acquiring the second type of image, the pose transformation parameters of the image acquisition device when acquiring the specified first type of image and when acquiring the second type of image, to obtain the first pose transformation parameter of the second type of image; to determine the image coordinate transformation relationship between the specified first type of image and the second type of image based on the first pose transformation parameter of the second type of image; and to determine the pose of the object to be labeled in the second type of image based on the image coordinate transformation relationship of the second type of image and the pose of the object to be labeled in the specified first type of image; or The third pose determination module is specifically used for each second type of image to determine, based on the movement parameters of the image acquisition device, the first world coordinates when the image acquisition device acquires the specified first type of image and the second world coordinates when acquiring the second type of image; to determine the third world coordinates of the object to be labeled based on the pose of the object to be labeled in the specified first type of image, the pose of the image acquisition device when acquiring the specified first type of image, and the first world coordinates; and to determine the pose of the object to be labeled in the second type of image based on the pose of the image acquisition device when acquiring the second type of image, the second world coordinates, and the third world coordinates.

13. An electronic device, characterized in that, Including processor and memory; The memory is used to store computer programs; When the processor executes the program stored in the memory, it implements the annotation information generation method according to any one of claims 1-5.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the annotation information generation method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Image labeling method, device and system and host

    CN111127422A

  • Planar marker pose determination method and device

    CN112215884A