Image labeling method and device, equipment and storage medium

By displaying the 3D scene corresponding to the 2D image on the image annotation interface, receiving and responding to object annotation instructions, and converting the position annotations in the 3D scene into position annotations in the 2D image, the problem of difficult object annotation on perspective view is solved, and accurate and convenient image annotation is achieved.

CN121962546APending Publication Date: 2026-05-01GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing image annotation methods struggle to accurately annotate targets that are far from the camera because they appear small in perspective views, making it difficult for annotators to identify and label them.

Method used

By displaying the target 2D image to be annotated and its corresponding 3D scene on the image annotation interface, receiving and responding to object annotation instructions, annotating the position of the target object in the 3D scene, and converting it into position annotation in the 2D image.

Benefits of technology

It enables accurate and convenient target object annotation in 2D images, avoiding visual errors such as objects appearing larger when closer and smaller when farther away. Furthermore, the accuracy of the annotation is improved through verification and fine-tuning in 3D scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962546A_ABST
    Figure CN121962546A_ABST
Patent Text Reader

Abstract

The invention provides an image annotation method and device, equipment and a storage medium, and the method comprises the steps: displaying a to-be-annotated target two-dimensional image and a three-dimensional scene corresponding to the target two-dimensional image on an image annotation interface, and enabling the three-dimensional scene to be obtained based on the three-dimensional scene reconstruction of the target two-dimensional image; an object labeling instruction acting on a target object in the three-dimensional scene is received, the target object is an object to be labeled, the object labeling instruction is associated with a first scene position label, and the first scene position label is a position label of the target object in the three-dimensional scene; in response to the object annotation instruction, adding and displaying a first image position annotation obtained by converting the first scene position annotation in the target two-dimensional image on the image annotation interface, the first image position annotation being a position annotation of the target object in the target two-dimensional image. According to the technical scheme, the convenience and accuracy of image annotation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Image annotation methods, apparatus, devices and storage media Technical Field

[0001] This application relates to the field of image annotation, and more particularly to image annotation methods, apparatus, devices, and storage media. Background Technology

[0002] Image annotation is an important component of computer vision technology. It can provide training data for machine learning algorithms, improve the accuracy of image recognition and classification, facilitate image retrieval and analysis, and drive the development of computer vision technology.

[0003] Currently, image annotation typically involves annotators identifying the target object from a perspective drawing and then labeling it on the perspective view. However, due to perspective effects, the same object will appear larger when closer and smaller when farther away from the camera. Objects farther from the camera will appear smaller in the perspective view, making them difficult for annotators to identify and label. Summary of the Invention

[0004] This application provides image annotation methods, apparatus, devices, and storage media to solve the technical problem that targets far from the camera appear small in the image, making them difficult for annotators to identify and annotate.

[0005] Firstly, an image annotation method is provided, including:

[0006] The image annotation interface displays the target 2D image to be annotated and the corresponding 3D scene, which is obtained by reconstructing the 3D scene from the target 2D image.

[0007] Receive an object annotation instruction applied to a target object in the three-dimensional scene, wherein the target object is an object to be annotated, and the object annotation instruction is associated with a first scene position annotation, wherein the first scene position annotation is the position annotation of the target object in the three-dimensional scene;

[0008] In response to the object annotation instruction, a first image location annotation converted from the first scene location annotation is added and displayed in the target two-dimensional image on the image annotation interface. The first image location annotation is the location annotation of the target object in the target two-dimensional image.

[0009] In this technical solution, the target 2D image to be annotated and a 3D scene obtained by 3D reconstruction of the target 2D image are displayed on the image annotation interface. Then, an object annotation instruction is received and applied to the target object to be annotated in the 3D scene. Finally, in response to the object annotation instruction, a first image position annotation, converted from the first scene position annotation associated with the object annotation instruction, is added and displayed in the target 2D image on the image annotation interface. The first image position annotation is the position annotation of the target object in the target 2D image, thus realizing the annotation of the object to be annotated in the 2D image, i.e., image annotation. Since the first scene position annotation is the position annotation of the target object in the 3D scene, and the first image position annotation is the position annotation of the target object in the target 2D image, it is equivalent to annotating the position annotation of the object to be annotated in the 2D image by annotating the object to be annotated in the 3D scene. The visual presentation of the object to be annotated in the 3D scene is consistent with its visual presentation in the real world, which avoids near-large image distortion. In a 3D scene, users can observe the object to be labeled from different angles, avoiding occlusion. This allows for accurate and convenient labeling of the object's location within the 3D scene, ensuring accurate first-scene labeling. Since the first-image labeling is derived from the first-scene labeling, it is also accurate. Therefore, it enables accurate and convenient labeling of the object's location in the 2D image, improving the convenience of image labeling. Furthermore, by simultaneously displaying the target 2D image and the 3D scene reconstructed from the target 2D image on the image labeling interface, users can verify the accuracy of the first-scene labeling by comparing its overlap with the target object in the target 2D image. This allows for fine-tuning of the labeling, resulting in more accurate labeling and thus improving the overall accuracy of image labeling.

[0010] In conjunction with the first aspect, in one possible implementation, after receiving the object annotation instruction applied to the target object in the three-dimensional scene, the method further includes: responding to the object annotation instruction, obtaining the pose information of the target object from the three-dimensional scene to obtain target pose information; and annotating and displaying the target pose information on the image annotation interface.

[0011] After receiving the object annotation instruction for the target object in the 3D scene, it also responds to the object annotation instruction, obtains the pose information of the target object from the 3D scene, and then annotates and displays the pose information of the target object on the image annotation interface, enriching the annotation content of the object to be annotated, thus making the image annotation data richer.

[0012] In conjunction with the first aspect, in one possible implementation, adding and displaying the first image location marker, obtained by converting the first scene location marker, in the target two-dimensional image on the image annotation interface includes: performing coordinate mapping transformation on the first location coordinates corresponding to the first scene location marker to obtain second location coordinates, wherein the first location coordinates are location coordinates in the three-dimensional scene and the second location coordinates are location coordinates in the two-dimensional image; superimposing an object location marker representing the target object onto the second location coordinates in the target two-dimensional image to obtain the first image location marker; and displaying the first image location marker in the target two-dimensional image on the image annotation interface.

[0013] By mapping the position coordinates of the target object's position annotation in the 3D scene to the position coordinates in the 2D image, a second position coordinate is obtained. The object position marker used to represent the object to be annotated is superimposed on the second position coordinate of the target 2D image, which can realize the conversion of position annotation and ensure the accuracy of the target object's position annotation in the target 2D image.

[0014] In conjunction with the first aspect, in one possible implementation, displaying the three-dimensional scene corresponding to the target two-dimensional image on the image annotation interface includes: reconstructing the three-dimensional scene from the target two-dimensional image to obtain a virtual three-dimensional scene corresponding to the target two-dimensional image; acquiring a planar image of the virtual three-dimensional scene from the target's viewing perspective to obtain a planar view of the target three-dimensional scene; and displaying the planar view of the target three-dimensional scene on the image annotation interface.

[0015] By reconstructing a 3D scene from a 2D image, a virtual 3D scene is obtained, and a planar image of the virtual 3D scene from a viewing angle is acquired to obtain a 3D scene planar diagram. The 3D scene planar diagram is then displayed on the image annotation interface, thus realizing the display of a 3D scene corresponding to a 2D image.

[0016] In conjunction with the first aspect, in one possible implementation, the step of reconstructing a three-dimensional scene from the target two-dimensional image to obtain a virtual three-dimensional scene corresponding to the target two-dimensional image includes: transforming the target two-dimensional image from a pixel coordinate system to a camera coordinate system based on the camera intrinsic parameters of the target camera to obtain a target projection image, wherein the target camera is the camera that captured the target two-dimensional image; and projecting the target projection image onto a virtual three-dimensional space based on the camera extrinsic parameters of the target camera to obtain the virtual three-dimensional scene.

[0017] Based on the intrinsic parameters of the camera, the two-dimensional image is first transformed from the pixel coordinate system to the camera coordinate system to obtain the projected image in the camera coordinate system. Then, based on the extrinsic parameters of the camera, the projected image is projected onto the virtual three-dimensional space, which can realize the reconstruction of the three-dimensional scene.

[0018] In conjunction with the first aspect, in one possible implementation, before obtaining a planar image of the virtual 3D scene from the target viewing perspective to obtain a planar view of the target 3D scene, the method includes: obtaining a view setting instruction for the virtual 3D scene; and setting the target viewing perspective according to the view setting instruction.

[0019] Adjusting the viewpoint according to the virtual 3D scene's perspective settings allows users to easily observe the object to be labeled from various angles, which is beneficial for accurate labeling.

[0020] In conjunction with the first aspect, in one possible implementation, after adding and displaying the first image position annotation converted from the first scene position annotation in the target two-dimensional image on the image annotation interface, the method further includes: receiving an annotation adjustment instruction acting on the first scene position annotation, the annotation adjustment instruction including a second scene position annotation, the second scene position annotation being the position annotation of the target object in the three-dimensional scene; responding to the annotation adjustment instruction, adjusting the first image position annotation in the target two-dimensional image on the image annotation interface to the second image position annotation converted from the second scene position annotation, and displaying the second image position annotation, the second image position annotation being the position annotation of the target object in the target two-dimensional image.

[0021] After displaying the position annotations of the object to be annotated on the 2D image, the system receives annotation adjustment instructions that apply to the position annotations of the object in the 3D scene, and responds to the annotation adjustment instructions by adjusting the position annotations on the 2D image to the image position annotations corresponding to the scene position annotations in the annotation adjustment instructions. This allows users to easily fine-tune the position annotations in the 2D image, thereby making the position annotations in the 2D image more accurate.

[0022] In conjunction with the first aspect, in one possible implementation, after displaying the three-dimensional scene corresponding to the target two-dimensional image on the image annotation interface, the method further includes: receiving a scene adjustment instruction acting on the three-dimensional scene, the scene adjustment instruction including a first camera extrinsic parameter, the first camera extrinsic parameter being a camera extrinsic parameter of a target camera, the target camera being a camera that captured the target two-dimensional image; and responding to the scene adjustment instruction by displaying the three-dimensional scene corresponding to the first camera extrinsic parameter on the image annotation interface.

[0023] By receiving and responding to scene adjustment commands applied to a 3D scene, and displaying the 3D scene corresponding to the camera extrinsic parameters in the scene adjustment command on the image annotation interface, the reconstructed 3D scene can be made more accurate.

[0024] In conjunction with the first aspect, in one possible implementation, the target two-dimensional image is a frame of a target video, and the target video includes multiple frames of two-dimensional images; the method further includes: adding image position annotations to the multiple frames of two-dimensional images respectively to obtain the position annotations of the target object in the multiple frames of two-dimensional images, wherein the image position annotations are the position annotations of the target object in the two-dimensional images, and the image position annotations are obtained based on the position annotations of the target object in the three-dimensional scene.

[0025] When the two-dimensional image to be labeled is an image in a video, image position labels converted from position labels in the three-dimensional scene are added to the multiple frames of two-dimensional images contained in the video. This gives the position label of the object to be labeled in the multiple frames of two-dimensional images contained in the video, which can save the step of labeling frame by frame and improve the labeling speed.

[0026] In conjunction with the first aspect, in one possible implementation, the target object is a dynamic object; the step of adding image position annotations to the multiple frames of two-dimensional images to obtain the position annotations of the target object in the multiple frames of two-dimensional images includes: determining a first target two-dimensional image and a second target two-dimensional image in the multiple frames of two-dimensional images, wherein both the first target two-dimensional image and the second target two-dimensional image are target two-dimensional images in the target video, and both the first target two-dimensional image and the second target two-dimensional image are motion keyframes in the target video; interpolating the position annotations of the target object in the first target two-dimensional image and the position annotations of the target object in the second target two-dimensional image to obtain... The interpolated position annotation is obtained by converting the position annotation of the target object in the first target 2D image into the position annotation of the target object in the corresponding 3D scene of the first target 2D image, and the position annotation of the target object in the second target 2D image into the position annotation of the target object in the corresponding 3D scene of the second target 2D image. The interpolated position annotation is added to the intermediate 2D image between the first target 2D image and the second target 2D image to obtain the position annotation of the target object in the intermediate 2D image. The intermediate 2D image is a 2D image whose timestamp is located between the timestamp of the first target 2D image and the timestamp of the second target 2D image.

[0027] When the object to be labeled is a dynamic object, the position of the object in the image will change. By interpolating the dynamically changing key images to infer and label the position of the object in the image, the occlusion problem in the video image can be overcome. Only a part of the images in the video need to be labeled to achieve the labeling of all images in the video.

[0028] Secondly, an image annotation apparatus is provided, comprising:

[0029] The interface display module is used to display the target two-dimensional image to be annotated and the corresponding three-dimensional scene on the image annotation interface, wherein the three-dimensional scene is obtained by reconstructing the target two-dimensional image into a three-dimensional scene;

[0030] The instruction receiving module is used to receive an object annotation instruction applied to a target object in the three-dimensional scene. The target object is an object to be annotated. The object annotation instruction is associated with a first scene position annotation, which is the position annotation of the target object in the three-dimensional scene.

[0031] The annotation display module is used to respond to the object annotation instruction, add and display the first image position annotation converted from the first scene position annotation in the target two-dimensional image on the image annotation interface, wherein the first image position annotation is the position annotation of the target object in the target two-dimensional image.

[0032] Thirdly, a computer device is provided, including a memory, a display panel, and a processor, wherein the memory and the display panel are connected to the processor, the display panel is used to display an image annotation interface, and the processor is used to execute one or more computer programs stored in the memory, wherein when the processor executes one or more computer programs, the computer device implements the image annotation method of the first aspect described above.

[0033] Fourthly, a computer-readable storage medium is provided, which stores a computer program, the computer program including program instructions, which, when executed by a processor, cause the processor to perform the image annotation method of the first aspect.

[0034] This application achieves the following technical effects: it realizes the annotation of the object to be annotated in a two-dimensional image, that is, it realizes image annotation; since the first scene position annotation is the position annotation of the target object in the three-dimensional scene, and the first image position annotation is the position annotation of the target object in the target two-dimensional image, it is equivalent to annotating the position annotation of the object to be annotated in the two-dimensional image by annotating the position of the object to be annotated in the three-dimensional scene. The visual presentation of the object to be annotated in the three-dimensional scene is consistent with its visual presentation in the real world, which can avoid the phenomenon of objects appearing larger when closer and smaller when farther away. In the three-dimensional scene, users can observe the object to be annotated in the three-dimensional scene from different angles, which can avoid the object to be annotated being occluded. Thus, users can accurately and conveniently annotate the object to be annotated in the three-dimensional scene. In 3D scene annotation, if the first scene position annotation is accurate, and the first image position annotation is converted from the first scene position annotation, then the first image position annotation is also accurate. Therefore, it is possible to accurately and conveniently annotate the position of the object to be annotated in the 2D image, improving the convenience of image annotation. In addition, by simultaneously displaying the target 2D image to be annotated and the 3D scene reconstructed from the target 2D image on the image annotation interface, after the user annotates the first scene position annotation in the 3D scene, they can also check the accuracy of the position annotation by comparing the overlap between the first image position annotation in the target 2D image and the object to be annotated in the target 2D image. This makes it easier for the user to fine-tune the position annotation, making the position annotation more accurate, thus improving the accuracy of image annotation. Attached Figure Description

[0035] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0036] Figures 1 and 2 are schematic diagrams of teaching video images provided in the embodiments of this application;

[0037] Figure 3 is a flowchart illustrating an image annotation method provided in an embodiment of this application;

[0038] Figures 4A-4H are schematic diagrams of the image annotation interface provided in the embodiments of this application;

[0039] Figure 5 is a flowchart illustrating another image annotation method provided in an embodiment of this application;

[0040] Figure 6 is a flowchart illustrating another image annotation method provided in an embodiment of this application;

[0041] Figure 7 is a schematic diagram of an image annotation device provided in an embodiment of this application;

[0042] Figure 8 is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0043] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0044] It should be noted that, unless there is a conflict, the various features in the embodiments of this application can be combined with each other, all of which are within the protection scope of this application. Furthermore, although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than the module division in the device or the order in the flowchart. Moreover, the terms "first," "second," and "third" used in this application do not limit the data or execution order, but only distinguish identical or similar items with essentially the same function and effect.

[0045] The technical solution of this application is applicable to image annotation scenarios. In image annotation scenarios, annotators need to use annotation tools to mark the target objects to be annotated on the image, and the annotated image can be used for deep learning.

[0046] A typical scenario for image annotation is the annotation of classroom teaching videos. To accurately estimate the 3D layout of a classroom in a teaching video, annotators need to label target objects such as desks, the teacher's podium, and aisles in each frame of the teaching video used for deep learning. The conventional annotation method involves annotators identifying and labeling each target object in each frame of the teaching video frame one by one. Teaching video images are perspective views. Due to the varying distances of objects from the camera, objects appear larger in perspective, while those farther away appear smaller. See Figure 1 for details. The teacher's podium and desks near it appear as near-lined planes in the teaching video image, making them difficult for annotators to identify and label. Furthermore, in addition to desks, the podium, and aisles, the teaching video image also includes people (teachers and / or students). These people obscure these objects, as shown in Figure 2, where the desks at the back of the classroom are almost completely obscured by the students in front, further complicating identification and labeling. Therefore, conventional annotation methods are unsuitable for annotating objects in images with many or complex content (such as teaching video images).

[0047] Some solutions propose first converting the perspective view into a top-down view (i.e., a top-down view), showing this view to the annotator, who then identifies and labels the target objects in the top-down view. The annotations from the top-down view are then converted back into the perspective view, thus completing the labeling of the target objects. In a top-down view, users can observe objects in the image from above. This solution is suitable when all objects to be labeled are on the same horizontal plane, such as labeling road areas. However, a top-down view is still a two-dimensional view and still exhibits the principle of objects appearing larger when closer and smaller when farther away. When the objects to be labeled are not all on the same horizontal plane—for example, when the objects to be labeled include desks and other objects on the ground—taller objects (such as desks) can obstruct shorter objects (such as objects on the ground shorter than desks). Therefore, converting the perspective view to a top-down view for labeling does not completely solve the problem.

[0048] In view of this, this application proposes a novel image annotation scheme. First, a 3D scene is obtained by reconstructing a 2D perspective view. The 2D perspective image and its corresponding 3D scene are then displayed to the annotator. The annotator identifies and annotates the target object within the 3D scene corresponding to the 2D perspective image. The annotations in the 3D scene are then converted into annotations in the 2D perspective image, and the converted annotations are displayed to the annotator. Because the annotator identifies and annotates the target object within the 3D scene, the visual representation of the object in the 3D scene is consistent with its visual representation in the real world, avoiding the phenomenon of objects appearing larger when closer and smaller when farther away. The annotator can observe and identify the target object from different angles in the 3D scene, avoiding occlusion and reducing the difficulty of identification and annotation, thus making it easier to identify and annotate the target object. Furthermore, by displaying the annotations in the converted 2D perspective image to the user, the user can check the accuracy of the 3D scene annotations based on the 2D perspective image annotations, and then fine-tune the 3D scene annotations in a timely manner to adjust the annotations in the 2D perspective image, thereby making the annotations in the 2D perspective image more accurate.

[0049] The technical solution of this application is described in detail below. The technical solution of this application is applicable to any computer device with display functionality. Computer devices with display functionality include, but are not limited to, desktop computers, laptops, etc.

[0050] Referring to Figure 3, which is a flowchart illustrating an image annotation method provided in an embodiment of this application, the method includes the following steps:

[0051] S101, on the image annotation interface, the target 2D image to be annotated and the corresponding 3D scene are displayed.

[0052] Here, the image annotation interface is a user-oriented display interface for users to annotate images. The image annotation interface may display various tools, toolbars, etc., for users to use for image annotation. For example, the image annotation interface can be as shown in Figure 4A.

[0053] The target 2D image can be any image to be labeled, such as an image from a teaching video. The target 2D image to be labeled can be acquired and displayed in the labeling display area of ​​the image labeling interface, thus displaying the target 2D image to be labeled. The labeling display area of ​​the image labeling interface is used to display the 2D image and labeling information. For example, displaying the target 2D image to be labeled in the image labeling interface can be as shown in Figure 4B, where area Q1 in Figure 4B is the labeling display area, and P1 in Figure 4B is the target 2D image to be labeled.

[0054] The 3D scene corresponding to the target 2D image is obtained based on the 3D reconstruction of the target 2D image. Specifically, after acquiring the target 2D image to be labeled, a 3D scene reconstruction can be performed on the target 2D image to obtain the corresponding 3D scene. This 3D scene is then displayed in the annotation interaction area of ​​the image annotation interface, realizing the display of the 3D scene corresponding to the target 2D image. The annotation interaction area of ​​the image annotation interface is used to display the 3D scene and to interact with the user, obtaining user commands. For example, displaying the 3D scene corresponding to the target 2D image in the image annotation interface can be as shown in Figure 4B. Area Q2 in Figure 4B is the annotation interaction area, and P2 in Figure 4B is the 3D scene corresponding to the target 2D image P1.

[0055] The specific implementation method for displaying the 3D scene corresponding to the target 2D image to be annotated on the image annotation interface will be described in detail in subsequent steps A1-A3, and will not be described in detail here.

[0056] S102, receive object annotation instructions applied to the target object in the three-dimensional scene corresponding to the target two-dimensional image.

[0057] Here, the target object is the object to be labeled, which refers to the object that the user needs to identify and label from the image. The target object is related to the specific image labeling task; different image labeling tasks result in different target objects. For example, if the image labeling task is to label static objects in a classroom space, the target object could be a desk, podium, or aisle in a teaching video image; or if the image labeling task is to label cars on a road, the target object would be the cars in a road image; and this is not an exhaustive list of examples.

[0058] Target objects can be identified and selected by the user within the 3D scene corresponding to the target 2D image. Object annotation instructions applied to target objects within the 3D scene corresponding to the target 2D image are instructions triggered by the user's identification and annotation of target objects within the 3D scene. Specifically, the user can execute annotation operations on target objects in the annotation interaction area of ​​the image annotation interface to trigger these object annotation instructions. The annotation operations performed by the user in the annotation interaction area of ​​the image annotation interface can be acquired, thereby receiving object annotation instructions applied to target objects within the 3D scene corresponding to the target 2D image. The annotation operations performed by the user in the annotation interaction area of ​​the image annotation interface refer to the user's selection and annotation of target objects within that area.

[0059] The object annotation instruction for the target object is associated with a first scene location annotation. The first scene location annotation is the position of the target object in the 3D scene corresponding to the target 2D image. This first scene location annotation is generated based on the annotation operation performed by the user in the annotation interaction area of ​​the image annotation interface; that is, it is obtained based on the object annotation instruction. The first scene location annotation is used to represent the position of the target object in the 3D scene corresponding to the target 2D image. After receiving the object annotation instruction applied to the target object in the 3D scene corresponding to the target 2D image, the first scene location annotation can be displayed in the 3D scene corresponding to the target 2D image on the image annotation interface.

[0060] For example, as shown in Figure 4B, the image annotation interface displays a 3D scene P2 in the annotation interaction area Q2 of the image annotation interface. Referring to Figure 4C, when the user clicks the button b1 for selecting the target object on the image annotation interface and uses the selection tool corresponding to button b1 to select the desk plane of desk d1 in Figure 4, the desk becomes the target object selected by the user. After the user clicks "confirm" in the pop-up window that prompts whether to annotate, the user receives the object annotation instruction applied to desk d1, as shown in Figure 4D. In the 3D scene P2 of the image annotation interface, the position annotation z1 of desk d1 is displayed. The position annotation z1 is the first scene position annotation.

[0061] The first scene location annotation includes multiple position coordinates of the target object in the 3D scene corresponding to the target 2D image. The number of position coordinates included in the first scene location annotation depends on the rendering style of the first scene location annotation in the 3D scene corresponding to the target 2D image. Taking the target object as desk d1 in Figure 4D as an example, assuming that the position annotation of desk d1 in the 3D scene is as shown by z1 in Figure 4D, the first scene location annotation can include the position coordinates of the four corners of the quadrilateral model shown by z1 in the 3D scene shown by P2, that is, the first scene location annotation includes 4 position coordinates.

[0062] S103, in response to the object annotation instruction applied to the target object in the three-dimensional scene corresponding to the target two-dimensional image, add and display the first image position annotation converted from the first scene position annotation in the target two-dimensional image on the image annotation interface.

[0063] Here, the first image position label is the position label of the target object in the target two-dimensional image. For example, if the target object is desk d1 in Figure 4D, then the first image position label is the position label of desk d1 in the target two-dimensional image shown by P1 in Figure 4D.

[0064] In one feasible implementation, the first image location annotation can be added and displayed in the target two-dimensional image on the image annotation interface through the following steps B1-B3:

[0065] B1. Perform coordinate mapping transformation on the first position coordinates corresponding to the first scene position annotation to obtain the second position coordinates.

[0066] Here, the first position coordinates are the position coordinates in the 3D scene corresponding to the target 2D image. The first position coordinates include multiple position coordinates of the target object in the 3D scene corresponding to the target 2D image. For example, if the target object is the desk d1 in Figure 4D, and the first scene position label is shown as z1 in Figure 4D, then the first position coordinates include the position coordinates of the four corners of the quadrilateral model shown in z1 in Figure 4D in the 3D scene shown in P2. The first position coordinates can be obtained by acquiring the position coordinates of the first scene position label in the 3D scene corresponding to the target 2D image.

[0067] The second position coordinates are the position coordinates in the two-dimensional image. The second position coordinates are obtained by performing a coordinate mapping transformation on the first position coordinates corresponding to the first scene position annotation. This means transforming the first position coordinates from the three-dimensional scene coordinate system to the pixel coordinate system. The three-dimensional scene coordinate system refers to the spatial coordinate system in the three-dimensional scene corresponding to the target two-dimensional image.

[0068] Since the three-dimensional scene corresponding to the target two-dimensional image is reconstructed from the target two-dimensional image, during the process of reconstructing the target two-dimensional image to obtain the three-dimensional scene corresponding to the target two-dimensional image, a mapping transformation relationship is established between the pixel coordinate system corresponding to the target two-dimensional image and the three-dimensional scene coordinate system corresponding to the target two-dimensional image. According to this mapping transformation relationship, the first position coordinates can be transformed from the three-dimensional scene coordinate system to the two-dimensional pixel coordinate system to obtain the second position coordinates.

[0069] B2. The object position marker used to represent the target object is superimposed on the second position coordinate in the target two-dimensional image to obtain the first image position marker.

[0070] Here, the object location marker used to represent the target object refers to a marker used to represent the position of the target object in the target 2D image. This object location marker represents the presentation style of the target object in the target 2D image. The presentation style of the target object in the 2D image can be the same as the presentation style of the target object in the corresponding 3D scene. For example, if the target object is the desk d1 in Figure 4D, and the presentation style of the target object is a quadrilateral model as shown by z1 in Figure 4D, the position coordinates of the four corners of the quadrilateral model shown by z1 in the 3D scene can be converted into position coordinates in the pixel coordinate system to obtain the second position coordinates. At the second position coordinates of the target 2D image, a quadrilateral model with the same style as z1 is added. This quadrilateral model with the same style as z1 is the object location marker, thus obtaining z2 in Figure 4D. z2 is the first image position marker converted from the first scene position marker z1. The presentation style of the target object in the 2D image can also be different from the presentation style of the target object in the corresponding 3D scene. This application does not impose any restrictions.

[0071] B3. Display the first image position annotation in the target two-dimensional image on the image annotation interface.

[0072] For example, in the target two-dimensional image on the image annotation interface, the first image position annotation can be displayed as shown in z2 of Figure 4D.

[0073] By mapping the position coordinates of the target object's position annotation in the 3D scene to the position coordinates in the 2D image, a second position coordinate is obtained. The object position marker used to represent the object to be annotated is superimposed on the second position coordinate of the target 2D image, which can realize the conversion of position annotation and ensure the accuracy of the target object's position annotation in the target 2D image.

[0074] In some possible cases, the target two-dimensional image involved in steps S101-S103 above can be a single frame of a target video, the target video includes multiple frames of two-dimensional images, and the target video can be images obtained from long-term recording of the same spatial scene; the method may further include: adding image position annotations to the multiple frames of two-dimensional images contained in the target video to obtain the position annotations of the target object in the multiple frames of two-dimensional images contained in the target video. Here, the image position annotations are the position annotations of the target object in the two-dimensional images, and the image position annotations added to the multiple two-dimensional images are all obtained by transforming the position annotations of the target object in the three-dimensional scene.

[0075] When the target object is a static object (meaning an object whose position in the video image does not change over time), since the position of a static object in the video image will not change, one frame of the two-dimensional image in the target video can be used as the target two-dimensional image. Steps S101-S103 above are executed to obtain the first image position label; then the first image position label is added to each frame of the two-dimensional image in the target video.

[0076] Referring to the principles described in steps B1-B2 above, in each frame of the target video's two-dimensional image, the object position marker used to represent the target object is superimposed onto the second position coordinate of each frame of the target video's two-dimensional image to obtain the position label of the target object in each frame of the target video's two-dimensional image.

[0077] For example, as shown in the target two-dimensional image P1 in Figure 4D, P1 is a frame of a teaching video image, and the target object is desk d1. A quadrilateral model as shown in z2 in Figure 4D can be added to each frame of the teaching video image to obtain the position label of desk d1 in each frame of the teaching video image.

[0078] When the target object is a dynamic object (referring to an object whose position in a video image changes over time), since the position of the dynamic object in the video image changes, motion keyframes can be determined from the target video. Motion keyframes are video frames in the target video that can provide diverse information. Using the motion keyframes as the target two-dimensional image, the above steps S101-S103 are executed to obtain the position annotation of the target object in the motion keyframes.

[0079] There are multiple motion keyframes in the target video. Two adjacent motion keyframes in the target video can be used as the first target 2D image and the second target 2D image. Steps S101-S103 are executed respectively to obtain the position annotation of the target object in the first target 2D image and the position annotation of the target object in the second target 2D image. The position annotation of the target object in the first target 2D image and the position annotation of the target object in the second target 2D image are interpolated to obtain the interpolated position annotation. The interpolated position annotation is added to the intermediate 2D image between the first target 2D image and the second target 2D image to obtain the position annotation of the target object in the intermediate 2D image. Thus, the image position annotation in the intermediate frame between the two adjacent motion keyframes is obtained. The intermediate 2D image is a 2D object whose timestamp is located between the timestamp of the first target 2D image and the timestamp of the target 2D image.

[0080] By processing the intermediate frames of all two adjacent motion keyframes in the target video according to the above principle, the image position label of each frame of the target video can be obtained, that is, the position label of the target object in each frame of the target video.

[0081] For example, if the target video consists of 10 frames of two-dimensional images, and assuming that frames 1, 5, and 10 are motion keyframes in the target video, then frames 1, 5, and 10 can be used as target two-dimensional images respectively. Steps S101-S103 are then executed to obtain the position labels of the target object in frame 1, frame 5, and frame 10. Interpolation is performed on the position labels of the target object in frame 1 and frame 5 to obtain the interpolated position labels between frames 1 and 5. The corresponding interpolated position labels are added to frame 2 to obtain the position labels of the target object in frame 2. Similarly, the corresponding interpolated position labels are added to frame 3 to obtain the position labels of the target object in frame 3. Position annotation is performed by adding the corresponding interpolated position annotation in frame 4 to obtain the position annotation of the target object in frame 4. Then, interpolation is performed on the position annotations of the target object in frame 5 and frame 10 to obtain the interpolated position annotation between frame 6 and frame 10. The corresponding interpolated position annotation in frame 6 is added to obtain the position annotation of the target object in frame 6. The corresponding interpolated position annotation in frame 7 is added to obtain the position annotation of the target object in frame 7. The corresponding interpolated position annotation in frame 8 is added to obtain the position annotation of the target object in frame 8. The corresponding interpolated position annotation in frame 9 is added to obtain the position annotation of the target object in frame 9. In this way, image position annotations are added to each of the 10 frames.

[0082] Among them, image frames in the target video where the position of the target object changes significantly can be selected as motion keyframes.

[0083] In the process of interpolating the position annotations of the target object in the first target 2D image and the target object in the second target 2D image to obtain the interpolated position annotations, the interpolation model corresponding to the target object can be determined. The interpolation model includes, but is not limited to, linear interpolation, spline interpolation, etc. The time interval of the interpolation points is determined according to the number of intermediate 2D images to be interpolated. Using the interpolation model corresponding to the target object and the time interval of the interpolation points, the position annotations of the target object in the first target 2D image and the target object in the second target 2D image are interpolated to obtain the interpolated position annotations corresponding to each intermediate 2D image between the first target 2D image and the second target 2D image.

[0084] When the object to be labeled is a dynamic object, the position of the object in the image will change. By interpolating the dynamically changing key images, the position of the object in the image can be inferred and labeled. This can overcome the occlusion problem in video images. Only a portion of the images in the video need to be labeled to achieve the labeling of all images in the video.

[0085] When the two-dimensional image to be labeled is an image in a video, image position labels converted from position labels in the three-dimensional scene are added to the multiple frames of two-dimensional images contained in the video. This gives the position label of the object to be labeled in the multiple frames of two-dimensional images contained in the video, which can save the step of labeling frame by frame and improve the labeling speed.

[0086] In the technical solution corresponding to Figure 3 above, by displaying the target 2D image to be annotated and the 3D scene obtained by 3D reconstruction of the target 2D image on the image annotation interface, and then receiving the object annotation instruction for the target object to be annotated in the 3D scene, and finally responding to the object annotation instruction, a first image position annotation converted from the first scene position annotation associated with the object annotation instruction is added and displayed in the target 2D image on the image annotation interface. The first image position annotation is the position annotation of the target object in the target 2D image, thus realizing the annotation of the object to be annotated in the 2D image, that is, realizing image annotation; since the first scene position annotation is the position annotation of the target object in the 3D scene, and the first image position annotation is the position annotation of the target object in the target 2D image, it is equivalent to annotating the position annotation of the object to be annotated in the 2D image by annotating the object to be annotated in the 3D scene. The visual presentation of the object to be annotated in the 3D scene is consistent with its visual presentation in the real world, which can avoid errors. Given the principle of perspective, objects appear larger when closer and smaller when farther away, users can observe the objects to be labeled in a 3D scene from different angles. This avoids occlusion of the objects, allowing users to accurately and conveniently label their positions in the 3D scene. Since the first scene position label is accurate, and the first image position label is derived from it, it is also accurate. Therefore, it is possible to accurately and conveniently label the position of the object in the 2D image, improving the convenience of image labeling. Furthermore, by simultaneously displaying the target 2D image and the 3D scene reconstructed from the target 2D image on the image labeling interface, users can check the accuracy of the position label by comparing the overlap between the first image position label in the target 2D image and the object to be labeled in the target 2D image. This allows users to fine-tune the position label, making it more accurate and thus improving the overall accuracy of image labeling.

[0087] The following describes the specific implementation method for displaying the 3D scene corresponding to the target 2D image to be annotated on the image annotation interface, including the following steps A1-A3:

[0088] A1. Reconstruct the three-dimensional scene of the target two-dimensional image to be labeled to obtain the virtual three-dimensional scene corresponding to the target two-dimensional image.

[0089] The following steps, A11-A12, can be used to reconstruct a 3D scene from the labeled 2D image of the target, thus obtaining a virtual 3D scene corresponding to the 2D image of the target:

[0090] A11. Based on the camera intrinsic parameters of the target camera, transform the two-dimensional image of the target from the pixel coordinate system to the camera coordinate system to obtain the target projection image.

[0091] Here, the target camera is the camera that captures a two-dimensional image of the target; the camera intrinsic parameters of the target camera refer to the parameters that reflect the physical characteristics and imaging process of the target camera. The camera intrinsic parameters of the target camera can be expressed as [w, h, fx, fy, cx, cy, k1, k2, k3, k4], where w and h represent the width and height of the target two-dimensional image, respectively; fx and fy represent the focal length of the camera in the x and y directions; cx and cy represent the coordinates of the image center of the target two-dimensional image, cx = w / 2, cy = h / 2; and k1, k2, k3, and k4 are the distortion coefficients of the camera, used to describe the radial and tangential distortion of the lens.

[0092] Specifically, the camera intrinsic parameter matrix of the target camera can be constructed based on the target camera's intrinsic parameters. The camera intrinsic parameter matrix of the target camera is represented as follows:

[0093]

[0094] Based on the target's camera intrinsic parameters, the camera's shooting area is constructed in the camera coordinate system. The camera's shooting area is a frustum, and the geometry of the frustum is determined by the field of view angle and the image plane. The formula for calculating the field of view angle of the frustum is as follows:

[0095]

[0096] α represents the horizontal field of view angle, and β represents the vertical field of view angle.

[0097] The coordinates of the four vertices of the image plane in the camera coordinate system are as follows:

[0098]

[0099] z represents the depth of the camera, which is the value along the z-axis of the camera coordinate system.

[0100] Finally, based on the camera intrinsic parameter matrix of the target camera, the pixels in the target 2D image are transformed into the image plane in the camera coordinate system to obtain the target projection image, which is located on the image plane of the view frustum.

[0101] The correspondence between pixels in the target 2D image and pixels in the target projected image (hereinafter referred to as the first correspondence) is as follows:

[0102]

[0103] (u, v) represents the position coordinates of the pixel in the target 2D image, and (x, y, z) represents the position coordinates of the pixel in the camera coordinate system.

[0104] For any pixel in the target 2D image (hereinafter referred to as the first pixel), based on the position coordinates of the first pixel in the target 2D image and the aforementioned first correspondence, the position coordinates of the first pixel in the camera coordinate system (hereinafter referred to as the first camera coordinates) are determined. The texture and color of the first pixel in the target 2D image are then applied to these first camera coordinates, thereby transforming the first pixel into the image plane in the camera coordinate system. Each pixel in the target 2D image is processed in the same way as the first pixel to obtain the target projection image.

[0105] After transforming the pixels in the target 2D image to the image plane in the camera coordinate system based on the camera intrinsic parameter matrix of the target camera, the image in the image plane can be subjected to anti-distortion processing based on the distortion coefficient of the camera to obtain the target projection image.

[0106] A12. Based on the camera extrinsic parameters of the target camera, project the target projection image onto a virtual three-dimensional space to obtain a virtual three-dimensional scene corresponding to the target two-dimensional image.

[0107] Here, the camera extrinsic parameters of the target camera refer to the parameters that reflect the position and attitude information of the camera in three-dimensional space. The camera extrinsic parameters of the target camera can be represented as [X0, Y0, Z0, q]. x q y q z q w ], (X0, Y0, Z0) represents the position of the target camera in three-dimensional space, [q x q y q z q w ] is a quaternion representing the pose of the target camera in three-dimensional space.

[0108] Specifically, the rotation matrix and translation vector between the camera coordinate system and the world coordinate system can be determined based on the target camera's extrinsic parameters. The rotation matrix is ​​represented as follows:

[0109]

[0110] The translation vector is represented as follows:

[0111]

[0112] Ground objects are created in the camera coordinate system, where their coordinates are represented as (x, y, 0). These ground objects are then transformed into the world coordinate system, which is the spatial coordinate system within the 3D scene. The correspondence between the position coordinates in the world coordinate system and the position coordinates in the camera coordinate system (hereinafter referred to as the second correspondence) is as follows:

[0113]

[0114] (X1, Y1, Z1) represents the position coordinates in the world coordinate system.

[0115] Finally, a spatial camera is created in the world coordinate system based on the camera extrinsic parameters of the target camera. The initial position and initial attitude of the spatial camera are consistent with the position and attitude of the target camera in three-dimensional space. Based on the spatial camera and the ground in the world coordinate system, the target projection image is projected onto the virtual three-dimensional space corresponding to the world coordinate system to obtain the virtual three-dimensional scene corresponding to the target two-dimensional image.

[0116] For any pixel in the target projection image (hereinafter referred to as the second pixel), based on the position coordinates of the second pixel in the camera coordinate system (hereinafter referred to as the second camera coordinates) and the aforementioned second correspondence, the position coordinates of the second pixel in the world coordinate system (hereinafter referred to as the first spatial coordinates) are determined. The texture and color of the second pixel in the target projection image are then applied to these first spatial coordinates, thus projecting the second pixel onto the virtual 3D space corresponding to the world coordinate system. Each pixel in the target projection image is processed in the same way as the second pixel, allowing the target projection image to be projected onto the virtual 3D space corresponding to the world coordinate system, resulting in a virtual 3D scene corresponding to the target 2D image.

[0117] It is understood that the first and second correspondences involved in steps A1 to A2 above are the mapping and transformation relationships between the pixel coordinate system corresponding to the target two-dimensional image and the three-dimensional scene coordinate system corresponding to the target two-dimensional image in step B1 above.

[0118] Based on the intrinsic parameters of the camera, the two-dimensional image is first transformed from the pixel coordinate system to the camera coordinate system to obtain the projected image in the camera coordinate system. Then, based on the extrinsic parameters of the camera, the projected image is projected onto the virtual three-dimensional space, which can realize the reconstruction of the three-dimensional scene.

[0119] A2. Obtain a planar image of the virtual 3D scene from the target's viewing perspective to obtain a planar view of the target 3D scene.

[0120] Here, the target viewing perspective is the user's viewing perspective, which can be a default perspective, such as a top-down perspective. After obtaining the virtual 3D scene corresponding to the target 2D image, the position and orientation of the space camera can be adjusted to the position and orientation corresponding to the target viewing perspective to obtain the image on the image plane under the space camera, which serves as the planar image of the virtual 3D scene under the target viewing perspective.

[0121] In some possible cases, before obtaining a planar image of the virtual 3D scene from the target's viewing perspective and thus obtaining a planar view of the target 3D scene, it is also possible to obtain a view setting command for the virtual 3D scene corresponding to the target's 2D image, and set the target's viewing perspective according to the view setting command.

[0122] The viewpoint setting command for the virtual 3D scene corresponding to the target 2D image is triggered by the user. The user can perform user operations on the 3D scene corresponding to the target 2D image on the image annotation interface, thereby triggering the viewpoint setting command; alternatively, the user can also trigger the viewpoint setting command through the toolbar on the image annotation interface. The viewpoint setting command for the virtual 3D scene corresponding to the target 2D image can be triggered by the user before or after the 3D scene corresponding to the target 2D image is displayed on the image annotation interface.

[0123] For example, as shown in Figure 4A, the image annotation interface allows users to select the "Viewpoint Settings" toolbar to set the observation viewpoint. When a user clicks "Viewpoint Settings" and sets a specific observation viewpoint, they receive a viewpoint setting instruction for the virtual 3D scene corresponding to the target 2D image. Based on this instruction, the target observation viewpoint is set as the viewpoint used for setting. Similarly, as shown in Figure 4B, the image annotation interface also allows users to manipulate the 3D scene displayed in P2 (such as rotating the 3D scene) to set the observation viewpoint. Each time the user performs this operation, they receive a viewpoint setting instruction for the virtual 3D scene corresponding to the target 2D image. This application does not limit the timing or method of triggering the viewpoint setting instruction.

[0124] The viewing angle can be set according to the perspective setting command of the virtual 3D scene to obtain a 3D scene planar view, which makes it convenient for users to observe the object to be annotated from various angles, which is conducive to accurate annotation.

[0125] A3. Display the target 3D scene plan view on the image annotation interface.

[0126] In steps A1-A3 above, a virtual three-dimensional scene is obtained by reconstructing a three-dimensional scene from a two-dimensional image, and a planar image of the virtual three-dimensional scene from a viewing angle is obtained to obtain a three-dimensional scene planar diagram. Then, the three-dimensional scene planar diagram is displayed on the image annotation interface, thus realizing the display of a three-dimensional scene corresponding to a two-dimensional image.

[0127] In some possible cases, after displaying the three-dimensional scene corresponding to the target two-dimensional image to be annotated on the image annotation interface through the above steps A1-A3, a scene adjustment command acting on the three-dimensional scene corresponding to the target two-dimensional image can also be received; in response to the scene adjustment command acting on the three-dimensional scene corresponding to the target two-dimensional image, the three-dimensional scene corresponding to the first camera extrinsic parameters is displayed on the image annotation interface.

[0128] The scene adjustment command applied to the 3D scene corresponding to the target 2D image refers to the command used to adjust the camera extrinsic parameters of the target camera. This scene adjustment command includes a first camera extrinsic parameter, which is the camera extrinsic parameter of the target camera. The 3D scene corresponding to the first camera extrinsic parameter refers to the virtual 3D scene obtained by reconstructing the target 2D image based on the first camera extrinsic parameter.

[0129] It can acquire user operations performed on the image annotation interface and receive scene adjustment instructions applied to the three-dimensional scene corresponding to the target two-dimensional image.

[0130] For example, the image annotation interface is shown in Figure 4B. Referring to Figure 4E, when the user clicks on the external parameter adjustment on the image annotation interface and sets the corresponding parameter value in the pop-up window shown in 4E, the user receives a scene adjustment command that applies to the three-dimensional scene corresponding to the target two-dimensional image.

[0131] After receiving the scene adjustment instruction, step A1 above can be re-executed to obtain the virtual 3D scene corresponding to the first camera. Then, steps A2-A3 above can be executed to display the 3D scene corresponding to the extrinsic parameters of the first camera on the image annotation interface.

[0132] By receiving and responding to scene adjustment commands applied to a 3D scene, and displaying the 3D scene corresponding to the camera extrinsic parameters in the scene adjustment command on the image annotation interface, the reconstructed 3D scene can be made more accurate.

[0133] In some possible cases, in addition to annotating the location of the target object, the three-dimensional pose of the target object can also be annotated. Referring to Figure 5, which is a flowchart illustrating another image annotation method provided in an embodiment of this application, the method includes the following steps:

[0134] S201, on the image annotation interface, the target 2D image to be annotated and the corresponding 3D scene are displayed.

[0135] S202, receive object annotation instructions applied to the target object in the three-dimensional scene corresponding to the target two-dimensional image.

[0136] The specific implementation methods of steps S201 to S202 are described here, and will not be repeated here.

[0137] S203, responding to the object annotation instruction applied to the target object in the three-dimensional scene corresponding to the target two-dimensional image, obtain the pose information of the target object from the three-dimensional scene corresponding to the target two-dimensional image, and obtain the target pose information; add and display the first image position annotation obtained by the first scene position annotation in the target two-dimensional image on the image annotation interface, and annotate and display the target pose information on the image annotation interface.

[0138] For details on how to add and display the first image position annotation in the target two-dimensional image on the image annotation interface, please refer to the relevant description of step S103 above, which will not be repeated here.

[0139] Target pose information is used to represent the position and orientation of a target object in the 3D scene corresponding to the target 2D image. The position of the target object in the 3D scene corresponding to the target 2D image can be represented in the form [X, Y, Z], where X, Y, and Z represent the coordinates of the target object's center of gravity along the three axes of the 3D scene coordinate system, respectively. The orientation of the target object in the 3D scene corresponding to the target 2D image can be represented in the form [pitch, roll, yaw], where pitch represents the angle of rotation around the Y-axis of the 3D scene coordinate system, roll represents the angle of rotation around the Z-axis of the 3D scene coordinate system, and yaw represents the angle of rotation around the Z-axis of the 3D scene coordinate system. The orientation of the target object in the 3D scene corresponding to the target 2D image can also be represented in the form [q]. x q y q z q w In the form of ], q x q y q z q represents the axis of rotation. w Indicates the angle of rotation.

[0140] Specifically, the three-dimensional model corresponding to the target object can be determined in the three-dimensional scene corresponding to the target two-dimensional image, and the pose information of the three-dimensional model corresponding to the target object in the three-dimensional scene corresponding to the target two-dimensional image can be determined as the target pose information.

[0141] For example, in the target 2D image on the image annotation interface, the first image position annotation obtained by converting the first scene position annotation is added and displayed, and the target pose information is annotated and displayed on the image annotation interface, as shown in Figure 4F. In Figure 4F, [Xd, Yd, Zd, q xd q yd q zd q wd The symbol ] represents the position and orientation of desk d1 in P1.

[0142] In the technical solution corresponding to Figure 5 above, after receiving the object annotation instruction for the target object in the three-dimensional scene, it not only responds to the object annotation instruction by adding and displaying the first image position annotation converted from the first scene position annotation in the target two-dimensional image on the image annotation interface, but also responds to the object annotation instruction by obtaining the pose information of the target object from the three-dimensional scene, and then annotates and displays the pose information of the target object on the image annotation interface, thus enriching the annotation content of the object to be annotated and making the image annotation data richer.

[0143] In some possible cases, users can also adjust the positional annotations in the 3D scene, thereby adjusting the positional annotations in the 2D image. Referring to Figure 6, which is a flowchart illustrating another image annotation method provided in an embodiment of this application, the method includes the following steps:

[0144] S301, on the image annotation interface, displays the target 2D image to be annotated and the corresponding 3D scene.

[0145] S302, receive object annotation instructions applied to the target object in the three-dimensional scene corresponding to the target two-dimensional image.

[0146] S303, in response to the object annotation instruction applied to the target object in the three-dimensional scene corresponding to the target two-dimensional image, adds and displays the first image position annotation converted from the first scene position annotation in the target two-dimensional image on the image annotation interface.

[0147] The specific implementation methods of steps S301 to S303 can be found in the description of steps S101 to S103, and will not be repeated here.

[0148] S304, receive a label adjustment command applied to the position label of the first scene.

[0149] Here, the annotation adjustment instruction applied to the first scene location annotation refers to the instruction to adjust the first scene location annotation. It can obtain the user operation performed by the user on the first scene location annotation and receive the annotation adjustment instruction applied to the first scene location annotation.

[0150] For example, the first scene location can be labeled as z1 in Figure 4D. Referring to Figure 4G, after the user selects z1 and drags to adjust the four vertices of the quadrilateral model of z1 in 4D, and clicks "confirm" in the pop-up window that prompts whether to change, the user will receive the label adjustment instruction applied to the first scene location label, and the icon labeling interface will change from Figure 4D to Figure 4H.

[0151] The annotation adjustment instruction includes a second scene position annotation. This second scene position annotation is the scene position annotation obtained after the user adjusts the first scene position annotation. The second scene position annotation is the position annotation of the target object in the 3D scene corresponding to the target 2D image. It is used to represent the adjusted position of the target object in the 3D scene corresponding to the target 2D image. The form of the second scene position annotation is the same as that of the first scene position annotation. For example, if the first scene position annotation is shown as z1 in Figure 4D, including the position coordinates of the four corners of the quadrilateral shown in z1 in the 3D scene shown in P2, then the second scene position annotation is shown as z3 in Figure 4H, including the position coordinates of the four corners of the quadrilateral shown in z3 in the 3D scene shown in P2.

[0152] S305, in response to the annotation adjustment command applied to the first scene position annotation, in the target two-dimensional image on the image annotation interface, the first image position annotation is adjusted to the second image position annotation obtained by converting the second scene position annotation, and the second image position annotation is displayed.

[0153] Here, the method of converting the second scene location annotation into the second image location annotation is the same as that of converting the first scene location annotation into the first image location annotation in step S103 above. The principle described in steps B1 to B2 above can be referred to.

[0154] The second image position annotation is the position annotation of the target object in the target 2D image. The second image position annotation is used to indicate the adjusted position of the target object in the target 2D image.

[0155] For example, if the location of the second scene is labeled as z3 in Figure 4H, then the location of the second image is labeled as z4 in Figure 4H.

[0156] In the technical solution corresponding to Figure 6 above, after displaying the position annotation of the object to be annotated on the two-dimensional image, the annotation adjustment instruction acting on the position annotation of the object to be annotated in the three-dimensional scene is received, and the annotation adjustment instruction is responded to, and the position annotation on the two-dimensional image is adjusted to the image position annotation corresponding to the scene position annotation in the annotation adjustment instruction. This allows users to easily fine-tune the position annotation in the two-dimensional image, thereby making the position annotation in the two-dimensional image more accurate.

[0157] The method of this application has been described above; the apparatus of this application will be described below.

[0158] Referring to Figure 7, which is a schematic diagram of an image annotation device provided in an embodiment of this application, the image annotation device 40 includes:

[0159] The interface display module 401 is used to display the target two-dimensional image to be annotated and the corresponding three-dimensional scene on the image annotation interface, wherein the three-dimensional scene is obtained by reconstructing the target two-dimensional image into a three-dimensional scene.

[0160] The instruction receiving module 402 is used to receive an object annotation instruction applied to a target object in the three-dimensional scene, wherein the target object is an object to be annotated, and the object annotation instruction is associated with a first scene position annotation, wherein the first scene position annotation is the position annotation of the target object in the three-dimensional scene.

[0161] The annotation display module 403 is used to respond to the object annotation instruction, add and display the first image position annotation converted from the first scene position annotation in the target two-dimensional image on the image annotation interface, wherein the first image position annotation is the position annotation of the target object in the target two-dimensional image.

[0162] In one possible design, the above-mentioned annotation display module 403 is further configured to: respond to the object annotation instruction, obtain the pose information of the target object from the three-dimensional scene, and obtain the target pose information; and annotate and display the target pose information on the image annotation interface.

[0163] In one possible design, the above-mentioned annotation display module 403 is specifically used to: perform coordinate mapping transformation on the first position coordinates corresponding to the first scene position annotation to obtain the second position coordinates, wherein the first position coordinates are the position coordinates in the three-dimensional scene and the second position coordinates are the position coordinates in the two-dimensional image; superimpose the object position mark used to represent the target object onto the second position coordinates in the target two-dimensional image to obtain the first image position annotation; and display the first image position annotation in the target two-dimensional image on the image annotation interface.

[0164] In one possible design, the above-mentioned annotation display module 403 is specifically used for: reconstructing a three-dimensional scene from the target two-dimensional image to obtain a virtual three-dimensional scene corresponding to the target two-dimensional image; acquiring a planar image of the virtual three-dimensional scene from the target's viewing perspective to obtain a planar view of the target three-dimensional scene; and displaying the planar view of the target three-dimensional scene on the image annotation interface.

[0165] In one possible design, the above-mentioned annotation display module 403 is specifically used to: transform the target two-dimensional image from the pixel coordinate system to the camera coordinate system based on the camera intrinsic parameters of the target camera to obtain a target projection image, wherein the target camera is the camera that captures the target two-dimensional image; and project the target projection image onto a virtual three-dimensional space based on the camera extrinsic parameters of the target camera to obtain the virtual three-dimensional scene.

[0166] In one possible design, the above-mentioned annotation display module 403 is further configured to: obtain a view setting instruction for the virtual three-dimensional scene; and set the target observation view according to the view setting instruction.

[0167] In one possible design, the instruction receiving module 402 is further configured to: receive an annotation adjustment instruction acting on the first scene position annotation, the annotation adjustment instruction including a second scene position annotation, the second scene position annotation being the position annotation of the target object in the three-dimensional scene; the annotation display module 403 is further configured to: respond to the annotation adjustment instruction, adjust the first image position annotation in the target two-dimensional image on the image annotation interface to a second image position annotation converted from the second scene position annotation, and display the second image position annotation, the second image position annotation being the position annotation of the target object in the target two-dimensional image.

[0168] In one possible design, the instruction receiving module 402 is further configured to: receive a scene adjustment instruction acting on the three-dimensional scene, the scene adjustment instruction including a first camera extrinsic parameter, the first camera extrinsic parameter being a camera extrinsic parameter of a target camera, the target camera being a camera that captures the target two-dimensional image; the interface display module 401 is further configured to: respond to the scene adjustment instruction and display the three-dimensional scene corresponding to the first camera extrinsic parameter on the image annotation interface.

[0169] In one possible design, the target two-dimensional image is a frame of a target video, and the target video includes multiple frames of two-dimensional images; the above-mentioned annotation display module 403 is further configured to: add image position annotations to the multiple frames of two-dimensional images respectively to obtain the position annotation of the target object in the multiple frames of two-dimensional images, wherein the image position annotation is the position annotation of the target object in the two-dimensional image, and the image position annotation is obtained based on the position annotation of the target object in the three-dimensional scene.

[0170] In one possible design, the target object is a dynamic object; the annotation display module 403 is specifically used to: determine a first target two-dimensional image and a second target two-dimensional image in the multi-frame two-dimensional images, where both the first target two-dimensional image and the second target two-dimensional image are target two-dimensional images in the target video, and both the first target two-dimensional image and the second target two-dimensional image are motion keyframes in the target video; interpolate the position annotation of the target object in the first target two-dimensional image and the position annotation of the target object in the second target two-dimensional image to obtain an interpolated position annotation, wherein the position annotation of the target object in the first target two-dimensional image is obtained by converting the position annotation of the target object in the three-dimensional scene corresponding to the first target two-dimensional image, and the position annotation of the target object in the second target two-dimensional image is obtained by converting the position annotation of the target object in the three-dimensional scene corresponding to the second target two-dimensional image; add the interpolated position annotation to an intermediate two-dimensional image between the first target two-dimensional image and the second target two-dimensional image to obtain the position annotation of the target object in the intermediate two-dimensional image, wherein the intermediate two-dimensional image is a two-dimensional image with a timestamp located between the timestamp of the first target two-dimensional image and the timestamp of the second target two-dimensional image.

[0171] It should be noted that any content not mentioned in the embodiment corresponding to Figure 7 can be found in the description of the aforementioned method embodiments, and will not be repeated here.

[0172] The aforementioned device displays a target 2D image to be annotated and a 3D scene reconstructed from the target 2D image on an image annotation interface. It receives object annotation instructions applied to the target object in the 3D scene. Finally, in response to the object annotation instructions, it adds and displays a first image position annotation, converted from a first scene position annotation associated with the object annotation instructions, to the target 2D image on the image annotation interface. This first image position annotation represents the position of the target object in the target 2D image, thus achieving image annotation of the object to be annotated in the 2D image. Since the first scene position annotation represents the position of the target object in the 3D scene, and the first image position annotation represents the position of the target object in the target 2D image, it is equivalent to annotating the position of the object in the 2D image by annotating the object in the 3D scene. The visual representation of the object in the 3D scene is consistent with its visual representation in the real world, avoiding near-large formatting issues. In a 3D scene, users can observe the object to be labeled from different angles, avoiding occlusion. This allows for accurate and convenient labeling of the object's location within the 3D scene. Since the first scene location label is accurate, and the first image location label is derived from it, it is also accurate. Therefore, it's possible to accurately and conveniently label the object's location in the 2D image, improving the convenience of image labeling. Furthermore, by simultaneously displaying the target 2D image and the 3D scene reconstructed from the target 2D image on the image labeling interface, users can verify the accuracy of the location label by checking the overlap between the first image location label and the object in the target 2D image after labeling the first scene location. This allows for fine-tuning of the location label, making it more accurate and thus improving the overall accuracy of image labeling.

[0173] Referring to Figure 8, which is a schematic diagram of the structure of a computer device provided in an embodiment of this application, the computer device 50 includes a processor 501, a memory 502, and a display panel 503. The memory 502 and the display panel 503 are connected to the processor 501, for example, via a bus.

[0174] Processor 501 is configured to support the computer device 50 in performing the corresponding functions in the methods described in the above method embodiments. Processor 501 may be a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof. The aforementioned hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The aforementioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0175] Display panel 503 is used to display the image annotation interface.

[0176] Memory 502 is used to store program code, etc. Memory 502 may include volatile memory (VM), such as random access memory (RAM); memory 502 may also include non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); memory 502 may also include combinations of the above types of memory.

[0177] The memory 502 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the image annotation method in the embodiments of this application. The core processing and graphics processor cooperate to execute various functional applications and data processing of the image annotation method by running the non-volatile software programs, instructions, and modules stored in the memory, thereby realizing the functions of the image annotation method provided in the above method embodiments.

[0178] The memory 502 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function. The data storage area may store data created based on the use of the image annotation device. In some embodiments, the memory may include memory remotely located relative to the processor, which may be connected to the image annotation device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0179] The one or more modules are stored in the memory. When executed by the one or more processors, they perform the image annotation method in any of the above method embodiments. For example, they perform the method steps described in the above method embodiments to realize the functions of the modules described in the above device embodiments.

[0180] This application also provides a computer-readable storage medium storing a computer program, the computer program including program instructions, which, when executed by a computer, cause the computer to perform the method described in the foregoing embodiments.

[0181] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0182] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.

Claims

1. An image annotation method, characterized in that, include: On the image annotation interface, a target 2D image to be annotated and a corresponding 3D scene are displayed. The 3D scene is obtained by reconstructing the target 2D image into a 3D scene. An object annotation instruction is received for a target object in the 3D scene. The target object is the object to be annotated. The object annotation instruction is associated with a first scene position annotation, which is the position annotation of the target object in the 3D scene. In response to the object annotation instruction, a first image position annotation converted from the first scene position annotation is added to and displayed on the target 2D image on the image annotation interface. The first image position annotation is the position annotation of the target object in the target 2D image.

2. The method according to claim 1, characterized in that, After receiving the object annotation instruction applied to the target object in the three-dimensional scene, the method further includes: responding to the object annotation instruction, obtaining the pose information of the target object from the three-dimensional scene to obtain target pose information; and annotating and displaying the target pose information on the image annotation interface.

3. The method according to claim 1, characterized in that, Adding and displaying the first image location marker, obtained by converting the first scene location marker, in the target two-dimensional image on the image annotation interface includes: performing coordinate mapping transformation on the first location coordinates corresponding to the first scene location marker to obtain the second location coordinates, wherein the first location coordinates are the location coordinates in the three-dimensional scene and the second location coordinates are the location coordinates in the two-dimensional image; superimposing the object location marker used to represent the target object onto the second location coordinates in the target two-dimensional image to obtain the first image location marker; and displaying the first image location marker in the target two-dimensional image on the image annotation interface.

4. The method according to claim 1, characterized in that, The step of displaying the three-dimensional scene corresponding to the target two-dimensional image on the image annotation interface includes: reconstructing the three-dimensional scene from the target two-dimensional image to obtain a virtual three-dimensional scene corresponding to the target two-dimensional image; acquiring a planar image of the virtual three-dimensional scene from the target's viewing perspective to obtain a planar view of the target three-dimensional scene; and displaying the planar view of the target three-dimensional scene on the image annotation interface.

5. The method according to claim 4, characterized in that, The step of reconstructing a three-dimensional scene from the target two-dimensional image to obtain a virtual three-dimensional scene corresponding to the target two-dimensional image includes: transforming the target two-dimensional image from a pixel coordinate system to a camera coordinate system based on the camera intrinsic parameters of the target camera to obtain a target projection image, wherein the target camera is the camera that captured the target two-dimensional image; and projecting the target projection image onto a virtual three-dimensional space based on the camera extrinsic parameters of the target camera to obtain the virtual three-dimensional scene.

6. The method according to claim 4, characterized in that, Before obtaining a planar image of the virtual 3D scene from the target observation perspective and obtaining a planar view of the target 3D scene, the process includes: obtaining a view setting instruction for the virtual 3D scene; and setting the target observation perspective according to the view setting instruction.

7. The method according to any one of claims 1-6, characterized in that, After adding and displaying the first image position annotation converted from the first scene position annotation in the target two-dimensional image on the image annotation interface, the method further includes: receiving an annotation adjustment instruction acting on the first scene position annotation, the annotation adjustment instruction including a second scene position annotation, the second scene position annotation being the position annotation of the target object in the three-dimensional scene; responding to the annotation adjustment instruction, adjusting the first image position annotation in the target two-dimensional image on the image annotation interface to the second image position annotation converted from the second scene position annotation, and displaying the second image position annotation, the second image position annotation being the position annotation of the target object in the target two-dimensional image.

8. The method according to any one of claims 1-6, characterized in that, After displaying the target 2D image to be annotated and the corresponding 3D scene on the image annotation interface, the method further includes: receiving a scene adjustment command applied to the 3D scene, the scene adjustment command including a first camera extrinsic parameter, the first camera extrinsic parameter being the camera extrinsic parameter of a target camera, the target camera being the camera that captured the target 2D image; and responding to the scene adjustment command by displaying the 3D scene corresponding to the first camera extrinsic parameter on the image annotation interface.

9. The method according to any one of claims 1-6, characterized in that, The target two-dimensional image is a frame of two-dimensional image in the target video, and the target video includes multiple frames of two-dimensional images; the method further includes: adding image position annotations to the multiple frames of two-dimensional images respectively to obtain the position annotation of the target object in the multiple frames of two-dimensional images, wherein the image position annotation is the position annotation of the target object in the two-dimensional image, and the image position annotation is obtained by converting the position annotation of the target object in the three-dimensional scene.

10. The method according to claim 9, characterized in that, The target object is a dynamic object; the step of adding image position annotations to the multiple frames of two-dimensional images to obtain the position annotations of the target object in the multiple frames of two-dimensional images includes: determining a first target two-dimensional image and a second target two-dimensional image in the multiple frames of two-dimensional images, where both the first target two-dimensional image and the second target two-dimensional image are target two-dimensional images in the target video, and both the first target two-dimensional image and the second target two-dimensional image are motion keyframes in the target video; interpolating the position annotations of the target object in the first target two-dimensional image and the position annotations of the target object in the second target two-dimensional image to obtain interpolated position annotations, wherein the position annotation of the target object in the first target two-dimensional image is obtained by converting the position annotation of the target object in the three-dimensional scene corresponding to the first target two-dimensional image, and the position annotation of the target object in the second target two-dimensional image is obtained by converting the position annotation of the target object in the three-dimensional scene corresponding to the second target two-dimensional image; adding the interpolated position annotations to an intermediate two-dimensional image between the first target two-dimensional image and the second target two-dimensional image to obtain the position annotations of the target object in the intermediate two-dimensional image, wherein the intermediate two-dimensional image is a two-dimensional image whose timestamp is located between the timestamp of the first target two-dimensional image and the timestamp of the second target two-dimensional image.

11. An image annotation device, characterized in that, include: The interface display module is used to display the target two-dimensional image to be annotated and the corresponding three-dimensional scene on the image annotation interface, wherein the three-dimensional scene is obtained by reconstructing the target two-dimensional image into a three-dimensional scene; The instruction receiving module is used to receive an object annotation instruction applied to a target object in the three-dimensional scene. The target object is an object to be annotated. The object annotation instruction is associated with a first scene position annotation, which is the position annotation of the target object in the three-dimensional scene. The annotation display module is used to respond to the object annotation instruction by adding and displaying a first image position annotation converted from the first scene position annotation in the target two-dimensional image on the image annotation interface. The first image position annotation is the position annotation of the target object in the target two-dimensional image.

12. A computer device, characterized in that, The device includes a memory, a display panel, and a processor, wherein the memory and the display panel are connected to the processor, the display panel is used to display an image annotation interface, and the processor is used to execute one or more computer programs stored in the memory, wherein when the processor executes the one or more computer programs, the computer device causes the computer device to implement the method as described in any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program including program instructions that, when executed by a processor, cause the processor to perform the method as described in any one of claims 1-10.