Target object labeling method and device, storage medium and electronic equipment

By transforming the vertex coordinates of virtual target object models in 3D scene images, the problems of insufficient diversity and quality of image datasets are solved, achieving efficient target object annotation and dataset generation, and improving the effectiveness of model training and testing.

CN114863071BActive Publication Date: 2026-01-16BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210503463.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-09
Publication Date
2026-01-16
Estimated Expiration
2042-05-09

AI Technical Summary

Technical Problem

In the field of target tracking and detection, existing technologies are limited by the diversity and quality of image datasets. Traditional annotation methods rely on manual or machine learning, which are inefficient and prone to errors. Furthermore, existing virtual image generation methods cannot accurately simulate the relationship between lighting and materials in the real environment, resulting in insufficient annotation information.

Method used

By obtaining the vertex coordinates of the virtual target object model in the target 3D scene image, and using 3D modeling software and rendering camera configuration information, the 3D coordinates are converted into position coordinates in the 2D image to determine the position annotation information of the target object, and to simulate the real environment to construct diverse and complex annotation situations.

Benefits of technology

It enables the generation of rich datasets, improves the quality and diversity of datasets, reduces the reliance on manual annotation, and improves the training and testing efficiency of target tracking and detection models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114863071B_ABST
    Figure CN114863071B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a target object labeling method and device, a storage medium and an electronic device. The method comprises: in response to receiving a labeling request for a target object, obtaining a target three-dimensional scene image, wherein the target three-dimensional scene image includes a virtual target object model created for the target object; obtaining first position coordinates of each vertex of the virtual target object model in a world coordinate system to which the virtual target object model belongs; determining second position coordinates of the target vertex in a two-dimensional image corresponding to a rendering camera according to configuration information of the rendering camera and the first position coordinates of each vertex in the target three-dimensional scene image; and determining position labeling information of the target three-dimensional scene image related to the target task according to the second position coordinates of the target vertex. In this way, labeling information of more diverse and complex situations can be constructed by simulating a real environment, the labeling information of the target object can be effectively obtained, and a rich data set can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer vision, in particular, to a target object labeling method and device, a storage medium and an electronic device. BACKGROUND

[0002] With the continuous development of artificial intelligence technology and the increasing demand for image processing, labeling technology is widely used in target tracking and target detection. The purpose is to input an image into a model to obtain the labeling result of the object in the image output by the model, and then realize target tracking or detection. For training and testing verification of image-based machine learning and neural network models, rich image datasets and a large number of target object labels are the basis for obtaining a model with high tracking or detection accuracy. Therefore, in the training and testing verification tasks of machine learning and neural network models, how to enrich the image dataset and the labeling information is the focus of research. SUMMARY

[0003] This section is provided to introduce the general concepts of the present disclosure in a simplified form, which will be described in detail in the following detailed description section. This section does not intend to identify key or essential features of the claimed technical solutions, nor to limit the scope of the claimed technical solutions.

[0004] In a first aspect, the present disclosure provides a target object labeling method, comprising:

[0005] In response to receiving a labeling request for a target object, a target three-dimensional scene image is obtained, wherein the target three-dimensional scene image includes a virtual target object model created for the target object;

[0006] The first position coordinates of each vertex of the virtual target object model in a world coordinate system to which the virtual target object model belongs are obtained.

[0007] According to the configuration information of the rendering camera in the target three-dimensional scene image and the first position coordinates of the vertices, the second position coordinates of the target vertices in the two-dimensional image corresponding to the rendering camera are determined.

[0008] According to the second position coordinates of the target vertices, the position labeling information of the target three-dimensional scene image related to the target task is determined.

[0009] In a second aspect, the present disclosure provides a target object labeling device, comprising:

[0010] The first obtaining module is configured to, in response to receiving a labeling request for a target object, obtain a target three-dimensional scene image, wherein the target three-dimensional scene image comprises a virtual target object model created for the target object;

[0011] The second obtaining module is configured to obtain first position coordinates of each vertex of the virtual target object model in a world coordinate system to which the virtual target object model belongs.

[0012] The first determining module is configured to determine, according to configuration information of a rendering camera in the target three-dimensional scene image and the first position coordinates of the vertices, second position coordinates of target vertices in a two-dimensional image corresponding to the rendering camera.

[0013] The second determining module is configured to determine, according to the second position coordinates of the target vertices, position labeling information of the target three-dimensional scene image related to a target task.

[0014] In a third aspect, the present disclosure provides a computer readable medium having a computer program stored thereon, wherein the program, when executed by a processing device, implements the steps of the method in the first aspect.

[0015] In a fourth aspect, the present disclosure provides an electronic device comprising:

[0016] A storage device having at least one computer program stored thereon;

[0017] At least one processing device configured to execute the at least one computer program in the storage device to implement the steps of the method in the first aspect.

[0018] According to the above technical solution, the first position coordinates of each vertex of the virtual target object model in the target three-dimensional scene image in the world coordinate system to which the virtual target object model belongs are converted into the second position coordinates of the target vertices in the two-dimensional image corresponding to the rendering camera, that is, through spatial position conversion, the position information of the target object in the three-dimensional scene under the two-dimensional image corresponding to the rendering camera is obtained, and then the position labeling information of the target object is determined according to the position information. In this way, a more diverse and complex labeling information can be constructed by simulating a real environment, the labeling information of the target object can be effectively obtained, and then a rich data set can be obtained.

[0019] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS

[0020] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings. The same or similar components have the same or similar reference labels. It should be understood that the drawings are schematic and elements in the drawings have not necessarily been drawn to scale. In the drawings:

[0021] Figure 1 is a flowchart of a target object labeling method according to an exemplary embodiment;

[0022] Figure 2 is a schematic diagram of a virtual target object model, a virtual environment object model, and a rendering camera in a three-dimensional scene image according to an exemplary embodiment;

[0023] Figure 3 is Figure 2 is a two-dimensional image corresponding to the rendering camera shown in FIG. 8B;

[0024] Figure 4 is a block diagram of a target object labeling apparatus according to an exemplary embodiment;

[0025] Figure 5 is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0026] In related art, there are the following challenges in enriching image datasets and labeling information. On the one hand, due to privacy protection issues or restrictions of non-universal scenes, the difficulty of image collection is increasing, which limits the diversity of image detection environments and the quality of datasets. On the other hand, traditional image labeling work requires a large amount of manual intervention, or pre-labeling by means of machine learning or deep learning. However, the pre-labeling by means of machine learning or deep learning is heavily dependent on the transfer performance of the used model, and for a completely new dataset, manual intervention is still needed for post-labeling correction, which has the disadvantages of low efficiency and large errors.

[0027] In addition, there is a method of generating a dataset by means of a 3D model plus a material map to make a virtual target image. The poses and angles of an object are modeled by a 3D software (for example, 3DMAX software, Maya software, Poser software, etc.), and then a series of pictures of the object are generated by rendering as a dataset.

[0028] This method improves the way of relying purely on traditional real image acquisition and can generate a relatively rich virtual image dataset. However, the limitation of the current method is that: on the one hand, only the target object is modeled, without labeling, and the dataset is directly obtained, which is single in purpose and weak in generality; on the other hand, the comprehensive relationship between light and material in the real environment cannot be well simulated, such as the reflection and scattering of ambient light, the brightness and clarity of the target object, and the perspective angle and occlusion degree of the target object in space, etc., thus there is a large feature deviation from the image data collected in the real environment. Therefore, the method cannot accurately obtain the labeling information of the target object, and thus cannot obtain a rich dataset.

[0029] Therefore, the present disclosure provides a target object labeling method and device, a storage medium and an electronic device to effectively obtain the labeling information of the target object, and thus obtain a rich dataset.

[0030] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, but rather these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.

[0031] It should be understood that each step described in the method embodiments of the present disclosure can be executed in different order and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this respect.

[0032] The term "comprising" and variations thereof as used herein are open-ended, that is "including but not limited to". The term "based on" is "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related definitions will be given in the description below.

[0033] It should be noted that the concepts of "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.

[0034] It should be noted that the modification of "one" or "multiple" mentioned in the present disclosure is illustrative and not limiting, and those skilled in the art should understand that, unless otherwise explicitly indicated in the context, it should be understood as "one or more".

[0035] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0036] All actions of obtaining signals, information or data in the present disclosure are carried out in compliance with the corresponding data protection regulations and policies of the country where the device is located, and with the authorization given by the corresponding device owner.

[0037] Figure 1 is a flowchart of a target object labeling method according to an example embodiment. As shown in Figure 1 The method can include the following steps.

[0038] In step S11, in response to receiving a labeling request for a target object, a target three-dimensional scene image is obtained. The target three-dimensional scene image includes a virtual target object model created for the target object.

[0039] First of all, it should be understood that with the iterative updates of hardware such as GPU and CPU, and the algorithmic capabilities of software such as three-dimensional graphics renderers, the rendering of three-dimensional virtual scenes has brought qualitative improvements in light, material, etc., so that the virtual scene can achieve a photo-level display effect. Secondly, it should be understood that in a three-dimensional scene image, a variety of three-dimensional scene images can be obtained by simply setting different behavior or posture parameters of the target object, different configuration information of the rendering camera, and different environmental information. Therefore, based on the above considerations, the target object is labeled based on the three-dimensional scene image in the present disclosure.

[0040] In step S12, the first position coordinates of each vertex of the virtual target object model in the world coordinate system to which the virtual target object model belongs are obtained.

[0041] In the present disclosure, the first position coordinates of each vertex of the virtual target object model can be obtained through the API (Application Program Interface) embedded in the three-dimensional modeling software. For example, a third script tool for traversing each vertex of the virtual target object model is obtained, and in the target three-dimensional scene image, each vertex of the virtual target object model is traversed according to the third script tool to obtain the first position coordinates of each vertex in the world coordinate system to which the virtual target object model belongs. For example, the Get Object All Vectors In World 3D Axis() function is called to obtain the first position coordinates of each vertex in the world coordinate system. The world coordinate system is the coordinate information of the target object in the global coordinate system of the three-dimensional scene, and the first position coordinates are three-dimensional coordinates.

[0042] In step S13, according to the configuration information of the rendering camera in the target three-dimensional scene image and the first position coordinates of each vertex, the second position coordinates of the target vertex in the two-dimensional image corresponding to the rendering camera are determined.

[0043] In an embodiment, the target vertex can be each vertex determined in step S12, and accordingly, the specific implementation of step S13 is: according to the configuration information of the rendering camera in the target three-dimensional scene image, a first script tool for converting three-dimensional coordinates into two-dimensional coordinates is obtained, and the first script tool is called to convert the first position coordinates of each vertex into the second position coordinates of each vertex in the two-dimensional image corresponding to the rendering camera.

[0044] In another embodiment, the target vertex is a visible vertex in the two-dimensional image corresponding to the rendering camera, i.e., a vertex that is not blocked by an environmental object, and accordingly, the specific implementation of step S13 is: according to the configuration information of the rendering camera in the target three-dimensional scene image, a second script tool for determining the first position coordinates of the visible vertex in the first position coordinates of each vertex is obtained, and the second script tool is called to obtain the first position coordinates of the visible vertex; the first script tool for converting three-dimensional coordinates into two-dimensional coordinates is called to convert the first position coordinates of the visible vertex into the second position coordinates in the two-dimensional image corresponding to the rendering camera.

[0045] In yet another embodiment, the target vertex is each vertex determined in step S12 and the visible vertex in the two-dimensional image corresponding to the rendering camera. For example, the pseudo code for converting the first position coordinates V(g) of each vertex into the second position coordinates V(c) of each vertex in the two-dimensional image corresponding to the rendering camera and into the second position coordinates V(c') of the visible vertex in the two-dimensional image corresponding to the rendering camera is as follows:

[0046] Plain Text

[0047] input: given vertex set V(g) in world axis;

[0048] output: vertex set V(c) in camera local 2D axis;

[0049] V(c)←Convert To Camera 2D Axis(c,V(g))

[0050] V(g')←Get Visible Vertices(c,V(g))

[0051] V(c')←Convert To Camera 2D Axis(c,V(g'))

[0052] wherein c represents a two-dimensional image corresponding to the rendering camera, and V(g') represents the first position coordinates of the visible vertex.

[0053] In step S14, the position annotation information of the target three-dimensional scene image related to the target task is determined according to the second position coordinates of the target vertex.

[0054] According to the technical solution, the first position coordinates of each vertex of the virtual target object model in the world coordinate system to which the virtual target object model belongs are converted into the second position coordinates of the target vertex in the two-dimensional image corresponding to the rendering camera, that is, the position information of the target object in the three-dimensional scene under the two-dimensional image corresponding to the rendering camera is obtained through spatial position conversion, and then the position annotation information of the target object is determined according to the position information. In this way, the annotation information of more diverse and complex situations can be simulated in a real environment, the annotation information of the target object can be effectively obtained, and then a rich data set can be obtained.

[0055] For example, before obtaining the target three-dimensional scene image, the target object annotation method can further include:

[0056] Firstly, in response to receiving a request for constructing a three-dimensional scene image of a target object, a virtual target object model is created for the target object, and a virtual environment object model is created for a preset environment object. For example, if the target object is an object in an outdoor environment, the environment object can include but is not limited to buildings, vegetation, static obstacles, vehicles and the like. Among them, a large number of built-in or available models are provided in existing three-dimensional content creation software and real-time rendering software, so in the present disclosure, a virtual target object model can be created for the target object and a virtual environment object model can be created for the preset environment object through the three-dimensional content creation software and the real-time rendering software. In this way, the workload of constructing the three-dimensional scene image can be reduced, and the efficiency of constructing the three-dimensional scene image can be improved.

[0057] Then, the animation of the virtual target object model and the virtual environment object model, and the configuration information of the rendering camera are obtained to obtain a plurality of three-dimensional scene images.

[0058] For example, the user can set the animation of the virtual target object model and the virtual environment object model and the configuration information of the rendering camera according to actual needs in the human-computer interaction interface, and then the device performing the target object labeling method can obtain the animation of the virtual target object model and the virtual environment object model and the configuration information of the rendering camera. The position of the virtual target object model and the virtual environment object model is changed through the animation, and rich target object position information is obtained, which is beneficial to the diversity of the data set and the expansion of the applicable range. The configuration information of the rendering camera can include the position, angle, focal length, and the like of the rendering camera. The configuration information of the rendering camera is related to the relative position of the target object and the environment object in the camera rendering picture, the perspective, and the light material calculation, occlusion, blur degree, and the like in the subsequently generated three-dimensional scene image. Therefore, in the present disclosure, a plurality of three-dimensional scene images can be obtained by setting the animation of the virtual target object model and the virtual environment object model and the configuration information of different rendering cameras.

[0059] It should be understood that each frame of three-dimensional scene image can include one or more rendering cameras, which is not specifically limited in the present disclosure. For ease of description, one rendering camera is included in each frame of three-dimensional scene image as an example in the following description.

[0060] For example, the positions of the virtual target object model and the virtual environment object model are fixed and unchanged, and different three-dimensional scene images can be obtained by adjusting the configuration information of the rendering camera. For another example, the configuration information of the rendering camera is fixed and unchanged, and different three-dimensional scene images can be obtained by changing the positions of the virtual target object model and / or the virtual environment object model. For another example, the configuration information of the rendering camera and the positions of the virtual target object model and the virtual environment object model are adjusted simultaneously to obtain different three-dimensional scene images.

[0061] In this way, a plurality of three-dimensional scene images containing the virtual target object model and the virtual environment object model can be obtained in the above manner, and the three-dimensional scene images can more ideally present the mutual influence between various different objects (target objects and objects) in the scene, such as light, material, occlusion, clarity, and the like, and further more closely approximate the target object effect in the real scene.

[0062] Accordingly, in response to receiving the labeling request for the target object, the specific manner of obtaining the target three-dimensional scene image is: in response to receiving the labeling request for the target object, each frame of the three-dimensional scene image is sequentially determined as the target three-dimensional scene image. That is, for each frame of the three-dimensional scene image, the position labeling information of the three-dimensional scene image related to the target task can be determined according to the target object labeling method provided by the present disclosure. In this way, the data of the target three-dimensional scene image is enriched, and thus more labeling information of the target object can be obtained.

[0063] The specific implementation of step S14 will be described below. Figure 1 The specific implementation of step S14 will be described below.

[0064] In an embodiment, the target task is a segmentation task. The specific implementation of step S14 is: determining the second position coordinates of the target boundary vertices of the virtual target object model in the two-dimensional image corresponding to the rendering camera according to the second position coordinates of the target vertices, and determining the second position coordinates of the target boundary vertices as the position labeling information of the target three-dimensional scene image related to the segmentation task.

[0065] It is worth noting that in this embodiment, the target vertices can be the vertices and / or the visible vertices in step S12, which are not limited by the present disclosure. When the target vertices are the vertices, the target boundary vertices are referred to as the boundary vertices, and when the target vertices are the visible vertices, the target boundary vertices are referred to as the visible boundary vertices. The following will be described by taking the target vertices as the vertices and the visible vertices, and the target boundary vertices as the boundary vertices and the visible boundary vertices as examples.

[0066] For example, the second position coordinates V(s) of the boundary vertices of the virtual target object model in the two-dimensional image corresponding to the rendering camera are calculated based on the second position coordinates V(c) of the vertices, and the second position coordinates V(s') of the visible boundary vertices of the virtual target object model in the two-dimensional image corresponding to the rendering camera are calculated based on the second position coordinates V(c') of the visible vertices. For example, the second position coordinates V(s) and V(s') of the target boundary vertices can be obtained by using a graph boundary (contour) finding algorithm. For example, the Basic Find Contours() algorithm for finding contours provided in the OpenCV library. The pseudo code for determining the second position coordinates of the target boundary vertices by the Basic Find Contours() algorithm is as follows:

[0067] Plain Text

[0068] input: vertex sets V(c) and V(c') in camera local 2D axis;

[0069] output: vertex set V(s) and V(s') for object contours;

[0070] V(s)←Basic Find Contours(V(c))

[0071] V(s')←Basic Find Contours(V(c'))

[0072] In this embodiment, the second position coordinates V(s) and V(s') of the determined target boundary vertices are position labeling information of the target three-dimensional scene image related to the segmentation task. That is, the second position coordinates V(s) and V(s') of the target boundary vertices can be used as position labeling information for training and testing and verifying the segmentation model, and subsequently, a sample data set for training and testing and verifying the segmentation model can be constructed according to the second position coordinates V(s) and V(s') of the target boundary vertices.

[0073] In another embodiment, the target task is a detection task, and the specific implementation of the above step S14 is: determining the minimum value and the maximum value of the first coordinate axis and the minimum value and the maximum value of the second coordinate axis according to the second position coordinates of the target vertices, respectively; determining the minimum value of the first coordinate axis and the minimum value of the second coordinate axis as the position coordinates of the minimum value vertex, and determining the maximum value of the first coordinate axis and the maximum value of the second coordinate axis as the position coordinates of the maximum value vertex; and determining the position coordinates of the minimum value vertex and the position coordinates of the maximum value vertex as position labeling information of the target three-dimensional scene image related to the detection task.

[0074] Similarly, in this embodiment, the target vertices can be the vertices and / or the visible vertices in the above step S12, which are not specifically limited by the present disclosure. The following takes the vertices and the visible vertices as examples for illustration.

[0075] For example, the second position coordinates V(b) of the extreme value vertices are calculated based on the second position coordinates V(c) of the vertices, and the second position coordinates V(b') of the visible extreme value vertices are calculated based on the second position coordinates V(c') of the visible vertices, the extreme value vertices including the position coordinates of the maximum value vertex and the position coordinates of the minimum value vertex. The pseudo code for determining the second position coordinates V(b) of the extreme value vertices and the second position coordinates V(b') of the visible extreme value vertices is as follows:

[0076] Plain Text

[0077] input: given contour vertex sets V(c) and V(c');

[0078] output: bounding boxes V(b) and V(b');

[0079] V(b) <- Basic Get Bounding Boxes(V(c))

[0080] V(b') <- Basic Get Bounding Boxes(V(c'))

[0081] It is explained that the second position coordinates V(b) of the maximum vertices and the second position coordinates V(b') of the visible maximum vertices are obtained in a similar manner. Hereinafter, only the second position coordinates V(b) of the maximum vertices are taken as an example to describe the Basic Get Bounding Boxes().

[0082]

[0083]

[0084] Similarly, if V<<X,Y>> in the above pseudo code is replaced by the second position coordinates V(c') of the visible vertices, the second position coordinates V(b') of the visible maximum vertices can be obtained according to the above pseudo code.

[0085] In this embodiment, the second position coordinates V(b) of the maximum vertices and the second position coordinates V(b') of the visible maximum vertices are position labeling information of the target three-dimensional scene image related to the detection task. That is, V(b) and V(b') can be used as position labeling information for training and testing and verifying the detection model, and subsequently, sample data sets for training and testing and verifying the detection model can be constructed according to V(b) and V(b').

[0086] It should be understood that in the present disclosure, different position coordinates can be obtained according to actual needs. For example, the first position coordinates V(g) of each vertex, the first position coordinates V(g') of the visible vertex, the second position coordinates V(c) of each vertex, the second position coordinates V(c') of the visible vertex, the second position coordinates V(s) and V(s') of the target boundary vertex, and the second position coordinates V(b) and V(b') of the extreme vertex can be obtained at the same time. For another example, part of the above-mentioned position coordinates can also be obtained, for example, only V(g), V(c'), V(s') and V(b') are obtained. The present disclosure does not make specific limitations on this.

[0087] In the above manner, the position labeling information of each target three-dimensional scene image related to the target task can be obtained, and then the position labeling information can be stored.

[0088] For example, first, for each target three-dimensional scene image, the basic information of the virtual target object model, the configuration information of the rendering camera and the frame information of the target three-dimensional scene image are obtained as label information, and a labeling set of the target three-dimensional scene image is generated according to the label information and the position labeling information of the target three-dimensional scene image to obtain a plurality of labeling sets.

[0089] The basic information of the virtual target object model can include but is not limited to the name, center position, scaling size and other basic information of the target object. The configuration information of the rendering camera can include but is not limited to the name, position, focal length, resolution and other basic information of the rendering camera. The frame information of the target three-dimensional scene image can include but is not limited to the frame number, frame interval, frame rate and other basic information. In the present disclosure, one labeling set can be obtained for each target three-dimensional scene image, and then when the target three-dimensional scene image is multiple, a plurality of labeling sets can be obtained.

[0090] Then, the three-dimensional scene identifier corresponding to the plurality of three-dimensional scene images and the plurality of labeling sets are stored as labeling information files.

[0091] It should be understood that in the present disclosure, the plurality of three-dimensional scene images belong to the same three-dimensional scene, and the plurality of three-dimensional scene images under the three-dimensional scene can be obtained by only changing the position of the target object and the environmental object and the configuration information of the rendering camera, that is, the plurality of three-dimensional scene images mentioned above belong to the same three-dimensional scene, that is, the plurality of three-dimensional scene images have the same three-dimensional scene identifier. Therefore, in the present disclosure, the three-dimensional scene identifier corresponding to the plurality of three-dimensional scene images is stored in association with the plurality of labeling sets determined above.

[0092] For example, the pseudocode for associating and storing the 3D scene identifiers corresponding to multiple frames of 3D scene images with the multiple annotation sets determined above is shown below:

[0093]

[0094] In the pseudocode above, `meta_info` represents the basic information of the 3D scene corresponding to multiple frames of 3D scene images, i.e., the 3D scene identifier. `objects_basic_info` represents the basic information of the virtual target object model, `cameras_basic_info` represents the configuration information of the rendering camera, and `frames_basic_info` represents the frame information of the target 3D scene image.

[0095] In this disclosure, to ensure sufficient flexibility and ease of maintenance of the stored annotation information files, text-based formats such as YAML, JSON, and XML can be used for saving. This method facilitates subsequent secondary operations such as loading, appending, deleting, and adjusting annotations within a specified range.

[0096] Figure 2 This is a schematic diagram illustrating a virtual target object model, a virtual environment object model, and a rendering camera in a three-dimensional scene image according to an exemplary embodiment. Figure 2 As shown, the target object is a cube, and the environment object is a sphere. Figure 3 yes Figure 2 The image shown is a 2D image corresponding to the rendered camera. The pseudocode for the annotation process of the target object cube is shown below, and the resulting annotation information file is saved in YAML text format.

[0097]

[0098]

[0099]

[0100]

[0101]

[0102] The annotation method provided in this disclosure can simulate real-world environments to generate more diverse and complex annotation information, greatly improving the richness and quality of subsequent dataset creation and providing significant assistance and convenience for model algorithm training and testing. Furthermore, this annotation method can be further expanded to create various annotation tools, model training and testing tools, pipeline tools, and other applications.

[0103] Based on the same concept, the disclosure provides a target object labeling device. Figure 4 is a block diagram of a target object labeling device according to an exemplary embodiment. As shown in Figure 4 The target object labeling device 500 can include:

[0104] A first acquisition module 501 is configured to, in response to receiving a labeling request for a target object, acquire a target three-dimensional scene image, wherein the target three-dimensional scene image includes a virtual target object model created for the target object.

[0105] A second acquisition module 502 is configured to acquire first position coordinates of each vertex of the virtual target object model in a world coordinate system to which the virtual target object model belongs.

[0106] A first determination module 503 is configured to determine second position coordinates of a target vertex in a two-dimensional image corresponding to a rendering camera according to configuration information of the rendering camera and the first position coordinates of the each vertex in the target three-dimensional scene image.

[0107] A second determination module 504 is configured to determine position labeling information of the target three-dimensional scene image related to a target task according to the second position coordinates of the target vertex.

[0108] Optionally, the device 500 further includes:

[0109] A creation module is configured to, in response to receiving a request for constructing a three-dimensional scene image of a target object, create a virtual target object model for the target object and a virtual environment object model for a preset environment object.

[0110] A third acquisition module is configured to acquire animations of the virtual target object model and the virtual environment object model and configuration information of a rendering camera to obtain a plurality of three-dimensional scene images.

[0111] The first acquisition module includes:

[0112] A first determination sub-module is configured to, in response to receiving a labeling request for a target object, sequentially determine each three-dimensional scene image as a target three-dimensional scene image.

[0113] Optionally, the device 500 further includes:

[0114] The fourth acquisition module is configured to acquire, for each of the target three-dimensional scene images, basic information of the virtual target object model, configuration information of the rendering camera, and frame information of the target three-dimensional scene image as label information, and generate a label set of the target three-dimensional scene image according to the label information and the position labeling information of the target three-dimensional scene image to obtain a plurality of label sets;

[0115] The storage module is configured to store the three-dimensional scene identifiers corresponding to the plurality of frames of three-dimensional scene images and the plurality of label sets as label information files.

[0116] Optionally, the target vertex is each vertex; and the first determination module 503 includes:

[0117] The first acquisition submodule is configured to acquire, according to the configuration information of the rendering camera in the target three-dimensional scene image, a first script tool for converting three-dimensional coordinates into two-dimensional coordinates;

[0118] The first calling submodule is configured to call the first script tool to convert the first position coordinates of each vertex into second position coordinates of each vertex in a two-dimensional image corresponding to the rendering camera.

[0119] Optionally, the target vertex is a visible vertex in a two-dimensional image corresponding to the rendering camera.

[0120] The first determination module 503 includes:

[0121] The second acquisition submodule is configured to acquire, according to the configuration information of the rendering camera in the target three-dimensional scene image, a second script tool for determining the first position coordinates of the visible vertex in the first position coordinates of each vertex, and call the second script tool to obtain the first position coordinates of the visible vertex.

[0122] The second calling submodule is configured to call the first script tool for converting three-dimensional coordinates into two-dimensional coordinates to convert the first position coordinates of the visible vertex into second position coordinates in a two-dimensional image corresponding to the rendering camera.

[0123] Optionally, the target task is a segmentation task; and the second determination module 504 includes:

[0124] The second determination submodule is configured to determine, according to the second position coordinates of the target vertex, second position coordinates of a target boundary vertex of the virtual target object model in a two-dimensional image corresponding to the rendering camera.

[0125] The third determination submodule is configured to determine the second position coordinates of the target boundary vertex as position labeling information of the target three-dimensional scene image related to the segmentation task.

[0126] Optionally, the target task is a detection task, and the second determining module 504 comprises:

[0127] a fourth determining sub-module, configured to determine minimum and maximum values of a first coordinate axis and minimum and maximum values of a second coordinate axis respectively according to the second position coordinates of the target vertex;

[0128] a fifth determining sub-module, configured to determine the minimum value of the first coordinate axis and the minimum value of the second coordinate axis as position coordinates of a minimum value vertex, and determine the maximum value of the first coordinate axis and the maximum value of the second coordinate axis as position coordinates of a maximum value vertex;

[0129] a sixth determining sub-module, configured to determine the position coordinates of the minimum value vertex and the position coordinates of the maximum value vertex as position labeling information of the target three-dimensional scene image related to the detection task.

[0130] Optionally, the second obtaining module 502 comprises:

[0131] a third obtaining sub-module, configured to obtain a third script tool for traversing each vertex of the virtual target object model;

[0132] a traversing sub-module, configured to traverse each vertex of the virtual target object model according to the third script tool in the target three-dimensional scene image, to obtain first position coordinates of each vertex in a world coordinate system to which the virtual target object model belongs.

[0133] As to the apparatus in the above-described embodiments, the specific manners in which various modules perform operations have been described in details in the embodiments of the method, and thus will not be described in details here.

[0134] Based on the same idea, the embodiments of the present disclosure further provide a computer readable medium, having a computer program stored thereon, which is executed by a processing apparatus to implement the steps of any of the above target object labeling methods.

[0135] Based on the same idea, the embodiments of the present disclosure further provide an electronic device, comprising:

[0136] a storage apparatus, having at least one computer program stored thereon;

[0137] at least one processing apparatus, configured to execute the at least one computer program in the storage apparatus to implement the steps of any of the above target object labeling methods.

[0138] Reference will be made to the following Figure 5The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0139] like Figure 5 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0140] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0141] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.

[0142] It is noted that the aforementioned computer-readable medium of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example and without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a computer-readable program code transmitted by a computer-readable storage medium or carried by a carrier wave in a baseband or as part of a carrier wave. Such a propagated computer-readable signal medium can take various forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that can be used to carry or store a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to wire, cable, RF (radio frequency), or the like, or any suitable combination of the foregoing.

[0143] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.

[0144] The aforementioned computer-readable medium can be included in the aforementioned electronic device; or can exist separately from the electronic device, and can not be assembled in the electronic device.

[0145] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: in response to receiving a labeling request for a target object, acquire a target three-dimensional scene image, wherein the target three-dimensional scene image includes a virtual target object model created for the target object; acquire first position coordinates of each vertex of the virtual target object model in a world coordinate system to which the virtual target object model belongs; determine second position coordinates of a target vertex in a two-dimensional image corresponding to a rendering camera according to configuration information of the rendering camera in the target three-dimensional scene image and the first position coordinates of the vertices; and determine position labeling information of the target three-dimensional scene image related to a target task according to the second position coordinates of the target vertex.

[0146] Computer program code for carrying out operations of the present disclosure can be written in any of one or more programming languages, including object oriented programming languages such as Java, Smalltalk, C++, or conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0147] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in a different order than that noted in the figures. For example, two blocks noted in succession can in fact be executed substantially concurrently or in the opposite order, depending on the functionality involved. It is also noted that each block in the block diagrams and / or flow diagrams, and combinations of blocks in the block diagrams and / or flow diagrams, can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0148] The modules described in the embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. In some cases, the name of the module does not constitute a limitation on the module itself.

[0149] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0150] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0151] According to one or more embodiments of the present disclosure, example 1 provides a target object labeling method comprising:

[0152] In response to receiving a labeling request for a target object, a target three-dimensional scene image is obtained, wherein the target three-dimensional scene image includes a virtual target object model created for the target object;

[0153] The first position coordinates of each vertex of the virtual target object model in a world coordinate system to which the virtual target object model belongs are obtained;

[0154] According to the configuration information of the rendering camera in the target three-dimensional scene image and the first position coordinates of the vertices, the second position coordinates of the target vertices in the two-dimensional image corresponding to the rendering camera are determined;

[0155] According to the second position coordinates of the target vertices, position labeling information of the target three-dimensional scene image related to a target task is determined.

[0156] According to one or more embodiments of the present disclosure, example 2 provides the method of example 1, the method further comprising:

[0157] in response to receiving a request for constructing a three-dimensional scene image of a target object, creating a virtual target object model for the target object, and creating a virtual environment object model for a preset environment object;

[0158] obtaining animations of the virtual target object model and the virtual environment object model, and configuration information of a rendering camera, to obtain a plurality of three-dimensional scene images;

[0159] The method further includes, in response to receiving a labeling request for the target object, obtaining a target three-dimensional scene image.

[0160] The method further includes, in response to receiving a labeling request for the target object, sequentially determining each three-dimensional scene image as a target three-dimensional scene image.

[0161] According to one or more embodiments of the present disclosure, example 3 provides the method of example 2, the method further comprising:

[0162] For each of the target three-dimensional scene images, obtaining basic information of the virtual target object model, configuration information of the rendering camera, and frame information of the target three-dimensional scene image as label information, and generating a labeling set of the target three-dimensional scene image according to the label information and the position labeling information of the target three-dimensional scene image to obtain a plurality of labeling sets.

[0163] The three-dimensional scene identifiers corresponding to the plurality of three-dimensional scene images are stored in association with the plurality of labeling sets as a labeling information file.

[0164] According to one or more embodiments of the present disclosure, example 4 provides the method of example 1, the target vertex being each vertex; and the method further comprises:

[0165] According to the configuration information of the rendering camera in the target three-dimensional scene image, a first script tool for converting three-dimensional coordinates into two-dimensional coordinates is obtained.

[0166] The first script tool is called to convert the first position coordinates of the vertices into second position coordinates of the vertices in the two-dimensional image corresponding to the rendering camera.

[0167] According to one or more embodiments of the present disclosure, example 5 provides the method of example 1, the target vertex being a visible vertex in the two-dimensional image corresponding to the rendering camera.

[0168] determining, according to the configuration information of the rendering camera in the target three-dimensional scene image and the first position coordinates of the vertices, second position coordinates of the target vertices in a two-dimensional image corresponding to the rendering camera, includes:

[0169] According to the configuration information of the rendering camera in the target three-dimensional scene image, a second script tool for determining the first position coordinates of the visible vertices in the first position coordinates of the vertices is obtained, and the second script tool is called to obtain the first position coordinates of the visible vertices.

[0170] A first script tool for converting three-dimensional coordinates into two-dimensional coordinates is called to convert the first position coordinates of the visible vertices into second position coordinates in the two-dimensional image corresponding to the rendering camera.

[0171] According to one or more embodiments of the present disclosure, example 6 provides the method of any one of examples 1-5, the target task is a segmentation task; determining, according to the second position coordinates of the target vertices, position labeling information of the target three-dimensional scene image related to the target task, includes:

[0172] According to the second position coordinates of the target vertices, determining second position coordinates of target boundary vertices of the virtual target object model in the two-dimensional image corresponding to the rendering camera;

[0173] The second position coordinates of the target boundary vertices are determined as the position labeling information of the target three-dimensional scene image related to the task.

[0174] According to one or more embodiments of the present disclosure, example 7 provides the method of any one of examples 1-5, the target task is a detection task, and determining, according to the second position coordinates of the target vertices, position labeling information of the target three-dimensional scene image related to the target task, includes:

[0175] According to the second position coordinates of the target vertices, respectively determining minimum and maximum values of a first coordinate axis and minimum and maximum values of a second coordinate axis;

[0176] The minimum value of the first coordinate axis and the minimum value of the second coordinate axis are determined as the position coordinates of the minimum value vertex, and the maximum value of the first coordinate axis and the maximum value of the second coordinate axis are determined as the position coordinates of the maximum value vertex;

[0177] The position coordinates of the minimum value vertex and the position coordinates of the maximum value vertex are determined as the position labeling information of the target three-dimensional scene image related to the detection task.

[0178] According to one or more embodiments of the present disclosure, example 8 provides the method of any one of examples 1-5, wherein the obtaining the first position coordinates of each vertex of the virtual target object model in a world coordinate system to which the virtual target object model belongs comprises:

[0179] obtaining a third script tool for traversing each vertex of the virtual target object model;

[0180] traversing each vertex of the virtual target object model according to the third script tool in the target three-dimensional scene image to obtain the first position coordinates of each vertex in the world coordinate system to which the virtual target object model belongs.

[0181] According to one or more embodiments of the present disclosure, example 9 provides a target object labeling device, the device comprising:

[0182] a first obtaining module configured to, in response to receiving a labeling request for a target object, obtain a target three-dimensional scene image, wherein the target three-dimensional scene image includes a virtual target object model created for the target object;

[0183] a second obtaining module configured to obtain first position coordinates of each vertex of the virtual target object model in a world coordinate system to which the virtual target object model belongs;

[0184] a first determining module configured to determine second position coordinates of a target vertex in a two-dimensional image corresponding to a rendering camera according to configuration information of the rendering camera in the target three-dimensional scene image and the first position coordinates of each vertex;

[0185] a second determining module configured to determine position labeling information of the target three-dimensional scene image related to a target task according to the second position coordinates of the target vertex.

[0186] According to one or more embodiments of the present disclosure, example 10 provides the device of example 9, the device further comprising:

[0187] a creating module configured to, in response to receiving a request for constructing a three-dimensional scene image of a target object, create a virtual target object model for the target object and a virtual environment object model for a preset environment object;

[0188] a third obtaining module configured to obtain animations of the virtual target object model and the virtual environment object model and configuration information of a rendering camera to obtain a plurality of three-dimensional scene images;

[0189] the first obtaining module comprises:

[0190] A first determining sub-module is configured to, in response to receiving a labeling request for a target object, determine each frame of three-dimensional scene image as a target three-dimensional scene image in sequence.

[0191] According to one or more embodiments of the present disclosure, example 11 provides the apparatus of example 10, further comprising:

[0192] A fourth obtaining module is configured to, for each of the target three-dimensional scene images, obtain basic information of the virtual target object model, configuration information of the rendering camera, and frame information of the target three-dimensional scene image as label information, and generate a labeling set of the target three-dimensional scene image according to the label information and the position labeling information of the target three-dimensional scene image to obtain a plurality of labeling sets.

[0193] A storage module is configured to store the three-dimensional scene identification corresponding to the plurality of frames of three-dimensional scene images and the plurality of labeling sets as a labeling information file.

[0194] According to one or more embodiments of the present disclosure, example 12 provides the apparatus of example 9, wherein the target vertex is each vertex; and the first determining module comprises:

[0195] A first obtaining sub-module is configured to obtain a first script tool for converting three-dimensional coordinates into two-dimensional coordinates according to configuration information of a rendering camera in the target three-dimensional scene image.

[0196] A first calling sub-module is configured to call the first script tool to convert the first position coordinates of each vertex into second position coordinates of each vertex in a two-dimensional image corresponding to the rendering camera.

[0197] According to one or more embodiments of the present disclosure, example 13 provides the apparatus of example 9, wherein the target vertex is a visible vertex in a two-dimensional image corresponding to the rendering camera; and the first determining module comprises:

[0198] A second obtaining sub-module is configured to obtain a second script tool for determining the first position coordinates of the visible vertex in the first position coordinates of each vertex according to the configuration information of the rendering camera in the target three-dimensional scene image, and call the second script tool to obtain the first position coordinates of the visible vertex.

[0199] A second calling sub-module is configured to call a first script tool for converting three-dimensional coordinates into two-dimensional coordinates to convert the first position coordinates of the visible vertex into second position coordinates in a two-dimensional image corresponding to the rendering camera.

[0200] According to one or more embodiments of the present disclosure, example 14 provides the apparatus of any one of examples 9-13, the target task being a segmentation task; the second determining module comprising:

[0201] a second determining sub-module, configured to determine, according to the second position coordinates of the target vertex, second position coordinates of target boundary vertices of the virtual target object model in a two-dimensional image corresponding to the rendering camera;

[0202] a third determining sub-module, configured to determine the second position coordinates of the target boundary vertices as position annotation information of the target three-dimensional scene image related to the segmentation task.

[0203] According to one or more embodiments of the present disclosure, example 15 provides the apparatus of any one of examples 9-13, the target task being a detection task, the second determining module comprising:

[0204] a fourth determining sub-module, configured to determine, according to the second position coordinates of the target vertex, minimum and maximum values of a first coordinate axis and minimum and maximum values of a second coordinate axis, respectively;

[0205] a fifth determining sub-module, configured to determine the minimum value of the first coordinate axis and the minimum value of the second coordinate axis as position coordinates of a minimum value vertex, and determine the maximum value of the first coordinate axis and the maximum value of the second coordinate axis as position coordinates of a maximum value vertex;

[0206] a sixth determining sub-module, configured to determine the position coordinates of the minimum value vertex and the position coordinates of the maximum value vertex as position annotation information of the target three-dimensional scene image related to the detection task.

[0207] According to one or more embodiments of the present disclosure, example 16 provides the apparatus of any one of examples 9-13, the second obtaining module comprising:

[0208] a third obtaining sub-module, configured to obtain a third script tool for traversing each vertex of the virtual target object model;

[0209] a traversing sub-module, configured to traverse each vertex of the virtual target object model according to the third script tool in the target three-dimensional scene image to obtain first position coordinates of each vertex in a world coordinate system to which the virtual target object model belongs.

[0210] According to one or more embodiments of the present disclosure, example 17 provides a computer readable medium having stored thereon a computer program, which, when executed by a processing apparatus, implements the steps of the method of any one of examples 1-8.

[0211] According to one or more embodiments of the present disclosure, example 18 provides an electronic device comprising:

[0212] a storage device having stored thereon at least one computer program;

[0213] at least one processing device configured to execute the at least one computer program in the storage device to implement the steps of the method of any one of examples 1-8.

[0214] The above description is only preferred embodiments of the present disclosure and a description of the technical principles of the application. It should be understood by those skilled in the art that the disclosure range involved in the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by replacing the above features with the technical features disclosed in the present disclosure (but not limited to) having similar functions.

[0215] In addition, although each operation is described in a particular order, this should not be understood as requiring the operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multi-tasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be separated and implemented in multiple embodiments.

[0216] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. As for the devices in the above embodiments, the specific manner in which the various modules perform operations has been described in detail in the embodiments related to the method, and will not be described here.

Claims

1. A target object labeling method, characterized by comprising: The method comprises: in response to receiving a request for constructing a three-dimensional scene image of a target object, creating a virtual target object model for the target object, and creating a virtual environment object model for a preset environment object; obtaining animations of the virtual target object model and the virtual environment object model, and configuration information of a rendering camera, to obtain a plurality of frames of three-dimensional scene images; in response to receiving a labeling request for the target object, sequentially determining each frame of the three-dimensional scene image as a target three-dimensional scene image, wherein the target three-dimensional scene image comprises the virtual target object model created for the target object; obtaining first position coordinates of each vertex of the virtual target object model in a world coordinate system to which the virtual target object model belongs; determining second position coordinates of a target vertex in a two-dimensional image corresponding to the rendering camera according to the configuration information of the rendering camera in the target three-dimensional scene image and the first position coordinates of the vertex; determining position labeling information of the target three-dimensional scene image related to a target task according to the second position coordinates of the target vertex; for each target three-dimensional scene image, obtaining basic information of the virtual target object model, configuration information of the rendering camera, and frame information of the target three-dimensional scene image as label information, and generating a labeling set of the target three-dimensional scene image according to the label information and the position labeling information of the target three-dimensional scene image, to obtain a plurality of labeling sets; storing three-dimensional scene identifiers corresponding to the plurality of frames of three-dimensional scene images and the plurality of labeling sets as labeling information files.

2. The method of claim 1, wherein, The target vertex is the vertex; and the determining of the second position coordinates of the target vertex in the two-dimensional image corresponding to the rendering camera according to the configuration information of the rendering camera in the target three-dimensional scene image and the first position coordinates of the vertex comprises: obtaining a first script tool for converting three-dimensional coordinates into two-dimensional coordinates according to the configuration information of the rendering camera in the target three-dimensional scene image; calling the first script tool to convert the first position coordinates of the vertex into the second position coordinates of the vertex in the two-dimensional image corresponding to the rendering camera.

3. The method of claim 1, wherein, The target vertex is a visible vertex in the two-dimensional image corresponding to the rendering camera; and the determining of the second position coordinates of the target vertex in the two-dimensional image corresponding to the rendering camera according to the configuration information of the rendering camera in the target three-dimensional scene image and the first position coordinates of the vertex comprises: obtaining a second script tool for determining the first position coordinates of the visible vertex in the first position coordinates of the vertex according to the configuration information of the rendering camera in the target three-dimensional scene image, and calling the second script tool to obtain the first position coordinates of the visible vertex; calling a first script tool for converting three-dimensional coordinates into two-dimensional coordinates to convert the first position coordinates of the visible vertex into the second position coordinates in the two-dimensional image corresponding to the rendering camera. ​ 4. The method according to any one of claims 1 to 3, characterized in that, The target task is a segmentation task; and the determining, according to the second position coordinates of the target vertex, of position labeling information of the target three-dimensional scene image related to the target task comprises: determining, according to the second position coordinates of the target vertex, second position coordinates of a target boundary vertex of the virtual target object model in a two-dimensional image corresponding to the rendering camera; and determining the second position coordinates of the target boundary vertex as the position labeling information of the target three-dimensional scene image related to the segmentation task.

5. The method according to any one of claims 1-3, characterized in that, The target task is a detection task, and the determining, according to the second position coordinates of the target vertex, of position labeling information of the target three-dimensional scene image related to the target task comprises: determining, according to the second position coordinates of the target vertex, minimum and maximum values of a first coordinate axis and minimum and maximum values of a second coordinate axis, respectively; determining the minimum value of the first coordinate axis and the minimum value of the second coordinate axis as position coordinates of a minimum value vertex, and determining the maximum value of the first coordinate axis and the maximum value of the second coordinate axis as position coordinates of a maximum value vertex; and determining the position coordinates of the minimum value vertex and the position coordinates of the maximum value vertex as the position labeling information of the target three-dimensional scene image related to the detection task.

6. The method according to any one of claims 1-3, characterized in that, The obtaining of the first position coordinates of each vertex of the virtual target object model in a world coordinate system to which the virtual target object model belongs comprises: obtaining a third script tool for traversing each vertex of the virtual target object model; traversing each vertex of the virtual target object model according to the third script tool in the target three-dimensional scene image to obtain the first position coordinates of each vertex in the world coordinate system to which the virtual target object model belongs. 7.An object labeling device, characterized by comprising: The apparatus comprises: a creating module configured to, in response to receiving a request for constructing a three-dimensional scene image of a target object, create a virtual target object model for the target object and create a virtual environment object model for a preset environment object; a first obtaining module configured to obtain animations of the virtual target object model and the virtual environment object model and configuration information of a rendering camera to obtain a plurality of three-dimensional scene images; a second obtaining module configured to, in response to receiving a labeling request for the target object, sequentially determine each three-dimensional scene image as a target three-dimensional scene image, wherein the target three-dimensional scene image comprises the virtual target object model created for the target object; a third obtaining module configured to obtain first position coordinates of each vertex of the virtual target object model in a world coordinate system to which the virtual target object model belongs; a first determining module configured to determine, according to the configuration information of the rendering camera and the first position coordinates of each vertex, second position coordinates of a target vertex in a two-dimensional image corresponding to the rendering camera; a second determining module configured to determine, according to the second position coordinates of the target vertex, position labeling information of the target three-dimensional scene image related to a target task. A fourth acquisition module is configured to acquire, for each target three-dimensional scene image, basic information of the virtual target object model, configuration information of the rendering camera, and frame information of the target three-dimensional scene image as label information, and generate a label set of the target three-dimensional scene image according to the label information and the position label information of the target three-dimensional scene image to obtain a plurality of label sets; A storage module is configured to store the three-dimensional scene identification corresponding to the plurality of frames of three-dimensional scene images and the plurality of label sets as label information files.

8. A computer readable medium having stored thereon a computer program, characterized in that, The program is executed by the processing device to implement the steps of the method of any one of claims 1-6.

9. An electronic device, comprising: The program is executed by the processing device to implement the steps of the method of any one of claims 1-6. The program is executed by the processing device to implement the steps of the method of any one of claims 1-6. The program is executed by the processing device to implement the steps of the method of any one of claims 1-6. The program is executed by the processing device to implement the steps of the method of any one of claims 1-6.

Citation Information

Patent Citations

  • Sample image processing method and device, electronic equipment and storage medium

    CN112132213A

  • Three-dimensional virtual scene generation method and device

    CN112396688A