Method for labeling at least one specific object by automatically creating at least one specific pseudo 3D bounding box and computing device using the same

The method automates the generation of pseudo 3D bounding boxes by matching lidar and image data, addressing manual effort and mislabeling issues in existing tools, enhancing accuracy and efficiency for training data.

JP2025100459AActive Publication Date: 2025-07-03STRADVISION

Patent Information

Application Number
JP2024221796
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-12
Filing Date
2024-12-18
Publication Date
2025-07-03
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

Existing labeling tools for generating 3D bounding boxes require significant manual effort and suffer from mislabeling due to operator fatigue, leading to increased man-hours and reduced accuracy, especially in setting size and rotation information for 3D bounding boxes.

Method used

A method using lidar, image, and calibration data to automatically generate pseudo 3D bounding boxes by matching and fitting 3D bounding boxes to 2D bounding boxes, and generating pseudo boxes when mismatches occur, utilizing regression techniques and average size and rotation information.

Benefits of technology

Reduces manual effort and mislabeling, improving accuracy and efficiency in generating 3D bounding boxes by automating the process and enhancing the quality of training data for autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025100459000001_ABST
    Figure 2025100459000001_ABST
Patent Text Reader

Abstract

To provide a method and device for labeling specific objects by automatically creating specific pseudo 3D bounding boxes.SOLUTION: A method includes the steps of: acquiring at least one projected 3D LiDAR bounding boxes on an image coordinate system using raw data including LiDAR data, image data and calibration data; determining whether each of the projected 3D LiDAR bounding boxes matches each of 2D GT bounding boxes on the image coordinate system; if matches, performing regression on a first 3D LiDAR bounding box to fit the first 3D LiDAR bounding box into a first 2D GT bounding box; and if not matches, generating specific pseudo 3D bounding boxes to perform regression.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for labeling at least one specific object by automatically generating at least one specific pseudo 3D bounding box and a computing device using the same.

Background Art

[0002] Recently, research has been conducted on methods for identifying objects using machine learning. As part of such machine learning, deep learning using a neural network having several hidden layers between an input layer and an output layer has high identification performance.

[0003] And, a neural network using deep learning generally performs learning through backpropagation using loss. However, in order to perform learning of a deep learning network, training data with tags, that is, labels added to objects by a labeling tool is required. In order to prepare such training data, conventionally, a labeling tool for adding a simple working function such as generating a 3D bounding box using a mouse pointer or the like has been implemented, and thus there is an advantage that the man-hours for developing the labeling tool are not large.

[0004] However, when working using an existing labeling tool, since an operator directly uses a mouse pointer to generate a 3D bounding box, due to a decrease in the operator's concentration generated by repeating the work, mislabeling frequently occurs. In fact, even if pre-work related education is implemented for the operator, labeling-related mistakes are still confirmed, and thus there is a problem that additional man-hours are generated for separately collecting the result of the labeling work.

[0005] In addition, when generating a 2D bounding box, the operator can generate a box that includes the object by dragging using the mouse pointer while confirming the object on the image. However, in order to generate a 3D bounding box, the size information (width, height, length) and rotation information (roll, pitch, yaw) of the 3D bounding box should be appropriately set while confirming the three-dimensional coordinate information (x, y, z) corresponding to the object. However, because the difficulty of the corresponding operation is high, there is a problem that the working time per object is long.

[0006] Therefore, there is a situation where an improvement plan for solving the above problems is required.

Summary of the Invention

Problems to be Solved by the Invention

[0007] An object of the present invention is to solve the above problems.

[0008] In addition, the present invention uses raw data including rider data, image data, and calibration data to obtain each of at least one projected rider 3D bounding box on an image coordinate system. (i) If it is determined that each of the first projected rider 3D bounding boxes, which is at least one of the projected rider 3D bounding boxes, and each of the first 2D bounding boxes for GT, which is at least one of the 2D bounding boxes for GT, are matched, regression is performed on each of the first projected rider 3D bounding boxes to fit each of the first projected rider 3D bounding boxes to each of the first 2D bounding boxes for GT. (ii) If it is determined that each of the second 2D bounding boxes for GT, which is at least one of the 2D bounding boxes for GT, is not matched to any of the projected rider 3D bounding boxes, each of the specific pseudo 3D bounding boxes corresponding to each of the specific objects is generated by referring to each of the second 2D bounding boxes for GT and the average size information and rotation information of each of the corresponding specific objects, and regression is performed on each of the specific pseudo 3D bounding boxes to fit each of the specific pseudo 3D bounding boxes to each of the second 2D bounding boxes for GT for other purposes.

Means for Solving the Problems

[0009] According to an embodiment of the present invention, in a method of labeling at least one specific object by automatically generating at least one specific pseudo 3D bounding box, (a) a computing device uses raw data including lidar data, image data, and calibration data to obtain each of at least one projected lidar 3D bounding box on an image coordinate system, and determines whether each of the projected lidar 3D bounding boxes and each of the 2D bounding boxes for GT on the image coordinate system are mutually matched;and (b)(b1) if it is determined that each of at least one of the first projected lidar 3D bounding boxes among the projected lidar 3D bounding boxes and each of at least one of the first 2D bounding boxes for GT are matched, the computing device performs a regression on each of the first projected lidar 3D bounding boxes to fit each of the first projected lidar 3D bounding boxes to each of the first 2D bounding boxes for GT, and (b2) if it is determined that each of at least one of the second 2D bounding boxes for GT among the 2D bounding boxes for GT is not matched to any of the projected lidar 3D bounding boxes, the computing device generates each of the specific pseudo 3D bounding boxes corresponding to each of the specific objects with reference to each of the second 2D bounding boxes for GT, the average size information and rotation information of each of the specific objects corresponding to each of the second 2D bounding boxes for GT, and performs a regression on each of the specific pseudo 3D bounding boxes to fit each of the specific pseudo 3D bounding boxes to each of the second 2D bounding boxes for GT; A method is provided that includes the steps of;

[0010] In one example, in the step (a), the computing device inputs the lidar data included in the load data into a 3D lidar model to generate respective ones of at least one lidar 3D bounding box on a vehicle coordinate system with the 3D lidar model, projects respective ones of the lidar 3D bounding box onto the image coordinate system by referring to respective ones of the lidar 3D bounding box and the calibration data included in the load data to obtain respective ones of the projected lidar 3D bounding box, generates respective ones of a lidar 2D bounding box by referring to respective minimum values and maximum values of coordinates including respective ones of the projected lidar 3D bounding box, and applies a Hungarian algorithm using an IOU (Intersection Over Union) metric to determine whether respective ones of the lidar 2D bounding box and respective ones of the GT 2D bounding box match each other.

[0011] In one example, in the (b1) sub-process, the computing device generates respective first lidar 2D bounding boxes by referring to respective minimum values and maximum values of coordinates each including the respective first projected lidar 3D bounding boxes, calculates a first IOU loss which is a loss with respect to a matching ratio obtained by applying an IOU metric by referring to the respective first lidar 2D bounding boxes and the respective first 2D bounding boxes for GT, calculates a first center loss which is a loss generated by referring to each difference value between each median value of the respective first projected 2D bounding boxes and each median value of the respective first 2D bounding boxes for GT, generates a first integrated loss by referring to the first IOU loss and the first center loss, performs the regression by fitting the respective first lidar 2D bounding boxes by referring to the first integrated loss, and updates parameters of a 3D lidar model (the 3D lidar model is a model that generates respective at least one lidar 3D bounding box on a vehicle coordinate system when the lidar data included in the raw data is input).

[0012] In one example, in the (b2) sub-process, the computing device refers to each of the second GT 2D bounding boxes to obtain respective x-axis coordinate values and respective y-axis coordinate values for the position information of the 2D bounding boxes on the image coordinate system of each of the second GT 2D bounding boxes, refers to the depth information of the position information of the 2D bounding boxes on the lidar data included in the load data to obtain the z-axis coordinate value for the position information of the 2D bounding boxes on the image coordinate system, refers to the position information of the 2D bounding boxes to determine the position prediction information of the 3D bounding boxes for each of the specific objects, refers to each of the specific object classes corresponding to each of the specific objects confirmed through each of the second GT 2D bounding boxes, and obtains respective average widths, average lengths, and average heights corresponding to each of the specific object classes already calculated by the 3D lidar model from the 3D lidar model (the 3D lidar model is a model that inputs the lidar data included in the load data to generate respective lidar 3D bounding boxes on the vehicle coordinate system), determines the average size information for each of the specific objects with reference to the respective average widths, average lengths, and average heights corresponding to each of the specific object classes, inputs the image data in the load data into a Yaw regression model, and outputs the first rotation angle for each of the specific objects on the camera coordinate system using the Yaw regression model, determines the second rotation angle as the rotation information for each of the specific objects by converting the first rotation angle into the second rotation angle for each of the specific objects on the vehicle coordinate system with reference to the first rotation angle and the calibration data included in the load data, and generates each of the specific pseudo 3D bounding boxes with reference to the position prediction information of the 3D bounding boxes, the average size information for each of the specific objects, and the rotation information.

[0013] In one example, in the (b2) sub-process, the computing device inputs the image data into the Yaw regression model to output each of a plurality of features using the Yaw regression model, and after sequentially applying concatenation operation and convolution operation to each of the plurality of features, it is classified into a specific multi-bin class corresponding to the direction pointed by the front of the specific object among a plurality of multi-bin classes through a Yaw class header, and among each of the preset reference angles corresponding to each of the plurality of multi-bin classes, a third rotation angle which is the angle by which the specific object is rotated from a specific preset reference angle corresponding to the specific multi-bin class is obtained, and a specific code value corresponding to the direction pointed by the front of the specific object is obtained by referring to each of the plurality of preset code values based on each of the reference angles of each of the plurality of multi-bin classes through the Yaw delta header, and a third adjusted rotation angle is obtained by giving the specific code value to the third rotation angle, and the first rotation angle with respect to the specific object is estimated by referring to the third adjusted rotation angle and the specific reference angle.

[0014] In one example, in the (b2) sub-process, once each of the specific pseudo-3D bounding boxes is generated, the computing device calculates a matching ratio between each of the specific pseudo-3D bounding boxes and each of the second GT 2D bounding boxes using the IOU metric, and if it is determined that the matching ratio is less than a preset threshold ratio, regression is performed on each of the specific pseudo-3D bounding boxes.

[0015] In one example, in the (b2) sub-process, the computing device generates respective specific pseudo 2D bounding boxes by referring to respective minimum values and maximum values of coordinates each including the respective specific pseudo 3D bounding box, calculates a second IOU loss which is a loss with respect to a matching ratio obtained by applying the IOU metric by referring to the respective specific pseudo 2D bounding boxes and the respective second GT 2D bounding boxes, calculates a second center loss which is a loss generated by referring to each difference value between each center value of the respective specific pseudo 2D bounding boxes and each center value of the respective second GT 2D bounding boxes, generates a second integrated loss by referring to the second IOU loss and the second center loss, and performs the regression by fitting the respective specific pseudo 2D bounding boxes by referring to the second integrated loss.

[0016] In one example, in the (b2) sub-process, if it is determined that each of the second projected lidar 3D bounding boxes, which are at least one of the projected lidar 3D bounding boxes, is not matched to any of the GT 2D bounding boxes, the computing device deletes each of the second projected lidar 3D bounding boxes.

[0017] In one example, in the step (a), the calibration data includes intrinsic parameters of a camera, extrinsic parameters, and conversion parameters for converting the lidar data into the image data.

[0018] In one example, the 2D bounding box for GT is obtained from data in which the image data generated by a predetermined camera is input, the specific object is detected through a deep learning model, and a 2D bounding box is generated for the specific object.

[0019] According to another embodiment of the present invention, in a computing device that labels at least one specific object by automatically generating at least one specific pseudo 3D bounding box, at least one memory for storing instructions; and at least one processor configured to execute the instructions, the processor (I) utilizes raw data including lidar data, image data, and calibration data to obtain each of at least one projected lidar 3D bounding box on an image coordinate system, and determines whether each of the projected lidar 3D bounding boxes and each of the 2D bounding boxes for GT on the image coordinate system are mutually matched;And (II) (II-1) If it is determined that each of the first projected lidar 3D bounding boxes, which is at least one of the projected lidar 3D bounding boxes, and each of the first 2D bounding boxes for GT, which is at least one of the 2D bounding boxes for GT, are matched, a sub-process of performing regression on each of the first projected lidar 3D bounding boxes to fit each of the first projected lidar 3D bounding boxes to each of the first 2D bounding boxes for GT, and (II-2) If it is determined that each of the second 2D bounding boxes for GT, which is at least one of the 2D bounding boxes for GT, is not matched to any of the projected lidar 3D bounding boxes, each of the second 2D bounding boxes for GT, each of the specific objects corresponding to each of the second 2D bounding boxes for GT, and the average size information and rotation information of each of the specific objects are referred to, and each of the specific pseudo-3D bounding boxes corresponding to each of the specific objects is generated, and a sub-process of performing the regression on each of the specific pseudo-3D bounding boxes to fit each of the specific pseudo-3D bounding boxes to each of the second 2D bounding boxes for GT; A computing device is provided that performs the process;

[0020] In one example, in the (I) process, the processor inputs the lidar data included in the load data into a 3D lidar model to generate each of at least one lidar 3D bounding box on a vehicle coordinate system, and projects each of the lidar 3D bounding boxes onto the image coordinate system with reference to each of the lidar 3D bounding boxes and the calibration data included in the load data, thereby obtaining each of the projected lidar 3D bounding boxes, generates each of the lidar 2D bounding boxes with reference to each of the minimum values and the maximum values of the coordinates including each of the projected lidar 3D bounding boxes, and applies a Hungarian algorithm using an IOU (Intersection Over Union) metric to determine whether each of the lidar 2D bounding boxes and each of the GT 2D bounding boxes match each other.

[0021] In one example, in the (II-1) sub-process, the processor generates each of the first lidar 2D bounding boxes by referring to each of the minimum values and each of the maximum values of the coordinates including each of the first projected lidar 3D bounding boxes, and calculates a first IOU loss, which is a loss with respect to the matching ratio obtained by applying the IOU metric by referring to each of the first lidar 2D bounding boxes and each of the first GT 2D bounding boxes. The processor calculates a first center loss, which is a loss generated by referring to each difference value between each median value of each of the first projected 2D bounding boxes and each median value of each of the first GT 2D bounding boxes, generates a first integrated loss by referring to the first IOU loss and the first center loss, performs the regression by fitting each of the first lidar 2D bounding boxes by referring to the first integrated loss, and updates the parameters of a 3D lidar model (the 3D lidar model is a model that generates each of at least one lidar 3D bounding box on a vehicle coordinate system when the lidar data included in the low data is input) through backpropagation using the first integrated loss.

[0022] In one example, in the (II-2) sub-process, the processor refers to each of the second GT 2D bounding boxes to obtain the x-axis coordinate values and y-axis coordinate values for the position information of the 2D bounding boxes on the image coordinate system of each of the second GT 2D bounding boxes, refers to the depth information of the position information of the 2D bounding boxes on the lidar data included in the load data to obtain the z-axis coordinate value for the position information of the 2D bounding boxes on the image coordinate system, refers to the position information of the 2D bounding boxes to determine the position prediction information of the 3D bounding boxes for each of the specific objects, refers to each of the specific object classes corresponding to each of the specific objects confirmed through each of the second GT 2D bounding boxes, and obtains, from a 3D lidar model (the 3D lidar model is a model that generates each of at least one lidar 3D bounding box on the vehicle coordinate system when the lidar data included in the load data is input), the average width, average length, and average height corresponding to each of the specific object classes that have been calculated by the 3D lidar model, refers to the average width, average length, and average height corresponding to each of the specific object classes to determine the average size information for each of the specific objects, inputs the image data in the load data into a Yaw regression model, and outputs, using the Yaw regression model, the first rotation angle for each of the specific objects on the camera coordinate system, and determines the second rotation angle as the rotation information for each of the specific objects by converting the first rotation angle to the second rotation angle for each of the specific objects on the vehicle coordinate system with reference to the first rotation angle and the calibration data included in the load data, and generates each of the specific pseudo 3D bounding boxes with reference to the position prediction information of the 3D bounding boxes, the average size information for each of the specific objects, and the rotation information.

[0023] In one example, in the (II-2) sub-process, the processor inputs the image data into the Yaw regression model to output each of a plurality of features using the Yaw regression model, and after sequentially applying a concatenation operation and a convolution operation to each of the plurality of features, the processor classifies the object into a specific multi-bin class corresponding to the direction pointed by the front of the specific object among a plurality of multi-bin classes through a Yaw class header, obtains a third rotation angle, which is the angle by which the specific object is rotated from a specific reference angle preset corresponding to the specific multi-bin class among each of the preset reference angles corresponding to each of the plurality of multi-bin classes, obtains a specific code value corresponding to the direction pointed by the front of the specific object by referring to each of the plurality of code values preset based on each of the reference angles of each of the plurality of multi-bin classes through the Yaw delta header, obtains a third adjusted rotation angle by giving the specific code value to the third rotation angle, and estimates the first rotation angle with respect to the specific object by referring to the third adjusted rotation angle and the specific reference angle.

[0024] In one example, in the (II-2) sub-process, when each of the specific pseudo 3D bounding boxes is generated, the processor calculates a matching ratio between each of the specific pseudo 3D bounding boxes and each of the second GT 2D bounding boxes using an IOU metric, and if it is determined that the matching ratio is less than a preset threshold ratio, the processor performs the regression on each of the specific pseudo 3D bounding boxes.

[0025] In one example, in the (II-2) sub-process, the processor generates respective specific pseudo 2D bounding boxes by referring to respective minimum values and maximum values of coordinates each including the respective specific pseudo 3D bounding box, calculates a second IOU loss which is a loss with respect to the matching ratio obtained by applying the IOU metric by referring to the respective specific pseudo 2D bounding boxes and the respective second GT 2D bounding boxes, calculates a second center loss which is a loss generated by referring to each difference value between each center value of each of the respective specific pseudo 2D bounding boxes and each center value of each of the respective second GT 2D bounding boxes, generates a second integrated loss by referring to the second IOU loss and the second center loss, and performs the regression by fitting each of the respective specific pseudo 2D bounding boxes by referring to the second integrated loss.

[0026] In one example, in the (II-2) sub-process, if the processor determines that each of the second projected lidar 3D bounding boxes, which is at least one of the projected lidar 3D bounding boxes, is not matched to any of the GT 2D bounding boxes, the processor deletes each of the second projected lidar 3D bounding boxes.

[0027] In one example, in the (I) process, the calibration data includes the intrinsic parameters of the camera, the extrinsic parameters, and the conversion parameters for converting the lidar data into the image data.

[0028] In one example, the 2D bounding box for GT is obtained from data in which the image data generated by a predetermined camera is input, the specific object is detected through a deep learning model, and the 2D bounding box is generated for the specific object.

Effect of the Invention

[0029] The present invention uses raw data including lidar data, image data, and calibration data to obtain each of at least one projected lidar 3D bounding box on an image coordinate system, and (i) if it is determined that each of the first projected lidar 3D bounding boxes, which is at least one of the projected lidar 3D bounding boxes, and each of the first 2D bounding boxes for GT, which is at least one of the 2D bounding boxes for GT, are matched, regression is performed on each of the first projected lidar 3D bounding boxes to fit each of the first projected lidar 3D bounding boxes to each of the first 2D bounding boxes for GT, and (ii) if it is determined that each of the second 2D bounding boxes for GT, which is at least one of the 2D bounding boxes for GT, is not matched to any of the projected lidar 3D bounding boxes, each of the specific pseudo 3D bounding boxes corresponding to each of the specific objects is generated with reference to each of the second 2D bounding boxes for GT and the average size information and rotation information of each of the corresponding specific objects, and regression is performed on each of the specific pseudo 3D bounding boxes to fit each of the specific pseudo 3D bounding boxes to each of the second 2D bounding boxes for GT.

Brief Description of the Drawings

[0030] The following drawings attached for use in the description of embodiments of the present invention are merely a part of the embodiments of the present invention, and for those having ordinary knowledge in the technical field to which the present invention pertains (hereinafter referred to as "ordinary technicians"), other drawings can be obtained based on these drawings without inventive work being performed.

[0031]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Mode for Carrying Out the Invention

[0032] The detailed description of the present invention to be described later refers to the accompanying drawings which show, as an example, specific embodiments in which the present invention can be implemented in order to clarify each object, each technical solution, and each advantage of the present invention. These embodiments are described in sufficient detail so that an ordinary technician can implement the present invention. Also, throughout the detailed description and each claim of the present invention, the word "comprising" and its variations are not intended to exclude other technical features, additional elements, components or steps. For an ordinary technician, other objects, advantages and characteristics of the present invention will become apparent partly from this specification and partly from the implementation of the present invention. The following examples and drawings are provided as examples and are not intended to limit the present invention.

[0033]

[0034] ​Furthermore, the present invention encompasses all possible combinations of the embodiments shown herein. It should be understood that the various embodiments of the present invention are different from each other but do not necessarily have to be mutually exclusive. For example, the specific shapes, structures, and characteristics described herein can be embodied in other embodiments without departing from the spirit and scope of the present invention in relation to one embodiment. Also, it should be understood that the position or arrangement of individual components within each disclosed embodiment can be changed without departing from the spirit and scope of the present invention. Therefore, the following detailed description is not to be taken in a limiting sense, and the scope of the present invention is limited only by the appended claims together with all scopes equivalent to what the claims claim, provided that they are properly described. Similar reference numerals in the drawings are the same or refer to similar functions across various aspects.

[0035] Hereinafter, for the purpose of enabling a person having ordinary skill in the art to which the present invention pertains to easily implement the present invention, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings.

[0036] FIG. 1 schematically shows a computing device that labels at least one specific object by automatically generating at least one specific pseudo-3D bounding box according to an embodiment of the present invention.

[0037] Referring to FIG. 1, the computing device 1000 can include a memory 1010 that stores instructions for labeling at least one specific object by automatically generating at least one specific pseudo-3D bounding box, and a processor 1020 that labels at least one specific object by automatically generating at least one specific pseudo-3D bounding box corresponding to the instructions stored in the memory 1010. At this time, the computing device 1000 can include a PC (Personal Computer), a mobile computer, and the like.

[0038] Specifically, the computing device 1000 may achieve the desired system performance by using a combination of a typical computing device (e.g., a device that can include a computer processor, memory, storage, input device, output device, and other components of an existing computing device; an electronic communication device such as a router or a switch; an electronic information storage system such as a network-attached storage (NAS) and a storage area network (SAN)) and computer software (i.e., instructions that cause the computing device to function in a specific manner).

[0039] In addition, the processor of the computing device may include hardware components such as an MPU (Micro Processing Unit) or a CPU (Central Processing Unit), a cache memory, and a data bus. The computing device may further include an operation system and a software configuration of an application for performing a specific purpose.

[0040] However, the case where the computing device includes an integrated processor in which the media processor and the memory for implementing the present invention are integrated is not excluded.

[0041] And, the computing device 1000 can be associated with a database 900 that includes information used to label at least one specific object by automatically generating at least one specific pseudo 3D bounding box. Here, the database 900 can include at least one type of storage medium among flash memory type, hard disk type, multimedia card micro type, card type memory (e.g., SD or XD memory), Random Access Memory (RAM), Static Random Access Memory (SRAM), Read Only Memory (ROM), Electrically Erasable Programmable Read Only Memory (EEPROM), Programmable Read Only Memory (PROM), magnetic memory, magnetic disk, optical disk, and is not limited thereto, and can include all media capable of storing data. Also, the database 900 can be installed separately from the computing device 100, or, conversely, can be installed inside the computing device 100 to transmit data or record the received data, and, different from what is shown in the figure, can also be embodied separately into two or more, which may vary depending on the implementation conditions of the invention.

[0042] FIG. 2 schematically shows a process of labeling at least one specific object by automatically generating at least one specific pseudo 3D bounding box according to an embodiment of the present invention.

[0043] First, referring to FIG. 2, the computing device 1000 can utilize raw data including image data, lidar data, and calibration data to obtain each of at least one projected lidar 3D bounding box on the image coordinate system, and determine (S100) whether each of the projected lidar 3D bounding boxes and each of the 2D bounding boxes for GT on the image coordinate system are mutually matched.

[0044] As an example, the computing device 1000 can input the lidar data included in the raw data into a 3D lidar model, and use the 3D lidar model to generate and output each of the lidar 3D bounding boxes for each of the objects included in the lidar data through calculation. The image data can be input into SVNet3 (a deep learning model that inputs image data generated by a camera to detect an object and generates a 2D bounding box for the object, which has high detection accuracy and can be utilized as GT), and SVNet3 can be used to generate and output each of the 2D bounding boxes for GT for each of the objects in the image coordinate system. However, the subject for generating each of the 2D bounding boxes for GT from the image data is not limited to SVNet3, and each of the 2D bounding boxes for GT can also be generated by other known methods that can generate a bounding box with higher accuracy than the lidar model, and this can be treated as GT.

[0045] As another example, in a state where the computing device 100 stores and manages in the database 900 each of the lidar 3D bounding boxes and each of the 2D bounding boxes for GT, which are obtained as a result of causing the lidar 3D model and SVNet3 to output each of the lidar 3D bounding boxes and each of the 2D bounding boxes for GT in advance, it is also possible to obtain from the database 900 each of the lidar 3D bounding boxes and each of the 2D bounding boxes for GT.

[0046] For reference, the calibration data may include the intrinsic and extrinsic parameters of the camera, parameters for converting lidar data into image data, etc. The image coordinate system may mean an ICS (Image Coordinate System), which is a two-dimensional coordinate system with the upper left corner of the image obtained by the camera as the origin, but is not limited thereto.

[0047] In addition, the computing device 1000 can confirm whether matching is performed by making a pair and checking each of the lidar 3D bounding boxes output from the 3D lidar model and each of the 2D bounding boxes for GT on the image data (that is, the data output from the SVNet3 model with the image data acquired from the camera input). At this time, in order for the computing device 1000 to confirm whether each of the lidar 3D bounding boxes is matched with the 2D bounding box for GT in a paired manner, by projecting the lidar 3D bounding box onto the image coordinate system, each of the lidar 3D bounding boxes projected onto the image coordinate system can be obtained, and it may be to confirm whether each of the lidar 3D bounding boxes projected in this way is matched with each of the 2D bounding boxes for GT in a paired manner, but it is not limited thereto.

[0048] At this time, an example regarding the lidar 3D bounding box output from the 3D lidar model will be described with reference to FIG. 3.

[0049] FIG. 3 briefly shows an example of the lidar 3D bounding box output from the 3D lidar model according to an embodiment of the present invention.

[0050] Referring to FIG. 3, the 3D bounding boxes 301, 302, 311, 313, 315, 321 shown in (a) to (c) may be in a state where the lidar 3D bounding boxes included in the lidar data output through the 3D lidar model are projected onto the image coordinate system using the calibration data to obtain the lidar 3D bounding boxes. For reference, it is clarified that in FIGS. 3(a) and 3(c), for the sake of convenience of explanation, the GT 2D bounding boxes that exactly match the projected lidar 3D bounding boxes are not shown.

[0051] In FIG. 3(a), it can be confirmed that the lidar 3D bounding boxes 301, 302 projected for each of the person and the vehicle on the image coordinate system are generated. In (b), since the calibration data included in the raw data is not accurate, the lidar 3D bounding boxes 311, 313 projected with a certain part separated from the GT 2D bounding boxes 312, 314 for each of the person and the vehicle on the image coordinate system are obtained. By checking the right side of the GT 2D bounding box 314 for the vehicle, it can be confirmed that the lidar 3D bounding box 315 projected for the misdetected (FP, False Positive) object is obtained. In (c), the lidar 3D bounding box 315 projected for the object located at a short distance is generated, but due to the characteristics of the lidar data where less point cloud data is shown for other objects located farther away, the vehicle 322 as an object is undetected (FN, False Negative), and it can be confirmed that the lidar 3D bounding box projected for the vehicle 322 is not obtained.

[0052] That is, in the case of the projected lidar 3D bounding boxes 311, 313, 315, 321 projected as shown in FIGS. 3(b) and 3(c), the 3D bounding boxes 311, 313 are displayed at a certain distance from the corresponding GT 2D bounding box, and a projected 3D bounding box 315 for a misdetected object is generated, or a problem occurs where a projected 3D bounding box is not generated because the object is not detected. In such a case, since it is not suitable for use as learning data for an autonomous vehicle, there is a need to solve this problem.

[0053] Referring again to FIG. 2, when the computing device 1000 checks the matching results for each of the projected lidar 3D bounding boxes and each of the GT 2D bounding boxes, three cases may occur. Describing the process for processing each of the three cases is as follows.

[0054] As an example, the computing device 1000 can determine (S200_1) that each of at least one first projected lidar 3D bounding box among the projected lidar 3D bounding boxes and each of at least one first GT 2D bounding box among the GT 2D bounding boxes are matched.

[0055] That is, there is a case where there is at least a partially overlapping portion between each of the first projected lidar 3D bounding boxes and each of the first 2D bounding boxes for GT, and the positions of each of the first projected lidar 3D bounding boxes and the positions of each of the first 2D bounding boxes for GT do not exactly match due to reasons such as inaccurate calibration data. However, it can be determined that each of the first objects corresponding to each of the first projected lidar 3D bounding boxes and each of the first 2D bounding boxes for GT is matched as the same object.

[0056] Therefore, if it is determined that each of the first projected lidar 3D bounding boxes and each of the first 2D bounding boxes for GT are matched, the computing device 1000 can perform regression (S210_1) on each of the first projected lidar 3D bounding boxes.

[0057] At this time, as described above, since the 2D bounding box for GT has high detection accuracy and can be used as GT, the computing device 1000 may perform regression on each of the first projected 3D bounding boxes so that each of the first projected 3D bounding boxes is fitted to each of the first 2D bounding boxes for GT. For reference, the process of performing regression on each of the first projected 3D bounding boxes will be specifically described later with reference to FIG. 5.

[0058] As another example, it can be determined (S200_2) that each of the second 2D bounding boxes for GT, which is at least one of the 2D bounding boxes for GT in the computing device 1000, does not match any of the projected lidar 3D bounding boxes.

[0059] That is, each specific object corresponding to each of the second 2D bounding boxes for GT has been detected, but at least a part of the projected lidar 3D bounding boxes that should be generated corresponding to each specific object has not been generated. In this case, among each specific object corresponding to each of the second 2D bounding boxes for GT, for an object for which the projected lidar 3D bounding box has not been generated, it can be determined that it is an undetected object. This is considered to occur because the ability to detect objects located at a long distance is insufficient due to the characteristics of the lidar data described in FIG. 3 above.

[0060] Therefore, the computing device 1000 can generate (S210_2) each of the pseudo 3D bounding boxes corresponding to each specific object by referring to each of the second 2D bounding boxes for GT, the average size information and the rotation information of each corresponding specific object.

[0061] Specifically, the computing device 1000 generates position prediction information of the 3D bounding box corresponding to each specific object by referring to the position information of the 2D bounding box corresponding to each of the second 2D bounding boxes for GT. After confirming the specific object class corresponding to each specific object, the computing device 1000 can obtain the average size information corresponding to each specific object that has been calculated by the 3D lidar model from the 3D lidar model, and obtain the rotation information corresponding to each specific object from the Yaw regression model.

[0062] That is, even if a lidar 3D bounding box corresponding to a specific object has not been generated by the 3D lidar model, by referring to the position prediction information, average size information, and rotation information of the 3D bounding box corresponding to each specific object (i.e., undetected object), it is possible to substitute by generating a specific pseudo 3D bounding box corresponding to each specific object. For reference, the meaning of referring to the average size information and rotation information does not mean directly applying the actual average size of a specific object to generate a specific pseudo 3D bounding box. Instead, by referring to the depth information, the position prediction information of the 3D bounding box corresponding to the position information of the 2D bounding box is obtained, and by referring to such depth information, the actual average size of a specific object is adjusted by perspective projection and treated as the average size information to generate a specific pseudo 3D bounding box.

[0063] Thereafter, the computing device 1000 can perform regression on each of the specific pseudo 3D bounding boxes (S220_2).

[0064] At this time, the computing device 1000 calculates the matching ratio between each of the specific pseudo 3D bounding boxes generated using the IOU (Intersection Over Union) metric and each of the second GT 2D bounding boxes, and refers to the matching ratio to grasp the degree to which the two (e.g., their centers) are separated. If it is determined that the ratio is less than a preset threshold ratio (e.g., it may be 100%, but not limited to this), regression may be performed on each of the specific pseudo 2D bounding boxes so that each of the specific pseudo 3D bounding boxes fits each of the second GT 2D bounding boxes. For reference, the process of performing regression on each of the specific pseudo 2D bounding boxes will be described later.

[0065] As another example, the computing device 1000 can determine (S200_3) that each of the second projected lidar 3D bounding boxes, which is at least one of the projected lidar 3D bounding boxes, does not match any of the GT 2D bounding boxes.

[0066] That is, each of the second projected lidar 3D bounding boxes is generated, but in a case where no GT 2D bounding box containing each of the same specific objects as each of the second projected lidar 3D bounding boxes is generated. In this case, each of the specific objects corresponding to each of the second projected lidar 3D bounding boxes can be determined to be a misdetected object. This is considered to occur because the calibration data included in the raw data described in FIG. 3 above is not accurate.

[0067] Therefore, by deleting (S210_3) each of the second projected lidar 3D bounding boxes, the computing device 1000 can improve the accuracy of labeling for generating learning data.

[0068] In this way, the process for solving the problems in FIGS. 3(b) and 3(c) will be described in more detail with reference to FIGS. 4 to 8.

[0069] FIG. 4 schematically shows a process of checking whether each of the lidar 3D bounding boxes output by the 3D lidar model according to an embodiment of the present invention matches each of the GT 2D bounding boxes in the image coordinate system.

[0070] Referring to FIG. 4, first, the computing device 1000 inputs the lidar data included in the raw data into the 3D lidar model to generate each of at least one lidar 3D bounding box on the vehicle coordinate system with the 3D lidar model (S110), and projects each of the lidar 3D bounding boxes onto the image coordinate system with reference to each of the lidar 3D bounding boxes and the calibration data included in the raw data, so as to obtain each of the projected lidar 3D bounding boxes (S120).

[0071] At this time, the vehicle coordinate system may mean a 3D coordinate system VCS (Vehicle Coordinate System) with the center position of the vehicle as the origin, but is not limited thereto.

[0072] In addition, the computing device 1000 can generate each of the lidar 2D bounding boxes corresponding to (e.g., surrounding) each of the projected lidar 3D bounding boxes with reference to the respective minimum values and maximum values of the coordinates including each of the projected lidar 3D bounding boxes (S130), and use the IoU metric for each of the lidar 2D bounding boxes and each of the GT 2D bounding boxes and apply it to the Hungarian algorithm to determine whether each of the lidar 2D bounding boxes and each of the GT 2D bounding boxes are mutually matched (S140).

[0073] The computing device 1000 can output a predetermined value between "0" and "1" as a result of matching each of the rider 2D bounding boxes and each of the 2D bounding boxes for GT. If it is "0", since there is no matching part, it can be determined that the object has been misdetected or not detected from the 3D rider model. If it is a predetermined value between "0.01" and "0.99", it can be determined that the same object is in a state where a certain part is matched. If it is "1", it can be determined that the object is completely matched (i.e., the positions coincide), but it is not limited thereto.

[0074] Among these, for the case where at least a part of each of the projected rider 3D bounding boxes and each of the 2D bounding boxes for GT is matched, the process of fitting each of the projected rider 3D bounding boxes to the 2D bounding boxes for GT will be described later with reference to FIG. 5. For the case where each of the projected rider 3D bounding boxes and each of the 2D bounding boxes for GT are not matched at all, the case where the object is misdetected, and the case where the object is not detected, the process of generating a pseudo 3D bounding box for the case where the object is not detected will be described later with reference to FIG. 6.

[0075] FIG. 5 schematically shows the process of regression such that each of the projected rider 3D bounding boxes matched according to an embodiment of the present invention is fitted to each of the 2D bounding boxes for GT in the image coordinate system.

[0076] Referring to FIG. 5, the computing device 1000 can generate each of the first rider 2D bounding boxes by referring to each of the minimum values and each of the maximum values of the coordinates corresponding to each of the first projected rider 3D bounding boxes (S211_1).

[0077] This is because it is easier to generate each of the first lidar 2D bounding boxes in the form of 2D bounding boxes by referring to each of the first projected lidar 3D bounding boxes, where each of the first projected lidar 3D bounding boxes is at least one of the projected lidar 3D bounding boxes in the projected lidar 3D bounding box, and fit each of the first lidar 2D bounding boxes to each of the first GT 2D bounding boxes, where each of the first GT 2D bounding boxes is at least one of the GT 2D bounding boxes in the GT 2D bounding box.

[0078] Also, the computing device 1000 can calculate a first IOU loss and a first center loss, generate a first integrated loss with reference to the first IOU loss and the first center loss (S212_1), and perform regression (S213_1) by fitting each of the first lidar 2D bounding boxes with reference to the first integrated loss. This is to complement the problem that when the calibration data included in the raw data is not accurate, each of the first projected lidar 3D bounding boxes, where each of the first projected lidar 3D bounding boxes is at least one of the projected lidar 3D bounding boxes generated by projecting each of the lidar 3D bounding boxes output from the 3D lidar model into the image coordinate system, is shown separated from each of the first GT 2D bounding boxes, where each of the first GT 2D bounding boxes is at least one of the GT 2D bounding boxes in the GT 2D bounding box.

[0079] Specifically, the computing device 1000 calculates an IOU loss, which is a loss with respect to a matching ratio obtained by applying an IOU metric with reference to each of the first lidar 2D bounding boxes and each of the first GT 2D bounding boxes, and calculates a center loss with reference to an L1 loss value using an error obtained by taking the absolute value of each difference value between each median value of each of the first lidar 2D bounding boxes and each median value of each of the first GT 2D bounding boxes. The first integrated loss can be generated by summing the IOU loss and the center loss, but is not limited thereto. Then, the computing device 1000 may perform regression by fitting each frame of each of the first lidar 2D bounding boxes to match each frame of each of the first GT 2D bounding boxes with reference to the first integrated loss.

[0080] Further, the computing device 1000 can update (S214_1) the parameters of the 3D lidar model through backpropagation using the first integrated loss, and can repeat the update of the parameters of the 3D lidar model at least N times (for example, 500 times) for learning.

[0081] FIG. 6 schematically shows a process of generating each specific pseudo 3D bounding box for each specific object for which each of the GT 2D bounding boxes was generated but the lidar 3D bounding box was not generated according to an embodiment of the present invention.

[0082] Referring to FIG. 6, the computing device 1000 can determine (S211_2_1) the x-axis coordinate value, y-axis coordinate value, and z-axis coordinate value corresponding to each of the second 2D bounding boxes for GT, which are at least one of the 2D bounding boxes for GT, as prediction information of the 3D bounding box for each specific object. The computing device 1000 can refer to the average width, average length, and average height for each specific object class corresponding to each specific object that has already been calculated by the 3D lidar model from the 3D lidar model to determine (S211_2_2) the average size information for each specific object. The computing device 1000 can use the Yaw regression model to estimate the first rotation angle for each specific object in the camera coordinate system, and determine (S211_2_3) the second rotation angle for each specific object in the vehicle coordinate system that is converted by referring to the first rotation angle and calibration data as the rotation information for each specific object.

[0083] For example, the computing device 1000 can refer to each of the second 2D bounding boxes for GT, which are at least one of the 2D bounding boxes for GT, to obtain the x-axis coordinate value and y-axis coordinate value for the position information of the 2D bounding box in the image coordinate system (i.e., the position information for each specific object). However, to generate a specific pseudo 3D bounding box, it is necessary to know the z-axis coordinate value. Therefore, the computing device 1000 can obtain the z-axis coordinate value for the position information of the 2D bounding box in the image coordinate system by referring to the depth information for the position information of the 2D bounding box on the lidar data included in the raw data. Then, the computing device 1000 can determine the position prediction information of the 3D bounding box for each specific object by referring to the x-axis coordinate value, y-axis coordinate value, and z-axis coordinate value corresponding to the position information of the 2D bounding box in the image coordinate system. At this time, each of the x-axis coordinate value, y-axis coordinate value, and z-axis coordinate value may be determined as a range value.

[0084] In addition, when the 3D lidar model repeats the process of generating each of the lidar 3D bounding boxes on the lidar data with the lidar data input, it may be associated with a database that calculates and manages, for each of the object classes (e.g., vehicle, truck, person, etc.) corresponding to each of the lidar 3D bounding boxes, the numerical value of the average width, the numerical value of the average length, and the numerical value of the average height.

[0085] Therefore, the computing device 1000 can refer to each of the specific object classes corresponding to each of the specific objects confirmed through each of the second GT-specific 2D bounding boxes, and obtain from the 3D lidar model the numerical value of the average width, the numerical value of the average length, and the numerical value of the average height corresponding to each of the specific object classes that have already been calculated by the 3D lidar model. Based on the numerical value of the average width, the numerical value of the average length, and the numerical value of the average height corresponding to each of the specific object classes, the average size information corresponding to each of the specific objects can be determined. As mentioned above, the average size information may be applied with perspective based on depth information.

[0086] At this time, even if the position prediction information and average size information of the 3D bounding box corresponding to each specific object corresponding to the 2D bounding box for GT are obtained, if the rotation information for each specific object is unknown, the specific pseudo 3D bounding box to be generated may be randomly generated in a form where each specific object is rotated by 45 degrees, 90 degrees, 180 degrees, etc. Therefore, in order to prevent this, the computing device 1000 inputs the image data included in the load data into the Yaw regression model, and uses the Yaw regression model to estimate and output the first rotation angle as the rotation information for each specific object on the camera coordinate system, and further applies calibration data to the first rotation angle to convert it into the second rotation angle for each specific object on the vehicle coordinate system, which can be determined as the rotation information for each specific object. For reference, the camera coordinate system may be a CCD (Camera Coordinate System) which is a 3D coordinate system with the position of the camera lens as the origin, but it is not limited to this. Thus, regarding the algorithm and architecture of the Yaw regression model for providing the rotation information for each specific object, it will be described with reference to FIGS. 7 and 8.

[0087] FIG. 7 briefly shows a process of estimating the rotation angle (Yaw) of a vehicle on a vehicle coordinate system from the rotation angle (Alpha) of the vehicle on a camera coordinate system according to an embodiment of the present invention.

[0088] Referring to FIG. 7, for example, when there are an x-axis and a y-axis perpendicular to each other in the vehicle coordinate system (VCS), the degree of rotation of a vehicle, which is a specific object to be observed, with reference to the x-axis or the y-axis is defined as Yaw for the vehicle, and the degree of rotation of the vehicle calculated with reference to two axes generated based on the direction in which the camera looks at the vehicle in the camera coordinate system can be defined as Alpha for the camera.

[0089] That is, how much the vehicle has rotated (i.e., Alpha) based on the direction in which the camera looks at the vehicle can be estimated through model learning. However, since it is very difficult to directly estimate the degree of rotation (Yaw) of the vehicle in the vehicle coordinate system, after estimating Alpha in the camera coordinate system, a method of applying calibration data to Alpha to convert it to Yaw will be used. For reference, Alpha can be regarded as the first rotation angle described in FIG. 6, and Yaw can be regarded as the second rotation angle.

[0090] FIG. 8 briefly shows the architecture of a Yaw regression model for estimating Alpha according to an embodiment of the present invention.

[0091] Referring to FIG. 8, the computing device 1000 inputs the image data included in the raw data into the feature extraction model 400 of the Yaw regression model 400, and the Yaw regression model 400 can cause the feature extraction layer 410 to output features_3 to features_5. At this time, the more the number of convolution operations increases, the more features from feature 1 to feature 5 can be generated. However, it may be assumed that features_3 to features_5 are selected among them to perform subsequent processes. Then, the computing device 1000 applies an average pooling operation to features_3 using the Yaw regression model 400, and applies an upsampling operation to features_5. After that, features_3 to features_5 are input into the Concatenation layer 420, and a Concatenation operation is applied to features_3 to features_5 through the Concatenation layer 420 to generate a fusion feature. The fusion feature can be sequentially input into the first Convolution layer 431 and the second Convolution layer 432 so that a Convolution operation is applied, and a mish function may be used as the activation function.

[0092] Thereafter, the computing device 1000 further applies a global average pooling (GAP) operation to the fusion feature to which the convolution operation is applied using the Yaw regression model 400, and then inputs the fusion feature to each of the Yaw class header 441, the Yaw delta header 442, and the class header 443, and can estimate Alpha as the first rotation angle with reference to the results output through each of the Yaw class header 441 and the Yaw delta header 442. For reference, the feature extraction layer 410 can include, but is not limited to, a VGG (Visual Geometry Group)-based convolutional neural network (CNN, Convolution Neural Network) such as RegVGG-A1.

[0093] Also, by checking the graph shown as the output value of the Yaw class header 441, it can be determined how much the vehicle, which is a specific object to be observed at 360 degrees, has rotated, and it can be seen that a plurality of multi-bin classes are assigned 0, 1, 2, and 3 to each of the four quadrants.

[0094] For example, if the vehicle's head is facing the first quadrant direction, the computing device 1000 can classify it into the "0" class (i.e., Alpha bin 0) using the Yaw regression model 400, and obtain the third rotation angle (expressed as Delta in FIG. 8) as the angle by which the front of the vehicle has rotated from the pre-set reference angle between 0 degrees and -90 degrees corresponding to the "0" class. However, the third rotation angle may be smaller than the angle range (e.g., 90-degree range) corresponding to each of the plurality of multi-bin classes.

[0095] At this time, if "1" is output through the Yaw class header 441, the computing device 1000 can determine that the vehicle's head is facing the fourth quadrant. If "2" is output through the Yaw class header 441, the computing device 1000 can determine that the vehicle's head is facing the third quadrant. If "3" is output through the Yaw class header 441, the computing device 1000 can determine that the vehicle's head is facing the second quadrant. For reference, each of the reference angles preset for each multi-bin class in the above example may indicate 45 degrees, but is not limited thereto.

[0096] Also, by checking the graph shown as the output value of the Yaw delta header 442, it can be confirmed that a predetermined signed value is set based on the reference angle within the multi-bin class towards which the vehicle's head is facing.

[0097] Specifically, the computing device 1000 uses the Yaw regression model 400 to refer to each of the plurality of signed values set based on each reference angle of each of the plurality of multi-bin classes through the Yaw delta header 442 to obtain the signed value corresponding to the direction indicated by the vehicle's head. Then, it gives the signed value obtained for the third rotation angle to obtain the third adjusted rotation angle, and referring to the third adjusted rotation angle and the reference angle, it can estimate Alpha, which is the first rotation angle with respect to the vehicle.

[0098] For example, if the head direction of the vehicle is classified into the "0" class through the Yaw class header 441, the reference angle within the "0" class is -45 degrees, and the delta which is the third rotation angle is 15 degrees, the Yaw regression model 400 can confirm that the sign corresponding to the direction pointed by the front of the vehicle is "-" through the output value output from the Yaw delta header 442, and can obtain -15 degrees as the third adjusted rotation angle. Then, the Yaw regression model 400 can be made to be able to estimate that Alpha as the first rotation angle with respect to the vehicle on the camera coordinate system is -60 degrees by referring to the reference angle of -45 degrees and the third adjusted rotation angle of -15 degrees. Therefore, if the first rotation angle is obtained through the Yaw regression model 400, the computing device 1000 can apply the calibration data included in the raw data to the first rotation angle to obtain Yaw which is the second rotation angle on the vehicle coordinate system.

[0099] Referring back to FIG. 6, the computing device 1000 generates each of the specific pseudo bounding boxes (S212_2) by referring to the prediction information, average size information, and rotation information of the 3D bounding boxes for each of the specific objects, calculates the second IOU loss and the second center loss by referring to the matching ratio between each of the generated specific pseudo 3D bounding boxes and each of the second GT 2D bounding boxes, generates the second integrated loss by referring to the second IOU loss and the second center loss (S221_2), and can perform regression (S222_2) by fitting each of the second lidar 2D bounding boxes by referring to the second integrated loss. That is, the computing device 1000 can assist in automatically generating pseudo 3D bounding boxes for objects not detected by the 3D lidar model so that each of the specific pseudo 3D bounding boxes is fitted to each of the second GT 2D bounding boxes. At the same time, it can increase the accuracy of matching and reduce the man-hours of actual learning data collection, which is effective.

[0100] For example, computing device 1000 generates respective specific pseudo 2D bounding boxes by referring to respective minimum values and respective maximum values of coordinates each including a respective specific pseudo 3D bounding box, calculates a second IOU loss which is a loss with respect to a matching ratio obtained by applying an IOU metric by referring to respective specific pseudo 2D bounding boxes and respective second GT 2D bounding boxes, calculates a second center loss which is an L1 loss generated by referring to each difference value between each center value of respective specific pseudo 2D bounding boxes and each center value of respective second GT 2D bounding boxes, generates a second integrated loss by referring to the second IOU loss and the second center loss, and performs regression by fitting so that the frame of each specific pseudo 2D bounding box matches the frame of each second GT 2D bounding box by referring to the second integrated loss. Here, generating respective specific pseudo 2D bounding boxes by referring to respective minimum values and respective maximum values of coordinates each including a respective specific pseudo 3D bounding box may mean generating respective specific pseudo 2D bounding boxes so as to surround (i.e., circumscribe) respective specific pseudo 3D bounding boxes, but is not limited thereto.

[0101] Thus, the difference before and after the labeling process for a specific object is performed by the computing device 1000 will be described with reference to FIG. 9.

[0102] FIG. 9 shows the generation results of projected lidar 3D bounding boxes before and after the process of the present invention is applied according to an embodiment of the present invention.

[0103] Referring to FIG. 9, (a) shows the state where the rider 3D bounding box output from the 3D lidar model before applying the present invention is projected onto the image data and output to the rider 3D bounding box. If we check the left side of the observation area 510 shown as a black rectangle after enlarging it, for the vehicle located on the far left, a projected rider 3D bounding box 511 was generated, but it is shown that a certain part is separated from the vehicle. For the other vehicles 512 to 515, as described in FIG. 3, it can be confirmed that no projected rider 3D bounding box was generated because the vehicle was not detected due to the characteristics of the lidar data.

[0104] On the other hand, (b) of FIG. 9 shows the result after applying the present invention. If we check the left side of the observation area 520 shown as a black rectangle after enlarging it, it can be seen that the position of the projected rider 3D bounding box 521 for the vehicle located on the far left has been adjusted, and for the remaining vehicles 522 to 525, it can be seen that respective pseudo 3D bounding boxes have been generated.

[0105] In addition, the embodiments according to the present invention described above can be embodied in the form of program instruction words that can be executed through various computer components and can be stored in a computer-readable recording medium. The computer-readable recording medium can include program instruction words, data files, data structures, etc. alone or in combination. The program instruction words stored in the computer-readable recording medium can be those specially designed and configured for the present invention or those known and usable by those skilled in the field of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program instruction words such as ROMs, RAMs, and flash memories. Examples of program instruction words include not only machine language codes such as those made by compilers but also high-level language codes that can be executed by a computer using an interpreter or the like. The hardware device can be configured to operate as one or more software modules for performing the processing according to the present invention, and vice versa.

[0106] As described above, the present invention has been described by way of specific examples and drawings limited to specific matters such as specific components. However, this is only provided to assist in a more general understanding of the present invention, and the present invention is not limited to the above embodiments. Those having ordinary knowledge in the technical field to which the present invention pertains can make various modifications and variations from such descriptions.

[0107] Therefore, the idea of the present invention should not be defined only by the embodiments described above. It can be said that not only the claims described below but also all those equivalently or equivalently modified fall within the scope of the idea of the present invention.

Claims

**Claim 1** A method of labeling at least one specific object by automatically generating at least one specific pseudo 3D bounding box, comprising: (a) a computing device obtaining each of at least one projected lidar 3D bounding box on an image coordinate system by using raw data including lidar data, image data, and calibration data, and determining whether each of the projected lidar 3D bounding boxes and each of the 2D bounding boxes for GT on the image coordinate system are matched with each other; and (b)(b1) If it is determined that each of the first projected lidar 3D bounding boxes, which is at least one of the projected lidar 3D bounding boxes, and each of the first 2D bounding boxes for GT, which is at least one of the 2D bounding boxes for GT, are matched, the computing device performs a regression on each of the first projected lidar 3D bounding boxes in order to fit each of the first projected lidar 3D bounding boxes to each of the first 2D bounding boxes for GT, and (b2) if it is determined that each of the second 2D bounding boxes for GT, which is at least one of the 2D bounding boxes for GT, is not matched to any of the projected lidar 3D bounding boxes, the computing device generates each of the specific pseudo 3D bounding boxes corresponding to each of the specific objects by referring to each of the second 2D bounding boxes for GT, the average size information and rotation information of each of the specific objects corresponding to each of the second 2D bounding boxes for GT, and performs a sub-process of performing the regression on each of the specific pseudo 3D bounding boxes in order to fit each of the specific pseudo 3D bounding boxes to each of the second 2D bounding boxes for GT; A method including the above.

2. In the step (a), The computing device inputs the lidar data included in the loaded data into a 3D lidar model to generate each of at least one lidar 3D bounding box on a vehicle coordinate system, and projects each of the lidar 3D bounding boxes onto the image coordinate system by referring to each of the lidar 3D bounding boxes and the calibration data included in the loaded data, thereby obtaining each of the projected lidar 3D bounding boxes. Each of the lidar 2D bounding boxes is generated by referring to each of the minimum values and maximum values of the coordinates including each of the projected lidar 3D bounding boxes. In order to determine whether each of the lidar 2D bounding boxes and each of the 2D bounding boxes for GT match each other, the Hungarian algorithm is applied using the IoU (Intersection Over Union) metric. The method according to claim 1, characterized in that.

3. In the sub-process (b1), The computing device generates respective first lidar 2D bounding boxes by referring to respective minimum values and respective maximum values of coordinates each including the respective first projected rider 3D bounding boxes, calculates a first IoU loss which is a loss with respect to a matching ratio obtained by applying an IoU metric by referring to the respective first lidar 2D bounding boxes and the respective first GT 2D bounding boxes, calculates a first center loss which is a loss generated by referring to respective difference values between respective central values of the respective first projected 2D bounding boxes and respective central values of the respective first GT 2D bounding boxes, generates a first integrated loss by referring to the first IoU loss and the first center loss, performs the regression by fitting the respective first lidar 2D bounding boxes by referring to the first integrated loss, and updates parameters of a 3D lidar model (the 3D lidar model is a model that generates respective at least one lidar 3D bounding box on a vehicle coordinate system when the lidar data included in the raw data is input) through backpropagation using the first integrated loss. The method according to claim 1, characterized in that.

4. In the sub-process (b2), The computing device refers to each of the second GT 2D bounding boxes to obtain respective x-axis coordinate values and respective y-axis coordinate values for the position information of the 2D bounding boxes on the image coordinate system of each of the second GT 2D bounding boxes, refers to the depth information of the position information of the 2D bounding boxes on the lidar data included in the load data to obtain the z-axis coordinate value for the position information of the 2D bounding boxes on the image coordinate system, refers to the position information of the 2D bounding boxes to determine the position prediction information of the 3D bounding boxes for each of the specific objects, refers to each of the specific object classes corresponding to each of the specific objects confirmed through each of the second GT 2D bounding boxes, and obtains, from a 3D lidar model (the 3D lidar model is a model that generates respective lidar 3D bounding boxes on the vehicle coordinate system when the lidar data included in the load data is input), respective average widths, average lengths, and average heights corresponding to each of the specific object classes that have already been calculated by the 3D lidar model, refers to the respective average widths, average lengths, and average heights corresponding to each of the specific object classes to determine the average size information for each of the specific objects, inputs the image data in the load data into a Yaw regression model to output a first rotation angle for each of the specific objects on the camera coordinate system using the Yaw regression model, and determines the second rotation angle as the rotation information for each of the specific objects by converting the first rotation angle to a second rotation angle for each of the specific objects on the vehicle coordinate system with reference to the first rotation angle and the calibration data included in the load data, and generates each of the specific pseudo 3D bounding boxes with reference to the position prediction information of the 3D bounding boxes, the average size information for each of the specific objects, and the rotation information. The method according to claim 1, characterized in that.

5. In the sub-process (b2), The computing device inputs the image data into the Yaw regression model to output each of a plurality of features using the Yaw regression model, and after sequentially applying a concatenation operation and a convolution operation to each of the plurality of features, it is classified into a specific multi-bin class corresponding to the direction pointed by the front of the specific object among a plurality of multi-bin classes through a Yaw class header. Among each of the preset reference angles corresponding to each of the plurality of multi-bin classes, a third rotation angle, which is the angle by which the specific object is rotated from the specific reference angle preset corresponding to the specific multi-bin class, is obtained. Through the Yaw delta header, a specific code value corresponding to the direction pointed by the front of the specific object is obtained by referring to each of the plurality of preset code values based on each of the reference angles of each of the plurality of multi-bin classes. The specific code value is given to the third rotation angle to obtain a third adjusted rotation angle, and the first rotation angle with respect to the specific object is estimated by referring to the third adjusted rotation angle and the specific reference angle. The method according to claim 4, characterized in that.

6. In the (b2) sub-process, Once each of the specific pseudo-3D bounding boxes is generated, the computing device calculates a matching ratio between each of the specific pseudo-3D bounding boxes and each of the second GT 2D bounding boxes using the IOU metric. If it is determined that the matching ratio is less than a preset threshold ratio, regression is performed on each of the specific pseudo-3D bounding boxes. The method according to claim 4, characterized in that.

7. In the (b2) sub-process, The computing device generates respective specific pseudo 2D bounding boxes by referring to respective minimum values and respective maximum values of coordinates including each of the specific pseudo 3D bounding boxes, calculates a second IoU loss which is a loss with respect to the matching ratio obtained by applying the IoU metric by referring to each of the specific pseudo 2D bounding boxes and each of the second GT 2D bounding boxes, calculates a second center loss which is a loss generated by referring to each difference value between each center value of each of the specific pseudo 2D bounding boxes and each center value of each of the second GT 2D bounding boxes, generates a second integrated loss by referring to the second IoU loss and the second center loss, and performs the regression by fitting each of the specific pseudo 2D bounding boxes by referring to the second integrated loss. The method according to claim 6, characterized in that.

8. In the (b2) sub-process, If it is determined that each of the second projected lidar 3D bounding boxes, which is at least one of the projected lidar 3D bounding boxes, is not matched to any of the GT 2D bounding boxes, the computing device deletes each of the second projected lidar 3D bounding boxes. The method according to claim 1, characterized in that.

9. In the step (a), The calibration data includes camera intrinsic parameters, extrinsic parameters, and conversion parameters for converting the lidar data into the image data. The method according to claim 1, characterized in that.

10. The GT 2D bounding box is obtained from data in which the image data generated by a predetermined camera is input, the specific object is detected through a deep learning model, and a 2D bounding box is generated for the specific object. The method according to claim 1, characterized in that.

11. In a computing device that labels at least one specific object by automatically generating at least one specific pseudo 3D bounding box, at least one memory storing instructions; and at least one processor configured to execute the instructions; comprising The processor performs the following processes: (I) using raw data including lidar data, image data, and calibration data to obtain each of at least one projected lidar 3D bounding box on an image coordinate system, and determining whether each of the projected lidar 3D bounding boxes and each of the GT 2D bounding boxes on the image coordinate system are mutually matched; and (II) (II-1) if it is determined that each of the first projected lidar 3D bounding boxes, which is at least one of the projected lidar 3D bounding boxes, and each of the first GT 2D bounding boxes, which is at least one of the GT 2D bounding boxes, are matched, performing regression on each of the first projected lidar 3D bounding boxes to fit each of the first projected lidar 3D bounding boxes to each of the first GT 2D bounding boxes, and (II-2) if it is determined that each of the second GT 2D bounding boxes, which is at least one of the GT 2D bounding boxes, is not matched to any of the projected lidar 3D bounding boxes, generating each of the specific pseudo 3D bounding boxes corresponding to each of the specific objects by referring to each of the second GT 2D bounding boxes, the average size information and rotation information of each of the specific objects corresponding to each of the second GT 2D bounding boxes, and performing the regression on each of the specific pseudo 3D bounding boxes to fit each of the specific pseudo 3D bounding boxes to each of the second GT 2D bounding boxes; a computing device that performs the processes.

12. The processor, in the process of (I), Input the lidar data included in the raw data into a 3D lidar model to generate at least one lidar 3D bounding box on a vehicle coordinate system using the 3D lidar model, and project each of the lidar 3D bounding boxes onto the image coordinate system with reference to each of the lidar 3D bounding boxes and the calibration data included in the raw data, thereby obtaining each of the projected lidar 3D bounding boxes. Generate each of the lidar 2D bounding boxes with reference to the minimum value and the maximum value of each coordinate including each of the projected lidar 3D bounding boxes, and apply the Hungarian algorithm using the Intersection over Union (IOU) metric to determine whether each of the lidar 2D bounding boxes and each of the 2D bounding boxes for GT match each other. The computing device according to claim 11, characterized in that.

13. The processor is In the (II-1) subprocess, Generating each of the first lidar 2D bounding boxes with reference to each of the minimum values and each of the maximum values of the coordinates including each of the first projected lidar 3D bounding boxes, calculating a first IoU loss which is a loss with respect to the matching ratio obtained by applying the IoU metric with reference to each of the first lidar 2D bounding boxes and each of the first GT 2D bounding boxes, calculating a first center loss which is a loss generated with reference to each difference value between each median value of each of the first projected 2D bounding boxes and each median value of each of the first GT 2D bounding boxes, generating a first integrated loss with reference to the first IoU loss and the first center loss, performing the regression by fitting each of the first lidar 2D bounding boxes with reference to the first integrated loss, and updating parameters of a 3D lidar model (the 3D lidar model is a model that generates each of at least one lidar 3D bounding box on a vehicle coordinate system when the lidar data included in the raw data is input) through backpropagation using the first integrated loss. The computing device according to claim 11, characterized in that.

14. The processor In the (II-2) subprocess With reference to each of the second GT 2D bounding boxes, obtain the x-axis coordinate values and the y-axis coordinate values for the position information of the 2D bounding boxes on the image coordinate system of each of the second GT 2D bounding boxes. Refer to the depth information of the position information of the 2D bounding boxes on the lidar data included in the load data to obtain the z-axis coordinate value for the position information of the 2D bounding boxes on the image coordinate system. Refer to the position information of the 2D bounding boxes to determine the position prediction information of the 3D bounding boxes for each of the specific objects. Refer to each of the specific object classes corresponding to each of the specific objects confirmed through each of the second GT 2D bounding boxes, and obtain, from a 3D lidar model (the 3D lidar model is a model that inputs the lidar data included in the load data and generates each of at least one lidar 3D bounding box on the vehicle coordinate system), the average width, average length, and average height corresponding to each of the specific object classes that have already been calculated by the 3D lidar model. Refer to the average width, average length, and average height corresponding to each of the specific object classes to determine the average size information for each of the specific objects. Input the image data in the load data into a Yaw regression model, and use the Yaw regression model to output a first rotation angle for each of the specific objects on the camera coordinate system. By referring to the first rotation angle and the calibration data included in the load data and converting the first rotation angle into a second rotation angle for each of the specific objects on the vehicle coordinate system, determine the second rotation angle as the rotation information for each of the specific objects. Generate each of the specific pseudo 3D bounding boxes by referring to the position prediction information of the 3D bounding boxes, the average size information for each of the specific objects, and the rotation information. The computing device according to claim 11, characterized in that

15. The processor In the (II-2) sub-process Input the image data into the Yaw regression model, and use the Yaw regression model to output each of a plurality of features. After sequentially applying concatenation operation and convolution operation to each of the plurality of features, classify it into a specific multi-bin class corresponding to the direction pointed by the front of the specific object among a plurality of multi-bin classes through the Yaw class header. Obtain a third rotation angle, which is the angle by which the specific object is rotated from a specific reference angle preset corresponding to the specific multi-bin class among each of the preset reference angles corresponding to each of the plurality of multi-bin classes. Through the Yaw delta header, obtain a specific code value corresponding to the direction pointed by the front of the specific object by referring to each of the plurality of preset code values based on each of the reference angles of each of the plurality of multi-bin classes. Give the specific code value to the third rotation angle to obtain a third adjusted rotation angle. Estimate the first rotation angle with respect to the specific object by referring to the third adjusted rotation angle and the specific reference angle. The computing device according to claim 14, characterized in that.

16. The processor is In the (II-2) sub-process, If each of the specific pseudo-3D bounding boxes is generated, calculate the matching ratio between each of the specific pseudo-3D bounding boxes and each of the second GT 2D bounding boxes using the IOU metric. If it is determined that the matching ratio is less than a preset threshold ratio, perform the regression on each of the specific pseudo-3D bounding boxes. The computing device according to claim 14, characterized in that.

17. The processor is In the (II-2) sub-process, Generating respective specific pseudo 2D bounding boxes by referring to respective minimum values and maximum values of coordinates each including the respective specific pseudo 3D bounding box, calculating a second IoU loss which is a loss with respect to a matching ratio obtained by applying the IoU metric by referring to the respective specific pseudo 2D bounding boxes and the respective second GT 2D bounding boxes, calculating a second center loss which is a loss generated by referring to each difference value between each center value of the respective specific pseudo 2D bounding boxes and each center value of the respective second GT 2D bounding boxes, generating a second integrated loss by referring to the second IoU loss and the second center loss, and performing the regression by fitting the respective specific pseudo 2D bounding boxes by referring to the second integrated loss. The computing device according to claim 16, characterized in that.

18. The processor In the (II-2) subprocess, If it is determined that each of at least one second projected lidar 3D bounding box among the projected lidar 3D bounding boxes is not matched to any of the GT 2D bounding boxes, deleting each of the second projected lidar 3D bounding boxes. The computing device according to claim 11, characterized in that.

19. The processor In the (I) process, The calibration data includes camera intrinsic parameters, extrinsic parameters, and conversion parameters for converting the lidar data into the image data. The computing device according to claim 11, characterized in that.

20. The GT 2D bounding box is obtained from data in which the image data generated by a predetermined camera is input, the specific object is detected through a deep learning model, and a 2D bounding box is generated for the specific object. The computing device according to claim 11, characterized in that.

Citation Information

Patent Citations

  • Database construction system for article recognition algorism machine-learning

    JP2017102838A

  • Information processing method, information processing apparatus and program

    JP2020021326A

  • Method and system generating correct answer data for machine learning of recognition unit

    JP2024000347A

  • Status determination device, status determination system, and status determination method

    WO2022097426A1

Cited By

  • Contact member, drying apparatus, and printing apparatus for treating liquid-coated substrates

    US12459276B2