METHOD FOR LABELING AT LEAST ONE SPECIFIC OBJECT BY AUTOMATICALLY CREATING AT LEAST ONE SPECIFIC PSEUDO 3D BOUNDING BOX AND COMPUTING DEVICE USING THE SAME
The method automates the generation of pseudo-3D bounding boxes by matching lidar 3D boxes with 2D boxes and using average size/rotation info, addressing inefficiencies and errors in existing tools, thus improving training data quality.
Patent Information
- Application Number
- JP2024221796
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2024-11-12
- Filing Date
- 2024-12-18
- Publication Date
- 2025-10-15
- Estimated Expiration
- 2044-12-18
AI Technical Summary
Existing labeling tools for generating 3D bounding boxes in machine learning, particularly using deep learning, suffer from high error rates due to repetitive tasks and require significant manual effort, making them inefficient and time-consuming.
A method for automatically generating pseudo-3D bounding boxes by matching lidar 3D bounding boxes with 2D bounding boxes using calibration data, and generating pseudo-3D boxes based on average size and rotation information when a match is not found, with regression processes to improve accuracy.
Reduces labeling errors and improves efficiency by automating the generation of accurate 3D bounding boxes, enhancing the quality of training data for deep learning models.
Smart Images

Figure 0007755035000001 
Figure 0007755035000002 
Figure 0007755035000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method for labeling at least one specific object by automatically generating at least one specific pseudo-3D bounding box, and a computing device using the same. [Background technology]
[0002] Recently, a method for identifying objects using machine learning As part of this type of machine learning, deep learning, which uses neural networks with several hidden layers between the input and output layers, has high discrimination performance.
[0003] In addition, neural networks that use deep learning generally learn through backpropagation using losses, but training data in which objects are tagged, i.e., labeled, using a labeling tool is required for deep learning network training. To prepare such training data, labeling tools have traditionally been implemented to add simple functions such as generating a 3D bounding box using a mouse pointer, which has the advantage of not requiring much man-hours to develop the labeling tool.
[0004] However, when using existing labeling tools, workers have to generate 3D bounding boxes directly using the mouse pointer, which can lead to frequent labeling errors due to a decrease in worker concentration caused by repetitive tasks. In fact, even if workers are trained in labeling in advance, labeling errors are still found, which results in additional work required to collect the labeled products.
[0005] In addition, when generating a 2D bounding box, the operator can check the object on the image and drag it using the mouse pointer to generate a box that includes the object. However, to generate a 3D bounding box, the operator must check the 3D coordinate information (x, y, z) corresponding to the object and then appropriately set the size information (width, height, length) and rotation information (roll, pitch, yaw) of the 3D bounding box. However, this task is difficult, and there is a problem that it takes a long time to complete each object.
[0006] Therefore, there is a need for an improved solution to the above problems. Summary of the Invention [Problem to be solved by the invention]
[0007] An object of the present invention is to solve the above-mentioned problems.
[0008] In addition, the present invention provides a method for obtaining at least one projected lidar 3D bounding box on an image coordinate system using raw data including lidar data, image data, and calibration data; and (i) if it is determined that at least one first projected lidar 3D bounding box among the projected lidar 3D bounding boxes and at least one first GT 2D bounding box among the GT 2D bounding boxes are matched, performing a calibration of the first projected lidar 3D bounding box to fit the first projected lidar 3D bounding box to the first GT 2D bounding box. (ii) if it is determined that at least one second 2D bounding box among the 2D bounding boxes for GT does not match any of the projected lidar 3D bounding boxes, generate specific pseudo 3D bounding boxes corresponding to each of the specific objects by referring to average size information and rotation information of each of the second 2D bounding boxes for GT and each of the corresponding specific objects, and perform the regression on each of the specific pseudo 3D bounding boxes to fit each of the specific pseudo 3D bounding boxes to each of the second 2D bounding boxes for GT. [Means for solving the problem]
[0009] According to one embodiment of the present invention, a method for labeling at least one specific object by automatically generating at least one specific pseudo-3D bounding box includes: (a) a computing device obtains at least one projected LIDAR 3D bounding box on an image coordinate system using raw data including LIDAR data, image data, and calibration data, and determines whether each of the projected LIDAR 3D bounding boxes and each of the GT 2D bounding boxes on the image coordinate system are mutually matched;and (b) (b1) if it is determined that at least one first projected rider 3D bounding box among the projected rider 3D bounding boxes and at least one first GT 2D bounding box among the GT 2D bounding boxes are matched, the computing device performs a regression on each of the first projected rider 3D bounding boxes to fit each of the first projected rider 3D bounding boxes to each of the first GT 2D bounding boxes; and (b2) if at least one of the GT 2D bounding boxes is matched, the computing device performs a regression on each of the first projected rider 3D bounding boxes to fit each of the first GT 2D bounding boxes to each of the first GT 2D bounding boxes. If it is determined that each of the second GT 2D bounding boxes, which are also one of the first GT 2D bounding boxes, does not match any of the projected lidar 3D bounding boxes, the computing device generates specific pseudo-3D bounding boxes corresponding to each of the specific objects by referring to each of the second GT 2D bounding boxes and average size information and rotation information of each of the second GT 2D bounding boxes and specific objects corresponding to each of the second GT 2D bounding boxes, and performs the sub-process of performing the regression on each of the specific pseudo-3D bounding boxes to fit each of the specific pseudo-3D bounding boxes to each of the second GT 2D bounding boxes.
[0010] In one example, in step (a), the computing device inputs the lidar data included in the raw data into a 3D lidar model and generates at least one lidar 3D bounding box on a vehicle coordinate system using the 3D lidar model; obtains each of the projected lidar 3D bounding boxes by projecting each of the lidar 3D bounding boxes onto the image coordinate system with reference to each of the lidar 3D bounding boxes and the calibration data included in the raw data; generates each of the lidar 2D bounding boxes with reference to minimum and maximum coordinate values including each of the projected lidar 3D bounding boxes; and applies a Hungarian algorithm using an Intersection Over Union (IOU) metric to determine whether each of the lidar 2D bounding boxes and each of the GT 2D bounding boxes match each other.
[0011] In one example, in the subprocess (b1), the computing device generates each of the first lidar 2D bounding boxes by referring to the minimum and maximum values of the coordinates including each of the first projected lidar 3D bounding boxes, calculates a first IOU loss, which is a loss relative to the matching ratio obtained by applying an IOU metric with reference to each of the first lidar 2D bounding boxes and each of the first GT 2D bounding boxes, and calculates a first IOU loss, which is a loss relative to the matching ratio obtained by applying an IOU metric with reference to each of the first projected 2D bounding boxes and each of the first GT 2D bounding boxes. the first integrated loss is generated by fitting each of the first LIDAR 2D bounding boxes to the first integrated loss, and parameters of a 3D LIDAR model (the 3D LIDAR model is a model that receives the LIDAR data included in the raw data and generates at least one LIDAR 3D bounding box on a vehicle coordinate system) are updated through backpropagation using the first integrated loss.
[0012] In one example, in the subprocess (b2), the computing device refers to each of the 2D bounding boxes for the second GT to obtain x-axis coordinate values and y-axis coordinate values for position information of the 2D bounding box on the image coordinate system for each of the 2D bounding boxes for the second GT, refers to depth information of the position information of the 2D bounding box on the LIDAR data included in the raw data to obtain z-axis coordinate values for the position information of the 2D bounding box on the image coordinate system, and determines position prediction information of a 3D bounding box for each of the specific objects by referring to the position information of the 2D bounding box, Each of the specific object classes identified through each of the 2D bounding boxes is referenced to a 3D lidar model (the 3D lidar model is a model in which the lidar data included in the raw data is input and at least one lidar 3D bounding box is generated on a vehicle coordinate system), and an average width, an average length, and an average height corresponding to each of the specific object classes calculated by the 3D lidar model are obtained from the 3D lidar model (the 3D lidar model is a model in which the lidar data included in the raw data is input and at least one lidar 3D bounding box is generated on a vehicle coordinate system), and the average size information for each of the specific objects is determined by referring to the average width, the average length, and the average height corresponding to each of the specific object classes, and the image data in the raw data is yawed. and inputting the raw data into a yaw regression model to output a first rotation angle for each of the specific objects on the camera coordinate system using the yaw regression model; converting the first rotation angle into a second rotation angle for each of the specific objects on the vehicle coordinate system by referring to the first rotation angle and the calibration data included in the raw data, thereby determining the second rotation angle as the rotation information for each of the specific objects; and generating each of the specific pseudo 3D bounding boxes by referring to position prediction information of the 3D bounding boxes, the average size information for each of the specific objects, and the rotation information.
[0013] In one example, in the subprocess (b2), the computing device inputs the image data into the Yaw regression model and outputs each of a plurality of features using the Yaw regression model, sequentially applies a concatenation operation and a convolution operation to each of the plurality of features, classifies the specific object into a specific multi-bin class among a plurality of multi-bin classes corresponding to a forward pointing direction thereof through a Yaw class header, obtains a third rotation angle by which the specific object is rotated from a specific reference angle pre-set for the specific multi-bin class among pre-set reference angles corresponding to each of the plurality of multi-bin classes, obtains a specific code value corresponding to a forward pointing direction of the specific object by referring to each of a plurality of pre-set code values based on each of the reference angles of the plurality of multi-bin classes through the Yaw delta header, obtains a third adjusted rotation angle by assigning the specific code value to the third rotation angle, and estimates the first rotation angle for the specific object by referring to the third adjusted rotation angle and the specific reference angle.
[0014] In one example, in the subprocess (b2), once each of the specific pseudo 3D bounding boxes is generated, the computing device calculates a matching ratio between each of the specific pseudo 3D bounding boxes and each of the second GT 2D bounding boxes using an IOU metric, and if it is determined that the matching ratio is less than a preset threshold ratio, performs the regression on each of the specific pseudo 3D bounding boxes.
[0015] In one example, in the subprocess (b2), the computing device generates each of the specific pseudo 2D bounding boxes by referring to the minimum and maximum values of coordinates including each of the specific pseudo 3D bounding boxes, calculates a second IOU loss, which is a loss for the matching ratio obtained by applying the IOU metric by referring to each of the specific pseudo 2D bounding boxes and each of the second GT 2D bounding boxes, calculates a second center loss, which is a loss generated by referring to each median of each of the specific pseudo 2D bounding boxes and each difference value between the medians of each of the second GT 2D bounding boxes, generates a second integrated loss by referring to the second IOU loss and the second center loss, and performs the regression by fitting each of the specific pseudo 2D bounding boxes by referring to the second integrated loss.
[0016] In one example, in the (b2) subprocess, if it is determined that at least one second projected rider 3D bounding box among the projected rider 3D bounding boxes does not match any of the GT 2D bounding boxes, the computing device deletes each of the second projected rider 3D bounding boxes.
[0017] In one example, in step (a), the calibration data includes intrinsic parameters and extrinsic parameters of the camera, and transformation parameters for converting the LIDAR data into the image data.
[0018] In one example, the 2D bounding box for GT is obtained from data in which the image data generated by a predetermined camera is input, the specific object is detected through a deep learning model, and a 2D bounding box is generated for the specific object.
[0019] According to another embodiment of the present invention, a computing device for labeling at least one specific object by automatically generating at least one specific pseudo-3D bounding box includes at least one memory for storing instructions; and at least one processor configured to execute the instructions, wherein the processor (I) performs a process of obtaining at least one projected LIDAR 3D bounding box on an image coordinate system using raw data including LIDAR data, image data, and calibration data, and determining whether each of the projected LIDAR 3D bounding boxes and each of the GT 2D bounding boxes on the image coordinate system are mutually matched;and (II) (II-1) a sub-process of performing regression on each of the first projected rider 3D bounding boxes to fit each of the first projected rider 3D bounding boxes to each of the first GT 2D bounding boxes if it is determined that at least one of the first projected rider 3D bounding boxes and at least one of the first GT 2D bounding boxes are matched; and (II-2) a sub-process of performing regression on each of the first projected rider 3D bounding boxes to fit each of the first GT 2D bounding boxes to each of the first GT 2D bounding boxes. If it is determined that each of the second GT 2D bounding boxes does not match any of the projected lidar 3D bounding boxes, a sub-process is provided that generates specific pseudo 3D bounding boxes corresponding to each of the specific objects by referring to each of the second GT 2D bounding boxes and average size information and rotation information of each of the second GT 2D bounding boxes and specific objects corresponding to each of the second GT 2D bounding boxes, and performs the regression on each of the specific pseudo 3D bounding boxes to fit each of the specific pseudo 3D bounding boxes to each of the second GT 2D bounding boxes.
[0020] In one example, in process (I), the processor inputs the lidar data included in the raw data into a 3D lidar model and generates at least one lidar 3D bounding box on a vehicle coordinate system using the 3D lidar model; obtains each of the projected lidar 3D bounding boxes by projecting each of the lidar 3D bounding boxes onto the image coordinate system with reference to each of the lidar 3D bounding boxes and the calibration data included in the raw data; generates each of the lidar 2D bounding boxes with reference to each of the minimum and maximum coordinate values including each of the projected lidar 3D bounding boxes; and applies a Hungarian algorithm using an Intersection Over Union (IOU) metric to determine whether each of the lidar 2D bounding boxes and each of the GT 2D bounding boxes match each other.
[0021] In one example, in the (II-1) subprocess, the processor generates each of the first rider 2D bounding boxes by referring to the minimum and maximum values of the coordinates including each of the first projected rider 3D bounding boxes, calculates a first IOU loss, which is a loss relative to a matching ratio obtained by applying an IOU metric with reference to each of the first rider 2D bounding boxes and each of the first GT 2D bounding boxes, and calculates a first IOU loss, which is a loss relative to a matching ratio obtained by applying an IOU metric with reference to each of the first projected 2D bounding boxes and each of the first GT 2D bounding boxes. the first integrated loss is generated by fitting each of the first LIDAR 2D bounding boxes to the first integrated loss; and the parameters of a 3D LIDAR model (the 3D LIDAR model is a model that receives the LIDAR data included in the raw data and generates at least one LIDAR 3D bounding box on a vehicle coordinate system) are updated through backpropagation using the first integrated loss.
[0022] In one example, in the (II-2) subprocess, the processor refers to each of the 2D bounding boxes for the second GT to acquire x-axis coordinate values and y-axis coordinate values for position information of the 2D bounding box on the image coordinate system for each of the 2D bounding boxes for the second GT, acquires z-axis coordinate values for position information of the 2D bounding box on the image coordinate system by referring to depth information of position information of the 2D bounding box on the LIDAR data included in the raw data, determines position prediction information of a 3D bounding box for each of the specific objects by referring to the position information of the 2D bounding box, and Each of the specific object classes identified through each of the bounding boxes is referenced, and an average width, an average length, and an average height corresponding to each of the specific object classes calculated by the 3D lidar model is obtained from the 3D lidar model (the 3D lidar model is a model in which the lidar data included in the raw data is input and generates at least one lidar 3D bounding box on a vehicle coordinate system) by referring to each of the specific object classes corresponding to each of the specific objects identified through each of the bounding boxes. The average width, the average length, and the average height corresponding to each of the specific object classes are determined by referring to the average width, the average length, and the average height corresponding to each of the specific object classes. The image data in the raw data is then yawed. and inputting the raw data into a Yaw regression model to output a first rotation angle for each of the specific objects on the camera coordinate system using the Yaw regression model; converting the first rotation angle into a second rotation angle for each of the specific objects on the vehicle coordinate system by referring to the first rotation angle and the calibration data included in the raw data, thereby determining the second rotation angle as the rotation information for each of the specific objects; and generating each of the specific pseudo-3D bounding boxes by referring to position prediction information of the 3D bounding boxes, the average size information for each of the specific objects, and the rotation information.
[0023] In one example, in the subprocess (II-2), the processor inputs the image data into the Yaw regression model and outputs each of a plurality of features using the Yaw regression model, sequentially applies concatenation and convolution operations to each of the plurality of features, classifies the specific object into a specific multi-bin class among a plurality of multi-bin classes corresponding to a forward pointing direction thereof through a Yaw class header, obtains a third rotation angle by which the specific object is rotated from a specific reference angle pre-set for the specific multi-bin class among pre-set reference angles corresponding to each of the plurality of multi-bin classes, obtains a specific code value corresponding to a forward pointing direction of the specific object by referring to each of a plurality of pre-set code values based on each of the reference angles of each of the plurality of multi-bin classes through the Yaw delta header, obtains a third adjusted rotation angle by assigning the specific code value to the third rotation angle, and estimates the first rotation angle for the specific object by referring to the third adjusted rotation angle and the specific reference angle.
[0024] In one example, in the subprocess (II-2), when each of the specific pseudo 3D bounding boxes is generated, the processor calculates a matching ratio between each of the specific pseudo 3D bounding boxes and each of the second GT 2D bounding boxes using an IOU metric, and if it is determined that the matching ratio is less than a preset threshold ratio, performs the regression on each of the specific pseudo 3D bounding boxes.
[0025] In one example, in the (II-2) subprocess, the processor generates each of the specific pseudo 2D bounding boxes by referring to the minimum and maximum values of coordinates including each of the specific pseudo 3D bounding boxes, calculates a second IOU loss that is a loss with respect to the matching ratio obtained by applying the IOU metric by referring to each of the specific pseudo 2D bounding boxes and each of the second GT 2D bounding boxes, calculates a second center loss that is a loss generated by referring to each median of each of the specific pseudo 2D bounding boxes and each difference value between the median of each of the second GT 2D bounding boxes, generates a second integrated loss by referring to the second IOU loss and the second center loss, and performs the regression by fitting each of the specific pseudo 2D bounding boxes by referring to the second integrated loss.
[0026] In one example, if the processor determines in the (II-2) subprocess that at least one second projected rider 3D bounding box among the projected rider 3D bounding boxes does not match any of the GT 2D bounding boxes, it deletes each of the second projected rider 3D bounding boxes.
[0027] In one example, the processor is characterized in that in process (I), the calibration data includes intrinsic parameters of the camera, extrinsic parameters, and transformation parameters for converting the lidar data into the image data.
[0028] In one example, the 2D bounding box for GT is obtained by inputting the image data generated by a predetermined camera, detecting the specific object through a deep learning model, and generating a 2D bounding box for the specific object from the input data. [Effects of the Invention]
[0029] The present invention provides a method for obtaining at least one projected lidar 3D bounding box on an image coordinate system using raw data including lidar data, image data, and calibration data, and (i) if it is determined that at least one first projected lidar 3D bounding box among the projected lidar 3D bounding boxes and at least one first GT 2D bounding box among the GT 2D bounding boxes are matched, performing a calibration on the first projected lidar 3D bounding box to fit the first projected lidar 3D bounding box to the first GT 2D bounding box. (ii) if it is determined that at least one of the 2D bounding boxes for GT, a second 2D bounding box, does not match any of the projected lidar 3D bounding boxes, a specific pseudo 3D bounding box corresponding to each of the specific objects is generated by referring to the average size information and rotation information of each of the second 2D bounding boxes for GT and each of the corresponding specific objects, and the regression is performed on each of the specific pseudo 3D bounding boxes to fit each of the specific pseudo 3D bounding boxes to each of the second 2D bounding boxes for GT. [Brief explanation of the drawings]
[0030] The following drawings, which are attached to be used for explaining the embodiments of the present invention, are merely a part of the embodiments of the present invention, and a person having ordinary knowledge in the technical field to which the present invention pertains (hereinafter referred to as "ordinary engineer") can derive each of the other drawings based on these drawings without performing any inventive work.
[0031] [Figure 1] 1 is a simplified diagram illustrating a computing device for labeling at least one specific object by automatically generating at least one specific pseudo-3D bounding box according to an embodiment of the present invention; [Figure 2] 1 is a simplified diagram illustrating a process of labeling at least one specific object by automatically generating at least one specific pseudo-3D bounding box according to an embodiment of the present invention; [Figure 3] 1 is a simplified illustration of an example of a lidar 3D bounding box output from a 3D lidar model according to an embodiment of the present invention; [Figure 4] 10 is a simplified diagram illustrating a process for checking whether each of the lidar 3D bounding boxes output by the 3D lidar model matches each of the 2D bounding boxes for GT on the image coordinate system according to an embodiment of the present invention; [Figure 5] 10 is a simplified diagram illustrating a regression process in which each matched projected lidar 3D bounding box is fitted to each GT 2D bounding box on the image coordinate system according to an embodiment of the present invention; [Figure 6] According to one embodiment of the present invention, a lidar 3D bounding box is not generated, but a specific pseudo 3D bounding box for each specific object for which a 2D bounding box for GT is generated is simply illustrated. [Figure 7]1 is a simplified diagram illustrating a process of estimating a rotation angle (Yaw) of a vehicle on a vehicle coordinate system using a rotation angle (Alpha) of a vehicle on a camera coordinate system according to an embodiment of the present invention; [Figure 8] 1 is a simplified diagram illustrating the architecture of a Yaw regression model for estimating Alpha according to an embodiment of the present invention; [Figure 9] 10 shows the results of generating a projected lidar 3D bounding box before and after the process of the present invention is applied according to one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0032] The detailed description of the present invention below will highlight the objectives, technical solutions, and advantages of the present invention. For clarity, reference is made to the accompanying drawings which show, by way of illustration, specific embodiments in which the invention may be practiced, these embodiments being described in sufficient detail to enable those of ordinary skill in the art to practice the invention.
[0033] Furthermore, throughout the detailed description of the present invention and the claims, the word "comprises" and variations thereof are not intended to exclude other technical features, additions, components, or steps. Other objects, advantages, and characteristics of the present invention will become apparent to those of ordinary skill in the art, in part from the description and in part from the practice of the present invention. The following examples and drawings are provided as illustrations and are not intended to limit the present invention.
[0034] Furthermore, the present invention covers all possible combinations of the embodiments described herein. It should be understood that the various embodiments of the present invention, although different from one another, are not necessarily mutually exclusive. For example, a specific shape, structure, and characteristic described herein may be embodied in other embodiments without departing from the spirit and scope of the present invention. It should also be understood that the location or arrangement of individual components within each disclosed embodiment may be modified without departing from the spirit and scope of the present invention. Therefore, the following detailed description is not intended to be taken in a limiting sense, and the scope of the present invention is limited only by the appended claims, along with the full scope of equivalents to which such claims are entitled, if properly interpreted. Like reference numerals in the drawings refer to the same or similar functionality throughout the various aspects.
[0035] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS In the following, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily carry out the present invention.
[0036] FIG. 1 is a simplified diagram illustrating a computing device for labeling at least one specific object by automatically generating at least one specific pseudo-3D bounding box according to one embodiment of the present invention.
[0037] 1, a computing device 1000 may include a memory 1010 that stores instructions for labeling at least one specific object by automatically generating at least one specific pseudo 3D bounding box, and a processor 1020 that labels at least one specific object by automatically generating at least one specific pseudo 3D bounding box in accordance with the instructions stored in the memory 1010. In this case, the computing device 1000 may include a personal computer (PC), a mobile computer, etc.
[0038] Specifically, computing device 1000 may typically utilize a combination of computing devices (e.g., devices that may include a computer processor, memory, storage, input and output devices, and other components of existing computing devices; electronic communication devices such as routers, switches, etc.; electronic information storage systems such as network attached storage (NAS) and storage area networks (SAN)) and computer software (i.e., instructions that cause a computing device to function in a particular manner) to achieve desired system performance.
[0039] The processor of a computing device may include hardware components such as a microprocessing unit (MPU) or central processing unit (CPU), cache memory, and data bus. The computing device may also include software components such as an operating system and applications that perform specific purposes.
[0040] However, this does not exclude the case where the computing device includes an integrated processor in which a medium processor and memory for implementing the present invention are integrated.
[0041] The computing device 1000 may be linked to a database 900 that includes information used to label at least one specific object by automatically generating at least one specific pseudo-3D bounding box. Here, the database 900 may include at least one type of storage medium selected from the group consisting of a flash memory type, a hard disk type, a multimedia card micro type, a card-type memory (e.g., SD or XD memory), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, and an optical disk, but is not limited thereto, and may include any medium capable of storing data. In addition, the database 900 may be installed separately from the computing device 100, or alternatively, it may be installed inside the computing device 100 to transmit data and record received data. Unlike the illustrated example, the database 900 may also be embodied as two or more separate components, which may vary depending on the implementation conditions of the invention.
[0042] FIG. 2 is a simplified diagram illustrating a process of labeling at least one specific object by automatically generating at least one specific pseudo-3D bounding box according to one embodiment of the present invention.
[0043] First, referring to FIG. 2, the computing device 1000 may obtain at least one projected lidar 3D bounding box on an image coordinate system using raw data including image data, lidar data, and calibration data, and determine whether each of the projected lidar 3D bounding boxes and each of the GT 2D bounding boxes on the image coordinate system are mutually matched (S100).
[0044] For example, the computing device 1000 may input lidar data included in the raw data into a 3D lidar model and generate and output lidar 3D bounding boxes for each object included in the lidar data through calculations using the 3D lidar model, or may input image data into SVNet3 (a deep learning model that receives image data generated by a camera, detects objects, and generates 2D bounding boxes for the objects, has high detection accuracy, and can be used in GT) and generate and output 2D bounding boxes for GT for each object in the image coordinate system using SVNet3. However, the entity for generating 2D bounding boxes for GT from image data is not limited to SVNet3, and 2D bounding boxes for GT may be generated using other known methods that can generate bounding boxes with higher accuracy than the lidar model, and the 2D bounding boxes may be treated as GT.
[0045] As another example, the computing device 100 can acquire each of the rider 3D bounding boxes and each of the 2D bounding boxes for the GT from the database 900 while storing and managing the rider 3D bounding boxes and each of the 2D bounding boxes for the GT obtained as a result of previously outputting each of the rider 3D bounding boxes and each of the 2D bounding boxes for the GT using a 3D rider model and SVNet3, respectively.
[0046] For reference, the calibration data may include intrinsic and extrinsic parameters of the camera, parameters for converting lidar data into image data, etc., and the image coordinate system may refer to, but is not limited to, an Image Coordinate System (ICS), which is a two-dimensional coordinate system with the upper left corner of the image acquired by the camera as its origin.
[0047] Additionally, the computing device 1000 may check whether each of the lidar 3D bounding boxes output from the 3D lidar model matches each of the 2D bounding boxes for GT in the image data (i.e., data obtained by inputting image data acquired from a camera and outputting the SVNet3 model) by comparing them in pairs. In this case, to check whether each of the lidar 3D bounding boxes matches each of the 2D bounding boxes for GT, the computing device 1000 may project the lidar 3D bounding boxes onto an image coordinate system to obtain each of the lidar 3D bounding boxes projected onto the image coordinate system, and then check whether each of the projected lidar 3D bounding boxes matches each of the 2D bounding boxes for GT by comparing them in pairs, but the method is not limited to this.
[0048] At this time, an example of the lidar 3D bounding box output from the 3D lidar model will be described with reference to FIG.
[0049] FIG. 3 is a simplified illustration of an exemplary lidar 3D bounding box output from a 3D lidar model according to one embodiment of the present invention.
[0050] 3, the 3D bounding boxes 301, 302, 311, 313, 315, and 321 shown in (a) through (c) may be obtained by projecting a lidar 3D bounding box included in the lidar data output through the 3D lidar model onto an image coordinate system using calibration data. For reference, it should be noted that for convenience of explanation, 2D bounding boxes for GT that exactly match the projected lidar 3D bounding boxes are not shown in (a) and (c) of FIG.
[0051] In (a) of FIG. 3, it can be seen that lidar 3D bounding boxes 301 and 302 projected for each person and vehicle in the image coordinate system are generated. In (b), due to inaccurate calibration data included in the raw data, lidar 3D bounding boxes 311 and 313 are obtained that are projected at a certain distance from the 2D bounding boxes 312 and 314 for the GT for each person and vehicle in the image coordinate system. By checking the right side of the 2D bounding box 314 for the vehicle, it can be seen that a lidar 3D bounding box 315 projected for an object that is falsely detected (FP, False Positive) is obtained. In (c), a lidar 3D bounding box 315 projected for an object located close by is generated, but due to the characteristics of the lidar data that show little point cloud data for other objects located further away, the object vehicle 322 is not detected (FN, False Positive). Negative) and it can be confirmed that this is the case where a projected lidar 3D bounding box for the vehicle 322 is not obtained.
[0052] That is, in the case of the projected lidar 3D bounding boxes 311, 313, 315, and 321 as shown in (b) and (c) of Figure 3, the 3D bounding boxes 311 and 313 are displayed at a certain distance from the corresponding 2D bounding boxes for GT, resulting in a problem in which a projected 3D bounding box 315 is generated for an incorrectly detected object, or a projected 3D bounding box is not generated because the object is not detected. In such cases, the data is not suitable for use as learning data for an autonomous vehicle, and a solution to this problem is needed.
[0053] Referring again to FIG. 2, when the computing device 1000 checks the matching results for each of the projected lidar 3D bounding boxes and each of the GT 2D bounding boxes, three cases may occur. The processing process for each of the three cases will be described as follows.
[0054] As an example, the computing device 1000 may determine (S200_1) that each of the first projected rider 3D bounding boxes, which is at least one among the projected rider 3D bounding boxes, and each of the first GT 2D bounding boxes, which is at least one among the GT 2D bounding boxes, are matched.
[0055] In other words, this is a case where there is at least a partial overlap between each of the first projected lidar 3D bounding boxes and each of the first 2D bounding boxes for the GT, and although the positions of each of the first projected lidar 3D bounding boxes and each of the first 2D bounding boxes for the GT do not exactly match due to inaccurate calibration data, it can be determined that each of the first objects corresponding to each of the first projected lidar 3D bounding boxes and each of the first 2D bounding boxes for the GT are matched as the same object.
[0056] Therefore, if the computing device 1000 determines that each of the first projected rider 3D bounding boxes matches each of the first GT 2D bounding boxes, it can perform regression on each of the first projected rider 3D bounding boxes (S210_1).
[0057] In this case, since the 2D bounding box for the GT has high detection accuracy and can be used as the GT as described above, the computing device 1000 may perform regression on each of the first projected 3D bounding boxes so that each of the first projected 3D bounding boxes is fitted to each of the first 2D bounding boxes for the GT. For reference, the process of performing regression on each of the first projected 3D bounding boxes will be described in detail below with reference to FIG. 5.
[0058] As another example, the computing device 1000 may determine (S200_2) that at least one of the second GT 2D bounding boxes does not match any of the projected lidar 3D bounding boxes.
[0059] That is, specific objects corresponding to the second GT 2D bounding boxes are detected, but at least some of the projected LIDAR 3D bounding boxes that should be generated for each specific object have not been generated. In this case, objects for which projected LIDAR 3D bounding boxes have not been generated among the specific objects corresponding to the second GT 2D bounding boxes can be determined to be undetected objects. This is believed to occur due to the lack of ability to detect objects located at long distances due to the characteristics of the LIDAR data described in Figure 3 above.
[0060] Therefore, the computing device 1000 can generate (S210_2) pseudo 3D bounding boxes corresponding to each specific object by referring to the average size information and rotation information of each of the 2D bounding boxes for the second GT and the corresponding specific objects.
[0061] Specifically, the computing device 1000 generates position prediction information of a 3D bounding box corresponding to each specific object by referring to position information of a 2D bounding box corresponding to each of the second GT 2D bounding boxes, identifies a specific object class corresponding to each specific object, and then obtains average size information corresponding to each specific object calculated by the 3D lidar model from the 3D lidar model and rotation information corresponding to each specific object from the Yaw regression model.
[0062] That is, even if a lidar 3D bounding box corresponding to a specific object is not generated using a 3D lidar model, a specific pseudo-3D bounding box corresponding to each specific object can be generated by referring to position prediction information, average size information, and rotation information of the 3D bounding box corresponding to each specific object (i.e., undetected object). For reference, referring to average size information and rotation information does not literally mean generating a specific pseudo-3D bounding box by directly applying the actual average size of a specific object, but rather refers to obtaining position prediction information of a 3D bounding box corresponding to the position information of a 2D bounding box by referring to depth information, and treating the actual average size of the specific object adjusted according to perspective by referring to this depth information as average size information to generate a specific pseudo-3D bounding box.
[0063] Thereafter, the computing device 1000 may perform regression on each particular pseudo-3D bounding box (S220_2).
[0064] In this case, the computing device 1000 may calculate a matching ratio between each of the specific pseudo 3D bounding boxes generated using an Intersection Over Union (IOU) metric and each of the 2D bounding boxes for the second GT, determine the degree of separation between them (e.g., their centers) by referring to the matching ratio, and if it is determined that the matching ratio is less than a preset threshold ratio (e.g., which may be, but is not limited to, 100%), perform regression on each of the specific pseudo 2D bounding boxes so that each of the specific pseudo 3D bounding boxes fits to each of the 2D bounding boxes for the second GT. For reference, the process of performing regression on each of the specific pseudo 2D bounding boxes will be described later.
[0065] As another example, the computing device 1000 may determine (S200_3) that at least one second projected rider 3D bounding box among the projected rider 3D bounding boxes does not match any of the GT 2D bounding boxes.
[0066] That is, each of the second projected lidar 3D bounding boxes is generated, but no GT 2D bounding boxes are generated that include the same specific objects as each of the second projected lidar 3D bounding boxes. In this case, each of the specific objects corresponding to each of the second projected lidar 3D bounding boxes can be determined to be a misdetected object. This is believed to occur due to inaccurate calibration data included in the raw data described in Figure 3 above.
[0067] Therefore, by the computing device 1000 removing each of the second projected lidar 3D bounding boxes (S210_3), the accuracy of labeling for generating training data can be improved.
[0068] Thus, the process for solving the problems in FIGS. 3(b) and 3(c) will be described in more detail with reference to FIGS.
[0069] FIG. 4 shows a simplified process for checking whether each of the lidar 3D bounding boxes output by the 3D lidar model matches each of the 2D bounding boxes for GT on the image coordinate system in accordance with one embodiment of the present invention.
[0070] Referring to FIG. 4, first, the computing device 1000 inputs the lidar data included in the raw data into a 3D lidar model and generates at least one lidar 3D bounding box on the vehicle coordinate system using the 3D lidar model (S110). Then, the computing device 1000 may obtain each projected lidar 3D bounding box by projecting each lidar 3D bounding box onto the image coordinate system with reference to each lidar 3D bounding box and the calibration data included in the raw data (S120).
[0071] In this case, the vehicle coordinate system may refer to a VCS (Vehicle Coordinate System), which is a three-dimensional coordinate system with the center position of the vehicle as the origin, but is not limited to this.
[0072] The computing device 1000 may also generate (S130) each of the rider 2D bounding boxes corresponding to each of the projected rider 3D bounding boxes (e.g., surrounding each of the projected rider 3D bounding boxes) by referencing each of the minimum and maximum coordinate values including each of the projected rider 3D bounding boxes, and may determine (S140) whether each of the rider 2D bounding boxes and each of the GT 2D bounding boxes match each other by applying the Hungarian algorithm using the IoU metric to the rider 2D bounding boxes and each of the GT 2D bounding boxes.
[0073] The computing device 1000 can output a predetermined value between "0" and "1" as a result of matching each of the lidar 2D bounding boxes and each of the GT 2D bounding boxes. If it is "0", there is no matching part, so it can be determined that the object was misdetected or not detected from the 3D lidar model. If it is a predetermined value between "0.01" and "0.99", it can be determined that a certain part of the same object matches. If it is "1", it can be determined that the objects match perfectly (i.e., the positions match), but is not limited to this.
[0074] In this regard, the process of fitting each of the projected rider 3D bounding boxes to the 2D bounding box for GT when there is at least a partial match between each of the projected rider 3D bounding boxes and each of the 2D bounding boxes for GT will be described later in FIG. 5. In addition, when there is no complete match between each of the projected rider 3D bounding boxes and each of the 2D bounding boxes for GT, between cases where an object is falsely detected and cases where an object is not detected, the process of generating a pseudo 3D bounding box will be described later in FIG. 6.
[0075] FIG. 5 shows a simplified representation of the regression process by which each matched projected lidar 3D bounding box is fitted to each GT 2D bounding box in the image coordinate system according to one embodiment of the present invention.
[0076] Referring to FIG. 5, the computing device 1000 may generate each of the first lidar 2D bounding boxes (S211_1) by referring to the minimum and maximum coordinate values corresponding to each of the first projected lidar 3D bounding boxes.
[0077] This is because it is easier to generate each of the first rider 2D bounding boxes in the form of a 2D bounding box by referencing each of the first projected 3D bounding boxes, and then fit each of the first rider 2D bounding boxes to match each of the first GT 2D bounding boxes, than to fit each of the first projected rider 3D bounding boxes, which are at least one among the projected rider 3D bounding boxes, to each of the first GT 2D bounding boxes, which are at least one among the GT 2D bounding boxes.
[0078] The computing device 1000 may also calculate a first IOU loss and a first center loss, generate a first integrated loss by referring to the first IOU loss and the first center loss (S212_1), and perform regression by fitting each of the first lidar 2D bounding boxes by referring to the first integrated loss (S213_1). This is to address the problem that, when the calibration data included in the raw data is inaccurate, at least one first projected lidar 3D bounding box among the projected lidar 3D bounding boxes generated by projecting each of the lidar 3D bounding boxes output from the 3D lidar model onto the image coordinate system is displayed apart from at least one first GT 2D bounding box among the GT 2D bounding boxes.
[0079] Specifically, the computing device 1000 may calculate an IOU loss, which is a loss relative to a matching ratio obtained by applying an IOU metric with reference to each of the first lidar 2D bounding boxes and each of the first GT 2D bounding boxes, calculate a center loss with reference to an L1 loss value using the error obtained by taking the absolute value of the difference between each median of each of the first lidar 2D bounding boxes and each median of the first GT 2D bounding box, and generate a first integrated loss by summing the IOU loss and the center loss, but is not limited to this.The computing device 1000 may then perform regression by fitting the frame of each of the first lidar 2D bounding boxes to match the frame of each of the first GT 2D bounding boxes with reference to the first integrated loss.
[0080] In addition, the computing device 1000 can update the parameters of the 3D LIDAR model through backpropagation using the first integrated loss (S214_1), and can repeat the updates to the parameters of the 3D LIDAR model at least N times (e.g., 500 times) to learn.
[0081] FIG. 6 shows a simplified process for generating specific pseudo-3D bounding boxes for specific objects for which 2D bounding boxes for GT have been generated, but no lidar 3D bounding boxes have been generated, according to one embodiment of the present invention.
[0082] Referring to FIG. 6, the computing device 1000 may determine x-axis coordinate values, y-axis coordinate values, and z-axis coordinate values corresponding to at least one second 2D bounding box for the GT among the 2D bounding boxes for the GT as predicted information of a 3D bounding box for each specific object (S211_2_1), determine average size information for each specific object by referring to the average width, average length, and average height for a specific object class corresponding to each specific object calculated by the 3D lidar model from the 3D lidar model (S211_2_2), estimate a first rotation angle for each specific object on the camera coordinate system using a Yaw regression model, and determine a second rotation angle for each specific object on the vehicle coordinate system converted with reference to the first rotation angle and calibration data as rotation information for each specific object (S211_2_3).
[0083] For example, the computing device 1000 may acquire the x-axis coordinate value and the y-axis coordinate value for the position information of the 2D bounding box on the image coordinate system (i.e., the position information for each specific object) by referring to each of the second 2D bounding boxes for GT, which is at least one of the 2D bounding boxes for GT. However, to generate a specific pseudo-3D bounding box, the z-axis coordinate value must be known. Therefore, the computing device 1000 may acquire the z-axis coordinate value for the position information of the 2D bounding box on the image coordinate system by referring to depth information for the position information of the 2D bounding box on the LIDAR data included in the raw data. Then, the computing device 1000 may determine the position prediction information of the 3D bounding box for each specific object by referring to the x-axis coordinate value, y-axis coordinate value, and z-axis coordinate value corresponding to the position information of the 2D bounding box on the image coordinate system. In this case, each of the x-axis coordinate value, y-axis coordinate value, and z-axis coordinate value may be determined as a range value.
[0084] In addition, the 3D lidar model may be linked to a database that calculates and manages the average width, average length, and average height values for each object class (e.g., vehicle, truck, person, etc.) corresponding to each lidar 3D bounding box as lidar data is input and the process of generating each lidar 3D bounding box on the lidar data is repeated.
[0085] Therefore, the computing device 1000 may acquire, from the 3D lidar model, the average width values, the average length values, and the average height values corresponding to each of the specific object classes calculated by the 3D lidar model, by referring to each of the specific object classes corresponding to each of the specific objects identified through each of the second GT 2D bounding boxes, and may determine average size information corresponding to each of the specific objects by referring to the average width values, the average length values, and the average height values corresponding to each of the specific object classes. As mentioned above, the average size information may be obtained by applying a perspective method based on depth information.
[0086] In this case, even if position prediction information and average size information of 3D bounding boxes corresponding to each specific object corresponding to the 2D bounding box for the GT are obtained, if rotation information for each specific object is unknown, the specific pseudo-3D bounding box to be generated may be randomly generated in a form in which each specific object is rotated 45 degrees, 90 degrees, 180 degrees, etc. To prevent this, the computing device 1000 inputs image data included in the raw data into a Yaw regression model and uses the Yaw regression model to estimate and output a first rotation angle as rotation information for each specific object in the camera coordinate system, and further applies calibration data to the first rotation angle to convert it into a second rotation angle for each specific object in the vehicle coordinate system, which can be determined as rotation information for each specific object. For reference, the camera coordinate system may be a CCD (Camera Coordinate System), which is a 3D coordinate system with the position of the camera lens as the origin, but is not limited thereto. The algorithm and architecture of the Yaw regression model for providing rotation information for each specific object will now be described with reference to FIGS.
[0087] FIG. 7 is a simplified diagram illustrating a process for estimating the rotation angle (Yaw) of the vehicle on the vehicle coordinate system from the rotation angle (Alpha) of the vehicle on the camera coordinate system according to an embodiment of the present invention.
[0088] Referring to FIG. 7, for example, when there are mutually perpendicular x-axis and y-axis in a vehicle coordinate system (VCS), the degree to which a vehicle, which is a specific object to be observed, is rotated with reference to the x-axis or y-axis can be defined as Yaw for the vehicle, and the degree to which the vehicle is rotated, calculated with reference to two axes generated based on the direction in which the camera views the vehicle in a camera coordinate system, can be defined as Alpha for the camera.
[0089] That is, how much the vehicle has rotated based on the direction the camera is viewing the vehicle (i.e., Alpha) can be estimated through model learning, but since it is very difficult to directly estimate the degree to which the vehicle has rotated in the vehicle coordinate system (Yaw), we will use a method in which Alpha is estimated in the camera coordinate system and then converted to Yaw by applying calibration data to Alpha. For reference, Alpha can be considered to be defined as the first rotation angle described in Figure 6, and Yaw can be considered to be defined as the second rotation angle.
[0090] FIG. 8 shows a simplified diagram of the Yaw regression model architecture for estimating Alpha according to one embodiment of the present invention.
[0091] 8, the computing device 1000 inputs image data included in raw data to a feature extraction model 400 of the Yaw regression model 400, and causes a feature extraction layer 410 to output feature_3 through feature_5 using the Yaw regression model 400. In this case, as the number of convolution operations increases, feature_1 through feature_5 can be generated, and it may be assumed that feature_3 through feature_5 are selected from among these to perform subsequent processes. The computing device 1000 uses the Yaw regression model 400 to apply an average pooling operation to feature_3 and an upsampling operation to feature_5, and then inputs features_3 to feature_5 to a concatenation layer 420, whereby a concatenation operation is applied to features_3 to feature_5 through the concatenation layer 420 to generate fusion features. The fusion features may be input sequentially to a first convolution layer 431 and a second convolution layer 432, whereby a convolution operation is applied. A mish function may be used as an activation function.
[0092] Thereafter, the computing device 1000 further applies a global average pooling (GAP) operation to the fusion features to which the convolution operation has been applied using the Yaw regression model 400, and then inputs the fusion features to each of the Yaw class header 441, the Yaw delta header 442, and the class header 443, and estimates Alpha as a first rotation angle by referring to the results output through each of the Yaw class header 441 and the Yaw delta header 442. For reference, the feature extraction layer 410 may include, but is not limited to, a VGG (Visual Geometry Group)-based convolution neural network (CNN) such as RegVGG-A1.
[0093] Also, by checking the graph shown as the output value of the Yaw class header 441, it is possible to determine how much the specific object being observed, the vehicle, has rotated in 360 degrees, and it can be seen that multiple multi-bin classes are assigned to 0, 1, 2, and 3 for each of the four quadrants.
[0094] For example, if the head of the vehicle is facing the first quadrant, the computing device 1000 can classify it into the “0” class (i.e., Alpha bin 0) using the Yaw regression model 400, and can obtain a third rotation angle (represented as Delta in FIG. 8) as the angle by which the front of the vehicle is rotated from a preset reference angle between 0 degrees and −90 degrees corresponding to the “0” class, but the third rotation angle may be smaller than the angle range (e.g., 90-degree range) corresponding to each of the multiple multi-bin classes.
[0095] In this case, if "1" is output through the Yaw class header 441, the computing device 1000 can determine that the head of the vehicle is heading toward the fourth quadrant, if "2" is output through the Yaw class header 441, the computing device 1000 can determine that the head of the vehicle is heading toward the third quadrant, and if "3" is output through the Yaw class header 441, the computing device 1000 can determine that the head of the vehicle is heading toward the second quadrant. For reference, each of the preset reference angles for each multi-bin class in the above example may indicate 45 degrees, but is not limited thereto.
[0096] Also, by checking the graph shown as the output value of the Yaw Delta Header 442, it can be confirmed that a predetermined code value is set based on the reference angle within the multi-bin class to which the vehicle head is heading.
[0097] Specifically, the computing device 1000 uses the Yaw regression model 400 to obtain a code value corresponding to the direction indicated by the head of the vehicle by referring to each of a plurality of code values set based on each reference angle of each of a plurality of multi-bin classes through the Yaw delta header 442, and obtains a third adjusted rotation angle by applying the obtained code value to a third rotation angle. Then, the computing device 1000 can estimate Alpha, which is the first rotation angle for the vehicle, by referring to the third adjusted rotation angle and the reference angle.
[0098] For example, if the head direction of the vehicle is classified into the “0” class through the Yaw class header 441, the reference angle within the “0” class is −45 degrees, and Delta, which is the third rotation angle, is 15 degrees, the Yaw regression model 400 may confirm that the sign corresponding to the direction the front of the vehicle is pointing is “-” through the output value output from the Yaw Delta header 442, and may acquire −15 degrees as the third adjusted rotation angle. The Yaw regression model 400 may then estimate that Alpha, which is the first rotation angle for the vehicle in the camera coordinate system, is −60 degrees by referring to the reference angle of −45 degrees and the third adjusted rotation angle of −15 degrees. Therefore, once the computing device 1000 acquires the first rotation angle through the Yaw regression model 400, it may acquire Yaw, which is the second rotation angle in the vehicle coordinate system, by applying the calibration data included in the raw data to the first rotation angle.
[0099] Referring again to FIG. 6, the computing device 1000 may generate specific pseudo-bounding boxes for each specific object by referring to prediction information, average size information, and rotation information of the 3D bounding box, calculate a second IOU loss and a second center loss by referring to the matching ratio between each of the generated specific pseudo-3D bounding boxes and each of the second GT 2D bounding boxes, generate a second integrated loss by referring to the second IOU loss and the second center loss (S221_2), and perform regression by fitting each of the second lidar 2D bounding boxes by referring to the second integrated loss (S222_2). That is, the computing device 1000 supports automatic generation of pseudo 3D bounding boxes for objects not detected by the 3D lidar model by fitting each specific pseudo 3D bounding box to each 2D bounding box for the second GT, thereby improving the accuracy of matching and reducing the amount of manual time and effort required to actually collect learning data.
[0100] For example, the computing device 1000 may generate each of the specific pseudo-2D bounding boxes by referencing the minimum and maximum values of the coordinates including each of the specific pseudo-3D bounding boxes, calculate a second IOU loss, which is a loss relative to a matching ratio obtained by applying an IOU metric by referencing each of the specific pseudo-2D bounding boxes and each of the second GT 2D bounding boxes, calculate a second center loss, which is an L1 loss generated by referencing each difference value between each median of each of the specific pseudo-2D bounding boxes and each of the second GT 2D bounding boxes, generate a second integrated loss by referencing the second IOU loss and the second center loss, and perform regression by fitting the frame of each of the specific pseudo-2D bounding boxes to match the frame of each of the second GT 2D bounding boxes by referencing the second integrated loss. Here, generating each of the specific pseudo-2D bounding boxes by referring to each of the minimum and maximum coordinate values that include each of the specific pseudo-3D bounding boxes may mean generating each of the specific pseudo-2D bounding boxes so as to surround each of the specific pseudo-3D bounding boxes (i.e., so as to circumscribe each of the specific pseudo-3D bounding boxes), but is not limited to this.
[0101] The difference between before and after the labeling process for a specific object is performed by the computing device 1000 will be described with reference to FIG.
[0102] FIG. 9 shows the results of generating a projected lidar 3D bounding box before and after applying the process of the present invention according to one embodiment of the present invention.
[0103] Referring to Figure 9, (a) shows the state in which the LIDAR 3D bounding box output from the 3D LIDAR model before applying the present invention is output as a LIDAR 3D bounding box projected onto image data. If we look at the enlarged left side of the observation area 510 displayed as a black rectangle, we can see that a projected LIDAR 3D bounding box 511 has been generated for the vehicle located at the far left, but that a certain portion is shown to be separated from the vehicle. For the other vehicles 512 to 515, it can be seen that no projected LIDAR 3D bounding box has been generated because the vehicles were not detected due to the characteristics of the LIDAR data, as described in Figure 3.
[0104] On the other hand, (b) of Figure 9 shows the results after the present invention has been applied. If you look at the enlarged left side of the observation area 520 displayed in the black rectangle, you can see that the position of the projected lidar 3D bounding box 521 for the vehicle located on the far left has been adjusted, and that pseudo 3D bounding boxes have been generated for the remaining vehicles 522 to 525.
[0105] The above-described embodiments of the present invention may be embodied in the form of program instructions that can be executed by various computer components and stored on a computer-readable recording medium. The computer-readable recording medium may include program instructions, data files, data structures, and the like, alone or in combination. The program instructions stored on the computer-readable recording medium may be specially designed and constructed for the present invention, or may be known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include not only machine language code, such as that produced by a compiler, but also high-level language code executed by a computer using an interpreter, etc. The hardware devices may be configured to operate as one or more software modules to perform processes according to the present invention, or vice versa.
[0106] Although the present invention has been described above using specific details such as specific components and limited examples and drawings, this is provided merely to aid in a more general understanding of the present invention, and the present invention is not limited to the above examples. Those skilled in the art will be able to make various modifications and variations from such descriptions.
[0107] Therefore, the concept of the present invention should not be limited to the embodiments described above, and all modifications equivalent to or equivalent to the scope of the claims, as well as the scope of the claims described below, are considered to fall within the scope of the concept of the present invention.
Claims
1. 1. A method for labeling at least one specific object by automatically generating at least one specific pseudo 3D bounding box, comprising: (a) a computing device obtains at least one projected lidar 3D bounding box on an image coordinate system using raw data including lidar data, image data, and calibration data, and determines whether each of the projected lidar 3D bounding boxes and each of the GT 2D bounding boxes on the image coordinate system are mutually matched; and (b) (b1) if it is determined that at least one first projected rider 3D bounding box among the projected rider 3D bounding boxes and at least one first GT 2D bounding box among the GT 2D bounding boxes are matched, the computing device performs a regression on each of the first projected rider 3D bounding boxes to fit each of the first projected rider 3D bounding boxes to each of the first GT 2D bounding boxes; and (b2) a regression of each of the GT 2D bounding boxes. If it is determined that at least one second GT 2D bounding box does not match any of the projected lidar 3D bounding boxes, the computing device generates specific pseudo 3D bounding boxes corresponding to each of the specific objects by referring to each of the second GT 2D bounding boxes and average size information and rotation information of each of the second GT 2D bounding boxes and specific objects corresponding to each of the second GT 2D bounding boxes, and performs the regression on each of the specific pseudo 3D bounding boxes to fit each of the specific pseudo 3D bounding boxes to each of the second GT 2D bounding boxes; A method comprising:
2. In the step (a), 2. The method of claim 1, wherein the computing device inputs the lidar data included in the raw data into a 3D lidar model and generates at least one lidar 3D bounding box on a vehicle coordinate system using the 3D lidar model; obtains each of the projected lidar 3D bounding boxes by projecting each of the lidar 3D bounding boxes onto the image coordinate system with reference to each of the lidar 3D bounding boxes and the calibration data included in the raw data; generates each of the lidar 2D bounding boxes by referencing minimum and maximum coordinate values including each of the projected lidar 3D bounding boxes; and applies a Hungarian algorithm using an Intersection Over Union (IOU) metric to determine whether each of the lidar 2D bounding boxes and each of the GT 2D bounding boxes match each other.
3. In the (b1) sub-process, The computing device generates each of the first rider 2D bounding boxes by referencing the minimum and maximum values of coordinates including each of the first projected rider 3D bounding boxes, calculates a first IOU loss, which is a loss relative to a matching ratio obtained by applying an IOU metric with reference to each of the first rider 2D bounding boxes and each of the first GT 2D bounding boxes, and calculates each difference value between each median value of each of the first projected 2D bounding boxes and each median value of each of the first GT 2D bounding boxes. the first integrated loss is generated by referring to the first IOU loss and the first center loss; the regression is performed by fitting each of the first lidar 2D bounding boxes to the first integrated loss; and parameters of a 3D lidar model are updated through backpropagation using the first integrated loss, wherein the 3D lidar model is a model that receives the lidar data included in the raw data and generates at least one lidar 3D bounding box on a vehicle coordinate system.
4. In the (b2) sub-process, The computing device acquires x-axis coordinate values and y-axis coordinate values for position information of the 2D bounding box on the image coordinate system for each of the second GT 2D bounding boxes, acquires z-axis coordinate values for position information of the 2D bounding box on the image coordinate system by referring to depth information of the position information of the 2D bounding box on the lidar data included in the raw data, determines position prediction information of a 3D bounding box for each of the specific objects by referring to the position information of the 2D bounding box, acquires average widths, average lengths, and average heights corresponding to each of the specific object classes calculated by the 3D lidar model from a 3D lidar model by referring to each of the specific object classes corresponding to each of the specific objects confirmed through each of the second GT 2D bounding boxes, determines the average size information for each of the specific objects by referring to the average widths, average lengths, and average heights corresponding to each of the specific object classes, and yaws the image data in the raw data.
2. The method of claim 1, further comprising: inputting the raw data into a yaw regression model and outputting a first rotation angle for each of the specific objects in a camera coordinate system using the yaw regression model; converting the first rotation angle into a second rotation angle for each of the specific objects in the vehicle coordinate system by referring to the first rotation angle and the calibration data included in the raw data, thereby determining the second rotation angle as the rotation information for each of the specific objects; and generating each of the specific pseudo-3D bounding boxes by referring to position prediction information of the 3D bounding boxes, the average size information, and the rotation information for each of the specific objects; and the 3D lidar model is a model that receives the lidar data included in the raw data and generates at least one lidar 3D bounding box on the vehicle coordinate system.
5. In the (b2) sub-process, The computing device inputs the image data into the Yaw regression model to obtain the Yaw 5. The method of claim 4, wherein: a plurality of features are output using a regression model; a concatenation operation and a convolution operation are sequentially applied to each of the plurality of features; a yaw class header is used to classify the plurality of features into a specific multi-bin class corresponding to a forward pointing direction of the specific object among a plurality of multi-bin classes; a yaw class header is used to obtain a third rotation angle by which the specific object is rotated from a specific reference angle pre-set for the specific multi-bin class among pre-set reference angles corresponding to each of the plurality of multi-bin classes; a yaw delta header is used to obtain a specific code value corresponding to a forward pointing direction of the specific object based on each of the reference angles of the plurality of multi-bin classes; a yaw delta header is used to obtain a third adjusted rotation angle by assigning the specific code value to the third rotation angle; and a yaw delta header is used to estimate the first rotation angle for the specific object based on the third adjusted rotation angle and the specific reference angle.
6. In the (b2) sub-process, 5. The method of claim 4, wherein, after each of the specific pseudo-3D bounding boxes is generated, the computing device calculates a matching ratio between each of the specific pseudo-3D bounding boxes and each of the second GT 2D bounding boxes using an IOU metric, and if it is determined that the matching ratio is less than a preset threshold ratio, performs the regression on each of the specific pseudo-3D bounding boxes.
7. In the (b2) sub-process, 7. The method of claim 6, wherein the computing device performs the regression by: generating each of the specific pseudo-2D bounding boxes by referring to a minimum value and a maximum value of coordinates including each of the specific pseudo-3D bounding boxes; calculating a second IOU loss, which is a loss with respect to the matching ratio obtained by applying the IOU metric by referring to each of the specific pseudo-2D bounding boxes and each of the second GT 2D bounding boxes; calculating a second center loss, which is a loss generated by referring to each median value of each of the specific pseudo-2D bounding boxes and each difference value between the respective median values of each of the second GT 2D bounding boxes; generating a second integrated loss by referring to the second IOU loss and the second center loss; and fitting each of the specific pseudo-2D bounding boxes by referring to the second integrated loss.
8. In the (b2) sub-process, 2. The method of claim 1, wherein if it is determined that at least one second projected lidar 3D bounding box among the projected lidar 3D bounding boxes does not match any of the GT 2D bounding boxes, the computing device deletes each of the second projected lidar 3D bounding boxes.
9. In the step (a), 2. The method of claim 1, wherein the calibration data includes camera intrinsic parameters, extrinsic parameters, and transformation parameters for transforming the lidar data into the image data.
10. The method of claim 1, wherein the 2D bounding box for the GT is obtained from data obtained by inputting the image data generated by a predetermined camera, detecting the specific object through a deep learning model, and generating a 2D bounding box for the specific object.
11. 1. A computing device for labeling at least one specific object by automatically generating at least one specific pseudo-3D bounding box, comprising: at least one memory for storing instructions; and at least one processor configured to execute said instructions; The processor (I) uses raw data including lidar data, image data, and calibration data to obtain at least one projected lidar 3D bounding box on an image coordinate system, and determines whether each of the projected lidar 3D bounding boxes and each of the GT 2D bounding boxes on the image coordinate system are mutually matched; and (II) (II-1) if it is determined that at least one first projected lidar 3D bounding box among the projected lidar 3D bounding boxes and at least one first GT 2D bounding box among the GT 2D bounding boxes are matched, fitting each of the first projected lidar 3D bounding boxes to each of the first GT 2D bounding boxes. and (II-2) if it is determined that at least one second GT 2D bounding box among the GT 2D bounding boxes does not match any of the projected rider 3D bounding boxes, generating specific pseudo 3D bounding boxes corresponding to each of the specific objects by referring to each of the second GT 2D bounding boxes and average size information and rotation information of each of the second GT 2D bounding boxes, and performing the regression on each of the specific pseudo 3D bounding boxes to fit each of the specific pseudo 3D bounding boxes to each of the second GT 2D bounding boxes.
12. the processor: In the process (I), 12. The computing device of claim 11, further comprising: inputting the lidar data included in the raw data into a 3D lidar model and generating at least one lidar 3D bounding box on a vehicle coordinate system using the 3D lidar model; acquiring each of the projected lidar 3D bounding boxes by projecting each of the lidar 3D bounding boxes onto the image coordinate system with reference to each of the lidar 3D bounding boxes and the calibration data included in the raw data; generating each of the lidar 2D bounding boxes with reference to minimum and maximum coordinate values including each of the projected lidar 3D bounding boxes; and applying a Hungarian algorithm using an Intersection Over Union (IOU) metric to determine whether each of the lidar 2D bounding boxes and each of the GT 2D bounding boxes match each other.
13. the processor: In the subprocess (II-1), generating each of the first rider 2D bounding boxes by referencing the minimum and maximum values of the coordinates including each of the first projected rider 3D bounding boxes; calculating a first IOU loss, which is a loss relative to a matching ratio obtained by applying an IOU metric with reference to each of the first rider 2D bounding boxes and each of the first GT 2D bounding boxes; and calculating a loss generated by referencing each difference value between each median value of each of the first projected 2D bounding boxes and each median value of each of the first GT 2D bounding boxes.
12. The computing device of claim 11, further comprising: calculating a first center loss where m is a vector of the lidar data included in the raw data; generating a first integrated loss by referring to the first IOU loss and the first center loss; performing the regression by fitting each of the first lidar 2D bounding boxes by referring to the first integrated loss; and updating parameters of a 3D lidar model through backpropagation using the first integrated loss, wherein the 3D lidar model is a model that receives the lidar data included in the raw data and generates at least one lidar 3D bounding box on a vehicle coordinate system.
14. the processor: In the subprocess (II-2), and acquiring x-axis coordinate values and y-axis coordinate values for position information of the 2D bounding box on the image coordinate system for each of the second GT 2D bounding boxes, acquiring z-axis coordinate values for position information of the 2D bounding box on the image coordinate system by referring to depth information of the position information of the 2D bounding box on the LIDAR data included in the raw data, determining position prediction information of a 3D bounding box for each of the specific objects by referring to the position information of the 2D bounding box, acquiring average widths, average lengths, and average heights corresponding to each of the specific object classes calculated by the 3D LIDAR model from a 3D LIDAR model by referring to each of the specific object classes corresponding to each of the specific objects confirmed through each of the second GT 2D bounding boxes, determining the average size information for each of the specific objects by referring to the average widths, average lengths, and average heights corresponding to each of the specific object classes, and y-axis coordinate values and x-axis coordinate values and y-axis coordinate values and z ...
12. The computing device of claim 11, wherein the 3D lidar model inputs the lidar data included in the raw data into a yaw regression model and outputs a first rotation angle for each of the specific objects in the camera coordinate system using the yaw regression model; converts the first rotation angle into a second rotation angle for each of the specific objects in the vehicle coordinate system by referring to the first rotation angle and the calibration data included in the raw data, thereby determining the second rotation angle as the rotation information for each of the specific objects; and generates each of the specific pseudo-3D bounding boxes by referring to position prediction information of the 3D bounding boxes, the average size information for each of the specific objects, and the rotation information; and the 3D lidar model is a model that receives the lidar data included in the raw data and generates at least one lidar 3D bounding box on the vehicle coordinate system.
15. the processor: In the subprocess (II-2), 15. The computing device of claim 14, wherein the image data is input to the Yaw regression model, and each of a plurality of features is output using the Yaw regression model; concatenation and convolution operations are sequentially applied to each of the plurality of features; the image data is classified into a specific multi-bin class corresponding to a forward pointing direction of the specific object among a plurality of multi-bin classes through a Yaw class header; a third rotation angle is obtained by rotating the specific object from a specific reference angle pre-set for the specific multi-bin class among pre-set reference angles corresponding to each of the plurality of multi-bin classes; a specific code value corresponding to a forward pointing direction of the specific object is obtained by referring to each of a plurality of pre-set code values based on each of the reference angles of each of the plurality of multi-bin classes through a Yaw delta header; a third adjusted rotation angle is obtained by assigning the specific code value to the third rotation angle; and the first rotation angle for the specific object is estimated by referring to the third adjusted rotation angle and the specific reference angle.
16. the processor: In the subprocess (II-2), 15. The computing device of claim 14, wherein, when each of the specific pseudo 3D bounding boxes is generated, a matching ratio between each of the specific pseudo 3D bounding boxes and each of the second GT 2D bounding boxes is calculated using an IOU metric, and when it is determined that the matching ratio is less than a preset threshold ratio, the regression is performed on each of the specific pseudo 3D bounding boxes.
17. the processor: In the subprocess (II-2), 17. The computing device of claim 16, wherein the regression is performed by: generating each of the specific pseudo-2D bounding boxes by referencing a minimum value and a maximum value of coordinates including each of the specific pseudo-3D bounding boxes; calculating a second IOU loss, which is a loss with respect to the matching ratio obtained by applying the IOU metric by referencing each of the specific pseudo-2D bounding boxes and each of the second GT 2D bounding boxes; calculating a second center loss, which is a loss generated by referencing each median value of each of the specific pseudo-2D bounding boxes and each difference value between the median values of each of the second GT 2D bounding boxes; generating a second integrated loss by referencing the second IOU loss and the second center loss; and fitting each of the specific pseudo-2D bounding boxes by referencing the second integrated loss.
18. the processor: In the subprocess (II-2), 12. The computing device of claim 11, wherein if it is determined that at least one second projected rider 3D bounding box among the projected rider 3D bounding boxes does not match any of the GT 2D bounding boxes, the computing device deletes each of the second projected rider 3D bounding boxes.
19. the processor: In the process (I), 12. The computing device of claim 11, wherein the calibration data includes camera intrinsic parameters, extrinsic parameters, and transformation parameters for transforming the lidar data into the image data.
20. 12. The computing device of claim 11, wherein the 2D bounding box for GT is obtained from data obtained by inputting the image data generated by a predetermined camera, detecting the specific object through a deep learning model, and generating a 2D bounding box for the specific object.
Citation Information
Patent Citations
Database construction system for article recognition algorism machine-learning
JP2017102838A
Information processing method, information processing apparatus and program
JP2020021326A
Method and system generating correct answer data for machine learning of recognition unit
JP2024000347A
Status determination device, status determination system, and status determination method
WO2022097426A1