Method for marking specific object and computing device using same
By using the computing device to automatically generate pseudo-3D bounding boxes using lidar and image data, and fit them with GT with 2D bounding boxes, the problem of cumbersome marking and high error rate in the prior art is solved, and a more efficient and accurate marking process is achieved.
Patent Information
- Application Number
- CN202411859604.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-11-12
- Filing Date
- 2024-12-17
- Publication Date
- 2025-06-24
AI Technical Summary
When generating 3D bounding boxes, existing marking tools require staff to manually set the size and rotation information, which leads to cumbersome work and prone to marking errors, and it is difficult to efficiently handle the matching and fit of 3D bounding boxes.
By using lidar data, image data and calibration data, the computing device automatically generates a specific pseudo-3D bounding box, uses regression technology to fit it with the GT with a 2D bounding box, generates a corresponding specific pseudo-3D bounding box, and adjusts it through average size information and rotation information.
It improves the efficiency and accuracy of 3D bounding box marking, reduces staff operation errors, simplifies the matching and fitting process of 3D bounding box, and reduces the working hours of subsequent inspections.
Smart Images

Figure CN120198746A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for labeling at least one specific object by automatically generating at least one specific pseudo 3D bounding box and a computing device using the method. Background Art
[0002] In recent years, research has been conducted on methods for identifying objects using machine learning. A neural network has multiple hidden layers between an input layer and an output layer, and deep learning using a neural network, as part of machine learning, has high recognition performance.
[0003] In addition, a neural network using deep learning learns by using backpropagation of loss. To perform the learning of a deep learning network, training data in which an object is labeled (i.e., a label is added) by a labeling tool is required. In order to prepare such training data, a labeling tool that originally implements a simple working function (e.g., generating a 3D bounding box using a mouse pointer, etc.) is used, and thus has the advantage of a short development man-hour for the labeling tool.
[0004] However, when performing work using an existing labeling tool, a worker directly generates a 3D bounding box using a mouse pointer. Due to repetitive work, the concentration of the worker decreases, and thus labeling errors frequently occur. Even if the worker is educated in advance about labeling, the problem of labeling errors still appears. Therefore, additional man-hours are required to inspect the completed work results.
[0005] In addition, when generating a 2D bounding box, a worker can drag using a mouse pointer while confirming an object in an image so that the generated box includes the object. However, in a state where three-dimensional coordinate information (x, y, z) corresponding to the object is confirmed to generate a 3D bounding box, it is necessary to appropriately set the size information (width, height, length) and rotation information (roll, pitch, yaw) of the 3D bounding box. Since this work is highly difficult, a large amount of time is required for each object.
[0006] Therefore, an improved solution that can solve the above problems is needed. Summary of the Invention
[0007] Technical Problem
[0008] An object of the present invention is to solve the above problems.
[0009] In addition, another object of the present invention is to use raw data including lidar data, image data, and calibration data to respectively obtain at least one projected lidar 3D bounding box projected onto an image coordinate system. i) When it is determined that each first projected lidar 3D bounding box matches each first GT 2D bounding box, perform regression on each of the first projected lidar 3D bounding boxes so that each of the first projected lidar 3D bounding boxes fits each of the first GT 2D bounding boxes, where each of the first projected lidar 3D bounding boxes is at least one of the projected lidar 3D bounding boxes, and each of the first GT 2D bounding boxes is at least one of the GT 2D bounding boxes; ii) When it is determined that each second GT 2D bounding box does not match any of the projected lidar 3D bounding boxes, respectively generate specific pseudo 3D bounding boxes corresponding to each of the specific objects by referring to each of the second GT 2D bounding boxes, the average size information and rotation information of each corresponding specific object, and perform the regression on each of the specific pseudo 3D bounding boxes so that each of the specific pseudo 3D bounding boxes fits each of the second GT 2D bounding boxes, where each of the second GT 2D bounding boxes is at least one of the GT 2D bounding boxes.
[0010] Technical solution
[0011] The features of the present invention are used to achieve the above object of the present invention and the feature effects of the present invention described hereinafter, and its structure is as follows.
[0012] According to an embodiment of the present invention, there is provided a method for labeling at least one specific object by automatically generating at least one specific pseudo 3D bounding box, which includes the following steps: a) The computing device uses raw data including lidar data, image data, and calibration data to respectively obtain at least one projected lidar 3D bounding box projected onto the image coordinate system, and determines whether each of the projected lidar 3D bounding boxes matches each GT 2D bounding box on the image coordinate system; and b) Perform the following sub-processes: b1) When it is determined that each first projected lidar 3D bounding box and each first GT 2D bounding box match, the computing device performs regression on each of the first projected lidar 3D bounding boxes to fit each of the first projected lidar 3D bounding boxes with each of the first GT 2D bounding boxes, where each of the first projected lidar 3D bounding boxes is at least one of the projected lidar 3D bounding boxes, and each of the first GT 2D bounding boxes is at least one of the GT 2D bounding boxes; and b2) When it is determined that each second GT 2D bounding box does not match any of the projected lidar 3D bounding boxes, the computing device respectively generates specific pseudo 3D bounding boxes corresponding to each of the specific objects by referring to each of the second GT 2D bounding boxes, the average size information and rotation information of each specific object corresponding to each of the second GT 2D bounding boxes, and performs the regression on each of the specific pseudo 3D bounding boxes to fit each of the specific pseudo 3D bounding boxes with each of the second GT 2D bounding boxes, where each of the second GT 2D bounding boxes is at least one of the GT 2D bounding boxes.
[0013] In step a), the computing device inputs the lidar data included in the raw data into a 3D lidar model to respectively generate at least one lidar 3D bounding box on the vehicle coordinate system by the 3D lidar model, projects each of the lidar 3D bounding boxes onto the image coordinate system by referring to each of the lidar 3D bounding boxes and the calibration data included in the raw data to respectively obtain the projected lidar 3D bounding boxes, generates each lidar 2D bounding box by referring to the minimum values and maximum values of the coordinates including each of the projected lidar 3D bounding boxes, and uses the Hungarian algorithm applicable to the IOU index to determine whether each of the lidar 2D bounding boxes matches each of the GT 2D bounding boxes.
[0014] In the sub - process of b1), the computing device generates each first lidar 2D bounding box by referring to the minimum and maximum values of the coordinates of each of the first projected lidar 3D bounding boxes, calculates the first IOU loss by referring to each of the first lidar 2D bounding boxes and each of the first GT 2D bounding boxes, calculates the first center loss, generates the first comprehensive loss by referring to the first IOU loss and the first center loss, performs the regression by fitting each of the first lidar 2D bounding boxes with reference to the first comprehensive loss, and updates the parameters of the 3D lidar model by using the backpropagation of the first comprehensive loss. Among them, the first IOU loss is the loss relative to the matching rate obtained by applying the IOU metric, the first center loss is generated by referring to the differences between the center values of each of the first projected 2D bounding boxes and the center values of each of the first GT 2D bounding boxes, and the 3D lidar model is a model that generates at least one lidar 3D bounding box on the vehicle coordinate system by receiving the lidar data included in the original data.
[0015] In the sub - process b2), the computing device obtains the x - axis coordinate values and y - axis coordinate values of the 2D bounding box position information on the image coordinates relative to each 2D bounding box for the second GTs by referring to each of the second GTs' 2D bounding boxes. It obtains the z - axis coordinate value of the 2D bounding box position information on the image coordinate system by referring to the depth information of the 2D bounding box position information included in the lidar data of the original data. It determines the 3D bounding box position prediction information relative to each specific object by referring to the 2D bounding box position information. It obtains each of the average width, average length, and average height corresponding to each specific object category pre - calculated by the 3D lidar model from the 3D lidar model by referring to each specific object category corresponding to each 2D bounding box confirmed by the second GTs. It determines the average size information relative to each specific object by referring to each of the average width, average length, and average height corresponding to each specific object category. It inputs the image data included in the original data into the yaw regression model to make the yaw regression model output the first rotation angle relative to each specific object on the camera coordinate system. It converts the first rotation angle to the second rotation angle relative to each specific object on the vehicle coordinate system by referring to the first rotation angle and the calibration data included in the original data, and determines the second rotation angle as the rotation information relative to each specific object. It generates each specific pseudo - 3D bounding box by referring to the 3D bounding box position prediction information, the average size information relative to each specific object, and the rotation information. Here, the 3D lidar model is a model that generates at least one lidar 3D bounding box on the vehicle coordinate system by receiving the lidar data included in the original data.
[0016] In the sub - process b2), the computing device inputs the image data into the yaw regression model, and the yaw regression model performs the following operations: outputs multiple each feature, after sequentially applying cascade calculation and convolution calculation to the multiple each feature, classifies them into a specific multi - interval class corresponding to the forward - pointing direction of the specific object among multiple multi - interval classes through a yaw class header, obtains a third rotation angle, obtains a specific symbol value corresponding to the forward - pointing direction of the specific object by referring to each of multiple symbol values preset by the yaw class header based on each standard angle of each of the multiple multi - interval classes, obtains a third adjusted rotation angle by assigning the specific symbol value to the third rotation angle, and estimates a first rotation angle relative to the specific object by referring to the third adjusted rotation angle and a specific standard angle, where the third rotation angle is the angle by which the specific object rotates from a specific standard angle preset corresponding to the specific multi - interval class among each standard angle preset corresponding to each of the multiple multi - interval classes.
[0017] In the sub - process b2), if each of the specific pseudo 3D bounding boxes is generated, the computing device calculates the matching rate between each of the specific pseudo 3D bounding boxes and each of the second GT 2D bounding boxes using the IOU metric, and when it is determined that the matching rate is less than a preset threshold ratio, performs the regression on each of the specific pseudo 3D bounding boxes.
[0018] In the sub - process b2), the computing device generates each specific pseudo 2D bounding box by referring to the minimum and maximum values of the coordinates of each of the specific pseudo 3D bounding boxes, calculates a second IOU loss by referring to each of the specific pseudo 2D bounding boxes and each of the second GT 2D bounding boxes, calculates a second center loss, generates a second comprehensive loss by referring to the second IOU loss and the second center loss, and performs regression by fitting each of the specific pseudo 2D bounding boxes with reference to the second comprehensive loss, where the second IOU loss is the loss relative to the matching rate obtained by applying the IOU metric, and the second center loss is the loss generated by referring to the differences between the center values of each of the specific pseudo 2D bounding boxes and the center values of each of the second GT 2D bounding boxes.
[0019] In the sub - process b2), when it is determined that each of the second projected lidar 3D bounding boxes does not match any of the GT 2D bounding boxes, the computing device deletes each of the second projected lidar 3D bounding boxes, where each of the second projected lidar 3D bounding boxes is at least one of the projected lidar 3D bounding boxes.
[0020] In the step a), the calibration data includes conversion parameters for converting the internal parameters, external parameters of the camera, and the lidar data into the image data.
[0021] The 2D bounding box of the GT is obtained from data that, after receiving the image data generated by a predetermined camera, detects the specific object through a deep learning model and generates a 2D bounding box for the specific object.
[0022] According to another embodiment of the present invention, there is provided a computing device for labeling at least one specific object by automatically generating at least one specific pseudo 3D bounding box, comprising: at least one memory storing instructions; and at least one processor configured to execute the instructions, wherein the processor performs the following processes: I) using raw data including lidar data, image data, and calibration data to respectively obtain at least one projected lidar 3D bounding box projected onto an image coordinate system, and determining whether each of the projected lidar 3D bounding boxes matches each GT 2D bounding box on the image coordinate system; and II) performing the following sub-processes: II-1) when it is determined that each first projected lidar 3D bounding box matches each first GT 2D bounding box, performing regression on each of the first projected lidar 3D bounding boxes to fit each of the first projected lidar 3D bounding boxes to each of the first GT 2D bounding boxes, wherein each of the first projected lidar 3D bounding boxes is at least one of the projected lidar 3D bounding boxes, and each of the first GT 2D bounding boxes is at least one of the GT 2D bounding boxes; and II-2) when it is determined that each second GT 2D bounding box does not match any of the projected lidar 3D bounding boxes, generating specific pseudo 3D bounding boxes corresponding to each of the specific objects by referring to each of the second GT 2D bounding boxes, average size information, and rotation information of each specific object corresponding to each of the second GT 2D bounding boxes, and performing the regression on each of the specific pseudo 3D bounding boxes to fit each of the specific pseudo 3D bounding boxes to each of the second GT 2D bounding boxes, wherein in the process of I), the processor inputs the lidar data included in the raw data into a 3D lidar model to cause the 3D lidar model to respectively generate at least one lidar 3D bounding box on a vehicle coordinate system, projects each of the lidar 3D bounding boxes onto the image coordinate system by referring to each of the lidar 3D bounding boxes and the calibration data included in the raw data to respectively obtain the projected lidar 3D bounding boxes, generates each lidar 2D bounding box by referring to each minimum value and each maximum value of the coordinates including each of the projected lidar 3D bounding boxes, and uses the Hungarian algorithm applicable to the IOU metric to determine whether each of the lidar 2D bounding boxes matches each of the GT 2D bounding boxes.
[0023] In the sub-process of II-1), the processor generates each first lidar 2D bounding box by referring to each minimum value and each maximum value of the coordinates of each of the first projected lidar 3D bounding boxes, calculates the first IOU loss by referring to each of the first lidar 2D bounding boxes and each of the first GT 2D bounding boxes, calculates the first center loss, generates the first comprehensive loss by referring to the first IOU loss and the first center loss, performs the regression by fitting each of the first lidar 2D bounding boxes by referring to the first comprehensive loss, and updates the parameters of the 3D lidar model by using the backpropagation of the first comprehensive loss. Among them, the first IOU loss is the loss relative to the matching rate obtained by applying the IOU metric, the first center loss is generated by referring to the differences between the center values of each of the first projected 2D bounding boxes and the center values of each of the first GT 2D bounding boxes, and the 3D lidar model is a model that generates at least one lidar 3D bounding box on the vehicle coordinate system by receiving the lidar data included in the original data.
[0024] In the sub-process II-2), the processor obtains each x-axis coordinate value and each y-axis coordinate value of the 2D bounding box position information on the image coordinates with respect to each 2D bounding box for the second GT by referring to each second GT, obtains the z-axis coordinate value of the 2D bounding box position information on the image coordinate system by referring to the depth information of the 2D bounding box position information included in the lidar data included in the original data, determines the 3D bounding box position prediction information with respect to each specific object by referring to the 2D bounding box position information, obtains each of the average width, average length, and average height pre-calculated by the 3D lidar model corresponding to each specific object category from the 3D lidar model by referring to each specific object category corresponding to each specific object confirmed by each 2D bounding box for the second GT, determines the average size information corresponding to each specific object by referring to each of the average width, average length, and average height corresponding to each specific object category, inputs the image data included in the original data into a yaw regression model to cause the yaw regression model to output a first rotation angle with respect to each specific object on the camera coordinate system, converts the first rotation angle into a second rotation angle with respect to each specific object on the vehicle coordinate system by referring to the first rotation angle and the calibration data included in the original data, and determines the second rotation angle as the rotation information with respect to each specific object, and generates each specific pseudo 3D bounding box by referring to the 3D bounding box position prediction information, the average size information with respect to each specific object, and the rotation information, wherein the 3D lidar model is a model that generates at least one lidar 3D bounding box on the vehicle coordinate system by receiving the lidar data included in the original data.
[0025] In the sub-process II-2), the processor inputs the image data into the yaw regression model, and the yaw regression model performs the following operations: outputting a plurality of each feature, after sequentially applying cascade calculation and convolution calculation to the plurality of each feature, classifying them into a specific multi-interval class corresponding to the forward direction of the specific object among a plurality of multi-interval classes through a yaw class header, obtaining a third rotation angle, obtaining a specific symbol value corresponding to the forward direction of the specific object by referring to each of a plurality of symbol values preset by the yaw class header based on each standard angle of each of the plurality of multi-interval classes, obtaining a third adjusted rotation angle by assigning the specific symbol value to the third rotation angle, and estimating a first rotation angle relative to the specific object by referring to the third adjusted rotation angle and a specific standard angle, where the third rotation angle is the angle by which the specific object rotates from a specific standard angle preset corresponding to the specific multi-interval class among each standard angle preset corresponding to each of the plurality of multi-interval classes.
[0026] In the sub-process II-2), if each specific pseudo 3D bounding box is generated, the processor calculates the matching rate between each specific pseudo 3D bounding box and each second GT 2D bounding box using the IOU metric, and when it is determined that the matching rate is less than a preset threshold ratio, the regression is performed on each specific pseudo 3D bounding box.
[0027] In the sub-process II-2), the processor generates each specific pseudo 2D bounding box by referring to the minimum values and maximum values of the coordinates including each specific pseudo 3D bounding box, calculates a second IOU loss by referring to each specific pseudo 2D bounding box and each second GT 2D bounding box, calculates a second center loss, generates a second comprehensive loss by referring to the second IOU loss and the second center loss, and performs regression by fitting each specific pseudo 2D bounding box with reference to the second comprehensive loss, where the second IOU loss is the loss relative to the matching rate obtained by applying the IOU metric, and the second center loss is the loss generated by referring to the differences between the center values of each specific pseudo 2D bounding box and the center values of each second GT 2D bounding box.
[0028] In the sub-process II-2), when it is determined that each second projected lidar 3D bounding box does not match any of the GT 2D bounding boxes, the processor deletes each second projected lidar 3D bounding box, where each second projected lidar 3D bounding box is at least one of the projected lidar 3D bounding boxes.
[0029] In the process of I), the calibration data includes conversion parameters for converting the internal parameters, external parameters of the camera, and the lidar data into the image data.
[0030] The GT 2D bounding box is obtained from data that, after receiving the image data generated by a predetermined camera, detects the specific object through a deep learning model and generates a 2D bounding box for the specific object.
[0031] Advantageous Effects
[0032] The effect of the present invention is that, using the original data including lidar data, image data, and calibration data, at least one projected lidar 3D bounding box projected onto the image coordinate system is obtained respectively. i) When it is determined that each first projected lidar 3D bounding box and each first GT 2D bounding box match, regression is performed on each of the first projected lidar 3D bounding boxes to fit each of the first projected lidar 3D bounding boxes with each of the first GT 2D bounding boxes, where each of the first projected lidar 3D bounding boxes is at least one of the projected lidar 3D bounding boxes, and each of the first GT 2D bounding boxes is at least one of the GT 2D bounding boxes; ii) When it is determined that each second GT 2D bounding box does not match any of the projected lidar 3D bounding boxes, specific pseudo 3D bounding boxes corresponding to each of the specific objects are respectively generated by referring to each of the second GT 2D bounding boxes, the average size information and rotation information of the corresponding specific objects, and regression is performed on each of the specific pseudo 3D bounding boxes to fit each of the specific pseudo 3D bounding boxes with each of the second GT 2D bounding boxes, where each of the second GT 2D bounding boxes is at least one of the GT 2D bounding boxes. Description of the Drawings
[0033] The following drawings for illustrating the embodiments of the present invention are only a part of the embodiments of the present invention. Those skilled in the art (hereinafter referred to as "skilled artisans") in the field to which the present invention pertains can obtain other drawings based on the following drawings without creative work.
[0034] Figure 1 It is a schematic diagram of a computing device for marking at least one specific object by generating at least one specific pseudo 3D bounding box according to an embodiment of the present invention.
[0035] Figure 2 It is a schematic diagram of a process for marking at least one specific object by generating at least one specific pseudo 3D bounding box according to an embodiment of the present invention.
[0036] Figure 3It is a schematic diagram of an example of a lidar 3D bounding box output from a 3D lidar model according to an embodiment of the present invention.
[0037] Figure 4 It is a schematic diagram of a process for confirming whether each lidar 3D bounding box output from a 3D lidar model according to an embodiment of the present invention matches each GT 2D bounding box on an image coordinate system.
[0038] Figure 5 It is a schematic diagram of a process for fitting each projected lidar 3D bounding box that matches according to an embodiment of the present invention to each GT 2D bounding box on an image coordinate system.
[0039] Figure 6 It is a schematic diagram of a process for generating a specific pseudo 2D bounding box for each characteristic object that has a GT 2D bounding box generated but does not generate a lidar 3D bounding box according to an embodiment of the present invention.
[0040] Figure 7 It is a schematic diagram of a process for estimating the rotation angle (yaw angle) of a vehicle in a vehicle coordinate system using the vehicle rotation angle (α) in a camera coordinate system according to an embodiment of the present invention.
[0041] Figure 8 It is a schematic diagram of the architecture of a yaw regression model for estimating α according to an embodiment of the present invention.
[0042] Figure 9 It is a diagram showing the generation results of projected lidar 3D bounding boxes before and after applying the present invention according to an embodiment of the present invention. Detailed Description of the Invention
[0043] The following detailed description of the present invention refers to the accompanying drawings shown as examples of specific embodiments capable of implementing the present invention, so that the purpose, technical solution, and advantages of the present invention are obvious. These embodiments will be described in detail so that those skilled in the art can fully implement the present invention.
[0044] In addition, throughout the description of the present invention and the claims, the term "comprising" and its variations are not intended to exclude other technical features, additional components, components, or steps. Other objects, advantages, and features of the present invention will be obvious to those skilled in the art through this specification and partly through the embodiments of the present invention. The examples and drawings provided below are for illustration and are not intended to limit the present invention.
[0045] In addition, the present invention includes all possible combinations of the embodiments in this specification. It should be understood that although the various embodiments in the present invention are different, they do not necessarily exclude each other. For example, an embodiment related to a specific shape, structure, and characteristics described in this specification can be implemented with another embodiment without departing from the spirit and scope of the present invention. In addition, it should be understood that the positions or configurations of the respective constituent elements in each disclosed embodiment can be changed without departing from the spirit and scope of the present invention. Therefore, the following detailed description is not limiting, and if the scope of the present invention can be appropriately described, it is limited to all scopes equivalent to the claims requested and the appended claims. Similar reference numerals in the drawings refer to the same or similar functions in all aspects.
[0046] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings so that those skilled in the art to which the present invention pertains can easily implement the present invention.
[0047] Figure 1 is a schematic diagram of a computing device that labels at least one specific object by generating at least one specific pseudo 3D bounding box according to an embodiment of the present invention.
[0048] Referring to Figure 1 , the computing device (1000) may include a memory (1010) and a processor (1020). The memory (1010) stores instructions for labeling at least one specific object by automatically generating at least one specific pseudo 3D bounding box. The processor (1020) automatically generates at least one specific pseudo 3D bounding box to label at least one specific object by corresponding to the instructions stored in the memory (1010). At this time, the computing device (1000) may include a personal computer (PC), a mobile computer, etc.
[0049] Specifically, the computing device (1000) can typically utilize a combination of a computing device (for example, a device including a computer processor, memory, storage, input devices, and output devices, and other existing computing device components; electronic communication devices such as routers and switches; electronic information storage systems such as network-attached storage (NAS) and storage area networks (SAN)) and computer software (that is, instructions that can enable the computing device to function in a specific manner) to achieve the required system performance.
[0050] In addition, the processor of the computing device may include hardware configurations such as an MPU (Micro Processing Unit) or a CPU (Central Processing Unit), a cache memory, and a data bus. In addition, the computing device may further include a software configuration including an operating system and an application program for a specific purpose.
[0051] However, it does not exclude that the computing device includes a medium, a processor, and a memory integrated form for implementing the present invention, that is, an integrated processor.
[0052] In addition, the computing device (1000) can be linked with the database (900) by automatically generating at least one specific pseudo 3D bounding box, and the database (900) includes information for marking at least one specific object. Among them, the database (900) can include at least one type of storage medium such as a flash memory type, a hard disk type, a multimedia card micro type, a cartridge memory (e.g., SD memory or XD memory), RAM (Random Access Memory), SRAM (Static Random Access Memory), ROM (ReadOnlyMemory), EEPROM (Electrically Erasable Programmable ReadOnly Memory), PROM (Programmable ReadOnly Memory), magnetic memory, magnetic disk, optical disc, and is not limited thereto, and it can include any medium capable of storing data. In addition, the database (900) can be separately provided from the computing device (1000), or differently arranged inside the computing device (1000), thereby transmitting data or recording the received data, or can be separated into two or more than two differently from the drawings to be implemented, which can be different according to the implementation conditions of the invention.
[0053] Figure 2 It is a schematic diagram of a process for marking at least one specific object by generating at least one specific pseudo 3D bounding box according to an embodiment of the present invention.
[0054] First, referring to Figure 2 , the computing device (1000) can use the raw data including image data, lidar data, and calibration data to respectively obtain at least one projected lidar 3D bounding box on the image coordinate system, and determine whether each of the projected lidar 3D bounding boxes matches each GT 2D bounding box on the image coordinate system (S100).
[0055] For example, a computing device (1000) may input lidar data included in the original data into a 3D lidar model, so that the 3D lidar model generates and outputs respective lidar 3D bounding boxes for each object included in the lidar data through calculation, and input image data into SVNet3 (a deep learning model that detects objects by receiving image data generated by a camera and generates 2D bounding boxes for the objects, having high detection accuracy and thus can be used as GT), so that SVNet3 generates and outputs respective 2D bounding boxes for each GT with respect to each object on the image coordinate system. However, the main body for generating the 2D bounding boxes for each GT from the image data is not limited to SVNet3, and the 2D bounding boxes for each GT can also be generated by other known methods that can generate bounding boxes more accurately than the lidar model and used as GT.
[0056] For another example, in a state where the respective lidar 3D bounding boxes and the 2D bounding boxes for each GT are stored in and managed by a database (900), the computing device (1000) may obtain the respective 3D lidar models and the 2D bounding boxes for each GT from the database (900), and the respective lidar 3D bounding boxes and the 2D bounding boxes for each GT are obtained through the result of pre-enabling the 3D lidar model and each SVNet3 to output the respective lidar 3D bounding boxes and the 2D bounding boxes for each GT.
[0057] As a reference, the calibration data may include the intrinsic parameters and extrinsic parameters of the camera, as well as parameters for converting lidar data into image data, etc. The image coordinate system may refer to the ICS (Image Coordinate System) that is a two-dimensional coordinate system with the upper left corner of the image obtained by the camera as the origin, but is not limited thereto.
[0058] In addition, the computing device (1000) confirms whether there is a match by pairing the respective lidar 3D bounding boxes output from the 3D lidar model and the 2D bounding boxes for each GT on the image data (i.e., the data output from the SVNet3 model by receiving the image data obtained by the camera). At this time, in order to confirm whether there is a match by pairing the lidar 3D bounding boxes and the 2D bounding boxes for GT, the computing device (1000) projects the lidar 3D bounding boxes onto the image coordinate system to obtain the respective lidar 3D bounding boxes projected onto the image coordinate system, and thus pairs the respective projected lidar 3D bounding boxes and the 2D bounding boxes for each GT to confirm whether there is a match, but is not limited thereto.
[0059] At this time, refer to Figure 3 Describe an example of the lidar 3D bounding box output from the 3D lidar model.
[0060] Figure 3 It is a schematic diagram of a lidar 3D bounding box instance output from a 3D lidar model according to an embodiment of the present invention.
[0061] Referring to Figure 3 , the 3D bounding boxes (301, 302, 311, 313, 315, 321) shown in (a) to (c) may be states in which the lidar 3D bounding boxes included in the lidar data output from the 3D lidar model are projected onto the image coordinate system by using calibration data to obtain the projected lidar 3D bounding boxes respectively. As a reference, for ease of description, Figure 3 (a) and Figure 3 (c) do not show the GT 2D bounding boxes that precisely match the projected lidar 3D bounding boxes.
[0062] In Figure 3 (a), it can be confirmed that the projected lidar 3D bounding boxes (301, 302) for each of the person and the vehicle on the image coordinate system have been generated; (b) shows that due to inaccurate calibration data included in the original data, the projected lidar 3D bounding boxes (311, 313) are at a certain distance from the GT 2D bounding boxes for each of the person and the vehicle on the image coordinate system, and when confirming the right side of the GT 2D bounding box 314 for the vehicle, it can be confirmed that the projected lidar 3D bounding box (315) for the false positive object has been obtained; in (c), the projected lidar 3D bounding box (315) for the close object has been generated, but due to the characteristic of the lidar data having less point cloud data, the vehicle (322) as an object for other distant objects is missed (false negative), so the projected lidar 3D bounding box for the vehicle 322 is not obtained.
[0063] That is, taking the projected lidar 3D bounding boxes (311, 313, 315, 321) such as in Figure 3 (b) and (c) as an example, the 3D bounding boxes (311, 313) are shown at a certain distance from their corresponding GT 2D bounding boxes, and the projected 3D bounding box (315) for the false positive object is generated or the projected 3D bounding box is not generated due to the missed object. This kind of problem is not suitable as learning data for autonomous vehicles, so this problem needs to be solved.
[0064] Referring to Figure 2 again, when the computing device (1000) confirms the matching results of each projected lidar 3D bounding box and each GT 2D bounding box, three situations may occur, and the processes for handling each of the three situations are described below.
[0065] For example, the computing device (1000) may determine that each first projected lidar 3D bounding box (which is at least one of the projected lidar 3D bounding boxes) and each first GT 2D bounding box (which is at least one of the GT 2D bounding boxes) match (S200_1).
[0066] That is, in the case where there is at least a partially overlapping portion between each first projected lidar 3D bounding box and each first GT 2D bounding box, although the positions of each first projected lidar 3D bounding box and the positions of each first GT 2D bounding box are not precisely consistent due to inaccurate calibration data or the like, it can be determined that each first object corresponding to each first projected lidar 3D bounding box and each first GT 2D bounding box matches as the same object.
[0067] Therefore, when it is determined that each first projected lidar 3D bounding box and each first GT 2D bounding box match, the computing device (1000) may perform regression on each first projected lidar 3D bounding box (S210_1).
[0068] At this time, as described above, the GT 2D bounding box can be used as the GT because of its high detection accuracy. Therefore, in order to fit each first projected lidar 3D bounding box and each first GT 2D bounding box, the computing device (1000) may perform regression on each first projected lidar 3D bounding box. For reference, the process of performing regression on each first projected lidar 3D bounding box is specifically described below in Figure 5 .
[0069] For another example, it can be determined that each second GT 2D bounding box (which is at least one of the GT 2D bounding boxes) does not match any of the projected lidar 3D bounding boxes (S200_2).
[0070] That is, it is the case where each specific object corresponding to each second GT 2D bounding box is detected, but at least a part of the projected lidar 3D bounding boxes generated corresponding to each specific object is not generated. In this case, the object for which the projected lidar 3D bounding box is not generated among each specific object corresponding to each second GT 2D bounding box can be determined as a missed detection object. The reason can be considered to be the characteristics of the lidar data as described in Figure 3 resulting in insufficient ability to detect distant objects.
[0071] Therefore, the computing device (1000) can generate respective pseudo 3D bounding boxes corresponding to respective specific objects (S210_2) by referring to the 2D bounding boxes for respective second GTs, the average size information, and the rotation information of respective corresponding specific objects.
[0072] Specifically, the computing device (1000) can generate 3D bounding box position prediction information corresponding to respective specific objects by referring to the 2D bounding box position information of the 2D bounding boxes for respective second GTs, obtain the average size information corresponding to respective specific objects from the 3D lidar model after confirming the specific object categories of respective specific objects, and obtain the rotation information corresponding to respective specific objects from the yaw regression model.
[0073] That is, even if the 3D lidar model does not generate a lidar 3D bounding box corresponding to a specific object, it is possible to generate and substitute respective specific pseudo 3D bounding boxes corresponding to respective specific objects (i.e., missed detected objects) by referring to the 3D bounding box position prediction information, the average size information, and the rotation information of respective corresponding specific objects. As a reference, referring to the average size information and the rotation information may not mean that the actual average size of a specific object is directly applied as it is to generate a specific pseudo 3D bounding box, but rather to obtain 3D bounding box position prediction information for the 2D bounding box position information by referring to depth information, and adjust the actual average size of a specific object according to perspective method with reference to such depth information and use it as the average size information, thereby generating a specific pseudo 3D bounding box.
[0074] Then, the computing device (1000) can perform regression on respective specific pseudo 3D bounding boxes (S220_2).
[0075] At this time, the computing device (1000) calculates the matching rate between respective specific pseudo 3D bounding boxes generated by the IOU (Intersection Over Union) metric and the 2D bounding boxes for respective second GTs. When determining the degree of separation between the two (e.g., their centers) with reference to the matching rate and determining that it is less than a preset threshold ratio (e.g., it can be 100%, but not limited thereto), regression can be performed on respective specific pseudo 2D bounding boxes to fit respective specific pseudo 3D bounding boxes to the 2D bounding boxes for respective second GTs. As a reference, the process of performing regression on respective specific pseudo 2D bounding boxes will be described later.
[0076] For another example, the computing device (1000) can determine that none of the respective second projected lidar 3D bounding boxes (which are at least one of the projected lidar 3D bounding boxes) match any of the 2D bounding boxes for GT (S200_3).
[0077] That is, although each second projected lidar 3D bounding box is generated, a GT 2D bounding box including each specific object identical to each second projected lidar 3D bounding box is not generated. In this case, each specific object corresponding to each second projected lidar 3D bounding box can be determined as a misdetected object. The reason is that, as described above Figure 3 in the description, the approved data included in the original data may be inaccurate.
[0078] Therefore, the computing device (1000) can improve the accuracy of the labels for generating learning data by deleting each second projected lidar 3D bounding box (S210_3).
[0079] As described above, refer to Figures 4 to 8 for a further detailed description of the process for solving the Figure 3 problems in (b) and (c).
[0080] Figure 4 is a schematic diagram of a process for confirming whether each lidar 3D bounding box output by a 3D lidar model according to an embodiment of the present invention matches each GT 2D bounding box on an image coordinate system.
[0081] Refer to Figure 4 , first, the computing device (1000) can input the lidar data included in the original data into the 3D lidar model to cause the 3D lidar model to respectively generate at least one lidar 3D bounding box on the vehicle coordinate system (S110), and project each lidar 3D bounding box onto the image coordinate system by referring to each lidar 3D bounding box and the calibration data included in the original data, thereby obtaining each projected lidar 3D bounding box (S120).
[0082] At this time, the vehicle coordinate system may refer to the VCS (Vehicle Coordinate System), that is, a three-dimensional coordinate system with the center position of the vehicle as the origin, but is not limited thereto.
[0083] In addition, the computing device (1000) can generate each lidar 2D bounding box corresponding to each projected lidar 3D bounding box (for example, surrounding each projected lidar 3D bounding box) by referring to each minimum value and each maximum value of the coordinates including each projected lidar 3D bounding box, and apply the Hungarian algorithm to each of the lidar 2D bounding box and the GT 2D bounding box using the IoU metric, thereby determining whether each lidar 2D bounding box and each GT 2D bounding box match (S140).
[0084] As a result of matching each lidar 2D bounding box and each GT 2D bounding box, the computing device (1000) may output a predetermined value between "0" and "1". If it is "0", there is no matching part, so it can be determined that the 3D lidar model misdetects or misses an object. If it is a predetermined value between "0.01" and "0.99", it can be determined that it is a state of partial matching of the same object. If it is "1", it can be determined that the object is completely matched (i.e., the positions are the same), but it is not limited to this.
[0085] Among them, the process of fitting each projected lidar 3D bounding box to the GT 2D bounding box when each projected lidar 3D bounding box and each GT 2D bounding box are partially matched will be Figure 5 described in, and as an example where each projected lidar 3D bounding box and each GT 2D bounding box do not match at all, the process of generating a pseudo 3D bounding box for the missed object among the misdetected object or the missed object will be Figure 6 described in.
[0086] Figure 5 is a schematic diagram of the process of fitting each projected lidar 3D bounding box matched according to an embodiment of the present invention to each GT 2D bounding box on the image coordinate system.
[0087] Referring to Figure 5 , the computing device (1000) can generate each first lidar 2D bounding box (S211_1) by referring to the minimum values and the maximum values of the coordinates corresponding to each first projected lidar 3D bounding box.
[0088] The reason is that compared with fitting each first projected lidar 3D bounding box (which is at least one of the projected lidar 3D bounding boxes) to each first GT 2D bounding box (which is at least one of the GT 2D bounding boxes), it is easier to generate each first lidar 2D bounding box in the form of a 2D bounding box by referring to each first projected lidar 3D bounding box and perform fitting to match each first lidar 2D bounding box with each first GT 2D bounding box.
[0089] In addition, the computing device (1000) can calculate a first IOU loss and a first center loss, generate a first comprehensive loss by referring to the first IOU loss and the first center loss (S212_1), and perform regression by fitting each first lidar 2D bounding box with reference to the first comprehensive loss. The purpose is to compensate for the problem that each first projected lidar 3D bounding box is separated from the first GT 2D bounding box (which is at least one of the GT 2D bounding boxes) when the calibration data included in the original data is inaccurate. Here, the first projected lidar 3D bounding box is at least one of the projected lidar 3D bounding boxes generated by projecting each lidar 3D bounding box output by the 3D lidar model onto the image coordinate system.
[0090] Specifically, the computing device (1000) calculates the IOU loss (which is the loss relative to the matching rate obtained by applying the IOU metric) by applying the IOU metric with reference to each first lidar 2D bounding box and each first GT 2D bounding box, calculates the center loss by referring to the L1 loss value, and generates the first comprehensive loss by adding the IOU loss and the center loss. Here, the L1 loss value is the error obtained by applying the absolute value to each difference between the center values of each first lidar 2D bounding box and the center values of each first GT 2D bounding box, but is not limited thereto. In addition, the computing device (1000) can fit the edges of each first lidar 2D bounding box with the edges of each first GT 2D bounding box by referring to the first comprehensive loss to perform regression.
[0091] In addition, the computing device (1000) can update the parameters of the 3D lidar model by backpropagation (which uses the first comprehensive loss) (S214_1), and perform learning by repeating the update of the parameters of the 3D lidar model N times (for example, 500 times).
[0092] Figure 6 It is a schematic diagram of a process of not generating a lidar 3D bounding box according to an embodiment of the present invention, but generating a specific pseudo 2D bounding box for each characteristic object for which each GT 2D bounding box is generated.
[0093] Refer to Figure 6, the computing device (1000) may determine the x-axis coordinate value, y-axis coordinate value, and z-axis coordinate value corresponding to each second GT with a 2D bounding box (which is at least one of the GT with 2D bounding boxes) as 3D bounding box prediction information (S211_2_1) relative to each specific object, determine average dimension information (S211_2_2) relative to each specific object by referring to the average width, average length, and average height of each specific object category pre-calculated by the 3D lidar model from the 3D lidar model, estimate the first rotation angle relative to each specific object on the camera coordinate system using a yaw regression model, and determine the second rotation angle (which is converted by referring to the first rotation angle and calibration data) relative to each specific object on the vehicle coordinate system as rotation information (S211_2_3) relative to each specific object.
[0094] For example, the computing device (1000) may obtain each x-axis coordinate value and each y-axis coordinate value of the 2D bounding box position information (i.e., the position information relative to each specific object) relative to the image coordinates by referring to each second GT with a 2D bounding box (which is at least one of the GT with 2D bounding boxes). However, to generate a specific pseudo 3D bounding box, the z-axis coordinate value needs to be known. Therefore, the computing device (1000) may obtain the z-axis coordinate value of the 2D bounding box position information relative to the image coordinate system by referring to the depth information included in the original data regarding the 2D bounding box position information on the lidar data. Then, the computing device (1000) may determine the 3D bounding box position prediction information relative to each specific object by referring to the x-axis coordinate value, y-axis coordinate value, and z-axis coordinate value corresponding to the 2D bounding box position information. At this time, each of the x-axis coordinate value, y-axis coordinate value, and z-axis coordinate value may be determined as a range value.
[0095] In addition, when the 3D lidar model repeatedly generates each lidar 3D bounding box on the lidar data by receiving lidar data, it may be linked with a database that calculates and manages the average width value, average length value, and average height value corresponding to each object category (e.g., vehicle, truck, person, etc.) corresponding to each lidar 3D bounding box.
[0096] Therefore, the computing device (1000) may obtain each of the average width value, average length value, and average height value pre-calculated by the 3D lidar model corresponding to each specific object category by referring to each specific object category corresponding to each specific object (confirmed by each second GT with a 2D bounding box), and determine the average dimension information corresponding to each specific object by referring to each of the average width value, average length value, and average height value corresponding to each specific object category. As described above, the average dimension information may be obtained by applying perspective based on depth information.
[0097] At this time, even if the 3D bounding box position prediction information and average size information corresponding to each specific object of the corresponding GT's 2D bounding box are obtained, if the rotation information with respect to each specific object is not determined, the generated specific pseudo 3D bounding box may be randomly generated in the form of each specific object rotating by 45 degrees, 90 degrees, 180 degrees, etc. To avoid this situation, the computing device (1000) can input the image data included in the original data into the yaw regression model to cause the yaw regression model to estimate and output the first rotation angle as the rotation information with respect to each specific object in the camera coordinate system, convert it to the second rotation angle with respect to each specific object in the vehicle coordinate system by additionally applying calibration data to the first rotation angle, and determine it as the rotation information with respect to each specific object. As a reference, the camera coordinate system may be a CCD (Camera Coordinate System), that is, a three-dimensional coordinate system with the lens position of the camera as the origin, but is not limited thereto. As described above, refer to Figure 7 and Figure 8 Describe the algorithm and architecture of the yaw regression model for providing the rotation information with respect to each specific object.
[0098] Figure 7 is a schematic diagram of the process of estimating the rotation angle (yaw angle) of the vehicle in the vehicle coordinate system from the vehicle rotation angle (α) in the camera coordinate system according to an embodiment of the present invention.
[0099] Refer to Figure 7 , assuming that there are an x-axis and a y-axis perpendicular to each other in the vehicle coordinate system (VCS), by referring to the x-axis and the y-axis, the degree of rotation of a specific object (i.e., the vehicle) to be observed can be defined as the yaw angle with respect to the vehicle, and the degree of rotation of the vehicle calculated by referring to two axes (blue lines) generated based on the direction in which the camera in the camera coordinate system faces the vehicle can be defined as α with respect to the camera.
[0100] That is, the degree of rotation of the vehicle based on the direction in which the camera faces the vehicle (i.e., α) can be estimated through model training, and it is difficult to directly estimate the degree of rotation of the vehicle in the vehicle coordinate system (yaw angle). Therefore, after estimating α in the camera coordinate system, a method of converting it to the yaw angle by applying calibration data to α is required. As a reference, α should be regarded as defined as Figure 6 the first rotation angle described in
[0101] Figure 8 is a schematic diagram of the architecture of the yaw regression model for estimating α according to an embodiment of the present invention.
[0102] Refer toFigure 8 The computing device (1000) can cause the feature extraction layer (410) of the yaw regression model (400) to output Feature_3 to Feature_5 by inputting the image data included in the original data into the feature extraction model (400) of the yaw regression model (400). At this time, as the number of convolutional calculations increases, Features 1 to 5 can be generated, but it can be assumed that the subsequent process is carried out by selecting Feature_3 to Feature_5. Then, the computing device (1000) can cause the yaw regression model (400) to apply average pooling calculation to Feature_3, apply upsampling calculation to Feature_5, and then input Feature_3 to Feature_5 into the concatenation layer (420) so that the concatenation calculation is applied to Feature_3 to Feature_5 by the concatenation layer (420) to generate a fused feature, and apply convolutional calculation by inputting the fused feature into the first convolutional layer (431) and the second convolutional layer (432) in sequence, and the Mish function can be used as the activation function.
[0103] Then, the computing device (1000) can cause the yaw regression model (400) to apply global average pooling (GAP) calculation to the fused feature to which convolutional calculation has been applied, and then input the fused feature into each of the yaw class header (441), the yaw increment header (442), and the class header (443), and estimate α as the first rotation angle with reference to the results output by each of the yaw class header (441) and the yaw increment header (442). As a reference, the feature extraction layer (410) can include, but is not limited to, a convolutional neural network (CNN) based on VGG (Visual Geometry Group) such as RegVGG-A1.
[0104] In addition, by confirming the graph shown as the output value of the yaw class header (441), the rotation degree of a specific object (i.e., a vehicle) to be observed can be determined in 360 degrees, and multiple multi-bin classes are assigned 0, 1, 2, 3 to the four quadrants.
[0105] For example, when the head of the vehicle faces the first quadrant, the computing device (1000) can cause the yaw regression model (400) to be classified into the "0" class (i.e., αbin0), and cause it to obtain the third rotation angle (represented by δ in Figure 8 ) as the rotation angle of the vehicle in front from the standard angle (preset between 0 degrees and -90 degrees corresponding to the "0" class), and the third rotation angle can be less than the angle range corresponding to each of the multiple multi-bin classes (e.g., 90-degree range).
[0106] At this time, if "1" is output through the yaw category header (441), the computing device (1000) can determine that the vehicle head is facing the 4th quadrant. If "2" is output through the yaw category header (441), it can be determined that the vehicle head is facing the 3rd quadrant. If "3" is output through the yaw category header (441), it can be determined that the vehicle head is facing the 2nd quadrant. For reference, each preset standard angle in each multi-interval category in the above example can represent 45 degrees, but it is not limited to this.
[0107] In addition, when confirming the graph of the output value of the yaw increment header (442), it can be confirmed that it is in a state where a predetermined symbol value is set based on the standard angle in the multi-interval category where the vehicle head is facing.
[0108] Specifically, the computing device (1000) can enable the yaw regression model (400) to obtain the symbol value corresponding to the direction pointed by the vehicle head by referring to each of the multiple symbol values set based on the standard angle of each of the multiple multi-interval categories through the yaw increment header (442). The third adjusted rotation angle is obtained by assigning the obtained symbol value to the third rotation angle. The first rotation angle (i.e., α) relative to the vehicle is estimated by referring to the third adjusted rotation angle and the standard angle.
[0109] If the vehicle head is classified as the "0" category through the yaw category header (441), the standard angle within the "0" category is -45 degrees, and δ as the third rotation angle is 15 degrees, the yaw regression model (400) confirms that the symbol of the direction pointed by the vehicle ahead corresponding to the output value output by the yaw increment header (442) is "-", thereby being able to obtain -15 degrees as the third adjusted rotation angle. In addition, the yaw regression model (400) can estimate that α as the first rotation angle relative to the vehicle in the camera coordinate system is -60 degrees by referring to the standard angle of -45 degrees and -15 degrees as the third adjusted rotation angle. Therefore, when the computing device (1000) obtains the first rotation angle through the yaw regression model (400), the yaw angle (Yaw) as the second rotation angle in the vehicle coordinate system can be obtained by applying the calibration data included in the original data to the first rotation angle.
[0110] Refer to again Figure 6, the computing device (1000) generates respective specific pseudo 3D bounding boxes (S212_2) by referring to the 3D bounding box prediction information, average size information, and rotation information with respect to each specific object, calculates the second IOU loss and the second center loss by referring to the matching rate between the generated respective specific pseudo 3D bounding boxes and the second GT 2D bounding boxes, generates a second comprehensive loss by referring to the second IOU loss and the second center loss (S221_2), and performs regression by fitting each second lidar 2D bounding box by referring to the second comprehensive loss (S222_2). That is, by fitting each specific pseudo 3D bounding box to each second GT 2D bounding box, the computing device (1000) can support the automatic generation of pseudo 3D bounding boxes for objects missed by the 3D lidar model while improving the matching accuracy, thereby reducing the man-hours for staff to inspect the actual training data.
[0111] For example, the computing device (1000) can generate respective specific pseudo 2D bounding boxes by referring to the respective minimum values and maximum values including the coordinates of the respective specific pseudo 3D bounding boxes, calculate the second IOU loss (which is the loss with respect to the matching rate obtained by applying the IOU metric) by referring to the respective specific pseudo 2D bounding boxes and the second GT 2D bounding boxes, calculate the second center loss as the L1 loss (which is generated by referring to the respective differences between the respective center values of the respective specific pseudo 2D bounding boxes and the respective center values of the second GT 2D bounding boxes), generate a second comprehensive loss by referring to the second IOU loss and the second center loss, and perform regression by fitting to suppress the edges of the respective specific pseudo 2D bounding boxes to the edges of the second GT 2D bounding boxes by referring to the second comprehensive loss. Herein, generating the respective specific pseudo 2D bounding boxes by referring to the respective minimum values and maximum values including the coordinates of the respective specific pseudo 3D bounding boxes means generating the respective specific pseudo 2D bounding boxes to enclose (i.e., circumscribe) the respective specific pseudo 3D bounding boxes, but is not limited thereto.
[0112] As described above, with reference to Figure 9 the differences before and after the computing device (1000) performs the marking process on a specific object are described.
[0113] Figure 9 is a diagram showing the generation results of the projected lidar 3D bounding boxes before and after applying the process in accordance with an embodiment of the present invention.
[0114] With reference to Figure 9, (a) shows the state where the lidar 3D bounding box output by the 3D lidar model before applying the present invention is output as a projected lidar 3D bounding box on the image data. When confirming the left part of the enlarged green frame observation area (510), a projected lidar 3D bounding box (511) is generated for the leftmost vehicle, but there is a certain distance from the vehicle, and the remaining vehicles (512 to 515) are missed due to the characteristics of the lidar data as described in Figure 3 , so it can be confirmed that no projected lidar 3D bounding box is generated.
[0115] On the other hand, Figure 9 (b) of
[0116] shows the result after applying the present invention. When confirming the left part of the enlarged green frame observation area (520), it can be seen that the position of the projected lidar 3D bounding box (521) for the leftmost vehicle is adjusted, and pseudo 3D bounding boxes are generated for the remaining vehicles (522 to 525).
[0116] In addition, the embodiments according to the present invention described above can be implemented in the form of program commands executed by various computer components and recorded on a computer-readable medium. The computer-readable recording medium may independently or combinatorially include program commands, data files, data structures, etc. The program instructions recorded on the computer-readable recording medium may be designed and configured specifically for the present invention, or may also be known and available to those skilled in the field of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs, DVDs, magneto-optical media such as floptical disks, and hardware devices specifically configured to store and execute program commands such as ROMs, RAMs, and flash memories. Examples of program commands include not only machine language codes generated by compilers but also high-level language codes executable by computers using interpreters, etc. The hardware device may be configured to operate as one or more software modules to perform the processing according to the present invention, and vice versa.
[0117] As above, the present invention has been described through specific matters such as specific components, limited embodiments, and drawings, but these are for the purpose of helping a more comprehensive understanding of the present invention. The present invention is not limited to these embodiments, and those skilled in the art to which the present invention pertains can attempt various modifications and variations based on these descriptions.
[0118] Therefore, the technical concept of the present invention should not be limited to the embodiments described above, and all contents of the claims of the present invention and their equivalent or equivalent transformations belong to the technical concept of the present invention.
Claims
1. A method for marking a specific object, wherein the method marks at least one specific object by automatically generating at least one specific pseudo 3D bounding box, the method comprising the following steps: a) the computing device uses raw data including lidar data, image data, and calibration data to respectively obtain at least one projected lidar 3D bounding box projected onto the image coordinate system, and determines whether each of the projected lidar 3D bounding boxes matches each GT 2D bounding box on the image coordinate system; and b) Execute the following sub-processes: b1) When it is determined that each first projected lidar 3D bounding box matches each first GT 2D bounding box, the computing device performs regression on each first projected lidar 3D bounding box to fit each first projected lidar 3D bounding box with each first GT 2D bounding box, wherein each first projected lidar 3D bounding box is at least one of the projected lidar 3D bounding boxes, and each first GT 2D bounding box is at least one of the GT 2D bounding boxes; and b2) When it is determined that each second GT 2D bounding box does not match any of the projected lidar 3D bounding boxes, the computing device generates specific pseudo 3D bounding boxes corresponding to each specific object by referring to each second GT 2D bounding box and average size information and rotation information of each specific object corresponding to each second GT 2D bounding box, and performs the regression on each specific pseudo 3D bounding box to fit each specific pseudo 3D bounding box with each second GT 2D bounding box, wherein each second GT 2D bounding box is at least one of the GT 2D bounding boxes.
2. The method according to claim 1, characterized in that: In the step a), The computing device inputs the lidar data included in the original data into the 3D lidar model so that the 3D lidar model generates at least one lidar 3D bounding box on the vehicle coordinate system, projects each lidar 3D bounding box onto the image coordinate system by referring to each lidar 3D bounding box and the calibration data included in the original data, and obtains the projected lidar 3D bounding boxes respectively, generates each lidar 2D bounding box by referring to each minimum value and each maximum value of the coordinates of each projected lidar 3D bounding box, and applies the Hungarian algorithm by using the IOU indicator to determine whether each lidar 2D bounding box matches each GT 2D bounding box.
3. The method according to claim 1, characterized in that: In the sub-process b1), The computing device generates each first laser radar 2D bounding box by referring to each minimum value and each maximum value of the coordinates of each first projected laser radar 3D bounding box, calculates a first IOU loss by referring to each first laser radar 2D bounding box and each first GT 2D bounding box, calculates a first center loss, generates a first comprehensive loss by referring to the first IOU loss and the first center loss, performs the regression by fitting each first laser radar 2D bounding box with reference to the first comprehensive loss, and updates the parameters of the 3D laser model by back propagation using the first comprehensive loss, wherein the first IOU loss is a loss relative to the matching rate obtained by applying the IOU indicator, the first center loss is generated by referring to each difference between each center value of each first projected 2D bounding box and each center value of each first GT 2D bounding box, and the 3D laser model is a model of at least one laser radar 3D bounding box generated in a vehicle coordinate system by receiving the laser radar data included in the original data.
4. The method according to claim 1, characterized in that: In the sub-process b2), The computing device obtains each x-axis coordinate value and each y-axis coordinate value of the 2D bounding box position information on the image coordinate system relative to each second GT 2D bounding box by referring to each second GT 2D bounding box, obtains the z-axis coordinate value of the 2D bounding box position information on the image coordinate system relative to the 2D bounding box position information on the image coordinate system by referring to the depth information of the 2D bounding box position information of the lidar data included in the original data, determines the 3D bounding box position prediction information relative to each of the specific objects by referring to the 2D bounding box position information, obtains each of the average width, average length, and average height corresponding to each of the specific object categories pre-calculated by the 3D lidar model by referring to each specific object category corresponding to each of the specific objects confirmed by each second GT 2D bounding box, and obtains each of the average width, average length, and average height corresponding to each of the specific object categories pre-calculated by the 3D lidar model by referring to the average width, average length, and average height corresponding to each of the specific object categories. The average size information corresponding to each of the specific objects is determined by measuring the average height of each of the specific objects, the image data included in the original data is input into a yaw regression model to enable the yaw regression model to output a first rotation angle relative to each of the specific objects in the camera coordinate system, the first rotation angle is converted into a second rotation angle relative to each of the specific objects in the vehicle coordinate system by referring to the first rotation angle and the calibration data included in the original data, and the second rotation angle is determined as the rotation information relative to each of the specific objects, each of the specific pseudo 3D bounding boxes is generated by referring to the 3D bounding box position prediction information, the average size information relative to each of the specific objects, and the rotation information, wherein the 3D lidar model is a model that generates at least one lidar 3D bounding box in the vehicle coordinate system by receiving the lidar data included in the original data.
5. The method according to claim 4, characterized in that: In the sub-process b2), The computing device inputs the image data into the yaw regression model, so that the yaw regression model performs the following operations: outputs a plurality of features, classifies the features into a specific multi-interval category corresponding to the forward pointing direction of the specific object among a plurality of multi-interval categories through a yaw category header after sequentially applying cascade calculation and convolution calculation to the plurality of features, obtains a third rotation angle, obtains a specific symbol value corresponding to the forward pointing direction of the specific object by referring to each of a plurality of symbol values preset by the yaw category header based on each standard angle of each of the plurality of multi-interval categories, obtains a third adjusted rotation angle by assigning the specific symbol value to the third rotation angle, and estimates the first rotation angle relative to the specific object by referring to the third adjusted rotation angle and the specific standard angle, wherein the third rotation angle is an angle by which the specific object is rotated from a specific standard angle preset corresponding to the specific multi-interval category among the standard angles preset corresponding to each of the plurality of multi-interval categories.
6. The method according to claim 4, characterized in that: In the sub-process b2), If each of the specific pseudo 3D bounding boxes is generated, the computing device calculates the matching rate between each of the specific pseudo 3D bounding boxes and each of the second GT 2D bounding boxes using the IOU indicator, and when it is determined that the matching rate is less than a preset threshold ratio, the regression is performed on each of the specific pseudo 3D bounding boxes.
7. The method according to claim 6, characterized in that: In the sub-process b2), The computing device generates each specific pseudo 2D bounding box by referring to each minimum value and each maximum value of the coordinates of each specific pseudo 3D bounding box, calculates a second IOU loss by referring to each specific pseudo 2D bounding box and each second GT 2D bounding box, calculates a second center loss, generates a second comprehensive loss by referring to the second IOU loss and the second center loss, and performs regression by fitting each specific pseudo 2D bounding box with reference to the second comprehensive loss, wherein the second IOU loss is a loss relative to the matching rate obtained by applying the IOU indicator, and the second center loss is a loss generated by referring to the differences between each center value of each specific pseudo 2D bounding box and each center value of each second GT 2D bounding box.
8. The method according to claim 1, characterized in that: In the sub-process b2), When it is determined that each second projected lidar 3D bounding box does not match any of the GT 2D bounding boxes, the computing device deletes each second projected lidar 3D bounding box, wherein each second projected lidar 3D bounding box is at least one of the projected lidar 3D bounding boxes.
9. The method according to claim 1, characterized in that: In the step a), The calibration data includes conversion parameters for converting the camera's internal parameters, external parameters, and the lidar data into the image data.
10. The method according to claim 1, characterized in that: The GT 2D bounding box is obtained from data of detecting the specific object through a deep learning model after receiving the image data generated by a predetermined camera and generating a 2D bounding box for the specific object.
11. A computing device, the computing device marking at least one specific object by automatically generating at least one specific pseudo 3D bounding box, the computing device comprising: at least one memory storing instructions; as well as at least one processor configured to execute said instructions, The processor performs the following process: I) using raw data including lidar data, image data, and calibration data to respectively obtain at least one projected lidar 3D bounding box projected onto an image coordinate system, and determining whether each of the projected lidar 3D bounding boxes matches each GT 2D bounding box on the image coordinate system; and II) Execute the following sub-processes: II-1) when it is determined that each first projected lidar 3D bounding box matches each first GT 2D bounding box, regression is performed on each of the first projected lidar 3D bounding boxes to fit each of the first projected lidar 3D bounding boxes with each of the first GT 2D bounding boxes, wherein each of the first projected lidar 3D bounding boxes is at least one of the projected lidar 3D bounding boxes, and each of the first GT 2D bounding boxes is at least one of the GT 2D bounding boxes; and II-2) When it is determined that each second GT 2D bounding box does not match any of the projected lidar 3D bounding boxes, specific pseudo 3D bounding boxes corresponding to each specific object are generated respectively by referring to each second GT 2D bounding box and the average size information and rotation information of each specific object corresponding to each second GT 2D bounding box, and the regression is performed on each specific pseudo 3D bounding box to fit each specific pseudo 3D bounding box with each second GT 2D bounding box, wherein each second GT 2D bounding box is at least one of the GT 2D bounding boxes.
12. The computing device according to claim 11, characterized in that: In the process 1), The processor inputs the lidar data included in the original data into the 3D lidar model so that the 3D lidar model generates at least one lidar 3D bounding box on the vehicle coordinate system, projects each lidar 3D bounding box onto the image coordinate system by referring to each lidar 3D bounding box and the calibration data included in the original data to obtain the projected lidar 3D bounding boxes, generates each lidar 2D bounding box by referring to the minimum values and maximum values of the coordinates of each projected lidar 3D bounding box, and applies the Hungarian algorithm using the IOU indicator to determine whether each lidar 2D bounding box matches each GT 2D bounding box.
13. The computing device according to claim 11, characterized in that: In the sub-process II-1), The processor generates each first laser radar 2D bounding box by referring to each minimum value and each maximum value of the coordinates of each first projected laser radar 3D bounding box, calculates a first IOU loss by referring to each first laser radar 2D bounding box and each first GT 2D bounding box, calculates a first center loss, generates a first comprehensive loss by referring to the first IOU loss and the first center loss, performs the regression by fitting each first laser radar 2D bounding box with reference to the first comprehensive loss, and updates the parameters of the 3D laser model by back propagation using the first comprehensive loss, wherein the first IOU loss is a loss relative to the matching rate obtained by applying the IOU indicator, the first center loss is generated by referring to each difference between each center value of each first projected 2D bounding box and each center value of each first GT 2D bounding box, and the 3D laser model is a model of generating at least one laser radar 3D bounding box in a vehicle coordinate system by receiving the laser radar data included in the original data.
14. The computing device according to claim 11, characterized in that: In the II-2) sub-process, The processor obtains each x-axis coordinate value and each y-axis coordinate value of the 2D bounding box position information on the image coordinates relative to each second GT 2D bounding box by referring to each second GT 2D bounding box, obtains the z-axis coordinate value of the 2D bounding box position information on the image coordinate system by referring to the depth information of the 2D bounding box position information of the lidar data included in the original data, determines the 3D bounding box position prediction information relative to each of the specific objects by referring to the 2D bounding box position information, obtains each of the average width, average length, and average height corresponding to each of the specific object categories pre-calculated by the 3D lidar model by referring to each specific object category corresponding to each of the specific objects confirmed by each second GT 2D bounding box, and obtains each of the average width, average length, and average height corresponding to each of the specific object categories pre-calculated by the 3D lidar model by referring to the average width, average length, and average height corresponding to each of the specific object categories. The average size information corresponding to each of the specific objects is determined by referring to each of the average heights, the image data included in the original data is input into a yaw regression model so that the yaw regression model outputs a first rotation angle relative to each of the specific objects in the camera coordinate system, the first rotation angle is converted into a second rotation angle relative to each of the specific objects in the vehicle coordinate system by referring to the first rotation angle and the calibration data included in the original data, and the second rotation angle is determined as the rotation information relative to each of the specific objects, each of the specific pseudo 3D bounding boxes is generated by referring to the 3D bounding box position prediction information, the average size information relative to each of the specific objects, and the rotation information, wherein the 3D lidar model is a model that generates at least one lidar 3D bounding box in the vehicle coordinate system by receiving the lidar data included in the original data.
15. The computing device according to claim 14, characterized in that: In the II-2) sub-process, The processor inputs the image data into the yaw regression model, so that the yaw regression model performs the following operations: outputs a plurality of features, classifies the features into a specific multi-interval category corresponding to the forward pointing direction of the specific object among a plurality of multi-interval categories through a yaw category header after sequentially applying cascade calculation and convolution calculation to the plurality of features, obtains a third rotation angle, obtains a specific symbol value corresponding to the forward pointing direction of the specific object by referring to each of a plurality of symbol values preset by the yaw category header based on each standard angle of each of the plurality of multi-interval categories, obtains a third adjusted rotation angle by assigning the specific symbol value to the third rotation angle, and estimates the first rotation angle relative to the specific object by referring to the third adjusted rotation angle and the specific standard angle, wherein the third rotation angle is an angle by which the specific object is rotated from a specific standard angle preset corresponding to the specific multi-interval category among the standard angles preset corresponding to each of the plurality of multi-interval categories.
16. The computing device according to claim 14, characterized in that: In the II-2) sub-process, If each of the specific pseudo 3D bounding boxes is generated, the processor calculates the matching rate between each of the specific pseudo 3D bounding boxes and each of the second GT 2D bounding boxes using the IOU indicator, and when it is determined that the matching rate is less than a preset threshold ratio, the regression is performed on each of the specific pseudo 3D bounding boxes.
17. The computing device according to claim 16, characterized in that: In the II-2) sub-process, The processor generates each specific pseudo 2D bounding box by referring to each minimum value and each maximum value of the coordinates of each specific pseudo 3D bounding box, calculates a second IOU loss by referring to each specific pseudo 2D bounding box and each second GT 2D bounding box, calculates a second center loss, generates a second comprehensive loss by referring to the second IOU loss and the second center loss, and performs regression by fitting each specific pseudo 2D bounding box with reference to the second comprehensive loss, wherein the second IOU loss is the loss relative to the matching rate obtained by applying the IOU indicator, and the second center loss is the loss generated by referring to the differences between the center values of each specific pseudo 2D bounding box and the center values of each second GT 2D bounding box.
18. The computing device according to claim 11, characterized in that: In the II-2) sub-process, When it is determined that each second projected lidar 3D bounding box does not match any of the GT 2D bounding boxes, the processor deletes each second projected lidar 3D bounding box, wherein each second projected lidar 3D bounding box is at least one of the projected lidar 3D bounding boxes.
19. The computing device according to claim 11, characterized in that: In the process 1), The calibration data includes conversion parameters for converting the camera's internal parameters, external parameters, and the lidar data into the image data.
20. The computing device of claim 11, wherein: The GT 2D bounding box is obtained from data of detecting the specific object through a deep learning model after receiving the image data generated by a predetermined camera and generating a 2D bounding box for the specific object.