METHOD FOR ANNOTATION BASED ON LIDAR MAP AND COMPUTING DEVICE USING THE SAME
By selecting key frames for annotation in lidar maps and using both lidar and camera images, the method addresses inefficiencies and inconsistencies in existing annotation methods, improving efficiency and accuracy of training data for AI systems.
Patent Information
- Application Number
- JP2024213351
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2024-11-05
- Filing Date
- 2024-12-06
- Publication Date
- 2025-07-28
- Estimated Expiration
- 2044-12-06
AI Technical Summary
Existing annotation methods for artificial intelligence training data are inefficient due to repetitive work and inconsistent results, particularly when annotating frames in lidar maps, which negatively impact learning processes.
A method that selects key frames from a plurality of frames to generate a lidar map and performs annotation only on these key frames, utilizing both lidar maps and camera images to ensure consistency and efficiency, and accumulates annotations for each object in a single map data set.
Improves annotation efficiency and accuracy by reducing repetitive work, ensuring consistent annotation across frames, and enabling management of annotated data in a single map, thereby enhancing the quality of training data for AI systems.
Smart Images

Figure 0007714103000001 
Figure 0007714103000002 
Figure 0007714103000003
Abstract
Description
Technical Field
[0001] The present invention relates to a method for annotating based on a rider map and a computing device using the same.
Background Art
[0002] Annotation means performing labeling on data to create training data so that artificial intelligence can understand the content of the data.
[0003] Generally, since annotation has to be performed for each frame, when annotation has to be repeatedly performed on the same target, the working speed is reduced and it is difficult to ensure the consistency of annotation. Therefore, there is a problem that efficient annotation is difficult to perform.
[0004] Also, there may be a problem that the annotation result in a previous frame is reset in the next frame by an annotation tool (for example, refer to the video at https: / / www.cvat.ai / post / 3d-point-cloud-annotation).
[0005] As such, inconsistent annotation has a negative impact on learning. Therefore, there is a need for a method that can ensure as consistent annotation as possible, improve the working speed of annotation, and efficiently increase the amount of training data.
Summary of the Invention
Problems to be Solved by the Invention
[0006] An object of the present invention is to solve the above-described problems.
[0007] Further, the present invention does not annotate all frames, but selects key frames from among the plurality of frames used to generate a lidar map, and another object is to perform annotation only on the key frames.
[0008] Also, another object of the present invention is to perform annotation while checking all of them by simultaneously displaying the lidar map and the camera image corresponding to each of the plurality of key frames.
[0009] Also, another object of the present invention is to prevent repeated annotation on the same object by accumulating and storing the annotation performed on each of the plurality of objects included in the generated lidar map.
[0010] Also, another object of the present invention is to obtain an annotated lidar map in which annotation has been completed for each of the plurality of objects included in the lidar map and be able to manage it with one map data.
Means for Solving the Problems
[0011] According to an embodiment of the present invention, in a method of annotating based on a rider map, (a) in a state where datasets related to corresponding travel routes are classified for each of a plurality of travel routes (the datasets include respective rider point cloud data corresponding to each of a plurality of rider frames acquired according to a predetermined standard while traveling on the corresponding travel route and respective camera image data corresponding to each of a plurality of camera frames) and stored in a database, when a specific dataset is acquired from among the plurality of datasets, a computing device generates a specific rider map and a specific keyframe trajectory using a plurality of specific rider point cloud data included in the specific dataset; and (b) the computing device assists in performing annotation on at least a part of specific keyframes among the plurality of keyframes included in the specific keyframe trajectory and stores this in the specific rider map; A method is provided, characterized by including.
[0012] In one example, in step (a), the computing device selects a part of the plurality of specific rider frames as the plurality of keyframes with reference to respective position information of each of the plurality of specific rider frames corresponding to each of the plurality of specific rider point cloud data, and generates a path of moving along the plurality of keyframes as the specific keyframe trajectory.
[0013] In one example, in the step (a), the computing device performs, for each of the plurality of specific lidar frames, at least one of: (i) a sub-process of determining whether a moving distance between a k-th point corresponding to the k-th specific lidar frame and a (k + 1)-th point corresponding to the (k + 1)-th specific lidar frame exceeds a preset threshold distance; and (ii) a sub-process of determining whether a pose change amount between a k-th pose corresponding to the k-th specific lidar frame and a (k + 1)-th pose corresponding to the (k + 1)-th specific lidar frame exceeds a preset threshold change amount, to select the plurality of key frames.
[0014] In one example, in the step (b), (b1) a step of the computing device assisting such that a t-th key frame is selected as the specific key frame among the plurality of key frames included in the specific key frame trajectory; (b2) if the t-th key frame is selected, a step of the computing device assisting to display a t-th lidar map and a t-th camera image corresponding to the position information of the t-th key frame using a display device; (b3) a step of the computing device assisting such that a t-th annotation for at least some of the plurality of objects included in the t-th lidar map is performed with reference to at least one of the t-th lidar map and the t-th camera image; and (b4) a step of the computing device converting the result information of the t-th annotation into a reference coordinate system of the t-th lidar map and then storing it in the specific lidar map.
[0015] In one example, in the step (b3), if the t-th object is selected from the plurality of objects, the computing device generates a t_1-annotation box having a predetermined range with respect to the t-th object, and supports the display device to display the t_1-annotation box, thereby supporting the execution of the t-th annotation for the t-th object.
[0016] In one example, the computing device supports a t_2-annotation box corresponding to the t_1-annotation box to be projected and appear on the area of the t-th camera image corresponding to the coordinate range of the t-th lidar map where the t_1-annotation box is located, and refers to the correlation between the t_1 shooting angle range in which the t-th lidar map is acquired and the t_2 shooting angle range in which the t-th camera image is acquired, so that the second form of the t_2-annotation box changes in conjunction with the first form of the t_1-annotation box.
[0017] In one example, if the t-th object is selected from the plurality of objects, the computing device supports the display device to further display a detailed area in which at least one of the planar portion, side portion, and front portion of the t-th object is enlarged and represented. The computing device generates at least one t_3-annotation box in which at least one of the planar portion, side portion, and front portion of the t_1-annotation box is enlarged and represented, and supports the display device to display the t_3-annotation box through the detailed area.
[0018] In one example, after the step (b4) is completed, if a (b5) t+1-th key frame is selected, the computing device converts the result information of the t-th annotation stored after being converted into the reference coordinate system of the t-th lidar map based on the reference coordinate system of the (t + 1)-th lidar map and applies it to the (t + 1)-th lidar map, and assists the display device to display the (t + 1)-th lidar map to which the result information of the t-th annotation is applied, and assists to display the result information of the t-th annotation on the (t + 1)-th camera image with reference to the matching relationship between the (t + 1)-th lidar map and the (t + 1)-th camera image; (b6) the computing device assists to perform a (t + 1)-th annotation on at least some of the plurality of objects included in the (t + 1)-th lidar map with reference to at least one of the (t + 1)-th lidar map and the (t + 1)-th camera image; and (b7) after the computing device converts the result information of the (t + 1)-th annotation into the reference coordinate system of the (t + 1)-th lidar map, stores it in the specific lidar map; It is characterized by including.
[0019] In one example, in the step (a), the computing device is characterized in that it applies a Lidar SLAM algorithm to the plurality of specific lidar point cloud data to generate the specific lidar map.
[0020] In one example, in the step (a), the position information of each of the plurality of specific lidar frames corresponding to each of the plurality of specific lidar point cloud data is interlocked, and the position information includes information regarding 6-DOF (Six degrees of freedom).
[0021] According to another embodiment of the present invention, in a computing device for annotating based on a rider map, it includes one or more memories for storing instructions; and one or more processors configured to execute the instructions. The processor: (I) in a state where a data set related to a corresponding travel route is classified and stored in a database for each of a plurality of travel routes (the data set includes respective rider point cloud data corresponding to each of a plurality of rider frames obtained according to a predetermined criterion while traveling on the corresponding travel route and respective camera image data corresponding to each of a plurality of camera frames), if a specific data set is obtained among the plurality of data sets, a process of generating a specific rider map and a specific key frame trajectory by using a plurality of specific rider point cloud data included in the specific data set, and (II) a process of assisting in performing annotation on at least a part of specific key frames among the plurality of key frames included in the specific key frame trajectory and storing this in the specific rider map is performed. A computing device is provided, which is characterized in that.
[0022] In one example, in the process (I), the processor selects a part of the plurality of specific rider frames as the plurality of key frames with reference to the respective position information of the plurality of specific rider frames corresponding to each of the plurality of specific rider point cloud data, and generates a path of moving along the plurality of key frames as the specific key frame trajectory.
[0023] In one example, in the (I) process, the processor performs at least one of the following sub-processes for each of the plurality of specific lidar frames: (i) a sub-process of determining whether a moving distance between a k-th point corresponding to the k-th specific lidar frame and a (k + 1)-th point corresponding to the (k + 1)-th specific lidar frame exceeds a preset threshold distance; and (ii) a sub-process of determining whether a pose change amount between a k-th pose corresponding to the k-th specific lidar frame and a (k + 1)-th pose corresponding to the (k + 1)-th specific lidar frame exceeds a preset threshold change amount, and selects the plurality of key frames accordingly.
[0024] In one example, in the (II) process, the processor performs the following processes: (II-1) a process of assisting in selecting a t-th key frame as the specific key frame among the plurality of key frames included in the specific key frame trajectory; (II-2) a process of assisting in displaying a t-th lidar map and a t-th camera image corresponding to the position information of the t-th key frame on a display device if the t-th key frame is selected; (II-3) a process of assisting in performing a t-th annotation on at least some of the plurality of objects included in the t-th lidar map with reference to at least one of the t-th lidar map and the t-th camera image; and (II-4) a process of storing the result information of the t-th annotation in the specific lidar map after converting it into the reference coordinate system of the t-th lidar map.
[0025] In one example, in the (II-3) process, if a t-th object is selected among the plurality of objects, the processor generates a t_1-annotation box having a predetermined range with reference to the t-th object, and assists in performing the t-th annotation on the t-th object by assisting in displaying the t_1-annotation box on the display device.
[0026] In one example, the processor supports projecting a t_2 annotation box corresponding to the t_1 annotation box onto an area of the t camera image corresponding to the coordinate range of the t lidar map in which the t_1 annotation box is located, and changes the second shape of the t_2 annotation box in conjunction with the first shape of the t_1 annotation box by referring to the correlation between the t_1 shooting angle range in which the t lidar map was acquired and the t_2 shooting angle range in which the t camera image was acquired.
[0027] In one example, when a t-th object is selected from the plurality of objects, the processor supports the display device to further display a detailed area in which at least one of the planar portion, side portion, and front portion of the t-th object is enlarged and represented, generate at least one t_3-th annotation box in which at least one of the planar portion, side portion, and front portion of the t_1 annotation box is enlarged and represented, and support the display device to display the t_3 annotation box through the detailed area.
[0028] In one example, after the (II-4) process is completed, if the (II-5) (t + 1)-th key frame is selected, the processor converts the result information of the t-th annotation stored after being converted into the reference coordinate system of the t-th lidar map, applies it to the (t + 1)-th lidar map based on the reference coordinate system of the (t + 1)-th lidar map, assists the display device to display the (t + 1)-th lidar map to which the result information of the t-th annotation is applied, and refers to the matching relationship between the (t + 1)-th lidar map and the (t + 1)-th camera image to assist in displaying the result information of the t-th annotation on the (t + 1)-th camera image; a process of assisting in performing a (t + 1)-th annotation on at least a part of a plurality of objects included in the (t + 1)-th lidar map with reference to at least one of the (t + 1)-th lidar map and the (t + 1)-th camera image; and a process of storing the result information of the (t + 1)-th annotation in the specific lidar map after converting it into the reference coordinate system of the (t + 1)-th lidar map.
[0029] In one example, in the (I) process, the processor is characterized by applying a Lidar SLAM algorithm to the plurality of specific lidar point cloud data to generate the specific lidar map.
[0030] In one example, in the (I) process, the position information of each of the plurality of specific lidar frames corresponding to each of the plurality of specific lidar point cloud data is linked, and the position information includes information related to 6-DOF (Six degrees of freedom).
Advantages of the Invention
[0031] The present invention does not annotate all frames. Instead, by selecting key frames from among a plurality of frames used to generate a lidar map and performing annotation only on the key frames, it is possible to improve work efficiency.
[0032] In addition, the present invention makes it possible to reduce errors and improve accuracy by simultaneously displaying the lidar map and camera image corresponding to each of the plurality of key frames so that annotation is performed while checking all of them.
[0033] In addition, the present invention has the effect of preventing repetitive annotation of the same object by accumulating and storing the annotation performed on each of the plurality of objects included in the generated lidar map.
[0034] In addition, the present invention has the effect of obtaining an annotated lidar map in which annotation has been completed for each of the plurality of objects included in the lidar map and being able to manage it with a single map data.
Brief Description of the Drawings
[0035]
Figure 1
Figure 2
Figure 3a
Figure 3b
Figure 4
Figure 5a
Figure 5b
Figure 5c
Figure 5d
Figure 6a
Figure 6b
Mode for Carrying Out the Invention
[0036] The detailed description of the present invention, which will be described later, refers to the accompanying drawings that illustrate specific embodiments in which the present invention can be implemented as examples. These embodiments are described in detail so that those skilled in the art can sufficiently implement the present invention. It should be understood that although various embodiments of the present invention are different from each other, they do not necessarily have to be mutually exclusive. For example, the specific shapes, structures, and characteristics described herein can be embodied in other embodiments without departing from the spirit and scope of the present invention in relation to one embodiment. Also, it should be understood that the position or arrangement of individual components within each disclosed embodiment can be changed without departing from the spirit and scope of the present invention. Therefore, the following detailed description is not intended to be taken in a limiting sense, and the scope of the present invention is limited only by the appended claims together with all scopes equivalent to what the claims claim when appropriately described. Similar reference numerals in the drawings are the same or refer to similar functions throughout various aspects.
[0037] Hereinafter, for the convenience of those having ordinary knowledge in the technical field to which the present invention pertains to easily implement the present invention, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings.
[0038] FIG. 1 is a drawing showing a simplified configuration of a computing device that performs a method of annotating based on a lidar map according to an embodiment of the present invention.
[0039] Referring to FIG. 1, the computing device 100 can include a memory 110 that stores instructions for performing a method of annotating based on a lidar map, and a processor 120 that preprocesses for annotating based on the lidar map corresponding to the instructions stored in the memory. At this time, the computing device 100 can include a PC (Personal Computer), a mobile computer, and the like.
[0040] Specifically, the computing device 100 may achieve the desired system performance by using a combination of a typical computing device (for example, a device that can include a computer processor, memory, storage, input devices, and output devices, and other components of existing computing devices; an electronic communication device such as a router or a switch; an electronic information storage system such as a network-attached storage (NAS) and a storage area network (SAN)) and computer software (that is, instructions for causing the computing device to function in a specific manner).
[0041] In addition, the processor of the computing device may include hardware configurations such as an MPU (Micro Processing Unit) or a CPU (Central Processing Unit), a cache memory, and a data bus. The computing device may further include an operation system and a software configuration of an application for performing a specific purpose.
[0042] However, it does not exclude the case where the computing device includes an integrated processor in which the medium, the processor, and the memory for implementing the present invention are integrated.
[0043] And the computing device 100 can be linked with a database 900 that contains information used for annotating based on a rider map. Here, the database 900 can include at least one type of storage medium among flash memory type, hard disk type, multimedia card micro type, card type memory (e.g., SD or XD memory), random access memory (RAM), static random access memory (SRAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), programmable read only memory (PROM), magnetic memory, magnetic disk, optical disk, and is not limited thereto, and can include all media capable of storing data. Also, the database 900 can be installed separately from the computing device 100, or, differently, can be installed inside the computing device 100 to transmit data or record the received data, and, differently from what is illustrated, can also be embodied separately into two or more, which may vary depending on the implementation conditions of the invention.
[0044] Also, the computing device 100 can be linked with a display device (not shown) that displays at least a part of the process in which annotation is performed based on a rider map.
[0045] Using the computing device 100 according to an embodiment of the present invention configured as described above, a method of annotating based on a rider map will be described with reference to FIG. 2 as follows.
[0046] FIG. 2 is a drawing for explaining a simplified procedure of a method of annotating based on a rider map according to an embodiment of the present invention.
[0047] First, in a state where data sets are classified for each of a plurality of travel routes and stored in the database 900, if a specific data set is selected from among the plurality of data sets, the computing device 100 can generate a specific lidar map and a specific key frame trajectory by using a plurality of specific lidar point cloud data included in the specific data set (S100). Then, the computing device 100 can assist in performing annotation on at least a part of specific key frames among the plurality of key frames included in the specific key frame trajectory, and store this in the specific lidar map (S200). Hereinafter, this will be described in more detail.
[0048] Generation of Lidar Map and Key Frame Trajectory Describing the step S100, which is a method for generating a lidar map and a specific key frame trajectory using the computing device 100 according to an embodiment of the present invention, is as follows.
[0049] First, the computing device 100 can acquire a specific data set from among the plurality of data sets stored in the database 900.
[0050] At this time, in the database 900, data sets related to the corresponding travel route can be classified and stored for each of a plurality of travel routes, and each data set can include respective lidar point cloud data corresponding to each of a plurality of lidar frames obtained according to a predetermined standard while an automobile equipped with a lidar and a camera travels on the corresponding travel route, and respective camera image data corresponding to each of a plurality of camera frames.
[0051] Here, the predetermined standard may be to take pictures at each time point of a preset period, or may be to take pictures at each preset moving distance, but is not limited thereto.
[0052] Also, since the multiple lidar point cloud data acquired by the lidar are 3D data and the multiple camera image data acquired by the camera are 2D data, the lidar and the camera may be calibrated to match their coordinate systems.
[0053] And, the computing device 100 can generate a specific lidar map and a specific keyframe trajectory by using the multiple specific lidar point cloud data included in a specific dataset.
[0054] At this time, the computing device 100 can generate a specific lidar map by applying a Lidar SLAM (Simultaneous Localization and Mapping) algorithm to the multiple specific lidar point cloud data.
[0055] Here, the Lidar SLAM algorithm can include, but is not limited to, the ICP (Iterative Closest Point) algorithm, the HDL GRAPH SLAM algorithm, etc.
[0056] Also, the computing device 100 can further improve the quality of a specific lidar map by generating the specific lidar map by using at least some of the multiple GPS data and at least some of the multiple IMU data acquired by at least some of the GPS sensors and IMU sensors mounted on the vehicle.
[0057] Here, when generating a lidar map by using GPS sensor data, there is an advantage that the previously acquired data can be reused even if the time and date of lidar shooting and camera shooting are different and the surrounding environment changes.
[0058] Further, the computing device 100 can generate a specific keyframe trajectory as a path that moves along a plurality of keyframes by referring to the position information of each of a plurality of specific lidar frames corresponding to each of the plurality of specific lidar point cloud data, and selecting at least a part of the plurality of specific lidar frames as the plurality of keyframes.
[0059] At this time, the computing device 100 can perform at least one of the following sub-processes for each of the plurality of specific lidar frames: (i) a sub-process of determining whether the moving distance between the k-th point corresponding to the k-th specific lidar frame and the (k + 1)-th point corresponding to the (k + 1)-th specific lidar frame exceeds a preset threshold distance, and (ii) a sub-process of determining whether the amount of pose change between the k-th pose corresponding to the k-th specific lidar frame and the (k + 1)-th pose corresponding to the (k + 1)-th specific lidar frame exceeds a preset threshold change amount, to select the plurality of keyframes. For example, when selecting a plurality of keyframes, it means selecting so that the positions where the plurality of keyframes are obtained are not too close to each other, and the poses of the positions where the plurality of keyframes are obtained are not too similar to each other.
[0060] Further, the position information of each of the plurality of specific lidar frames is interlocked, and the position information can include information related to 6-DOF (Six degrees of freedom) (X, Y, Z, Roll, Pitch, Yaw) information.
[0061] For example, the position information may be information regarding the location where the vehicle is located and the pose of the vehicle on a specific lidar map, or information regarding the location where the lidar is located and the pose of the lidar on a specific lidar map, but is not limited thereto. Here, the location may mean Position in 6-DOF, and the pose may mean Orientation in 6-DOF, but is not limited thereto.
[0062] FIG. 3a and FIG. 3b are drawings showing an example of a generated specific lidar map and a specific keyframe trajectory.
[0063] As can be seen in FIGS. 3a and 3b, the computing device 100 can assist in representing and displaying specific keyframe trajectories 315 and 325 in specific lidar maps 310 and 320, but is not limited thereto.
[0064] Annotation FIG. 4 is a drawing for explaining a simplified procedure of the S200 step, which is a method of annotating at least some of a plurality of keyframes included in a specific keyframe trajectory using the computing device 100 according to an embodiment of the present invention and storing the annotation in a specific lidar map.
[0065] Referring to FIG. 4, the computing device 100 can assist in selecting a t-th keyframe, which is one of a plurality of keyframes included in a specific keyframe trajectory, as a specific keyframe (S210).
[0066] Here, t is an integer greater than or equal to 1 and less than or equal to n, and n may be the number of a plurality of keyframes.
[0067] Then, if the t-th keyframe is selected, the computing device 100 can assist in displaying the t-th lidar map information and the t-th camera image data corresponding to the position information of the t-th keyframe using a display device (S220).
[0068] Next, the computing device 100 can assist in performing a t-th annotation on at least some of the plurality of lidar objects included in the t-th lidar map by referring to at least one of the t-th lidar map and the t-th camera image (S230).
[0069] At this time, if the t-th object is selected from among the plurality of objects, the computing device 100 can generate a t_1-th annotation box having a predetermined range based on the t-th object, and assist in displaying the t_1-th annotation box on the display device, thereby assisting in performing the t-th annotation on the t-th object.
[0070] For example, when the t-th object is a 3D object having the same volume as a vehicle, the t-th annotation can be assisted in being performed by generating and displaying a cuboid, which is a 3D box that can include the corresponding object, as the t_1-th annotation box. When the t-th object is a 2D object or area having only width, such as a parking slot, the t-th annotation can be assisted in being performed by generating and displaying a bounding box, which is a 2D box that can include the corresponding object or area, as the t_1-th annotation box. However, the present invention is not limited thereto. For reference, the form of the bounding box referred to in this specification can be understood as a concept including not only a rectangle but also a quadrangle in which the angle of the rectangle is deformed.
[0071] Here, the computing device 100 can cause the process information and result information of the t-th annotation on the t-th object included in the t-th lidar map to be projected and appear on the t-th camera image.
[0072] Specifically, the computing device 100 can support the t_2th annotation box corresponding to the t_1th annotation box being projected and displayed on the area of the tth camera image corresponding to the coordinate range of the tth lidar map in which the t_1th annotation box is located.
[0073] Here, the computing device 100 can refer to the correlation between the t_1 shooting angle range in which the t lidar map was acquired and the t_2 shooting angle range in which the t camera image was acquired, and change the second shape of the t_2 annotation box in conjunction with the first shape of the t_1 annotation box.
[0074] In addition, when the t-th object is selected from among the plurality of objects, the computing device 100 may further support displaying a detailed area in which at least one of the plane portion, side portion, and front portion of the t-th object is enlarged using the display device.
[0075] Here, the computing device 100 generates at least one t_3 annotation box in which at least one of the planar portion, side portion, and front portion of the t_1 annotation box is enlarged and represented, and can support the display device to display the t_3 annotation box through a detail area.
[0076] Next, the computing device 100 may transform the t-th annotation result information into the reference coordinate system of the t-th lidar map and then store it in a particular lidar map (S240).
[0077] Here, the specific rider map may include first through nth rider maps corresponding to the position information of the first through nth key frames, respectively.
[0078] Therefore, the meaning of storing in a specific rider map at the S240 stage may mean storing in the specific rider map itself, or may mean storing in the t-th rider map included in the specific rider map, but is not limited thereto.
[0079] On the other hand, after the S240 stage is completed, the computing device 100 can assist in selecting the (t + 1)-th keyframe, which is one of the plurality of keyframes included in the specific keyframe trajectory, as the specific keyframe.
[0080] Then, if the (t + 1)-th keyframe is selected, the computing device 100 converts the result information of the t-th annotation stored after being converted into the reference coordinate system of the t-th rider map based on the reference coordinate system of the (t + 1)-th rider map and applies it to the (t + 1)-th rider map, and assists in displaying the (t + 1)-th rider map to which the result information of the t-th annotation is applied using the display device, and can assist in displaying the result information of the t-th annotation on the (t + 1)-th camera image with reference to the matching relationship between the (t + 1)-th rider map and the (t + 1)-th camera image.
[0081] Next, the computing device 100 can assist in performing the (t + 1)-th annotation on at least some of the plurality of objects included in the (t + 1)-th rider map by referring to at least one of the displayed (t + 1)-th rider map and the (t + 1)-th camera image.
[0082] Here, the (t + 1)-th object to be the target of the (t + 1)-th annotation may be a new object not included in the target of the first annotation to the t-th annotation performed before the (t + 1)-th annotation, but is not limited thereto.
[0083] Next, after converting the result information of the (t + 1)-th annotation into the reference coordinate system of the (t + 1)-th lidar map, the computing device 100 can store it in a specific lidar map.
[0084] In this way, the computing device 100 supports the execution of the first annotation, the second annotation, …, the n-th annotation for each of the first key frame, the second key frame, …, the n-th key frame, accumulates the result information of the executed annotations, and stores it in a specific lidar map, thereby obtaining an annotated lidar map in which the annotation for a plurality of objects included in the specific lidar map is completed, and it is characterized in that it can be managed with one map data.
[0085] As an example, a method of annotating with reference to FIGS. 5a to 5d will be described. Assuming that the t-th key frame is selected, the computing device 100 can assist in displaying the t-th lidar map 510, the t-th camera images 520 and 530 corresponding to the position information of the t-th key frame as shown in FIG. 5a. Then, as shown in FIG. 5b, if a specific automobile 511, which is the t-th object among a plurality of objects included in the t-th lidar map 510, is selected, a cuboid 512, which is a 3D box, is generated as the t_1-th annotation box having a predetermined range based on the specific automobile 511, and the computing device 100 can assist in displaying the cuboid 512. Here, it can be seen that a predetermined box 521 corresponding to the cuboid 512 is projected and appears as the t_2-th annotation box on the area of the t-th camera image 520 corresponding to the coordinate range of the t-th lidar map 510 where the cuboid 512 is located. For reference, in FIG. 5b, the predetermined box 521 is shown in an opaque manner, but it is not limited thereto. Also, the computing device 100 assists in further displaying a detailed area 540 in which the planar portion, the side portion, and the front portion of the specific automobile 511 are each enlarged, and generates a bounding box 541 in which the planar portion, the side portion, and the front portion of the cuboid 512 are each enlarged as the t_3-th annotation box, and by assisting in displaying the bounding box 541 directionally through the detailed area 540, more accurate annotation can be performed. Next, as shown in FIG. 5c, if the cuboid 512 generated for the specific automobile 511 is fitted to the specific automobile 511 and the t-th annotation is completed, the computing device 100 can convert the result information of the t-th annotation into the reference coordinate system of the t-th lidar map 510 and then store it.Here, the cuboid 512 that has been fitted in the t-th lidar map 510 is represented as a 3D box, and it can be seen that a predetermined box 521 corresponding to the cuboid 512 is also represented as a 3D box in the t-th camera image 520. Subsequently, assuming that the (t + 1)-th keyframe is selected, the computing device 100, as shown in FIG. 5d, converts the result information of the t-th annotation based on the reference coordinate system of the (t + 1)-th lidar map 550 and applies it to the (t + 1)-th lidar map 550, so that the cuboid 512 fitted to a specific vehicle 511, which is the result information of the t-th annotation, is displayed on the (t + 1)-th lidar map 550, and a predetermined box 521 corresponding to the cuboid 512 can be displayed on the (t + 1)-th camera image 560.
[0086] As another example, referring to FIGS. 6a and 6b, as shown in FIG. 6a, assuming that as a result of the 7th annotation performed on the parking slot at the 7th keyframe, a 7_2 annotation box displayed in gray is stored in the 7th key camera image, as shown in FIG. 6b, it can be seen that a 7_2 annotation box displayed in gray, which is the result of the 7th annotation, is also displayed in the 51st camera image corresponding to the position information of the 51st keyframe.
[0087] As described above, the embodiments according to the present invention can be embodied in the form of program instruction words that can be executed through various computer components and can be stored in a computer-readable recording medium. The computer-readable recording medium can include program instruction words, data files, data structures, etc. alone or in combination. The program instruction words stored in the computer-readable recording medium can be those specially designed and configured for the present invention or those known and usable by those skilled in the field of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program instruction words such as ROMs, RAMs, and flash memories. Examples of program instruction words include not only machine language codes such as those created by compilers but also high-level language codes that can be executed by a computer using an interpreter or the like. The hardware device can be configured to operate as one or more software modules to perform the processing according to the present invention, and vice versa.
[0088] As described above, the present invention has been described by way of specific examples and drawings limited to specific matters such as specific components. However, this is only provided to assist in a more general understanding of the present invention, and the present invention is not limited to the above embodiments. Those having ordinary knowledge in the technical field to which the present invention pertains can make various modifications and variations from such descriptions.
[0089] Therefore, the idea of the present invention should not be defined only by the embodiments described above. It can be said that not only the claims described below but also all those equivalently or equivalently modified belong to the scope of the idea of the present invention.
Claims
1. In a method of annotating based on a rider map, (a) In a state where datasets corresponding to respective driving routes (the datasets include respective lidar point cloud data corresponding to respective ones of a plurality of lidar frames acquired according to a predetermined standard while driving on the corresponding driving route and respective camera image data corresponding to respective ones of a plurality of camera frames) are classified and stored in a database for each of a plurality of driving routes, when a specific dataset is acquired from among the plurality of datasets, a computing device generates a specific lidar map and a specific keyframe trajectory using a plurality of specific lidar point cloud data included in the specific dataset; and (b) The computing device performs annotation on at least some of a plurality of keyframes included in the specific keyframe trajectory, and stores the specifically annotated keyframe in the specific lidar map; including in the step (a), the computing device refers to respective position information of a plurality of specific lidar frames corresponding to respective ones of the plurality of specific lidar point cloud data, selects some of the plurality of specific lidar frames as the plurality of keyframes, generates a path of moving along the plurality of keyframes as the specific keyframe trajectory, and displays the specific keyframe trajectory on the specific lidar map, the method being characterized by this.
2. In the step (a), the computing device performs at least one of a sub-process of determining whether a moving distance between a k-th point corresponding to a k-th specific lidar frame and a (k + 1)-th point corresponding to a (k + 1)-th specific lidar frame exceeds a preset threshold distance for each of the plurality of specific lidar frames, and a sub-process of determining whether a pose change amount between a k-th pose corresponding to the k-th specific lidar frame and a (k + 1)-th pose corresponding to the (k + 1)-th specific lidar frame exceeds a preset threshold change amount, to select the plurality of keyframes, the method according to claim 1, characterized by this.
3. In the step (b), (b1) a step of causing the computing device to select the t-th keyframe as the specific keyframe among the plurality of keyframes included in the specific keyframe trajectory; (b2) a step of assisting, if the t-th keyframe is selected, the computing device to display, using a display device, a t-th lidar map and a t-th camera image corresponding to the position information of the t-th keyframe; (b3) a step of causing the computing device to perform a t-th annotation on at least some of the plurality of objects included in the t-th lidar map with reference to at least one of the t-th lidar map and the t-th camera image; and (b4) a step of causing the computing device to convert the result information of the t-th annotation into a reference coordinate system of the t-th lidar map and then store it in the specific lidar map; The method according to claim 1, characterized by including the above.
4. In the step (b3), If a t-th object is selected among the plurality of objects, the computing device generates a t_1-annotation box having a predetermined range with reference to the t-th object, and assists the display device to display the t_1-annotation box, thereby assisting the performance of the t-th annotation on the t-th object. The method according to claim 3, characterized by this.
5. The computing device assists such that a t_2-annotation box corresponding to the t_1-annotation box is projected and appears on a region of the t-th camera image corresponding to a coordinate range of the t-th lidar map where the t_1-annotation box is located, With reference to a correlation relationship between a t_1 shooting angle range of a predetermined lidar device that acquired the t-th lidar map and a t_2 shooting angle range of a predetermined camera device that acquired the t-th camera image, a second form that is the form of the t_2-annotation box is changed in conjunction with a first form that is the form of the t_1-annotation box. The method according to claim 4, characterized by this.
6. If the t-th object is selected from the plurality of objects, the computing device assists the display device to further display a detailed area in which at least one of a planar portion, a side surface portion, and a front surface portion of the t-th object is enlarged and represented. The method according to claim 4, wherein the computing device generates at least one t_3 annotation box in which at least one of a planar portion, a side surface portion, and a front surface portion of the t_1 annotation box is enlarged and represented, and assists the display device to display the t_3 annotation box through the detailed area.
7. After the step (b4) is completed, (b5) If the (t + 1)-th key frame is selected, the computing device converts the result information of the t-th annotation stored after being converted into the reference coordinate system of the t-th lidar map with reference to the reference coordinate system of the (t + 1)-th lidar map and applies it to the (t + 1)-th lidar map, and assists the display device to display the (t + 1)-th lidar map to which the result information of the t-th annotation is applied, and refers to the matching relationship between the (t + 1)-th lidar map and the (t + 1)-th camera image to assist in displaying the result information of the t-th annotation on the (t + 1)-th camera image; (b6) The computing device refers to at least one of the (t + 1)-th lidar map and the (t + 1)-th camera image to perform a (t + 1)-th annotation on at least a part of a plurality of objects included in the (t + 1)-th lidar map; and (b7) After the computing device converts the result information of the (t + 1)-th annotation into the reference coordinate system of the (t + 1)-th lidar map, it stores it in the specific lidar map. The method according to claim 3, comprising the above steps.
8. In the step (a), The method according to claim 1, wherein the computing device applies a Lidar SLAM algorithm to the plurality of specific lidar point cloud data to generate the specific lidar map.
9. In the step (a), For each of the plurality of specific lidar point cloud data, the position information of each of the plurality of specific lidar frames is interlocked, and the position information includes information regarding 6-DOF (Six degrees of freedom). The method according to claim 1, characterized in that.
10. In a computing device for annotating based on a lidar map, One or more memories for storing instructions; and Including one or more processors configured to execute the instructions, The processor, (I) in a state where a dataset related to a corresponding travel route for each of a plurality of travel routes (the dataset includes respective lidar point cloud data corresponding to each of a plurality of lidar frames obtained according to a predetermined standard while traveling on the corresponding travel route and respective camera image data corresponding to each of a plurality of camera frames) is classified and stored in a database, if a specific dataset is obtained from among the plurality of datasets, a process of generating a specific lidar map and a specific keyframe trajectory using the plurality of specific lidar point cloud data included in the specific dataset, and (II) performing annotation on at least a part of the specific keyframes included in the specific keyframe trajectory, and storing the specifically keyframe on which the annotation has been performed in the specific lidar map. The processor is In the process of (I), Referring to the respective position information of each of the plurality of specific lidar frames corresponding to each of the plurality of specific lidar point cloud data, selecting a part of the plurality of specific lidar frames as the plurality of keyframes, generating a path moving along the plurality of keyframes as the specific keyframe trajectory, and displaying the specific keyframe trajectory on the specific lidar map. A computing device, characterized in that.
11. The processor is In the process of (I), For each of the plurality of specific lidar frames, at least one of the following sub - processes is performed to select the plurality of key frames: (i) a sub - process of determining whether the moving distance between the k - th point corresponding to the k - th specific lidar frame and the (k + 1)-th point corresponding to the (k + 1)-th specific lidar frame exceeds a preset threshold distance; and (ii) a sub - process of determining whether the amount of pose change between the k - th pose corresponding to the k - th specific lidar frame and the (k + 1)-th pose corresponding to the (k + 1)-th specific lidar frame exceeds a preset threshold amount of change. The computing device according to claim 10, characterized in that.
12. The processor is, In the process (II), (II - 1) A process of enabling the t - th key frame to be selected as the specific key frame among the plurality of key frames included in the specific key - frame trajectory; (II - 2) If the t - th key frame is selected, a process of assisting to display the t - th lidar map and the t - th camera image corresponding to the position information of the t - th key frame using a display device; (II - 3) A process of enabling the t - th annotation to be performed on at least some of the plurality of objects included in the t - th lidar map by referring to at least one of the t - th lidar map and the t - th camera image; and (II - 4) After converting the result information of the t - th annotation into the reference coordinate system of the t - th lidar map, a process of storing it in the specific lidar map. The computing device according to claim 10, characterized in that.
13. The processor is, In the process (II - 3), If the t - th object is selected among the plurality of objects, a t_1 annotation box having a predetermined range is generated based on the t - th object, and the t - th annotation for the t - th object is assisted to be performed by assisting to display the t_1 annotation box using the display device. The computing device according to claim 12, characterized in that.
14. The processor is, Supporting the t_2th annotation box corresponding to the t_1 annotation box to be projected and displayed on an area of the t camera image corresponding to the coordinate range of the t lidar map in which the t_1 annotation box is located; 14. The computing device of claim 13, wherein a second form, which is the form of the t_2 annotation box, is changed in conjunction with a first form, which is the form of the t_1 annotation box, by referring to a correlation between a t_1th shooting angle range of a predetermined LIDAR device that acquired the t lidar map and a t_2th shooting angle range of a predetermined camera device that acquired the t camera image.
15. The processor, When a t-th object is selected from the plurality of objects, the display device is supported to further display a detailed area in which at least one of a plane portion, a side portion, and a front portion of the t-th object is enlarged and represented; The computing device of claim 13, further comprising: generating at least one t_3rd annotation box by enlarging and representing at least one of a plane portion, a side portion, and a front portion of the t_1 annotation box; and supporting the display device to display the t_3rd annotation box through the detail area.
16. The processor, After the process (II-4) is completed, If the (t + 1)-th key frame is selected, the result information of the t-th annotation that has been converted and stored in the reference coordinate system of the t-th lidar map is converted based on the reference coordinate system of the (t + 1)-th lidar map and applied to the (t + 1)-th lidar map, and the display device is used to assist in displaying the (t + 1)-th lidar map to which the result information of the t-th annotation is applied, and the result information of the t-th annotation is assisted to be displayed on the (t + 1)-th camera image by referring to the matching relationship between the (t + 1)-th lidar map and the (t + 1)-th camera image, (II - 6) a process of performing a (t + 1)-th annotation on at least a part of a plurality of objects included in the (t + 1)-th lidar map by referring to at least one of the (t + 1)-th lidar map and the (t + 1)-th camera image, and (II - 7) a process of storing the result information of the (t + 1)-th annotation in the specific lidar map after converting it into the reference coordinate system of the (t + 1)-th lidar map. The computing device according to claim 12, characterized in that the processes are performed.
17. The processor is In the process of (I), The computing device according to claim 10, characterized in that in the process of (I), a Lidar SLAM algorithm is applied to the plurality of specific lidar point cloud data to generate the specific lidar map.
18. In the process of (I), The position information of each of the plurality of specific lidar frames corresponding to each of the plurality of specific lidar point cloud data is interlocked, and the position information includes information regarding 6 - DOF (Six degrees of freedom). The computing device according to claim 10, characterized in that it is included.
Citation Information
Patent Citations
Video-based localization and mapping method and system
JP2020516853A
Ground truth data generation for deep neural network perception in autonomous driving applications
JP2022132075A
Method and system for annotation of sensor data
JP2023548749A
System and method for vehicle-to-everything (V2X) collaborative perception
US20230059897A1
Teacher data collection device
WO2019116423A1