Annotation support device, method, and program using images and point clouds

The annotation support device integrates image and point cloud data with AI-assisted quality assurance to address visibility and spatial resolution challenges, enhancing annotation precision and efficiency for autonomous driving applications.

JP7792176B1Active Publication Date: 2025-12-25FASTLABEL CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025138720
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-12-25
Estimated Expiration
2045-08-22

AI Technical Summary

Technical Problem

Existing annotation tools struggle with efficiently integrating image and point cloud data for high-precision annotation tasks, particularly in autonomous driving, due to visibility issues, spatial resolution challenges, and the lack of mechanisms for intuitive and efficient operation in 3D space, leading to inconsistent labels and inefficient work processes.

Method used

An annotation support device and method that integrates image and point cloud data through a recording unit, output units, and execution units to facilitate display, correction, and recording, utilizing AI models for quality assurance, allowing flexible switching between single and overall overview, and supporting collaboration among multiple users.

Benefits of technology

Enables efficient, high-precision annotation work by ensuring consistency between image and point cloud data, improving annotation quality and efficiency, and supporting flexible collaboration for large-scale data generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007792176000001_ABST
    Figure 0007792176000001_ABST
Patent Text Reader

Abstract

Annotation work using point cloud data and image data has presented challenges in terms of accuracy, workload, and quality assurance due to the difficulty of recognizing objects, data variability, and the complexity of time-series data. In particular, creating large-scale, highly accurate training data for autonomous driving and drones has been difficult with conventional tools. [Solution] This invention relates to an apparatus, method, and program that supports annotation by linking point clouds and images. Annotation class information, calibration information, point cloud / image data, and annotation information are recorded in the recording unit, and a work screen is provided by the first and second output units. Furthermore, the system achieves highly accurate, efficient, and high-quality annotation work through point cloud / image synchronization processing, automatic annotation, quality confirmation using VLM (AutoQA), and a multi-user compatible logging mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a technique for annotating images and point cloud data when creating datasets for machine learning, and more particularly to an apparatus, method, and program that support annotation work based on the association of camera images with three-dimensional point clouds. [Background technology]

[0002] In recent years, with the development of sensing technology in autonomous driving, robotics, and smart cities, the use of 3D point cloud data from LiDAR etc. has increased in addition to 2D image data from cameras. As a result, there is a growing demand for annotation that assigns object positions and attributes to both image data and point cloud data.

[0003] Conventionally, annotation work on point cloud data has been a heavy workload and difficult to ensure accuracy due to low visibility and difficulty in distinguishing objects. For this reason, methods that link camera images with point cloud data and reflect annotation information added on images on the point cloud have attracted attention.

[0004] However, in conventional technologies, accurately mapping the rectangle information and attribute information attached to camera images onto point clouds often requires tedious manual work and advanced skills, resulting in issues with annotation quality and efficiency. Furthermore, generating annotations that accurately reflect depth information based on point clouds requires the appropriate handling of a variety of information, such as calibration information and shape characteristics for each object class. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] International Publication No. 2022 / 185363 [Patent Document 2] Japanese Patent Publication No. 2025-017225 [Non-patent literature]

[0006] [Non-Patent Document 1] Open source annotation tool: Cloud Compare, URL: https: / / www.danielgm.net / cc / (accessed August 4, 2025) [Non-patent document 2] Open source annotation tool CVAT, URL: https: / / www.cvat.ai / (Retrieved August 4, 2025) [Non-patent document 3] Commercial annotation tool Annofab URL https: / / annofab.com / (Retrieved August 4, 2025) Summary of the Invention [Problem to be solved by the invention]

[0007] In recent years, the importance of developing AI models using point cloud data has rapidly increased in fields such as autonomous driving, drones, and infrastructure inspection. While these applications require the mass generation of highly accurate training data, annotation work on point clouds has issues with visibility and spatial resolution, making it extremely inefficient compared to traditional image annotation. Furthermore, traditional methods of annotating images and point clouds separately make it difficult to ensure consistency between the two, leading to frequent rework and inconsistent labels.

[0008] Previous tools included the lightweight and multi-functional Cloud Compare, the open-source CVAT, and the commercial service Annofab, but these were primarily focused on point cloud-based tasks and were not adequately able to create integrated training data with camera images.In particular, for the high-precision annotation required for autonomous driving, it is difficult to identify objects using point clouds alone due to factors such as point cloud density, noise, and the influence of obstructions, resulting in significant issues such as missing annotations in point clouds and insufficient accuracy.

[0009] In addition, attempts have been made to introduce automatic assistance using AI to support annotation using point clouds and images, but due to limitations in the accuracy of point cloud data itself, conventional automatic inference models have not been able to ensure sufficient reliability in terms of accuracy. Furthermore, while the generation of large-scale training data requires collaboration among multiple people, there has been a lack of mechanisms that allow intuitive and efficient operation in 3D space.

[0010] In addition, because each operation in annotation work directly affects cumulative work efficiency when dealing with huge amounts of data, it is essential to have a system that allows users to flexibly switch between "checking a single annotation" and "overall overview" during review work, as well as a system that provides feedback to users on important information such as manual corrections and reference frame history.

[0011] The present invention aims to solve the above problems by providing an annotation mechanism that utilizes both images and point clouds in an integrated manner and consistently supports display, operation, correction, and recording. [Means for solving the problem]

[0012] An annotation support device according to a first embodiment of the present invention includes: a recording unit that records setting information including annotation class information, calibration information, point cloud data, and image data, and annotation information; a first output unit that refers to the recording unit and displays the point cloud data to enable annotation work; a second output unit that refers to the recording unit and displays a corresponding calibration image to support an annotation task; a first execution unit that uses the calibration information to receive a region designation from a user in the point cloud data displayed by the first output unit, and adds a target region to the image data based on the region designation; a second execution unit that executes various screen control processes related to the display of the first output unit and the second output unit; a third execution unit that refers to the recording unit and the first execution unit, and executes an automatic quality assurance process that supports review and quality confirmation of annotation results using an AI model including one or more of an object detection model, a region detection model, a visual language model, and a large-scale language model; and a fourth execution unit that executes annotation processing for adding a target area to the image data and correction processing for a rectangle on the point cloud data; The present invention is characterized by comprising:

[0013] An annotation support device according to a second embodiment of the present invention is the annotation support device according to the first embodiment, wherein the recording unit records data at a plurality of points in time, including time-series point cloud data and image data; The first execution unit performs a process of synchronizing the point cloud data and the image data based on the time series.

[0014] An annotation support device according to a third embodiment of the present invention is an annotation support device according to the first or second embodiment, characterized in that the first output unit provides a user interface that displays an annotation target area based on point cloud data and accepts correction operations in order to support annotation work on an object.

[0015] An annotation support device according to a fourth embodiment of the present invention is an annotation support device according to either the first or third embodiment, characterized in that the second output unit switches between or simultaneously displays calibration images and annotation data corresponding to multiple image displays according to the number of camera sensors, and provides a user interface that allows the user to visually confirm consistency.

[0016] An annotation support device according to a fifth embodiment of the present invention is an annotation support device according to any of the first to fourth embodiments, characterized in that the third execution unit executes a review function that evaluates omissions or consistency of the object based on the recorded annotation information.

[0017] An annotation support device according to a sixth embodiment of the present invention is the annotation support device according to any one of the first to fifth embodiments, wherein the fourth execution unit, when an object exists in the image data, specifies an area including at least one of a rectangle and a point by user input, or specifies an area by inputting annotation data, or extracts the object area as a rectangle using a pre-trained model, associates the extracted area with point cloud data, and displays it on the output unit or records it on the recording unit; The fourth execution unit is characterized by realizing highly accurate initial annotation by automatically correcting the orientation and size of a rectangle using at least one of calibration information or point cloud information when manually or automatically adding a rectangle to point cloud data.

[0018] An annotation support device according to a seventh embodiment of the present invention is an annotation support device according to any of the first to sixth embodiments, characterized in that the second execution unit highlights point clouds inside and near the annotation area, thereby making rectangular correction work more efficient.

[0019] An annotation support device according to an eighth embodiment of the present invention is an annotation support device according to any of the first to seventh embodiments, characterized in that the second execution unit makes annotation work more efficient by switching between parallel projection and perspective projection, filtering the displayed point cloud using a rectangular prism or cylinder, and embodying point cloud information using a mesh display.

[0020] An annotation support device according to a ninth embodiment of the present invention is an annotation support device according to any of the first to eighth embodiments, characterized in that the first output unit is capable of displaying point cloud annotation information using sub-views from each of the x, y and z directions in addition to the main view in order to support annotation work on an object, and provides a user interface that accepts correction operations in each view.

[0021] An annotation support method according to a tenth embodiment of the present invention includes: recording setting information and annotation information, including annotation class information, calibration information, point cloud data, and image data; a step of referencing the recorded setting information, displaying the point cloud data, and enabling annotation work; a step of referring to the recorded setting information and displaying a calibration image corresponding to a multiple image display according to the number of camera sensors to support an annotation work; executing a synchronization process based on a correspondence relationship between the point cloud data and the image data; executing various screen control processes related to the display of the point cloud data and the image data; performing an automated quality assurance process to assist in reviewing and quality checking of annotation results using AI models, including one or more of an object detection model, a region detection model, a visual language model, and a large-scale language model; a step of performing annotation processing to add a target area to the image data and correction processing for a rectangle on the point cloud data; The present invention is characterized by comprising:

[0022] An annotation support method according to an eleventh embodiment of the present invention is the annotation support method according to the tenth embodiment, recording data at multiple points in time including time series point cloud data and image data; executing a process for synchronizing the point cloud data and the image data based on the recorded time series; The present invention is characterized by comprising:

[0023] An annotation support method according to a twelfth embodiment of the present invention is an annotation support method according to the tenth or eleventh embodiment, characterized in that it includes a step of displaying an annotation target area based on point cloud data and providing a user interface that accepts correction operations in order to support annotation work on an object.

[0024] An annotation support method according to a thirteenth embodiment of the present invention is an annotation support method according to any of the tenth to twelfth embodiments, characterized in that it includes a step of switching between or simultaneously displaying calibration images and annotation data corresponding to multiple image displays according to the number of camera sensors, and providing a user interface that allows the user to visually confirm consistency.

[0025] An annotation support method according to a fourteenth embodiment of the present invention is an annotation support method according to any of the tenth to thirteenth embodiments, characterized in that it includes a step of executing a review function that evaluates omissions or consistency of an object based on recorded annotation information.

[0026] An annotation support method according to a fifteenth embodiment of the present invention is an annotation support method according to any of the tenth to fourteenth embodiments, and includes a step of, when an object exists in image data, specifying an area including at least one of a rectangle and a point by user input, or specifying an area by inputting annotation data, or extracting the object area as a rectangle using a pre-trained model, and displaying or recording the extracted area in association with point cloud data; When manually or automatically adding a rectangle to point cloud data, this method achieves high-precision initial annotation by automatically correcting the orientation and size of the rectangle using at least one of calibration information or point cloud information.

[0027] An annotation support method according to a 16th embodiment of the present invention is an annotation support method according to any of the 10th to 15th embodiments, characterized in that the method includes a step of recording annotation work logs by multiple users and making it possible to display each work history when displaying point cloud data and calibration images.

[0028] An annotation support method according to a 17th embodiment of the present invention is an annotation support method according to any of the 10th to 16th embodiments, characterized in that the method highlights point clouds inside and near the outside of the annotation area, thereby making the rectangular correction process more efficient.

[0029] An annotation support method according to an 18th embodiment of the present invention is an annotation support method according to any of the 10th to 17th embodiments, characterized in that the method makes annotation work more efficient by switching between parallel projection and perspective projection, filtering the displayed point cloud using a rectangular prism or cylinder, and embodying point cloud information using a mesh display.

[0030] An annotation support method according to a 19th embodiment of the present invention is an annotation support method according to any of the 10th to 18th embodiments, characterized in that the method is capable of displaying point cloud annotation information using sub-views from each of the xyz directions in addition to the main view in order to support annotation work on an object, and provides a user interface that accepts correction operations in each view.

[0031] An annotation support program according to a twentieth embodiment of the present invention is characterized in that it causes a computer to execute the annotation support method according to any of the tenth to nineteenth embodiments. [Effects of the Invention]

[0032] According to the present invention, it is possible to provide an annotation support device, a support method, and a support program that enable efficient execution of high-precision annotation work that links point cloud data and image data. Specifically, by centrally recording annotation class information, calibration information, target point cloud data, image data, and annotation information in a recording unit, it is possible to create a flexible work environment even for complex targets including time-series data.

[0033] Furthermore, the first output unit enables annotation work centered on point clouds, and the second output unit checks consistency with camera images, providing visual support for areas with variations in point clouds or areas that are difficult to see.Furthermore, the system is equipped with an automatic review function (AutoQA) that utilizes AI models including an automatic rectangle detection function based on images, an object detection model, an area detection model, a visual language model, and / or a large-scale language model, thereby enabling significant efficiency improvements in annotation work and ensuring quality.

[0034] In addition, the ability to record and display work logs from multiple users allows for flexible collaboration and review work. This solves the challenges that were difficult to address with conventional tools when creating large-scale, highly accurate AI training data using point clouds and images for autonomous driving, drones, robotics, and other applications, and provides practical annotation support. [Brief explanation of the drawings]

[0035] [Figure 1] 1 is a schematic diagram of a system including an annotation support device according to an embodiment of the present invention. [Figure 2] 1 is a block diagram showing a hardware configuration of an annotation support device (computer system) according to the present invention. [Figure 3] 1 is a flowchart showing a basic annotation workflow in the annotation support device. [Figure 4] 10 is a flowchart showing a point cloud-based annotation workflow in the annotation support device. [Figure 5] 10 is a flowchart showing an example of a processing flow of annotation work using a camera image as a base point in the annotation support device. [Figure 6] 1 is a flowchart showing an example of an image-based annotation workflow utilizing an AI model in an annotation support device. [Figure 7]10 is a flowchart illustrating an example of a processing flow for automatically arranging rectangles on a point cloud based on rectangle information added to an image in the annotation support device. [Figure 8] FIG. 10 is a diagram illustrating an example of annotation class information used in the annotation support device. [Figure 9] FIG. 10 is a diagram illustrating an example of a database structure showing the correspondence between point cloud data, camera image data, and a calibration file. [Figure 10] FIG. 10 is a diagram showing an example of calibration data, and is a diagram showing the structure of camera calibration information in YAML format conforming to the kitti format. [Figure 11] FIG. 10 is a diagram illustrating an example of an output of a user interface for performing annotation work on point cloud data. [Figure 12] FIG. 10 is a diagram showing an example of output from a user interface (second output unit) for checking and editing the position and attributes of an annotation object using a camera image. [Figure 13] FIG. 10 is a diagram showing an example of a user interface in an integrated mode in which annotation work is performed by integrating the displays of the first output unit (point cloud data) and the second output unit (image data) in the annotation support device. [Figure 14] FIG. 10 is a diagram showing an example of an interface in a first output unit in the annotation support device that allows comments to be added during annotation work on point cloud data, enabling collaboration among multiple users. [Figure 15] FIG. 10 is a diagram illustrating an example of an algorithm for adding a rectangle to an object from image data and generating an annotation on point cloud data. [Figure 16] FIG. 10 is a diagram illustrating an example of a process for automatically adding a rectangle to image data using an AI model. [Figure 17] FIG. 1 is a diagram illustrating an example of an automatic quality assurance (AutoQA) process utilizing a visual language model (VLM). [Figure 18]10A and 10B are diagrams illustrating an example in which the efficiency of annotation work is improved by various screen controls executed by a second execution unit. [Figure 19] FIG. 10 is a diagram showing a display control function of a gizmo that maintains visibility when displaying an enlarged or reduced image, as an example of various controls performed by the second execution unit. [Figure 20] 10A and 10B are diagrams showing how the camera image magnification control function operates in the ON / OFF state on the annotation support screen. [Figure 21] FIG. 10 is a diagram showing a frame display UI for making annotation and review work in chronological order more efficient. [Figure 22] FIG. 10 is a diagram illustrating a control function for automatically focusing on the photographing vehicle and the annotation target in accordance with frame movement during point cloud annotation work. [Figure 23] FIG. 10 is a diagram showing a display UI for improving the efficiency of annotation and review work on point cloud data. DETAILED DESCRIPTION OF THE INVENTION

[0036] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS An embodiment of the present invention will now be described with reference to the accompanying drawings. The individual embodiments of the present invention are not independent and can be appropriately combined with each other for implementation.

[0037] Fig. 1 is a block diagram showing the hardware configuration of a system including an annotation support device according to one embodiment of the present invention. As shown in Fig. 1, the system S includes terminals 1-1 to 1-N (N is a natural number) used by users, and a computer system 2 connected to these terminals 1-1 to 1-N via a communication network CN.

[0038] The terminals 1-1 to 1-N are devices used by users who perform annotation work, and include, for example, multi-function mobile phones (so-called smartphones), tablets, notebook computers, desktop computers, etc. These terminals are collectively referred to as terminals 1.

[0039] The computer system 2 functions as an annotation support device by executing the annotation support program of the present invention. For example, it is used or managed by an administrator of an organization (including a company) that manages annotation target data. The computer system 2 executes processing in response to a request from the terminal 1 and provides the results to the terminal 1. The computer system 2 may be configured as a single computer or may be configured as multiple computers. An example configured as a single computer will be described below.

[0040] 2 is a block diagram showing the hardware configuration of an annotation support device (computer system) according to the present invention. As shown in FIG. 2, the computer system 2 includes an input interface 21, a communication module 22, a storage device 23, a memory 24, an output interface 25, and a processor 26. The input interface 21 accepts an operation input from an administrator of the computer system 2 and outputs a signal corresponding to the accepted input to the processor 26. The communication module 22 is connected to a communication line network CN and performs data communication with the terminal 1. Note that this communication may be wired or wireless, but the present embodiment will be described assuming a wired connection.

[0041] The storage device 23 is, for example, a storage device, and stores various data and programs that are read and executed by the processor 26. The memory 24 is a storage area for temporarily storing data and programs, and is configured as a volatile memory, for example, RAM (Random Access Memory). The output interface 25 enables connection to an external device and has the function of outputting signals to the external device. The processor 26 loads a program stored in the storage device 23 into the memory 24 and executes a series of instructions contained in the program, thereby operating as the following functional blocks: That is, they are a recording unit 261, a first output unit 262, a second output unit 263, a first execution unit 264, a second execution unit 265, a third execution unit 266, and a fourth execution unit 267.

[0042] The recording unit 261 is a component that plays a core role in the annotation support device according to the present invention, and systematically stores and manages annotation target information such as point cloud data and image data, as well as various setting information required for annotation work, work results, user logs, etc. Specifically, the recording unit 261 records the following information: (1) Annotation class information Classification information set for each annotation object, including label name, attributes, color, drawing style, recommended rectangle size (default length), etc. This allows annotators to mark objects in a consistent format, ensuring accuracy and uniformity. (2) Calibration information It holds the external calibration and internal calibration information required to associate image data with point cloud data. For example, the position and orientation of each camera, focal length, distortion coefficients, perspective transformation matrix, etc. are stored in a YAML file format (e.g., KITTI format). (3) Point cloud data (LiDAR, etc.) Point cloud data that captures a target scene in three dimensions is stored. In particular, in the present invention, point cloud data corresponding to multiple points in time are managed as time-series frames, and each frame is managed together with the corresponding image group. (4) Image data This is image data acquired by multiple cameras, and the time-series and spatial correspondence with point cloud data is recorded. Calibration information is linked to each image and used for projection correspondence with the point cloud. (5) Annotation information The annotation information assigned by the user or the fourth execution unit 267 is stored. This includes point clouds, rectangles on images, attribute labels, annotation reliability, revision history, etc. Version management is also possible for review purposes. (6) Work logs and user information The system records each user's annotation operation history, comments, corrections, review history, etc. in chronological order, making it possible to assign responsibility and manage the progress of work in collaborative work. (7) Estimated Metadata The estimation method type (clustering / AI model / combined) for cuboid (oriented bounding box: OBB) estimation, key parameters (cluster threshold, neighborhood distance, minimum number of points, quantile range, etc.), reliability indicators (e.g., cluster density, reprojection error), and time series smoothing status (e.g., estimation error covariance) are recorded. This provides traceability that is useful for subsequent reviews (third execution unit 266), re-estimation, and audits.

[0043] As described above, the recording unit 261 serves as the operating platform for the entire annotation support device, performing a variety of functions, such as permanently storing data, providing consistent reference, tracking work history, and maintaining evidence for quality assurance, thereby significantly contributing to improving the accuracy, efficiency, and traceability of annotation.

[0044] The first output unit 262 is a component in the annotation support device according to the present invention that visually displays the point cloud data stored in the recording unit 261 and provides a user interface for enabling annotation work to be performed. The first output unit 262 has the following functions: (1) Visualization function for point cloud data The first output unit 262 renders and displays the point cloud data in three-dimensional space. Through operations such as changing the viewpoint, rotating, and zooming in and out, the user can intuitively grasp the shape and positional relationship of the target object. (2) Display and edit function for annotation areas This output unit can overlay annotations such as rectangular areas on the point cloud data. This allows users to check and modify existing annotation information, as well as add new annotations. The shape and coordinate information of the rectangle are saved in the recording unit 261 and are also used for linking with other modules (e.g., the fourth execution unit 267). (3) Providing a correction interface The first output unit 262 provides a GUI (Graphical User Interface) that allows the user to correct the position and size of the displayed annotation by directly operating it. For example, it is possible to correct an incorrectly placed rectangle by operating the mouse, which contributes to improving the accuracy of the annotation. (4) Point cloud annotation review support By working in cooperation with the third execution unit 266, it is possible to overlay the quality check results (e.g., warnings of missing labels, indications of inconsistency, etc.) for the annotation content using AI models (including one or more of object detection models, area detection models, visual language models, and large-scale language models) on the GUI displayed by the first output unit 262. (5) Collaboration support The first output unit 262 supports simultaneous work and review work by multiple users, and can display work logs, comments, etc. in real time, thereby improving the efficiency of the review and revision process within a team.

[0045] As such, the first output unit 262 is an extremely important component for visually supporting and streamlining annotation work on point cloud data, achieving both intuitive user operation and highly accurate annotation processing. To enable users to perform point cloud annotation more accurately, the first output unit 262 provides sub-views based on viewpoints from the X-axis, Y-axis, and Z-axis directions in addition to the main view. These views allow users to check the point cloud area to be annotated from multiple angles, and correction operations can be performed in each view. This improves the accuracy of annotation based on the three-dimensional shape of the object, enabling fast and accurate work even on objects with complex structures.

[0046] The second output unit 263 is a component in the annotation support device according to the present invention that displays the calibration image and annotation data stored in the recording unit 261 and provides a user interface to assist in annotation work. The second output unit 263 mainly performs the following functions: (1) Image data display function A calibration image is created using the calibration information (2) and image data (4) stored in the recording unit 261. The still images of each frame constituting the video are corrected using the focal length and distortion coefficient of the camera used to capture the video to be annotated. Many in-vehicle cameras use multiple cameras to capture images with small lenses so the driver can check the surroundings. Some of these cameras use fisheye lenses, making it impossible for AI to align them with point cloud data as they are. Because the images and the lidar that acquires the point cloud are not synchronized, the images are captured and data is acquired at different time constants, and the frame rates do not match. Furthermore, because the camera and sensor are mounted in different positions, the three-dimensional positions of the captured data differ. Therefore, to synchronize the image data and point cloud data, the second output unit 263 displays the calibrated camera image using the camera image, point cloud data, and calibration information, allowing the user to identify the object in the image. Images are obtained from the data for each camera stored in the recording unit 261, and an appropriate view is selected according to the scene. (2) Switching and comparing calibration images and annotation data The second output unit 263 has a function to switch between displaying calibration images and annotation data or displaying them in parallel (simultaneously). It also supports multiple calibration images according to the number of camera sensors, and has a function to display them separately on a separate browser screen and a function to change the image layout to optimize the work area. This makes it easier for users to visually check the consistency between images, helping to prevent misrecognition and annotation errors. (3) Overlaying annotation information It provides a user interface that displays annotation information (rectangular areas, labels, etc.) on images saved in the recording unit 261 as an overlay and enables editing and correction. In particular, in addition to manual annotations, the results of automatic annotations (by the fourth execution unit 267) are also displayed, allowing comparison and examination. (4) Zoom and magnification control The display image has a zoom-in / zoom-out function to enable detailed image confirmation, allowing for detailed annotation confirmation and fine correction with high precision. (5) Consistency confirmation support It is possible to check the consistency between point cloud data and image data in cooperation with the first output unit 262. For example, it supports the task of visually verifying the correspondence between the rectangular position of an object annotated on an image and the target area on the point cloud side. (6) Customizable user interface The software also includes a function that allows users to dynamically adjust image brightness, contrast, whether or not to display a grid, etc., depending on their needs, improving visibility and work efficiency.

[0047] In this way, the second output unit 263 has a variety of functions to support image-based annotation work and consistency checking, and is a component that plays an important role in supporting multimodal annotation in conjunction with point cloud data.

[0048] The first execution unit 264 has a function of executing synchronization processing of point cloud data and image data based on the correspondence between these data. The point cloud data and image data are collected in chronological order, and it is required to accurately match the correspondence between the data at each point in time (for example, data corresponding to the same time or the same frame number). The first execution unit 264 analyzes the correspondence between the point cloud data and image data using metadata such as calibration information, shooting time, and sensor ID stored in the recording unit 261, and performs conversion processing and correction processing for synchronous display.

[0049] For example, if the point cloud is a spatial 3D coordinate group acquired by a laser scanner and the image is visual information simultaneously captured by a camera, the first execution unit 264 maps the coordinate system of the point cloud data to the image coordinate system based on the calibration matrix (internal parameters and external parameters). For time-series data, the first execution unit 264 dynamically switches the corresponding point cloud-image pair for each frame, taking into account sensor synchronization between multiple frames.

[0050] This type of synchronization processing enables annotation work in the first output unit 262 and the second output unit 263, which will be described later, to be performed while maintaining consistency between the point cloud and the image, which contributes to intuitive user operation, improved visibility, and improved annotation accuracy.

[0051] The second execution unit 265 has a function of executing various control processes related to the display screens of the first output unit 262 and the second output unit 263. That is, in order to improve operability on the user interface, the second execution unit 265 has a variety of display control functions that support annotation work, such as switching control in the screen display of the output unit, zoom-in / zoom-out control, switching between parallel display and integrated display, and overlay display.

[0052] For example, the second execution unit 265 displays calibration images and annotation data in parallel, corresponding to the number of camera sensors, allowing the user to visually compare and confirm them as they work. Also, by overlaying point cloud and image displays in a single integrated view, a work mode can be realized in which spatial consistency can be visually confirmed while annotating. Furthermore, the system incorporates features to enhance visual review and operational efficiency, such as a function to control the magnification of camera images, a function to automatically or manually switch between frames in chronological order, focus control for the target object, and automatic switching between calibration images corresponding to the target annotation.

[0053] The second execution unit 265 also has a function to highlight the inside of the annotation rectangle and the point clouds existing in its vicinity in order to clarify the area of ​​the annotation target. This makes it easier for the worker to visually grasp the outline of the target area and its boundary with the surroundings, allowing for efficient rectangle correction of the annotation. Even when the outline of the target becomes unclear due to the density of the point cloud or the influence of noise, the highlighting allows the worker to accurately grasp the area to be corrected.

[0054] Furthermore, the second execution unit has the ability to switch display methods and filter information. Specifically, it supports switching between parallel projection and perspective projection, allowing users to view point clouds from different perspectives with different depths and sizes. Shape filters, such as rectangular prisms and cylinders, can be used to limit the point clouds to be displayed, making it easier to distinguish between target objects and other data. Furthermore, displaying the point cloud as a mesh makes it easier to visually recognize the continuity of the shape's edges and surfaces, supporting accurate annotation work.

[0055] These control functions allow users to flexibly switch display modes depending on the task at hand, enabling them to intuitively and efficiently carry out annotation work, resulting in improved annotation quality and reduced work time.

[0056] The third execution unit 266 has the function of performing automatic quality assurance processing that supports review and quality confirmation (AutoQA) of annotation results using AI models (including one or more of an object detection model, a region detection model, a visual language model, and a large-scale language model).

[0057] Specifically, the third execution unit 266 refers to the annotation information, point cloud data, image data, annotation class information, etc. stored in the recording unit 261, and automatically detects and points out problems such as missing, overlapping, and inconsistent annotations using an AI model. Furthermore, to make it easier for the user to check, annotations that are determined to be problematic can be graphically highlighted or feedback can be provided in natural language.

[0058] The third execution unit 266 also has a function to save the review results as a log and a function to support comparison with the user's manual correction history, thereby improving the reliability of quality control and work records. This reduces human error in annotation work and supports the generation of highly accurate data.

[0059] In this way, the third execution unit 266 functions as a multimodal quality assurance support mechanism that utilizes AI, and realizes the efficiency and automation of review work that has traditionally been performed manually.

[0060] The fourth execution unit 267 has a function to automatically detect a target area in image data and perform annotation processing to assign it as a rectangle. This makes it possible to significantly improve the efficiency of annotation work on images, which has previously been done manually.

[0061] The fourth execution unit 267 has a function of automatically recognizing objects contained in image data using a pre-trained AI model (e.g., Segment Anything Model, etc.) and generating annotation data that surrounds the object with a rectangle. This rectangular annotation is accurately added based on the boundary of the object, so it can be used as a high-quality initial annotation. Furthermore, the fourth execution unit 267 has a function of projecting rectangular information obtained from an image onto the point cloud data based on the calibration information and point cloud data to generate a corresponding rectangle in three-dimensional space, and automatically estimating the posture (orientation) and dimensions of the rectangle and correcting them as necessary.

[0062] The fourth execution unit 267 extracts a point cloud corresponding to the projection area of ​​the target rectangle obtained from the image data, and separates a point set of target candidates by applying clustering (e.g., density-based clustering (DBSCAN) or cluster extraction based on Euclidean distance) to the point cloud. It estimates the principal axes of the obtained point set using principal component analysis (PCA) of the covariance matrix, and calculates a rectangular parallelepiped (oriented bounding box: OBB) that is aligned with the estimated principal axes, thereby automatically inferring the orientation and dimensions of the target.

[0063] In OBB dimension estimation, robust edge length estimation can be performed based on the quantiles of the point cloud distribution (e.g., 2.5 to 97.5 percentile) to reduce the influence of outliers. In addition, consistency can be improved by applying dimension constraints to the estimation results using prior information on the predetermined dimension range or aspect ratio stored in the recording unit 261 for each target class.

[0064] When the point cloud is sparse or the influence of occlusion is significant, instead of or in combination with clustering, a trained AI model (e.g., a 3D object detection model or a pose estimation model) can be used to directly estimate the three-dimensional pose and dimensions of the object, and adopt them as the OBB or use them as initial values. These processes can be applied as automatic correction not only to automatically generated rectangles based on images, but also to rectangles manually placed on the point cloud by the user.

[0065] In addition, the fourth execution unit 267 also has a function of assigning a correspondence between the image annotation result and the point cloud data, which allows the rectangular information obtained from the image to be projected onto the point cloud data and the corresponding annotation in three-dimensional space to be recorded in the recording unit 261. Furthermore, manual correction by the user and cooperation with subsequent automatic review (third execution unit 266) are also taken into consideration, making it possible to implement a hybrid operation that combines automation of the annotation process with manual quality assurance.

[0066] The fourth execution unit 267 extracts a reference plane corresponding to the ground surface (for example, ground estimation by RANSAC), stabilizes the pitch and roll components of the OBB using normal direction constraints on the reference plane, and determines the yaw component by principal axis estimation, thereby realizing attitude estimation that is consistent with the real environment.

[0067] Furthermore, if time-series data is recorded, we use a time-series estimation method such as a Kalman filter to smooth the position, orientation, and dimensions of the OBB between consecutive frames, ensuring temporal consistency of the estimated values. This allows us to obtain stable initial annotations even in scenes with varying density or partial occlusions.

[0068] In this way, the fourth execution unit 267 is an important component that contributes to improving work efficiency and accuracy as an image-based automatic annotation mechanism and a high-quality initial annotation mechanism.

[0069] 3 is a flowchart showing a series of annotation processing flows (basic workflow) in the annotation support device according to the present invention. This processing starts in step S101, when point cloud data recorded in the recording unit is displayed via the first output unit 262. In step S102, in addition to displaying the point cloud, basic layout adjustments suited to the work are performed, such as displaying a calibration image, displaying a camera image in conjunction with the image, and dividing the display area. The second output unit 263 and the second execution unit 265 can be involved in these adjustments.

[0070] In step S103, all of the annotation target data recorded in the recording unit 261 is read as the target for processing. In the following step S104, the user or the fourth execution unit 267 adds or modifies rectangular annotations. This includes processing to automatically extract rectangles based on image data and fine adjustments by manual operation.

[0071] Next, in step S105, attribute information (category information, reliability, etc.) associated with the annotation is assigned or corrected. After that, in step S106, all annotation target data is checked again, and the processes from steps S103 to S105 can be repeated until all annotation target data has been processed.

[0072] In step S107, the third execution unit 266 performs an AutoQA (automated quality assurance) check on the entire annotation before submission. In this process, an AI model including one or more of an object detection model, a region detection model, a visual language model, and a large-scale language model is used to check for omissions and consistency. In step S108, it is determined whether there are any check items, and if there are (Yes in S108), the process returns to the correction process in step S104. If there are no check items (No in S108), the process proceeds to step S109, where the annotation data is finally stored in the recording unit 261 and, if necessary, is submitted to an external system or the like.

[0073] In this way, according to the present invention, annotation processing for point clouds and image data can be carried out efficiently and with high accuracy as a series of flows including optimization of work layout, automatic rectangle assignment, and automatic quality inspection.

[0074] 4 is a flowchart showing a detailed flow of annotation processing starting from point cloud data in the annotation support device according to the present invention. This processing starts in step S201 by reading all point cloud data to be annotated that is recorded in the recording unit 261. In step S202, an annotation target is searched for while moving and enlarging / reducing the display state of the point cloud. This processing is supported by the first output unit 262 and the second execution unit 265, making it easier for the user to grasp the position and shape of the target.

[0075] Next, in step S203, the camera image associated with the point cloud data is displayed by the second output unit 263. In step S204, the user selects an appropriate annotation class based on the objects captured in the camera image. By taking advantage of the high visibility on the image, accurate classification is possible even for objects that are difficult to determine from the point cloud alone.

[0076] In step S205, a rectangular annotation is added to the point cloud area corresponding to the selected object, and the size is adjusted as necessary. Furthermore, in step S206, attributes (type, direction, label, etc.) of the object are added based on the camera image. Then, in step S207, all annotation target data is checked again, and the processes from step S202 to S207 can be repeated until all annotation target data has been processed.

[0077] In step S208, a process is performed to detect and correct missing or incomplete annotations that are likely to be overlooked using only the point cloud, while referencing the camera image. This process is performed using the automatic quality assurance (AutoQA) function of the third execution unit 266 or by visual confirmation by the user, and is important for compensating for missed detections and lack of consistency of objects that cannot be captured in annotation work that relies solely on point cloud data. This process prevents overlooking "invisible parts" and "distant objects," which are likely to occur in point cloud-based work, and significantly improves the comprehensiveness and accuracy of annotations.

[0078] If step S208 determines that additional corrections are necessary, the process returns to step S202, where the process can be repeated to search for the object and display the point cloud data and camera images, etc. This creates a loop structure that allows for flexible and repeated application and correction of annotations to objects with omissions or errors, improving the quality of the annotation results.

[0079] In this way, the annotation support device of the present invention, while based on work using point clouds as a base, realizes efficient and highly accurate annotation work in terms of both visibility and operability by displaying the corresponding camera images in conjunction with the images.

[0080] FIG. 5 is a flowchart showing an example of the processing flow of annotation work based on camera images in the annotation support device according to the present invention. Unlike conventional point cloud-based annotation work, this image-based annotation work allows objects to be directly viewed and manipulated on the image, significantly reducing the burden of searching for objects. Furthermore, adding rectangles on highly visible images reduces the occurrence of missed annotations, achieving an overall efficient annotation work. First, in step S301, all image data to be annotated is acquired. In the following step S302, an annotation class is selected by the user or automatically estimated based on the objects captured in the camera image. In step S303, a rectangle is added to the image for the object.

[0081] In step S304, a corresponding rectangle is automatically added to the point cloud position based on the coordinate information of the rectangle added to the image and the default size information associated with the annotation class. This automatic addition is performed by automatic annotation processing by the fourth execution unit 267. As a subsequent process of step S304, OBB estimation processing can be performed on the point cloud side to automatically correct the orientation and size of the rectangle. If the clustering result is insufficient, it can be complemented by a trained AI model.

[0082] Then, in step S305, all annotation target data is checked again, and steps S301 to S305 can be repeated until all annotation target data has been processed. In step S306, all image data is acquired for all annotation rectangles. If depth size adjustments are required on the point cloud, fine adjustments are made in step S307. Furthermore, in step S308, attribute information (e.g., vehicle type, color, direction, etc.) is assigned based on the characteristics of the object captured in the camera image. In this way, annotations based on image information make it easier to grasp the overall image of the object, allowing for more accurate and simple attribute setting. Finally, in step S309, all annotation rectangles are checked as targets, and steps S306 to S309 can be repeated until all data has been processed. The process then proceeds to the correction and saving process.

[0083] In this way, in an image-based workflow, by starting the annotation from the image, workers basically no longer need to search for the target object on the point cloud, and the process of adding rectangles is also automated, making the annotation process significantly more efficient than conventional point cloud-based methods.

[0084] 6 is a flowchart showing an example of image-based annotation work using an AI model in the annotation support device according to the present invention. The present invention utilizes advances in deep learning technology in the image processing field to significantly improve the efficiency and accuracy of annotation work.

[0085] In step S401, an AI model is used to automatically add rectangles to the image. The AI ​​model here uses a deep learning-based algorithm for object detection (e.g., YOLO, Faster R-CNN, SAM (Segment Anything Model), etc.), and automatically draws rectangles for objects detected in the image. At this rectangle addition stage, the position on the point cloud is calculated based on camera calibration information, etc., and automatic addition to the corresponding point cloud position is also performed at the same time.

[0086] Next, in step S402, the attributes of the annotation object are automatically estimated based on the assigned rectangle. Attribute information includes the object type, moving / still status, driving lane, vehicle direction, etc., and is either estimated by AI or determined based on the settings for each annotation class.

[0087] In step S403, all annotated rectangles are displayed as a list, allowing the user to confirm their contents. This allows for visual identification of errors such as overlapping or missing rectangles and misidentification of classes. Next, in step S404, the depth-direction size of the rectangles is adjusted based on their correspondence with the point cloud data. Since the initial depth size assigned in step S401 uses a default value for each class, the user can fine-tune it to match the actual object in the point cloud. Following step S404, OBB estimation based on clustering or an AI model can be performed to correct the 3D orientation and dimensions of the image-based annotated rectangles. If time-series data is available, smoothing processing (e.g., a Kalman filter) can be applied to ensure frame-to-frame stability.

[0088] In step S405, attribute information is corrected based on the information about the object captured in the camera image. Here, the attributes automatically assigned in step S402 are visually checked by a human and corrected if there are any errors. This process ensures the accuracy of the AI-based estimation results while improving the quality of the final annotation.

[0089] Finally, in step S406, all annotation rectangles are reconfirmed, and the confirmed data is saved in the recording unit 261. The processes from steps S403 to S406 can be repeated until all data has been processed. At this time, the history of each annotation, worker information, version information, etc. are also recorded, making it possible to handle future quality control, difference checks, etc.

[0090] As described above, the image-based annotation workflow of the present invention starts with rectangle assignment and attribute estimation using an AI model, and by combining this with point cloud integration and manual correction processing, it is possible to achieve significant improvements in work efficiency and accuracy compared to conventional methods.

[0091] 7 is a flowchart showing an example of a processing flow for automatically arranging a rectangle on a point cloud based on rectangle information added to an image in the annotation support device according to the present invention. That is, in image-based annotation work, by using the rectangle added on the image as a reference and three-dimensionally converting and projecting the rectangle onto the corresponding point cloud area, initial arrangement becomes possible, which eliminates the need for annotation work on the point cloud.

[0092] In step S501, a user or an AI model inputs a rectangle on an image, and the annotation support device receives information about the rectangle (such as coordinate information). This rectangle encloses the annotation target and includes at least information about the height and width directions in the image coordinate system.

[0093] Next, in step S502, the coordinates of the point cloud corresponding to each vertex of the rectangle are identified using the coordinate information of the rectangle and the calibration information of the camera (internal and external parameters). This defines a two-dimensional plane on the point cloud corresponding to the height and width directions of the image. Clustering is applied to the candidate point cloud identified in step S502 to extract major clusters, and principal axis estimation is performed on the extracted clusters to derive the center, orientation, and dimensions of the OBB.

[0094] In step S503, a projection vector is calculated from the photographing vehicle toward the point cloud coordinate system, starting from the center of the image rectangle, and the position where this vector intersects with the point cloud data is searched for. This identifies the starting point of the depth direction (Z direction) in the point cloud space, i.e., the front position of the rectangle.

[0095] Finally, in step S504, the end point in the depth direction is interpolated by referring to the standard depth dimension (default size) set for the annotation class. This completes the initial placement of a 3D rectangle (voxel) in the point cloud area from the 2D rectangle information added to the image. After the depth is interpolated by referring to the default dimension in step S504, a consistency check with the OBB estimation result (reprojection error, volume difference, class-specific dimension constraints) is performed, and if a threshold is exceeded, the OBB estimation side can be prioritized or integrated using a weighted average.

[0096] In this way, the present invention can automatically convert an image into a point cloud into a rectangle based on geometric processing, eliminating the need for tasks such as searching for objects on the point cloud and drawing rectangles. As a result, the annotation workload can be significantly reduced, and the variability between workers can be reduced.

[0097] 8 is a diagram showing an example of attribute information stored by default for each annotation class in the annotation support device according to the present invention. The present invention includes a mechanism for automatically complementing the size in the depth direction (Z direction) when a rectangle is added to an image and a three-dimensional rectangle on a point cloud is generated based on the information about the rectangle. This complementation process makes it possible to accurately achieve the initial placement of the rectangle on the point cloud by referring to a default depth length set in advance for each annotation class.

[0098] As shown in Fig. 8, annotation classes are set as "car," "truck," "bus," etc., and these are associated with default lengths (depth direction), such as "car: 4000 mm," "truck: 6500 mm," and "bus: 12000 mm." This data is stored in the database as annotation class information and is referenced when generating a 3D rectangle on the point cloud.

[0099] By retaining such annotation class information, it becomes possible to automatically and efficiently assign 3D rectangles to point clouds based on rectangles obtained from images, which will greatly contribute to automating and reducing the labor required for annotation work.

[0100] 9 is a diagram showing the correspondence between point cloud data, camera image data, and calibration files stored in a database in the annotation support device. As shown in Fig. 9, image data acquired by multiple cameras (e.g., cameraA.png, cameraB.png, cameraC.png) are linked to point cloud data corresponding to each scene (e.g., 001.pcd, 002.pcd), and calibration files (e.g., file1.yaml, file2.yaml, file3.yaml) corresponding to each camera image data are also associated and stored.

[0101] This configuration allows for the systematic management of camera image information from multiple viewpoints and the corresponding calibration information for a single point cloud, which serves as a foundation for accurately projecting image-based annotations onto the point cloud and achieving highly accurate rectangular placement and attribute correction.

[0102] In particular, by accurately understanding the correspondence between point cloud data and image data, it becomes possible to efficiently reflect the rectangles annotated for each image in point cloud space, and to efficiently perform confirmation work by switching between multiple viewpoints during point cloud review. Therefore, a structured data storage method such as the one shown in this figure greatly contributes to improving the accuracy and efficiency of annotation work.

[0103] FIG. 10 shows the data format of the calibration information used in the present invention, illustrating a specific example of the external and internal parameters of a camera written in YAML format. In this embodiment, a calibration file in YAML format conforming to the KitTi format is used to accurately establish the correspondence between camera images and point clouds. The calibration file is defined for each camera used for shooting and contains the following information:

[0104] First, the extrinsic parameter matrix defined as CameraExtrinsicMat realizes the transformation from the camera coordinate system to the world coordinate system and is composed of a 3x4 rotation matrix and a translation vector. Also, the intrinsic parameter matrix (CameraMat) describes the focal length and optical center (principal point) of the camera as a 3x3 matrix.

[0105] In addition, the distortion coefficient defined as DistCoeff describes the parameters (k1, k2, p1, p2, k3) for lens distortion correction, which enables high-precision image processing even under special optical conditions such as fisheye lenses. Furthermore, the ImageSize field indicates the image resolution (width and height), and the Reprojection Error indicates a value that represents the accuracy of the calibration process.

[0106] In this way, by using calibration information accurately described for each camera, it becomes possible to perform highly accurate projection transformation between an image and a point cloud, thereby improving the accuracy and reliability of processes such as rectangle position calculation, attribute estimation, and review correction in the annotation support device of the present invention.

[0107] Fig. 11 is a diagram showing an example of an annotation screen for point cloud data output by the first output unit 262 in the annotation support device according to the present invention. The user interface shown in Fig. 11 supports annotation work on a three-dimensional point cloud, and allows a user to intuitively and efficiently perform annotation work by placing a rectangle (bounding box) for an object displayed in the point cloud.

[0108] On the left side of the screen, a rectangle is superimposed on the point cloud data, and attribute information (such as the class name "Car" or an identifier) ​​is assigned to the selected annotation. In addition, on the screen on the right side of the center, the shape and accuracy of the rectangle as viewed from each direction of the point cloud (top, back, side) can be confirmed, and the user can adjust and check the rectangle in three dimensions using these views.

[0109] Additionally, the right side of the screen displays items for directly editing the coordinate information (X / Y / Z coordinates) and size information (total length X, total width Y, total height Z) of the selected annotation by entering values, allowing for highly accurate numerical adjustments. Furthermore, the app also provides a timeline display, tagging function, annotation class selection and display switching function, and settings for reflecting data in past and future frames, enabling efficient annotation work for multi-frame and time-series data.

[0110] In this way, the user interface of the present invention provides integrated visualization and editing functions for 3D point clouds, making it possible to comprehensively improve the accuracy, work efficiency, and operability of annotation work.

[0111] 12 is a diagram showing an example of an annotation screen for image data output by the second output unit 263 in the annotation support device according to the present invention. This screen is an interface for selecting and displaying an arbitrary viewpoint from images captured by multiple cameras and performing annotation work on an object. An image of the selected viewpoint (e.g., "back_camera__00.jpg") is displayed large in the center of the screen, and a rectangle (bounding box) can be added to an object such as a car.

[0112] Additionally, multiple-viewpoint thumbnail images are displayed at the bottom, and users can switch viewpoints by clicking on these thumbnails. This allows annotation work to be performed while viewing the object from multiple directions, enabling more accurate position estimation and attribute determination.

[0113] With this configuration, the second output unit displays the image data to be annotated and the auxiliary information based on multiple viewpoints in an integrated manner, helping the user to efficiently add rectangular annotations to images. Furthermore, the annotation information is a core component of the image-based annotation support mechanism of the present invention, as it is also used in subsequent processes such as reflecting the annotation information in point cloud data and modifying attributes.

[0114] Fig. 13 is a diagram showing an example of a user interface in a mode in which the first output unit 262 (point cloud-based annotation screen) and the second output unit 263 (image-based annotation screen) of the annotation support device are displayed in an integrated manner. As shown in this figure, a view based on 3D point cloud data (point cloud view) is displayed in the center, and the user can add rectangles to the point cloud to perform annotation. In the point cloud view, a rectangular annotation (e.g., a vehicle) is drawn three-dimensionally, and detailed information such as coordinate values, angles, and dimensions (total length, total width, total height) can be edited and displayed in the panel on the right.

[0115] Furthermore, camera images synchronized with the point cloud are displayed in the upper left of the screen, allowing users to visually confirm how annotations added on the point cloud view are associated with the corresponding camera images. Multiple camera images are displayed in a list at the bottom of the screen, either in chronological order or by viewpoint, allowing users to switch between them as they work.

[0116] In this way, in the integrated mode shown in Figure 13, the first output unit 262 and the second output unit 263 are configured to allow both point cloud data and image data to be viewed and manipulated simultaneously, thereby improving the visibility, operability, and accuracy of annotation work.

[0117] FIG. 14 is a diagram showing how annotations are pointed out and shared using the comment function and collaboration function on the first output unit 262 (annotation screen based on point cloud data) in the annotation support device. As shown in this figure, a comment box that enables communication between users is displayed for a rectangular annotation added to 3D point cloud data. For example, a user can write specific instructions to another user, such as "Please add a rectangle without including the side mirrors." This comment is linked to the corresponding annotation and saved, and can be referenced later when reviewing or correcting it.

[0118] The right panel displays a list of comments in thread format, allowing for centralized management of feedback given to the target frame or annotation. In addition, by selecting a comment, the UI / UX allows you to instantly jump to the corresponding annotation location in 3D space and intuitively check the point of criticism, improving the efficiency of the review workflow.

[0119] This allows multiple users to collaborate on the same data simultaneously or asynchronously, improving annotation accuracy and work efficiency.One aspect of the present invention is characterized by the interface configuration shown in this figure, which allows visual comments to be added to point cloud annotations and efficient feedback circulation within the team.

[0120] 15 is a diagram showing an example of a process in which a rectangle on an image is automatically projected and converted onto point cloud data in the annotation support device to generate a rectangular annotation. That is, the process shows a series of processing flows in which, when a user adds a rectangle to image data (second output unit 263), a rectangle is automatically generated on point cloud data using information about the rectangle and displayed on the first output unit 262.

[0121] In the first step, a user draws a rectangle around an object to be annotated (e.g., a vehicle) on an image (e.g., an image captured by an on-board camera), and the coordinates (x, y) of each vertex of the rectangle are acquired. At this time, the rectangle can be assigned either manually or automatically.

[0122] Next, in the second step, the first execution unit 264 refers to the calibration information and converts the two-dimensional coordinates on the image into the height and width coordinates on the corresponding point cloud data, thereby identifying the positions of the upper and lower ends of the object on the point cloud.

[0123] In the third step, the position in the depth direction (i.e., the Z coordinate of the starting point of the rectangle (origin in the depth direction) is identified from the acquired height and width information and its distribution on the point cloud. Subsequently, in the fourth step, the first execution unit 264 constructs a rectangular annotation (height, width, length, and position coordinates) on the point cloud space using standard length dimension information corresponding to the object class (e.g., "Car", "Truck", etc.) stored in the recording unit 261.

[0124] Then, in the fifth step, the 3D rectangle generated as described above is automatically displayed on the user interface of the first output unit 262, allowing the user to immediately visually confirm the annotation results in the point cloud space. In this way, with the configuration shown in Fig. 15, the user can semi-automatically generate accurate 3D annotations in the point cloud space simply by adding a rectangle to the image, thereby achieving both a reduction in the burden of annotation work and an improvement in accuracy.

[0125] Fig. 16 is an explanatory diagram showing an example of a process in which rectangular annotations are automatically added to camera images in the annotation support device. In the first step, all camera image information associated with a point cloud is input. This includes images acquired from multiple cameras mounted on a vehicle, and each image is spatially associated with the point cloud and calibration information.

[0126] Next, in the second step, the first execution unit automatically detects objects in the image based on the annotation class information and mapping information acquired from the recording unit, and automatically assigns rectangular annotations to the objects. This process is realized using a machine learning (AI) model, such as an object detection model (e.g., YOLO or Faster R-CNN), and assigns different class labels to each object type (e.g., vehicle, sign, pedestrian, etc.).

[0127] In the third step, the automatically annotated rectangles are reflected on all camera images at once. This eliminates the need for users to individually annotate images from multiple viewpoints, enabling extremely efficient annotation work. As noted in the note at the bottom right of the figure, after this process, the "algorithm for adding rectangles from images to point clouds" shown in the previous figure (Fig. 15) connects to a series of steps in which the annotation information on the images is reflected on the point cloud data. The automatic annotation function shown in this figure has the advantage of automating the initial annotation process for large amounts of image data, significantly reducing the burden of subsequent manual confirmation and point cloud conversion processing.

[0128] FIG. 17 is an explanatory diagram showing an example of an annotation support device performing automatic quality assurance (Auto QA) processing using a visual language model (VLM). As shown in this diagram, the annotation support device can automatically verify the quality of the type (class) of an object annotated in an image. The input section shown on the left receives a pre-annotated image and a corresponding prompt (definition of classification rules, etc.). The prompt explicitly states the classification criteria, such as "Objects to be classified are those enclosed in a red BBOX (thick line in the image)," "Classification is into six categories: car, bus, truck, pedestrian, bicycle, and motorcycle," and "People riding bicycles are classified as bicycles, and all others are classified as pedestrians."

[0129] This input information is provided to the VLM. While this example uses a large-scale multimodal model such as Qwen, other VLMs are also applicable. The VLM analyzes the provided prompt and annotation image to determine whether the object in the image conforms to the specified class. The output section shown on the right displays the Auto QA results for the target object (in this example, a vehicle annotated as "truck"), displaying a message stating, "The annotation class may be incorrect (expected: bus, current: truck)." In this way, the VLM can automatically check the consistency of the annotation content and output a warning if there is any doubt. This mechanism enables objective and efficient review of annotation results, improving annotation quality and reducing the burden of confirmation work.

[0130] Furthermore, the annotation support device of the present invention not only automatically evaluates annotations (AutoQA) using a Visual Language Model (VLM), but also has various automatic check rules to ensure annotation quality, making it possible to automatically verify quality from a variety of perspectives, such as the following: [List of rules that can be automatically checked by AutoQA] Point cloud checks Checking ground contact status (detecting ungrounded objects) -Detection of raised or recessed points in the point cloud (rectangular mismatch) -Detecting unfilled areas in segmentation -Detection of overlaps and distance anomalies between annotation cuboids -Detection of abnormal density based on the number of points inside a rectangular parallelepiped Video check -Detection of annotations that extend outside the image area -Detecting unfilled areas in segmentation - Check for discontinuity and separation of annotation areas (separation of areas that should be one) ·Checking common rules · Check the consistency of multiple attribute information (e.g. category, car model, color, etc.) · Check for missing required attribute information These check items, combined with label consistency verification by VLM, enable comprehensive quality evaluation of the entire annotation, significantly contributing to the efficiency of review work and the reduction of human error. These processes are performed by the third execution unit.

[0131] 18 shows a specific example of how the annotation support device can improve the accuracy and efficiency of annotation work by using various display control functions provided by the second execution unit 265. As shown in the upper left diagram, the annotation support device of the present invention can highlight point clouds near an annotation rectangle in a color different from that of the rectangle itself. This function makes it possible to intuitively grasp and prevent point clouds whose correspondence with the object is unclear, so-called "floating" or "sinking," and supports accurate positioning of the rectangle.

[0132] In addition, as shown in the upper right image, users can freely switch between parallel projection mode and perspective projection mode, enabling more intuitive annotation work by utilizing both geometric confirmation that is independent of viewpoint distortion (parallel projection) and highly visible confirmation that is closer to the actual field of view (perspective projection).

[0133] Furthermore, as shown in the bottom left image, a function is provided that allows users to filter the point cloud to be displayed. By specifying a spatial filter using a rectangular or cylindrical area, only the point cloud to be worked on can be extracted, eliminating visual noise and contributing to work efficiency.

[0134] The bottom right figure shows an example of mesh data of the ground. This mesh display makes it easy to check whether the installation surface of a rectangular annotation is consistent with the ground, and is particularly effective for verifying the accuracy of whether an object such as a car is installed on the ground surface. As described above, the various visualization, projection, filter, and mesh control functions performed by the second execution unit 265 greatly improve both the accuracy and efficiency of the annotation worker's work.

[0135] 19 shows an example of display control provided by the second execution unit 265 in the annotation support device, in which the "rotation gizmo" in the user interface is displayed independently of the magnification ratio. In the present invention, the rotation gizmo, which is used to perform operations such as angle adjustment on the annotation target, is scaled independently of the display magnification ratio so that it is always displayed at a size that is easy to see. As shown in the left figure, the gizmo is displayed at a sufficient size even for small objects when reduced in size, and is displayed at a similar size when enlarged in size as shown in the right figure, so that it does not become excessively large.

[0136] This dynamic scaling mechanism allows users to intuitively perform editing tasks such as rotation while maintaining consistent visibility and operability. This prevents users from losing track of the target to be edited, significantly improving work efficiency and accuracy, especially in annotation tasks involving high-density display of point cloud data and frequent viewpoint movement. This scaling-independent display of the rotation gizmo is one component of the user interface control realized by the second execution unit 265, and plays an important role in improving the operability of the annotation support device.

[0137] FIG. 20 shows an example of work efficiency improvement by image display control executed by the second execution unit 265 in the annotation support device, contrasting the control ON state (left diagram) with the control OFF state (right diagram). In the control ON state, the magnification of the corresponding camera image is automatically adjusted depending on the annotation target (in this example, a car). This makes it easier for the user to view and confirm the target, significantly improving the efficiency of the work of checking and correcting the position and shape of the rectangle.

[0138] On the other hand, when the control is OFF, the camera image is displayed at an arbitrary scale, and the annotation target may appear small within the image. In such a case, the user must manually perform the zoom operation, which takes time to check and correct, and may lead to operational errors. In this way, the automatic zoom control function provided by the second execution unit 265 optimizes the display of the camera image of the target, thereby visually and intuitively supporting the user's annotation work and improving the efficiency and accuracy of the work.

[0139] Figure 21 shows an example of a user interface (UI) for the annotation support device to streamline annotation and review work for time-series data. The UI shown in this figure supports the work for the time-series data (frame by frame) to be annotated using the following visual guides: Frames whose bounding box size is the reference frame are highlighted in dark blue on the timeline (the dark color in the image), allowing you to immediately identify the reference frame. Frames with annotations are displayed in light blue (a light color in the image) on the timeline, allowing you to intuitively understand the range of annotations. Frames where the worker has manually corrected the annotation are indicated with a "·" symbol, making it easy to see the correction history.

[0140] This display UI allows the progress of annotations in chronological order to be grasped at a glance, significantly improving the efficiency of checking differences from the reference frame and reviewing work. It also contributes to preventing the accumulation of errors and omissions. This UI function is controlled by the second execution unit 265 and is used in conjunction with the linked display of multiple views (Top / Back / Side), enabling more accurate chronological annotations while checking changes and deviations in three-dimensional space.

[0141] 22 shows an example of automatic focus control when moving frames based on time-series data in the annotation support device. The present invention provides a control mechanism (screen control by the second execution unit 265) that automatically displays a noteworthy object at the center of the screen when switching target frames during annotation or review work using time-series data, thereby reducing the effort required for the worker to move their viewpoint and improving work efficiency.

[0142] The left side of the figure shows an example of selecting and displaying a specific annotation target (car) in the 25th frame, while the right side shows how the target annotation is automatically centered on the screen when the frame is moved to the 31st frame. This control includes the following two functions: (1) Automatic focus control function centered on the shooting vehicle: The camera view is automatically adjusted when the frame is moved to maintain a viewpoint centered on the shooting vehicle (this function can be switched ON / OFF). (2) Automatic focus control function for the annotation target: When the frame is moved while an annotation target is selected, the target is automatically corrected so that it is always displayed in the center of the screen. These control functions improve the continuity and accuracy of annotation review and correction work over time, dramatically improving the efficiency of inspection and correction work, especially over long periods of time.

[0143] FIG. 23 is a diagram showing an example of a display UI for streamlining annotation and review work on point cloud data. In the annotation support device according to this embodiment, when visualizing an annotation target area in point cloud data, a color-coded display reflecting the distinction between annotation classes can be performed. Specifically, point clouds contained within an annotation box (rectangular parallelepiped) are clearly highlighted in the color of the class corresponding to the annotation. On the other hand, point clouds that exist outside the box but are nearby are displayed in a secondary color different from the annotation class color, facilitating boundary determination near the contour of the annotation box.

[0144] This visual support allows users to instantly see whether the point cloud is floating or sinking, or whether it contains incorrect ranges, greatly improving the accuracy and efficiency of annotation work. Furthermore, as shown in Figure 23, the annotation support device has the function of simultaneously displaying multiple views (Top, Back, Side), making it possible to perform annotation while intuitively grasping the three-dimensional layout relationship. This allows for precise annotation that appropriately considers the geometric relationship of the point cloud.

[0145] According to the present invention, by configuring the system to perform annotation by combining image data and point cloud data, it is possible to realize more precise and efficient annotation work while taking advantage of the characteristics of various sensing data. Furthermore, by configuring the system to also handle time-series data, it is possible to provide consistent annotations within consecutive scenes, and the system can be effectively applied to annotation work targeting videos, etc.

[0146] Furthermore, a mechanism for automatically mapping rectangular areas extracted from images onto point cloud data and annotating them eliminates the need for the cumbersome manual operations previously required, reducing the burden on workers. In addition, the system is equipped with an automated function that assists in the annotation of point clouds, significantly reducing the work time and stabilizing the quality of annotation.

[0147] By automatically performing a composite rule check (AutoQA) based on both images and point clouds on the annotation results, the effort required for human review can be significantly reduced, improving the reliability and reproducibility of the work. Furthermore, a collaborative work environment can be provided that enables multiple people to simultaneously annotate and review point cloud data, making it possible to build an efficient annotation pipeline for large-scale data. As described above, this invention enables highly accurate, efficient, and reliable annotation processing for high-dimensional data, including multimodal and time-series data. [Industrial Applicability]

[0148] The annotation support device, annotation support method, and annotation support program according to the present invention are applicable to annotation work on image data and point cloud data, and contribute to the advancement of 3D recognition and environmental understanding technologies in fields such as autonomous driving, robotics, smart cities, construction, logistics, and disaster prevention infrastructure. In particular, the present invention is configured to achieve highly accurate annotation processing by combining point cloud data acquired by sensors such as LiDAR with camera images, and will greatly contribute to improving the efficiency and quality of the process of creating training data for AI model training.

[0149] In addition, through features such as simultaneous collaboration by multiple workers, automatic annotation assistance, multi-modal AutoQA, and support for time-series data, annotation work on large amounts of data can be performed on an industrial scale in a scalable manner, contributing to improved productivity in AI development and verification, making it useful in a wide range of industrial fields.

[0150] As a result, this invention is extremely useful as a means of improving the efficiency and quality of the training data generation process in a variety of industrial fields where image recognition AI is being increasingly introduced, such as manufacturing, construction, agriculture, medicine, distribution and logistics, urban transportation, and security. [Explanation of symbols]

[0151] 1: Terminal 1-1~1-N: Terminal 2: Computer Systems 21: Input interface 22: Communication module 23: Storage device 24: Memory 25: Output interface 26: Processor 261: Recording Department 262: First output section 263: Second output section 264: First Executive Division 265: Second Executive Division 266: Third Executive Division 267: 4th Executive Division CN: communication line network

Claims

1. a recording unit that records setting information including annotation class information, calibration information, point cloud data, and image data, and annotation information; a first output unit that refers to the recording unit, displays the point cloud data, and enables annotation work to be performed; a second output unit that refers to the recording unit and displays a corresponding calibration image to support an annotation task; a first execution unit that uses the calibration information to accept a region designation by a user in the point cloud data displayed by the first output unit, and adds the target region to the image data based on the region designation; a second execution unit that executes various screen control processes related to the display of the first output unit and the second output unit; a third execution unit that refers to the recording unit and the first execution unit, and executes an automatic quality assurance process that supports review and quality confirmation of annotation results using an AI model including one or more of an object detection model, a region detection model, a visual language model, and a large-scale language model; and a fourth execution unit that executes annotation processing for adding a target area to the image data and correction processing for a rectangle on the point cloud data; Equipped with The first execution unit is an annotation support device that performs a process of synchronizing point cloud data and image data.

2. The recording unit records data at multiple points in time including time-series point cloud data and image data. The annotation support device according to claim 1 .

3. the first output unit displays an annotation target area based on point cloud data and provides a user interface for accepting a correction operation in order to support an annotation work on the object; The annotation support device according to claim 1 or 2.

4. the second output unit switches between or simultaneously displays calibration images and annotation data corresponding to a multiple image display according to the number of camera sensors, and provides a user interface that allows a user to visually confirm consistency. The annotation support device according to claim 1 or 2.

5. the third execution unit executes a review function to evaluate omissions or consistency of the object based on the recorded annotation information; The annotation support device according to claim 1 or 2.

6. the fourth execution unit, when an object exists in the image data, specifies an area including at least one of a rectangle and a point by user input, or specifies an area by inputting annotation data, or extracts the object area as a rectangle using a pre-trained model, associates the extracted area with point cloud data, and displays it on the output unit or records it on the recording unit; the fourth execution unit, when manually or automatically adding a rectangle to the point cloud data, realizes a highly accurate initial annotation by automatically correcting the orientation and size of the rectangle using at least one of calibration information and point cloud information; The annotation support device according to claim 1 or 2.

7. The annotation support device according to claim 1 , wherein the second execution unit highlights point clouds inside the annotation area and outside the area adjacent to the annotation area, thereby improving the efficiency of a rectangle correction operation.

8. The annotation support device according to claim 1 or 2, wherein the second execution unit makes annotation work more efficient by switching between parallel projection and perspective projection, filtering the displayed point cloud using a rectangular prism or a cylinder, and embodying point cloud information using a mesh display.

9. 3. The annotation support device according to claim 1, wherein the first output unit is capable of displaying point cloud annotation information using sub-views from each of the x, y and z directions in addition to a main view to support annotation work on an object, and provides a user interface that accepts correction operations in each view.

10. recording setting information and annotation information, including annotation class information, calibration information, point cloud data, and image data; a step of referencing the recorded setting information, displaying the point cloud data, and enabling annotation work; a step of referring to the recorded setting information and displaying a corresponding calibration image to support an annotation task; executing a synchronization process based on a correspondence relationship between the point cloud data and the image data; executing various screen control processes related to the display of the point cloud data and the image data; performing an automated quality assurance process that utilizes AI models, including one or more of an object detection model, a region detection model, a visual language model, and a large-scale language model, to assist in reviewing and quality checking the annotation results; a step of performing annotation processing to add a target area to the image data and correction processing for a rectangle on the point cloud data; An annotation assistance method comprising:

11. recording data at multiple points in time including time series point cloud data and image data; executing a process for synchronizing the point cloud data and the image data based on the recorded time series; The annotation assistance method according to claim 10, comprising:

12. providing a user interface that displays an annotation target area based on point cloud data and accepts a correction operation, in order to support annotation work on the object; The annotation support method according to claim 10 or 11.

13. a step of providing a user interface that switches between or simultaneously displays calibration images and annotation data corresponding to a multiple image display according to the number of camera sensors, and enables a user to visually confirm consistency; The annotation support method according to claim 10 or 11.

14. performing a review function to evaluate the object for omissions or consistency based on the recorded annotation information; The annotation support method according to claim 10 or 11.

15. When an object exists in the image data, the method includes a step of specifying an area including at least one of a rectangle and a point by user input, or specifying an area by inputting annotation data, or extracting the object area as a rectangle using a pre-trained model, and displaying or recording the extracted area in association with point cloud data, 12. The annotation support method according to claim 10 or 11, wherein when a rectangle is manually or automatically added to point cloud data, highly accurate initial annotation is achieved by automatically correcting the orientation and size of the rectangle using at least one of calibration information and point cloud information.

16. a step of recording annotation work logs by a plurality of users and displaying each work history in displaying the point cloud data and the calibration image; The annotation support method according to claim 10 or 11.

17. The annotation support method according to claim 10 or 11, wherein the step of executing various screen control processes highlights point clouds inside the annotation area and outside the area in the vicinity thereof, thereby improving the efficiency of a rectangle correction operation.

18. The annotation support method according to claim 10 or 11, wherein the step of executing the various screen control processes makes annotation work more efficient by switching between parallel projection and perspective projection, filtering the displayed point cloud using a rectangular prism or a cylinder, and embodying point cloud information using a mesh display.

19. 12. The annotation support method according to claim 10 or 11, wherein the step of displaying the point cloud data and enabling annotation work to be performed provides a user interface that can display point cloud annotation information using sub-views from each of x, y and z directions in addition to a main view to support annotation work on an object, and accepts correction operations in each view.

20. An annotation support program for causing a computer to execute the annotation support method according to claim 10 or 11.

Citation Information

Patent Citations

  • Image processing device, image processing method, and recording medium

    WO2020179065A1

  • Point cloud annotation device, method and program

    WO2020225889A1

  • Recognition device, method, and program

    WO2025120694A1

  • Information processing device and annotation program

    JP2025017225A

  • Label assignment assistance device, label assignment assistance method, and program

    WO2022185363A1