Information processing device, information processing method, and non-transitory computer-readable storage medium
The proposed information processing apparatus integrates 3D annotation and 2D region division to address inefficiencies in creating high-quality learning data for image recognition, improving data quality and annotation consistency.
Patent Information
- Application Number
- PCT/JP2024/043506
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-19
- Filing Date
- 2024-12-10
- Publication Date
- 2025-07-24
AI Technical Summary
Existing image recognition technologies face challenges in efficiently creating high-quality learning data for pixel-level classification tasks due to shape errors in 3D models and the need for manual annotation of large datasets, leading to inefficiencies and inconsistent annotation results.
An information processing apparatus and method that acquires camera pose information, integrates 3D annotation and 2D region division information to streamline the annotation process, reducing shape errors and improving data quality by using a control unit to integrate annotation information and region division results.
Facilitates the easy acquisition of high-quality learning data by reducing shape errors and unifying annotation results across multiple views, enhancing the efficiency and accuracy of image recognition tasks.
Smart Images

Figure JP2024043506_24072025_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and computer-readable non-transitory storage medium
[0001] The present disclosure relates to an information processing device, an information processing method, and a computer-readable non-transitory storage medium.
[0002] In recent years, advances in artificial intelligence (AI) technology have led to active research and application of image processing tasks that involve pixel-level classification. For example, in the field of autonomous driving, heuristic classification of roadways, sidewalks, oncoming vehicles, pedestrians, etc., is useful for estimating the drivable area of a vehicle. In addition, some high-performance smartphone camera application software is equipped with a function that applies the aforementioned pixel-level classification to extract only areas such as people captured in camera images.
[0003] Such applications are implemented using, for example, deep learning, but require training of an image recognizer (neural network) using a large amount of training data in advance. However, the annotation process of creating training data classified into classes on a pixel-by-pixel basis by adding annotation information such as class labels to each pixel takes a long time.
[0004] For example, in-vehicle video and video of a person have significantly different distances from the subject and the classes of objects and backgrounds that appear, so separate training data must be prepared for each task. Also, even though a group of images taken in the same environment contains images with almost identical angles of view, each image must be annotated independently.
[0005] To address this issue, a technique has been proposed for easily creating learning data, which combines a computer-aided design (CAD) model of a known physical object with an augmented reality (AR) camera (see, for example, Patent Document 1). This technique acquires a three-dimensional (3D) model of a physical object placed in a scene, generates a virtual object corresponding to the object based on the model, and substantially overlays the virtual object on the physical object within the field of view of the AR camera, thereby improving the efficiency of annotation work.
[0006] As another method, a technique has been proposed that reduces the cost of annotation work by using an automatically segmented 3D model (see, for example, Non-Patent Document 1).
[0007] Japanese Patent Application Laid-Open No. 2020-087440
[0008] Angela Dai and five others, "ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes," Stanford University, Princeton University, Technical University of Munich, [online], [Retrieved January 15, 2024], Internet <URL: https: / / openaccess.thecvf.com / content_cvpr_2017 / papers / Dai_ScanNet_Richly-Annotated_3D_CVPR_2017_paper.pdf>
[0009] However, the above-described conventional techniques have room for further improvement in terms of enabling high-quality learning data to be easily obtained using 3D data such as 3D models.
[0010] For example, when using a 3D model representing real space as in the above-mentioned conventional technology, it is currently difficult to prepare a 3D model that is completely free of shape errors for each physical object in real space. Therefore, when using the conventional technology, there is a problem in that shape errors in the 3D model are reflected in the quality of the training data.
[0011] Therefore, the present disclosure proposes an information processing device, an information processing method, and a computer-readable non-transitory storage medium that enable high-quality learning data to be easily obtained while using 3D data.
[0012] In order to solve the above problem, an information processing device according to one embodiment of the present disclosure includes a control unit. The control unit acquires camera posture information for presenting 3D data representing a real space. The control unit also acquires annotation information for the 3D data corresponding to the camera posture information. The control unit also acquires 2D data representing the real space corresponding to the camera posture information. The control unit also acquires 2D region division information by dividing the 2D data into a plurality of regions. The control unit also acquires integrated annotation information by integrating the annotation information and the 2D region division information.
[0013] FIG. 1 is a diagram illustrating an overview of an image processing method according to an embodiment of the present disclosure. FIG. 2 is a block diagram illustrating an example configuration of a learning device and an image processing device according to an embodiment of the present disclosure. FIG. 3 is a diagram illustrating an example of a multi-view image DB. FIG. 4 is a flowchart illustrating a processing procedure executed by a learning device according to an embodiment of the present disclosure. FIG. 5 is a flowchart illustrating a processing procedure of integration processing. FIG. 6 is a diagram illustrating an overview of an image processing method according to a modified example. FIG. 7 is a diagram illustrating an example operation using a 2D viewer. FIG. 8 is a block diagram illustrating an example configuration of a 3D annotation unit and 3D annotation information. FIG. 9 is a flowchart illustrating a processing procedure of 3D annotation processing executed by the 3D annotation unit. FIG. 10 is a diagram illustrating the processing procedure of face detection and separation processing. FIG. 11 is a diagram illustrating the processing procedure of camera position calculation processing. FIG. 12 is a diagram illustrating the effect of an annotation work support method according to an embodiment of the present disclosure (part 1). FIG. 13 is a diagram illustrating the effect of an annotation work support method according to an embodiment of the present disclosure (part 2). FIG. 14 is a hardware configuration diagram illustrating an example of a computer that realizes the functions of an image processing device.
[0014] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted.
[0015] In the following description, an information processing device according to an embodiment of the present disclosure (hereinafter referred to as "the present embodiment") is assumed to be an image processing device 20 configured as a part of the learning device 10 shown in Figure 2 and subsequent drawings. Furthermore, an information processing method according to the present embodiment is assumed to be an image processing method executed by the image processing device 20.
[0016] The present disclosure will be described in the following order: 1. Overview 1-1. Issues with existing technology 1-2. Overview of image processing method according to embodiment of the present disclosure 2. Configuration examples of learning device and image processing device 3. Processing procedure 3-1. Processing procedure executed by learning device 3-2. Processing procedure of integration processing 4. Image processing method according to modified example 5. Annotation work support method according to embodiment of the present disclosure 5-1. Issues with existing technology 5-2. Configuration examples of 3D annotation unit and 3D annotation information 5-3. Processing procedure 5-4. Effects of annotation work support method according to embodiment of the present disclosure 6. Other modified examples 7. Hardware configuration 8. Conclusion
[0017] <<1. Overview>> <1-1. Issues with Existing Technology> Prior to describing the image processing method according to this embodiment, issues with existing technology will first be described in more detail. In recent years, image recognition processing using image recognizers that utilize AI technology has become increasingly widespread, but image recognizers must be trained in advance using large amounts of training data. However, the annotation process of creating training data that is classified into classes on a pixel-by-pixel basis by assigning annotation information such as class labels to each pixel takes a significant amount of time.
[0018] To address this issue, there is an existing technology that combines a CAD model of a known physical object with an AR camera. This existing technology acquires a 3D model of the physical object placed in a scene, generates a virtual object corresponding to the object based on the model, and virtually overlays the virtual object on the physical object within the field of view of the AR camera, thereby improving the efficiency of the annotation work.
[0019] However, when using this existing technology, if there are shape errors in the 3D models of physical objects that can be placed in a scene, the errors are rendered as they are in the two-dimensional (2D) images that serve as training data. Furthermore, preparing CAD models corresponding to all physical objects in a given scene requires a very high operational cost, and in order to apply the models to classification tasks, annotation information such as class labels must be added to the models in advance.
[0020] A different, more general-purpose approach exists: annotating 2D images using a combination of manual and machine learning models. This technology involves manually specifying sparse annotation information, such as the center position of the annotation target, and then inputting this sparse annotation information into a machine learning model to convert it into dense, pixel-level annotation information. While manual annotation work can take more than an hour per image, this technology has the advantage of significantly reducing the time required.
[0021] However, when annotating images individually, it is wasteful to annotate overlapping angles of view in multiple images separately.In addition, the inference results of the image recognizer may differ depending on how the same object is viewed, making it impossible to spatially unify the annotation results.
[0022] Another existing technique reduces the cost of annotation by using automatically segmented 3D models. This technique automatically segments a 3D model of an indoor environment into a large number of small parts, and then the user assigns a class label to each part to generate a class-labeled 3D model. This model is then rendered onto the image plane of 2D images that serve as training data, generating a training dataset for pixel-level classification. This existing technique has the advantage that annotating the 3D model eliminates the need to annotate each image independently.
[0023] However, this existing technology renders 3D models with geometric errors as they are, resulting in misalignment between the region boundaries in the image plane of the 2D image and the region boundaries defined by annotations on the 3D model. It is difficult to accurately train pixel-level classification tasks using training datasets adversely affected by such geometric errors.
[0024] <1-2. Overview of Image Processing Method According to an Embodiment of the Present Disclosure> In order to overcome these problems with existing technology and easily obtain high-quality learning data while using 3D data, in the image processing method according to the present embodiment, an image processing device 20 acquires camera posture information for presenting 3D data representing a real space. The image processing device 20 also acquires annotation information for the 3D data corresponding to the camera posture information. The image processing device 20 also acquires 2D data representing the real space corresponding to the camera posture information. The image processing device 20 also acquires 2D region division information by dividing the 2D data into a plurality of regions. The image processing device 20 also acquires integrated annotation information by integrating the annotation information and the 2D region division information.
[0025] FIG. 1 is a diagram illustrating an overview of an image processing method according to an embodiment of the present disclosure. As shown in FIG. 1, in the image processing method according to this embodiment, an image processing device 20 first receives, as input, multi-viewpoint images obtained by capturing a 3D space to be classified from multiple viewpoints. The multi-viewpoint images are acquired, for example, based on image data captured by a user of the 3D space to be classified. The multi-viewpoint images are a group of 2D images captured from various positions of the 3D space to be annotated, and each image includes camera posture information corresponding to the image. The 2D images are RGB images.
[0026] The image processing device 20 then acquires a 3D model corresponding to the acquired multi-viewpoint images (step S1). Possible methods for acquiring this 3D model include 3D reconstruction using the input 2D images and the camera posture information, or 3D reconstruction using depth images corresponding to the 2D images. Another possible acquisition method is to match the position information of an existing 3D model, such as a known CAD model or a model on a geographic information system (e.g., PLATEAU, hosted by the Ministry of Land, Infrastructure, Transport and Tourism), with the camera postures of the multi-viewpoint images.
[0027] The image processing device 20 then executes a 3D annotation process to add annotation information to the acquired 3D model (step S2). Possible methods for this annotation include manual annotation in 3D structural units such as points, surfaces, and voxels, annotation using a machine learning model, and a combination of manual and machine learning methods. In this embodiment, a method is proposed to support manual annotation work by classifying 3D models into planar objects and non-planar objects and presenting these objects to the user at an optimal camera position that makes it easy for the user to annotate them. Details of this method will be described later using Figures 8 to 16.
[0028] Next, the image processing device 20 performs a rendering process to render the annotation information assigned to the 3D model on the image plane corresponding to each image of the 2D image group of multi-viewpoint images (step S3), thereby obtaining a rendering result that includes shape errors of the 3D model.
[0029] In parallel with steps S2 and S3, the image processing device 20 performs a region segmentation process for segmenting each image in the group of 2D multi-viewpoint images using a machine learning model (step S4). This allows for a region segmentation result that does not include annotation information but has highly accurate region boundaries.
[0030] Finally, the image processing device 20 executes an integration process (step S5) to integrate the rendering result of step S3 with the region segmentation result of step S4. In the integration process, regions included in the region segmentation result are selected in descending order of area, and the most frequent value of the rendering result corresponding to the selected region is adopted as annotation information to be assigned to the selected region. Details of this integration process will be described later using FIGS. 5 and 6.
[0031] The image processing device 20 then outputs the integration result of this integration process (hereinafter referred to as "integrated annotation information" as appropriate), and the aforementioned learning device 10 executes a learning process using the group of 2D images classified by this integrated annotation information as a learning dataset.
[0032] As described above, in the image processing method according to this embodiment, the image processing device 20 acquires camera posture information for presenting 3D data representing a real space. The image processing device 20 also acquires annotation information for the 3D data corresponding to the camera posture information. The image processing device 20 also acquires 2D data representing the real space corresponding to the camera posture information. The image processing device 20 also acquires 2D region division information by dividing the 2D data into a plurality of regions. The image processing device 20 also acquires integrated annotation information by integrating the annotation information and the 2D region division information.
[0033] Therefore, according to the image processing method of this embodiment, it is possible to easily obtain high-quality learning data while using 3D data. Below, a more specific description will be given of example configurations of the learning device 10 and the image processing device 20 to which the image processing method of this embodiment is applied.
[0034] 2. Configuration Examples of Learning Device and Image Processing Device Fig. 2 is a block diagram showing configuration examples of the learning device 10 and image processing device 20 according to an embodiment of the present disclosure. Note that Fig. 2 and Fig. 9 shown later show functional blocks of only components necessary for explaining this embodiment, and descriptions of general components are omitted.
[0035] In addition, in the description using FIGS. 2 and 9, the description of components that have already been described will be appropriately simplified or omitted.
[0036] As shown in Figure 2, study device 10 includes a memory unit 11 and a control unit 12. Study device 10 is also connected to an operation unit 3 and a display unit 5. Operation unit 3 is a component that accepts input operations from a user. Operation unit 3 is realized by, for example, a keyboard, a pointing device, or the like.
[0037] Display unit 5 is a component that displays visual information such as images output from study device 10. Display unit 5 is realized by a display or the like. Note that operation unit 3 and display unit 5 may be realized as a single unit, for example, by a touch panel display or the like.
[0038] The storage unit 11 is realized by a storage device such as a random access memory (RAM), a read only memory (ROM), a flash memory, or a hard disk drive (HDD).
[0039] In the example of Figure 2, the memory unit 11 stores a multi-viewpoint image DB (Database) 11a, a 3D model 11b, 3D annotation information 11c, 2D rendering information 11d, an area division model 11e, 2D area division information 11f, integrated annotation information 11g, and a recognition model 11h.
[0040] The multi-viewpoint image DB 11a is a database that stores multi-viewpoint images acquired by the acquisition unit 12a, which will be described later. FIG. 3 is a diagram showing an example of the multi-viewpoint image DB 11a. As shown in FIG. 3, the multi-viewpoint image DB 11a individually stores each of the 2D images -1, -2, -3, ... in the acquired multi-viewpoint images. Each of the 2D images -1, -2, -3, ... is associated with and stored with camera position and angle -1, -2, -3, ..., which are corresponding camera posture information. In other words, the camera posture information includes the camera position and angle corresponding to each 2D image.
[0041] Returning to the description of Fig. 2, the 3D model 11b is a model of the 3D space to be subjected to classification. The 3D model 11b is acquired by the acquisition unit 12a (described later) using the various acquisition methods in step S1 described above.
[0042] The 3D annotation information 11c stores the annotation results from the 3D annotation process in step S2 described above. The 3D annotation process is executed by a 3D annotation unit 12b, which will be described later.
[0043] The 2D rendering information 11d stores the rendering results of the rendering process in step S3 described above. The rendering process is executed by a rendering unit 12c, which will be described later.
[0044] The region division model 11e is a machine learning model used by the region division unit 12d (described later) when executing the region division process of step S4 described above. The region division model 11e is, for example, a deep neural network (DNN) model trained for region division purposes using a deep learning algorithm. The 2D region division information 11f stores the region division results obtained by the region division process of step S4 described above.
[0045] The integrated annotation information 11g stores the integration results of the integration process in step S5 described above. The integration process is executed by an integration unit 12e, which will be described later.
[0046] The recognition model 11h is an image recognizer that is trained using a group of class-labeled 2D images included in the integrated annotation information 11g as a training data set. In this embodiment, the recognition model 11h is a DNN model based on a deep learning algorithm, but the learning algorithm of the image recognizer is not limited thereto.
[0047] The control unit 12 controls each unit of the learning device 10. The control unit 12 is realized by, for example, a central processing unit (CPU), a micro processing unit (MPU), a graphical processing unit (GPU), or the like executing a program according to this embodiment (not shown) stored in the storage unit 11 using RAM as a work area. The control unit 12 can also be realized by, for example, an integrated circuit such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA).
[0048] The control unit 12 has an acquisition unit 12a, a 3D annotation unit 12b, a rendering unit 12c, an area division unit 12d, an integration unit 12e, a display control unit 12f, and a learning unit 12g, and realizes or executes the functions and actions of information processing described below.
[0049] The acquisition unit 12a acquires the multi-viewpoint images and the camera posture information of each 2D image in the multi-viewpoint images via a network, a recording medium, etc. (not shown) The acquisition unit 12a also stores the acquired 2D images and the corresponding camera posture information in the multi-viewpoint image DB 11a.
[0050] The acquisition unit 12a also acquires and stores the 3D model 11b. For example, the acquisition unit 12a acquires the 3D model 11b by generating the 3D model 11b through 3D reconstruction using the acquired 2D image group and the camera posture information.
[0051] The 3D annotation unit 12b assigns annotation information specified by the user to the 3D model 11b via the operation unit 3. For example, the 3D annotation unit 12b displays a 2D image of the 3D model 11b viewed from an arbitrary position on the display unit 5. When the user selects an arbitrary object and assigns a class label to the 2D image displayed on the display unit 5 via the operation unit 3, the 3D annotation unit 12b assigns the information specified by the user as annotation information. Then, the 3D annotation unit 12b projects the annotation information assigned to the 2D image displayed on the display unit 5 onto the 3D model 11b and stores the projected annotation result in the 3D annotation information 11c.
[0052] The rendering unit 12c renders the 3D annotation information 11c on the image plane of each 2D image in the multi-viewpoint image in accordance with the camera posture of each image, and stores the rendering results in the 2D rendering information 11d.
[0053] The region dividing unit 12d divides each image in the 2D image group of the multi-viewpoint images into regions using the region division model 11e in parallel with the 3D annotation process by the 3D annotation unit 12b and the rendering process by the rendering unit 12c. The region dividing unit 12d also stores the region division results in 2D region division information 11f.
[0054] The integrating unit 12e executes an integration process to integrate the 2D rendering information 11d and the 2D region division information 11f. The integrating unit 12e selects each region included in the 2D region division information 11f in descending order of area, and adopts the most frequent value in the 2D rendering information 11d corresponding to the selected region as the annotation information to be assigned to the selected region. The integrating unit 12e also stores the integration result in the integrated annotation information 11g.
[0055] The display control unit 12f appropriately displays visual information such as each 2D image to be annotated, the annotation result, the rendering result, the region segmentation result, and the integration result on the display unit 5. The learning unit 12g learns the recognition model 11h by deep learning using the group of class-labeled 2D images stored in the integrated annotation information 11g as a learning dataset for class classification.
[0056] Note that Figure 2 shows an example in which the image processing device 20 is configured as part of the learning device 10, but the image processing device 20 can be realized as a separate device independent of the learning device 10 by having each unit and memory information other than the learning unit 12g as components.
[0057] <<3. Processing Procedure>> <3-1. Processing Procedure Executed by the Learning Device> Next, the processing procedure executed by the learning device 10 will be described with reference to Fig. 4. Fig. 4 is a flowchart showing the processing procedure executed by the learning device 10 according to an embodiment of the present disclosure.
[0058] As shown in FIG. 4, the control unit 12 of the learning device 10 first inputs a group of 2D images in the multi-viewpoint image and camera posture information corresponding to each image in the group of 2D images (step S101).
[0059] The control unit 12 then acquires the 3D model 11b (step S102) and assigns annotation information to the 3D model 11b (step S103). Subsequently, the control unit 12 renders the annotation information assigned to the 3D model 11b onto an image plane corresponding to each image in the group of 2D images of the multi-viewpoint images (step S104).
[0060] Meanwhile, in parallel with steps S102 to S104, the control unit 12 divides each image in the 2D image group into regions (step S105).
[0061] The control unit 12 then performs an integration process (step S106). The control unit 12 then uses the integrated annotation information 11g, which is the integration result of the integration process, i.e., the group of 2D images with class labels, as a training dataset to train the recognition model 11h (step S107). The process then ends.
[0062] 3-2. Integration Process Procedure Next, the integration process procedure of step S106 will be described with reference to Fig. 5 and Fig. 6. Fig. 5 is a flowchart showing the integration process procedure. Fig. 6 is an explanatory diagram of the integration process.
[0063] In the integration process, as shown in FIG. 5, the control unit 12 receives the 2D rendering information 11d and the 2D region division information 11f (step S201).
[0064] The control unit 12 then selects the largest unintegrated region in the 2D region division information 11f (step S202). Next, the control unit 12 obtains rendering results for all pixels included in the selected region from the 2D rendering information 11d (step S203).
[0065] Then, the control unit 12 adopts the most frequent value of the obtained rendering result as annotation information for the selected region (step S204).
[0066] When the integration process for one selected region is completed, the control unit 12 determines whether or not there is an unintegrated region in the 2D region division information 11f (step S205). If there is an unintegrated region (step S205, Yes), the control unit 12 repeats the process from step S202.
[0067] If there are no unintegrated regions (No at step S205), the control unit 12 outputs integrated annotation information 11g, which is the integration result for all regions in the 2D region division information 11f, and ends the integration process.
[0068] 6 shows a schematic diagram of the integration process. The upper left shows the result of segmenting the 2D image using the 2D region segmentation information 11f, and to the right of that shows the result of rendering the annotation information assigned to the 3D model 11b onto the image plane corresponding to the 2D image.
[0069] The segmentation result does not contain any annotation information, and what should be a single chair has been over-segmented into regions showing many parts. Furthermore, while the rendering result includes annotation information for the 3D model 11b, it is distorted due to shape errors in the 3D model 11b.
[0070] In this example, the control unit 12 first selects the largest area of the chair back from the segmentation results, acquires the annotation information of the corresponding rendering result, and then, if the most frequent value of this annotation information is "chair," integrates the annotation information of the entire chair back area into "chair."
[0071] Next, the control unit 12 selects the chair seat area, which has the largest area among the unintegrated areas, and similarly assigns the annotation information of "chair." The control unit 12 then repeats this process for the areas of each leg of the chair, thereby finally obtaining an integrated result with annotation information that eliminates shape errors of the 3D model 11b.
[0072] 4. Image Processing Method According to a Modified Example Next, an image processing method according to a modified example will be described. Fig. 7 is a diagram illustrating an outline of the image processing method according to the modified example.
[0073] In the image processing method according to the modified example, even for camera poses for which multi-viewpoint images have not been acquired in advance, it is possible to generate learning data by combining the method with existing technology for synthesizing new viewpoint images.
[0074] Specifically, in the image processing method according to the modified example, a machine learning model such as Neural Radiance Fields (NeRF) is prepared that uses multi-viewpoint images acquired in advance to synthesize new viewpoint images (see "NeRF model" in the figure).
[0075] Then, using this model, a 2D image is synthesized for an arbitrary camera posture (see "new viewpoint image" in the figure), and the region dividing unit 12d divides this image into regions.
[0076] The rendering unit 12c also renders annotation information of the 3D model 11b for a camera pose similar to that of the new viewpoint image. The integrating unit 12e then integrates the rendering result with the region segmentation result of the new viewpoint image, thereby generating learning data for a camera pose for which no multi-viewpoint image has been acquired in advance.
[0077] <<5. Annotation Work Support Method According to an Embodiment of the Present Disclosure>> <5-1. Issues with Existing Technology> When manually annotating 3D data such as the above-described 3D model 11b, the annotation is generally performed via a 2D viewer that displays the 3D data as a 2D image. In this case, the user segments the object through the 2D viewer and assigns annotation information such as a class label to each segment, but the user must perform the segmentation by moving the camera position to prevent occlusion from occurring.
[0078] 8 is a diagram showing an example of operation using a 2D viewer. In this example, the camera position needs to be moved twice. First, the user uses the 2D viewer to move to camera position A where there is no occlusion, and performs segmentation of the object (here, a cylinder) to be annotated (step S11).
[0079] In step S11, no matter what camera position is moved to, there are other objects including wall surfaces before and after the cylinder to be annotated, and complete segmentation cannot be performed. In the example of Figure 8, in addition to the cylinder, part of the wall surface is cut out.
[0080] Therefore, the user operates the 3D data segmented in step S11 from another camera position B (step S12), thereby completing complete segmentation. Note that while annotation of simple 3D data has been described here, with complex 3D data such as a room, occlusion occurs due to objects such as walls, floors, and ceilings, and therefore more complex camera position operations are required.
[0081] There is an existing technology that detects and removes walls, floors, and ceilings from 3D point clouds in 3D data of a room. This existing technology first identifies the XY planes parallel to the ceiling and floor of the point cloud. Next, the space is sliced parallel to the identified XY planes, and high-density areas are identified as the ceiling and floor. The same process is performed on the XZ and YZ planes to identify walls.
[0082] However, this existing technology cannot detect planar objects other than walls, floors, and ceilings. A room is filled with a variety of furniture, including various planar objects that can cause occlusions. Examples of such objects include tabletops and shelf sides. For example, if a chair is positioned so that at least a portion of it is hidden below the tabletop, segmentation that separates the desk and chair from each other is not possible, regardless of the camera position.
[0083] Furthermore, there may be various planar objects in a room that are in contact with the walls, floor, or ceiling, such as a smartphone or remote control placed on a table, a photo hanging on the wall, or a carpet on the floor. If existing technology is used that simply deletes the walls, floor, and ceiling, annotation for these objects cannot be performed.
[0084] Furthermore, while the above-mentioned walls, floors, ceilings, table tops, etc. are planar objects, indoors occlusion can occur due to non-planar objects such as lighting fixtures hanging from the ceiling. Existing technologies cannot handle occlusion due to such non-planar objects.
[0085] Therefore, in the annotation work support method according to this embodiment, the 3D annotation unit 12b classifies the 3D model 11b into at least one planar object and multiple non-planar objects, and acquires annotation information representing the multiple non-planar objects. Below, we will explain in more detail exemplary configurations of the 3D annotation unit 12b and the 3D annotation information 11c to which the annotation work support method according to this embodiment is applied.
[0086] 9 is a block diagram showing an example of the configuration of the 3D annotation portion 12b and the 3D annotation information 11c. Note that Fig. 9 corresponds to Fig. 2, and differences from Fig. 2 will be mainly described here.
[0087] 9, the 3D annotation unit 12b includes a surface detection / separation unit 12ba and a camera position calculation unit 12bb. The 3D annotation information 11c is stored with planar objects and non-planar objects separated from each other.
[0088] The surface detection and separation unit 12ba inputs the 3D point cloud of the 3D model 11b, classifies all of the point cloud into planar objects and non-planar objects, and performs surface detection and separation processing to separate the objects based on this classification. The term "separation" here includes enabling changes to the positional relationship between planar objects and non-planar objects. Changing the positional relationship includes changing the camera position corresponding to each object.
[0089] Specifically, the surface detection and separation unit 12ba receives the 3D point cloud of the 3D model 11b, temporarily stores all of the point cloud as non-planar objects, and then detects and stores planar objects by solving a plane equation from the sampled point cloud using the least squares method or the like while repeating the point cloud sampling a predetermined number of times, and then deletes the point cloud that has been determined to be a planar object from the non-planar object.
[0090] The camera position calculation unit 12bb calculates an optimal camera position for a planar object and a non-planar object to be annotated, making it easy for the user to annotate. Specifically, the camera position calculation unit 12bb inputs a 3D point cloud of each object to be annotated, and calculates a sphere that contains all of the point clouds from the center of the point cloud.
[0091] The camera position calculation unit 12bb also arranges multiple virtual cameras at equal intervals on the calculated spherical surface, converts the point cloud from a plane tangent to the spherical surface at each camera position into 2D coordinates by parallel projection, and calculates covariance from the 2D coordinates.The camera position calculation unit 12bb then selects the camera position that maximizes the covariance as the optimal camera position.
[0092] The display control unit 12f presents the object to be annotated at the optimal camera position calculated by the camera position calculation unit 12bb on the 2D viewer displayed on the display unit 5 during manual annotation work.
[0093] 5-3. Processing Procedure Next, more specific processing procedures executed by the 3D annotation unit 12b will be described with reference to Fig. 10 to Fig. 14. Fig. 10 is a flowchart showing the processing procedures of the 3D annotation process executed by the 3D annotation unit 12b. Fig. 11 is an explanatory diagram of the 3D annotation process.
[0094] As shown in FIG. 10, the 3D annotation unit 12b first inputs the 3D model 11b (step S301), and causes the surface detection and separation unit 12ba to execute surface detection and separation processing based on the input 3D model 11b (step S302).
[0095] Then, the 3D annotation unit 12b inputs the planar objects and non-planar objects classified and separated by the surface detection and separation unit 12ba (step S303), and determines whether there are any unlabeled objects that have not been assigned a class label for each of these objects (step S304).
[0096] If there is an unlabeled object (Yes at step S304), it is determined whether the object is a non-planar object or a planar object (step S305).
[0097] If the object is a non-planar object (step S305, non-planar), the 3D annotation unit 12b performs clustering on each non-planar object (step S306). In this process, the 3D annotation unit 12b divides the non-planar object into several clusters based on its geometric structure.
[0098] The 3D annotation unit 12b then accepts a cluster selection by the user via the operation unit 3 (step S307) and deletes clusters other than the selected cluster (step S308). Note that the "deletion" here is an example of separating multiple non-planar objects.
[0099] Here, "separation" includes changing the positional relationship of multiple non-planar objects. The change in the positional relationship includes changing the camera position. Alternatively, the change in the positional relationship includes hiding all clusters other than the one selected by the user on the 2D viewer displayed on the display unit 5.
[0100] On the other hand, if the object is a plane object (step S305, plane), the process proceeds to step S309. In step S309, the 3D annotation unit 12b causes the camera position calculation unit 12bb to execute a camera position calculation process (step S309).
[0101] The 3D annotation unit 12b then causes the display control unit 12f to present the object at the optimal camera position calculated in the camera position calculation process (step S310).The 3D annotation unit 12b then accepts segmentation and labeling by the user via the operation unit 3 (step S311).
[0102] Then, the 3D annotation unit 12b updates the 3D annotation information 11c with the content received in step S311 (step S312), and repeats the process from step S304.
[0103] If there are no unlabeled objects in step S304 (step S304, No), the 3D annotation unit 12b outputs and saves the 3D annotation information 11c, and ends the 3D annotation process.
[0104] Fig. 11 is a schematic diagram illustrating the flow of the processing procedure in Fig. 10. In the example of Fig. 11, the 3D annotation unit 12b inputs a 3D model 11b representing a certain real space, and detects Wall A, Wall B, the floor, and the table top as planar objects by the surface detection and separation process in step S302. The 3D annotation unit 12b also detects chairs and table legs as non-planar objects.
[0105] When a planar object is to be annotated, an optimal camera position is calculated by the camera position calculation process in step S309, and in the example of Fig. 11, for example, wall B is presented at a camera position where the user views it from the front (step S310). At this time, non-planar objects separated from the planar object are hidden.
[0106] On the other hand, when a non-planar object is set as the annotation target, the planar object separated from the non-planar object is hidden. Then, the non-planar object is clustered in step S306 and divided into several clusters (here, two clusters: a group of chairs and table legs) based on its geometric structure. Then, when a cluster selection is accepted in step S307, all clusters other than the selected cluster (here, chairs) are hidden in step S308.
[0107] Then, the optimal camera position is calculated by the camera position calculation process in step S309, and in the example of FIG. 11, a cluster corresponding to a group of chairs is presented at a camera position where there is no occlusion by the chairs (step S310).
[0108] Next, the processing procedure of the face detection and separation processing in step S302 will be described with reference to Fig. 12. Fig. 12 is a flowchart showing the processing procedure of the face detection and separation processing.
[0109] In the surface detection and separation process, as shown in FIG. 12, the surface detection and separation unit 12ba receives the 3D point cloud of the 3D model 11b (step S401).
[0110] The surface detection and separation unit 12ba then stores all of the point clouds as non-planar objects (step S402), and then determines whether the number of times the point clouds have been sampled is equal to or less than a specified number (step S403).
[0111] If the number of sampling times is equal to or less than the specified number (Yes at step S403), the surface detection and separation unit 12ba samples the point cloud (step S404). The range of the point cloud sampled at one time is assumed to be specified in advance.
[0112] Then, the surface detection and separation unit 12ba calculates a plane from the sampled point group using the least squares method or the like (step S405). The surface detection and separation unit 12ba calculates the plane by solving an equation of the plane using the least squares method or the like.
[0113] Next, the surface detection / separation unit 12ba calculates the error between the calculated plane and the sampled point cloud (step S406).Then, the surface detection / separation unit 12ba determines whether the calculated error is equal to or smaller than a predetermined threshold (step S407).If it is equal to or smaller than the threshold (step S407, Yes), the surface detection / separation unit 12ba resamples the point cloud whose error is equal to or smaller than the threshold (step S408).
[0114] The surface detection and separation unit 12ba then stores the resampled point cloud as a plane object (step S409), and deletes the point cloud that has been converted into the plane object from the non-plane object (step S410), thereby separating the plane object from the non-plane object.
[0115] Then, the surface detection and separation unit 12ba repeats the process from step S403. If the error exceeds the threshold value in step S407 (No in step S407), the surface detection and separation unit 12ba also repeats the process from step S403.
[0116] If the sampling count exceeds the specified count in step S403 (No in step S403), the surface detection and separation unit 12ba outputs and stores the planar object and the non-planar object, and ends the surface detection and separation process.
[0117] Next, the processing procedure of the camera position calculation processing in step S309 will be described with reference to Fig. 13 and Fig. 14. Fig. 13 is a flowchart showing the processing procedure of the camera position calculation processing. Fig. 14 is an explanatory diagram of the camera position calculation processing.
[0118] 13, the camera position calculation unit 12bb inputs a 3D point cloud of each object to be annotated (step S501), and then calculates a spherical surface that contains all of the point clouds from the center of the point clouds (step S502).
[0119] Next, the camera position calculation unit 12bb arranges virtual cameras at equal intervals on the calculated spherical surface (step S503). Then, the camera position calculation unit 12bb derives a point cloud (X i , Y i , Z i ) into 2D coordinates (x i , y i ) (step S504).
[0120] Then, the camera position calculation unit 12bb calculates the covariance σ from the converted 2D coordinates using the following formula (a): xy is calculated (step S505).
[0121]
[0122] The camera position calculation unit 12bb then selects the camera position with the maximum covariance as the optimal camera position (step S506) and outputs it (step S507). Note that Fig. 14 shows a schematic diagram of the difference in relative camera positions when the covariance is large or small.
[0123] The optimal camera position is the one that allows the user to perform segmentation and annotation with minimal occlusion and minimal camera movement.
[0124] As shown in the left diagram of Figure 14, the larger the covariance calculated in step S505, the less likely the camera position is to be occluded by other chairs in a cluster corresponding to, for example, a group of chairs. On the other hand, as shown in the right diagram of Figure 14, the smaller the covariance, the more likely the camera position is to be occluded by other chairs in a cluster corresponding to, for example, a group of chairs. Therefore, the camera position with the largest covariance can be determined to be the optimal camera position.
[0125] 5-4. Effects of the annotation work support method according to the embodiment of the present disclosure> Fig. 15 is a diagram (part 1) showing the effects of the annotation work support method according to the embodiment of the present disclosure. Fig. 16 is a diagram (part 2) showing the effects of the annotation work support method according to the embodiment of the present disclosure.
[0126] 15 shows an example in which the annotation support method according to this embodiment is applied to the annotation of a room containing chairs and a table. First, the input 3D model 11b of this room is classified into planar objects and non-planar objects and then separated.
[0127] The table top is classified and separated as a planar object and is presented on the 2D viewer of the display unit 5 at the optimal camera position. Chairs are classified and separated as non-planar objects, and if there are multiple chairs, they are grouped into the same cluster as nearby chairs.
[0128] When the user selects this cluster, objects other than the selected cluster are hidden, and the selected cluster is presented on the 2D viewer of the display unit 5 at an optimal camera position. At this time, the chair cluster is presented without occlusion, allowing the user to easily perform segmentation and annotation information assignment for the presented cluster. The table legs are also presented as non-planar objects in the same way as the chair, allowing the user to easily perform segmentation and annotation information assignment.
[0129] FIG. 16 shows an example in which the annotation work support method according to this embodiment is applied to the annotation of a room in which a smartphone 50 is placed on a table.
[0130] First, the input 3D model 11b of this room is classified and separated into planar objects and non-planar objects. Here, the desk and smartphone 50 are classified and separated as a single planar object, the optimal camera position is calculated, and the optimal camera position is presented on the 2D viewer of the display unit 5. The user can individually segment and add annotation information to the desk and smartphone 50 presented on the 2D viewer.
[0131] 11 can be cited as an example similar to that of Fig. 16. In this case, wall B and the door attached to wall B are classified and separated as a single planar object and presented on the 2D viewer of the display unit 5 at an optimal camera position that allows the user to view them from the front. The user can individually segment and add annotation information to wall B and the door presented on the 2D viewer.
[0132] <<6. Other Modifications>> Although the embodiments of the present disclosure have been described up to this point, the information processing method according to the present embodiment can be modified in several other ways.
[0133] For example, in this embodiment, planes are detected by solving plane equations using the least squares method, but this is merely an example and other methods may be used, such as a plane detection function implemented in Open3D, an OSS (Open Source Software) library for 3D data processing.
[0134] Furthermore, in this embodiment, the region segmentation of the 2D image is performed using the region segmentation model 11e, which is a machine learning model. However, it is not necessary to use a machine learning model, and the region segmentation may be performed by image processing such as edge detection, for example.
[0135] Furthermore, among the processes described in the above-described embodiments of the present disclosure, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using known methods. In addition, the information including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.
[0136] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.
[0137] The above-described embodiments of the present disclosure can be combined as appropriate within the scope of the present disclosure without causing any contradiction in the processing content. The order of the steps shown in the sequence diagrams or flowcharts of the present embodiments can be changed as appropriate.
[0138] <<7. Hardware Configuration>> The learning device 10 and image processing device 20 according to the above-described embodiments of the present disclosure are realized by, for example, a computer 1000 configured as shown in FIG. 17 . The image processing device 20 will be described as an example. FIG. 17 is a hardware configuration diagram showing an example of a computer 1000 that realizes the functions of the image processing device 20. The computer 1000 includes a CPU 1100, a RAM 1200, a ROM 1300, a secondary storage device 1400, a communication interface 1500, and an input / output interface 1600. The components of the computer 1000 are connected via a bus 1050.
[0139] The CPU 1100 operates and controls each component based on programs stored in the ROM 1300 or the secondary storage device 1400. For example, the CPU 1100 loads the programs stored in the ROM 1300 or the secondary storage device 1400 into the RAM 1200 and executes processing corresponding to the various programs.
[0140] The ROM 1300 stores boot programs such as a Basic Input Output System (BIOS) that is executed by the CPU 1100 when the computer 1000 is started, and programs that depend on the hardware of the computer 1000 .
[0141] The secondary storage device 1400 is a computer-readable recording medium that non-temporarily records programs executed by the CPU 1100 and data used by such programs. Specifically, the secondary storage device 1400 is a recording medium that records the programs according to this embodiment.
[0142] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550. For example, the CPU 1100 receives data from other devices and transmits data generated by the CPU 1100 to other devices via the communication interface 1500.
[0143] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from an input device such as a keyboard or a mouse via the input / output interface 1600. The CPU 1100 also transmits data to an output device such as a display, a speaker, or a printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs and the like recorded on a predetermined recording medium. Examples of media include optical recording media such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), magneto-optical recording media such as an MO (Magneto-Optical Disk), tape media, magnetic recording media, and semiconductor memories.
[0144] For example, when the computer 1000 functions as the image processing device 20, the CPU 1100 of the computer 1000 executes a program loaded onto the RAM 1200 to realize the functions of the control unit 12. The secondary storage device 1400 stores the program according to the present disclosure and data in the storage unit 11. The CPU 1100 reads and executes the program data 1450 from the secondary storage device 1400, but as another example, the CPU 1100 may obtain these programs from another device via an external network 1550.
[0145] <<8. Conclusion>> As described above, according to an embodiment of the present disclosure, the image processing device 20 (corresponding to an example of an “information processing device”) includes the control unit 12. The control unit 12 acquires camera posture information for presenting a 3D model 11b (corresponding to an example of “3D data”) representing a real space. The control unit 12 also acquires 3D annotation information 11c (corresponding to an example of “annotation information”) of the 3D model 11b corresponding to the camera posture information. The control unit 12 also acquires a 2D image (corresponding to an example of “2D data”) representing the real space corresponding to the camera posture information. The control unit 12 also acquires 2D region division information 11f by dividing the 2D image into a plurality of regions. The control unit 12 also acquires integrated annotation information 11g by integrating the 3D annotation information 11c and the 2D region division information 11f. This makes it possible to easily obtain high-quality learning data using 3D data.
[0146] Although the embodiments of the present disclosure have been described above, the technical scope of the present disclosure is not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present disclosure. Furthermore, components of different embodiments and modifications may be combined as appropriate.
[0147] Furthermore, the effects of each embodiment described in this specification are merely examples and are not intended to be limiting, and other effects may also be obtained.
[0148] Note that the present technology can also be configured as follows. (1) An information processing device including a control unit that acquires camera posture information for presenting 3D data representing a real space, acquires annotation information of the 3D data corresponding to the camera posture information, acquires 2D data representing the real space corresponding to the camera posture information, acquires 2D region division information by dividing the 2D data into a plurality of regions, and acquires integrated annotation information by integrating the annotation information and the 2D region division information. (2) The information processing device described in (1), wherein the control unit acquires 2D rendering information by rendering the annotation information onto an image plane corresponding to the 2D data, and integrates each region in the 2D region division information with the 2D rendering information corresponding to each region. (3) The information processing device described in (2), wherein the control unit integrates the annotation information and the 2D region division information by assigning, to each region in the 2D region division information, a mode of the annotation information corresponding to each region in the 2D rendering information. (4) The information processing device according to (1), (2), or (3), wherein the control unit synthesizes the 2D data for any viewpoint, acquires the 2D region division information of the synthesized 2D data, and integrates the 2D region division information with the annotation information. (5) The information processing device according to any one of (1) to (4), wherein the control unit classifies the 3D data into at least one planar object and a plurality of non-planar objects, and acquires the annotation information representing the plurality of non-planar objects. (6) The information processing device according to (5), wherein the control unit separates the planar object and the non-planar object based on the classification. (7) The information processing device according to (6), wherein the separation includes changing a positional relationship between the planar object and the non-planar object.(8) The information processing device according to (7), wherein the control unit presents the camera posture information so as to correspond to at least one of the plurality of non-planar objects that does not have the annotation information, based on a positional relationship between the plurality of non-planar objects. (9) The information processing device according to (6), (7), or (8), wherein the separation includes processing to hide the non-planar objects. (10) An information processing method comprising: acquiring camera posture information for presenting 3D data representing a real space; acquiring annotation information for the 3D data corresponding to the camera posture information; acquiring 2D data representing the real space that corresponds to the camera posture information; dividing the 2D data into a plurality of regions to acquire 2D region division information; and acquiring integrated annotation information by integrating the annotation information and the 2D region division information. (11) A computer-readable non-transitory storage medium storing a program that causes a computer to perform the following operations: acquire camera posture information for presenting 3D data representing a real space; acquire annotation information of the 3D data corresponding to the camera posture information; acquire 2D data representing the real space corresponding to the camera posture information; acquire 2D region division information by dividing the 2D data into a plurality of regions; and acquire integrated annotation information by integrating the annotation information and the 2D region division information.
[0149] 3 Operation unit 5 Display unit 10 Learning device 11 Storage unit 11a Multi-viewpoint image DB 11b 3D model 11c 3D annotation information 11d 2D rendering information 11e Region division model 11f 2D region division information 11g Integrated annotation information 11h Recognition model 12 Control unit 12a Acquisition unit 12b 3D annotation unit 12ba Surface detection and separation unit 12bb Camera position calculation unit 12c Rendering unit 12d Region division unit 12e Integration unit 12f Display control unit 12g Learning unit 20 Image processing device
Claims
1. An information processing apparatus comprising a control unit that acquires camera pose information for presenting 3D data representing a real space, acquires annotation information of the 3D data corresponding to the camera pose information, acquires 2D data representing the real space corresponding to the camera pose information, acquires 2D region division information by dividing the 2D data into a plurality of regions, and acquires integrated annotation information by integrating the annotation information and the 2D region division information.
2. The information processing apparatus according to claim 1, wherein the control unit acquires 2D rendering information by rendering the annotation information onto an image plane corresponding to the 2D data, and integrates each region in the 2D region division information with the 2D rendering information corresponding to each region.
3. The information processing apparatus according to claim 2, wherein the control unit integrates the annotation information and the 2D region division information by assigning the most frequent value of the annotation information corresponding to each region in the 2D rendering information to each region in the 2D region division information.
4. The information processing apparatus according to claim 1, wherein the control unit synthesizes the 2D data for an arbitrary viewpoint, acquires the 2D region division information of the synthesized 2D data, and integrates the 2D region division information and the annotation information.
5. The information processing apparatus according to claim 1, wherein the control unit classifies the 3D data into at least one planar object and a plurality of non-planar objects, and acquires the annotation information representing the plurality of non-planar objects.
6. The information processing apparatus according to claim 5, wherein the control unit separates the planar object and the non-planar object based on the classification.
7. The information processing apparatus according to claim 6, wherein the separation includes changing the positional relationship between the planar object and the non-planar object.
8. The information processing apparatus according to claim 7, wherein the control unit presents the camera pose information so as to correspond to at least one of the plurality of non-planar objects having no annotation information based on the positional relationship of the plurality of non-planar objects.
9. The information processing apparatus according to claim 6, wherein the separation includes a non-display process of the non-planar object.
10. Obtaining camera pose information for presenting 3D data representing a real space, obtaining annotation information of the 3D data corresponding to the camera pose information, obtaining 2D data representing the real space corresponding to the camera pose information, obtaining 2D region division information by dividing the 2D data into a plurality of regions, and obtaining integrated annotation information by integrating the annotation information and the 2D region division information. An information processing method having the above.
11. A computer-readable non-transitory storage medium storing a program that causes a computer to obtain camera pose information for presenting 3D data representing a real space, obtain annotation information of the 3D data corresponding to the camera pose information, obtain 2D data representing the real space corresponding to the camera pose information, obtain 2D region division information by dividing the 2D data into a plurality of regions, and obtain integrated annotation information by integrating the annotation information and the 2D region division information.
Citation Information
Patent Citations
Method and system generating correct answer data for machine learning of recognition unit
JP2024000347A