Recognition device, method, and program
The recognition device and method address the high annotation costs for point clouds by extracting and aligning feature amounts from point clouds and images, enabling accurate class discrimination and recognition without additional annotation.
Patent Information
- Application Number
- PCT/JP2023/043279
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-04
- Publication Date
- 2025-06-12
AI Technical Summary
The high cost of annotating point clouds for urban environments makes it challenging to create sufficient learning data for accurate recognition, especially when compared to image datasets.
A recognition device and method that extracts feature amounts from both point cloud and image data, using a feature extraction unit and a class identification unit, allowing for appropriate recognition of point clouds without additional annotation, by learning model parameters to align point cloud and image features within the same class.
Enables effective class discrimination of point clouds lacking annotation information, allowing for accurate recognition and model learning without the need for additional annotation, thereby reducing costs and improving data efficiency.
Smart Images

Figure JP2023043279_12062025_PF_FP_ABST
Abstract
Description
Recognition device, method and program
[0001] FIELD Embodiments of the present invention relate to a recognition device, a method, and a program.
[0002] In recent years, high-precision 3D maps of urban environments have been created using 3D point clouds measured by LiDAR (Light Detection and Ranging) installed in automobiles, and it is expected that these maps will be used for urban planning or disaster prevention simulations.
[0003] However, the work cost of annotating measured point clouds to utilize them is relatively high, and compared to image datasets based on images for which the cost of annotating is relatively low, point cloud datasets have a relatively small number of data points or classes, making it difficult to train point clouds to recognize the real world with a sufficient number of classes and accuracy.
[0004] Therefore, in order to compensate for the lack of data volume caused by the relatively high cost of annotating point clouds, Non-Patent Document 1 discloses a method for recognizing point clouds for new classes where point cloud annotations are insufficient, by utilizing not only point clouds but also image knowledge.
[0005] Xu, Chenfeng, et al. "Image2point: 3d point-cloud understanding with 2d image pretrained models." European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2022.
[0006] However, in order to achieve a recognition rate comparable to that of the same class in image recognition, the above method requires the creation of point cloud training data using annotations and fine-tuning based on the annotation data.
[0007] The present invention has been made in light of the above circumstances, and its object is to provide a recognition device, method, and program that are capable of appropriately recognizing point clouds.
[0008] A recognition device according to one aspect of the present invention includes a feature extraction unit that extracts features from point cloud data or image data, and a class identification unit that identifies a class to which the data from which the features extracted by the feature extraction unit were extracted belongs, and the feature extraction unit extracts mutually similar features from multiple data belonging to the same class.
[0009] A recognition method according to one aspect of the present invention is a method performed by a recognition device, and includes: extracting features of point cloud data or image data using a feature extraction unit of the recognition device; and identifying, using a class identification unit of the recognition device, a class to which the data from which the features extracted by the feature extraction unit were extracted belongs; and the feature extraction unit extracts mutually similar features from multiple data belonging to the same class.
[0010] According to the present invention, a point cloud can be appropriately recognized.
[0011] FIG. 1 is a diagram showing an application example of a recognition device according to an embodiment of the present invention. FIG. 2 is a diagram explaining an example of inference by the recognition device according to an embodiment of the present invention. FIG. 3 is a diagram explaining a first example of learning by the recognition device according to an embodiment of the present invention. FIG. 4 is a diagram explaining a second example of learning by the recognition device according to an embodiment of the present invention. FIG. 5 is a diagram explaining a third example of learning by the recognition device according to an embodiment of the present invention. FIG. 6 is a flowchart showing an example of a learning processing procedure by the recognition device according to an embodiment of the present invention. FIG. 7 is a flowchart showing an example of an inference processing procedure by the recognition device according to an embodiment of the present invention. FIG. 8 is a block diagram showing an example of the hardware configuration of the recognition device according to an embodiment of the present invention.
[0012] An embodiment of the present invention will be described below with reference to the drawings. In this embodiment, with regard to the recognition of objects present in a three-dimensional point cloud and the training of a model related to said recognition, attention is paid to the fact that the subject in an image obtained by photographing the same area of space as that measured by the point cloud is the same subject, that is, the class to which the point cloud belongs and the class to which the image belongs are the same class.
[0013] In this embodiment, the point cloud feature values and the point cloud and image from which the image feature values are extracted, obtained by inputting point cloud and image data for the same subject, are considered to belong to the same class in class identification, and the parameters of the model related to feature extraction are learned so that the respective data are close to each other in feature space.
[0014] Then, for point cloud data or image data for which annotations have been obtained, model parameters related to feature extraction are trained so that data belonging to the same class are closer in feature space and data belonging to different classes are farther apart in feature space. For example, in this embodiment, for image data depicting the same subject as the subject associated with the point cloud data, the point cloud class and the image class are considered to be the same class, and the point cloud features and image features are approximated. Furthermore, using a large amount of annotated image data, parameters of a model related to feature extraction and a model related to class identification are trained so that classes can be identified from various features. This makes it possible to appropriately identify the class of a point cloud with insufficient annotation information and to train a model that achieves this identification, for example, by inputting features extracted from the point cloud into a class identifier, without having to perform additional annotations on the measured point cloud.
[0015] Fig. 1 is a diagram showing an application example of a recognition device according to an embodiment of the present invention. As shown in Fig. 1, the recognition device 100 according to this embodiment includes a point cloud feature extraction unit 10, an image feature extraction unit 20, a class identification unit 30, and a feature extraction learning unit 40. Below, an example of the processing of each unit will be described in order.
[0016] 2 is a diagram illustrating a first example of inference by a recognition device according to an embodiment of the present invention. In the example shown in FIG. 2, a point cloud feature extraction unit 10 inputs a point cloud (point cloud data) A, extracts point cloud features, and outputs them. To extract the point cloud features, any point cloud class classifier based on a deep learning model such as PointNet can be used. A method using the output of the fully connected (FC) layer of this classifier as the point cloud features, or any other point cloud recognition method may be used.
[0017] The image feature extraction unit 20 receives an image (image data) B, extracts image features, and outputs them. To extract the image features, any image class classifier based on a deep learning model such as ResNet can be used. In addition to a method of using the output of the FC layer of this classifier as the image features, any image recognition method may also be used.
[0018] In the example shown in Figure 2, when the point cloud feature extraction unit 10 inputs a point cloud A that belongs to a certain class, and the image feature extraction unit 20 inputs an image B that belongs to the same class as the point cloud A, the point cloud feature extraction unit 10 and the image feature extraction unit 20 output feature amounts that are similar to each other.
[0019] The class identification unit 30 receives image features or point group features, infers the class to which the input data belongs based on the received features, and outputs the class as a recognition result.
[0020] 3 is a diagram illustrating a first example of learning by a recognition device according to an embodiment of the present invention. In the example shown in FIG. 3, learning is described in which feature quantities of data belonging to the same class, which are extracted by the point cloud feature extraction unit 10 and the image feature extraction unit 20, are approximated.
[0021] First, point cloud data and image data containing the same subject without annotations are prepared. Next, the point cloud feature extraction unit 10 and the image feature extraction unit 20 input the point cloud data and image data containing the same subject in a one-to-one correspondence. The point cloud feature extraction unit 10 and the image feature extraction unit 20 extract point cloud feature quantities and image feature quantities in a one-to-one correspondence from this input.
[0022] For these two extracted features, the feature extraction learning unit 40 updates (learns) the parameters of the model related to point cloud feature extraction in the point cloud feature extraction unit 10 (sometimes referred to as the point cloud feature extraction model) and the parameters of the model related to image feature extraction in the image feature extraction unit 20 (sometimes referred to as the image feature extraction model) using, for example, the cross-modal center loss method disclosed in "Jing, Longlong, et al. "Cross-modal center loss for 3d cross-modal retrieval." Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2021." or any loss function.
[0023] At this time, the feature extraction learning unit 40 updates the parameters of the models related to feature extraction in the point cloud feature extraction unit 10 and the image feature extraction unit 20 so that, for input data that has annotations and belongs to the same class, the features output from the point cloud feature extraction unit 10 and the features output from the image feature extraction unit 20 are similar to each other, and for input data that belongs to different classes, the output features are different from each other.
[0024] For example, first, if the set of training point cloud data is point cloud data belonging to the same class, such as the “training point cloud (class α)” or the “training point cloud (class β)” shown in FIG. 3, the feature extraction learning unit 40 learns the parameters of the above model of the point cloud feature extraction unit 10 so that mutually similar point cloud features are extracted for each point cloud data.
[0025] Second, if the set of image data for training belongs to the same class, such as "training images (class α)" or "training images (class β)" shown in Figure 3, the feature extraction learning unit 40 learns the parameters of the model of the image feature extraction unit 20 so that mutually similar image features are extracted for each image data.
[0026] Third, the feature extraction learning unit 40 updates the parameters of the above models in the point cloud feature extraction unit 10 and the image feature extraction unit 20 so that the feature values output from the point cloud feature extraction unit 10 when point cloud data belonging to a certain class is input are similar to the feature values output from the image feature extraction unit 20 when image data belonging to the same class is input.
[0027] The class identification unit 30 has a learning function for the parameters of a model related to its own class identification. This function may be realized by a learning unit separate from the class identification unit 30. The class identification unit 30 inputs an arbitrary image data set containing a class to be recognized from a point cloud, inputs the output layer of the image feature extraction unit 20, and updates the parameters of the model related to class identification (sometimes referred to as a class identification model). The class identification unit 30 also outputs the trained point cloud feature extraction model, image feature extraction model, and class identification model.
[0028] The order of learning the models related to the point cloud feature extraction unit 10 and the image feature extraction unit 20 and the model related to the class identification unit 30 can be arbitrary. For example, the learning of the models related to the image feature extraction unit 20 and the point cloud feature extraction unit 10 can be performed before the learning of the model related to the class identification unit 30, and it is also possible to perform the learning of the model related to the image feature extraction unit 20 and the model related to the point cloud feature extraction unit 10 in parallel with the learning of the model related to the class identification unit 30.
[0029] 4 is a diagram illustrating a second example of learning by a recognition device according to an embodiment of the present invention. In the example shown in Fig. 4, feature extraction learning unit 40 inputs pairs of point clouds and images obtained by photographing the same object but whose class is unknown, into point cloud feature extraction unit 10 and image feature extraction unit 20 on a one-to-one basis, and learns the model parameters of point cloud feature extraction unit 10 and image feature extraction unit 20 so that the feature quantities extracted from the point clouds and images are similar to each other.
[0030] Fig. 5 is a diagram illustrating a third example of learning by a recognition device according to an embodiment of the present invention. In the example shown in Fig. 5, feature extraction learning unit 40 inputs first and second images belonging to different classes, for example, "learning image (class α)" and "learning image (class β)" shown in Fig. 5, to image feature extraction unit 20, and learns the parameters of the model of image feature extraction unit 20 so that the feature quantities extracted from these images are different from each other.
[0031] In addition, the feature extraction learning unit 40 inputs the first and second point clouds, which belong to different classes, into the point cloud feature extraction unit 10, and learns the parameters of the above model of the point cloud feature extraction unit 10 so that the feature quantities extracted from these point clouds are separated from each other.
[0032] 6 is a flowchart showing an example of a learning procedure performed by a recognition device according to an embodiment of the present invention. When a set of annotated point cloud data or image data is input, the feature extraction learning unit 40 learns the parameters of the model corresponding to the input data, out of the models of the point cloud feature extraction unit 10 and the image feature extraction unit 20, so as to bring feature quantities of input data belonging to the same class closer to each other and move feature quantities of input data belonging to different classes farther apart (S11).
[0033] When a set of pairs of point cloud data and image data obtained by measuring the same object is input, the feature extraction learning unit 40 learns the parameters of the models of at least one of the point cloud feature extraction unit 10 and the image feature extraction unit 20 so that the image features extracted from the image data and the point cloud features extracted from the point cloud data paired with this image data are similar to each other (S12).
[0034] When feature quantities for a set of annotated point cloud data or image data are input, the class identification unit 30 learns the model parameters related to the class identification unit 30 so that the correct class can be identified in accordance with this input (S13). The class identification unit 30 outputs the trained point cloud feature extraction model, image feature extraction model, and class identification model (S14).
[0035] 7 is a flowchart showing an example of an inference processing procedure by a recognition device according to an embodiment of the present invention. Here, an example of inference when the input data is point cloud data is shown, but inference when the input data is image data is also possible. The point cloud feature extraction unit 10 inputs point cloud data to be classified as a class (S21). The point cloud feature extraction unit 10 extracts features from the input point cloud data and outputs the extracted features to the class classification unit 30 (S22). The class classification unit 30 recognizes the class to which the point cloud data related to the features output in S22 belongs and outputs the classified point cloud class (S23).
[0036] 8 is a block diagram showing an example of the hardware configuration of a recognition device 100 according to an embodiment of the present invention. In the example shown in FIG. 8, the recognition device 100 according to the embodiment is configured, for example, by a server computer or a personal computer, and has a hardware processor 111A such as a CPU (Central Processing Unit). A program memory 111B, a data memory 112, an input / output interface 113, and a communication interface 114 are connected to this hardware processor 111A via a bus 115.
[0037] The communication interface 114 includes, for example, one or more wireless communication interface units, and enables transmission and reception of information to and from a communication network NW. As the wireless interface, for example, an interface that adopts a low-power wireless data communication standard such as a wireless LAN (Local Area Network) is used.
[0038] An input device 200 and an output device 300, which are attached to the recognition device 100 and used by a user or the like, are connected to the input / output interface 113. The input / output interface 113 receives operation data input by a user or the like through the input device 200, such as a keyboard, a touch panel, a touchpad, or a mouse, and outputs output data to an output device 300, which includes a display device using a liquid crystal or an organic electroluminescence (EL) display, for display. The input device 200 and the output device 300 may be devices built into the recognition device 100, or may be input devices and output devices of other information terminals that can communicate with the recognition device 100 via a network NW.
[0039] The program memory 111B is a non-transitory tangible storage medium that is a combination of a non-volatile memory that can be written to and read from at any time, such as a hard disk drive (HDD) or a solid state drive (SSD), and a non-volatile memory such as a read only memory (ROM), and stores programs necessary to execute various control processes, etc., according to one embodiment.
[0040] The data memory 112 is a tangible storage medium that is a combination of, for example, the above-mentioned nonvolatile memory and a volatile memory such as a RAM (Random Access Memory), and is used to store various data acquired and created during various processes performed by the recognition device 100.
[0041] The recognition device 100 according to an embodiment of the present invention may be configured as a data processing system or information processing device having a software-based processing function unit. A storage system used as a work memory or the like by the recognition device 100 may be configured by using the data memory 112 shown in FIG. 8. However, these configured storage areas are not essential components within the recognition device 100, and may be areas provided in a storage system such as an external storage medium such as a USB (Universal Serial Bus) memory, or a database server located in the cloud.
[0042] The processing function unit can be realized by having the hardware processor 111A read and execute a program stored in the program memory 111B, but the processing function unit may also be realized in various other forms, including an integrated circuit such as an application specific integrated circuit (ASIC) or a field-programmable gate array (FPGA).
[0043] The techniques described in the above embodiments can be stored as a program (software means) that can be executed by a computer on a recording medium such as a magnetic disk (e.g., a floppy disk, a hard disk, etc.), an optical disk (e.g., a CD-ROM, a DVD, an MO, etc.), or a semiconductor memory (e.g., a ROM, a RAM, a flash memory, etc.), and can be distributed by transmission via a communication medium. The program stored on the medium also includes a configuration program that configures the software means (including not only execution programs but also tables and data structures) that the computer executes. The computer that realizes this device reads the program stored on the recording medium and, in some cases, configures the software means using the configuration program, and executes the above-mentioned processing by controlling the operation of this software means. The term "recording medium" as used herein is not limited to a storage medium for distribution, but also includes a storage medium such as a magnetic disk or semiconductor memory installed inside the computer or in a device connected via a network.
[0044] The present invention is not limited to the above-described embodiments, and various modifications can be made in the implementation stage without departing from the spirit of the invention. Furthermore, the embodiments may be implemented in appropriate combinations, in which case the combined effects can be obtained. Furthermore, the above-described embodiments include various inventions, and various inventions can be extracted by combining selected elements from the disclosed elements. For example, if the problem can be solved and the desired effect can be obtained even if some elements are deleted from all elements shown in the embodiments, the configuration from which these elements are deleted can be extracted as an invention.
[0045] 100... Recognition device 10... Point cloud feature extraction unit 20... Image feature extraction unit 30... Class identification unit 40... Feature extraction learning unit
Claims
1. A recognition device comprising: a feature extraction unit that extracts feature amounts of point group data or image data; and a class identification unit that identifies a class to which the data from which the feature amounts extracted by the feature extraction unit belong, wherein the feature extraction unit extracts mutually similar feature amounts for a plurality of data belonging to the same class.
2. The recognition device according to claim 1, further comprising a learning unit that learns parameters of a model related to extraction of the feature amounts by the feature extraction unit such that the feature extraction unit extracts mutually similar feature amounts for a plurality of data belonging to the same class and extracts mutually different feature amounts for a plurality of data belonging to mutually different classes.
3. A recognition method performed by a recognition device, the method comprising: extracting, by a feature extraction unit of the recognition device, feature amounts of point group data or image data; and identifying, by a class identification unit of the recognition device, a class to which the data from which the feature amounts extracted by the feature extraction unit belong, wherein the feature extraction unit extracts mutually similar feature amounts for a plurality of data belonging to the same class.
4. A recognition processing program that causes a processor to function as each unit of the recognition device according to claim 1 or 2.
Citation Information
Patent Citations
Systems and methods for image processing
JP2021534523A
Cited By
Annotation support device, method, and program using images and point clouds
JP7792176B1