Point cloud data recognition device, method, and program
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-03
- Publication Date
- 2026-08-14
AI Technical Summary
【0011】 すなわちこの発明の一態様によれば、点群データを画像化する際の仮想的なカメラの視線を、剛体変形に止まらず非剛体変形まで考慮して最適化できるようにした技術を提供することができる。
Smart Images

Figure 2026131497000001_ABST
Abstract
Description
Technical Field
[0001] One aspect of this invention relates to a point cloud data recognition device, method, and program that recognize class labels of three-dimensional point cloud data using, for example, a machine learning model.
Background Art
[0002] Point cloud data measured using a depth camera capable of measuring depth in a three-dimensional space, a laser radar (Light Detection and Ranging: LiDAR), etc. is expected to be used for creating high-precision three-dimensional maps, urban planning, simulations, etc. used in autonomous vehicles and autonomous mobile robots. In addition, research and development on technologies for recognizing objects from measured point cloud data are also widely carried out.
[0003] By the way, in recent years, many methods have been proposed for extracting features from point cloud data using a machine learning model to which deep learning is applied and recognizing the class label of the point cloud based on the extracted features. However, point cloud data often records a specific area or object, and it is difficult to obtain diverse data compared to modalities such as images and texts. In addition, the work of attaching learning annotations to the obtained point cloud data also requires a great deal of cost.
[0004] Therefore, for example, in Non-Patent Documents 1-3, methods have been proposed to utilize an image recognition model that has been pre-trained using a large amount of data in order to recognize the class label of a point cloud.
Prior Art Documents
Non-Patent Documents
[0005]
Non-Patent Document 1
[0006] However, the methods described in Non-Patent Documents 1-3 involve converting point cloud data into an image via a virtual camera, then using a machine learning model to extract features from the image, and finally identifying class labels based on the extracted features. Therefore, depending on the orientation and position of the virtual camera, the image generated from the point cloud data may change, resulting in images that are unsuitable for class label recognition.
[0007] Furthermore, methods have been proposed to optimize the arrangement of virtual cameras, that is, the virtual line of sight used when creating images. However, conventionally proposed optimization methods only consider so-called rigid body deformation, which involves rotating the point cloud without deformation. As a result, non-rigid body deformations such as distortion and shear of the virtual camera lens are not taken into consideration, which has been an obstacle to further improving the accuracy of class label recognition.
[0008] This invention was made in view of the above circumstances, and aims to provide a technology that enables the optimization of the line of sight of a virtual camera when imaging point cloud data, taking into account not only rigid body deformation but also non-rigid body deformation. [Means for solving the problem]
[0009] To solve the above problems, one embodiment of the point cloud data recognition device or point cloud data recognition method according to the present invention acquires point cloud data of a target space and generates an image by projecting the acquired point cloud data according to pre-set projection parameters. Then, it extracts feature quantities from the image and recognizes the class of the imaged point cloud data based on the feature quantities. Furthermore, as the projection parameters, rigid body transformation parameters that transform the arrangement of a virtual camera onto which the point cloud data is projected and non-rigid body transformation parameters that transform the projection plane of the virtual camera are used.
[0010] According to one aspect of this invention, point cloud data is imaged using rigid transformation parameters that transform the arrangement of a virtual camera and non-rigid transformation parameters that transform the projection plane, such as the distortion of the virtual camera's lens. Therefore, it becomes possible to identify the class of point cloud data based on an image generated that takes into account not only rigid deformations such as the attitude and position of the virtual camera, but also non-rigid deformations such as the distortion characteristics and shear of the virtual camera's lens. Consequently, it becomes possible to recognize class labels with higher accuracy. [Effects of the Invention]
[0011] In other words, according to one aspect of this invention, it is possible to provide a technology that enables the optimization of the line of sight of a virtual camera when imaging point cloud data, taking into account not only rigid body deformation but also non-rigid body deformation. [Brief explanation of the drawing]
[0012] [Figure 1] Figure 1 shows an example of the configuration of a system equipped with a point cloud data recognition device according to the first embodiment of this invention. [Figure 2] Figure 2 is a block diagram showing an example of the hardware configuration of a point cloud data recognition device according to the first embodiment of this invention. [Figure 3] Figure 3 is a block diagram showing an example of the software configuration of a point cloud data recognition device according to the first embodiment of this invention. [Figure 4] Figure 4 is a block diagram showing the software configuration of the point cloud image processing unit of the point cloud data recognition device shown in Figure 3. [Figure 5] Figure 5 is a flowchart showing an example of the processing procedure and content of the projection parameter optimization process performed by the control unit of the point cloud data recognition device shown in Figure 3 during the learning phase. [Figure 6] Figure 6 is a flowchart showing an example of the processing procedure and content of the solution candidate generation process, which is part of the point cloud class label identification process shown in Figure 5. [Figure 7] Figure 7 is a flowchart showing the processing procedure and processing content of the point cloud class label identification process performed by the control unit of the point cloud data recognition device shown in Figure 3 during the estimation phase, as in the first embodiment. [Figure 8] Figure 8 is a flowchart showing an example of the processing procedure and content of the image conversion process for point cloud data, which is part of the point cloud class label identification process shown in Figure 7. [Figure 9] Figure 9 is a diagram illustrating an example of the operation of the point cloud class label identification process shown in Figure 7. [Figure 10]FIG. 10 is a flowchart showing an example of the processing procedure and processing content of the projection parameter optimization process executed by the control unit of the point group data recognition device according to the second embodiment of the present invention in the learning phase. [Figure 11] FIG. 11 is a flowchart showing an example of the processing procedure and processing content of the solution candidate generation process among the point group class label identification processes shown in FIG. 10.
Embodiments for Carrying Out the Invention
[0013] Hereinafter, embodiments of the present invention will be described with reference to the drawings.
[0014] [First Embodiment] (Configuration Example) (1) System FIG. 1 is a diagram showing an example of the configuration of a system including a point group data recognition device according to the first embodiment of the present invention.
[0015] The system of the first embodiment connects a point group data measurement device MS and a terminal device US to a point group data recognition device SM via a network NW.
[0016] The point group data measurement device MS is composed of, for example, a depth camera or a lidar (LiDAR), and transmits three-dimensional point group data (hereinafter simply referred to as point group data) obtained by measuring a target range including an object that is a target object to the point group data recognition device SV via the network NW.
[0017] The terminal device US is used by a system administrator or a user and is composed of, for example, a personal computer. The terminal device US is used to input learning point group data, verification data, and point cloud imaging conditions to the point group data recognition device SV, and to receive data representing the class label of the point group identified by the point group data recognition device SV. Note that the terminal device US is not limited to a personal computer, and for example, a smartphone or a tablet terminal may be used.
[0018] A network (NW) comprises, for example, a wide-area network centered on the internet, and an access network for accessing this wide-area network. The access network may, but is not limited to, a public data communication network using wireless technology or a LAN (Local Area Network).
[0019] The point cloud data recognition device SV, the point cloud data measurement device MS, and the terminal device US may be connected via a wired LAN, a wireless interface using a low-power data communication standard such as Wi-Fi (registered trademark) or Bluetooth (registered trademark), or directly via a signal cable.
[0020] (2) Point cloud data recognition device SV The point cloud data recognition device SV is installed, for example, in a server computer located on the Web or in the cloud. Alternatively, the point cloud data recognition device SV may be installed in a terminal device, such as a personal computer used by a system administrator.
[0021] Figures 2 and 3 are block diagrams showing an example of the hardware and software configuration of the point cloud data recognition device SV.
[0022] The point cloud data recognition device SV includes a control unit 1 that uses a hardware processor such as a Central Processing Unit (CPU), and a storage unit having a program storage unit 2 and a data storage unit 3, and an input / output interface (hereinafter referred to as I / F) unit 4 are connected to this control unit 1 via a bus 5.
[0023] The input / output interface unit 4 transmits and receives various types of data between the point cloud data measurement device MS and the terminal device US in accordance with the communication protocol defined in the network NW.
[0024] The program storage unit 2 is configured, for example, as a storage medium, by combining a non-volatile memory that can be written to and read at any time, such as an HDD (Hard Disk Drive) or SSD (Solid State Drive), with a non-volatile memory such as ROM (Read Only Memory). In addition to middleware such as an OS (Operating System), it stores various programs necessary to execute various control processes according to the first embodiment of this invention.
[0025] The data storage unit 3 is configured, for example, as a storage medium, by combining a non-volatile memory that can be written to and read at any time, such as an HDD or SSD, with a volatile memory such as RAM (Random Access Memory). The storage area is provided with a projection parameter storage unit 31, an imaging condition storage unit 32, and an image feature extraction model storage unit 33, which are necessary storage units for carrying out the first embodiment.
[0026] The projection parameter storage unit 31 stores projection parameters used to image point cloud data using a machine learning model. The projection parameters include rigid transformation parameters for transforming the orientation and position of a virtual camera, and non-rigid transformation parameters for performing transformations according to the distortion model of the virtual camera's lens. The rigid transformation parameters are given by a rotation matrix representing the orientation of the virtual camera and a position vector. The non-rigid transformation parameters are given as a transformation matrix defining the distortion model of the virtual camera's lens.
[0027] The image conversion condition storage unit 32 stores various conditions used for image correction processes such as point cloud coloring, pixelation, sharpening, or denoising, which are performed during the image conversion process of point cloud data. These image conversion conditions are registered in advance, for example, from the system administrator's terminal device US.
[0028] The image feature extraction model storage unit 33 stores the weight parameters of a trained machine learning model used to extract features from an image generated based on point cloud data.
[0029] The control unit 1 includes, as processing functions for carrying out the first embodiment, a point cloud data acquisition processing unit 11, a point cloud image processing unit 12, an image feature extraction processing unit 13, a point cloud class label identification processing unit 14, a class label output processing unit 15, and a learning processing unit 16.
[0030] Each of these processing units 11 to 16 is implemented by having the hardware processor of the control unit 1 execute the application program stored in the program storage unit 2. Note that some or all of the above processing units 11 to 16 may be implemented using hardware such as LSI (Large Scale Integration) or ASIC (Application Specific Integrated Circuit).
[0031] In the projection parameter learning phase, the point cloud data acquisition processing unit 11 acquires multiple point cloud data for learning and corresponding validation class labels from the system administrator's terminal device US. In the estimation phase, the point cloud data acquisition processing unit 11 acquires point cloud data output from the point cloud data measurement device MS.
[0032] The point cloud image processing unit 12 uses a machine learning model, for example, composed of a neural network, to image the point cloud data acquired by the point cloud data acquisition processing unit 11 based on the rigid body transformation parameters and non-rigid body transformation parameters included in the projection parameters stored in the projection parameter storage unit 31. For example, it is configured as follows.
[0033] Figure 4 is a block diagram showing the functional configuration of the point cloud imaging processing unit 12. Specifically, the point cloud imaging processing unit 12 includes a virtual camera placement processing unit 121, a point cloud coloring processing unit 122, a rasterization processing unit 123, and an image correction processing unit 124 as its processing functions.
[0034] The virtual camera placement processing unit 121 sets the placement and characteristics of the virtual camera based on the rigid body transformation parameters and non-rigid body transformation parameters stored as projection parameters in the projection parameter storage unit 31, and generates images by projecting point cloud data using the virtual camera in this state.
[0035] The point cloud coloring processing unit 122 colors the point cloud of the image based on the conditions that specify the illumination position in the target space, which are stored in the imaging condition storage unit 32.
[0036] The rasterization processing unit 123 pixels the image after the colorization process based on the conditions specifying the resolution and point cloud radius stored in the image conversion condition storage unit 32.
[0037] The image correction processing unit 124 performs sharpening and denoising on the image pixelated by the rasterization processing unit 123 based on the conditions for sharpening and denoising stored in the image conversion condition storage unit 32.
[0038] The image feature extraction processing unit 13 uses a trained machine learning model, whose weight parameters are stored in the image feature extraction model storage unit 33, to extract features from each image generated by the point cloud imaging processing unit 12 for each point cloud class. The machine learning model for feature extraction is, for example, composed of a neural network.
[0039] The point cloud class label identification processing unit 14 estimates (identifies) the class label of the point cloud based on the feature quantities extracted from the images corresponding to each class by the image feature extraction processing unit 13.
[0040] The class label output processing unit 15 calculates the confidence level of the class labels identified for each class from the image features by the point cloud class label identification processing unit 14, selects the class label with the highest confidence level, and transmits the selected class label from the input / output I / F unit 4 to the requesting terminal device US.
[0041] During the learning phase, the learning processing unit 16 optimizes the rigid and non-rigid body transformation parameters used as projection parameters by the point cloud imaging processing unit 12 when given multiple point cloud data for learning and corresponding class labels for verification. The learning processing unit 16 then stores the optimized rigid and non-rigid body transformation parameters for each point cloud class in the projection parameter storage unit 31. An example of the learning process will be explained in the operation example.
[0042] (Example of operation) Next, we will explain an example of the operation of the point cloud data recognition device SV configured as described above.
[0043] (1) Learning Phase The point cloud data recognition device SV first performs a learning process in the point cloud imaging processing unit 12 to optimize the projection parameters used to image the point cloud data for each point cloud class.
[0044] Figure 5 is a flowchart showing an example of the processing procedure and content of the projection parameter optimization process performed by the control unit 1 of the point cloud data recognition device SV during the learning phase.
[0045] In step S10, the control unit 1 of the point cloud data recognition device SV first acquires, under the control of the point cloud data acquisition processing unit 11, multiple training point cloud data sets and corresponding verification class labels from the terminal device US, which have been prepared in advance, for example by the system administrator.
[0046] Once the above-mentioned training point cloud data and validation class labels are acquired, the control unit 1 of the point cloud data recognition device SV performs the following process to optimize the projection parameters under the control of the learning processing unit 16.
[0047] For example, suppose we are given a set of data consisting of training point cloud data and validation class labels (P_1, y_1), (P_2, y_2), ..., (P_N, y_N). Then the learning processing unit 16 first uses the attitude transformation matrix, such as Euler angles or quaternions, and the position vector as design variables, and uses the classification accuracy or cross-entropy error of the class labels identified by the point cloud class label identification processing unit 14 as the target variable to optimize the rigid body transformation parameters.
[0048] In addition, the learning processing unit 16 optimizes the non-rigid body transformation parameters, using, for example, a matrix for affine transformation of a two-dimensional image of the projection plane of a virtual camera as a design variable, and the classification accuracy or cross-entropy error of the class labels identified by the point cloud class label identification processing unit 14 as the target variable.
[0049] The optimized rigid-body and non-rigid-body transformation parameters are represented as different components in a single, unified transformation matrix.
[0050] Furthermore, the above optimization may use gradient-based optimization algorithms such as the steepest descent method or Newton's method, or heuristic optimization algorithms such as genetic algorithms, evolutionary strategy algorithms, or particle swarm optimization.
[0051] The above learning process is executed in the control unit 1 of the point cloud data recognition device SV as follows. Specifically, first in step S11, one of several point cloud classes is selected. Then, in step S12, a process is performed to randomly generate P sets of candidate solutions for the components of the transformation matrix representing the projection parameters, as follows.
[0052] Figure 6 is a flowchart showing an example of the processing steps and contents of this solution candidate generation process.
[0053] Specifically, in step S121, the point cloud imaging processing unit 12 generates an image by projecting the training point cloud data according to the transformation matrix of projection parameters. Next, in step S122, the image feature extraction processing unit 13 extracts features from the image using a feature extractor composed of a trained machine learning model stored in the image feature extraction model storage unit 33. For extracting image features, a feature extractor based on edge patterns or grayscale may be used, or a feature extractor pre-trained by a deep learning method such as ResNet or Vision Transformer may be used.
[0054] Next, in step S123, the point cloud class label discrimination processing unit 14 estimates the class labels of the given point cloud data based on the extracted image features. As an estimation method, for example, statistical machine learning algorithms such as the nearest neighbor method, Support Vector Machine (SVM), or Random Forest may be used, or a method that estimates the classification result using a fully connected layer of a neural network may be used.
[0055] The estimated class labels are passed from the point cloud class label identification processing unit 14 to the learning processing unit 16 in step S124.
[0056] In step S13, the learning processing unit 16 obtains a class label for the point cloud data of one of the candidate solutions for the projection parameters of group P. In step S14, it calculates the classification accuracy (identification accuracy) for the candidate solution for the projection parameters using the verification class label and stores it in the working memory area of the data storage unit 3.
[0057] Then, in step S15, the learning processing unit 16 determines whether the calculation and determination of classification accuracy for all solution candidates in group P has been completed. If it has not been completed, step S16 selects the next solution candidate and executes the calculation and determination of classification accuracy in steps S13 to S15. Thereafter, the learning processing unit 16 repeatedly performs the calculation and determination of classification accuracy for all solution candidates in group P in the same manner.
[0058] In response to this, once the calculation and determination of classification accuracy for all solution candidates in group P is complete, the learning processing unit 16 determines in step S17 whether or not the selection of all classes has been completed. If there are any unselected classes remaining, the learning processing unit 16 returns to step S11 to select the next class, and for the selected class, in steps S12 to S17, it performs the calculation and determination of classification accuracy for each solution candidate in group P of the components of the transformation matrix representing the projection parameters mentioned earlier. Similarly, the learning processing unit 16 repeatedly performs the calculation and determination of classification accuracy for each solution candidate in group P for each unselected class.
[0059] Once the calculation and determination of the classification accuracy for P sets of solution candidates representing the components of the transformation matrix that display the projection parameters has been completed for all classes, the learning processing unit 16 stores the rigid body transformation parameters and non-rigid body transformation parameters corresponding to the components of the transformation matrix that have the highest classification accuracy for each class in the projection parameter storage unit 31, associating them with the corresponding classes.
[0060] (2) Estimation Phase Once the rigid body transformation parameters and non-rigid body transformation parameters, which have been optimized during the learning phase, have been set, the control unit 1 of the point cloud data recognition device SV performs the process of recognizing class labels from the measured point cloud data as follows.
[0061] Figure 7 is a flowchart showing an example of the processing procedure and processing content of the class label recognition process performed by the control unit 1 of the point cloud data recognition device SV.
[0062] In other words, in step S20, the control unit 1 of the point cloud data recognition device SV first acquires point cloud data measured for the target space by the point cloud data measurement device MS under the control of the point cloud data acquisition processing unit 11 via the input / output I / F unit 4. The target space includes the object to be recognized.
[0063] In step S21, the control unit 1 of the point cloud data recognition device SV then performs the following process to image the point cloud data according to the projection parameters stored in the projection parameter storage unit 31, under the control of the point cloud imaging processing unit 12.
[0064] Figure 8 is a flowchart showing an example of the processing procedure and processing content of the image processing performed by the point cloud image processing unit 12 described above.
[0065] In other words, in step S211, the point cloud imaging processing unit 12 first determines the arrangement of the virtual camera in the point cloud coordinate system and the distortion characteristics of the camera lens based on the rigid body transformation parameters and non-rigid body transformation parameters stored as projection parameters for each class in the projection parameter storage unit 31. That is, it sets the arrangement of the virtual camera and the distortion characteristics of the camera lens to an optimized state corresponding to the point cloud class. Then, the point cloud imaging processing unit 12 generates point cloud images projected by the virtual camera for each point cloud class.
[0066] Next, in step S212, the point cloud imaging processing unit 12 reads conditions specifying the position of illumination in the target space from the imaging condition storage unit 32, and performs point cloud coloring processing on the generated point cloud image based on the read conditions. Subsequently, in step S213, the point cloud imaging processing processing unit 12 reads conditions specifying the resolution and point cloud radius from the imaging condition storage unit 32, and performs pixelation processing on the image after the coloring processing based on the read conditions. Then, in step S214, the point cloud imaging processing processing unit 12 reads conditions specifying the targets for sharpening and denoising from the imaging condition storage unit 32, and performs correction processing for sharpening and denoising on the pixelated image according to the read conditions, and outputs the corrected image to the image feature extraction processing unit 13.
[0067] In step S22, the control unit 1 of the point cloud data recognition device SV, under the control of the image feature extraction processing unit 13, inputs the point cloud images output by the point cloud imaging processing unit 12 for each point cloud class into the feature extractor, which is stored in the image feature extraction model storage unit 33 and consists of a trained machine learning model. The control unit then obtains the feature quantities of the point cloud images for each point cloud class from this feature extractor.
[0068] As mentioned in the learning phase, feature extractors based on edge patterns or grayscale may be used for extracting image features, or feature extractors pre-trained using deep learning methods such as ResNet or Vision Transformer may be used.
[0069] The control unit 1 of the point cloud data recognition device SV then, in step S23, under the control of the point cloud class label identification processing unit 14, estimates (identifies) the class label of the measured point cloud data based on the feature quantities of the point cloud image obtained for each point cloud class, and calculates the confidence level of the estimated class label.
[0070] As for the estimation method, as mentioned in the learning phase, statistical machine learning algorithms such as the nearest neighbor method, Support Vector Machine (SVM), and Random Forest may be used, or a method that estimates the classification result using the fully connected layers of a neural network may be used.
[0071] In step S24, the control unit 1 of the point cloud data recognition device SV, under the control of the class label output processing unit 15, calculates, for example, the average confidence score based on the confidence scores of the class labels of each class obtained by the point cloud class label identification processing unit 14, and selects the class label of the class with the highest value. The selected class label is then transmitted from the input / output I / F unit 4 to, for example, the system administrator's or user's terminal device US.
[0072] Figure 9 is used to explain the process of estimating and outputting the class label with the highest confidence level from the point cloud data described above.
[0073] This example demonstrates the process of acquiring point cloud data of a target space containing an object called "chair" and recognizing its class label. Specifically, C classes, Class1 to ClassC, are defined as point cloud classes, and the rigid and non-rigid body transformation parameters corresponding to each of these C classes, Class1 to ClassC, are pre-optimized and set as projection parameters R1 to Rc during the learning phase. Then, images are generated by projecting the point cloud data using these projection parameters R1 to Rc, and features are extracted from each generated image. Based on these features, the class label is identified. Finally, the average confidence score Pi (i=1 to C) of each identified class label is calculated. Pi = 1 / C(Pi R1 +Pi R2 +…+Pi Rc ) Each is calculated using the following method, and the argmaxPi that has the largest average confidence score Pi of the calculated class label is selected and output.
[0074] (effect) As described above, in the first embodiment, during the learning phase, multiple training point cloud data and corresponding validation class labels are used to set rigid body transformation parameters and non-rigid body transformation parameters optimized for each class so as to minimize evaluation metrics such as cross-entropy loss, and these are stored as projection parameters in the projection parameter storage unit 31. Then, during the estimation phase, point cloud data measured for the target space is acquired, and images are generated by projecting the acquired point cloud data according to the rigid body transformation parameters and non-rigid body transformation parameters set as projection parameters for each class. Features are extracted from each generated image, and class labels are estimated based on the extracted features. The class label with the highest confidence level among the estimated class labels is then output.
[0075] Therefore, images are generated in which the point cloud data is projected using rigid and non-rigid transformation parameters optimized for each class during the learning phase. In other words, the projection plane, including the arrangement of the virtual camera and lens distortion, is optimized for each class, and the point cloud data is imaged using the line of sight of this virtual camera. Then, feature quantities that appropriately reflect the features of the point cloud are extracted from the generated image, and the class labels of the point cloud data are accurately recognized based on the extracted feature quantities.
[0076] As a result, it becomes possible to recognize class labels in point cloud data based on images generated by considering not only rigid body deformations such as the pose and position of a virtual camera, but also non-rigid body deformations such as lens distortion characteristics and shear. Therefore, it becomes possible to recognize class labels with higher accuracy. Furthermore, as a result, it can be expected that object recognition accuracy will be improved when recognizing objects in three-dimensional space while changing the line of sight, for example, in autonomous mobile robots and drones.
[0077] [Second Embodiment] In the first embodiment, we described a case in which point cloud data is projected and imaged using projection parameters in which rigid body transformation parameters and non-rigid body transformation parameters are expressed as different components of the transformation matrix.
[0078] In contrast, the second embodiment of this invention optimizes rigid body transformation parameters and non-rigid body transformation parameters individually during the learning phase and stores them in the projection parameter storage unit 31.
[0079] Figure 10 is a flowchart showing an example of the processing procedure and processing content of the projection parameter optimization process performed by the control unit 1 of the point cloud data recognition device SV according to the second embodiment of this invention. In Figure 10, steps whose processing content is the same as those described in Figure 5 in the first embodiment are denoted by the same reference numerals.
[0080] In step S10, the control unit 1 of the point cloud data recognition device SV acquires training point cloud data and validation class labels. Under the control of the learning processing unit 16, it first selects one of several point cloud classes in step S11. Then, in step S31, the control unit 1 of the point cloud data recognition device SV randomly generates P sets of candidate solutions for rigid body transformations and images obtained by non-rigid body transformations, as follows.
[0081] Figure 11 is a flowchart illustrating an example of the processing procedure and content of this solution candidate generation process. Steps in Figure 11 that perform the same processing as in Figure 6 are denoted by the same reference numerals and explained accordingly.
[0082] Specifically, in step S311, the control unit 1 of the point cloud data recognition device SV first projects the training point cloud data onto an image using a rigid body transformation such as Euler angles or quaternions, under the control of the point cloud image processing unit 12. Subsequently, in step S312, the image generated by the above rigid body transformation is deformed using a non-rigid body transformation such as the affine transformation of a two-dimensional image.
[0083] Next, in step S122, the control unit 1 of the point cloud data recognition device SV uses the image feature extraction processing unit 13 to extract features from the images to which the rigid body transformation and non-rigid body transformation have been applied, using a feature extractor composed of a trained machine learning model stored in the image feature extraction model storage unit 33. Subsequently, in step S123, the point cloud class label identification processing unit 14 estimates the class label of the point cloud data based on the extracted image features. Then, in step S124, the estimated class label is passed to the learning processing unit 16.
[0084] In step S13, the learning processing unit 16 obtains a class label for the point cloud data of one of the solution candidates in group P. In step S14, it calculates the classification accuracy for the solution candidate using the verification class label and stores it in the working memory area of the data storage unit 3.
[0085] Then, in step S15, the learning processing unit 16 determines whether the calculation and determination of classification accuracy for all solution candidates in group P has been completed. If it has not been completed, step S16 selects the next solution candidate and executes the calculation and determination of classification accuracy in steps S13 to S15. Thereafter, the learning processing unit 16 repeatedly performs the calculation and determination of classification accuracy for all solution candidates in group P in the same manner.
[0086] In response to this, once the calculation and determination of classification accuracy for all solution candidates in group P is complete, the learning processing unit 16 determines in step S17 whether or not the selection of all classes has been completed. If there are any unselected classes remaining, the learning processing unit 16 returns to step S11 to select the next class, and for the selected class, in steps S12 to S17, it performs the calculation and determination of classification accuracy for each solution candidate in group P of the images to which the rigid body transformation and non-rigid body transformation described above have been applied. Similarly, the learning processing unit 16 repeatedly performs the calculation and determination of classification accuracy for each solution candidate in group P for each unselected class.
[0087] Then, once the calculation and determination of the classification accuracy for the P sets of solution candidates for all classes to which rigid and non-rigid transformations have been applied is completed, the learning processing unit 16 stores the rigid and non-rigid transformation parameters corresponding to the solution candidate with the highest classification accuracy for each class in the projection parameter storage unit 31, linked to the corresponding class, in step S18.
[0088] Subsequently, in the estimation phase, similar to the first embodiment, the point cloud data is imaged using the rigid body transformation parameters and non-rigid body transformation parameters stored in the projection parameter storage unit 31, and the class label of the point cloud data is identified based on the features extracted from this image.
[0089] According to the second embodiment, when projecting point cloud data to create an image, rigid body transformation parameters that transform the attitude and position of the virtual camera and non-rigid body transformation parameters that non-rigidize the projection plane of the virtual camera can be optimized separately.
[0090] [Other embodiments] In the first and second embodiments, the cases in which each function of the control unit 1 of the point cloud data recognition device SV is provided on a single server computer or personal computer were described as examples. However, each processing function of the control unit 1 of the point cloud data recognition device SV may be distributed across multiple server computers or personal computers.
[0091] Furthermore, the types and configurations of machine learning models that realize each processing function of the control unit 1 of the point cloud data recognition device SV, the processing procedures and content of the image processing, image feature extraction processing, and class label identification processing, the types of objects, etc., can be modified in various ways without departing from the spirit of this invention.
[0092] Although embodiments of this invention have been described in detail above, the above description is merely illustrative in all respects. It goes without saying that various improvements and modifications can be made without departing from the scope of this invention. In other words, when implementing this invention, specific configurations may be adopted as appropriate depending on the embodiment.
[0093] In short, this invention is not limited to the embodiments described above, and in the implementation stage, the components can be modified and materialized without departing from the gist of the invention. Furthermore, various inventions can be formed by appropriately combining the multiple components disclosed in the embodiments. For example, some components may be deleted from all the components shown in the embodiments. Moreover, components from different embodiments may be appropriately combined. [Explanation of Symbols]
[0094] SV…Point Cloud Data Recognition Device MS... Point Cloud Data Measurement Device US…terminal device NW...Network 1…Control Unit 2…Program memory 3…Data storage unit 4…Input / Output I / F section 5... Bus 11…Point cloud data acquisition processing unit 12…Point cloud image processing unit 13…Image Feature Extraction Processing Unit 14…Point cloud class label identification processing unit 15...Class label output processing unit 16…Learning Processing Unit 31...Projection parameter storage unit 32…Image conversion condition storage unit 33…Image Feature Extraction Model Memory Unit
Claims
1. A first processing unit that acquires point cloud data of the target space, A second processing unit that holds pre-set projection parameters and generates an image by projecting the point cloud data according to the projection parameters, A third processing unit that extracts features from the aforementioned image, A fourth processing unit recognizes the class of the imaged point cloud data based on the aforementioned feature quantities. Equipped with, The projection parameters consist of rigid transformation parameters that transform the arrangement of the virtual camera that projects the point cloud data, and non-rigid transformation parameters that transform the projection plane of the virtual camera. Point cloud data recognition device.
2. The point cloud data recognition device according to claim 1, further comprising a fifth processing unit that optimizes the rigid body transformation parameters and non-rigid body transformation parameters based on a plurality of training point cloud data prepared corresponding to the target space and a verification class set corresponding to the training point cloud data.
3. A point cloud data recognition method performed by an information processing device, The process of acquiring point cloud data of the target space, A process of generating an image by projecting the point cloud data according to the projection parameters, while maintaining pre-set projection parameters, The process of extracting features from the aforementioned image, A process for recognizing the class of the imaged point cloud data based on the aforementioned features, Equipped with, The process of generating the aforementioned image uses, as projection parameters, rigid transformation parameters that transform the arrangement of the virtual camera onto which the point cloud data is projected, and non-rigid transformation parameters that transform the projection plane of the virtual camera. Point cloud data recognition method.
4. A program that causes a processor in a point cloud data recognition device to perform the processing performed by each processing unit in the point cloud data recognition device described in claim 1 or 2.