Point cloud merging device and point cloud merging method

The point cloud combining device and method address the inefficiencies of conventional alignment techniques by using constrained alignment processes, ensuring accurate alignment of point cloud data without physical markers, thus enhancing efficiency and precision.

WO2026004279A1PCT designated stage Publication Date: 2026-01-02PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/012958
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-28
Filing Date
2025-03-28
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Conventional point cloud data alignment techniques require tedious preparations such as placing markers at the target location and measuring their installation positions, leading to potential inaccuracies and inefficiencies, especially when aligning data from large or complex locations.

Method used

A point cloud combining device and method that uses a processor to perform rough and fine alignment processes with constraints on the search space, eliminating the need for physical markers by determining alignment conditions through processes like RANSAC and Point-to-Plane ICP, ensuring accurate alignment without cumbersome preparations.

Benefits of technology

This approach reduces processing load and prevents issues like tilted floor surfaces, enabling easy and accurate alignment of point cloud data without the need for advance marker placement, thus improving efficiency and precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025012958_02012026_PF_FP_ABST
    Figure JP2025012958_02012026_PF_FP_ABST
Patent Text Reader

Abstract

[Problem] To easily and accurately perform registration of point cloud data without requiring tiresome preparation, such as installing a marker in a subject place in advance and measuring the installation position, for registration when merging a plurality of point cloud data sets obtained by three-dimensional reconstruction processing of the subject place. [Solution] The present invention executes: approximate registration processing (ST204) for acquiring an approximate registration condition while limiting a search space, on two extracted point cloud data sets obtained by extracting a common area from each of two point cloud data sets; precise registration processing (ST205) for acquiring a precise registration condition while limiting the search space, on the two extracted point cloud data sets registered using the approximate registration condition; and point cloud merging processing for merging the two point cloud data sets registered using the precise registration condition.
Need to check novelty before this filing date? Find Prior Art

Description

Point cloud combining device and point cloud combining method

[0001] The present disclosure relates to a point cloud merging device and a point cloud merging method that align and merge two pieces of point cloud data acquired by three-dimensional reconstruction processing of a target location.

[0002] A three-dimensional reconstruction technique is known that generates point cloud data as three-dimensional spatial information about a target location based on photographed images of the target location. Among these three-dimensional reconstruction techniques, the SLAM (Simultaneous Localization and Mapping) method has recently attracted attention. The SLAM method allows a mobile object (e.g., a worker) to carry a photographing device and capture images of the target location while moving around the target location, thereby acquiring point cloud data of the target location.

[0003] On the other hand, when the target location is large, such as a factory, or when the target location is divided into complex sections, such as a house, it is not possible to capture the entire target location at once. In such cases, the user must capture the target location multiple times. Furthermore, a process is performed in which multiple point cloud data generated by capturing partial images of the target location are combined to generate combined point cloud data representing the entire target location. In this case, since the positional relationships between the multiple point cloud data cannot be accurately recognized, combining the multiple point cloud data requires alignment to align common areas where the same subject is captured in the multiple point cloud data.

[0004] A known technique for aligning multiple point cloud data is to display a point cloud image that visualizes the point cloud data on a screen and allow a user to manipulate the point cloud image so that the common areas of the two point cloud images overlap (see Patent Document 1). In this technique, multiple markers (targets) that serve as references for alignment are placed at target locations, and the markers are displayed on the point cloud image, allowing the user to perform alignment operations based on the markers.

[0005] Japanese Patent Application Laid-Open No. 2018-147065

[0006] However, conventional techniques have the problem that it is necessary to place multiple markers (targets) that serve as alignment references at the target location in advance and measure the installation positions of the markers using surveying equipment (such as a total station), which is extremely time-consuming. Also, if alignment is not performed with sufficient precision, point cloud data may be combined in a state where the floor surface of the target location is tilted. However, conventional techniques do not take into consideration how to prevent such problems, and therefore have the problem of being unable to prevent problems caused by insufficient alignment precision.

[0007] Therefore, the main objective of the present disclosure is to provide a point cloud combining device and a point cloud combining method that can easily and accurately align point cloud data when combining multiple point cloud data acquired by 3D reconstruction processing of the target location, without the need for tedious preparations such as placing markers at the target location in advance and measuring their installation positions.

[0008] The point cloud combining device disclosed herein is a point cloud combining device that uses a processor to execute a process of combining two point cloud data acquired by a three-dimensional reconstruction process of a target location after aligning them with each other, and the processor is configured to execute a rough alignment process on two extracted point cloud data obtained by extracting a common area from each of the two point cloud data, to obtain rough alignment conditions while imposing constraints on a search space, execute a fine alignment process on the two extracted point cloud data that have been aligned using the rough alignment conditions, to obtain fine alignment conditions while imposing constraints on the search space, and execute a point cloud combining process to combine the two point cloud data that have been aligned using the fine alignment conditions.

[0009] Furthermore, the point cloud combining method of the present disclosure is a point cloud combining method that causes a processor to perform a process of combining two point cloud data acquired by a three-dimensional reconstruction process of a target location after mutual alignment, and is configured to perform a rough alignment process on two extracted point cloud data obtained by extracting a common area from each of the two point cloud data, to obtain rough alignment conditions while imposing constraints on a search space, perform a fine alignment process on the two extracted point cloud data that have been aligned using the rough alignment conditions, to obtain fine alignment conditions while imposing constraints on the search space, and perform a point cloud combining process that combines the two point cloud data that have been aligned using the fine alignment conditions.

[0010] According to the present disclosure, the rough alignment conditions and the fine alignment conditions are determined with constraints imposed on the search space, which reduces the processing load and prevents problems such as combining point cloud data when the floor surface of the target location is tilted. This eliminates the need for cumbersome preparations such as installing markers at the target location in advance and measuring their installation positions, and enables easy and accurate alignment of point cloud data.

[0011] FIG. 1 is an explanatory diagram showing the status of a photographing operation performed by a user using a photographing system according to the present embodiment; FIG. 1 is an explanatory diagram showing an example of point cloud joining processing performed by a control device;

[0012] The first invention made to solve the above problem is a point cloud combining device that uses a processor to execute a process of combining two point cloud data acquired by a three-dimensional reconstruction process of a target location after aligning them with each other, wherein the processor is configured to execute a rough alignment process for two extracted point cloud data obtained by extracting a common area from each of the two point cloud data, to obtain rough alignment conditions while imposing constraints on a search space, to execute a fine alignment process for the two extracted point cloud data that have been aligned using the rough alignment conditions, to obtain fine alignment conditions while imposing constraints on the search space, and to execute a point cloud combining process that combines the two point cloud data that have been aligned using the fine alignment conditions.

[0013] This allows the rough alignment conditions and the fine alignment conditions to be determined with restrictions placed on the search space, thereby reducing the processing load and preventing problems such as combining point cloud data when the floor of the target location is tilted. This eliminates the need for cumbersome preparations such as placing markers at the target location in advance and measuring their positions, and allows for easy and accurate alignment of point cloud data.

[0014] In addition, in a second invention, the processor is configured to perform the following preprocessing of the approximate alignment process: estimating normals for each point included in each of the two extracted point cloud data, or re-estimating normals if normal information is already available; calculating feature values ​​for each point included in each of the two extracted point cloud data based on the normals of the points; and comparing feature values ​​for each point included in each of the two extracted point cloud data to obtain corresponding points that associate points with similar features, and then perform the approximate alignment process based on the corresponding points.

[0015] This allows the rough alignment process to be performed appropriately.

[0016] In a third aspect of the present invention, the processor acquires the rough alignment conditions using a RANSAC method in the rough alignment process.

[0017] This makes it possible to eliminate the influence of outliers and obtain appropriate rough alignment conditions.

[0018] In a fourth aspect of the present invention, the processor is configured to acquire a rotation element of the rough alignment condition in the rough alignment process, with the rotation being limited to a horizontal plane as a constraint on the search space.

[0019] According to this, the rotation element of the approximate alignment condition is limited to rotation on a horizontal plane (around a vertical axis), which reduces the processing load and also prevents the problem of combining point cloud data when the floor surface of the target location is tilted, because alignment is performed so that the vertical direction of the point cloud does not change. In this case, the rotation element may be obtained by projecting corresponding points of the two extracted point cloud data onto a horizontal plane and comparing the corresponding points of the two extracted point cloud data.

[0020] In a fifth aspect of the present invention, the processor acquires the precise alignment conditions using a Point-to-Plane ICP technique in the precise alignment process.

[0021] This allows appropriate precision alignment conditions to be obtained efficiently.

[0022] In a sixth aspect of the present invention, the processor is configured to acquire a rotation element of the precise registration condition in the precise registration process, with the rotation being limited to a horizontal plane as a constraint on the search space.

[0023] This reduces the processing load because the rotation element of the precision alignment condition is limited to rotation on a horizontal plane (around the vertical axis), and alignment is performed so that the vertical direction of the point cloud does not change, preventing the problem of point cloud data being combined when the floor of the target location is tilted.

[0024] A seventh invention is a point cloud merging method that causes a processor to perform a process of merging two point cloud data acquired by a three-dimensional reconstruction process of a target location after aligning them with each other, and is configured to perform a rough alignment process on two extracted point cloud data obtained by extracting a common area from each of the two point cloud data, to obtain rough alignment conditions while restricting a search space, to perform a fine alignment process on the two extracted point cloud data that have been aligned using the rough alignment conditions, to obtain fine alignment conditions while restricting a search space, and to perform a point cloud merging process that merges the two point cloud data that have been aligned using the fine alignment conditions.

[0025] As with the first invention, this eliminates the need for tedious preparations such as placing markers at the target location in advance and measuring their positions in order to align multiple point cloud data obtained by 3D reconstruction processing of the target location when combining them, and allows for easy and accurate alignment of the point cloud data.

[0026] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.

[0027] FIG. 1 is an explanatory diagram showing the situation of a photographing operation performed by a user using a photographing system according to this embodiment.

[0028] As shown in FIG. 1, the photography system includes a photography device 1 and a control device 2 (point cloud combining device).

[0029] The photographing device 1 includes a sensor unit 11, a display / input panel unit 12, and a rod-shaped support member 13. The sensor unit 11 and the display / input panel unit 12 of the photographing device 1 are connected via a first cable 18. The display / input panel unit 12 of the photographing device 1 and the control device 2 are connected via a second cable 19.

[0030] The control device 2 is configured as a laptop or tablet PC that can be carried by a user (worker). In the example shown in Fig. 1, the control device 2 can be stored in a shoulder bag and carried by the user, but the manner in which the user carries the control device 2 is not limited to this.

[0031] A user performs a photographing task by walking around a target location while holding the photographing device 1 in his / her hand and causing the photographing device 1 to photograph the target location. At this time, the user can change the position (height) of the sensor unit 11 by moving the arm holding the photographing device 1.

[0032] Next, a description will be given of the point cloud combining process performed by the control device 2. Fig. 2 is an explanatory diagram showing an example of the point cloud combining process.

[0033] When the target location is large, such as a factory, or when the target location is divided into complex sections, such as a house, it is difficult to capture the entire target location at once. In such cases, the user takes multiple photographs of the target location. The control device 2 also performs a process of combining multiple point cloud data generated by partially photographing the target location to generate combined point cloud data representing the entire target location.

[0034] In the example shown in FIG. 2, the target location is a house. The house is equipped with a living room, kitchen, bedroom, toilet, bathroom, etc. In this example, the interior of the house is photographed four times, and four point cloud data #1 to #4 are generated. The point cloud data #1 and #2 are combined to generate combined point cloud data #S1, and the point cloud data #3 and #4 are combined to generate combined point cloud data #S2. The combined point cloud data #S1 and #S2 are further combined to generate combined point cloud data #S3, which represents the entire target location. Note that the order in which multiple point cloud data are combined is not limited to this example.

[0035] Next, a description will be given of the general configuration of the photographing device 1 and the control device 2. Fig. 3 is a block diagram showing the general configuration of the photographing device 1 and the control device 2. Fig. 4 is a block diagram showing an overview of the processing performed by the processor 46 of the control device 2.

[0036] As shown in FIG. 3, the sensor unit 11 of the image capturing device 1 includes a visible camera 21, a depth sensor 22, an IMU (Inertial Measurement Unit) 23, and an input / output interface 24.

[0037] The visible light camera 21 (color camera) takes color photographs and outputs color photographed images.

[0038] The depth sensor 22 outputs depth information (distance image) as a detection result based on images captured by left and right infrared cameras (not shown).

[0039] The IMU 23 detects the motion state of the device itself, specifically, three-dimensional angular velocity and acceleration. Based on the detection results of the IMU 23, the position, orientation, and velocity of the sensor unit 11 can be detected.

[0040] The input / output interface 24 inputs and outputs data to and from the control device 2 via the display / input panel unit 12. Specifically, the input / output interface 24 transmits detection data from the visible camera 21, the depth sensor 22, and the IMU 23. The input / output interface 24 may be based on the USB (registered trademark) standard.

[0041] The visible camera 21, the depth sensor 22, and the IMU 23 may not be integrated. Alternatively, the depth sensor 22 and the IMU 23 may be omitted, and only the visible camera 21 may be provided. Alternatively, the visible camera 21 may be provided together with either the depth sensor 22 or the IMU 23.

[0042] The display / input panel unit 12 of the photographing device 1 includes a touch panel display 31 , a repeater 32 , and an input / output interface 33 .

[0043] The touch panel display 31 displays a screen that assists the user in taking photographs, etc., under the control of the control device 2.

[0044] The repeater 32 repeats data communication between the control device 2 and the sensor unit 11. The repeater 32 may be a hub based on the USB (registered trademark) standard.

[0045] The input / output interface 33 inputs and outputs data to and from the control device 2. Specifically, the input / output interface 33 receives display information such as a screen that assists the user in taking photographs from the control device 2. The input / output interface 33 may be based on the USB (registered trademark) standard.

[0046] The control device 2 includes an input / output interface 41 , a display 42 , an input device 43 , a memory 44 , a storage device 45 , and a processor 46 .

[0047] The input / output interface 41 inputs and outputs data to and from the image capturing device 1. The input / output interface 41 may be based on the USB (registered trademark) standard.

[0048] The display 42 displays a screen for managing previously acquired photographic data, a screen for setting the operating conditions of the photographic device 1, a screen for combining point cloud data, and the like.

[0049] The input device 43 is used by the user to perform input operations. The input device 43 may be a keyboard, a mouse, a touchpad, a touch panel, etc. If the control device 2 is configured as a tablet PC, a touch panel display is provided in which the touch panel as the input device 43 and the display panel as the display 42 are integrated.

[0050] The memory 44 stores programs executed by the processor 46 and the like.

[0051] The storage device 45 stores the shooting data (shooting information) acquired from the image capturing device 1. The shooting data includes the image captured by the visible camera 21, the detection results (distance information) of the depth sensor 22, the detection results of the IMU 23, the shooting time, etc. The storage device 45 also stores point cloud data generated by the processor 46 as the three-dimensional measurement results.

[0052] The processor 46 performs various processes by executing programs stored in the memory 44. In this embodiment, the processor 46 performs a three-dimensional reconstruction process P1.

[0053] In the three-dimensional reconstruction process P1, the processor 46 generates point cloud data (environment map) as three-dimensional spatial information related to the target location using the SLAM (Simultaneous Localization And Mapping) method based on images captured by the visible camera 21. In the three-dimensional reconstruction process P1, self-position estimation is performed in addition to the generation of the point cloud data, and the self-position for each time, i.e., the position of the photographing point, is acquired.

[0054] As shown in FIG. 4, the three-dimensional reconstruction process P1 includes a feature extraction process P11, a tracking process P12, a position and orientation correction process P13, and a point cloud generation process P14.

[0055] In the feature extraction process P11, the processor 46 extracts feature information (feature points, etc.) from the image (frame) captured by the visible camera 21.

[0056] In tracking processing P12, the processor 46 compares the currently extracted feature points with previously extracted feature points to estimate the amount of transition related to the position and orientation of the image capture device 1, and updates the position and orientation trajectory data based on the amount of transition. Note that the position and orientation trajectory data is obtained by tracking the position and orientation of the image capture device 1, and includes the tracking results of the position and orientation at the time of capturing each captured image (frame), i.e., information regarding the position and orientation of the image capture device 1 at each time of capturing.

[0057] In the position and orientation correction process P13, the processor 46 corrects the position and orientation trajectory data acquired in the tracking process P12 based on the detection data of the IMU 23. Here, the position and orientation trajectory data is corrected so as to compensate for the measurement results of locations with few features, such as walls and ceilings.

[0058] In the point cloud generation process P14, the processor 46 generates point cloud data based on the distance information from the depth sensor 22 and the position and orientation trajectory data acquired in the tracking process P12 and the position and orientation correction process P13. The point cloud generation process P14 includes a process (standard point cloud generation process) for generating regular point cloud data acquired as a 3D measurement result, and a process (simplified point cloud generation process) for generating simplified point cloud data that allows the user to check the shooting conditions (point cloud generation conditions). The simplified point cloud generation process must be performed in real time during shooting, but the standard point cloud generation process may be performed after shooting.

[0059] The processor 46 also performs a point cloud alignment process P2, a point cloud joining process P3, a drawing generation process P4, a display information generation process P5, and a display process P6.

[0060] In the point cloud alignment process P2, the processor 46 aligns the common areas of the multiple point cloud data to be combined. Specifically, as an alignment condition for aligning the common areas of the two point cloud data on the source side (alignment source) and the target side (alignment destination), a transformation matrix for coordinate transformation of each point of the source side point cloud data is calculated, and the transformation matrix is ​​applied to the source side point cloud data to perform coordinate transformation of the source side point cloud data.

[0061] In the point cloud combining process P3, the processor 46 combines the point cloud data that have been aligned in the point cloud alignment process P2 to generate a single combined point cloud data that represents the entire target location.

[0062] In the drawing generation process P4, the processor 46 generates a 3D model representing the entire target location based on the point cloud data representing the entire target location generated by the 3D reconstruction process P1, and generates a layout drawing representing the entire target location based on the 3D model. At this time, a 2D layout drawing (plan view) representing the entire target location is generated by projecting the 3D model representing the entire target location onto a horizontal plane. In addition, a 3D layout drawing representing the entire target location is generated by projecting the 3D model representing the entire target location based on a predetermined line of sight.

[0063] In the display information generation process P5, the processor 46 generates display information for a screen to be displayed on the display 42 of the control device 2. The display 42 displays a screen on which the user performs operations such as managing the imaging data, issuing instructions for post-imaging processes (point cloud generation process P14, point cloud alignment process P2, point cloud joining process P3, and drawing generation process P4), and setting processing conditions. The processor 46 also generates display information for a screen to be displayed on the touch panel display 31 of the imaging device 1. The touch panel display 31 displays a screen to assist the user in the imaging operation.

[0064] In the display process P6, the processor 46 displays a screen on the display 42 of the control device 2 based on the display information generated in the display information generation process P5, and also displays a screen on the touch panel display 31 of the photographing device 1.

[0065] In this embodiment, the processing related to the merging of point cloud data is performed on a control device 2 that can be carried by a user, but the processing related to the merging of point cloud data may also be performed on another device (e.g., a server device) that can communicate with the control device 2.

[0066] Next, a description will be given of the procedure of the process performed by the control device 2. FIG.

[0067] First, the processor 46 acquires the point cloud data to be combined in response to a user's operation to specify a point cloud file in which the point cloud data to be combined is stored (ST101). The point cloud data generated by the 3D reconstruction process P1 is stored in the storage device 45.

[0068] Next, the processor 46 performs alignment to align the common areas of the multiple point cloud data to be combined (point cloud alignment process) (ST102). At this time, as an alignment condition for aligning the common areas of the two point cloud data on the source side and the target side, a transformation matrix for transforming the coordinates of each point of the source side point cloud data is calculated, and the transformation matrix is ​​applied to the source side point cloud data to perform coordinate transformation of the source side point cloud data.

[0069] Next, the processor 46 combines the point cloud data that have been aligned in the point cloud alignment process (ST102) to generate one combined point cloud data that represents the entire target location (point cloud combination process) (ST103). At this time, in addition to combining two point cloud data, if necessary, combining one combined point cloud data with another point cloud data or combining two combined point cloud data is also performed.

[0070] Next, a description will be given of the procedure of the point cloud alignment process performed by the control device 2. Fig. 6 is a flow diagram showing the procedure of the point cloud alignment process. Fig. 7 is a block diagram showing an outline of the point cloud alignment process.

[0071] Since multiple point cloud data sets generated by taking multiple photographs of the target location cannot accurately recognize their relative positions, the control device 2 performs a process (point cloud alignment process) to align the common areas of the multiple point cloud data sets when combining the multiple point cloud data sets.

[0072] As shown in Figures 6 and 7, first, the processor 46 displays a point cloud image on the screen of the display 42, which visualizes each of the multiple point cloud data to be combined, and acquires information (preliminary alignment information) regarding the relative positional relationship between the two overlapping point cloud images in response to a user operation to overlap the common areas of the two point cloud images (preliminary alignment process) (ST201).

[0073] Next, based on the preliminary alignment information acquired in the preliminary alignment process, the processor 46 extracts point clouds contained in the common area from each of the two point cloud data on the source side and the target side, and generates two extracted point cloud data on the source side and the target side (common area extraction process) (ST202).

[0074] Next, the processor 46 performs preprocessing (ST203) on the two extracted point cloud data on the source side and the target side extracted in the common area extraction process (ST202) to properly perform the next rough alignment process (ST204).

[0075] Next, the processor 46 determines the rough alignment conditions (transformation matrix) for roughly aligning the two extracted point cloud data on the source side and the target side, and applies the rough alignment conditions to the extracted point cloud data on the source side to perform coordinate transformation of the extracted point cloud data on the source side (rough alignment process) (ST204). By performing the rough alignment process in this way prior to the fine alignment process, it is possible to avoid the problem of getting stuck in a local solution in the fine alignment process and obtain appropriate fine alignment conditions.

[0076] Next, the processor 46 uses the extracted point cloud data that has been aligned using the rough alignment conditions (transformation matrix) acquired in the rough alignment process (ST204) to determine fine alignment conditions (transformation matrix) for precisely aligning the two point cloud data on the source side and the target side, and applies these fine alignment conditions to the point cloud data on the source side to perform coordinate transformation of the point cloud data on the source side (fine alignment process) (ST205).

[0077] Next, a description will be given of the preliminary alignment processing performed by the control device 2. Fig. 8 is a flow diagram showing the procedure of the preliminary alignment processing. Fig. 9 is an explanatory diagram showing a canvas screen 51 displayed on the display 42 of the control device 2 during the preliminary alignment processing.

[0078] As shown in Fig. 8, in the preliminary alignment process (ST201 in Fig. 6), the processor 46 first generates a point cloud image that visualizes the point cloud data in a bird's-eye view (point cloud image generation process) (ST301). Specifically, the point cloud image is generated by projecting each point of the point cloud data onto a two-dimensional horizontal plane. At this time, multiple points in the point cloud data may be compressed into one point to reduce the size of the point cloud image.

[0079] Next, the processor 46 displays the canvas screen 51 (point cloud operation screen) (see FIG. 9) on the display 42 (ST302).

[0080] At this time, as shown in FIG. 9 , a point cloud image 52 that visualizes the multiple point cloud data to be combined in a bird's-eye view is displayed on the canvas screen 51. In the example shown in FIG. 9 , point cloud images 52 #1 to #4 that visualize four point cloud data #1 to #4 to be combined are displayed. On the canvas screen 51, the user can manipulate the point cloud images 52 so that the common areas of the two point cloud images 52 overlap each other. Specifically, the user can manipulate the source-side point cloud image 52 so that the source-side point cloud image 52 is superimposed on the target-side point cloud image 52. At this time, the user can perform an operation to translate the point cloud image 52 and an operation to rotate the point cloud image 52 as a positioning operation.

[0081] Next, the processor 46 acquires information (preliminary alignment information) regarding the relative positional relationship between the two overlapping point cloud images in response to a user operation on the canvas screen 51 (ST303).

[0082] In addition, the canvas screen 51 is provided with a "Merge" button 53. When the user operates the "Merge" button 53, the point cloud alignment process (ST102 in FIG. 5) performs the preliminary alignment process (ST201 in FIG. 6) followed by the common area extraction process, pre-processing, rough alignment process, and fine alignment process (ST202 to ST205 in FIG. 6), as well as the point cloud merging process (ST103 in FIG. 5), and a merged point cloud image that visualizes the merged point cloud data as a result of the processing is displayed on the canvas screen 51. This allows the user to visually check whether the point cloud data has been merged appropriately.

[0083] The canvas screen 51 may also be provided with an operation unit (such as a button) for instructing enlargement and reduction of the point cloud image 52, or an operation unit for instructing a process of canceling the point cloud combining process and returning to the original state.

[0084] Furthermore, when the user operates the “Execute Merge” button 53, a simple point cloud merging process may be executed, and a merged point cloud image for the user to check the processing results may be displayed on the canvas screen 51. Furthermore, the canvas screen 51 may be provided with, for example, a “Save merged point cloud file” button (not shown) as an operation unit for the user to approve the merged results and instruct to save the data, and when the user operates this button, a regular point cloud merging process may be executed, and a merged point cloud file in which merged point cloud data is stored may be generated and saved in the storage device 45.

[0085] In this embodiment, the preliminary alignment process is performed based on a user operation, but the preliminary alignment process may also be performed without a user operation. Specifically, the processor 46 may perform a process such as image matching on the two point cloud images to detect similar areas in the two point cloud images as a common area (common area detection process). In this case, if there are three or more point cloud data that include a common area, the user may be able to select only two point cloud data to be subjected to the preliminary alignment process.

[0086] Next, a description will be given of the pre-processing performed by the control device 2. Fig. 10 is a flow chart showing the procedure of the pre-processing. Fig. 11 is an explanatory diagram showing an outline of the corresponding point acquisition process.

[0087] The control device 2 performs pre-processing (ST203 in Figure 6) on the two extracted point cloud data on the source side and the target side extracted in the common area extraction process to properly perform the next rough alignment process.

[0088] Specifically, first, processor 46 estimates normals for each point of the two extracted point cloud data on the source and target sides, or re-estimates normals if normal information already exists (normal re-estimation process) (ST401). In point cloud generation process P14 (see FIG. 4), a provisional normal is estimated for each point, and point cloud data including normal information for each point is generated. In the normal re-estimation process, re-estimating normals allows for the setting of highly accurate normals.

[0089] Next, the processor 46 calculates feature amounts (feature amount calculation process) (ST402) for each point of the two extracted point cloud data on the source side and the target side based on the normal information acquired in the normal re-estimation process (ST401). Specifically, FPFH (Fast Point Feature Histograms) feature amounts are calculated.

[0090] Next, the processor 46 compares the feature amounts of each point in the extracted point cloud data on the source side with those of each point in the extracted point cloud data on the target side, and finds a combination of points (corresponding points) that have the most similar features (corresponding point acquisition process) (ST403), as shown in Fig. 11. Note that an index (serial number) is assigned to each point as identification information for distinguishing each point in the extracted point cloud data.

[0091] Next, a description will be given of the rough alignment processing performed by the control device 2. Fig. 12 is a flow diagram showing the procedure of the rough alignment processing. Fig. 13 is an explanatory diagram showing an overview of the rough alignment processing.

[0092] The control device 2 performs a rough alignment process (ST204 in FIG. 6). In the rough alignment process, a rough alignment condition (transformation matrix) for roughly aligning the two extracted point cloud data on the source side and the target side is obtained.

[0093] The rough registration process uses the RANSAC (RANdom Sample Consensus) method. Specifically, when estimating the rough registration conditions (transformation matrix), a process of randomly selecting corresponding points on the source and target sides is repeated to select the rough registration conditions (transformation matrix) that yield the best results, thereby eliminating outliers related to the corresponding points on the source and target sides.

[0094] In the coarse alignment process, the coarse alignment conditions are determined with restrictions imposed on the search space. Specifically, as shown in Fig. 13, corresponding points on the source and target sides in the three-dimensional space of the XYZ coordinate system are projected onto the XY plane (a two-dimensional horizontal plane), and the corresponding points on the source and target sides are compared on the XY plane to determine the rotation elements (rotation components of the transformation matrix) of the coarse alignment conditions on the XY plane. This limits the rotation elements to rotation on the XY plane (around the vertical axis).

[0095] Furthermore, once the rotational elements (rotational components of the transformation matrix) of the approximate alignment conditions are determined, the translational elements (translation components of the transformation matrix) of the approximate alignment conditions are determined by taking into account the rotational elements and comparing corresponding points on the source and target sides in three-dimensional space.

[0096] As shown in FIG. 12, in the rough alignment process, the processor 46 first randomly selects a combination of a predetermined number of corresponding points (e.g., three) from the corresponding points acquired in the corresponding point acquisition process (corresponding point selection process) (ST501).

[0097] Next, processor 46 calculates the center of gravity of the corresponding points on each of the source side and the target side (center of gravity calculation process) (ST502).

[0098] Next, processor 46 translates the corresponding points on the source side and the target side so that the centers of gravity of the corresponding points on each side coincide with the origin of the coordinate axes (center of gravity matching process) (ST503).

[0099] Next, processor 46 projects the corresponding points on the source and target sides in the three-dimensional space of the XYZ coordinate system onto the XY plane (two-dimensional horizontal plane) (point cloud projection process) (ST504). Specifically, the Z-axis component of the coordinates of the corresponding points on the source and target sides is set to 0.

[0100] Next, processor 46 estimates rotation elements (rotation components of the transformation matrix) of the approximate alignment conditions (rotation element estimation process) (ST505). The rotation elements are used to correct deviations in the rotation direction (rotation direction around the vertical axis) of the corresponding points on the source side relative to the corresponding points on the target side on the XY plane. In the rotation element estimation process, the corresponding points on the source side are compared with the corresponding points on the target side, and the rotation elements of the approximate alignment conditions are found so as to shorten the distance between the corresponding points on the target side and the corresponding points on the source side.

[0101] The rotation element estimation process uses, for example, a method called Singular Value Decomposition (SVD). Specifically, a covariance matrix between a predetermined number (e.g., three) of corresponding points on the source and target sides is calculated, and singular vectors required to estimate the rotation components (rotation matrix) of the transformation matrix are calculated from the covariance matrix between the predetermined number of corresponding points.

[0102] Next, the translational elements (translation components of the transformation matrix) of the approximate alignment conditions are estimated (translation element estimation process) (ST506). The translational elements are used to correct deviations in the X, Y, and Z axes of the corresponding points on the source side relative to the corresponding points on the target side. In the translational element estimation process, the corresponding points on the source side and the corresponding points on the target side are compared, taking into account rotational components, to determine the translational elements of the approximate alignment conditions so as to shorten the distance between the corresponding points on the target side and the corresponding points on the source side.

[0103] Next, processor 46 applies the approximate alignment conditions (transformation matrix) consisting of rotation elements (rotation components of the transformation matrix) and translation elements (translation components of the transformation matrix) to the corresponding points on the source side, and aligns (coordinate transformation) the corresponding points on the source side (coordinate transformation process) (ST507).

[0104] Next, the processor 46 evaluates the currently acquired rough alignment conditions and determines whether the currently acquired rough alignment conditions have the best evaluation result among the rough alignment conditions acquired so far (alignment condition evaluation process) (ST508). At this time, the evaluation is performed using the degree to which the corresponding points on the source and target sides approach each other when the rough alignment conditions are applied to the corresponding points on the source side as the evaluation criterion.

[0105] If the currently acquired rough alignment conditions have the best evaluation result among the previously acquired rough alignment conditions (Yes in ST508), the processor 46 updates the estimation result with the currently acquired rough alignment conditions, i.e., adopts the currently acquired rough alignment conditions (ST509).On the other hand, if the currently acquired rough alignment conditions have not the best evaluation result among the previously acquired rough alignment conditions (No in ST508), the processor 46 discards the currently acquired rough alignment conditions.

[0106] Next, processor 46 determines whether a sufficiently accurate result can be obtained using the estimated rough alignment conditions (ST510). If a sufficiently accurate result can be obtained using the estimated rough alignment conditions (Yes in ST510), the process ends. At this time, if a predetermined number of corresponding points (e.g., three pairs) on the source and target sides match, it is determined that a sufficiently accurate result has been obtained.

[0107] On the other hand, if a result of sufficient accuracy is not obtained (No in ST510), processor 46 then determines whether or not a predetermined number of repetitions has been reached (ST511). If the predetermined number of repetitions has been reached (Yes in ST511), the process ends. On the other hand, if the predetermined number of repetitions has not been reached (No in ST511), the process returns to ST501, where processor 46 reselects a predetermined number of corresponding points and repeats the same process (ST501 to ST510).

[0108] In this way, in the rough alignment process, random sampling is repeated based on the RANSAC method, in which a predetermined number of corresponding points (e.g., three sets) are randomly selected, and rough alignment conditions (transformation matrices) that produce the best evaluation results, i.e., rough alignment conditions based on corresponding points that do not fall under outliers, can be obtained.

[0109] In the approximate alignment process, the approximate alignment conditions are estimated under constraints on the search space, and the approximate alignment conditions are estimated by rotation on the XY plane (around the vertical axis) and translation in the XYZ directions (coordinate transformation). This reduces the processing load and prevents the problem of combining point cloud data when the floor of the target location is tilted.

[0110] In this embodiment, in the 3D reconstruction process of the target location performed by the control device 2, the vertical direction (direction of gravity) can be recognized based on the detection results of the IMU 23 provided in the sensor unit 11 of the image capture device 1, so the vertical direction is known in the point cloud data acquired by the 3D reconstruction process, and the Z-axis direction (vertical direction) of multiple point cloud data acquired individually by the 3D reconstruction process matches. Therefore, even if the rough alignment conditions are estimated with restrictions placed on the search space, specifically, even if the rotation component of the transformation matrix is ​​limited to rotation on the XY plane (rotation around the vertical axis), appropriate alignment is performed, and therefore it is possible to generate combined point cloud data that accurately reproduces the target location by the point cloud combination process.

[0111] Next, a description will be given of the precision alignment processing performed by the control device 2. Fig. 14 is a flow diagram showing the procedure of the precision alignment processing. Fig. 15 is an explanatory diagram showing an overview of the precision alignment processing.

[0112] The control device 2 performs a fine alignment process (ST205 in FIG. 6 ). In the fine alignment process, the fine alignment conditions (transformation matrices) for precisely aligning the two point cloud data on the source and target sides are determined for the two extracted point cloud data aligned using the rough alignment conditions (transformation matrices) acquired in the rough alignment process (ST204 in FIG. 6 ).

[0113] The precise registration process uses a point-to-plane iterative closest point (ICP) technique. Specifically, as shown in FIG. 15 , the perpendicular distance between a source point and a target surface, i.e., the length of a perpendicular line drawn from the source point to the target surface, is measured, and a transformation matrix that performs a minute transformation to reduce the perpendicular distance is estimated. The process of alignment is repeated to determine the final precise registration condition (transformation matrix). Here, the source point refers to a corresponding point in the extracted point cloud data on the source side. The target surface refers to a plane that passes through the corresponding point in the extracted point cloud data on the target side and is perpendicular to the normal to the corresponding point.

[0114] Furthermore, in the precise alignment process, precise alignment conditions are determined with constraints imposed on the search space. Specifically, the optimum precise alignment conditions (transformation matrix) are determined by repeatedly setting candidates for precise alignment conditions (transformation matrices) that perform alignment (coordinate transformation) by rotation on the XY plane (around the vertical axis) and translation in the XYZ axis directions, and evaluating the results based on the inter-point distances between each point in the source point cloud data and the nearest points in the target point cloud data when the candidate precise alignment conditions are applied.

[0115] 14, in the precise alignment process, first, the processor 46 acquires corresponding points to be used in the precise alignment process (corresponding point acquisition process) (ST601). This corresponding point acquisition process differs from the corresponding point acquisition process (ST403 in FIG. 10) performed as pre-processing in that the nearest points in the target-side point cloud data to each point in the source-side point cloud data are determined as corresponding points. In other words, combinations of points with short inter-point distances are determined as corresponding points.

[0116] Next, processor 46 measures the vertical distance between each pair of corresponding points on the source and target sides (vertical distance measurement process) (ST602). At this time, the vertical distance between the source point and the target surface, i.e., the length of the perpendicular line drawn from the source point to the target surface, is measured (see FIG. 15).

[0117] Next, processor 46 estimates the fine alignment conditions (transformation matrix) while imposing constraints on the search space, and acquires the fine alignment conditions (ST603). At this time, to impose constraints on the search space, Jacobians for rotation on the XY plane and translation (parallel movement) in the XYZ axis directions are calculated. Then, with the objective of minimizing the perpendicular distance between corresponding points, the problem of minimizing the perpendicular distance is linearly approximated using the Jacobian, and the transformation matrix is ​​obtained by finding the solution.

[0118] Next, the processor 46 applies the precise alignment conditions (transformation matrix) to each point of the extracted point cloud data on the source side to align (transform coordinates) the extracted point cloud data on the source side (coordinate transformation process) (ST604).

[0119] Next, the processor 46 acquires an evaluation value for the alignment using the newly acquired precise alignment conditions (transformation matrix) (ST605). At this time, the nearest point in the extracted point cloud data on the target side to each point in the aligned extracted point cloud data on the source side is obtained. The inter-point distance between each point on the source side and the nearest point on the target side is calculated, and the inter-point distance is compared with a predetermined threshold value to obtain an evaluation value. The nearest point on the target side, whose inter-point distance is less than the threshold value, and its corresponding point on the source side, become the corresponding points to be used in the next processing.

[0120] Next, processor 46 determines whether the difference between the current evaluation value and the previous evaluation value is less than a predetermined threshold value (ST606).

[0121] Here, if the difference between the current evaluation value and the previous evaluation value is less than the threshold value (Yes in ST606), the processor 46 then determines the final fine alignment conditions that reflect all of the transformation matrices that have been repeatedly acquired up to that point, and applies the fine alignment conditions to the point cloud data on the source side to align (coordinate transformation) the point cloud data on the source side (coordinate transformation process) (ST607).

[0122] On the other hand, if the difference between the current evaluation value and the previous evaluation value is equal to or greater than the threshold value (No in ST606), processor 46 then determines whether or not a predetermined number of repetitions has been reached (ST608). If the predetermined number of repetitions has been reached (Yes in ST608), the process proceeds to ST607.

[0123] On the other hand, if the predetermined number of repetitions has not been reached (No in ST608), the process returns to ST602, where processor 46 reselects corresponding points and repeats the same processes (ST501 to ST510). At this time, the closest point on the target side, whose inter-point distance obtained in the process of ST605 is less than the threshold, and its corresponding point on the source side are used as corresponding points in this process.

[0124] In this way, the precise alignment process uses the Point-to-Plane ICP method to find precise alignment conditions (transformation matrices) that reduce vertical distances. Furthermore, the precise alignment process estimates the precise alignment conditions while restricting the search space, and estimates the precise alignment conditions for alignment (coordinate transformation) by rotation on the XY plane (around the vertical axis) and translation in the XYZ directions. This reduces the processing load and prevents the problem of combining point cloud data when the floor of the target location is tilted.

[0125] 15, the point-to-plane ICP method is used, but the point-to-point ICP method may also be used. In the point-to-plane ICP method, an alignment condition (transformation matrix) is obtained that reduces the perpendicular distance between a point on the source side and a surface on the target side, while in the point-to-point ICP method, an alignment condition is obtained that reduces the inter-point distance between a point on the source side and a point on the target side.

[0126] As described above, the embodiments have been described as examples of the technology disclosed in this application. However, the technology in this disclosure is not limited to these, and can be applied to embodiments in which modifications, substitutions, additions, omissions, etc. are made. Furthermore, it is also possible to combine the components described in the above embodiments to create new embodiments.

[0127] The point cloud combining device and point cloud combining method disclosed herein have the advantage of being able to easily and accurately align point cloud data without the need for tedious preparations such as placing markers at the target location in advance and measuring their positions in order to align multiple point cloud data obtained by 3D reconstruction processing of the target location, and are useful as a point cloud combining device and point cloud combining method that aligns two point cloud data obtained by 3D reconstruction processing of the target location with each other and then combines them.

[0128] 1: Image capture device 2: Control device (point cloud connection device) 46: Processor P1: 3D reconstruction processing P2: Point cloud alignment processing P3: Point cloud connection processing

Claims

1. A point cloud merging device that uses a processor to execute a process of merging two point cloud data acquired by a three-dimensional reconstruction process of a target location after mutual alignment, wherein the processor: executes a rough alignment process for two extracted point cloud data obtained by extracting a common area from each of the two point cloud data, to obtain rough alignment conditions while imposing constraints on the search space; executes a fine alignment process for the two extracted point cloud data that have been aligned using the rough alignment conditions, to obtain fine alignment conditions while imposing constraints on the search space; and executes a point cloud merging process to merge the two point cloud data that have been aligned using the fine alignment conditions.

2. The point cloud combining device described in claim 1, characterized in that the processor performs the following preprocessing for the rough alignment process: estimating normals for each point included in each of the two extracted point cloud data, or re-estimating normals if normal information is already available; calculating feature values ​​for each point included in each of the two extracted point cloud data based on the normals of the points; and comparing feature values ​​for each point included in each of the two extracted point cloud data to obtain corresponding points that associate points with similar features.

3. The point cloud combining device according to claim 1, characterized in that the processor acquires the rough alignment conditions using a RANSAC method in the rough alignment process.

4. The point cloud combining device according to claim 1, characterized in that the processor, in the rough alignment process, obtains the rotation element of the rough alignment condition while limiting the rotation on a horizontal plane as a constraint on the search space.

5. The point cloud combining device according to claim 1, characterized in that the processor acquires the precise alignment conditions using a Point-to-Plane ICP method in the precise alignment process.

6. The point cloud combining device according to claim 1, characterized in that the processor, in the fine alignment process, obtains the rotation element of the fine alignment condition while limiting the rotation on a horizontal plane as a constraint on the search space.

7. A point cloud merging method that causes a processor to perform a process of merging two point cloud data acquired by a three-dimensional reconstruction process of a target location after mutual alignment, characterized in that the method comprises the steps of: performing a rough alignment process on two extracted point cloud data obtained by extracting a common area from each of the two point cloud data, to obtain rough alignment conditions while restricting the search space; performing a fine alignment process on the two extracted point cloud data that have been aligned using the rough alignment conditions, to obtain fine alignment conditions while restricting the search space; and performing a point cloud merging process that merges the two point cloud data that have been aligned using the fine alignment conditions.

Citation Information

Patent Citations

  • Urban point cloud coarse-to-fine automatic registration method based on multi-source dimension decomposition

    CN111899291A

  • Calculation device, method and program

    JP2015156136A

  • Image processing device, image processing program, and image processing method

    JP2018088209A