Method for matching two point clouds
By employing (N+M)-dimensional point clouds with additional features through a neural network, the method improves alignment and transformation estimation in point cloud comparison, addressing inefficiencies in existing methods.
Patent Information
- Application Number
- DE102023213268
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-26
AI Technical Summary
Existing methods for comparing partially overlapping point clouds are inefficient and lack the ability to accurately align and estimate transformations due to limited dimensionality and lack of additional features.
Utilizing (N+M)-dimensional point clouds with additional features, processed by a neural network, particularly a convolutional neural network, to extract feature descriptors and align point clouds, incorporating features like semantic annotations and odometry information.
Enhances alignment accuracy and efficiency by increasing the dimensionality of convolutional kernels, allowing for faster training and more accurate transformation estimation between point clouds.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The invention relates to a method for comparing two at least partially overlapping point clouds, a device, a computer program and a machine-readable storage medium. State of the art
[0002] The publication DE 11 2019 000 049 T5 of the international application with the publication number WO 2019 / 161300 discloses an object detection and detection reliability suitable for autonomous driving.
[0003] The published patent application US 2022 / 0180548 A1 discloses a method for estimating an object pose.
[0004] The published patent application WO 2022 / 138123 A1 discloses a device for identifying a free parking space. Disclosure of the invention
[0005] The object underlying the invention is to provide a concept for the efficient comparison of two at least partially overlapping point clouds.
[0006] This object is achieved by means of the respective subject matter of the independent claims. Advantageous embodiments of the invention are the subject matter of the respective dependent subclaims.
[0007] According to a first aspect, a method for matching two at least partially overlapping point clouds is provided, comprising the following steps: Receiving a first N-dimensional point cloud with M additional point-by-point features, wherein the first point cloud represents a first road section, Receiving a second N-dimensional point cloud with M additional point-by-point features, wherein the second point cloud represents a second road section at least partially overlapping the first road section, where M is greater than 1, Extracting first feature descriptors from the first point cloud and second feature descriptors from the second point cloud using a neural network, wherein the first point cloud and the second point cloud are applied to an input layer of the neural network and wherein the first and second feature descriptors are output to an output layer of the neural network, Matching the first point cloud with the second point cloud based on the first feature descriptors and on the second feature descriptors.
[0008] According to a second aspect, a device is provided which is configured to carry out all the steps of the method according to the first aspect.
[0009] According to a third aspect, a computer program is provided, comprising instructions which, when the computer program is executed by a computer, for example by the device according to the second aspect, cause the computer to carry out a method according to the first aspect.
[0010] According to a fourth aspect, a machine-readable storage medium is provided on which the computer program according to the third aspect is stored.
[0011] The invention is based on and incorporates the finding that the above-mentioned problem is solved by using (N+M)-dimensional point clouds as input data for the neural network to compare them. This means that n-dimensional point clouds with M additional, point-by-point features are applied to an input layer of the neural network.
[0012] Previously, one-, two- or three-dimensional point clouds were compared with each other, whereby according to the concept described here, in addition to the point coordinates, further features are included in the point clouds, which accordingly increase the dimensionality of the points.
[0013] This results in the technical advantage, for example, that the dimensionality of the convolutional kernels also increases in the depth dimension. The additional features used, i.e., the M additional features, are included in the calculation of the feature descriptors.
[0014] Based on the feature descriptors determined in this way, the two point clouds can be efficiently aligned. In particular, the alignment process can be efficiently improved compared to the current state of the art, which uses only point clouds with a maximum of three dimensions. This makes it particularly easy to establish a correspondence assumption between the two point clouds during the alignment process, enabling a more accurate estimation of the alignment translation between the two point clouds, which can ultimately lead to a high-quality map.
[0015] Furthermore, using (N+M)-dimensional point clouds results in faster training and convergence due to the use of discriminatory features: the M additional features to the data points of the point cloud.
[0016] Thus, according to the concept described here, significantly more information, the M additional features, is available compared to the known state of the art, which only uses N-dimensional point clouds without additional point-by-point features, so that this information can be used efficiently for the comparison.
[0017] In one embodiment of the method, it is provided that the neural network is a convolutional neural network, in particular a fully convolutional neural network.
[0018] This results in the technical advantage, for example, of using a neural network that is particularly suitable for comparing two point clouds.
[0019] In one embodiment of the method, it is provided that a dimensionality of convolutional kernels of the convolutional neural network is set to M.
[0020] This results in the technical advantage, for example, that the M additional features are actually included in the calculation of the feature descriptors.
[0021] In one embodiment of the method, it is provided that the features are selected from the following group of features: semantic annotation of a road marking, semantic annotation of a traffic sign, quality assessment of one or more points of the point cloud, odometry information that contributed to the creation of a point coordinate of the point cloud.
[0022] This provides the technical advantage, for example, that particularly suitable features can be provided for comparison.
[0023] In one embodiment of the method, it is provided that a correspondence assumption is made between the first feature descriptors and the second feature descriptors, based on which the first and second point clouds are compared.
[0024] This provides the technical advantage, for example, that the two point clouds can be compared efficiently.
[0025] In one embodiment of the method, it is provided that, based on the correspondence assumption, a transformation is determined which adjusts the first and the second point cloud, on the basis of which the first and the second point cloud are adjusted.
[0026] This provides a technical advantage, for example, that the two point clouds can be compared efficiently.
[0027] In one embodiment of the method, it is provided that determining the transformation comprises estimating the transformation, in particular using a RANSAC algorithm.
[0028] This provides a technical advantage, for example, that the transformation can be determined efficiently.
[0029] The abbreviation “RANSAC” stands for “Random Sample Consensus”, which can be translated into English as “agreement with a random sample”.
[0030] In one embodiment of the method, it is provided that the correspondence assumption is established using a convolutional neural network, in particular a deep global registration network.
[0031] This, for example, provides the technical advantage that the correspondence presumption can be established efficiently.
[0032] Process characteristics arise analogously from corresponding device characteristics. This means that technical functionalities and features of the process arise from corresponding technical functionalities of the device, and vice versa.
[0033] The method is carried out, for example, by means of the device.
[0034] For example, the device is programmed to execute the computer program.
[0035] The device is, for example, a computer.
[0036] For example, the method is a computer-implemented method.
[0037] The embodiments and exemplary embodiments described here can be combined with one another in any way, even if this is not explicitly described.
[0038] For example, matching between the two point clouds involves determining a relative delta pose between the two point clouds.
[0039] For example, determining a relative delta pose includes determining a relative rotation and a relative translation.
[0040] A point cloud within the meaning of the description was or is generated, for example, using environmental sensor data from one or more environmental sensors of a motor vehicle.
[0041] An environmental sensor within the meaning of the description is, for example, one of the following environmental sensors: radar sensor, lidar, image sensor, in particular image sensor of a video camera, ultrasonic sensor, infrared sensor and magnetic field sensor.
[0042] Thus, a point cloud within the meaning of the description can be generated or created, for example, using lidar data and / or radar data and / or video camera data and / or infrared data and / or magnetic field data and / or ultrasound data.
[0043] Environment sensor data as defined in the description describes the environment of the motor vehicle.
[0044] For example, “N” in “N-dimensional point cloud” can take one of the following values: 1, 2, or 3.
[0045] “M” in “M additional pointwise features” is > 1.
[0046] The invention is explained in more detail below using preferred embodiments. In the following: Fig. 1 a flowchart of a method according to the first aspect, Fig. 2 a device according to the second aspect, Fig. 3 a machine-readable storage medium according to the fourth aspect, Fig. 4 a visualization of a convolutional model extended by additional features, and Fig. 5 is a block diagram that exemplifies the concept described here.
[0047] Fig. 1 shows a flowchart of a method for matching two at least partially overlapping point clouds, comprising the following steps: Receiving 101 a first N-dimensional point cloud 101 with M additional point-by-point features, wherein the first point cloud represents a first road section, Receiving 103 a second N-dimensional point cloud 103 with M additional point-by-point features, wherein the second point cloud represents a second road section at least partially overlapping the first road section, where M is greater than 1, Extracting 105 first feature descriptors from the first point cloud and second feature descriptors from the second point cloud using a neural network, wherein the first point cloud and the second point cloud are applied 107 to an input layer of the neural network and wherein the first and second feature descriptors are output 109 to an output layer of the neural network, Matching 111 the first point cloud with the second point cloud based on the first feature descriptors and on the second feature descriptors.
[0048] Fig. 2 shows a device 201 which is configured to carry out all steps of the method according to the first aspect.
[0049] For example, device 201 includes an input configured to receive the two point clouds. Device 201 includes, for example, a processor device comprising one or more processors configured to perform the extraction and matching steps.
[0050] For example, the processor device is configured to determine a delta pose between the two point clouds, i.e., in particular, a relative translation and a relative rotation. For example, device 201 comprises an output configured to output the determined delta pose, in particular, the relative rotation and the relative translation.
[0051] Fig. 3 shows a machine-readable storage medium 301 on which a computer program 303 is stored. The computer program 303 includes instructions that, when executed by a computer, cause the computer program 303 to execute a method according to the first aspect.
[0052] Fig. 4 shows an example visualization of a sparse convolution extended by additional features.
[0053] An N-dimensional point cloud with M additional point-by-point features is provided as input data for an input layer of a convolutional neural network. This point cloud is visualized, for example, by M two-dimensional pixel grids, these pixel grids being identified by a curly bracket with the reference symbol 401. The execution of the sparse convolution is symbolically represented by a corresponding plurality of convolutional kernels, identified by a curly bracket with the reference symbol 403. At an output layer of the convolutional neural network, M feature descriptors are output, which are identified by a curly bracket with the reference symbol 405.
[0054] It is intended that M > 1.
[0055] Fig. 5 shows a block diagram 501 which exemplifies the concept described here.
[0056] Input data 503 is provided for an input layer of a neural network 505. The input data 503 comprises a first N-dimensional point cloud 507 with M additional point-by-point features. The input data 503 comprises a second N-dimensional point cloud 509 with M additional point-by-point features.
[0057] The first point cloud 507 represents a first road section 511. The second point cloud 509 represents a second road section 513.
[0058] The second road section 513 at least partially overlaps the first road section 511. In this respect, the two point clouds 507, 509 at least partially overlap.
[0059] At an output layer of the neural network 505, output data 515 is output, which comprises first feature descriptors 517 and second feature descriptors 519.
[0060] Specifically, the first feature descriptors 517 were extracted from the first point sequence 507. The second feature descriptors were extracted from the second point cloud 509.
[0061] For example, the feature descriptors 517, 519 are each an abstracted representation of points and lines of a digital map.
[0062] Subsequently, based on the feature descriptors 517, 519, a correspondence estimation 521 is performed, which results in an estimated correspondence between the two road sections 511, 513. Based on the correspondence estimation 521, a direct transformation estimation can be performed, for example, according to a function block 523, in particular using the RANSAC algorithm. A relative rotation and a relative translation are determined as the result 525 of the direct transformation estimation.
[0063] Alternatively or additionally, transformation estimation may be performed using a Deep Global Registration Network 529.
[0064] In detail, according to a function block 511, a correspondence evaluation and a comparison between the two road sections 511, 513 is carried out.
[0065] Using the Deep Global Registration Network 529, a correspondence classifier is determined according to a function block 531. Based on this correspondence classifier, a transformation estimation based on a regularization of the weighted Procrustes problem is performed according to a function block 523. As a result of this transformation estimation, initial values for the relative rotation and for the relative translation between the two road sections 511, 513 are determined. The initial value for the relative translation is symbolically represented by an arrow with the reference numeral 535. The initial value for the relative rotation is symbolically represented by a curved arrow with the reference numeral 537.
[0066] Based on these two initial values, a more precise estimation or optimization of the relative translation or relative rotation can be performed. This is done according to a function block 539, so that a relative translation and a relative rotation between the two road sections 511, 513 are determined as a result 441.
[0067] In summary, the concept described here is based on using (N+M)-dimensional point clouds as input data for the neural network to match the two point clouds. This allows the neural network to compute richer feature descriptors. These feature descriptors now contain more information. This leads to more accurate and robust correspondence estimation.
[0068] Thus, the point clouds used as input data for the neural network are supplemented with additional data features. These additional features can be determined, for example, using additional sensors on the vehicle and / or using additional perception methods, such as semantic annotations of the lane markings, quality assessments of the data points, or, for example, orometry information that contributed to the creation of the point coordinates.
[0069] The current state of the art typically used N-dimensional coordinates of point clouds generated by various sensors, such as resistance sensors or envelope sensors from a vehicle's cameras. This creates, in particular, an N-dimensional grid with dimensions X x Z x M, where M = 1. This means that each subsequent point is naively entered with a fixed value, in this case 1.
[0070] According to the concept described here, additional features corresponding to the respective points of the point cloud are also taken into account. This creates, in particular, an N-dimensional grid with M > 1, in particular M >> 1.
[0071] This advantageously increases the dimensionality of the convolutional kernel in the depth dimension to M. This means that the additional features actually enter into the calculation of the feature descriptors.
[0072] As already explained above, the additional features can be determined using additional sensor technology of the motor vehicle and / or other perception methods, for example, semantic annotations of the lane markings, quality assessments of the data points, or, for example, odometry information that contributed to the creation of the point coordinates. Aggregated combinations of such features can also be used or provided. QUOTES CONTAINED IN THE DESCRIPTION
[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited patent literature
[0000] DE 11 2019 000 049 T5
[0002] WO 2019 / 161300
[0002] US 2022 / 0180548 A1
[0003] WO 2022 / 138123 A1
[0004]
Claims
[1] Method for matching two at least partially overlapping point clouds, comprising the following steps: Receiving a first N-dimensional point cloud (101) with M additional point-by-point features, wherein the first point cloud (507) represents a first road section (511), Receiving a second N-dimensional point cloud (103) with M additional point-by-point features, wherein the second point cloud (509) represents a second road section at least partially overlapping the first road section (513), where M is greater than 1, Extracting first feature descriptors from the first point cloud (507) and second feature descriptors from the second point cloud (509) using a neural network, wherein the first point cloud (507) and the second point cloud (509) are applied to an input layer of the neural network (505) and wherein the first and second feature descriptors are output at an output layer of the neural network (505), Matching the first point cloud (507) with the second point cloud (509) based on the first feature descriptors and on the second feature descriptors. [2] Method according to claim 1, wherein the neural network (50) is a convolutional neural network, in particular a fully convolutional neural network. [3] The method of claim 2, wherein a dimensionality of convolutional kernels of the convolutional neural network is set to M. [4] Method according to one of the preceding claims, wherein the features are selected from the following group of features: semantic annotation of a road marking, semantic annotation of a traffic sign, quality assessment of one or more points of the point cloud, odometry information that contributed to the creation of a point coordinate of the point cloud. [5] Method according to one of the preceding claims, wherein a correspondence assumption is made between the first feature descriptors and the second feature descriptors, based on which the first and second point clouds (507, 509) are compared. [6] Method according to claim 5, wherein based on the correspondence assumption a transformation is determined which adjusts the first and the second point cloud (507, 509), based on which transformation the first and the second point cloud (507, 509) are adjusted. [7] The method of claim 6, wherein determining the transformation comprises estimating the transformation, in particular using a RANSAC algorithm. [8] Method according to one of claims 5 to 7, wherein the correspondence assumption is made using a convolutional neural network, in particular a deep global registration network (529). [9] Device (201) which is arranged to carry out all the steps of the method according to one of the preceding claims. [10] Computer program (303) comprising instructions which, when the computer program (303) is executed by a computer, cause the computer to carry out a method according to one of claims 1 to 8. [11] Machine-readable storage medium (301) on which the computer program (303) according to claim 10 is stored.
Citation Information
Patent Citations
OBJECT DETECTION AND DETECTION SECURITY SUITABLE FOR AUTONOMOUS DRIVING
DE112019000049T5
Method and apparatus with object pose estimation
US20220180548A1
Detecting objects and determining confidence scores
WO2019161300A1
Available parking space identification device, available parking space identification method, and program
WO2022138123A1