A multi-modal point cloud segmentation method, system, device and storage medium
By introducing multimodal fusion technology into the point cloud segmentation algorithm, multiple characterization features and image features of point clouds are dynamically fused, the problem of poor lighting changes and calibration errors in the existing technology is solved, segmentation accuracy and performance are improved, and unified multi-task processing is achieved.
Patent Information
- Application Number
- CN202310174842.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-27
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-02-27
AI Technical Summary
The existing point cloud segmentation algorithm is not robust enough in terms of lighting changes and calibration errors, and fails to fully explore the fusion between different point cloud representations and the fusion of image and point cloud representations.
The multimodal point cloud segmentation method is adopted to convert the point cloud into depth map and voxel features through spheric projection, and the image features are learned cross-modal correlations with depth map features and voxel features, dynamically fuse multiple characterization features of point clouds, and finally segmentation is performed through the task head.
The performance and accuracy of point cloud segmentation are improved, the tolerance for calibration matrix errors is enhanced, and the 3D geometric information of point clouds and the color and texture information of the image are fully utilized, so as to realize the unified processing of semantic segmentation, panoramic segmentation and 4D panoramic segmentation.
Smart Images

Figure CN116310326B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of point cloud segmentation, and in particular, to a multi-modal point cloud segmentation method, system, device, and storage medium. Background Art
[0002] Image segmentation tasks have been widely explored in various fields such as medical treatment, defect detection, and lane line detection, and very good results have been achieved. Images can provide rich information such as colors and textures. For applications like autonomous driving that require very high-precision perception results to assist downstream tasks in making better decisions, segmentation algorithms relying solely on images cannot solve the influence of factors such as lighting and cannot provide accurate 3D geometric information. In recent years, to obtain more robust segmentation effects, lidar has been widely applied to various scene understanding tasks such as autonomous driving and service robots. The goal of the point cloud segmentation task is to assign corresponding semantic labels to the input point cloud point by point.
[0003] Currently, the mainstream algorithms for point cloud segmentation use different representations of the point cloud as input signals. For example, in existing very mature 2D semantic segmentation networks, the point cloud is converted into a depth map through sphere projection and then fed into the 2D network, and finally a point-by-point predicted semantic segmentation result is obtained; or the input point cloud can be divided into cylindrical partitions, and the cylindrical features are input into a 3D network to obtain the point cloud semantic segmentation result. However, the fusion between different representations of the point cloud and the fusion between different representations of the point cloud and images have not been comprehensively explored. Currently, multi-modal fusion methods only stay between images and depth maps. At the same time, the way of matching the pixels of the image and the points in the point cloud depends on a fixed calibration matrix, which makes these fusion algorithms not robust enough to calibration errors. Summary of the Invention
[0004] The embodiments of the present application provide a multi-modal point cloud segmentation method, system, device, and storage medium, which improve the performance and segmentation accuracy of point cloud segmentation.
[0005] To solve the above technical problems, in a first aspect, an embodiment of the present application provides a multi-modal point cloud segmentation method, including the following steps: First, based on the point cloud and the multi-layer perceptron, point features are obtained; the point cloud is converted to a depth map through spherical projection, and the depth map is input into the encoder of the depth map to extract depth map features; the features after voxelizing the point cloud are input into the voxel encoder to obtain voxel features; the image is input into a network based on ResNet to extract image features; Next, the image features are respectively fused with the depth map features and the voxel features to obtain the finally fused depth map features and the finally fused voxel features; the finally fused depth map features and the finally fused voxel features are subjected to point feature conversion to obtain the converted depth map features and the converted voxel features; the point features, the converted depth map features, and the converted voxel features are concatenated to obtain multi-view features; the multi-view features are weighted by learnable weights to learn global and cross-view information to obtain the finally fused features; Finally, the finally fused features are fed into the task head for segmentation processing.
[0006] In some exemplary embodiments, fusing the image features with the voxel features includes: based on the internal parameter matrix and the external parameter matrix of the camera, obtaining the correspondence between the voxel features and the pixel points; based on the correspondence, calculating the reference image features; sampling multiple image features, and based on the reference image features, obtaining the offsets of the multiple image features; fusing the voxel features with the sampled image features to obtain the voxel features enhanced by the image; fusing the voxel features enhanced by the image with the voxel features to obtain the finally fused voxel features; the correspondence is shown in formula (1):
[0007] (1)
[0008] where T is the external parameter matrix of the camera, and the external parameter matrix of the camera includes a rotation matrix and a translation matrix; S is the internal parameter matrix of the camera; (x i , y i , z i )represents the voxel center, representing the position of the voxel; (u i , v i )represents the pixel point corresponding to the voxel center.
[0009] In some exemplary embodiments, the voxel features are fused with the sampled image features through formula (2):
[0010] (2)
[0011] where is the sampled image feature, is the voxel feature enhanced by the image, and are learnable weights, m indexes the attention head, M is the number of attention heads, and L is the number of sampled image features; and respectively represent the sampling offset and attention weight of the l-th sampled image feature in the m-th attention head; i is the sampled point, and p i is the coordinate of the sampling point, is the reference image feature.
[0012] In some exemplary embodiments, fusing the image features with the depth map features includes: based on the reference image features, fusing the depth map features with a plurality of sampled image features to obtain target image features; fusing the depth map features with the target image features to obtain the finally fused depth map features.
[0013] In some exemplary embodiments, the transformed depth map features and the transformed voxel features are obtained by trilinear interpolation and linear interpolation projection onto the point space respectively.
[0014] In some exemplary embodiments, after the multi-view features are weighted by learnable weights to learn global and cross-view information to obtain the finally fused features, it includes: using learnable parameters to learn the weights among three features, which are the depth map features, the voxel features, and the point features respectively; after the multi-view features are weighted by learnable weights, inputting them into a multi-layer perceptron layer to learn global information from each view to obtain globally enhanced features; inputting the globally enhanced features into a multi-layer perceptron to learn the interaction between different views to obtain enhanced features; splicing the enhanced features with the depth map features and the voxel features in a residual manner respectively to obtain the finally fused features.
[0015] In some exemplary embodiments, the learning process is represented by formula (3):
[0016] (3)
[0017] Wherein, is the enhanced feature, is the learnable parameter, 、 、 are the voxel features, the depth map features, and the point features that have been transformed into the point cloud space respectively; represents global; represents the multi-layer perceptron for extracting global point features; represents the multi-layer perceptron for learning the interaction between different views.
[0018] In a second aspect, the embodiments of the present application further provide a multi-modal point cloud segmentation system, including: a feature extraction module, a learnable cross-modal association module, a learnable cross-view association module, and a task head module connected in sequence; the feature extraction module is used to obtain point features according to the point cloud and a multi-layer perceptron; the point cloud is converted into a depth map through spherical projection, and the depth map is input into an encoder of the depth map to extract depth map features; the features after voxelizing the point cloud are input into a voxel encoder to obtain voxel features; the image is input into a network based on ResNet to extract image features; the learnable cross-modal association module is used to fuse the image features with the depth map features and the voxel features respectively to obtain the finally fused depth map features and the finally fused voxel features; the learnable cross-view association module is used to perform point feature conversion on the finally fused depth map features and the finally fused voxel features to obtain the converted depth map features and the converted voxel features; the point features, the converted depth map features, and the converted voxel features are concatenated to obtain multi-view features; after the multi-view features are weighted by learnable weights, global and cross-view information is learned to obtain finally fused features; the task head module is used to perform segmentation processing according to the finally fused features output by the learnable cross-view association module.
[0019] In addition, the present application further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above multi-modal point cloud segmentation method.
[0020] In addition, the present application further provides a computer-readable storage medium storing a computer program, and the computer program implements the above multi-modal point cloud segmentation method when executed by a processor.
[0021] The technical solutions provided by the embodiments of the present application have at least the following advantages:
[0022] An embodiment of the present application provides a multi-modal point cloud segmentation method, system, device, and storage medium. The method includes the following steps: First, based on the point cloud and a multi-layer perceptron, point features are obtained; the point cloud is converted into a depth map through spherical projection, and the depth map is input into an encoder of the depth map to extract depth map features; the features after voxelizing the point cloud are input into a voxel encoder to obtain voxel features; the image is input into a network based on ResNet to extract image features; Next, the image features are respectively fused with the depth map features and the voxel features to obtain the finally fused depth map features and the finally fused voxel features; the finally fused depth map features and the finally fused voxel features are subjected to point feature conversion to obtain the converted depth map features and the converted voxel features; the point features, the converted depth map features, and the converted voxel features are concatenated to obtain multi-view features; Then, after the multi-view features are weighted by learnable weights, global and cross-view information is learned to obtain the finally fused features; the finally fused features are fed into a task head for segmentation processing.
[0023] The multi-modal point cloud segmentation method and system provided by the present application use four signals, namely point cloud, voxel, depth map, and image, as inputs. Through a learnable cross-modal association module (LMA), the network dynamically fuses voxel features and image features, as well as depth map features and image features respectively. And through a learnable cross-view association module (LVA), the network dynamically fuses the features of points, voxels, and depths in the point cloud Figure 3 features of the representations. The present application makes full use of the 3D geometric information of different representations of the point cloud and the color, texture, etc. information of the image, and shows sufficient robustness to calibration errors, improving the performance and accuracy of segmentation. In addition, the present application can implement three tasks, namely semantic segmentation, panoramic segmentation, and 4D panoramic segmentation, in the same network. In addition, the multi-modal point cloud segmentation method and system provided by the present application improve the tolerance to calibration matrix errors, effectively fuse the features rich in semantic information of the image, and further improve the segmentation performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not constitute a limitation on the embodiments. Unless otherwise stated, the figures in the drawings do not constitute a scale limitation.
[0025] Figure 1 It is a schematic flow chart of a multi-modal point cloud segmentation method provided by an embodiment of the present application;
[0026] Figure 2 It is a schematic structural diagram of a multi-modal point cloud segmentation system provided by an embodiment of the present application;
[0027] Figure 3Schematic diagram of the overall framework of a multi-modal point cloud segmentation system provided by an embodiment of the present application;
[0028] Figure 4 Schematic diagram of the structure of a learnable cross-modal association module provided by an embodiment of the present application;
[0029] Figure 5 Schematic diagram of the structure of a learnable cross-view association module provided by an embodiment of the present application;
[0030] Figure 6 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0031] As can be seen from the background art, the existing multi-modal fusion methods currently only stay between images and depth maps. The way of matching pixels in images with points in point clouds depends on a fixed calibration matrix, and there is a problem that the fusion algorithm is not robust enough to calibration errors.
[0032] Point cloud semantic segmentation is an indispensable part in tasks such as autonomous driving, digital cities, and service robots. Point clouds and images are used as two common modal inputs in the point cloud segmentation task. Point clouds can be processed into different representations, such as the following three representations: voxels, depth maps, and points. Different modalities, and different point cloud representations have their own advantages and disadvantages, and they complement each other pairwise. However, the fusion between different representations of point clouds, as well as the fusion between different representations of point clouds and images, has not been comprehensively explored. Currently, multi-modal fusion methods only stay between images and depth maps, and the fusion methods between other point cloud representations and images are ignored by everyone. At the same time, the way of matching pixels in images with points in point clouds depends on a fixed calibration matrix, which makes these fusion algorithms not robust enough to calibration errors.
[0033] Most current mainstream point cloud segmentation algorithms only use different representations of point clouds as input signals. For example, in existing very mature 2D semantic segmentation networks, the point cloud is converted into a depth map through sphere projection and then fed into the 2D network, and finally a per-point predicted semantic segmentation result is obtained. There is also a related technology that divides the input point cloud into cylindrical partitions using the sparsity characteristics of outdoor point clouds, and inputs the cylindrical features into a 3D network to obtain the point cloud semantic segmentation result. In addition, some research has fused the features of the voxel and point representations of the point cloud, adding more fine-grained features of points on the basis of voxel features. Obviously, single-modal has its own disadvantages and does not utilize the rich color, texture and other information provided by images, so multi-modal fusion algorithms have gradually attracted everyone's attention. There is also a research that projects the point cloud into a perspective view, and the 2D perspective view and the image are propagated forward through two 2D networks. The two different modal features obtained are fused in a residual manner to obtain multi-modal features. In the field of point cloud segmentation, many algorithms are not open source and difficult to reproduce, and there is currently no relatively large and comprehensive code library.
[0034] Existing technologies often only explore single-representation algorithms or feature fusion algorithms between two representations such as voxels and points, and do not fully utilize the advantages of each representation to improve the segmentation performance. To solve the above technical problems, an embodiment of the present application provides a multi-modal point cloud segmentation method, including the following steps: First, based on the point cloud and a multi-layer perceptron, point features are obtained; the point cloud is converted into a depth map through sphere projection, and the depth map is input into the encoder of the depth map to extract depth map features; the features after voxelizing the point cloud are input into the voxel encoder to obtain voxel features; the image is input into a network based on ResNet to extract image features; Next, the image features are respectively fused with the depth map features and the voxel features to obtain the finally fused depth map features and the finally fused voxel features; the finally fused depth map features and the finally fused voxel features are subjected to point feature conversion to obtain the converted depth map features and the converted voxel features; the point features, the converted depth map features, and the converted voxel features are concatenated to obtain multi-view features; after the multi-view features are weighted by learnable weights, global and cross-view information is learned to obtain the finally fused features; Finally, the finally fused features are fed into the task head for segmentation processing. The present application fully utilizes the respective advantages of the three basic representations of the point cloud to improve the segmentation accuracy.
[0035] The following will elaborate on each embodiment of the present application with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in each embodiment of the present application, many technical details are proposed to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions required to be protected by the present application can still be implemented.
[0036] Refer to Figure 1 , an embodiment of the present application provides a multi-modal point cloud segmentation method, including the following steps:
[0037] Step S1: Based on the point cloud and the multi-layer perceptron, obtain point features; convert the point cloud to a depth map through sphere projection, and input the depth map into the encoder of the depth map to extract depth map features; input the features after voxelizing the point cloud into the voxel encoder to obtain voxel features; input the image into a network based on ResNet to extract image features.
[0038] Step S2: Fuse the image features with the depth map features and voxel features respectively to obtain the finally fused depth map features and the finally fused voxel features.
[0039] Step S3: Perform point feature conversion on the finally fused depth map features and the finally fused voxel features to obtain the converted depth map features and the converted voxel features; splice the point features, the converted depth map features, and the converted voxel features to obtain multi-view features.
[0040] Step S4: After the multi-view features are weighted by learnable weights, learn global and cross-view information to obtain the finally fused features; finally, send the finally fused features into the task head for segmentation processing.
[0041] Specifically, the point features in Step S1 are obtained through a series of multi-layer perceptrons (MLPs). The depth map features are obtained by converting the point cloud to a depth map through sphere projection and inputting the depth map into the encoder of the depth map. The voxel features are obtained by inputting the features after voxelizing the point cloud into the voxel encoder. The image features are obtained by inputting the image into a network based on ResNet. The representation of the point cloud includes three representations: voxel, depth map, and point. Each of the three representations has its own advantages and disadvantages and complements each other. The present application fully utilizes the respective advantages of the three basic representations of the point cloud to improve the segmentation accuracy. The present application fully utilizes the 3D geometric information of different representations of the point cloud and the color, texture, etc. of the image, and shows sufficient robustness to calibration errors. The present application realizes robust point cloud segmentation by dynamically fusing images and points, voxels, and depth maps, improving the performance and segmentation accuracy of point cloud segmentation.
[0042] Refer to Figure 2, embodiments of the present application also provide a multi-modal point cloud segmentation system, including: a feature extraction module 101, a learnable cross-modal association module 102, a learnable cross-view association module 103, and a task head module 104 connected in sequence; the feature extraction module 101 is used to obtain point features according to the point cloud and a multi-layer perceptron; the point cloud is converted into a depth map through spherical projection, and the depth map is input into an encoder of the depth map to extract depth map features; the features after voxelizing the point cloud are input into a voxel encoder to obtain voxel features; the image is input into a ResNet-based network to extract image features; the learnable cross-modal association module 102 is used to fuse the image features with the depth map features and voxel features respectively to obtain the finally fused depth map features and the finally fused voxel features; the learnable cross-view association module 103 is used to perform point feature conversion on the finally fused depth map features and the finally fused voxel features to obtain the converted depth map features and the converted voxel features; the point features, the converted depth map features, and the converted voxel features are concatenated to obtain multi-view features; after the multi-view features are weighted by learnable weights, global and cross-view information is learned to obtain finally fused features; the task head module 104 is used to perform segmentation processing according to the finally fused features output by the learnable cross-view association module.
[0043] Figure 3 Schematically shows the overall framework diagram of the multi-modal point cloud segmentation system. As Figure 3 shown, the present application uses four signals, namely point cloud (Point Cloud), voxel (Voxel), depth map (Range Image), and image (Image), as inputs to implement three tasks, namely semantic segmentation, panoptic segmentation, and 4D panoptic segmentation, in the same network. First, the present application converts the point cloud into a depth map through spherical projection, and the depth map is input into an encoder of the depth map (Range Encoder) to extract depth map features; the point features are obtained through a series of multi-layer perceptrons (MLPs); the voxel features are obtained by inputting the features after voxelizing the point cloud into a voxel encoder (Voxel Encoder); the image is input into a ResNet-based network to extract image features.
[0044] It should be noted that semantic segmentation is to assign a class label to each pixel in an image (such as cars, buildings, ground, sky, etc.), and different colors are used to represent the labels in the image. Panoramic segmentation is a combination of semantic segmentation and instance segmentation. Each pixel is divided into a category. If there are multiple instances in a category, different colors are used for differentiation, so that it can be known which pixel belongs to which instance in which category. For example, if two colors belong to a certain category, but the two colors belong to different instances respectively, different instances can be easily distinguished by the colors of the labels. Point cloud semantic segmentation is an indispensable part of tasks such as autonomous driving, digital cities, and service robots. In recent years, in order to obtain more robust segmentation results, lidar has been widely applied to various scene understanding tasks such as autonomous driving and service robots. The goal of the cloud segmentation task is to assign corresponding semantic labels to the input point cloud point by point.
[0045] Please continue to refer to Figure 3 , in this application, the proposed learnable cross-modal association module (LMA) fuses the image features with the voxel features and the depth map features respectively ( Figure 3 the VI and RI in ). Through the depth map to point conversion relationship and the voxel to point conversion relationship Figure 3 , both the depth map features and the voxel features are converted into point features. The point features, the depth map features and the voxel features that have been converted into point features are dynamically fused together (
[0046] the RPV in
[0047] ). (1)
[0048] where T is the external parameter matrix of the camera, and the external parameter matrix of the camera includes a rotation matrix and a translation matrix; S is the internal parameter matrix of the camera; (x i , y i , z i ) represents the voxel center, representing the position of the voxel; (ui , v i represents the pixel corresponding to the voxel center.
[0049] In some embodiments, the voxel feature and the sampled image feature are fused by formula (2):
[0050] (2)
[0051] Wherein, is the acquired image feature, is the voxel feature enhanced by the image, and are learnable weights, m indexes the attention head, M is the number of attention heads, and L is the number of sampled image features; and respectively represent the sampling offset and attention weight of the l-th sampled image feature in the m-th attention head; i is the sampled point, p i is the coordinate of the sampling point, is the reference image feature.
[0052] As Figure 4 shown, the present application dynamically fuses the image feature and the depth map feature and fuses the image feature and the depth map feature through a learnable cross-modal association module (LMA). F V , F I are the voxel feature and the image feature respectively. For each voxel feature , we first calculate the reference image feature based on the voxel center and the internal and external parameter matrices of the camera. Then, we use the learned offset to sample L image features. This voxel feature is used as a query, and the sampled image features are used as keys and values. The voxel feature and the sampled features are fed into a Multi-Head Cross-Attention (MHCA) to obtain the voxel feature enhanced by the image, and then concatenated with the original voxel feature through a multi-layer perceptron layer (MLP) to obtain the finally fused voxel feature.
[0053] Please continue to refer to Figure 4 , since the voxel and the image are in different coordinate systems, it is very important to match the voxel and the image pixels. We first use the internal and external parameter matrices of the camera to find the connection between the voxel and the pixel. We use the center point of each voxel to represent the position of the current voxel. For each voxel center (x i , y i , z i ), the corresponding pixel point (ui , v i ) is obtained through formula (1).
[0054] This application will use voxel features Based on the reference image features Dynamically fuse L image features. For these L image features, this application estimates their offsets from the reference image features, and then we find the selected image features And fuse them with the corresponding voxel features through formula (2) to obtain the image-enhanced voxel features .
[0055] After obtaining the image-enhanced voxel features , by concatenating the image-enhanced voxel features And the original voxel features to obtain the final fused voxel features , where C f Is the number of channels of the voxel features. In this state, the voxel features can automatically find the most relevant image features to fuse. Voxel features without matching image features will be concatenated with all-zero vectors to ensure the same length as the image features.
[0056] In some embodiments, step S2 fuses the image features with the depth map features, including: based on the reference image features, fusing the depth map features with multiple acquired image features to obtain target image features; fusing the depth features with the target image features to obtain the final fused depth map features.
[0057] For the fusion between the depth map and the image, we let the depth map features fuse L image features based on the reference image features. This application uses formula (2) to find the image features we want (target image features), and then fuses the depth map and the image features to generate the final fused depth map features.
[0058] In some embodiments, the converted depth map features and the converted voxel features in step S3 are respectively obtained by trilinear interpolation and linear interpolation projected into the point space.
[0059] See Figure 5 , this application enables the network to dynamically fuse points, voxels, and depths in the point cloud through a learnable cross-view association module (LVA) Figure 3The features of the three representations. The voxel features and depth map features are projected into the point space through trilinear interpolation and linear interpolation respectively. Then, we concatenate the voxel features, depth map features, and point features that have been transformed into the point plane, and use learnable parameters to automatically learn the weights between the three features. The weighted features are passed through multi-layer perceptrons (MLPs) to learn global and cross-view information. The finally obtained features that fuse the three representations are transformed into the voxel space, depth map space, and point space.
[0060] In some embodiments, after the multi-view features are weighted by learnable weights in step S4, global and cross-view information is learned to obtain the final fused features, including: using learnable parameters to learn the weights between the three features, which are the depth map feature, voxel feature, and point feature respectively; after the multi-view features are weighted by learnable weights, they are input into the multi-layer perceptron layer to learn global information from each view to obtain globally enhanced features; the globally enhanced features are input into the multi-layer perceptron to learn the interaction between different views to obtain enhanced features; the enhanced features are concatenated with the depth map feature and voxel feature in a residual manner respectively to obtain the final fused features.
[0061] In some embodiments, the learning process is represented by formula (3):
[0062] (3)
[0063] Where, is the enhanced feature, are learnable parameters, , , are the voxel feature, depth map feature, and point feature that have been transformed into the point space respectively; represents global; represents the multi-layer perceptron for extracting global point features; represents the multi-layer perceptron for learning the interaction between different views.
[0064] In this application, through the learnable cross-modal association module (LMA), we obtain the enhanced voxel features and depth map features of the image, and then through the learnable cross-view association module (LVA), the network dynamically fuses the features of points, voxels, and depths in the point cloud Figure 3 The features of the three representations. For the fusion of the three representations of the depth map, point, and voxel, we fuse the three representations in the point space through , transformation. However, instead of directly adding or concatenating the three representations, we let the network automatically learn the weights of each representation in the input features. As Figure 5 shows, first we pass the voxel features of the depth map feature through , Converting to point space, since the number of voxel features and depth map features is much smaller than the number of points in the point cloud, directly concatenating vectors of all zeros for voxel and depth map features will result in low performance. To address this problem of magnitude mismatch, we use trilinear interpolation and linear interpolation to interpolate voxel features and depth map features with the same number as the points in the point cloud respectively. After we obtain the voxel features , depth map features , and point cloud features , we first concatenate these three features to generate multi-view features . Then, we use learnable parameters to automatically learn the weights between different views. After the multi-view features are weighted by the learnable weights, they are first input into a 2-layer multi-layer perceptron to learn global information from each view. This globally enhanced feature is then input into a series of multi-layer perceptrons to learn the interactions between different views. The entire learning process is represented by the above formula (3).
[0065] After the multi-view features are weighted by the learnable weights, global and cross-view information is learned to obtain the final fused feature. Specifically, in this application, the enhanced feature and the original feature are concatenated in a residual manner to obtain the final fused feature. Finally, we transform the final fused feature through and into voxel space and depth map space respectively to become voxel features and depth map features.
[0066] After obtaining the final fused feature, the final fused feature is fed into the task head for segmentation processing. In this application, the final fused feature is obtained through the learnable cross-modal association module (LMA) and the learnable cross-view association module (LVA). The final fused feature is fed into the classifier to obtain the prediction of point cloud semantic segmentation. The prediction of point cloud semantic segmentation is fed into the panoramic segmentation head to estimate the center and offset of the category (thing) of the moving object to generate the prediction of panoramic segmentation. The panoramic segmentation predictions of multiple frames are all transformed into the coordinate system of the corresponding first frame, and then the same objects in different frames are given the same serial number through the intersection over union (IoU) association. In this case, point cloud semantic segmentation, panoramic segmentation, and 4D panoramic segmentation are processed under the same network.
[0067] In this application, the learnable cross-modal association module (LMA) enables the network to dynamically fuse voxel features and image features, as well as depth map features and image features respectively, and the learnable cross-view association module (LVA) enables the network to dynamically fuse points, voxels, and depths in the point cloud Figure 3The characteristics of the representation. This application can complete the three tasks of point cloud semantic segmentation, panoramic segmentation, and 4D panoramic segmentation in the same network.
[0068] In the prior art research on multi-modal algorithms in the field of point cloud segmentation, the matching of point cloud features and image features strongly depends on a fixed calibration matrix. When the calibration matrix is incorrect, the accuracy of the segmentation algorithm is greatly reduced. In point cloud segmentation, designing an algorithm that reduces the dependence on the calibration matrix has not been effectively explored. This application improves the tolerance to calibration matrix errors, effectively integrates the features of rich semantic information in images, and further improves the segmentation performance.
[0069] In addition, the multi-modal point cloud segmentation method and system provided in this application constructs the largest and most comprehensive outdoor point cloud segmentation code library OpenPCSeg, covering 18 outdoor point cloud segmentation algorithms. Compared with the prior art, many point cloud segmentation algorithms are not open source or difficult to reproduce. This application open sources a reproducible code library to promote the research iteration in the field of point cloud segmentation. Since most of the current mainstream point cloud segmentation algorithms are not open source and difficult to reproduce, this application constructs the largest and most comprehensive outdoor point cloud segmentation code library OpenPCSeg. The OpenPCSeg code library supports distributed data parallel (DDP) training. It includes algorithms based on distance images, points, voxels, fusion, enhancement, and bird's-eye view (BEV).
[0070] This application proposes a novel learnable association module (LA), which includes a learnable cross-modal association module (LMA) and a learnable cross-view association module (LVA) to dynamically fuse images and points, voxels, and depth maps to achieve robust point cloud segmentation. The learnable cross-modal association module (LMA) proposed in this application explores the fusion between voxel features and image features, as well as the fusion between depth map features and image features. Compared with the prior art, this application improves the tolerance to fixed calibration matrix errors; the learnable cross-view association module (LVA) proposed in this application explores the fusion between the three basic representations of point cloud, namely points, voxels, and depth maps, compared with the existing technology; finally, relying on this technology, this application achieves the first performance in the leaderboards of point cloud semantic segmentation and panoramic segmentation on the SemanticKITTI dataset, and wins the first place in the point cloud semantic segmentation challenge on the nuScenes dataset.
[0071] There have been relevant numerical experiments to verify that this application has achieved good results in the point cloud segmentation tasks of the SemanticKITTI, nuScenes, and Waymo datasets, and has achieved the first performance in the leaderboards of point cloud semantic segmentation and panoramic segmentation of the SemanticKITTI dataset, and has won the first place in the point cloud semantic segmentation challenge of the nuScenes dataset, demonstrating the effectiveness of the method.
[0072] Reference Figure 6 , Another embodiment of this application provides an electronic device, including: at least one processor 110; and a memory 111 communicatively connected to the at least one processor; wherein, the memory 111 stores instructions executable by the at least one processor 110, and the instructions are executed by the at least one processor 110 to enable the at least one processor 110 to execute any of the above method embodiments.
[0073] Among them, the memory 111 and the processor 110 are connected in a bus manner. The bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors 110 and the memory 111 together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be one element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices on the transmission medium. The data processed by the processor 110 is transmitted over the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor 110.
[0074] The processor 110 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. And the memory 111 can be used to store the data used by the processor 110 when performing operations.
[0075] Another embodiment of this application relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the above method embodiments are implemented.
[0076] That is, those skilled in the art can understand that all or part of the steps in the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods in the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0077] With the above technical solutions, the embodiments of the present application provide a multi-modal point cloud segmentation method, system, device, and storage medium. The method includes the following steps: First, based on the point cloud and a multi-layer perceptron, point features are obtained; the point cloud is converted into a depth map through spherical projection, and the depth map is input into an encoder of the depth map to extract depth map features; the features after voxelizing the point cloud are input into a voxel encoder to obtain voxel features; the image is input into a network based on ResNet to extract image features; Next, the image features are respectively fused with the depth map features and the voxel features to obtain the finally fused depth map features and the finally fused voxel features; the finally fused depth map features and the finally fused voxel features are subjected to point feature conversion to obtain the converted depth map features and the converted voxel features; the point features, the converted depth map features, and the converted voxel features are concatenated to obtain multi-view features; Then, after the multi-view features are weighted by learnable weights to learn global and cross-view information, the finally fused features are obtained; the finally fused features are fed into a task head for segmentation processing.
[0078] The multi-modal point cloud segmentation method and system provided by the present application use four signals, namely point cloud, voxel, depth map, and image, as inputs. Through a learnable cross-modal association module (LMA), the network dynamically fuses voxel features and image features, and depth map features and image features respectively. And through a learnable cross-view association module (LVA), the network dynamically fuses the features of points, voxels, and depths in the point cloud. Figure 3 features of several representations. The present application makes full use of the 3D geometric information of different representations of the point cloud and the color, texture, etc. information of the image, and shows sufficient robustness to calibration errors, improving the performance and accuracy of segmentation. In addition, the present application can implement three tasks, namely semantic segmentation, panoramic segmentation, and 4D panoramic segmentation, in the same network. Moreover, the multi-modal point cloud segmentation method and system provided by the present application improve the tolerance to calibration matrix errors, effectively fuse the features rich in semantic information of the image, and further improve the segmentation performance.
[0079] Those of ordinary skill in the art can understand that the above-described embodiments are specific examples for implementing the present application. In actual applications, various changes can be made in form and details without departing from the spirit and scope of the present application. Any person skilled in the art can make their own changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be subject to the scope defined by the claims.
Claims
1. A multimodal point cloud segmentation method, characterized in that, Including: Based on point cloud and multi-layer perceptron, point features are obtained; The point cloud is converted to a depth map through sphere projection, and the depth map is input into the encoder of the depth map to extract depth map features; the features after voxelizing the point cloud are input into the voxel encoder to obtain voxel features; the image is input into a network based on ResNet to extract image features; The image features are respectively fused with the depth map features and the voxel features to obtain the finally fused depth map features and the finally fused voxel features; The finally fused depth map features and the finally fused voxel features are subjected to point feature conversion to obtain the converted depth map features and the converted voxel features; the point features, the converted depth map features and the converted voxel features are concatenated to obtain multi-view features; After the multi-view features are weighted by learnable weights, global and cross-view information is learned to obtain the finally fused features; The finally fused features are fed into the task head for segmentation processing.
2. The multimodal point cloud segmentation method according to claim 1, wherein Fusing the image features with the voxel features includes: Based on the intrinsic matrix and extrinsic matrix of the camera, the corresponding relationship between the voxel features and pixel points is obtained; Based on the corresponding relationship, the reference image features are calculated; Multiple image features are sampled, and based on the reference image features, the offsets of the multiple image features are obtained; The voxel features are fused with the sampled image features to obtain voxel features enhanced by the image; The voxel features enhanced by the image are fused with the voxel features to obtain the finally fused voxel features; The corresponding relationship is shown in formula (1): (1) Among them, T is the external parameter matrix of the camera, and the external parameter matrix of the camera includes a rotation matrix and a translation matrix; S is the internal parameter matrix of the camera; (x i , y i , z i ) represents the voxel center, representing the position of the voxel; (u i , v i ) represents the pixel point corresponding to the voxel center.
3. The multimodal point cloud segmentation method according to claim 2, wherein The voxel features are fused with the sampled image features through formula (2): (2) Among them, is the sampled image feature, is the voxel feature enhanced by the image, and are learnable weights, m indexes the attention head, M is the number of attention heads, and L is the number of sampled image features; and respectively represent the sampling offset and attention weight of the l-th sampled image feature in the m-th attention head; i is the sampled point, and p i is the coordinate of the sampling point, is the reference image feature.
4. The multimodal point cloud segmentation method according to claim 3, characterized in that Fusing the image features with the depth map features includes: Based on the reference image features, the depth map features are fused with multiple sampled image features to obtain target image features; The depth map features are fused with the target image features to obtain the finally fused depth map features.
5. The multimodal point cloud segmentation method according to claim 1, wherein The converted depth map features and the converted voxel features are respectively projected onto the point space through trilinear interpolation and linear interpolation.
6. The multimodal point cloud segmentation method according to claim 1, wherein After the multi-view features are weighted by learnable weights, global and cross-view information is learned to obtain the finally fused features, including: Using learnable parameters to learn the weights among three features, the three features are depth map features, voxel features, and point features respectively; After the multi-view features are weighted by learnable weights, they are input into the multi-layer perceptron layer to learn global information from each view to obtain globally enhanced features; The globally enhanced features are input into the multi-layer perceptron to learn the interaction between different views to obtain enhanced features; The enhanced features are respectively concatenated with the depth map features and the voxel features in a residual manner to obtain the finally fused features.
7. The multimodal point cloud segmentation method according to claim 6, characterized in that, The learning process is represented by formula (3): (3) Among them, is an enhanced feature, is a learnable parameter, , , are the voxel feature, depth map feature, and point feature that have been transformed into the point space respectively; represents global; represents a multi-layer perceptron for extracting global point features; represents a multi-layer perceptron for learning the interaction between different perspectives.
8. A multi-modal point cloud segmentation system, characterized in that, Including: A feature extraction module, a learnable cross-modal association module, a learnable cross-view association module, and a task head module connected in sequence; The feature extraction module is used to obtain point features based on the point cloud and the multi-layer perceptron; convert the point cloud to a depth map through spherical projection, and input the depth map into the encoder of the depth map to extract depth map features; input the features after voxelizing the point cloud into the voxel encoder to obtain voxel features; input the image into the network based on ResNet to extract image features; the learnable cross-modal association module is used to fuse the image features with the depth map features and the voxel features respectively to obtain the finally fused depth map features and the finally fused voxel features; The learnable cross-view association module is used to perform point feature conversion on the finally fused depth map features and the finally fused voxel features to obtain the converted depth map features and the converted voxel features; splice the point features, the converted depth map features and the converted voxel features to obtain multi-view features; after the multi-view features are weighted by learnable weights, learn global and cross-view information to obtain the finally fused features; The task head module is used to perform segmentation processing based on the finally fused features output by the learnable cross-view association module.
9. An electronic device, characterized in that, Comprising: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the multi-modal point cloud segmentation method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the multi-modal point cloud segmentation method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Indoor scene modeling method based on visual angle generation
CN110458939A
Indoor mobile robot three-dimensional semantic map construction method
CN115035260A