Point Cloud Instance Segmentation Model Based on Dynamic Convolution and SLAM Map Construction Method
Through the point cloud instance segmentation model based on dynamic convolution, dynamic target mask is generated, and the problems of insufficient accuracy and poor robustness of traditional lidar SLAM algorithms in dynamic environments are solved, and high-precision positioning and mapping are realized in complex environments.
Patent Information
- Application Number
- CN202510312481.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-17
AI Technical Summary
Traditional lidar SLAM algorithms face the problems of insufficient accuracy and poor robustness in dynamic environments, especially in complex and rapidly changing scenarios, which are prone to misjudgment and error removal of static objects, resulting in reduced positioning and mapping accuracy.
A point cloud instance segmentation model based on dynamic convolution is adopted, and a dynamic target mask is generated through a sparse convolution network, masking, offset header, segmentation header, dynamic weight generator and instance decoder, and a dynamic target mask is generated to accurately describe the outline of the dynamic object, reducing the interference of dynamic objects to the static environment mapping.
Improves the robustness and reliability of the lidar SLAM system in complex and rapidly changing dynamic environments, ensuring the accuracy of positioning and mapping.
Smart Images

Figure CN119832254B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of marking scene content, and specifically relates to a point cloud instance segmentation model based on dynamic convolution and a SLAM map construction method. Background Art
[0002] When traditional lidar SLAM algorithms perform map construction in dynamic environments, they often face a series of challenges, including insufficient accuracy, poor robustness, and poor performance in complex environments. These limitations affect the reliability of lidar SLAM systems in rapidly changing scenarios. Especially in dynamic environments with multiple dynamic objects of different shapes and sizes, misjudgment and incorrect removal of static objects are likely to occur, resulting in reduced accuracy of positioning and mapping, and reducing the robustness and reliability of lidar SLAM systems in complex and rapidly changing dynamic environments. Summary of the Invention
[0003] Aiming at the deficiencies of the existing technology, the present invention proposes a point cloud instance segmentation model based on dynamic convolution and a SLAM map construction method, which can improve the recognition accuracy of dynamic objects in complex environments, and thus ensure the robustness and reliability of lidar SLAM systems in complex and rapidly changing dynamic environments. The specific technical solutions are as follows:
[0004] In the first aspect, a point cloud instance segmentation model based on dynamic convolution is provided. In the first implementable manner of the first aspect, it includes:
[0005] A sparse convolution network configured to obtain point cloud data and extract global features from the point cloud data;
[0006] A mask head, an offset head, and a segmentation head, respectively configured to separate mask features, centroid offset features, and semantic features from the global features extracted by the sparse convolution;
[0007] A dynamic weight generator configured to determine position information embedding, class masks, and instance weights according to the centroid offset features, semantic features separated by the offset head and segmentation head, and the global features;
[0008] An instance decoder configured to generate a dynamic target mask corresponding to the point cloud data according to the instance weights and the instance features combined with mask features, class masks, and position information embedding.
[0009] Combined with the first implementable manner of the first aspect, in the second implementable manner of the first aspect, the sparse convolution network is a submanifold sparse convolution network.
[0010] Combined with the first implementation manner of the first aspect, in the third implementation manner of the first aspect, a lightweight Transformer block is configured between the encoder and the decoder of the sparse convolutional network.
[0011] Combined with the first implementation manner of the first aspect, in the fourth implementation manner of the first aspect, the offset head moves the points in the point cloud to the centroid of the corresponding instance to obtain the centroid offset feature, and clusters the points with similar geometric centroids and the same category prediction results.
[0012] Combined with the first implementation manner of the first aspect, in the fifth implementation manner of the first aspect, the dynamic weight generator clusters the points with similar geometric centroids and the same category prediction results according to the centroid offset feature.
[0013] Combined with the first implementation manner of the first aspect, in the sixth implementation manner of the first aspect, by grouping the same points within the cluster and using a sub-network to aggregate the large-scale context, instance-aware instance weights are generated.
[0014] In a second aspect, a method for constructing a SLAM map based on dynamic convolution is provided. In the first implementation manner of the second aspect, it includes:
[0015] Construct the point cloud instance segmentation model and train the point cloud instance segmentation model through a pre-constructed training set;
[0016] Obtain the point cloud image data of the environment to be constructed in real time, and generate a dynamic target mask corresponding to the point cloud image data through the trained cloud instance segmentation model;
[0017] Separate the dynamic point cloud and the static point cloud in the point cloud image data through the dynamic target mask;
[0018] Construct a static map based on the separated static point cloud, construct a dynamic map based on the separated dynamic point cloud, and align and combine the dynamic map and the static map to form a global map.
[0019] Combined with the first implementation manner of the second aspect, in the second implementation manner of the second aspect, obtaining the point cloud image data includes:
[0020] Obtain the environmental scan data collected by the lidar, and perform noise removal processing, ground segmentation processing, and / or clustering segmentation processing on the environmental scan data to obtain the point cloud image data.
[0021] Combined with the first implementation manner of the second aspect, in the third implementation manner of the second aspect, constructing a static map based on the separated static point cloud includes:
[0022] Extract features from the static point cloud to obtain edge feature points and plane feature points corresponding to the current frame;
[0023] Construct a local edge map and a local plane map based on the edge feature points and plane feature points respectively, and construct a local map corresponding to the current frame through the local edge map and the local plane map;
[0024] Perform odometry estimation based on the extracted edge feature points, plane feature points and the constructed local map to determine key frames, and construct a static map based on the static point cloud corresponding to the key frames.
[0025] Combined with the third implementation manner of the second aspect, in the fourth implementation manner of the second aspect, performing odometry estimation based on edge feature points, plane feature points and a local map includes:
[0026] Search for the nearest edge feature points from the local edge map, and for the coordinates of each edge feature point in the local map, select two nearest edge feature points to construct a residual equation from the starting point to the edge;
[0027] Select three nearest plane feature points from the local plane map to construct a residual equation from the starting point to the plane;
[0028] Based on the residual equation from a point to an edge and the residual equation from a point to a plane, determine the transformation relationship between the current frame and the local map, and determine whether the current frame is a key frame according to the transformation relationship.
[0029] Advantageous effects: By using the point cloud instance segmentation model and the SLAM map construction method based on dynamic convolution of the present invention, the dynamic weight generator is set to configure corresponding instance weights for instances of different shapes and sizes, so as to generate a dynamic target mask that can accurately describe the contours of dynamic objects, for precise segmentation of diverse dynamic objects, reduce the interference of dynamic objects on static environment mapping, thus ensuring the accuracy of positioning and mapping, and further improving the robustness and reliability of the lidar SLAM system in complex and rapidly changing dynamic environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the specific embodiments of the present invention, the drawings required for use in the specific embodiments will be briefly introduced below. In all the drawings, the components or parts do not necessarily draw according to the actual ratio.
[0031] Figure 1 It is a schematic diagram of the framework structure of the point cloud instance segmentation model based on dynamic convolution provided by an embodiment of the present invention;
[0032] Figure 2 It is a flowchart of the SLAM map construction method based on dynamic convolution provided by an embodiment of the present invention. Detailed implementation manners
[0033] The embodiments of the technical solution of the present invention will be described in detail below with reference to the accompanying drawings. The following embodiments are only used to illustrate the technical solution of the present invention more clearly, so they are only examples and cannot be used to limit the protection scope of the present invention.
[0034] As Figure 1 shown in the schematic framework diagram of the point cloud instance segmentation model based on dynamic convolution, the segmentation model includes:
[0035] A sparse convolution network configured to obtain point cloud data and extract global features from the point cloud data;
[0036] A mask head, an offset head, and a segmentation head, respectively configured to separate mask features, centroid offset features, and semantic features from the global features extracted by the sparse convolution;
[0037] A dynamic weight generator configured to determine position information embedding, class masks, and instance weights according to the centroid offset features, semantic features separated by the offset head and the segmentation head, and the global features;
[0038] An instance decoder configured to generate a dynamic target mask corresponding to the point cloud data according to the instance weights and the instance features combined with the mask features, class masks, and position information embedding.
[0039] Specifically, the point cloud instance segmentation model includes a sparse convolution network, a mask head, an offset head, a segmentation head, a dynamic weight generator, and an instance decoder. Among them, the sparse convolution network can receive the point cloud data collected by the lidar scanning the surrounding environment and extract the global features of the surrounding environment from the obtained point cloud data. The mask head, the offset head, and the segmentation head can respectively separate mask features, centroid offset features, and semantic features from the global features extracted by the sparse convolution network. The dynamic weight generator can generate corresponding position information embedding, class masks, and instance weights for perceiving different instances in the point cloud data based on the centroid offset features, semantic features separated by the offset head and the segmentation head, and the global features extracted by the sparse convolution network. The instance decoder can generate a dynamic target mask for representing dynamic targets according to the instance features generated by adding position information embedding and class masks to the mask features and the instance weights corresponding to different instances determined by the dynamic weight generator.
[0040] The dynamic target mask can accurately describe the contours of dynamic objects, perform accurate segmentation on diverse dynamic objects in the point cloud data, reduce the interference of dynamic objects on static environment mapping, thereby ensuring the accuracy of positioning and mapping, and further improving the robustness and reliability of the lidar SLAM system in complex and rapidly changing dynamic environments.
[0041] In this embodiment, optionally, the sparse convolutional network is a submanifold sparse convolutional network. Specifically, a submanifold sparse convolutional network can be used to improve the segmentation efficiency, thereby improving the real-time performance and efficiency of the lidar SLAM system in a dynamic environment. Specifically, traditional sparse convolution allows activation positions to spread to all neighboring positions during the convolution process, which may lead to the expansion of the computational region and thus increase the computational amount. Submanifold sparse convolution restricts the convolution operation so that calculations are only performed on the activated positions, and no new activation points are generated in the non-activated regions, effectively reducing computational redundancy. Submanifold sparse convolution does not create new activation points in non-activated regions, so deep feature extraction can be performed without destroying the sparse structure. It maintains the spatial consistency of the point cloud data, avoids introducing noise, and thus improves the accuracy of segmentation. This enables the network to maintain high-quality feature extraction with less computational amount.
[0042] In this embodiment, optionally, a lightweight Transformer block is configured between the encoder and decoder of the sparse convolutional network.
[0043] Specifically, due to the small number of convolutional layers and channels, sparse convolution is often limited by a limited receptive field and representation ability. Therefore, a lightweight Transformer block can be configured in the sparse convolutional network to enhance the long-range interaction between the encoder and decoder.
[0044] In this embodiment, optionally, the offset head obtains the centroid offset feature by moving the points in the point cloud to the centroid of the corresponding instance.
[0045] Specifically, due to the sparsity and non-uniformity of 3D point clouds, the variability of the cross-surface point distribution makes it complex and unstable to aggregate context information. Therefore, the offset head can simplify the distance relationship between different surfaces by moving the points in the global feature to the centroid of the corresponding instance. To supervise the centroid offset prediction of points, the following loss function can be used to minimize the distance between the predicted offset points and the true centroid:
[0046] ;
[0047] where is the coordinate of the th point, represents the th term of the offset, represents the geometric centroid of the corresponding instance, is the effective point representing is the effective point of centroid prediction, represents the total number of effective points.
[0048] In this embodiment, optionally, the dynamic weight generator clusters the mass points with similar geometric centroids and the same class prediction results according to the centroid offset feature.
[0049] Specifically, due to the differences in the shapes and features of different instances, directly using a fixed convolution kernel may lead to inaccurate instance segmentation. Therefore, the weight generator needs to adaptively generate different convolution weights for each instance.
[0050] At the instance level, the sub-network is applied to aggregate the features within the candidate instance, so that the point cloud features within the instance can better capture the overall information of the instance. Through a lightweight MLP (Multi-Layer Perceptron) network, the dynamic weight at the instance level is calculated. The weight generator generates the dynamic weight at the instance level based on the global features, offset prediction, and semantic segmentation results.
[0051] In this embodiment, optionally, by grouping the mass points within the cluster and using the sub-network to aggregate the large-scale context, the instance-aware instance weight is generated.
[0052] The dynamic weight generator can use the geometric centroid of each instance as a reference to calculate the position information embedding of each point in the global feature relative to the centroid, and splice it with the mask feature to enhance the discrimination ability of the instance, so as to effectively distinguish different instances and improve the instance segmentation accuracy in complex environments. The position information embedding is expressed as follows:
[0053] ;
[0054] where represents the geometric centroid of the th instance.
[0055] Specifically, the generation process of the class mask depends on the input of the weight generator, including the global feature, the centroid prediction feature of the offset head, and the semantic segmentation feature of the segmentation head. Through the centroid prediction feature, the geometric centroid offset of each point is calculated, so that the similar points can gather around the centroid of the same instance. Combining the class prediction of the segmentation head ensures that the gathered points belong to the same semantic class, thus forming an instance candidate.
[0056] Calculate the semantic class distribution of the points in each instance candidate, and count the class with the largest number of points. If the class label of the instance candidate meets a certain confidence requirement, it is used as a valid instance. For each instance candidate, a class mask is constructed, which records the class corresponding to the instance and is used for subsequent instance feature extraction and weight application.
[0057] In this embodiment, the instance decoder can consist of three convolution kernels with a size of It consists of convolutional layers, and the activation function of each convolutional layer is ReLU. The loss function of the entire instance segmentation model can be expressed as:
[0058] ;
[0059] where, is the total number of clusters, is the semantic prediction of the -th point, is the semantic label of the -th cluster, is the binary cross-entropy loss function, is the indicator function, indicating that the loss is only calculated on the points that have the same semantic label as the cluster . represents the ground truth mask, represents the predicted mask.
[0060] As shown in Figure 2 is the flowchart of the SLAM map construction method based on dynamic convolution. The construction method includes:
[0061] Step 1: Construct the above-mentioned point cloud instance segmentation model, and train the point cloud instance segmentation model with a pre-constructed training set;
[0062] Step 2: Real-time obtain the point cloud image data of the environment to be constructed, and generate the dynamic target mask corresponding to the point cloud image data through the trained point cloud instance segmentation model;
[0063] Step 3: Separate the dynamic point cloud and the static point cloud in the point cloud image data through the dynamic target mask;
[0064] Step 4: Construct a static map based on the separated static point cloud, and construct a dynamic map based on the separated dynamic point cloud, and align and combine the dynamic map and the static map to form a global map.
[0065] Specifically, first, the above-mentioned point cloud instance segmentation model can be constructed, and the point cloud instance segmentation model can be trained with a pre-constructed training set to obtain a trained point cloud instance segmentation model. Then, the point cloud data of the surrounding environment can be scanned in real time by a lidar, and the dynamic target mask corresponding to the point cloud image data can be generated through the trained point cloud instance segmentation model. After that, the dynamic point cloud and the static point cloud in the point cloud image data can be separated through the dynamic target mask. Finally, a static map can be constructed based on all the separated static point clouds, and at the same time, a dynamic map can be constructed based on the separated dynamic point clouds, and the dynamic map and the static map can be aligned and combined to form a global map.
[0066] The dynamic object mask generated by the point cloud instance segmentation model can accurately describe the contour of the dynamic object, so as to accurately segment the diverse dynamic objects in the point cloud data, reduce the interference of the dynamic objects in the complex dynamic environment to the static environment mapping, ensure the accuracy of positioning and mapping, and further improve the robustness and reliability of the lidar SLAM system in the complex and rapidly changing dynamic environment.
[0067] In this embodiment, optionally, obtaining the point cloud image data includes:
[0068] Obtaining the environmental scan data collected by the lidar, and performing noise removal processing, ground segmentation processing, and / or clustering segmentation processing on the environmental scan data to obtain the point cloud image data.
[0069] Specifically, after the lidar obtains the environmental scan data of the surrounding environment, the point cloud image data of the surrounding environment can be obtained through preprocessing such as noise removal processing, ground segmentation processing, and / or clustering segmentation processing on the environmental scan data. Through noise removal processing, ground segmentation processing, and / or clustering segmentation processing, the quality of the point cloud image data can be improved, providing a solid foundation for subsequent analysis. Similarly, when constructing the training set, the same preprocessing can be performed on the obtained sample data to improve the quality of the samples in the training set.
[0070] In this embodiment, optionally, constructing a static map based on the separated static point cloud includes:
[0071] Performing feature extraction on the static point cloud to obtain edge feature points and plane feature points corresponding to the current frame;
[0072] Constructing a local edge map and a local plane map respectively according to the edge feature points and the plane feature points, and constructing a local map corresponding to the current frame through the local edge map and the local plane map;
[0073] Performing odometry estimation based on the extracted edge feature points, plane feature points, and the constructed local map to determine key frames, and constructing a static map according to the static point cloud corresponding to the key frames.
[0074] Specifically, when constructing the static map, first, feature extraction can be performed on all the separated static point clouds, so as to obtain the edge feature points and plane feature points in the point cloud image data of the current frame. Specifically, based on the set search radius and curvature threshold, the edge feature points and plane feature points can be extracted from all the static point clouds through the constructed KD tree. Specifically, for each point in the static point cloud, according to the search radius, the neighboring points of each point are searched in a spherical shape with the radius, and the curvature between the point and the neighboring points is calculated. If the calculated curvature is greater than the set curvature threshold, the point is an edge feature point, otherwise the point is a plane feature point.
[0075] Then, a local edge map and a local plane map can be constructed respectively based on the extracted edge feature points and plane feature points, and a local map can be constructed by combining the local edge map and the local plane map. After that, based on the constructed local map and the extracted edge feature points and plane feature points, the transformation relationship between the current frame and the local map can be determined through an odometry estimation method, and the transformation relationship includes the rotation change amount and the translation change amount from the current frame to the local map. When the translation change amount or the rotation change amount is greater than a preset threshold, the current frame can be determined as a key frame, and the local map can be updated through a certain number of updated key frames. When a new key frame is added, the oldest key frame is removed to re-form the local map.
[0076] In this embodiment, optionally, odometry estimation based on edge feature points, plane feature points, and the local map includes:
[0077] Search for the nearest edge feature points from the local edge map, and for the coordinates of each edge feature point in the local map, select two nearest edge feature points to construct a residual equation from the starting point to the edge;
[0078] Select three nearest plane feature points from the local plane map to construct a residual equation from the starting point to the plane;
[0079] Based on the residual equation from the point to the edge and the residual equation from the point to the plane, determine the transformation relationship from the current frame to the local map, and determine whether the current frame is a key frame according to the transformation relationship.
[0080] Specifically, the nearest edge feature points can be searched from the local edge map , and for the coordinates of each searched edge feature point in the local map , select two nearest edge feature points and to construct a residual equation from the starting point to the edge, specifically:
[0081]
[0082] Three nearest plane feature points can also be selected from the local plane map , , . The constructed residual equation from the point to the plane is:
[0083]
[0084] The total residual is: . After obtaining the minimum of this residual, the transformation relationship from the point to the local map can be obtained .
[0085] The construction method of the dynamic map is the same as that of the static map, which will not be elaborated here. Each frame of point cloud data contains a timestamp. The timestamp can be used to determine the static map that matches the dynamic map constructed from the dynamic point cloud segmented from the point cloud data. Finally, the matched dynamic map is added to the already constructed static map to reduce the interference of dynamic objects on the mapping of the static environment, thereby ensuring the accuracy of positioning and mapping, further improving the robustness and reliability of the lidar SLAM system in a complex and rapidly changing dynamic environment, and obtaining a static map containing dynamic objects.
[0086] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered by the scope of the claims and the description of the present invention.
Claims
1. A point cloud instance segmentation model based on dynamic convolution, characterized in that Comprising: A sparse convolutional network configured to obtain point cloud data and extract global features from the point cloud data; A mask head, an offset head, and a segmentation head, respectively configured to separate mask features, centroid offset features, and semantic features from the global features extracted by the sparse convolution; A dynamic weight generator configured to determine position information embedding, class masks, and instance weights based on the centroid offset features, semantic features separated by the offset head and segmentation head, and the global features; An instance decoder configured to generate a dynamic target mask corresponding to the point cloud data based on the instance weights and instance features combined with mask features, class masks, and position information embedding; The dynamic weight generator clusters the mass points with similar geometric centroids and the same class prediction results according to the centroid offset features, groups the homogeneous mass points within the cluster, and aggregates the large-scale context using a sub-network to generate instance-aware instance weights.
2. The point cloud instance segmentation model according to claim 1, wherein The sparse convolutional network is a sub-manifold sparse convolutional network.
3. The point cloud instance segmentation model according to claim 1, characterized in that, A lightweight Transformer block is configured between the encoder and decoder of the sparse convolutional network.
4. The point cloud instance segmentation model according to claim 1, wherein The offset head obtains the centroid offset features by moving the points in the point cloud to the centroid of the corresponding instance.
5. A method for constructing a SLAM map based on dynamic convolution, characterized in that, Comprising: Construct a point cloud instance segmentation model as described in any one of claims 1-4, and train the point cloud instance segmentation model through a pre-constructed training set; Obtain the point cloud image data of the environment to be constructed in real time, and generate a dynamic target mask corresponding to the point cloud image data through the trained cloud instance segmentation model; Separate the dynamic point cloud and static point cloud in the point cloud image data through the dynamic target mask; Construct a static map based on the separated static point cloud, construct a dynamic map based on the separated dynamic point cloud, and align and combine the dynamic map and the static map to form a global map; Constructing a static map based on the separated static point cloud includes: Extract features from the static point cloud to obtain edge feature points and plane feature points corresponding to the current frame; Construct a local edge map and a local plane map according to the edge feature points and plane feature points respectively, and construct a local map corresponding to the current frame through the local edge map and the local plane map; Perform odometry estimation based on the extracted edge feature points, plane feature points, and the constructed local map to determine key frames, and construct a static map according to the static point cloud corresponding to the key frames. The construction method of the dynamic map is the same as that of the static map.
6. The SLAM map construction method according to claim 5, wherein Obtaining the point cloud image data includes: Obtain the environmental scan data collected by the lidar, and perform noise removal processing, ground segmentation processing, and / or clustering segmentation processing on the environmental scan data to obtain the point cloud image data.
7. The SLAM map construction method according to claim 5, wherein Performing odometry estimation based on edge feature points, plane feature points, and local maps includes: Search for the nearest edge feature points from the local edge map, and for the coordinates of each edge feature point in the local map, select two nearest edge feature points to construct a residual equation from the starting point to the edge; Select three nearest plane feature points from the local plane map to construct a residual equation from the starting point to the plane; Based on the residual equations from points to edges and from points to planes, determine the transformation relationship from the current frame to the local map, and determine whether the current frame is a key frame according to the transformation relationship.
Citation Information
Patent Citations
Indoor environment 3D semantic map construction method based on point cloud deep learning
CN111798475A
Point cloud instance segmentation method based on weak supervision dynamic convolution
CN119360383A