Point cloud completion method and system based on up-sampling Mamba model
Through the point cloud completion method based on the upsampling Mamba model, the problems of missing point cloud data and broken topological structure are solved, efficient and stable point cloud reconstruction is achieved, and the safety and accuracy of industrial inspection and autonomous driving are improved.
Patent Information
- Application Number
- CN202510856994.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-19
AI Technical Summary
Existing three-dimensional point cloud data has problems such as data missing, uneven density and topological structure fracture during the acquisition process, resulting in a high missed detection rate in industrial inspection and reduced real-time performance of autonomous driving path planning. Traditional completion methods have technical limitations such as high manual intervention costs, low completion efficiency and large errors.
A point cloud completion method based on the upsampling Mamba model is adopted. Through technical means such as farthest point sampling, K-nearest neighbor algorithm, multi-layer perceptron, bidirectional Mamba encoder, feature fusion layer and upsampling layer, feature-preserving downsampling from dense to sparse and dense point cloud reconstruction are achieved.
It improves the integrity and accuracy of point cloud data, enhances the stability of industrial inspection and the safety of autonomous driving, reduces the risk of misjudgment, and meets the quality requirements of industrial applications.
Smart Images

Figure CN120673214A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electronic information technology, and in particular to a point cloud completion method and system based on an upsampling Mamba model. Background Art
[0002] With their high-precision spatial representation capabilities and adaptability to complex scenarios, 3D point clouds have become a core tool driving digital transformation in sectors such as manufacturing, construction, transportation, and industrial inspection. However, industry applications place extremely high demands on the integrity of point cloud data. Incomplete data can lead to multiple cascading problems: significantly increased missed defect detection rates, incomplete reconstruction of building models, and reduced real-time performance of autonomous driving path planning. In real-world projects, due to factors such as sensor occlusion, environmental noise interference, and the physical limitations of acquisition equipment, acquired point cloud data commonly suffers from defects such as missing data, uneven density, and topological structure discontinuities. These technical bottlenecks severely restrict the effectiveness of practical applications. Currently, mainstream traditional completion methods rely on hand-crafted rules or completion strategies based on prior knowledge. These methods not only suffer from high manual intervention costs and low completion efficiency, but are also prone to systematic errors such as geometric distortion and false details. When applied to high-security scenarios such as industrial inspection and autonomous driving, these errors can lead to misjudgments or omissions of critical information, resulting in significant economic losses and even safety hazards, highlighting the technical limitations of traditional methods in complex scenarios.
[0003] With the rapid development of computer graphics, deep learning-based point cloud completion methods have gradually become a research hotspot in the industry. These methods use neural networks to repair the structure and add details to incomplete 3D point cloud data, achieving topological reconstruction and geometric optimization. As a key task in the field of 3D vision, this technology has been widely used in fields such as intelligent manufacturing, autonomous driving, architectural modeling, industrial inspection, and cultural heritage protection. Although breakthroughs in deep learning technology have significantly promoted algorithm innovation and expanded the boundaries of application, technical bottlenecks still exist in areas such as model accuracy, data dependency, and generalization capabilities for complex scenarios. There is still ample room for future development in algorithm optimization, cross-domain adaptation, and real-time performance improvement, providing important research value for promoting intelligent 3D data processing and the digital transformation of the industry. Summary of the Invention
[0004] The present invention is proposed to address the above-mentioned problems, and aims to provide a point cloud completion method and system based on an upsampling Mamba model.
[0005] The present invention provides a point cloud completion method based on an upsampling Mamba model, which has the following characteristics, including the following steps: Step 1, point cloud downsampling: first, the input dense point cloud is downsampled based on the farthest point sampling algorithm, and then the dense point cloud is grouped using the K nearest neighbor algorithm, and finally the grouped point cloud features are aggregated into the sparse points after downsampling through a multi-layer perceptron, completing the feature-preserving downsampling process from dense to sparse to obtain point cloud data; Step 2, point cloud data encoding: in the point cloud data encoding module, the point cloud data is feature encoded based on the bidirectional Mamba encoder as the core network, and a residual connection mechanism of the original features before encoding and the abstract features after encoding is established through a cascade structure to obtain the encoded point cloud data; Step 3, seed point generation: first, the point cloud data is generated based on the Mamba architecture. The feature fusion layer performs spatial clustering and local feature grouping on the encoded point cloud data, then performs multi-scale fusion of global context features and local area features, and combines the position embedding module to encode and integrate the spatial coordinate information of the point cloud, and finally uses the transposed convolution layer to generate seed points with complete geometric structure; Step 4, point cloud upsampling: First, the residual point cloud and the seed point features are aligned through the nearest neighbor search, and then the current features and historical features are cross-layer fused and nonlinearly transformed using the multi-layer perceptron, and finally the fused features are spatially expanded and the geometric details are reconstructed through the upsampling layer based on the Mamba architecture to generate a dense point cloud reconstruction result; Step 5, model training: The point cloud completion model is iteratively optimized by using the preprocessed enhanced dataset to obtain the optimized point cloud completion model.
[0006] The point cloud completion method based on the upsampling Mamba model provided by the present invention may also have the following features: wherein, in step 1, the specific process is: step 1-1, enhancing the data set by adopting methods including random rotation, translation, flipping, jittering and random point masking for the complete point cloud data, and at the same time, preprocessing the point cloud sample set, removing invalid samples, and then dividing the data set into a training set and a validation set; step 1-2, further, performing feature extraction on the preprocessed point cloud data, and adopting the farthest point sampling method to downsample the dense point cloud to obtain a uniformly distributed sparse point cloud; step 1-3, using the K nearest neighbor algorithm, grouping the dense point cloud with the downsampled sparse point cloud as the center, and finally, aggregating the features within the group to the center point through a multi-layer perceptron to complete the feature extraction process.
[0007] The point cloud completion method based on the upsampling Mamba model provided by the present invention may also have the following features: wherein, in step 1-2, the implementation process of the farthest point sampling method is as follows: the farthest point sampling method first randomly selects a starting point as the first point of the sampling sequence; then, the point farthest from the currently selected point set is selected in turn as the next sampling point until the preset number of sparse point clouds is reached. In step 1-3, the specific implementation process is as follows: for each selected sparse point, the K nearest neighbor algorithm is used to search for K nearest neighbor points in its original dense point cloud to form a local neighborhood N i ={X i1 ,X i2 ,...,X iK}, the coordinates and features of the neighborhood points are input into the multi-layer perceptron, local geometric features are extracted through the fully connected layer and the nonlinear activation layer, and finally, the neighborhood features are aggregated to the sparse point X through the maximum pooling operation i , generate the final aggregate feature f' i .
[0008] The point cloud completion method based on the upsampling Mamba model provided by the present invention may also have the following features: wherein, in step 2, the feature encoding of the features extracted after aggregation is specifically implemented by using a bidirectional Mamba encoder as a backbone network to perform feature encoding on the point cloud data, thereby extracting point cloud features, and then selecting point cloud features of different resolutions as input to the seed generation module. The bidirectional Mamba encoder overcomes the limitations of the original model by introducing a bidirectional state space model instead of the traditional unidirectional processing method, making the model more efficient when updating redundant layer parameters. The formula is as follows:
[0009] Bi-SSM(F)=F+[L + SSM(F L+ )+C - SSM(F C- )] (1)
[0010] Where F represents the input point cloud feature; L + SSM is a standard forward state space model.
[0011] The point cloud completion method based on the upsampling Mamba model provided by the present invention may also have the following features: wherein, in step 3, a feature fusion layer based on Mamba is used to fuse global and local features according to the grouping of the point cloud data, and a transposed convolution layer is used to generate seed points for structure completion. The specific process is: the feature fusion layer based on Mamba constructs a local neighborhood through the K-nearest neighbor algorithm, and calculates the local true value of each point to obtain absolute local features; then, a Mamba encoder is used to model these local features and global context features; then, the weights after convolution processing are weighted summed with the modeled features, and then the seed points for missing structure reconstruction are generated through nonlinear mapping of multiple fully connected layers.
[0012] The point cloud completion method based on the upsampling Mamba model provided by the present invention may also have the following features: wherein, in step 4, the specific implementation method of upsampling the generated seed points is as follows: step 4-1, first, perform nearest neighbor interpolation on the residual point cloud and the seed point to fuse historical features, and extract high-order features through a multi-layer perceptron; step 4-2, aggregate local and global context information through two layers of Mamba-based upsampling layers, and then perform residual connection on the features before and after aggregation through a residual fully connected layer to avoid the gradient explosion problem.
[0013] The point cloud completion method based on the upsampling Mamba model provided by the present invention may also have the following features: wherein, in step 4-2, each upsampling layer includes a local truth calculation module and a position embedding module, and the calculated local truth, position information, and upsampling features are fused through the fully connected layer, and then the fused features are selectively screened and encoded with attention features through the bidirectional Mamba encoder, and the distribution of feature vectors is learned through the forward state space model and the feature reverse state space model to obtain unique and more reliable feature information, and finally, the final upsampled point cloud is obtained by three upsamplings of feature upsampling, center point upsampling, and residual result upsampling, and the final results are fused.
[0014] The point cloud completion method based on the upsampling Mamba model provided by the present invention may also have the following features: wherein, in step 5, the point cloud completion model is trained by using a training set in the processed residual point cloud data, and the model is optimized using a loss function to obtain an optimized point cloud completion model. The specific process is as follows: step 5-1, the expansion penalty loss divides the input point cloud data into multiple subsets, and calculates the minimum spanning tree between the points in each subset, and optimizes the model by minimizing the average minimum spanning tree length; step 5-2, the Chamfer Distance L1 loss measures the similarity between two groups of point clouds by calculating the bidirectional nearest neighbor distance between them, and optimizes the model by minimizing the bidirectional nearest neighbor distance between the two groups. The formula is as follows:
[0015]
[0016] Here, p and q represent points in two point clouds, respectively. The distance between each point p in one point cloud and the closest point q in the other point cloud is calculated and used as the distance parameter. The point cloud completion model is evaluated using the test set, and the model parameters are adjusted and the training process is repeated until the loss function falls below a set threshold. This results in an optimized point cloud completion model.
[0017] The present invention also provides a point cloud completion system based on the upsampling Mamba model, which has the following features: a point cloud downsampling module: downsampling the input dense point cloud based on the farthest point sampling algorithm, and then grouping the dense point cloud using the K nearest neighbor algorithm, and finally aggregating the grouped point cloud features into the sparse points after downsampling through a multi-layer perceptron, completing the feature-preserving downsampling process from dense to sparse, and obtaining point cloud data; a point cloud data encoding module: in the point cloud data encoding module, feature encoding is performed on the point cloud data based on the bidirectional Mamba encoder, and a residual connection mechanism of the original features before encoding and the abstract features after encoding is established through a cascade structure to obtain the encoded point cloud data; a seed point generation module: first, the feature fusion layer based on the Mamba architecture is used to generate the point cloud data. The encoded point cloud data is spatially clustered and grouped with local features, and then the global context features and local area features are multi-scale fused, and the spatial coordinate information of the point cloud is encoded and integrated in combination with the position embedding module. Finally, the transposed convolution layer is used to generate seed points with complete geometric structure; point cloud upsampling module: first, the residual point cloud and the seed point features are aligned through the nearest neighbor search, and then the current features and historical features are cross-layer fused and nonlinearly transformed using a multi-layer perceptron. Finally, the fused features are spatially expanded and the geometric details are reconstructed through the upsampling layer based on the Mamba architecture to generate a dense point cloud reconstruction result; model training module: the point cloud completion model is iteratively optimized by using the preprocessed enhanced dataset to obtain the optimized point cloud completion model. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a flow chart of the point cloud completion method based on the upsampling Mamba model in this embodiment;
[0019] Figure 2 This is a structural diagram of the bidirectional Mamba encoder in this embodiment;
[0020] Figure 3 Schematic diagram of a seed point generation module in this embodiment;
[0021] Figure 4 Schematic diagram of the Mamba-based upsampling module in this embodiment. DETAILED DESCRIPTION
[0022] The present invention will be described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several changes and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.
[0023] This embodiment provides a point cloud completion method based on the upsampling Mamba model to efficiently and conveniently provide high-quality complete point clouds for related fields to meet the point cloud quality requirements in industrial applications and ensure the stability of subsequent tasks.
[0024] This embodiment addresses the shortcomings of existing point cloud completion technology in the industrial and security fields. Through research and analysis of current challenges, it proposes targeted solutions. First, this technology improves data integrity in critical scenarios. Point cloud completion technology can repair missing point clouds due to sensor occlusion or equipment limitations. During power line inspections, it can complete the structural data of key transmission tower components, reducing the risk of equipment failure. Second, point cloud completion technology helps identify potential safety hazards. In architectural scanning, by completing detailed data of ancient buildings missing due to viewing angle limitations, cracks in beams and columns or unstable areas of decorative structures can be identified in advance, preventing the risk of collapse during restoration projects. Third, point cloud completion technology enhances environmental perception and safety protection capabilities in complex scenarios. In the field of autonomous driving, by completing point cloud data fragments of road boundaries or pedestrians due to occlusion by dynamic obstacles, the probability of path planning errors can be reduced. Furthermore, by completing point cloud data in the monitored area in real time, it can help identify unauthorized intrusions or equipment abnormalities, providing a basis for timely response for security management.
[0025] In summary, the point cloud completion technology plays a crucial role in industrial scenarios such as power, construction, and transportation by repairing data loss, enhancing environmental perception, and warning of potential risks, thus ensuring operation safety, preventing equipment hidden dangers, and supporting the stable operation of equipment.
[0026] Figure 1 It is the flowchart of the point cloud completion method based on the upsampling Mamba model in this embodiment.
[0027] As Figure 1 shown, the point cloud completion method based on the upsampling Mamba model in this embodiment includes the following steps:
[0028] Step S1, point cloud downsampling: First, downsample the input dense point cloud based on the farthest point sampling algorithm (FarthestPoint Sampling), then group the dense point cloud using the K-nearest neighbor algorithm (KNN), and finally aggregate the point cloud features after grouping into the downsampled sparse points through a multi-layer perceptron (MLP), completing the feature-preserving downsampling process from dense to sparse, thereby achieving efficient extraction and expression of features. The specific process is as follows:
[0029] Step S1-1, enhance the dataset by using methods including random rotation, translation, flipping, jittering, and random point masking on the complete point cloud data. At the same time, preprocess the point cloud sample set, remove invalid samples, and then divide the dataset into a training set and a validation set. This improves the robustness and adaptability of the model in the face of various situations.
[0030] Step S1-2, further, perform feature extraction on the preprocessed point cloud data, and use the farthest point sampling method to downsample the dense point cloud to obtain a uniformly distributed sparse point cloud. The specific implementation process is as follows:
[0031] The farthest point sampling method first randomly selects a starting point as the first point in the sampling sequence; then, sequentially selects the point farthest from the current selected point set as the next sampling point until the preset number of sparse point clouds is reached. (For example, downsample the original number of point clouds from N to M, where M < N). This sampling strategy maximizes the spatial distance between points, ensuring that the sparse point cloud is uniform in geometric distribution and retains the main structural information of the original point cloud. This not only improves the computational efficiency but also enhances the model's ability to capture the main features of the point cloud.
[0032] Step S1-3, use the K-nearest neighbor algorithm to group the dense point cloud centered on the downsampled sparse point cloud, and finally aggregate the intra-group features to the center point through MLP to complete the feature extraction process. The specific implementation process is as follows:
[0033] For each selected sparse point, the K nearest neighbor algorithm is used to search for K nearest neighbor points in its original dense point cloud to form a local neighborhood N i ={X i1 ,X i2 ,...,X iK The coordinates and features of the neighborhood points are input into the MLP, and the local geometric features are extracted through the fully connected layer and the nonlinear activation layer. Finally, the neighborhood features are aggregated to the sparse point X through the maximum pooling operation. i , generate the final aggregate feature f' i .
[0034] Figure 2 This is a structural diagram of the bidirectional Mamba encoder in this embodiment.
[0035] Step S2, point cloud data encoding: In the point cloud data encoding module, the bidirectional Mamba encoder is used as the core network to encode the features of the point cloud data. Subsequently, the encoded features are combined with the original point cloud features via residual connections to retain more original information and enhance model performance. Figure 2 The specific implementation method is:
[0036] In step S2-1, a bidirectional Mamba encoder is used as the backbone network to perform feature encoding on the point cloud data, thereby extracting point cloud features. The network consists of two state space streams, forward and backward, which not only avoids the gradient vanishing phenomenon in backpropagation, but also ensures the efficiency of the model when updating redundant layer parameters.
[0037] Step S2-2: Perform a residual connection between the encoded features and the original point cloud features to enhance feature expression and alleviate the gradient vanishing problem. Subsequently, point cloud features of different resolutions are selected as input to the seed generation module.
[0038] In this embodiment, Mamba is a state-space model (SSM) that achieves linear-time sequence learning by selecting a state space and possesses learning capabilities similar to the Transformer. It implements a selective attention mechanism by replacing the matrix in the traditional state-space model with a learnable neural network.
[0039] Specifically, the Mamba model parameterizes the input matrix D, output matrix M, and attenuation matrix F of the traditional state-space model (SSM) into an input-driven form, enabling the state-space model to dynamically adjust according to the input data. Furthermore, the semi-parallel computing method solves the problem of decreased computational efficiency caused by the destruction of the original SSM convolution characteristics. Specifically, the Mamba model uses scanning computing instead of the traditional time-step recursive computing and utilizes the parallel acceleration characteristics of matrix multiplication to improve computational efficiency. The formula is as follows:
[0040]
[0041] Formula (3) shows the implementation of semi-parallel computing. Where D and M are the dynamic input and output matrices to be calculated, and h t 、s t is the hidden state parameter, through the hidden state parameter h of the previous time state t-1 、s t-1 The calculation of D and M is dynamically updated. Formula (4) is a continuous time state space model, which consists of two parts: the state equation and the observation equation, where F x is the state transfer matrix, D and M are the input matrix and output matrix respectively, ω represents process noise, and V represents observation noise.
[0042] Using a bidirectional Mamba encoder as the backbone network largely solves the long-term dependency problem faced by deep networks when processing sequential data. By introducing a bidirectional state-space model (Bi-SSM) instead of the traditional unidirectional processing method, the limitations of the original model are overcome, making the model more efficient when updating redundant layer parameters. The formula is as follows:
[0043] Bi-SSM(F)=F+[L + SSM(F L+ )+C - SSM(F C- )] (5)
[0044] Where F represents the input point cloud feature; L + The SSM is a standard forward state space model (SSM). Different from the horizontally flipped inverse SSM used in the traditional Mamba encoder, the bidirectional Mamba encoder in this embodiment introduces an inverse SSM with vertical channel flipping (C - The bidirectional connection design ensures the effective fusion of forward and backward information while avoiding the problem of gradient vanishing during back propagation.
[0045] Figure 3 Schematic diagram of the seed point generation module in this embodiment.
[0046] Step S3, seed point generation: First, the encoded point cloud data is spatially clustered and grouped with local features through the feature fusion layer based on the Mamba architecture, and then the global context features and local area features are multi-scale fused, and the spatial coordinate information of the point cloud is encoded and integrated in combination with the position embedding module, and finally the transposed convolution layer is used to generate seed points with complete geometric structure. This can not only effectively capture the detailed information in the point cloud data, but also improve the accuracy and structural integrity of the generated seed points. Seed point generation is as follows Figure 3 , the specific implementation process is:
[0047] In step S3-1, the feature fusion layer based on Mamba constructs a local neighborhood through the K-nearest neighbor algorithm, and calculates the local true value of each point to obtain the absolute local features; then, the Mamba encoder is used to model these local features and global context features.
[0048] In this embodiment, a feature fusion layer based on Mamba is used to fuse global and local features according to the grouping of point cloud data. global Local feature F i local Combined with position embedding technology, fusion features are generated, and the formula is as follows:
[0049]
[0050] In step S3-2, the weights after convolution processing and the modeled features are weighted and summed, and then the seed points for missing structure reconstruction are generated through nonlinear mapping of multiple fully connected layers.
[0051] In this embodiment, the fused feature F final The spatial upsampling is performed through the transposed convolution layer to restore the resolution of the point cloud. The transposed convolution maps the low-resolution features to the high-resolution space by learning the weights of the deconvolution kernel, thereby restoring the detail information. Combined with the coordinate offset prediction, the generated seed point X seed Not only does feature enhancement complete the missing geometric structure in the original point cloud, but it also improves the overall integrity and accuracy of the point cloud.
[0052] Figure 4 Schematic diagram of the Mamba-based upsampling module in this embodiment.
[0053] Step S4, point cloud upsampling: First, the residual point cloud is aligned with the features of the seed point through the nearest neighbor search, and then the current features and historical features are cross-layer fused and nonlinearly transformed using MLP, such as Figure 4As shown in the figure, the fused features are finally spatially expanded and geometric details are reconstructed through the upsampling layer based on the Mamba architecture to generate a dense point cloud reconstruction result. The specific implementation method is:
[0054] In step S4-1, first, nearest neighbor interpolation is performed on the residual defect cloud and the seed point to fuse historical features, and then the fused features are processed using a multi-layer perceptron (MLP).
[0055] In step S4-2, two Mamba-based upsampling layers are used to aggregate local and global context information. Subsequently, a residual fully connected layer is used to perform residual connections on the aggregated features to avoid gradient explosion, thereby obtaining dense point cloud data. This process not only enhances the integrity of the point cloud but also improves the accuracy of detail preservation. The specific implementation is as follows:
[0056] The processed features are upsampled using the upsampling Mamba layer to generate a dense point cloud. This process includes three upsampling operations: feature upsampling, center point upsampling, and residual upsampling, ultimately generating a dense point cloud. The formula is as follows:
[0057] M up =Upsample(F+[L+SSM](F L+ )+[C-SSM](F C- )) (7)
[0058] C up =agg(M up ,Upsample(F center )) (8)
[0059]
[0060] Among them, Upsample is an upsampling interpolation layer using nearest neighbor interpolation, SSM is a standard Mamba encoder, and F L+ is the forward embedding feature, F C- is the backward embedding feature obtained after the channel is vertically flipped. center is the center point P obtained by grouping the dense point cloud P previously center Features. F c residua Representative C up The enhanced features obtained after residual connection.
[0061] The structural integrity and distribution uniformity of the dense point cloud are ensured by upsampling the point cloud data in different dimensions three times.
[0062] In this embodiment, the upsampling layer is different from the standard point cloud upsampling layer structure. Each upsampling layer includes a local truth calculation module and a position embedding module. The calculated local truth, position information, and upsampling features are fused through a fully connected layer. The fused features are then selectively screened and encoded using a bidirectional Mamba encoder. The distribution of feature vectors is learned through a forward state space model and a feature reverse state space model to obtain unique and more reliable feature information. Finally, the final upsampled point cloud is obtained by fusing the final results through three upsampling steps: feature upsampling, center point upsampling, and residual result upsampling. This makes the entire description more coherent while maintaining the accuracy and completeness of the technical details.
[0063] Step S5, model training: iteratively optimize the point cloud completion model by using the preprocessed enhanced dataset to obtain an optimized point cloud completion model.
[0064] In this embodiment, the point cloud completion model is trained by using the training set in the processed residual point cloud data, and the model is optimized using the loss function to obtain the optimized point cloud completion model. The specific process is as follows:
[0065] In step S5-1, the expansion penalty loss divides the input point cloud data into multiple subsets, and calculates the minimum spanning tree between points in each subset, and optimizes the model by minimizing the average minimum spanning tree length.
[0066] In step S5-2, the Chamfer Distance L1 loss measures the similarity between two sets of point clouds by calculating the bidirectional nearest neighbor distance between them, and optimizes the model by minimizing the bidirectional nearest neighbor distance between the two. The formula is as follows:
[0067]
[0068] Where p and q represent points in two point clouds respectively. The distance from each point p in the point cloud to the point q closest to its coordinate position in the other point cloud is calculated and used as the distance parameter.
[0069] Use the test set to evaluate the point cloud completion model, adjust the model parameters and repeat the training process until the value of the loss function is lower than the set threshold, and finally obtain the optimized point cloud completion model.
[0070] This embodiment also provides a point cloud completion system based on an upsampling Mamba model, including:
[0071] The point cloud downsampling module performs point cloud downsampling using the method in step S1 of this embodiment to obtain point cloud data.
[0072] The point cloud data encoding module uses the method in step S2 of this embodiment to encode the point cloud data to obtain encoded point cloud data.
[0073] The seed point generation module generates seed points using the method in step S3 of this embodiment.
[0074] The point cloud upsampling module performs point cloud upsampling using the method in step S4 of this embodiment to generate a dense point cloud reconstruction result.
[0075] The model training module uses the method in step S5 of this embodiment to iteratively optimize the point cloud completion model, thereby obtaining an optimized point cloud completion model.
[0076] Beneficial effects of this embodiment:
[0077] (1) This embodiment uses a large number of publicly available general point cloud completion datasets and preprocesses the data through transformation methods such as rotation, translation, flipping, random dithering, and random point masking. The general point cloud completion model constructed by training with these enhanced data provides high-quality complete point clouds for related fields, meeting the requirements for point cloud quality in industrial applications and ensuring the stability of subsequent tasks. This not only improves the robustness and generalization ability of the model, but also provides reliable guarantees for practical applications.
[0078] (2) This embodiment uses the Mamba model as the backbone network to extract point cloud features, and adopts the farthest point sampling method and the K-nearest neighbor algorithm to downsample the point cloud, which effectively reduces the computational complexity and improves the computational speed. The bidirectional Mamba encoder is used to encode the point cloud features, providing the features learned by the selective attention mechanism for the subsequent seed generation and upsampling layer, thereby enhancing the stability of the model. The Mamba-based feature fusion layer is used to perform global-local feature fusion on the grouped point cloud, thereby enhancing the global structural perception ability of the seed point. Then, by performing nearest neighbor interpolation on the residual point cloud and the seed point, and using the multi-layer perceptron to fuse the historical features, the fused features are upsampled by the Mamba-based upsampling layer, and finally a dense point cloud with a complete structure is obtained. This process not only improves the reliability of the system, but also further ensures the stability of point cloud completion.
[0079] (3) This embodiment aims to address the shortcomings of existing point cloud completion technologies. Through in-depth research and analysis of current problems, a series of variants based on the Mamba model have been developed for point cloud completion scenarios, including key technologies such as feature encoding, feature fusion, and feature upsampling. These improvements provide high-quality, complete point clouds for related fields, meeting the point cloud quality requirements of industrial applications and ensuring the stability of subsequent tasks.
[0080] The above embodiments are preferred examples of the present invention and are not intended to limit the scope of protection of the present invention.
Claims
1. A point cloud completion method based on an upsampling Mamba model, characterized in that: The following steps are involved: Step 1, point cloud downsampling: First, the input dense point cloud is downsampled based on the farthest point sampling algorithm, then the dense point cloud is grouped using the K-nearest neighbor algorithm, and finally the grouped point cloud features are aggregated into the sparse points after downsampling through the multi-layer perceptron, completing the feature-preserving downsampling process from dense to sparse to obtain point cloud data; Step 2, point cloud data encoding: In the point cloud data encoding module, the bidirectional Mamba encoder is used as the core network to perform feature encoding on the point cloud data, and a residual connection mechanism is established between the original features before encoding and the abstract features after encoding through a cascade structure to obtain the encoded point cloud data; Step 3, seed point generation: First, the encoded point cloud data is spatially clustered and grouped by local features using a feature fusion layer based on the Mamba architecture. Global context features and local region features are then multi-scale fused, and the spatial coordinate information of the point cloud is encoded and integrated using a position embedding module. Finally, a transposed convolutional layer is used to generate seed points with a complete geometric structure. Step 4, point cloud upsampling: First, the residual point cloud is aligned with the features of the seed point through nearest neighbor search. Then, the current features and historical features are fused and nonlinearly transformed across layers using a multi-layer perceptron. Finally, the fused features are spatially expanded and geometric details are reconstructed using an upsampling layer based on the Mamba architecture to generate a dense point cloud reconstruction result. Step 5, model training: The point cloud completion model is iteratively optimized using the preprocessed enhanced dataset to obtain an optimized point cloud completion model.
2. A point cloud completion method based on an upsampling Mamba model according to claim 1, Its characteristics are: in, In step 1, the specific process is: Step 1-1: Enhance the dataset by applying random rotation, translation, flipping, jittering, and random point masking to the complete point cloud data. At the same time, preprocess the point cloud sample set to remove invalid samples, and then divide the dataset into a training set and a validation set. Step 1-2: further, feature extraction is performed on the preprocessed point cloud data, and the dense point cloud is downsampled using the farthest point sampling method to obtain a uniformly distributed sparse point cloud; In steps 1-3, the K-nearest neighbor algorithm is used to group the dense point cloud with the downsampled sparse point cloud as the center. Finally, the features within the group are aggregated to the center point through the multi-layer perceptron to complete the feature extraction process.
3. The point cloud completion method based on the upsampling Mamba model according to claim 2, characterized in that: in, In step 1-2, the implementation process of the farthest point sampling method is: The farthest point sampling method first randomly selects a starting point as the first point of the sampling sequence; then, the point farthest from the currently selected point set is selected as the next sampling point until the preset number of sparse point clouds is reached. In steps 1-3, the specific implementation process is as follows: For each selected sparse point, the K nearest neighbor algorithm is used to search for K nearest neighbor points in its original dense point cloud to form a local neighborhood N i ={X i1 ,X i2 ,...,X iK }, the coordinates and features of the neighborhood points are input into the multi-layer perceptron, local geometric features are extracted through the fully connected layer and the nonlinear activation layer, and finally, the neighborhood features are aggregated to the sparse point X through the maximum pooling operation i , generate the final aggregate feature f' i .
4. The point cloud completion method based on the upsampling Mamba model according to claim 1, characterized in that: in, In step 2, the specific implementation method of feature encoding the features extracted after aggregation is: The bidirectional Mamba encoder is used as the backbone network to encode the point cloud data to extract the point cloud features. Subsequently, point cloud features of different resolutions are selected as the input of the seed generation module. The bidirectional Mamba encoder overcomes the limitations of the original model by introducing a bidirectional state space model to replace the traditional unidirectional processing method, making the model more efficient in updating the redundant layer parameters. The formula is as follows: Bi-SSM(F)=F+[L + SSM(F L+ )+C - SSM(F C- )] (1) Where F represents the input point cloud feature; L + SSM is a standard forward state space model.
5. The point cloud completion method based on the upsampling Mamba model according to claim 1, characterized in that: in, In step 3, a feature fusion layer based on Mamba is used to fuse global and local features according to the grouping of point cloud data, and a transposed convolutional layer is used to generate seed points for structure completion. The specific process is as follows: The Mamba-based feature fusion layer constructs a local neighborhood using the K-nearest neighbor algorithm and calculates the local true value of each point to obtain absolute local features. Subsequently, the Mamba encoder is used to model these local features and global context features. Next, the convolutional weights are weighted and summed with the modeled features, and then nonlinear mapping is performed through multiple fully connected layers to generate seed points for reconstructing missing structures.
6. The point cloud completion method based on the upsampling Mamba model according to claim 1, Its characteristics are: in, In step 4, the specific implementation method of upsampling the generated seed points is: In step 4-1, first, perform nearest neighbor interpolation on the residual defect cloud and the seed point to fuse historical features, and extract high-order features through a multi-layer perceptron. In step 4-2, local and global context information is aggregated through two layers of Mamba-based upsampling layers. Subsequently, residual connections are made to the features before and after aggregation through the residual fully connected layer to avoid the gradient explosion problem.
7. The point cloud completion method based on the upsampling Mamba model according to claim 6, characterized in that: in, In step 4-2, each upsampling layer includes a local truth calculation module and a position embedding module. The calculated local truth, position information, and upsampling features are fused through the fully connected layer. Subsequently, the fused features are selectively screened and encoded with attention features through the bidirectional Mamba encoder. The distribution of feature vectors is learned through the forward state space model and the feature reverse state space model to obtain unique and more reliable feature information. Finally, the final upsampled point cloud is obtained by three upsampling steps of feature upsampling, center point upsampling, and residual result upsampling and the final results are fused.
8. The point cloud completion method based on the upsampling Mamba model according to claim 1, characterized in that: in, In step 5, the point cloud completion model is trained by using the training set in the processed residual point cloud data, and the model is optimized using the loss function to obtain the optimized point cloud completion model. The specific process is as follows: Step 5-1, the expansion penalty loss divides the input point cloud data into multiple subsets, and calculates the minimum spanning tree between the points in each subset, and optimizes the model by minimizing the average minimum spanning tree length; In step 5-2, the Chamfer Distance L1 loss measures the similarity between two sets of point clouds by calculating the bidirectional nearest neighbor distance between them, and optimizes the model by minimizing the bidirectional nearest neighbor distance between the two. The formula is as follows: In the formula, p and q represent points in two point clouds respectively. The distance from each point p in the point cloud to the point q closest to its coordinate position in the other point cloud is calculated and used as the distance parameter. Use the test set to evaluate the point cloud completion model, adjust the model parameters and repeat the training process until the value of the loss function is lower than the set threshold, and finally obtain the optimized point cloud completion model.
9. A point cloud completion system based on an upsampling Mamba model, characterized in that: include: Point cloud downsampling module: Downsamples the input dense point cloud based on the farthest point sampling algorithm, then groups the dense point cloud using the K-nearest neighbor algorithm. Finally, the multi-layer perceptron aggregates the grouped point cloud features into the downsampled sparse points, completing the feature-preserving downsampling process from dense to sparse to obtain point cloud data. Point cloud data encoding module: In the point cloud data encoding module, feature encoding is performed on the point cloud data based on a bidirectional Mamba encoder, and a residual connection mechanism is established between the original features before encoding and the abstract features after encoding through a cascade structure to obtain the encoded point cloud data; Seed point generation module: First, the encoded point cloud data is spatially clustered and grouped by local features through a feature fusion layer based on the Mamba architecture. Then, global context features are fused with local region features at multiple scales. The spatial coordinate information of the point cloud is encoded and integrated in combination with the position embedding module. Finally, a transposed convolutional layer is used to generate seed points with a complete geometric structure. Point cloud upsampling module: First, the residual point cloud is aligned with the features of the seed point through nearest neighbor search. Then, the current features and historical features are fused and nonlinearly transformed across layers using a multi-layer perceptron. Finally, the fused features are spatially expanded and geometric details are reconstructed through an upsampling layer based on the Mamba architecture to generate a dense point cloud reconstruction result. Model training module: The point cloud completion model is iteratively optimized using the preprocessed augmented dataset to obtain the optimized point cloud completion model.
Citation Information
Cited By
Convolution, Mama and Transform hybrid efficient point cloud saliency target detection method
CN121459339A