Multi-Stage Mamba Point Cloud Completion Method and Device Based on Semantic and Geometric Guidance
By introducing a multi-stage Mamba point cloud completion method based on semantic and geometric guidance in the point cloud completion task, the problem of complexity and detail loss in the existing Transformer encoder-decoder structure is solved, and efficient point cloud completion and fine-grained detail reconstruction is achieved.
Patent Information
- Application Number
- CN202510261796.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-06
AI Technical Summary
The existing Transformer encoder-decoder structure faces the problems of quadratic complexity and local details loss in point cloud completion tasks, and it is difficult to generate a complete point cloud with fine-grained local details.
A multi-stage Mamba point cloud completion method based on semantic and geometric guidance is proposed. By constructing a point cloud local feature encoding unit, a sparse point cloud generation unit, a multi-sorting strategy Mamba decoder unit and a point cloud upsampling unit combined with Transformer-Mamba, multi-stage completion of point cloud is achieved.
By introducing the Mamba framework and semantic and geometrically guided point cloud sorting strategy, the global attention and detail reconstruction capabilities of point cloud completion have been improved, and the scalability and prediction accuracy of the completion network have been significantly improved.
Smart Images

Figure CN119762721B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of point cloud processing, and particularly to a multi-stage Mamba point cloud completion method and device based on semantic and geometric guidance. Background Art
[0002] Point cloud completion is one of the key tasks in computer vision and point cloud processing, which aims to generate complete and high-quality point clouds from incomplete and low-quality point clouds. Since Convolutional Neural Network (CNN) demonstrates its powerful feature representation, end-to-end trainable paradigm, and excellent performance in 2D tasks, early point cloud completion work attempted to transfer mature methods from 2D completion tasks to 3D point clouds through voxelization and 3D convolution. However, these methods face huge computational costs, which increase cubically with the increase of spatial resolution. With the emergence of PointNet and PointNet++, directly processing 3D point cloud coordinates has become the mainstream of point cloud analysis tasks.
[0003] In recent years, Transformer has become the mainstream framework for current point cloud completion tasks because it can better learn the structural features and long-range correlations between local parts of point clouds. PoinTr regards point cloud completion as a set-to-set transformation problem and proposes a Transformer encoder-decoder structure for point cloud completion. By representing the point cloud as a set of unordered points with positional embeddings, the point cloud can be transformed into a series of local combinations. At the same time, in order to facilitate the Transformer to better utilize the inductive bias of the 3D geometric structure of the point cloud, PoinTr further designs a geometric perception module that can explicitly simulate local geometric relationships.
[0004] However, the quadratic complexity of the attention mechanism in Transformer hinders the scalability of the completion network in long-sequence tasks, and the pooling operation of the encoder will cause the loss of local details, making it difficult to obtain a complete point cloud with fine-grained local details. Summary of the Invention
[0005] The purpose of this application is to propose a multi-stage Mamba point cloud completion method and device based on semantic and geometric guidance for the above-mentioned technical problems, and overcome the problems of quadratic complexity and local detail loss in the existing Transformer encoder-decoder structure.
[0006] In the first aspect, the present invention provides a multi-stage Mamba point cloud completion method based on semantic and geometric guidance, including the following steps:
[0007] Construct and train a multi-stage Mamba point cloud completion model based on semantic and geometric guidance to obtain a trained multi-stage Mamba point cloud completion model; the multi-stage Mamba point cloud completion model includes a point cloud local feature encoding unit combined with Transformer-Mamba, a sparse point cloud generation unit, a multi-sorting strategy Mamba decoder unit, and a point cloud upsampling unit connected in sequence; the sparse point cloud generation unit includes a skeleton point generation and calibration layer with several layers connected in sequence; the multi-sorting strategy Mamba decoder unit includes multi-sorting strategy Mamba decoders in several stages connected in sequence, and each stage of the multi-sorting strategy Mamba decoder includes 4 bidirectional Mamba layers connected in sequence.
[0008] Obtain the incomplete point cloud to be completed and input it into the trained multi-stage Mamba point cloud completion model. The incomplete point cloud passes through the point cloud local feature encoding unit combined with Transformer-Mamba to obtain encoded features. The encoded features are input into the sparse point cloud generation unit to obtain a sparse point cloud; the sparse point cloud is input into the multi-sorting strategy Mamba decoder unit. Each stage of the multi-sorting strategy Mamba decoder adopts 4 sorting strategies, and the 4 sorting strategies are randomly and non-repeatedly assigned to 4 bidirectional Mamba layers to learn the spatial information at different angles in the sparse point cloud, obtaining decoded features. The decoded features pass through the point cloud upsampling unit to obtain the predicted complete point cloud.
[0009] Preferably, the point cloud local feature encoding unit combined with Transformer-Mamba includes a farthest point sampling layer, a K-nearest neighbor query layer, an intra-block Transformer encoding unit, an importance score prediction layer, and an inter-block Mamba encoding unit connected in sequence. The intra-block Transformer encoding unit includes intra-block Transformer encoders connected in sequence, and the inter-block Mamba encoding unit includes inter-block Mamba encoders connected in sequence. The importance score prediction layer includes a non-linear projection layer and a score prediction layer connected in sequence. The non-linear projection layer is composed of a first convolutional layer, a first batch normalization layer, a first ReLU activation function layer, a second convolutional layer, and a layer normalization layer connected in cascade. The score prediction layer is composed of a third convolutional layer, a second batch normalization layer, a second ReLU activation function layer, and a fourth convolutional layer connected in cascade; the calculation process of the point cloud local feature encoding unit combined with Transformer-Mamba is as follows:
[0010] Input the incomplete point cloud into the farthest point sampling layer to select M center points and construct a center point set , and select the K points closest to each center point through the K-nearest neighbor query layer to construct a local point cloud set, obtaining a local point cloud block , where represents the set of real numbers, represents the m-th local point cloud patch after downsampling, and M represents the total number of center points, represents the k-th nearest neighbor point in the local point cloud patch, and K represents the total number of nearest neighbor points, represents the coordinates of the k-th nearest neighbor point in the m-th local point cloud patch, as shown in the following formula:
[0011] ;
[0012] where, represents the function corresponding to the farthest point sampling layer, represents the function corresponding to the K-nearest neighbor point query layer, represents the incomplete point cloud coordinates;
[0013] Input the local point cloud patch into the in-block Transformer encoding unit to obtain the local point cloud patch feature , and the calculation expression of the i-th in-block Transformer encoder is as follows:
[0014]
[0015]
[0016] where, C represents the feature dimension, represents the i-th in-block Transformer encoder, represents the number of in-block Transformer encoders, represents the layer normalization operation, represents the function corresponding to the multi-head self-attention layer, represents the function corresponding to the feed-forward network, represents the output feature of the multi-head self-attention unit in the i-th in-block Transformer encoder, represents the point cloud patch feature output by the (i - 1)-th in-block Transformer encoder, represents the point cloud patch feature output by the i-th in-block Transformer encoder, at when, is the position encoding of, repeat the above steps until obtaining the point cloud patch feature output by the last in-block Transformer encoder and use it as the local point cloud patch feature ;
[0017] Input the local point cloud patch feature into the importance score prediction layer to obtain the importance score of the local point cloud patch, as shown in the following formula:
[0018] ;
[0019] ;
[0020] Among them, represents the non - linear projection layer, represents the importance score of the local point cloud block, represents the function corresponding to the score prediction layer, represents the feature of the projected local point cloud block;
[0021] Sort the local point cloud block features according to the importance score of the local point cloud block to obtain the sorted local point cloud block features and input them into the inter - block Mamba encoding unit to model the dependencies between local point cloud blocks. The expression of the inter - block Mamba encoder is as follows:
[0022] ;
[0023] ;
[0024] Among them, represents the i'-th inter - block Mamba encoder, represents the number of inter - block Mamba encoders, represents the output feature of the bidirectional Mamba unit in the i'-th inter - block Mamba encoder, represents the function corresponding to the bidirectional Mamba layer, represents the local point cloud block feature enhanced by the i'-th inter - block Mamba encoder, represents the local point cloud block feature enhanced by the (i'-1)-th inter - block Mamba encoder. When i' = 1, is the sorted local point cloud block feature ;
[0025] The bidirectional Mamba unit includes a layer normalization operation, a first branch, a second branch, and a first linear projection layer. The first branch includes a second linear projection layer and a first SiLU activation function layer connected in sequence. The second branch includes a third linear projection layer, a depth - wise separable convolutional layer, a second SiLU activation function layer, and a cascade of a forward state - space equation unit and a backward state - space equation unit in parallel. The forward state - space equation unit includes a forward state - space equation and a layer normalization operation connected in sequence. The backward state - space equation unit includes a backward state - space equation and a layer normalization operation connected in sequence. The first linear projection layer, the second linear projection layer, and the third linear projection layer all adopt multi - layer perceptrons. The calculation process of the bidirectional Mamba unit is as follows:
[0026] ;
[0027] ;
[0028] Among them, represents the hidden state, represents the function corresponding to the depthwise separable convolutional layer, represents the function corresponding to the multi-layer perceptron layer, represents the forward state space equation, represents the backward state space equation, represents the SiLU activation function;
[0029] Take the enhanced local point cloud block features output by the last inter-block Mamba encoder as the encoded features .
[0030] Preferably, the skeleton point generation and calibration layer of each layer includes a semantically guided skeleton point cloud generation module and a geometrically guided skeleton point cloud calibration module connected in sequence. The input feature of the skeleton point generation and calibration layer of the current layer is the encoded feature or the feature formed by connecting the input feature of the skeleton point generation and calibration layer of the previous layer and the calibrated skeleton point feature obtained by the skeleton point cloud calibration module of the previous layer, as shown in the following formula:
[0031] ;
[0032] Among them, represents the connection operation, , represents the total number of layers of the skeleton point generation and calibration layer; when , is the encoded feature, represents the input feature of the skeleton point generation and calibration layer of the j-th layer, represents the input feature of the skeleton point generation and calibration layer of the (j - 1)-th layer, represents the calibrated skeleton point feature obtained by the skeleton point cloud calibration module of the (j - 1)-th layer;
[0033] In the skeleton point generation and calibration layer of the current layer, the input features of the skeleton point generation and calibration layer of the current layer first pass through the semantic-guided skeleton point cloud generation module of the current layer to obtain the first skeleton point prediction features of the current layer and the skeleton point coordinates of the current layer. The skeleton point coordinates of the current layer, the input point set of the previous layer, the first skeleton point prediction features of the current layer, and the input features of the skeleton point generation and calibration layer of the current layer are input into the geometry-guided skeleton point cloud calibration module of the current layer to obtain the second skeleton point prediction features of the current layer and the calibrated skeleton point coordinates of the current layer; the calibrated skeleton point coordinates of the current layer are connected to the input point set of the previous layer to obtain the input point set of the current layer and input it into the skeleton point generation and calibration layer of the next layer, and the skeleton point prediction features of the current layer are connected to the calibrated skeleton point features of the previous layer to obtain the calibrated skeleton point features of the current layer and input it into the skeleton point generation and calibration layer of the next layer. Repeat the above steps until the skeleton point coordinates generated by the skeleton point generation and calibration layer of the last layer are obtained and used as the sparse point cloud coordinates, and the calibrated skeleton point features generated by the skeleton point generation and calibration layer of the last layer are obtained and used as the sparse point cloud features. The sparse point cloud coordinates and the sparse point cloud features constitute the sparse point cloud.
[0034] Preferably, the input features of the skeleton point generation and calibration layer of the current layer first pass through the semantic-guided skeleton point cloud generation module of the current layer to obtain the first skeleton point prediction features of the current layer and the skeleton point coordinates of the current layer. The skeleton point coordinates of the current layer, the input point set of the previous layer, the first skeleton point prediction features of the current layer, and the input features of the skeleton point generation and calibration layer of the current layer are input into the geometry-guided skeleton point cloud calibration module of the current layer to obtain the second skeleton point prediction features of the current layer and the calibrated skeleton point coordinates of the current layer, specifically including:
[0035] In the semantic-guided skeleton point cloud generation module of the j-th layer, first perform a pooling operation on the input features of the skeleton point generation and calibration layer of the j-th layer to obtain the global features and input the feature difference between the input features of the skeleton point generation and calibration layer of the j-th layer and the global features into the fourth linear projection layer for linear projection to obtain skeleton point features of the current layer , and its expression is as follows:
[0036] ;
[0037] ;
[0038] where Denote the function corresponding to the pooling layer in the semantic-guided skeleton point cloud generation module of the j-th layer. The pooling layer in the semantic-guided skeleton point cloud generation module of the first layer is the importance-aware pooling layer, and the importance-aware pooling layer is used to re-weight and sum according to the importance scores predicted by the importance score prediction layer. The pooling layers in the semantic-guided skeleton point cloud generation modules of the remaining layers are max pooling layers;
[0039] The input features of the skeleton point generation calibration layer of the j-th layer and the skeleton point features of the current layer After connection, perform farthest point sampling of features and nearest point sampling sorting of features, and input them into two Mamba units to model the skeleton point features and the input features The semantic correlation between them, and obtain the first skeleton point prediction feature of the j-th layer with semantic interaction information , and its expression is as follows:
[0040] ;
[0041] Among them, Denote farthest point sampling of features, Denote nearest point sampling of features, Denote the connection operation, Denote two Mamba units, and the calculation process of the two Mamba units is as follows:
[0042] ;
[0043] ;
[0044] ;
[0045] Among them, Is the function corresponding to the root mean square normalization layer, Denote the input features of the skeleton point generation calibration layer of the j-th layer and the skeleton point features of the current layer The input features after farthest point sampling sorting after connection , Denote the input features of the skeleton point generation calibration layer of the j-th layer and the skeleton point features of the current layer The input features after nearest point sampling sorting after connection , Denote the state space equation;
[0046] Input the first skeleton point prediction feature of the j-th layer into the fifth linear projection layer, and predict to obtain Skeleton point coordinates of the j-th layer , and its expression is as follows:
[0047] ;
[0048] Among them, the fourth linear projection layer and the fifth linear projection layer adopt multi-layer perceptrons;
[0049] In the geometric-guided skeleton point cloud calibration module of the j-th layer, the skeleton point coordinates of the j-th layer and the input point set of the (j - 1)-th layer are connected and then farthest point sampling and nearest point sampling sorting are performed respectively, and the first skeleton point prediction feature of the j-th layer and the input feature of the skeleton point generation calibration layer of the j-th layer are sorted according to the sorting results, and the sorted connection results are input into two Mamba units to model the geometric correlation between them, and the second skeleton point prediction feature of the j-th layer containing spatial geometric information is obtained , and its expression is as follows:
[0050] ;
[0051] Among them, represents farthest point sampling, represents nearest point sampling, and respectively represent sorting according to farthest point sampling and nearest point sampling. When , the input point set is the center point set ;
[0052] Secondly, the second skeleton point prediction feature of the j-th layer is added to the global feature and then input into the sixth linear projection layer to obtain the point offset , and it is added to the skeleton point coordinates of the j-th layer obtained in the semantic-guided skeleton point cloud generation unit to obtain the calibrated skeleton point coordinates of the j-th layer , and its expression is as follows:
[0053] ;
[0054] ;
[0055] Among them, the sixth linear projection layer adopts multi-layer perceptrons.
[0056] Preferably, the calculation expression of the decoded feature is:
[0057] ;
[0058] ;
[0059] Among them, represents random allocation, represents the sorting method, , , and are the Z-curve sorting strategy, the inverse Z-curve sorting strategy, the Hilbert sorting strategy, and the inverse Hilbert sorting strategy respectively, represents the function corresponding to the bidirectional Mamba layer executed according to the nth sorting method in the mth stage, , , represents the total number of stages in the multi-sorting strategy Mamba decoder unit, represents the sparse point cloud feature, represents the decoded feature;
[0060] The decoded feature passes through the point cloud upsampling unit to obtain the predicted complete point cloud, specifically including:
[0061] Using the decoded feature to calculate the first affine parameter and the second affine parameter , and guiding the two-dimensional grid to deform with the affine function, and obtaining the three-dimensional offset through three deformations. At the same time, each point coordinate in the sparse point cloud coordinates is copied O times and added to the three-dimensional offset to finally obtain the predicted complete point cloud , and its expression is as follows:
[0062] ;
[0063] ;
[0064] ;
[0065] ;
[0066] ;
[0067] Among them, represents the max pooling operation, represents the function corresponding to the multi-layer perceptron, and represent the mean and standard deviation of the two-dimensional grid , represents the copy operation, represents the sparse point cloud coordinates, Represents the th sparse point in the sparse point cloud, represents the decoded feature of the th sparse point, represents the ReLU activation function, represents the second affine parameter corresponding to the th sparse point, represents the point coordinates of the th sparse point, represents the upsampled point generated from the th sparse point.
[0068] Preferably, the total loss function used in the training process of the multi-stage Mamba point cloud completion model is the sum of the importance loss and the chamfer loss. The calculation process of the importance loss is shown as follows:
[0069] ;
[0070] ;
[0071] ;
[0072] ;
[0073] where represents the function corresponding to the pooling layer. Here, the pooling layer is the importance-aware pooling layer, and the importance-aware pooling layer is used to re-weight and sum according to the importance scores predicted by the importance score prediction layer, represents the global encoded feature, represents the projected global encoded feature, represents the true importance score calculated by cosine similarity, represents the L1 smooth loss, represents the importance loss;
[0074] The calculation process of the chamfer loss is shown as follows:
[0075] ;
[0076] where represents the chamfer loss, represents the true complete point cloud, represents the number of points in the predicted complete point cloud, represents the number of points in the true complete point cloud, represents the predicted complete point cloud, and respectively represent one point in the predicted complete point cloud and one point in the true complete point cloud, Represents the square of the Euclidean distance.
[0077] In a second aspect, the present invention provides a multi-stage Mamba point cloud completion device based on semantic and geometric guidance, including:
[0078] A model construction module configured to construct and train a multi-stage Mamba point cloud completion model based on semantic and geometric guidance to obtain a trained multi-stage Mamba point cloud completion model; the multi-stage Mamba point cloud completion model includes a point cloud local feature encoding unit combined with Transformer-Mamba, a sparse point cloud generation unit, a multi-sorting strategy Mamba decoder unit, and a point cloud upsampling unit connected in sequence; the sparse point cloud generation unit includes a skeleton point generation and calibration layer with several layers connected in sequence; the multi-sorting strategy Mamba decoder unit includes multi-sorting strategy Mamba decoders in several stages connected in sequence, and each stage of the multi-sorting strategy Mamba decoder includes 4 bidirectional Mamba layers connected in sequence;
[0079] A completion module configured to obtain an incomplete point cloud to be completed and input it into the trained multi-stage Mamba point cloud completion model. The incomplete point cloud passes through the point cloud local feature encoding unit combined with Transformer-Mamba to obtain encoded features, and the encoded features are input into the sparse point cloud generation unit to obtain a sparse point cloud; the sparse point cloud is input into the multi-sorting strategy Mamba decoder unit. Each stage of the multi-sorting strategy Mamba decoder adopts 4 sorting strategies, and the 4 sorting strategies are randomly and non-repeatedly assigned to 4 bidirectional Mamba layers to learn the spatial information at different angles in the sparse point cloud to obtain decoded features. The decoded features pass through the point cloud upsampling unit to obtain a predicted complete point cloud.
[0080] In a third aspect, the present invention provides an electronic device, including one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect.
[0081] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.
[0082] In a fifth aspect, the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.
[0083] Compared with the prior art, the present invention has the following beneficial effects:
[0084] (1) The multi-stage Mamba point cloud completion method based on semantic and geometric guidance proposed by the present invention introduces the Mamba framework with improved global attention and linear complexity, and serializes the input point cloud through a semantic and geometric jointly guided point cloud sorting strategy, giving full play to the modeling ability of the Mamba framework for causal sequences.
[0085] (2) The multi-stage Mamba point cloud completion method based on semantic and geometric guidance proposed by the present invention designs a skeleton point generation calibration module, converts point cloud completion into a stage-by-stage generation process, encourages the interaction between local features, and significantly improves the detail reconstruction ability of the point cloud completion method.
[0086] (3) The multi-stage Mamba point cloud completion method based on semantic and geometric guidance proposed by the present invention constructs a total loss function by combining importance loss and chamfer loss, and uses the total loss function to train the multi-stage Mamba point cloud completion model, which not only restricts the generation of importance scores, but also restricts the final point cloud generation, greatly improving the prediction accuracy of the trained multi-stage Mamba point cloud completion model. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0088] Figure 1 It is a schematic flowchart of the multi-stage Mamba point cloud completion method based on semantic and geometric guidance for the embodiments of this application;
[0089] Figure 2 It is a schematic diagram of the multi-stage Mamba point cloud completion model of the multi-stage Mamba point cloud completion method based on semantic and geometric guidance for the embodiments of this application;
[0090] Figure 3 It is a schematic diagram of the point cloud local feature encoding unit of the Transformer-Mamba combination of the multi-stage Mamba point cloud completion method based on semantic and geometric guidance for the embodiments of this application;
[0091] Figure 4 It is a schematic diagram of the bidirectional Mamba unit of the multi-stage Mamba point cloud completion method based on semantic and geometric guidance for the embodiments of this application;
[0092] Figure 5Schematic diagram of the semantic-guided skeleton point cloud generation module of the multi-stage Mamba point cloud completion method based on semantic and geometric guidance according to the embodiments of the present application;
[0093] Figure 6 Schematic diagram of the geometric-guided skeleton point cloud calibration module of the multi-stage Mamba point cloud completion method based on semantic and geometric guidance according to the embodiments of the present application;
[0094] Figure 7 Schematic diagram of the multi-sorting strategy Mamba decoder of the multi-stage Mamba point cloud completion method based on semantic and geometric guidance according to the embodiments of the present application;
[0095] Figure 8 Schematic diagram of the point cloud upsampling unit of the multi-stage Mamba point cloud completion method based on semantic and geometric guidance according to the embodiments of the present application;
[0096] Figure 9 Schematic diagram of the calculation process of the three-dimensional offset of the multi-stage Mamba point cloud completion method based on semantic and geometric guidance according to the embodiments of the present application;
[0097] Figure 10 Schematic diagram of the calculation process of the importance loss of the multi-stage Mamba point cloud completion method based on semantic and geometric guidance according to the embodiments of the present application;
[0098] Figure 11 Schematic diagram of the multi-stage Mamba point cloud completion device based on semantic and geometric guidance according to the embodiments of the present application;
[0099] Figure 12 Schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present invention. Detailed implementation manners
[0100] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0101] Figure 1 An embodiment of the present application provides a multi-stage Mamba point cloud completion method based on semantic and geometric guidance, including the following steps:
[0102] S1. Construct and train a multi-stage Mamba point cloud completion model based on semantic and geometric guidance to obtain a trained multi-stage Mamba point cloud completion model. The multi-stage Mamba point cloud completion model includes a point cloud local feature encoding unit combined with Transformer-Mamba, a sparse point cloud generation unit, a multi-sorting strategy Mamba decoder unit, and a point cloud upsampling unit connected in sequence. The sparse point cloud generation unit includes a series of skeleton point generation and calibration layers connected in sequence. The multi-sorting strategy Mamba decoder unit includes multi-sorting strategy Mamba decoders in several stages connected in sequence, and each stage of the multi-sorting strategy Mamba decoder includes 4 bidirectional Mamba layers connected in sequence.
[0103] Specifically, referring to Figure 2 , an embodiment of the present application proposes a multi-stage Mamba point cloud completion model, which includes a point cloud local feature encoding unit combined with Transformer-Mamba, a sparse point cloud generation unit, a multi-sorting strategy Mamba decoder unit, and a point cloud upsampling unit connected in sequence.
[0104] In a specific embodiment, the point cloud local feature encoding unit combined with Transformer-Mamba includes a farthest point sampling layer, a K-nearest neighbor query layer, an intra-block Transformer encoding unit, an importance score prediction layer, and an inter-block Mamba encoding unit connected in sequence. The intra-block Transformer encoding unit includes intra-block Transformer encoders connected in sequence, and the inter-block Mamba encoding unit includes inter-block Mamba encoders connected in sequence. The importance score prediction layer includes a non-linear projection layer and a score prediction layer connected in sequence. The non-linear projection layer is composed of a first convolutional layer, a first batch normalization layer, a first ReLU activation function layer, a second convolutional layer, and a layer normalization layer cascaded. The score prediction layer is composed of a third convolutional layer, a second batch normalization layer, a second ReLU activation function layer, and a fourth convolutional layer cascaded. The calculation process of the point cloud local feature encoding unit combined with Transformer-Mamba is as follows:
[0105] Input the incomplete point cloud into the farthest point sampling layer to select M center points and construct a center point set , and select the K points closest to each center point through the K-nearest neighbor query layer to construct a local point cloud set, obtaining local point cloud blocks , where represents the set of real numbers, represents the m-th downsampled local point cloud block, M represents the total number of center points, denotes the k-th nearest neighbor point in the local point cloud patch, and K represents the total number of nearest neighbor points. denotes the coordinates of the k-th nearest neighbor point in the m-th local point cloud patch, as shown in the following formula:
[0106] ;
[0107] where, denotes the function corresponding to the farthest point sampling layer, denotes the function corresponding to the K-nearest neighbor point query layer, denotes the incomplete point cloud coordinates;
[0108] Input the local point cloud patch into the in-block Transformer encoding unit to obtain the local point cloud patch feature , and the calculation expression of the i-th in-block Transformer encoder is as follows:
[0109]
[0110]
[0111] where, C represents the feature dimension, denotes the i-th in-block Transformer encoder, denotes the number of in-block Transformer encoders, denotes the layer normalization operation, denotes the function corresponding to the multi-head self-attention layer, denotes the function corresponding to the feed-forward network, denotes the output feature of the multi-head self-attention unit in the i-th in-block Transformer encoder, denotes the point cloud patch feature output by the (i - 1)-th in-block Transformer encoder, denotes the point cloud patch feature output by the i-th in-block Transformer encoder. At time, is the position encoding. Repeat the above steps until the point cloud patch feature output by the last in-block Transformer encoder is obtained and used as the local point cloud patch feature ;
[0112] Input the local point cloud patch feature into the importance score prediction layer to obtain the importance score of the local point cloud patch, as shown in the following formula:
[0113] ;
[0114] ;
[0115] Among them, represents the non-linear projection layer, represents the importance score of the local point cloud patch, represents the function corresponding to the score prediction layer, represents the feature of the projected local point cloud patch;
[0116] Sort the local point cloud patch features according to the importance score of the local point cloud patch to obtain the sorted local point cloud patch features and input them into the inter-block Mamba encoding unit to model the dependencies between local point cloud patches. The expression of the inter-block Mamba encoder is as follows:
[0117] ;
[0118] ;
[0119] Among them, represents the i'-th inter-block Mamba encoder, represents the number of inter-block Mamba encoders, represents the output feature of the bidirectional Mamba unit in the i'-th inter-block Mamba encoder, represents the function corresponding to the bidirectional Mamba layer, represents the local point cloud patch feature enhanced by the i'-th inter-block Mamba encoder, represents the local point cloud patch feature enhanced by the (i'-1)-th inter-block Mamba encoder. When i' = 1, is the sorted local point cloud patch feature ;
[0120] The bidirectional Mamba unit includes a layer normalization operation, a first branch, a second branch, and a first linear projection layer. The first branch includes a second linear projection layer and a first SiLU activation function layer connected in sequence. The second branch includes a third linear projection layer, a depthwise separable convolution layer, a second SiLU activation function layer, and a cascade of a forward state space equation unit and a backward state space equation unit connected in parallel. The forward state space equation unit includes a forward state space equation and a layer normalization operation connected in sequence. The backward state space equation unit includes a backward state space equation and a layer normalization operation connected in sequence. The first linear projection layer, the second linear projection layer, and the third linear projection layer all use multi-layer perceptrons. The calculation process of the bidirectional Mamba unit is as follows:
[0121] ;
[0122] ;
[0123] Among them, represents the hidden state, represents the function corresponding to the depthwise separable convolutional layer, represents the function corresponding to the multi-layer perceptron layer, represents the forward state space equation, represents the backward state space equation, represents the SiLU activation function;
[0124] Take the enhanced local point cloud block features output by the last inter-block Mamba encoder as the encoded features .
[0125] Specifically, referring to Figure 3 , the local feature encoding unit of Transformer-Mamba is constructed by adding three key designs on the basis of the standard Mamba model: (1) the intra-block Transformer encoding unit; (2) the importance score prediction layer; (3) the inter-block Mamba encoding unit.
[0126] Specifically, by introducing the intra-block Transformer encoding unit to locally aggregate local point cloud blocks, local point cloud block features are obtained, solving the problem of insufficient local modeling ability of the Mamba model. Then, the local point cloud block features pass through the importance score prediction layer to calculate and sort the importance scores of the local point clouds, reducing the adverse impact of random sorting on the Mamba model. Finally, the sorted local point cloud block features pass through the inter-block Mamba encoding unit to scan the local point cloud block sequence in both forward and backward directions, realizing the bidirectional interaction of the local point cloud block sequence. Therefore, the proposed intra-block Transformer encoding unit and importance score prediction layer can provide a strengthened representation of the serialized local point cloud blocks for the inter-block Mamba encoding unit, ensuring that each enhanced local point cloud block feature fully aggregates the information from all other local point cloud blocks. Referring to Figure 4 , the bidirectional Mamba unit is improved on the basis of the Transformer encoder, replacing the multi-head self-attention layer and layer normalization in the multi-head self-attention unit in the Transformer with the bidirectional Mamba layer.
[0127] In a specific embodiment, the skeleton point generation and calibration layer of each layer includes a semantic-guided skeleton point cloud generation module and a geometry-guided skeleton point cloud calibration module connected in sequence. The input feature of the skeleton point generation and calibration layer of the current layer is the encoded feature or the feature formed by connecting the input feature of the skeleton point generation and calibration layer of the previous layer and the calibrated skeleton point feature obtained by the skeleton point cloud calibration module of the previous layer, as shown in the following formula:
[0128] ;
[0129] Among them, represents a connection operation, , represents the total number of layers of the skeleton point generation calibration layer; when then, is the encoded feature, represents the input feature of the j-th layer of the skeleton point generation calibration layer, represents the input feature of the (j - 1)-th layer of the skeleton point generation calibration layer, represents the calibrated skeleton point feature obtained by the skeleton point cloud calibration module of the (j - 1)-th layer;
[0130] In the skeleton point generation calibration layer of the current layer, the input feature of the current layer's skeleton point generation calibration layer first passes through the semantic-guided skeleton point cloud generation module of the current layer to obtain the first skeleton point prediction feature of the current layer and the skeleton point coordinates of the current layer. The skeleton point coordinates of the current layer, the input point set of the previous layer, the first skeleton point prediction feature of the current layer, and the input feature of the current layer's skeleton point generation calibration layer are input into the geometric-guided skeleton point cloud calibration module of the current layer to obtain the second skeleton point prediction feature of the current layer and the calibrated skeleton point coordinates of the current layer; the calibrated skeleton point coordinates of the current layer are connected with the input point set of the previous layer to obtain the input point set of the current layer and input it into the skeleton point generation calibration layer of the next layer, and the skeleton point prediction feature of the current layer is connected with the calibrated skeleton point feature of the previous layer to obtain the calibrated skeleton point feature of the current layer and input it into the skeleton point generation calibration layer of the next layer. Repeat the above steps until the skeleton point coordinates generated by the skeleton point generation calibration layer of the last layer are obtained and used as the sparse point cloud coordinates, and the calibrated skeleton point features generated by the skeleton point generation calibration layer of the last layer are obtained and used as the sparse point cloud features. The sparse point cloud coordinates and the sparse point cloud features constitute the sparse point cloud.
[0131] In a specific embodiment, the input feature of the current layer's skeleton point generation calibration layer first passes through the semantic-guided skeleton point cloud generation module of the current layer to obtain the first skeleton point prediction feature of the current layer and the skeleton point coordinates of the current layer. The skeleton point coordinates of the current layer, the input point set of the previous layer, the first skeleton point prediction feature of the current layer, and the input feature of the current layer's skeleton point generation calibration layer are input into the geometric-guided skeleton point cloud calibration module of the current layer to obtain the second skeleton point prediction feature of the current layer and the calibrated skeleton point coordinates of the current layer, specifically including:
[0132] In the semantic-guided skeleton point cloud generation module of the j-th layer, first perform a pooling operation on the input feature of the j-th layer's skeleton point generation calibration layer to obtain the global feature , and the input feature and the global features The feature difference between them is input into the fourth linear projection layer for linear projection to obtain the skeleton point features of the current layer , and its expression is as follows:
[0133] ;
[0134] ;
[0135] Among them, represents the function corresponding to the pooling layer in the semantic-guided skeleton point cloud generation module of the j-th layer. The pooling layer in the semantic-guided skeleton point cloud generation module of the first layer is the importance-aware pooling layer. The importance-aware pooling layer is used to re-weight and sum according to the importance score predicted by the importance score prediction layer. The pooling layers in the semantic-guided skeleton point cloud generation modules of the remaining layers are the max pooling layers;
[0136] The input features of the skeleton point generation calibration layer of the j-th layer and the skeleton point features of the current layer are connected and then subjected to farthest point sampling of features and nearest point sampling sorting of features, and are input into two-way Mamba units to model the semantic correlation between the skeleton point features and the input features to obtain the first skeleton point prediction feature of the j-th layer with semantic interaction information , and its expression is as follows:
[0137] ;
[0138] Among them, represents farthest point sampling of features, represents nearest point sampling of features, represents the connection operation, represents two-way Mamba units. The calculation process of the two-way Mamba units is as follows:
[0139] ;
[0140] ;
[0141] ;
[0142] Among them, is the function corresponding to the root mean square normalization layer, represents the input features of the skeleton point generation calibration layer of the j-th layer and the skeleton point features of the current layer the input features after farthest point sampling sorting after connection , denotes the input feature of the skeleton point generation calibration layer for the j-th layer and the skeleton point feature of the current layer After connection, the input feature after nearest point sampling sorting of the features , represents the state space equation;
[0143] Input the first skeleton point prediction feature of the j-th layer into the fifth linear projection layer, and predict to obtain the skeleton point coordinates of the j-th layer , and its expression is as follows:
[0144] ;
[0145] Among them, the fourth linear projection layer and the fifth linear projection layer adopt multi-layer perceptrons;
[0146] In the geometric-guided skeleton point cloud calibration module of the j-th layer, the skeleton point coordinates of the j-th layer and the input point set of the (j - 1)-th layer After connection, farthest point sampling and nearest point sampling sorting are performed respectively, and according to the sorting results, the first skeleton point prediction feature of the j-th layer and the input feature of the skeleton point generation calibration layer of the j-th layer The connection results are sorted, and after sorting, they are input into two Mamba units to model the geometric correlation between them, and the second skeleton point prediction feature of the j-th layer containing spatial geometric information is obtained , and its expression is as follows:
[0147] ;
[0148] Among them, represents farthest point sampling, represents nearest point sampling, and respectively represent sorting according to farthest point sampling and nearest point sampling, and when the input point set is the center point set ;
[0149] Secondly, add the second skeleton point prediction feature of the j-th layer to the global feature , and then input it into the sixth linear projection layer to obtain the point offset , and add it to the skeleton point coordinates of the j-th layer obtained in the semantic-guided skeleton point cloud generation unit to obtain the calibrated skeleton point coordinates of the j-th layer , and its expression is as follows:
[0150] ;
[0151] ;
[0152] Among them, the sixth linear projection layer adopts a multi-layer perceptron.
[0153] Specifically, referring to Figure 5 , the skeleton point generation and calibration layer composed of the semantic-guided skeleton point cloud generation module and the geometry-guided skeleton point cloud calibration module has a total of L layers. In each layer, the semantic-guided skeleton point cloud generation module and the geometry-guided skeleton point cloud calibration module correspond one by one. The input feature of the current layer's skeleton point generation and calibration layer is composed of the input feature of the previous layer's skeleton point generation and calibration layer and the second skeleton point prediction feature output by the geometry-guided skeleton point cloud calibration module of the previous layer. Among them is the encoded feature . The skeleton point coordinates Figure 6 of the j-th layer and the first skeleton point prediction feature of the j-th layer are output through the skeleton point generation and calibration layer. Referring to , the connection result of the first skeleton point prediction feature of the j-th layer and the input feature of the skeleton point generation and calibration layer of the j-th layer is further input into the geometry-guided skeleton point cloud calibration module, and combined with the skeleton point coordinates of the j-th layer and the input point set of the (j - 1)-th layer for calibration to obtain the calibrated skeleton point coordinates and the second skeleton point feature . The calibrated skeleton point coordinates are connected to the input coordinates to obtain , and are input into the geometry-guided skeleton point cloud calibration module in the skeleton point generation and calibration layer of the next layer. The second skeleton point feature is connected to the input feature to obtain
[0154] , and is input into the skeleton point generation and calibration layer of the next layer. . Finally, a complete but coarse-grained sparse point cloud is obtained through multi-stage generation and calibration. The sparse point cloud coordinates are
[0155] In a specific embodiment, the calculation expression of the decoded feature is:
[0156] ;
[0157] ;
[0158] Among them, represents random assignment, represents the sorting method, , , and are the Z-curve sorting strategy, the anti-Z-curve sorting strategy, the Hilbert sorting strategy, and the anti-Hilbert sorting strategy respectively, represents the function corresponding to the bidirectional Mamba layer executed according to the nth sorting method in the mth stage, , , represents the total number of stages in the multi-sorting strategy Mamba decoder unit, represents the sparse point cloud feature, represents the decoded feature;
[0159] The decoded feature passes through the point cloud upsampling unit to obtain the predicted complete point cloud, which specifically includes:
[0160] Using the decoded feature to calculate the first affine parameter and the second affine parameter , and guiding the two-dimensional grid to deform with the affine function, and obtaining the three-dimensional offset through three deformations. At the same time, each point coordinate in the sparse point cloud coordinates is copied O times and added to the three-dimensional offset to finally obtain the predicted complete point cloud , and its expression is as follows:
[0161] ;
[0162] ;
[0163] ;
[0164] ;
[0165] ;
[0166] Among them, represents the max pooling operation, represents the function corresponding to the multi-layer perceptron, and represent the mean and standard deviation of the two-dimensional grid , represents the copy operation, represents the sparse point cloud coordinates, represents the th sparse point in the sparse point cloud, represents the decoded feature of the th sparse point, represents the ReLU activation function, represents the second affine parameter corresponding to the th sparse point, represents the point coordinates of the th sparse point, represents the upsampled point generated from the
[0167] th sparse point. Figure 7 Specifically, referring to , the multi-sorting strategy Mamba decoder unit has a total of
[0168] stages, and each stage of the multi-sorting strategy Mamba decoder consists of 4 bidirectional Mamba layers. Four sorting strategies are adopted in the multi-sorting strategy Mamba decoder, namely the Z-curve sorting strategy, the inverse Z-curve sorting strategy, the Hilbert sorting strategy, and the inverse Hilbert sorting strategy. The four sorting strategies are randomly and non-repeatedly assigned to the 4 bidirectional Mamba layers to learn the spatial information of different angles of the sparse point cloud and finally obtain the decoded features. Figure 8 and 9 , the decoded features are input into the point cloud upsampling unit to generate upsampled points, and the upsampled points are combined with the sparse point cloud coordinates to obtain a complete point cloud with complete fine-grainedness . Both the seventh linear projection layer and the eighth linear projection layer in the point cloud upsampling unit adopt multi-layer perceptrons.
[0169] In a specific embodiment, the total loss function used in the training process of the multi-stage Mamba point cloud completion model is the sum of the importance loss and the chamfer loss. The calculation process of the importance loss is shown in the following formula:
[0170] ;
[0171] ;
[0172] ;
[0173] ;
[0174] where represents the function corresponding to the pooling layer. Here, the pooling layer is an importance-aware pooling layer, and the importance-aware pooling layer is used to re-weight and sum according to the importance scores predicted by the importance score prediction layer, represents the global encoded feature, Represents the globally encoded features after projection, Represents the true importance score calculated by cosine similarity, Represents the L1 smooth loss, Represents the importance loss;
[0175] The calculation process of the chamfer loss is shown as follows:
[0176] ;
[0177] Among them, Represents the chamfer loss, Represents the true complete point cloud, Represents the number of points in the predicted complete point cloud, Represents the number of points in the true complete point cloud, Represents the predicted complete point cloud, and Represent one point in the predicted complete point cloud and one point in the true complete point cloud respectively, Represents the square of the Euclidean distance.
[0178] Specifically, referring to Figure 10 , first calculate the importance loss for constraining the importance score prediction module. During the calculation of the importance loss, first project the projected local point cloud patch features and the globally encoded features to the same feature space through a non-linear projection layer, then calculate the cosine similarity between the projected local point cloud patch features and the globally encoded features, and finally calculate the L1 smooth loss between the importance score and the cosine similarity to improve the prediction accuracy.
[0179] Finally, calculate the chamfer loss between the true complete point cloud and the predicted complete point cloud, and add it to the importance loss as the final loss function as shown in the following formula:
[0180] .
[0181] S2. Obtain the incomplete point cloud to be completed and input it into the trained multi-stage Mamba point cloud completion model. The incomplete point cloud passes through the point cloud local feature encoding unit of Transformer-Mamba to obtain the encoded features. The encoded features are input into the sparse point cloud generation unit to obtain the sparse point cloud. The sparse point cloud is input into the multi-sorting strategy Mamba decoder unit. Each stage of the multi-sorting strategy Mamba decoder adopts 4 sorting strategies, randomly and non-repeatedly assigns the 4 sorting strategies to 4 bidirectional Mamba layers to learn the spatial information at different angles in the sparse point cloud, obtains the decoded features, and the decoded features pass through the point cloud upsampling unit to obtain the predicted complete point cloud.
[0182] Specifically, deploy the trained multi-stage Mamba point cloud completion model, and input the incomplete point cloud to be completed into the trained multi-stage Mamba point cloud completion model, then the predicted complete point cloud can be obtained.
[0183] For further reference Figure 11 , as an implementation of the methods shown in the above figures, an embodiment of a multi-stage Mamba point cloud completion device based on semantic and geometric guidance is provided in this application. This device embodiment corresponds to Figure 1 the method embodiment shown, and this device can be specifically applied to various electronic devices.
[0184] An embodiment of a multi-stage Mamba point cloud completion device based on semantic and geometric guidance is provided in this application embodiment, including:
[0185] A model construction module 1, configured to construct and train a multi-stage Mamba point cloud completion model based on semantic and geometric guidance to obtain a trained multi-stage Mamba point cloud completion model; the multi-stage Mamba point cloud completion model includes a point cloud local feature encoding unit combined with Transformer-Mamba, a sparse point cloud generation unit, a multi-sorting strategy Mamba decoder unit, and a point cloud upsampling unit connected in sequence; the sparse point cloud generation unit includes a skeleton point generation and calibration layer with several layers connected in sequence; the multi-sorting strategy Mamba decoder unit includes multi-sorting strategy Mamba decoders of several stages connected in sequence, and each stage's multi-sorting strategy Mamba decoder includes 4 bidirectional Mamba layers connected in sequence;
[0186] A completion module 2, configured to obtain the incomplete point cloud to be completed and input it into the trained multi-stage Mamba point cloud completion model. The incomplete point cloud passes through the point cloud local feature encoding unit combined with Transformer-Mamba to obtain an encoded feature, and the encoded feature is input into the sparse point cloud generation unit to obtain a sparse point cloud; the sparse point cloud is input into the multi-sorting strategy Mamba decoder unit. Each stage's multi-sorting strategy Mamba decoder adopts 4 sorting strategies, and the 4 sorting strategies are randomly and non-repeatedly assigned to 4 bidirectional Mamba layers to learn the spatial information of different angles in the sparse point cloud to obtain a decoded feature, and the decoded feature passes through the point cloud upsampling unit to obtain the predicted complete point cloud.
[0187] Figure 12 It is a schematic diagram of the hardware structure of the electronic device provided in the embodiment of the present invention. As Figure 12As shown in the figure, the electronic device of this embodiment includes: a processor 1201 and a memory 1202; wherein the memory 1202 is used to store computer-executable instructions; the processor 1201 is used to execute the computer-executable instructions stored in the memory to implement each step executed by the electronic device in the above embodiment. For details, please refer to the relevant descriptions in the foregoing method embodiments.
[0188] Optionally, the memory 1202 can be either independent or integrated with the processor 1201.
[0189] When the memory 1202 is independently provided, the electronic device further includes a bus 1203 for connecting the memory 1202 and the processor 1201.
[0190] An embodiment of the present invention also provides a computer storage medium, in which computer-executable instructions are stored. When the processor 1201 executes the computer-executable instructions, the above method is implemented.
[0191] An embodiment of the present invention also provides a computer program product, including a computer program. When the computer program is executed by the processor 1201, the above method is implemented.
[0192] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in electrical, mechanical or other forms.
[0193] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.
[0194] In addition, in each embodiment of the present invention, the various functional modules can be integrated in one processing unit, or each module can exist physically alone, or two or more modules can be integrated in one unit. The units formed by the above modules can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0195] The integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above-mentioned software functional modules stored in a storage medium include several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor 1201 to execute some steps of the methods according to various embodiments of the present application.
[0196] It should be understood that the above-mentioned processor 1201 may be a central processing unit (CPU for short), or may also be other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), etc. The general-purpose processor may be a microprocessor, or the processor 1201 may also be any conventional processor 1201, etc. The steps of the method disclosed in combination with the invention can be directly implemented by the execution of the hardware processor 1201, or can be implemented by the combination of the hardware and software modules in the processor 1201.
[0197] The memory 1202 may include a high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.
[0198] The bus 1203 may be an Industry Standard Architecture (ISA for short), a Peripheral Component Interconnect (PCI for short) bus, or an Extended Industry Standard Architecture (EISA for short) bus, etc. The bus 1203 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, the bus 1203 in the drawings of the present application is not limited to only one bus 1203 or one type of bus 1203.
[0199] The above-mentioned storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk, or an optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0200] An exemplary storage medium is coupled to the processor 1201, enabling the processor 1201 to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor 1201. The processor 1201 and the storage medium can be located in an Application Specific Integrated Circuit (ASIC). Of course, the processor 1201 and the storage medium can also exist as discrete components in an electronic device or a master control device.
[0201] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0202] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-stage Mamba point cloud completion method based on semantics and geometry guidance, characterized in that: The following steps are involved: A multi-stage Mamba point cloud completion model based on semantics and geometry guidance is constructed and trained to obtain a trained multi-stage Mamba point cloud completion model; the multi-stage Mamba point cloud completion model includes a sequentially connected Transformer-Mamba joint point cloud local feature encoding unit, a sparse point cloud generation unit, a multi-sorting strategy Mamba decoder unit and a point cloud upsampling unit; the sparse point cloud generation unit includes a skeleton point generation calibration layer of several layers connected in sequence; the multi-sorting strategy Mamba decoder unit includes a multi-sorting strategy Mamba decoder of several stages connected in sequence, and the multi-sorting strategy Mamba decoder of each stage includes 4 bidirectional Mamba layers connected in sequence; the bidirectional Mamba unit is improved on the basis of the Transformer encoder, and the Transformer The multi-head self-attention layer and layer normalization in the multi-head self-attention unit in the former are replaced with a bidirectional Mamba layer; the bidirectional Mamba unit includes a layer normalization operation, a first branch, a second branch and a first linear projection layer, the first branch includes a second linear projection layer and a first SiLU activation function layer connected in sequence, the second branch includes a third linear projection layer, a depth-separable convolution layer, a second SiLU activation function layer and a cascade of parallel forward state space equation units and reverse state space equation units, the forward state space equation unit includes a forward state space equation and a layer normalization operation connected in sequence, the reverse state space equation unit includes a reverse state space equation and a layer normalization operation connected in sequence, and the first linear projection layer, the second linear projection layer and the third linear projection layer all use multi-layer perceptrons; The incomplete point cloud to be completed is obtained and input into the trained multi-stage Mamba point cloud completion model. The incomplete point cloud is passed through the Transformer-Mamba joint point cloud local feature encoding unit to obtain the encoding feature, and the encoding feature is input into the sparse point cloud generation unit to obtain the sparse point cloud; each layer of the skeleton point generation calibration layer includes a semantically guided skeleton point cloud generation module and a geometrically guided skeleton point cloud calibration module connected in sequence, and the input feature of the skeleton point generation calibration layer of the current layer is the encoding feature or the input feature of the skeleton point generation calibration layer of the previous layer and the calibrated skeleton point feature obtained by the skeleton point cloud calibration module of the previous layer. The feature is shown in the following formula: Where Concat[·,·] represents the concatenation operation, j = {1,2,...,L}, L represents the total number of skeleton point generated calibration layers; when j = 1, is the encoding feature, Represents the input features of the j-th layer skeleton point generation calibration layer, represents the input features of the j-1th layer skeleton point generation calibration layer, X gj-1 Represents the calibration skeleton point features obtained by the skeleton point cloud calibration module of the j-1th layer; The sparse point cloud is input into the multi-sorting strategy Mamba decoder unit. The multi-sorting strategy Mamba decoder in each stage adopts 4 sorting strategies, which are randomly and non-repetitively assigned to 4 bidirectional Mamba layers to learn spatial information at different angles in the sparse point cloud and obtain decoding features. The decoding features are passed through the point cloud upsampling unit to obtain a predicted complete point cloud.
2. The multi-stage Mamba point cloud completion method based on semantics and geometry guidance according to claim 1, characterized in that: The Transformer-Mamba joint point cloud local feature encoding unit includes a farthest point sampling layer, a K nearest neighbor query layer, an intra-block Transformer encoding unit, an importance score prediction layer and an inter-block Mamba encoding unit connected in sequence, and the intra-block Transformer encoding unit includes N Atten Transformer encoder in each block, N Atten represents the number of Transformer encoders in the block, and the inter-block Mamba encoding unit includes N sequentially connected Mam inter-block Mamba encoders, N Mam represents the number of inter-block Mamba encoders, the importance score prediction layer includes a nonlinear projection layer and a score prediction layer connected in sequence, the nonlinear projection layer is composed of a first convolutional layer, a first batch of normalization layers, a first ReLU activation function layer, a second convolutional layer, and a layer normalization layer cascaded, and the score prediction layer is composed of a third convolutional layer, a second batch of normalization layers, a second ReLU activation function layer, and a fourth convolutional layer cascaded; the calculation process of the Transformer-Mamba joint point cloud local feature encoding unit is as follows: Input the incomplete point cloud into the farthest point sampling layer to select M center points and construct a center point set And through the K nearest neighbor query layer, the K points closest to each center point are selected to construct a local point cloud set to obtain a local point cloud block in, represents a set of real numbers, m∈{1,2,...,M} represents the mth local point cloud block of the downsampled data, M represents the total number of center points, k∈{1,2...,K} represents the kth nearest neighbor point in the local point cloud block, and K represents the total number of nearest neighbor points. Represents the coordinates of the kth nearest neighbor point in the mth local point cloud block, as shown in the following formula: Among them, FPS(·) represents the function corresponding to the farthest point sampling layer, KNN(·) represents the function corresponding to the K nearest neighbor query layer, Indicates the coordinates of incomplete point cloud; The local point cloud block Input the Transformer encoding unit in the block to obtain the local point cloud block features The calculation expression of the Transformer encoder in the i-th block is as follows: Where C represents the feature dimension, i∈{1,2,...,N Atten } represents the Transformer encoder in the ith block, LN(·) represents the layer normalization operation, MultiAtten(·) represents the function corresponding to the multi-head self-attention layer, FF(·) represents the function corresponding to the feedforward network, represents the output features of the multi-head self-attention unit in the Transformer encoder in the i-th block, represents the point cloud block features output by the Transformer encoder in the i-1th block, represents the point cloud block features output by the Transformer encoder in the i-th block. When i=1, for Repeat the above steps until the point cloud block features output by the Transformer encoder in the last block are obtained and used as the local point cloud block features The local point cloud block feature Input to the importance score prediction layer to obtain the importance score of the local point cloud block, as shown in the following formula: Among them, Proj(·) represents the nonlinear projection layer, represents the importance score of the local point cloud block, SP(·) represents the function corresponding to the score prediction layer, Represents the local point cloud block features after projection; The local point cloud block features are sorted according to the importance score of the local point cloud block. Sort and get the sorted local point cloud block features And input the inter-block Mamba encoding unit to model the dependencies between local point cloud blocks. The expression of the inter-block Mamba encoder is as follows: Where i'={1,2,...,N Mam } represents the i'th inter-block Mamba encoder, represents the output feature of the bidirectional Mamba unit in the i'th inter-block Mamba encoder, BiMamba(·) represents the function corresponding to the bidirectional Mamba layer, represents the local point cloud block features enhanced by the i'th inter-block Mamba encoder, represents the local point cloud block feature after being enhanced by the i'-1th inter-block Mamba encoder. When i'=1, is the sorted local point cloud block feature The calculation process of the bidirectional Mamba unit is as follows: Among them, z i' represents the hidden state, DWConv(·) represents the function corresponding to the depthwise separable convolutional layer, MLP(·) represents the function corresponding to the multi-layer perceptron layer, SSM forward (·) represents the forward state space equation, SSM backward (·) denotes the inverse state space equation, σ(·) denotes the SiLU activation function; The enhanced local point cloud block features output by the last inter-block Mamba encoder are used as encoding features 3. The multi-stage Mamba point cloud completion method based on semantics and geometry guidance according to claim 2 is characterized in that: In the skeleton point generation calibration layer of the current layer, the input features of the skeleton point generation calibration layer of the current layer first pass through the semantic-guided skeleton point cloud generation module of the current layer to obtain the first skeleton point prediction features of the current layer and the skeleton point coordinates of the current layer, and the skeleton point coordinates of the current layer, the input point set of the previous layer, the first skeleton point prediction features of the current layer and the input features of the skeleton point generation calibration layer of the current layer are input into the geometric-guided skeleton point cloud calibration module of the current layer to obtain the second skeleton point prediction features of the current layer and the calibrated skeleton point coordinates of the current layer; the calibrated skeleton point coordinates of the current layer are compared with the calibrated skeleton point coordinates of the previous layer. The input point set of the current layer is connected to obtain the input point set of the current layer and input it into the skeleton point generation calibration layer of the next layer, the skeleton point prediction features of the current layer are connected with the calibration skeleton point features of the previous layer to obtain the calibration skeleton point features of the current layer and input them into the skeleton point generation calibration layer of the next layer, and the above steps are repeated until the skeleton point coordinates generated by the skeleton point generation calibration layer of the last layer are obtained and used as sparse point cloud coordinates, and the calibration skeleton point features generated by the skeleton point generation calibration layer of the last layer are obtained and used as sparse point cloud features, and the sparse point cloud coordinates and sparse point cloud features constitute the sparse point cloud.
4. The multi-stage Mamba point cloud completion method based on semantics and geometry guidance according to claim 3 is characterized in that: The input features of the skeleton point generation calibration layer of the current layer are first passed through the semantically guided skeleton point cloud generation module of the current layer to obtain the first skeleton point prediction features of the current layer and the skeleton point coordinates of the current layer, and the skeleton point coordinates of the current layer, the input point set of the previous layer, the first skeleton point prediction features of the current layer and the input features of the skeleton point generation calibration layer of the current layer are input into the geometrically guided skeleton point cloud calibration module of the current layer to obtain the second skeleton point prediction features of the current layer and the calibrated skeleton point coordinates of the current layer, specifically including: In the semantically guided skeleton point cloud generation module of the jth layer, the input features of the calibration layer are first generated for the skeleton points of the jth layer. Perform pooling operation to obtain global features The skeleton points of the jth layer are used to generate the input features of the calibration layer. and global features The feature difference between them is input to the fourth linear projection layer for linear projection, and N l Skeleton point features of the current layer Its expression is as follows: Among them, Pool j (·) represents the function corresponding to the pooling layer in the semantically guided skeleton point cloud generation module of the jth layer, the pooling layer in the semantically guided skeleton point cloud generation module of the first layer is an importance-aware pooling layer, and the importance-aware pooling layer is used to re-weight and sum the importance scores predicted by the importance score prediction layer, and the pooling layers in the semantically guided skeleton point cloud generation modules of the remaining layers are maximum pooling layers; Generate the input features of the calibration layer from the skeleton points of the jth layer and the skeleton point features of the current layer After connection, the farthest point sampling and the closest point sampling of the features are sorted, and the skeleton point features are modeled in two Mamba units. and input features The semantic correlation between them is used to obtain the first skeleton point prediction feature of the jth layer with semantic interaction information. Its expression is as follows: Among them, F-FPS(·) represents the farthest feature point sampling, F-NPS(·) represents the closest feature point sampling, Concat[·,·] represents the connection operation, Dual-Mamba[·,·] represents the two-way Mamba unit, and the calculation process of the two-way Mamba unit is as follows: Among them, RMSNorm(·) is the function corresponding to the root mean square normalization layer, Represents the input features of the j-th layer skeleton point generation calibration layer and the skeleton point features of the current layer Input features after sampling and sorting the farthest feature points after connection Represents the input features of the j-th layer skeleton point generation calibration layer and the skeleton point features of the current layer Input features after connection and feature nearest point sampling and sorting SSM stands for state space equation; The first skeleton point prediction feature of the jth layer Input to the fifth linear projection layer, and predict N l The coordinates of the skeleton points at the jth layer Its expression is as follows: Wherein, the fourth linear projection layer and the fifth linear projection layer adopt a multi-layer perceptron; In the j-th layer geometry-guided skeleton point cloud calibration module, the j-th layer skeleton point coordinates and the input point set of the j-1th layer After connection, the farthest point sampling and the nearest point sampling are sorted respectively, and the first skeleton point of the jth layer is predicted according to the sorting results. And the skeleton points of the jth layer generate the input features of the calibration layer The connection results are sorted and input into two Mamba units to model the geometric correlation between the two, and the second skeleton point prediction feature of the jth layer containing spatial geometric information is obtained. Its expression is as follows: Among them, FPS(·) represents the farthest point sampling, NPS(·) represents the closest point sampling, Forder(·) and Norder(·) represent the sorting by farthest point sampling and the closest point sampling, respectively. When j=1, the input point set The center point set Secondly, the second skeleton point prediction feature of the jth layer With global features After adding, input the sixth linear projection layer to get the point offset The skeleton point coordinates of the jth layer obtained in the semantically guided skeleton point cloud generation unit are Add together to get the calibrated skeleton point coordinates of the jth layer Its expression is as follows: Wherein, the sixth linear projection layer adopts a multi-layer perceptron.
5. The multi-stage Mamba point cloud completion method based on semantics and geometry guidance according to claim 3, characterized in that: The calculation expression of the decoding feature is: Order=RA(Z,Trans-Z,Hilbert,Trans-Hilbert); Among them, RA(·) represents random assignment, Order represents the sorting method, Z, Trans-Z, Hilbert and Trans-Hilbert represent the Z-curve sorting strategy, the reverse Z-curve sorting strategy, the Hilbert sorting strategy and the reverse Hilbert sorting strategy respectively. represents the function corresponding to the bidirectional Mamba layer executed in the mth stage according to the nth sorting method, {m=1,2,...,N De }, {n=1,2,3,4}, N De represents the total number of stages in the multi-order strategy Mamba decoder unit, represents sparse point cloud features, represents the decoding feature; The decoded features are passed through the point cloud upsampling unit to obtain a predicted complete point cloud, which specifically includes: Using decoding features Compute the first affine parameter and the second affine parameter And guide the two-dimensional grid with affine function Deformation, three-dimensional offset obtained by three-dimensional deformation At the same time, each point coordinate in the sparse point cloud is copied O times and added to the three-dimensional offset to finally get the predicted complete point cloud. Its expression is as follows: α=MLP(MaxPool(F De )); Among them, MaxPool represents the maximum pooling operation, MLP represents the function corresponding to the multi-layer perceptron, and μ and σ represent the two-dimensional grid G 2D The mean and standard deviation of , Dup(·) represents the copy operation, Represents the sparse point cloud coordinates, n c ∈{1,2,...,(M+N l ×L)} represents the nth c A sparse point, Indicates the nth c The decoding features of sparse points, ReLU(·) represents the ReLU activation function, Indicates the nth c The second affine parameter corresponding to the sparse points, Indicates the nth c The point coordinates of sparse points, Indicates that the nth c The upsampled points generated by the sparse points.
6. The multi-stage Mamba point cloud completion method based on semantics and geometry guidance according to claim 2, characterized in that: The total loss function used in the training process of the multi-stage Mamba point cloud completion model is the sum of the importance loss and the chamfer loss. The calculation process of the importance loss is shown in the following formula: Wherein, Pool(·) represents the function corresponding to the pooling layer. The pooling layer here is the importance-aware pooling layer, which is used to re-weight and sum the importance scores predicted by the importance score prediction layer. represents the global encoding feature, represents the global encoding feature after projection, represents the true importance score calculated by cosine similarity, L Smooth represents L1 smoothing loss, Indicates loss of importance; The calculation process of the chamfer loss is shown in the following formula: in, represents the chamfer loss, Represents the true complete point cloud, N fine Represents the number of points in the predicted complete point cloud, N GT Represents the number of points of the true complete point cloud, represents the predicted complete point cloud, x and y represent one point in the predicted complete point cloud and one point in the real complete point cloud respectively. Represents the square of the Euclidean distance.
7. A multi-stage Mamba point cloud completion device based on semantics and geometry guidance, characterized in that: include: The model construction module is configured to construct and train a multi-stage Mamba point cloud completion model based on semantics and geometry guidance to obtain a trained multi-stage Mamba point cloud completion model; the multi-stage Mamba point cloud completion model includes a sequentially connected Transformer-Mamba joint point cloud local feature encoding unit, a sparse point cloud generation unit, a multi-sorting strategy Mamba decoder unit and a point cloud upsampling unit; the sparse point cloud generation unit includes a skeleton point generation calibration layer of several layers connected in sequence; the multi-sorting strategy Mamba decoder unit includes a multi-sorting strategy Mamba decoder of several stages connected in sequence, and the multi-sorting strategy Mamba decoder of each stage includes 4 bidirectional Mamba layers connected in sequence; the bidirectional Mamba unit is improved on the basis of the Transformer encoder, The multi-head self-attention layer and layer normalization in the multi-head self-attention unit in the Transformer are replaced with a bidirectional Mamba layer; the bidirectional Mamba unit includes a layer normalization operation, a first branch, a second branch and a first linear projection layer, the first branch includes a second linear projection layer and a first SiLU activation function layer connected in sequence, the second branch includes a third linear projection layer, a depth-separable convolution layer, a second SiLU activation function layer, and a cascade of parallel forward state space equation units and reverse state space equation units, the forward state space equation unit includes a forward state space equation and a layer normalization operation connected in sequence, the reverse state space equation unit includes a reverse state space equation and a layer normalization operation connected in sequence, and the first linear projection layer, the second linear projection layer and the third linear projection layer all use multi-layer perceptrons; The completion module is configured to obtain an incomplete point cloud to be completed and input it into the trained multi-stage Mamba point cloud completion model, the incomplete point cloud passes through the Transformer-Mamba joint point cloud local feature encoding unit to obtain a coding feature, and the coding feature is input into the sparse point cloud generation unit to obtain a sparse point cloud; each layer of the skeleton point generation calibration layer includes a semantically guided skeleton point cloud generation module and a geometrically guided skeleton point cloud calibration module connected in sequence, and the input feature of the skeleton point generation calibration layer of the current layer is the coding feature or the input feature of the skeleton point generation calibration layer of the previous layer and the calibrated skeleton point feature obtained by the skeleton point cloud calibration module of the previous layer. The feature is shown in the following formula: Where Concat[·,·] represents the concatenation operation, j = {1,2,...,L}, L represents the total number of skeleton point generated calibration layers; when j = 1, is the encoding feature, Represents the input features of the j-th layer skeleton point generation calibration layer, represents the input features of the j-1th layer skeleton point generation calibration layer, X gj-1 Represents the calibration skeleton point features obtained by the skeleton point cloud calibration module of the j-1th layer; The sparse point cloud is input into the multi-sorting strategy Mamba decoder unit. The multi-sorting strategy Mamba decoder in each stage adopts 4 sorting strategies, which are randomly and non-repetitively assigned to 4 bidirectional Mamba layers to learn spatial information at different angles in the sparse point cloud and obtain decoding features. The decoding features are passed through the point cloud upsampling unit to obtain a predicted complete point cloud.
8. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Point cloud super-resolution method and system based on multi-stage deep learning, and electronic equipment
CN116342387A
Point cloud matching method and system based on derivative-free optimization
CN118314180A