Three-dimensional point cloud reconstruction method and application thereof in target detection
By employing an encoder-decoder architecture that incorporates multi-scale feature extraction and dynamic neighbor selection, the problem of insufficient multi-scale feature extraction and neighbor selection in point cloud reconstruction is solved, achieving more accurate point cloud completion and object detection.
Patent Information
- Application Number
- CN202511452781.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies for point cloud completion suffer from limited multi-scale feature extraction capabilities, a single neighbor selection method, and a lack of geometric structure rationality constraints, resulting in poor point cloud reconstruction performance.
An encoder-decoder architecture employing a multi-scale feature extraction module, dynamic neighbor selection, and geometric structure constraints generates multi-scale features through farthest point sampling, kNN, and KAN modules, and combines semantic-geometric scoring and dynamic queries to generate a complete point cloud.
It significantly enhances the multi-scale representation capability and generalization performance of point clouds, achieving more accurate point cloud completion and target detection.
Smart Images

Figure CN120953516A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a three-dimensional point cloud reconstruction method and its application in target detection. Background Technology
[0002] With the continuous development of 3D sensing technologies such as RGB-D sensors and LiDAR, 3D point clouds, as an important data format, have been widely used in computer vision, robot navigation, and building surveying due to their efficient 3D shape representation capabilities and low storage requirements. Point cloud data can effectively convey the spatial structure information of objects; however, due to unavoidable factors such as limited sensor resolution, self-occlusion, light reflection, and environmental noise, the acquired point clouds are often incomplete, sparse, or missing. These problems severely restrict the performance of subsequent 3D reconstruction, object detection, and semantic segmentation tasks, thus point cloud completion technology has become a core research problem in this field.
[0003] Early research on point cloud completion largely relied on voxelized 3D convolutional networks, processing point clouds by transforming them into regular 3D mesh voxels. However, as spatial resolution increases, the computational cost of such methods grows cubically, making it difficult to efficiently handle large-scale, high-precision point clouds. With the introduction of PointNet and PointNet++, deep learning methods that directly use the 3D coordinates of the point cloud as input have become mainstream. Encoder-decoder architectures are widely used for point cloud completion tasks, generating complete point clouds through global feature vectors. However, the single global features generated during the encoding stage of these methods often fail to fully capture local details, limiting the completion results.
[0004] To overcome the disordered and unstructured nature of point clouds, learning the local structure and long-range semantic correlations of point clouds has become crucial for improving the quality of point cloud completion. In recent years, Transformer-based models have demonstrated powerful global relation modeling capabilities in point cloud completion. For example, the PoinTr model employs an encoder-decoder architecture, treating point cloud completion as a set-to-set transformation task. It models pairwise interactions between points through a self-attention mechanism, explicitly models 3D spatial structure information using geometrically perceptual transformer blocks, and achieves coarse-to-fine completion of missing parts through dynamic query generation and multi-scale point cloud generation modules. Based on this, AdaPoinTr further proposes an adaptive query generation mechanism and an auxiliary denoising task, effectively improving training efficiency and completion accuracy, and significantly enhancing point cloud completion performance in real-world environments.
[0005] Despite the significant progress made in the Transformer architecture described above, existing technologies still have the following shortcomings: First, the multi-scale feature extraction capability is limited. Traditional methods use multilayer perceptrons (MLPs) with fixed parameters to uniformly map point cloud features at all scales, which is difficult to adapt to the feature representation requirements under different spatial resolutions, resulting in the inability to fully capture complex geometric structures and local details.
[0006] Secondly, the neighbor selection method is too simplistic and fails to fully utilize semantic information. Existing k-nearest neighbor (kNN) methods based on geometric distance only consider single-hop neighbors, ignoring the semantic features of the point cloud, which limits the expressive power and coverage of the neighbor set and affects the discriminative ability of local features.
[0007] Third, there is a lack of constraints on the rationality of geometric structure. In tasks such as predicting 3D polygonal buildings from point clouds, there is a lack of explicit geometric constraints on the smoothness of side lengths, the rationality of angles, and the closure of polygons, which leads to unstable structure and insufficient accuracy in the prediction results. Summary of the Invention
[0008] In view of the above-mentioned deficiencies of the prior art, the present invention provides a three-dimensional point cloud reconstruction method and its application in target detection, so as to solve the technical problem of poor point cloud completion effect in the prior art.
[0009] To achieve the above and other related objectives, this invention provides a three-dimensional point cloud reconstruction method, comprising: acquiring an incomplete original point cloud; inputting the original point cloud into a trained point cloud completion model to obtain a complete point cloud, wherein the point cloud completion model comprises: a feature extraction module for extracting multi-scale features of the original point cloud; a feature generation module for constructing a set of center points and calculating a point proxy for each center point in the set based on the multi-scale features; an encoder for modeling the global semantic dependencies and local geometric relationships of the point proxies to obtain global features of each center point; a query generator for generating multiple dynamic queries based on the global features of all center points, each dynamic query containing a predicted center point and its query vector; a decoder for obtaining a point proxy for each predicted center point based on the global features of all center points and the multiple dynamic queries; and a reconstruction head for generating a predicted missing point cloud based on each predicted center point and its point proxy, and fusing the predicted missing point cloud with the original point cloud to obtain a complete point cloud.
[0010] In one embodiment of the present invention, extracting multi-scale features of the original point cloud includes: calculating the center points of the original point cloud at L different resolutions using a farthest point sampling algorithm, thereby obtaining a set of center points at different resolutions. Where, l∈{1,2,…,L}, N inLet F be the total number of points in the original point cloud; use kNN to collect the neighbors of each center point at each resolution to obtain the local point patch corresponding to that point; use a KAN module with L independent parameters to perform nonlinear mapping on the local point patches corresponding to all center points at different resolutions to obtain the feature representation F at different resolutions. (l) ; Fusing feature representations at different resolutions (l) The fused multi-scale features are obtained. .
[0011] In one embodiment of the present invention, feature representations at different resolutions are fused using the following formula: , In the formula, , l∈{2,3,…,L}, Up represents the feature interpolation upsampling operation based on the distance between points, Concat represents feature concatenation, Conv is the convolution operation, and σ is the non-linear activation function.
[0012] In one embodiment of the present invention, constructing a center point set and calculating the point proxy for each center point in the center point set based on the multi-scale features includes: calculating N center points of the original point cloud using a farthest point sampling algorithm to obtain the center point set. ; center point C i The multi-scale features of k neighboring points in the original point cloud are aggregated to obtain the local semantic information g corresponding to the center point. i ; center point C i By mapping to a high-dimensional space using the function Φ1, we obtain the first position embedding vector Φ1(c i ); the local semantic information g i and the first position embedding vector Φ1(c i Adding them together, we get the center point C. i Point agent.
[0013] In one embodiment of the present invention, modeling the global semantic dependencies and local geometric relationships of the point proxies to obtain the global features of each center point includes: calculating the center point C from the set of center points. i Neighbor sets that take into account both semantic and geometric information Based on center point C i and its neighbor set N i The point proxy obtains the center point C. i Global features.
[0014] In one embodiment of the present invention, the center point C is calculated. i The neighbor set N that takes into account both semantic and geometric information i This includes: using kNN to collect center points C iFind the neighbors of the point to obtain the initial set corresponding to that point. Perform h jumps for expansion, and merge and deduplicate all the neighbor points obtained from these jumps to obtain the center point C. i The corresponding candidate point set; calculate the center point C. i The semantic-geometric score of each point in the candidate point set and its corresponding semantic-geometric score; based on the semantic-geometric score, the center point C is selected from the candidate point set. i The neighbor set N is obtained by selecting a predetermined number of neighbor points with the highest scores. i .
[0015] In one embodiment of the present invention, the center point C is calculated. i The semantic-geometric score of each point in the candidate point set and its corresponding score includes: based on the center point C i The coordinates of each point in the candidate point set and the corresponding points are used to calculate the Euclidean distance between them to obtain the geometric affinity; based on the center point C i The cosine similarity between the feature vectors of each point in the candidate point set and the geometric affinity is calculated to obtain the semantic similarity. Based on the geometric affinity and the semantic similarity, the semantic-geometric score is obtained.
[0016] In one embodiment of the present invention, based on the center point C i and its neighbor set N i The point proxy obtains the center point C. i The global features include: based on the center point C i and its neighbor set N i For each point in the matrix, perform point proxying and calculate the feature difference between the two; then, for the center point C... i The point proxy is concatenated with the feature difference to obtain a combined vector; a linear transformation is applied to each of the combined vectors to obtain the center point C. i Interaction characteristics with its neighbors; using max pooling operation, for center point C i The center point C is obtained by processing the interaction characteristics with all its neighbors. i Local features; center point C i The point proxy and its local features are fused to obtain the center point C. i Global features.
[0017] In one embodiment of the present invention, multiple dynamic queries are generated based on the global features of all central points, including: based on the global features V of all central points... i The M predicted center points are obtained using the following formula: , In the formula, For the j-th prediction center point, MLP posFor a multilayer perceptron used to predict the center points of missing regions, Pool is globally pooled; a query vector for each predicted center point is generated according to the following formula: , In the formula, Let be the query vector for the j-th dynamic query, MLP be a multilayer perceptron, and Φ2 be used to calculate the predicted center point. The second position embedding vector; Used to extract features surrounding the predicted center point from global features of all center points. Contextual information of the local region.
[0018] In one embodiment of the present invention, generating a predicted missing point cloud based on each predicted center point and its point proxy includes: generating a point cloud block corresponding to each predicted center point according to the following formula: , In the formula, Q i For the i-th point cloud block, H i For the i-th predicted center point, G is a predefined regular grid, and f is a folding operation mapping function; merge the point cloud blocks corresponding to all predicted center points to obtain the predicted missing point cloud.
[0019] To achieve the above and other related objectives, the present invention also provides a target detection method, comprising: obtaining a complete point cloud of the target using the above-described three-dimensional point cloud reconstruction method; and obtaining the volume of the target based on the complete point cloud of the target.
[0020] The beneficial effects of this invention are as follows: This invention proposes a three-dimensional point cloud reconstruction method and its application in target detection. This method achieves automatic completion of incomplete point clouds by constructing a point cloud completion model. When performing feature extraction and generation, it takes into account both global consistency and local detail representation, significantly enhancing the multi-scale representation ability and generalization performance of point clouds. Furthermore, through an encoder-decoder structure, it quickly generates features for predicting the center point and performs prediction of missing point clouds and fusion of all point clouds based on the reconstruction head. Through this method, point cloud completion prediction can be performed more accurately. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. The accompanying drawings are incorporated in and constitute a part of this specification, illustrating embodiments consistent with this application, and are used together with the description to explain the principles of this application. Obviously, the drawings described below are merely some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of a point cloud completion method provided in an embodiment of the present invention; Figure 2 This is an architecture diagram of a point cloud completion model provided in an embodiment of the present invention; Figure 3 This is a flowchart of multi-scale feature extraction provided in an embodiment of the present invention; Figure 4 A flowchart of point proxy generation provided in an embodiment of the present invention; Figure 5 This is a flowchart of the encoder processing according to an embodiment of the present invention; Figure 6 A flowchart of multi-hop nearest neighbor selection provided in an embodiment of the present invention; Figure 7 A flowchart of scoring calculation provided in an embodiment of the present invention; Figure 8 A flowchart for generating global features is provided in one embodiment of the present invention; Figure 9 This is a flowchart for generating dynamic queries provided in an embodiment of the present invention; Figure 10 This is a flowchart of missing point cloud prediction provided in an embodiment of the present invention.
[0023] Figure labeling: 201, Feature extraction module; 202, Feature generation module; 203, Encoder; 204, Query generator; 205, Decoder; 206, Reconstruction head. Detailed Implementation
[0024] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. It should be noted that, unless otherwise specified, the following embodiments and features can be combined with each other. In addition to the specific methods, equipment, and materials used in the embodiments, based on the knowledge of the prior art and the description of the present invention by those skilled in the art, any prior art methods, equipment, and materials similar to or equivalent to the methods, equipment, and materials in the embodiments of the present invention can be used to implement the present invention.
[0025] It should be understood that the terminology used in the embodiments of this invention is for describing specific implementations and not for limiting the scope of protection of this invention. Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art.
[0026] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In some embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.
[0027] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functions, and operations that may be implemented in the methods and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0028] Please see Figure 1 , Figure 1 An embodiment of the present invention provides a method for reconstructing three-dimensional point clouds, comprising: acquiring an incomplete original point cloud; and inputting the original point cloud into a trained point cloud completion model to obtain a complete point cloud.
[0029] Understandably, the trained point cloud completion model can be obtained by following these steps: (1) First, construct a dataset containing many samples. Each sample is a point cloud pair consisting of an incomplete point cloud and its corresponding complete point cloud. To reduce annotation work, part of the complete point cloud is usually removed to obtain an incomplete point cloud. Depending on the area removed, a complete point cloud can yield multiple incomplete point clouds. These incomplete point clouds and their corresponding complete point clouds constitute a sample. (2) Construct a point cloud completion model. The structure of this model will be described in detail later, and it is the focus of this invention. (3) Train the point cloud completion model using the constructed dataset. Once training is complete, the trained point cloud completion model can be obtained.
[0030] In a specific embodiment of the present invention, the expression for the loss function L during the above training process is as follows: In the formula, λ cd , λ emd , λ geom , λ reg These are the weight hyperparameters for each item; θ represents the network parameters; the four parts are responsible for global point cloud alignment, detailed point matching, geometry-structure consistency, and regularization, respectively.
[0031] Global point cloud alignment loss L CD To ensure that the predicted point cloud and the actual point cloud are as consistent as possible in overall distribution, the calculation formula is as follows: , In the formula, P pred and P gt Let L represent the predicted point and the ground truth point, respectively. Global point cloud alignment loss L... CD For each point in the predicted point cloud, find the nearest point in the real point cloud and calculate the squared distance; for each point in the real point cloud, find the nearest point in the predicted point cloud and calculate the squared distance; then average the distances and sum them.
[0032] Detailed matching loss L EMD To ensure a globally optimal match between predicted and actual points, rather than a local nearest neighbor, its calculation formula is as follows: , In the formula, ϕ is a bijective matching function that maps predicted points to real points one-to-one; the optimal match is found by minimizing the total distance.
[0033] Geometric-structural consistency loss L geom_struct This involves imposing constraints on the local geometric properties of the point cloud, ensuring that the completed point cloud not only approximates the real points in location but also preserves their geometric structure. The expression is as follows: , This loss is also composed of multiple sub-items, and the corresponding coefficients are the weight hyperparameters of each sub-item.
[0034] Local side length consistency L edge Used to keep the local point distance distribution in the predicted point cloud consistent with the true point distance: , In the formula, μ pred (x) represents the average neighborhood distance of the predicted point x; μ gt (π(x)) is the average neighborhood distance of the corresponding real point; π(x) is the nearest neighbor of the predicted point x in the real point cloud.
[0035] Local plane / curvature consistency L plane To ensure the predicted points "fit" to the local geometry of the true points, the planar / curvature structure of the point cloud on the local surface must be maintained. , In the formula, R j G is a local neighborhood of the real point cloud; μ_(G_j) is the center point of this neighborhood; n_(G_j) is the principal plane normal vector calculated by PCA.
[0036] Laplacian smoothing regularization L lap Suppress noise points, ensure surface smoothness, maintain the local smoothness of the predicted point cloud, and avoid isolated noise points: , In the formula, P represents the predicted point cloud, |P| is the total number of points in point cloud P, the summation symbol indicates that the loss is calculated by traversing each point in P, and N(x) is the set of neighboring points of point x.
[0037] Please see Figure 2In a specific embodiment of the present invention, the point cloud completion model includes a feature extraction module 201, a feature generation module 202, an encoder 203, a query generator 204, a decoder 205, and a reconstruction head 206. Specifically, the feature extraction module 201 extracts multi-scale features from the original point cloud; the feature generation module 202 constructs a set of center points and calculates a point proxy for each center point in the set based on the multi-scale features; the encoder 203 models the global semantic dependencies and local geometric relationships of the point proxies to obtain the global features of each center point; the query generator 204 generates multiple dynamic queries based on the global features of all center points, each dynamic query containing a predicted center point and its query vector; the decoder 205 obtains a point proxy for each predicted center point based on the global features of all center points and the multiple dynamic queries; and the reconstruction head 206 generates a predicted missing point cloud based on each predicted center point and its point proxy, and fuses the predicted missing point cloud with the original point cloud to obtain a complete point cloud.
[0038] Please see Figure 3 Extracting multi-scale features from the original point cloud, including steps S301 to S304.
[0039] Step 301: Calculate the center points of the original point cloud at L different resolutions using the farthest point sampling algorithm (FPS), thus obtaining the center point sets at different resolutions. , l∈{1,2,…,L}, N in This represents the total number of points in the original point cloud. Where P... (1) For the lowest resolution (corresponding to large-scale features), P (L) This represents the lowest resolution (corresponding to small-scale features).
[0040] Step S302: Use kNN to collect the neighbors of each center point at each resolution to obtain the local point patch corresponding to that point. kNN is an abbreviation for k-Nearest Neighbors, and its purpose is to find the k nearest points to a given center point in three-dimensional space.
[0041] Step S303: Using a KAN module with L independent parameters, perform nonlinear mapping on the local point blocks corresponding to all center points at different resolutions to obtain the feature representation F at different resolutions. (l) It can be expressed by the formula: , In the formula, This represents the l-th independent KAN module (Kolmogorov–Arnold Network). This function performs non-linear mapping and is an emerging neural network architecture. It is important to note that the parameters of the KAN module are different at different resolutions; C is the feature dimension.
[0042] Step S304: Fuse feature representations F at different resolutions (l) The fused multi-scale features are obtained. .
[0043] In a specific embodiment of the present invention, feature representations at different resolutions are fused using the following formula: , In the formula, Let l ∈ {2, 3, ..., L}, Up represent a feature interpolation upsampling operation based on distance between points, Concat represent feature concatenation, Conv represent a convolution operation, and σ represent a non-linear activation function. In this step, features at the low-resolution scale are progressively mapped to the high-resolution scale through interpolation upsampling, and then concatenated and fused with features at the high-resolution scale. After the final fusion, multi-scale features are obtained. .
[0044] Please see Figure 4 In a specific embodiment of the present invention, a set of center points is constructed, and a point proxy for each center point in the set of center points is calculated based on multi-scale features, including steps S401 to S404.
[0045] Step S401: Calculate the N center points of the original point cloud using the farthest point sampling algorithm to obtain the center point set. This step still uses the FPS algorithm to calculate N center points, where N is a preset number. Since the incomplete original point cloud may contain a large number of points, processing each point individually would significantly increase computation time. The N center points extracted in this step are equivalent to removing some representative points for subsequent encoding and decoding. The predicted center points generated later are also representative points, so subsequent steps will need to generate a missing point cloud (multiple points) based on the predicted center points.
[0046] Step S402: Place the center point C i The multi-scale features of the k neighboring points in the original point cloud are aggregated to obtain the local semantic information g corresponding to the center point. i Here, k is also a preset value; this step is to obtain local semantic information.
[0047] Step S403: Place the center point C iBy mapping to a high-dimensional space using the function Φ1, we obtain the first position embedding vector Φ1(c i A point cloud is an unordered set of points. For neural networks (especially the subsequent decoder 205), the inputs {c1,c2,c3} and {c2,c3,c1} are completely identical because the order doesn't matter. However, the spatial location information of the points is crucial; directly shuffling the order will result in the loss of all spatial relationships. Φ1(ci) represents each unique center point C. i Generating a unique high-dimensional feature vector is equivalent to assigning a unique "address code" to each point. Even if the input order of the points is shuffled, the positional information of each point is still preserved and fixed through this embedding vector, enabling the network to stably understand the spatial structure.
[0048] Step S403: Transfer local semantic information g i and the first position embedding vector Φ1(c i Adding them together, we get the center point C. i Point agent P i After processing, the point proxy sequence can be obtained from N center points. (Length N, each is C-dimensional).
[0049] Please see Figure 5 In a specific embodiment of the present invention, the global semantic dependencies and local geometric relationships of the modeling point proxies are used to obtain the global features of each center point, including steps S501 and S502.
[0050] Step S501: Calculate the center point C from the center point set C. i Neighbor sets that take into account both semantic and geometric information .
[0051] In a specific embodiment of the present invention, the center point C is calculated. i The neighbor set N that takes into account both semantic and geometric information i This includes steps S601 to S604.
[0052] Step S601: Collect center points C using kNN. i Find the neighbors of the point to obtain the initial set corresponding to that point. .
[0053] Step S602: Perform h jumps for expansion, and merge and deduplicate all the neighbor points obtained from these jumps to obtain the center point C. i The corresponding set of candidate points. Jump expansion can be expressed by the formula: , In the formula, h is the maximum number of hops, and the final set of candidate points is obtained. (Duplicates need to be removed).
[0054] Step S603: Calculate the center point C i The semantic-geometric score of each point in the corresponding candidate point set.
[0055] Please see Figure 7 In a specific embodiment of the present invention, step S603 includes steps S701 to S703.
[0056] Step S701: Based on center point C i Given the coordinates of each point in the candidate point set and the corresponding points, calculate the Euclidean distance between them to obtain the geometric affinity, which can be expressed by the formula: , In the formula, g ij Center point C i The geometric affinity between the j-th point in the corresponding candidate point set and the j-th point. It is a hyperparameter representing a distance scaling factor or bandwidth parameter that controls the distance at point C. i With point C j The rate at which the "influence" of distance decreases as distance increases.
[0057] Step S702, based on center point C i and the feature vector of each point in the corresponding candidate point set (i.e., the point proxy P) i ), calculate the cosine similarity between the two to obtain the semantic similarity, which can be expressed by the formula: , In the formula, s ij Center point C i The semantic similarity between the j-th point in the corresponding candidate point set and s i Center point C i Point agent.
[0058] Step S703: Obtain the semantic-geometric score based on geometric affinity and semantic similarity.
[0059] In a specific embodiment of the present invention, geometric affinity and semantic similarity can be fused according to the following formula: , In the formula, α is the preset weight coefficient, ReLU is the activation function, and aij is the unnormalized semantic-geometric score. Then, softmax normalization is performed according to j: In the formula, w ij This is the normalized semantic-geometric score.
[0060] Step S604: Based on the semantic-geometric score, select the center point C from the candidate point set. i The neighbor set N is obtained by selecting a predetermined number of neighboring points with the highest scores. i The final neighbor set can be selected by choosing the top-K neighbors (preset number) or by using a weight threshold, selecting the neighbors with the highest weight or whose weight exceeds the threshold, for subsequent feature extraction and task processing. This neighbor set takes into account both spatial proximity and semantic relevance, significantly improving the richness and discriminative power of local feature representation in point clouds. Through this method, neighbor selection not only breaks through the limitations of traditional single-hop neighbors but also integrates semantic information, effectively improving the representation quality of point cloud features and the robustness of the model. It has significant performance improvement effects on various downstream tasks such as point cloud classification, segmentation, and reconstruction.
[0061] Step S502, based on center point C i and its neighbor set N i The point proxy obtains the center point C. i Global features.
[0062] Please see Figure 8 In a specific embodiment of the present invention, step S502 includes steps S801 to S805.
[0063] Step S801: Based on center point C i and its neighbor set N i For each point in the model, perform point proxying and calculate the feature difference Δ between the two. ij It can be expressed by the formula: , In the formula, P ij For the neighbor set N i Point agent of the j-th point, P i Center point C i Point agent.
[0064] Step S802: Set the center point C i The point proxy and feature difference are concatenated to obtain the combined vector X. ij It can be expressed by the formula: , In this way, each neighbor is connected to the central point C. i A connection has been established.
[0065] Step S803: Apply a linear transformation to each combined vector to obtain the center point C. i Interaction characteristics with its neighbors It can be expressed by the formula: , In the formula, W and b are both learnable parameters, representing the transformed feature set. Represents the center point C i With neighbor set N i The local geometric interaction representation between each point in the graph.
[0066] Step S804: Perform max pooling operation on center point C. i The center point C is obtained by processing the interaction characteristics with all its neighbors. i Local features It can be expressed by the formula: , In the formula, k is the neighbor set N i The total number of midpoints, where M is the maximum pooling operation.
[0067] Step S805: Place the center point C i The point proxy and its local features are fused to obtain the center point C. i The global features. In this step, the local geometric features output in step S804 are... The input feature of step S801 (for the first layer of encoder 203, this input feature is the center point C) is the input feature of step S801. i Point agent P i For subsequent layers, this input feature is the global feature output from the previous layer. The center point C is obtained by fusing the components using the Transformer module (e.g., through residual connections). i Global features at layer l The above steps S801~S805 are executed L times in total, and the global features output by the Lth layer are... The final global feature V, centered at point Ci. i That is, the output of encoder 203.
[0068] Please see Figure 9 In a specific embodiment of the present invention, multiple dynamic queries are generated based on the global features of all center points, including steps S901 and S902.
[0069] Step S901: Based on the global features V of all center points i The M predicted center points are obtained using the following formula: , In the formula, For the j-th prediction center point, MLP pos For a multilayer perceptron used to predict the center point of a missing region, Pool is a global pooling method.
[0070] Step S902: Generate the query vector for each predicted center point according to the following formula: , In the formula, Let be the query vector for the j-th dynamic query, MLP be a multilayer perceptron, and Φ2 be used to calculate the predicted center point. The second position embedding vector; Used to extract features surrounding the predicted center point from global features of all center points. Contextual information of the local region.
[0071] By using global pooling and MLP direct regression to predict the center point in step S901, the most likely core location of the missing region can be efficiently inferred from complete known context information, ensuring the global rationality and accuracy of the predicted center point. Secondly, in step S902, when generating the query vector, the absolute position information of the predicted center point (obtained through position embedding Φ2) is creatively combined with its local context information (extracted from the features of encoder 203 through the LocalFeatures function). This design ensures that each generated query vector simultaneously contains the key information of "where it is" and "what's around it," providing decoder 205 with a highly informative initial query. This not only greatly enhances the targeting and efficiency of the decoding process but also helps the model generate detailed missing content that seamlessly connects with the known parts, which is crucial for achieving high-precision completion.
[0072] Please see Figure 10 In a specific embodiment of the present invention, a predicted missing point cloud is generated based on each predicted center point and its point proxy, including steps S1001 and S1002.
[0073] Step S1001: Based on each predicted center point and its point proxy, generate the point cloud block corresponding to each predicted center point according to the following formula: , In the formula, Q i For the i-th point cloud block, H i Let G be a point proxy for the i-th predicted center point, G be a predefined regular grid, and f be a folding operation mapping function. The folding operation is usually implemented using a multilayer perceptron (MLP), which folds high-dimensional features into three-dimensional coordinate increments, and then superimposes the center point coordinates to generate a local point cloud. This method can generate detailed local geometry around each center point, thereby recovering missing regions and preserving the overall point cloud shape.
[0074] Step S1002: Merge the point cloud blocks corresponding to all predicted center points to obtain the predicted missing point cloud. After obtaining the predicted missing point cloud, it can be fused with the incomplete original point cloud to obtain the final complete point cloud.
[0075] This invention also provides a target detection method, comprising: obtaining a complete point cloud of the target using the aforementioned three-dimensional point cloud reconstruction method; and obtaining the target's volume based on the complete point cloud. Once the complete point cloud of the target is obtained, the target's volume can be accurately calculated based on it. Furthermore, it can accurately describe the target's size information, shape information, etc., all of which are applications based on the three-dimensional point cloud reconstruction method.
[0076] It should be noted that the steps of the various methods described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they contain the same logical relationship, they are all within the scope of protection of this application. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, but without changing the core design of the algorithm and process, are also within the scope of protection of this patent.
[0077] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.
Claims
1. A method for reconstructing three-dimensional point clouds, characterized in that, include: Obtain incomplete raw point clouds; The original point cloud is input into a trained point cloud completion model to obtain a complete point cloud, wherein the point cloud completion model includes: The feature extraction module is used to extract multi-scale features from the original point cloud; The feature generation module is used to construct a set of center points and calculate the point proxy for each center point in the set of center points based on the multi-scale features. An encoder is used to model the global semantic dependencies and local geometric relationships of the point proxies to obtain the global features of each center point; A query generator is used to generate multiple dynamic queries based on the global features of all centroids, each of which contains a predicted centroid and its query vector. A decoder is used to obtain a point proxy for each predicted center point based on the global features of all center points and multiple dynamic queries. A reconstruction head is used to generate a predicted missing point cloud based on each predicted center point and its point proxy, and to fuse the predicted missing point cloud with the original point cloud to obtain a complete point cloud.
2. The three-dimensional point cloud reconstruction method according to claim 1, characterized in that, Extracting multi-scale features from the original point cloud includes: The farthest point sampling algorithm is used to calculate the center points of the original point cloud at L different resolutions, thus obtaining the center point sets at different resolutions. Where, l∈{1,2,…,L}, N in The total number of points in the original point cloud; By using kNN to collect the neighbors of each center point at each resolution, the local point patch corresponding to that point is obtained. Using a KAN module with L independent parameters, nonlinear mapping is performed on the local point patches corresponding to all center points at different resolutions to obtain the feature representation F at different resolutions. (l) ; Fusing feature representations at different resolutions (l) The fused multi-scale features are obtained. .
3. The three-dimensional point cloud reconstruction method according to claim 2, characterized in that, Feature representations from different resolutions are fused using the following formula: , In the formula, , l∈{2,3,…,L}, Up represents the feature interpolation upsampling operation based on the distance between points, Concat represents feature concatenation, Conv is the convolution operation, and σ is the non-linear activation function.
4. The three-dimensional point cloud reconstruction method according to claim 1, characterized in that, Constructing a set of center points and calculating a point proxy for each center point in the set based on the multi-scale features, including: The farthest point sampling algorithm is used to calculate N center points of the original point cloud, thus obtaining the center point set. ; Center point C i The multi-scale features of k neighboring points in the original point cloud are aggregated to obtain the local semantic information g corresponding to the center point. i ; Center point C i By mapping to a high-dimensional space using the function Φ1, we obtain the first position embedding vector Φ1(c i ); The local semantic information g i and the first position embedding vector Φ1(c i Adding them together, we get the center point C. i Point agent.
5. The three-dimensional point cloud reconstruction method according to claim 1, characterized in that, Modeling the global semantic dependencies and local geometric relationships of the point proxies yields the global features of each center point, including: Calculate center point C from the set of center points. i Neighbor sets that take into account both semantic and geometric information ; Based on center point C i and its neighbor set N i The point proxy obtains the center point C. i Global features.
6. The three-dimensional point cloud reconstruction method according to claim 5, characterized in that, Calculate the center point C i The neighbor set N that takes into account both semantic and geometric information i , include: Collect center point C using kNN i Find the neighbors of the point to obtain the initial set corresponding to that point. ; Perform h jumps for expansion, and merge and deduplicate all the neighbor points obtained from these jumps to obtain the center point C. i The corresponding set of candidate points; Calculate the center point C i and the semantic-geometric score of each point in the corresponding candidate point set; Based on the semantic-geometric score, the center point C is selected from the candidate point set. i The neighbor set N is obtained by selecting a predetermined number of neighbor points with the highest scores. i .
7. The three-dimensional point cloud reconstruction method according to claim 5, characterized in that, Based on center point C i and its neighbor set N i The point proxy obtains the center point C. i Global features include: Based on center point C i and its neighbor set N i For each point in the network, perform point proxying and calculate the feature difference between the two. Center point C i The point proxy is concatenated with the feature difference to obtain a combined vector; Applying a linear transformation to each of the combined vectors yields the center point C. i Interaction characteristics with its neighbors; Max pooling is used for the center point C. i The center point C is obtained by processing the interaction characteristics with all its neighbors. i Local features; Center point C i The point proxy and its local features are fused to obtain the center point C. i Global features.
8. The three-dimensional point cloud reconstruction method according to claim 1, characterized in that, Based on the global characteristics of all central points, multiple dynamic queries are generated, including: Based on the global feature V of all center points i The M predicted center points are obtained using the following formula: , In the formula, For the j-th prediction center point, MLP pos For a multilayer perceptron used to predict the center point of a missing region, Pool is a global pooling mechanism; The query vector for each predicted centroid is generated using the following formula: , In the formula, Let be the query vector for the j-th dynamic query, MLP be a multilayer perceptron, and Φ2 be used to calculate the predicted center point. The second position embedding vector; Used to extract features surrounding the predicted center point from global features of all center points. Contextual information of the local region.
9. The three-dimensional point cloud reconstruction method according to claim 8, characterized in that, Based on each prediction center point and its point proxy, generate a prediction missing point cloud, including: Based on each prediction center point and its point proxy, generate the point cloud block corresponding to each prediction center point according to the following formula: , In the formula, Q i For the i-th point cloud block, H i Let G be the point proxy for the i-th predicted center point, G be the predefined regular grid, and f be the folding operation mapping function; Merge all point cloud blocks corresponding to the predicted center points to obtain the predicted missing point cloud.
10. A target detection method, characterized in that, include: A complete point cloud of a target is obtained using the three-dimensional point cloud reconstruction method according to any one of claims 1 to 9; The volume of the target is obtained from the complete point cloud of the target.
Citation Information
Cited By
Robot smart operating system, method and equipment for transparent object and medium
CN121613825A
Robotic dexterous manipulation system, method, apparatus and media for transparent objects
CN121613825B