Method for constructing three-dimensional shape from incomplete observation data
By combining the adaptive freeze-mix sampling module and the cross-resolution Transformer module, the problems of insufficient selection of key geometric feature points and insufficient multi-scale feature interaction in existing methods are solved, and high-precision 3D shape reconstruction is achieved.
Patent Information
- Application Number
- CN202511248762.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-11-18
AI Technical Summary
Existing methods lack the selection and preservation of key geometric feature points when reconstructing 3D shapes from incomplete observation data, and lack effective interaction mechanisms between multi-scale features, resulting in loss of structural information and insufficient detail restoration.
The Adaptive Freeze Hybrid Sampling (AFHS) module is used to preserve key structural points, and the feature interaction fusion is performed by combining the cross-resolution multi-scale Transformer module (MS-CRAFT). The complete shape is restored through a layer-by-layer cascaded point cloud refinement generator.
It significantly improves the accuracy and detail restoration capabilities of 3D shape reconstruction, especially in complex scenes where it demonstrates higher completion accuracy and geometric restoration capabilities.
Smart Images

Figure CN120976474A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a method for constructing a three-dimensional shape from incomplete observation data, and relates to the field of three-dimensional reconstruction. BACKGROUND
[0002] The three-dimensional point cloud completion task aims to reconstruct the complete three-dimensional structure from the partially observed point cloud, which has important application value in the fields of autonomous driving, augmented reality, robot perception, etc. However, due to the limitations of sensor accuracy, occlusion and view changes, the actual collected point cloud often has serious missing, which poses a challenge to the downstream tasks.
[0003] In order to overcome this problem, researchers have proposed various point cloud-based completion methods. PCN, as an early pioneering work, first adopts an encoder-decoder framework for global feature extraction and point cloud reconstruction, and subsequent methods such as TopNet, FoldingNet and SnowflakeNet further explore the layer-by-layer generation and multi-stage refinement strategy of point cloud structure. In addition, the Transformer architecture is also introduced into the completion process to enhance the modeling ability of global association between points. For example, PoinT and ProxyFormer model the point cloud completion as a set-to-set conversion problem, and introduce proxy points or query vectors to improve the generation quality. SeedFormer further maintains local consistency through Patch-based query and achieves leading performance on multiple datasets.
[0004] Although the above methods have made significant progress, they still face two key challenges: on the one hand, there is a lack of selection mechanism for key structure points in the encoding stage, which easily loses important geometric information; on the other hand, the fusion between feature layers is insufficient, which affects the completion ability of local details. SUMMARY
[0005] The application proposes a method for constructing a three-dimensional shape from incomplete observation data, named Freeze AdaCRAFT Point Cloud Network (FAC-PCN). First, an adaptive frozen hybrid sampling module (AFSM) is designed to adaptively retain key geometric points according to structural sensitivity, effectively improving the stability and expression ability of feature encoding. Second, a cross-resolution multi-scale attention enhanced transformer (MS-CRAFT) is introduced to interact and fuse features in multiple semantic spaces, enhancing the consistency of cross-scale context information and the ability to express details. Finally, a cascaded point cloud refinement generator is used to gradually reconstruct high-resolution point clouds from bottom to top, recovering complete structures and local details layer by layer from sparse seed points. Through comparative and ablation experiments on the PCN dataset, the results show that the application is superior to existing mainstream methods in multiple categories, especially in handling complex shapes and high-missing-rate point clouds, with higher completion accuracy and geometric restoration ability.
[0006] The three-dimensional point cloud completion task aims to reconstruct complete three-dimensional structures from partially observed point clouds, which has important application value in automatic driving, augmented reality, robot perception, etc. However, due to limitations such as sensor accuracy, occlusion, and changes in viewing angle, the actual collected point clouds often have serious missing, which poses a challenge to downstream tasks. To overcome this problem, researchers have proposed a method for constructing a three-dimensional shape from incomplete observation data, which is characterized by the following steps:
[0007] Step S1: Extract key geometric features by freezing the sampling hierarchical feature encoder to generate high-quality multi-scale feature representations;
[0008] Step S2: The point cloud seed generator uses these features to predict a rough seed point cloud to provide an initial structure for completion;
[0009] Step S3: The refinement generator gradually upsamples and enhances the details of the seed points to restore complete and delicate point cloud shapes.
[0010] In some possible embodiments, step S1 further includes:
[0011] Step S11: Obtain key points that take into account density and structure through the adaptive frozen hybrid sampling module;
[0012] Step S12: Use the PointNet+-based multi-scale structure to extract local and global information layer by layer, constantly compressing the space and improving the feature dimension;
[0013] Step S13: The cross-resolution module fuses features at different levels to strengthen semantic transmission between scales;
[0014] Step S14: integrate the multi-layer output to obtain a global feature vector.
[0015] In some possible embodiments, the adaptive frozen hybrid sampling module is composed of a curvature estimation module, a key point scoring network, a feature scoring network, a frozen mask generation and score fusion, and a Top-M sampling executor.
[0016] In some possible embodiments, step S2 further includes:
[0017] Step S21: generate seed features by fusing local spatial features and global semantic information through an upsampling module;
[0018] Step S22: obtain three-dimensional coordinates of the initial seed points by a multi-layer perception (MLP) mapping in combination with a global shape vector;
[0019] Step S23: output the seed point cloud.
[0020] In some possible embodiments, step S21 specifically includes:
[0021] Step S211: expand and fuse the point cloud features of the current scale by a multi-layer perception;
[0022] Step S212: a multi-scale feature enhanced cross-resolution module interacts and enhances the features between different scales and within the same scale;
[0023] Step S213: a deconvolution operation is used for upsampling in the spatial dimension, and the position of the point is adjusted in combination with the predicted displacement.
[0024] In some possible embodiments, the multi-scale feature enhanced cross-resolution module includes an Interpolation recursive multi-scale feature aggregation module, a Vector Attention (VA), a Residual Block residual feature extraction unit, and a Cat&MLP feature splicing and mapping structure.
[0025] In some possible embodiments, the multi-scale feature enhanced cross-resolution module is used to improve the feature interaction capability of the point cloud in the multi-scale scene, and is composed of an Inter-level CRAFT (inter-level enhancement) and an Intra-level CRAFT (intra-level enhancement), which respectively process the feature fusion and enhancement between different resolutions and within the same resolution.
[0026] The prior art method makes a breakthrough in modeling capability, but still has three deficiencies: one is the lack of selection and reservation of key geometric feature points, which easily leads to the loss of structure information in the completion process; two is the lack of effective interaction mechanism between multi-scale features, which affects the detail restoration of completion; three is the uneven distribution of points or structural distortion in complex scenes.
[0027] To overcome the above problems, the present application proposes a method for constructing a three-dimensional shape from incomplete observation data, which combines an adaptive frozen sampling mechanism and a cross-scale Transformer module to optimize each stage from feature coding, semantic enhancement to point cloud generation, and comprehensively improves the completion performance.
[0028] The present application proposes a method for constructing a three-dimensional shape from incomplete observation data, which aims to generate a high-fidelity complete three-dimensional shape from partial observation data. The method is designed in three stages of coding, enhancement and generation: first, through the adaptive frozen hybrid sampling module, the key structure points are effectively reserved in the coding stage, and the stability of feature expression is improved; second, the cross-resolution multi-scale Transformer (MS-CRAFT) is introduced to realize the deep interaction and semantic fusion between features of different scales; finally, through the point cloud refinement generator cascaded layer by layer, the complex geometric details are gradually recovered by upsampling. Experiments on the PCN dataset verify the effectiveness of FAC-PCN, which is superior to existing mainstream methods in accuracy, structure restoration and generalization ability.
[0029] The present application proposes a method for constructing a three-dimensional shape from incomplete observation data, which mainly includes three core modules: 1) adaptive frozen sampling encoder, which effectively reserves key structure points and improves feature expression ability and sampling fidelity; 2) cross-resolution multi-scale Transformer module (MS-CRAFT), which realizes the collaborative enhancement of interlayer and intralayer features; 3) point cloud generator with layer-by-layer upsampling, which integrates semantic and spatial information to gradually refine and restore complete point cloud. Experimental results show that FAC-PCN is superior to existing methods on the PCN dataset, especially in complex structure categories, showing stronger detail restoration ability and generalization performance.
[0030] Other aspects and advantages of the present application will be described in detail in the specific embodiments, and those skilled in the art can realize other aspects and advantages not explicitly stated in the present disclosure from the specific description below. BRIEF DESCRIPTION OF DRAWINGS
[0031] Figure 1 : FAC-PCN algorithm block diagram;
[0032] Figure 2 : frozen sampling hierarchical feature encoder;
[0033] Figure 3 Adaptive freeze-mix sampling module;
[0034] Figure 4 Inter-layer cross-resolution feature enhancement module (a) Intra-layer local feature enhancement module (b);
[0035] Figure 5 Point cloud refinement generator;
[0036] Figure 6 Upsampling module;
[0037] Figure 7 Comparison of visualization results of different methods on the PCN dataset. Detailed Implementation
[0038] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0039] The components of the embodiments of the invention described and illustrated in the accompanying drawings can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. Based on the embodiments of the invention, those skilled in the art will understand...
[0040] All other embodiments obtained with inventive effort are within the scope of protection of this invention. In the following, the terms "comprising," "having," and their cognates, used in the various embodiments of this invention, are intended only to indicate a particular feature, number, step, operation, element, component, or combination thereof, and should not be construed as primarily excluding the presence of one or more other features, numbers, steps, operations, elements, components, or combinations thereof, or adding the possibility of one or more combinations thereof.
[0041] Furthermore, the terms "first," "second," and "third" are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.
[0042] Unless otherwise specified, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of the invention pertain. Terms (such as those defined in commonly used dictionaries) shall be interpreted as having the same meaning as in their contextual meaning in the relevant technical field and shall not be interpreted as having an idealized or overly formal meaning, unless clearly defined in the various embodiments of the invention.
[0043] Point cloud completion methods. Point cloud, as a non-structured representation of three-dimensional data, is widely used in three-dimensional reconstruction and perception tasks due to its efficient storage and strong geometric expression. However, due to limitations such as sensor accuracy and occlusion of view angle, the collected point cloud is often incomplete. Therefore, a large number of studies have been carried out around point cloud completion.
[0044] PCN, as a pioneering work, first proposed a point cloud completion process based on an encoder-decoder structure, using a global feature regression method to predict the complete point set. Subsequent work has continuously optimized the encoding structure and decoding strategy. TopNet generates point cloud through a tree structure, and FoldingNet generates three-dimensional points through a two-dimensional grid folding method, improving the structural expression ability of completion.
[0045] Further, SnowflakeNet introduces a point level generation mechanism to generate point cloud from coarse to fine, effectively enhancing local detail expression. The PMP-Net series uses a method to predict point movement paths to complete point cloud expansion, enhancing geometric consistency. SeedFormer constructs the completion process through a local patch seed mechanism and combines Transformer to improve feature modeling capability. These methods have made some progress in improving completion quality, but most lack key point selection mechanism or multi-scale semantic modeling ability, making it difficult to accurately restore complex structures.
[0046] Application of Transformer in point cloud completion. With the success of Transformer in computer vision, its powerful global modeling capability has also been introduced into the point cloud completion task. PoinTr models the completion process as a set-to-set conversion problem and uses a Transformer decoder to enhance feature generation capability. ProxyFormer further introduces a proxy point mechanism to alleviate the problem of insufficient feature distribution caused by sparse input. SeedFormer improves the consistency and detail accuracy of the completed structure by explicitly constructing seed points as queries.
[0047] Although these methods have made breakthroughs in modeling capability, there are still three shortcomings: first, the selection and retention of key geometric feature points are lacking, which can easily lead to the loss of structural information in the completion process; second, there is a lack of effective interaction mechanism between multi-scale features, which affects the detail restoration of completion; third, the generated point distribution is uneven or the structure is distorted in complex scenes.
[0048] To overcome the above problems, the present invention proposes a method for constructing three-dimensional shapes from incomplete observation data, which integrates an adaptive frozen sampling mechanism and a cross-scale Transformer module to collaboratively optimize each stage from feature encoding, semantic enhancement to point cloud generation, and comprehensively improves the completion performance.
[0049] A method for constructing three-dimensional shapes from incomplete observation data, characterized in that it comprises the following steps:
[0050] Step S1: Key geometric features are extracted by freezing the sampling hierarchical feature encoder to generate a high-quality multi-scale feature representation;
[0051] Step S2: The point cloud seed generator predicts a rough seed point cloud using these features to provide an initial structure for completion;
[0052] Step S3: The refinement generator gradually upsamples and enhances the details of the seed points to restore the complete and delicate point cloud shape.
[0053] Specifically, as Figure 1 The overall framework of the method for constructing three-dimensional shapes from incomplete observation data is shown, which consists of three modules: adaptive frozen sampling feature encoder, seed point cloud generator, and point cloud refinement generator.
[0054] The input incomplete point cloud is first extracted by the frozen sampling hierarchical feature encoder (FHFE) to extract key geometric features, generating a high-quality multi-scale feature representation; then, the point cloud seed generator predicts a rough seed point cloud using these features to provide an initial structure for completion; finally, the refinement generator gradually upsamples and enhances the details of the seed points to restore the complete and delicate point cloud shape.
[0055] In some possible embodiments, step S1 further comprises:
[0056] Step S11: Obtain key points that take into account density and structure through the adaptive frozen hybrid sampling module;
[0057] Step S12: Use PointNet+-based multi-scale structure to extract local and global information layer by layer, constantly compressing space and improving feature dimension;
[0058] Step S13: The cross-resolution module fuses features at different levels to strengthen semantic transfer between scales;
[0059] Step S14: Integrate the multi-layer output to obtain a global feature vector.
[0060] In some possible embodiments, the adaptive frozen hybrid sampling module consists of a curvature estimation module, a key point scoring network, a feature scoring network, a frozen mask generation and score fusion, and a Top-M sampling executor.
[0061] Specifically, the Freeze-sampling Hierarchical Feature Encoder (FHFE) is used in FAC-PCN to extract point cloud features with multi-scale semantics. Unlike traditional encoders, FHFE introduces an adaptive freeze-sampling strategy, effectively preserving key geometric structure points and improving the accuracy and robustness of feature extraction.
[0062] Specifically, to achieve efficient and fine sampling and feature aggregation of three-dimensional point cloud key regions, this paper proposes an Adaptive Freeze-based Hybrid Sampling (AFHS) module. AFHS fully considers the geometric and semantic features of point clouds, and comprehensively utilizes the local geometric structure and global semantic information of point clouds, significantly improving the sensitivity of the model to important regions and the representativeness of sampling.
[0063] The overall structure of AFHS is shown in Figure 3 , which consists of a curvature estimation module, a key point scoring network, a feature scoring network, a freeze mask generation and score fusion, and a Top-M sampling executor.
[0064] First, the curvature estimation identifies the region with prominent structure and calculates as follows:
[0065]
[0066] where λ min is the smallest eigenvalue of the three eigenvalues of the covariance matrix, λ1, λ2, λ3 are the eigenvalues in descending order, and ε is a small constant to prevent division by zero error.
[0067] Subsequently, the key point scoring network predicts the importance probability of the point, and the feature scoring network evaluates the semantic expression ability, which are calculated by MLP respectively:
[0068] S key (p i )=σ(Conv1d(ReLU(BN(Conv1d(p i ))))) (2)
[0069] Here, p i represents the coordinates of the i-th point, σ is the Sigmoid activation function, BN is the batch normalization layer, Conv1d represents the one-dimensional convolution layer, and the final output S key is the key point probability distribution of the point cloud.
[0070] S feat (f i )=Conv1d(ReLU(BN(Conv1d(fi )))) (3)
[0071] where f i ∈R C represents the feature vector of the i-th point, C is the feature dimension, S feat represents the feature score.
[0072] After fusing the above information, a frozen mask is generated, and the key points are given higher priority to ensure their retention in sampling:
[0073]
[0074] where S combined = w c ·Curvature(p i )+w k ·S key (p i ), is the fusion score of curvature and key point prediction. The weights w c , w k control the proportion of curvature and key point probability in fusion respectively.
[0075] S adaptive (p i )=(1-M freeze (p i ))·S feat (p i )+M freeze (p i )·C max (5)
[0076] where C max is a very large constant, ensuring that the frozen points have the highest priority in the final sampling, thus realizing the retention of key structural information.
[0077] Finally, the Top-M sampler is used to select the points with the highest scores as the final sampling results for subsequent feature encoding:
[0078] Idx selected =Top-M(S adaptive ) (6)
[0079] Apply the final selected point index set to the original input point cloud coordinates to obtain the final adaptive sampled point cloud subset P selected , that is:
[0080]
[0081] Through the above process, AFHS achieves efficient mining and preservation of key structural regions in the original point cloud, improves the robustness and high-fidelity accuracy of model feature extraction, and thus lays a solid foundation for subsequent point cloud completion tasks.
[0082] In some possible embodiments, step S2 further includes:
[0083] Step S21: Generate seed features by fusing local spatial features and global semantic information through the upsampling module;
[0084] Step S22: Combine the global shape vector and obtain the three-dimensional coordinates of the initial seed point through multilayer perceptron (MLP) mapping;
[0085] Step S23: Output the seed point cloud.
[0086] Specifically, such as Figure 2 As shown, FHFE adopts a "layer-by-layer sampling + feature aggregation" design. The input point cloud first obtains key points that balance density and structure through the AFHS module; then, a multi-scale structure (SA-module) based on PointNet++ is used to extract local and global information layer by layer, continuously compressing the space and increasing the feature dimension; next, the cross-resolution Transformer module (MS-CRAFT) fuses features from different levels to enhance semantic transmission between scales; finally, the multi-layer outputs are integrated to obtain a global feature vector, providing a solid feature foundation for the completion task.
[0087] In some possible embodiments, step S21 specifically includes:
[0088] Step S211: Expand and fuse the point cloud features at the current scale using a multilayer perceptron;
[0089] Step S212: The multi-scale feature enhancement cross-resolution module interacts and enhances features between different scales and within the same scale;
[0090] Step S213: The deconvolution operation is used for upsampling in the spatial dimension and combined with the predicted displacement to adjust the position of the point.
[0091] In some possible embodiments, the multi-scale feature enhancement cross-resolution module includes: a recursive multi-scale feature aggregation module (Interpolation), a vector attention mechanism (VA), a residual feature extraction unit (Residual Block), and a feature concatenation and mapping structure (Cat&MLP).
[0092] In some possible embodiments, the multi-scale feature enhanced cross-resolution module is used to improve the feature interaction capability of the point cloud in a multi-scale scene, and is composed of an Inter-level CRAFT (inter-level enhancement) and an Intra-level CRAFT (intra-level enhancement), which are respectively used for processing feature fusion and enhancement between different resolutions and within the same resolution.
[0093] Specifically, in order to further improve the cross-scale information interaction capability in the three-dimensional point cloud completion task, the application proposes a multi-scale feature enhanced cross-resolution Transformer structure, named Multi-Scale Cross-Resolution Feature Augmented Transformer (MS-CRAFT for short). In the cross-scale feature fusion process, the MS-CRAFT introduces multiple effective feature enhancement mechanisms, including an interpolation multi-scale feature aggregation module (Interpolation), a vector attention mechanism (Vector Attention, VA), a residual feature extraction unit (Residual Block), and a feature concatenation and mapping structure (Cat & MLP), which significantly enhances the perception ability and expression robustness of the model to features of different resolutions.
[0094] The MS-CRAFT is used to improve the feature interaction capability of the point cloud in a multi-scale scene, and is mainly composed of an Inter-level CRAFT (inter-level enhancement) and an Intra-level CRAFT (intra-level enhancement), which are respectively used for processing feature fusion and enhancement between different resolutions and within the same resolution.
[0095] The structure of the inter-level cross-resolution feature enhancement module is as shown in Figure 4 The left half shows that this module uses a vector attention mechanism to realize feature enhancement of a high-resolution query point on a low-resolution support point. The following is the main process:
[0096] Given a query point cloud (P q ,F q ) and a support point cloud (P s ,F s ), hierarchical down-sampling is performed to obtain a scale set:
[0097]
[0098] At each scale, first, the vector attention mechanism is used to calculate the attention relationship between a point and its neighborhood points :
[0099]
[0100] where, is the relative position offset based on coordinate difference coding, a l (·) is the nonlinear mapping function based on MLP, denote the query, key, value vectors obtained by linear projection, respectively.
[0101] Subsequently, normalize the attention coefficients and perform local feature aggregation:
[0102]
[0103] The aggregated features will be mapped to the lower level scale by interpolation, and fused with the original features. This process is recursively executed from top to bottom, and finally outputs the enhanced features of the 0th layer
[0104] To make up for the shortcomings of Inter-level CRAFT in dealing with local geometric differences, this paper introduces the Intra-level CRAFT module to enhance the features within the same resolution. When the Query and Support point clouds are the same, the Inter-level module degenerates into the Intra-level structure, which performs self-attention, interpolation reconstruction and residual fusion, as shown in the right half of Figure 4 .
[0105] Intra-level CRAFT combines the local structure of point clouds to enhance feature consistency and detail expression through multi-scale residual enhancement. MS-CRAFT is based on recursive multi-scale design, which fuses inter-level and intra-level features to enhance the modeling ability of different resolution information. This module is embedded in each upsampling stage of FAC-PCN, effectively improving the geometric accuracy and detail restoration ability of point cloud reconstruction.
[0106] Specifically, the point cloud seed generator is used to generate an initial low-resolution seed point cloud from the encoded features, providing a basic structure for subsequent refinement and completion. This module first fuses local spatial features and global semantic information through an upsampling Transformer to generate seed features; then, combined with the global shape vector, it maps the three-dimensional coordinates of the initial seed points through a multi-layer perceptron (MLP). The final output of the seed point cloud can better restore the overall shape and lay the foundation for detail reconstruction.
[0107] As Figure 5 shown, the point cloud seed generator generates an initial low-resolution seed point cloud from the encoded features, first fuses local and global features through an upsampling Transformer, and then uses an MLP to predict the three-dimensional coordinates to preliminarily restore the overall structure.
[0108] The generator takes the initial point cloud seed position features and corresponding seed features generated by the point cloud seed generator as input, and sequentially passes through multiple up-sampling modules to gradually improve the resolution and detail quality, and finally outputs a dense and complete high-fidelity point cloud result.
[0109] The up-sampling process, as shown in Figure 6 , is composed of multiple cascaded up-sampling modules (Upsample Block), each module containing three core steps: feature fusion, cross-resolution Transformer enhancement (MS-CRAFT), and spatial up-sampling. First, the point cloud features at the current scale are expanded and fused using an MLP; then, the MS-CRAFT module interacts and enhances the features between different scales and within the same scale to improve global consistency and local detail expression ability; finally, the deconvolution operation is used for spatial dimension up-sampling, and the predicted displacement is used to adjust the position of the points, realizing the gradual improvement of point cloud density and the fine restoration of geometric structure.
[0110] In some possible embodiments, a loss function is also included, and the optimization of the point cloud completion algorithm relies on a proper loss function to effectively guide the model to generate more accurate completion results. The present application uses the Chamfer Distance (CD) of the L1 norm as the main loss function component. Chamfer distance is a commonly used loss function in point cloud completion tasks, which can comprehensively and effectively evaluate the overall spatial consistency and detail accuracy of the generated point cloud and the target point cloud.
[0111] Specifically, let the real complete target point cloud be The generated completion point cloud is The Chamfer distance is defined as:
[0112]
[0113] Where B is the number of batch samples; N pred is the number of points in the generated completion point cloud; N gt is the number of points in the target real complete point cloud; x is any point in the generated completion point cloud; y is any point in the target real point cloud; represents the square of the Euclidean distance between two points; min represents the minimum value in the set, used to find the nearest neighbor to calculate the minimum distance between points and points.
[0114] During training, the algorithm uses Chamfer distance as the only loss function to optimize network parameters, and continuously reduces the distance between the generated point cloud and the real complete point cloud through iteration, thereby achieving high-precision point cloud completion effect.
[0115] In order to verify the advantages of the method of the present application, specific examples are used for comparison test as follows.
[0116] We first systematically verify the effectiveness of the proposed modules through ablation experiments, focusing on analyzing the independent contribution and combined effect of each component in the multi-scale feature enhancement structure (MS-CRAFT), and exploring its impact on feature expression, geometric restoration, and overall completion performance. Subsequently, we conduct comparative experiments with existing advanced methods on multiple mainstream point cloud completion benchmark datasets, comprehensively evaluating the performance advantages and generalization ability of the proposed method from both quantitative indicators and visual effects.
[0117] Dataset
[0118] The PCN dataset, a widely recognized benchmark in the field of point cloud completion, is adopted. This dataset consists of 30,974 3D shapes selected from ShapeNet, covering eight categories such as airplanes, cabinets, cars, chairs, lamps, sofas, tables, and ships, providing a unified evaluation platform for point cloud completion methods.
[0119] Among them, the complete point cloud is obtained by uniformly sampling the original mesh surface, with a total of 16,384 points; the partial point cloud contains 2,048 points, generated by back-projection from a 2.5D depth map, simulating the acquisition scenario under real occlusion. This design effectively reflects the problem of incomplete data caused by sensor limitations in practical applications.
[0120] The PCN dataset uses the L1 version of the Chamfer distance (CD-l1) as the evaluation metric, measuring the bidirectional distance between the predicted results and the true point cloud, with lower values indicating better completion results. This dataset has been widely used in the evaluation of multiple completion methods, including FoldingNet, TopNet, and the recent FCS.
[0121] Experiments show that advanced methods perform more stably on PCN, generating smoother results with fewer noise points, especially on complex categories such as chairs, lamps, and tables. The PCN dataset continues to play a key role in promoting point cloud completion research and standardizing evaluation, serving as an important foundation for developing efficient and accurate algorithms.
[0122] Experimental environment
[0123] The method of the application is implemented based on a PyTorch framework, and all experiments are performed in a hardware and software environment as shown in Table 1. The experimental hyperparameter settings are as follows: the batch size is 16, the training epoch is 400, the initial learning rate is set to 0.001, and the cosine annealing strategy is used for dynamic adjustment. The Adam optimizer is used to train the model, wherein β1=0.9 and β2=0.999. In the training process, in order to enhance the generalization ability of the model, data enhancement techniques such as random rotation, jitter and scaling are applied. In the AFSM module, the weight coefficients of the key point freezing mechanism are set to α=0.5, β=0.3 and γ=0.2, and the number of K-neighbors is set to 32. In order to comprehensively evaluate the performance of the model, the Chamfer distance (CD) is used as the main evaluation index, which can effectively measure the shape similarity between the generated point cloud and the real point cloud.
[0124] TABLE I Soft and hardware environment configuration
[0125]
[0126] Comparative experiment
[0127] In order to comprehensively evaluate the performance of FAC-PCN in the point cloud completion task, we compared it with many methods, including FoldingNet, PCN, TopNet, GRNet, FBNet, PoinTr, SnowflakeNet and FSC. These methods represent different technical routes in the field of point cloud completion: FoldingNet and PCN are point-based methods; GRNet is a voxel-based method; and PoinTr, SnowflakeNet and the like represent Transformer-based methods. In order to ensure the fairness of the comparison, we strictly followed the parameter settings and training strategies in the original papers of each method to perform experiments on the same data set division. The evaluation index uses the Chamfer distance (CD-l1) in the L1 version, which measures the average minimum distance between the predicted point cloud and the real point cloud, and the smaller the value, the better the completion effect.
[0128] TABLE II Chamfer distance x 10 on IPCN dataset 3 (smaller is better) results
[0129]
[0130] Table 2 shows the chamfer distance results of each method on the eight classes of the PCN dataset, where the bold numbers represent the best performance in each column. From the results, it is clear that our proposed FAC-PCN method achieves the best performance on all eight classes, with an average chamfer distance reduction of 2.1% compared to the second place SeedFormer.
[0131] To more intuitively demonstrate the completion effect of FAC-PCN, Figure 7 A visualization comparison of the completion results of different methods on the PCN dataset is presented. From the visualization results, it can be clearly seen that the point cloud generated by FAC-PCN not only maintains the integrity of the overall shape, but also significantly outperforms other methods in detail processing.
[0132] Ablation experiments
[0133] To better understand the contribution of each component of FAC-PCN to the overall performance, we designed a series of ablation experiments to verify the effect of the adaptive frozen sampling module, the cross-resolution non-local collaborative attention aggregator, and the combination of the two. The experimental results are shown in Table III, where "baseline model" refers to the network structure using standard farthest point sampling (FPS) and traditional attention mechanism; "AFSM" indicates the addition of the adaptive frozen sampling module; "MS-CRAFT" indicates the addition of the cross-resolution multi-scale feature enhancement Transformer; and "FAC-PCN" is the complete model.
[0134] TABLE III FAC-PCN Ablation Experiment Data Table
[0135]
[0136] From Table III, we can see that after adding the AFSM module, the average CD decreases from 7.52 to 6.93, indicating that the frozen sampling strategy effectively preserves key geometric features; after adding the MS-CRAFT module, the CD further decreases to 6.87, indicating that cross-scale feature enhancement improves the point cloud expression ability. The complete model FAC-PCN integrates both modules, with CD reaching the lowest value of 6.60, which is 12.2% higher than the baseline model, verifying the synergistic effect of the combination of the two modules.
[0137] TABLE IV Effect of Different Freezing Ratios (CD x 10-3)
[0138]
[0139] As shown in Table IV, the freezing ratio has a significant impact on the completion effect. When the freezing ratio increases from 0% to 30%, the CD continuously decreases, reaching the lowest value of 6.60 at 30%, indicating that moderately preserving key points can achieve a good balance between geometric expression and uniformity of distribution; further increasing the ratio slightly decreases the performance.
[0140] Table V. Impact of number of multi-scale feature levels (CD x 10-3)
[0141]
[0142] Table V shows the impact of different number of layers on performance and parameter amount in MS-CRAFT. As the number of layers increases from 1 to 3, CD decreases to 6.60, indicating that multi-layer feature fusion helps to improve the completion effect; when the number of layers is 4, CD decreases slightly to 6.58, but the parameter amount increases significantly, so finally a 3-layer structure is adopted to balance the performance and efficiency.
[0143] The method proposed by the application aims to generate a complete three-dimensional shape with high fidelity from partial observation data. The method is systematically designed from three stages of encoding, enhancement and generation: first, through the adaptive frozen hybrid sampling module, the key structure points are effectively reserved in the encoding stage, and the stability of feature expression is improved; secondly, the cross-resolution multi-scale Transformer (MS-CRAFT) is introduced to realize the deep interaction and semantic fusion between features of different scales; finally, through the point cloud refinement generator cascaded layer by layer, the complex geometric details are gradually recovered by upsampling. Experiments on the PCN dataset verify the effectiveness of FAC-PCN, which is superior to existing mainstream methods in accuracy, structure restoration and generalization ability.
[0144] It should be understood that, although the application has been specifically disclosed through preferred embodiments and optional features, modifications, improvements and changes to the application disclosed herein can be made by those skilled in the art, which are considered to be within the scope of the application. The materials, methods and examples provided herein are representative and exemplary of the preferred embodiments and are not intended to limit the scope of the application.
Claims
1. A method of constructing a three-dimensional shape from incomplete observation data, characterized by Comprising the following steps: Step S1: Extract key geometric features by freezing the sampling hierarchical feature encoder to generate high-quality multi-scale feature representation; Step S2: The point cloud seed generator predicts a rough seed point cloud using these features to provide an initial structure for completion; Step S3: The refinement generator gradually upsamples and enhances the details of the seed points to restore the complete and delicate point cloud shape.
2. The method of constructing a three-dimensional shape from incomplete observation data according to claim 1, wherein Step S1 further comprises: Step S11: Obtain key points that take into account density and structure by an adaptive freezing mixed sampling module; Step S12: Use a multi-scale structure based on PointNet+ to extract local and global information layer by layer, constantly compressing the space and improving the feature dimension; Step S13: Cross-resolution module fuses features of different levels to strengthen semantic transmission between scales; Step S14: Integrate multi-layer output to obtain a global feature vector.
3. A method of constructing a three-dimensional shape from incomplete observation data according to claim 2, characterized in that Adaptive freezing The mixed sampling module consists of a curvature estimation module, a key point scoring network, a feature scoring network, a freezing mask generation and score fusion, and a Top-M sampling executor.
4. A method of constructing a three-dimensional shape from incomplete observation data according to claim 1 or 2, characterized in that Step S2 further comprises: Step S21: Fuse local spatial features and global semantic information through the upsampling module to generate seed features; Step S22: Combine the global shape vector to obtain the three-dimensional coordinates of the initial seed points through a multi-layer perception (MLP) mapping; Step S23: Output the seed point cloud.
5. A method of constructing a three-dimensional shape from incomplete observation data according to claim 4, characterized by Step S21 specifically comprises: Step S211: Use a multi-layer perception to expand and fuse the point cloud features of the current scale; Step S212: The multi-scale feature enhanced cross-resolution module interacts and enhances the features between different scales and within the same scale; Step S213: Deconvolution operation is used for spatial dimension upsampling, and the position of the point is adjusted in combination with the predicted displacement.
6. The method of constructing a three-dimensional shape from incomplete observation data according to claim 5, wherein The multi-scale feature enhanced cross-resolution module includes: recursive multi-scale feature aggregation module (Interpolation), vector attention mechanism (Vector Attention, VA), residual feature extraction unit (Residual Block), and feature concatenation and mapping structure (Cat&MLP).
7. The method of constructing a three-dimensional shape from incomplete observation data according to claim 5 or 6, characterized in that the multi-scale feature enhanced cross-resolution module is used to improve the feature interaction capability of the point cloud in the multi-scale scene, and is composed of inter-layer enhancement and intra-layer enhancement, which respectively process the feature fusion and enhancement between different resolutions and within the same resolution.