Robust three-dimensional surface reconstruction method based on multi-scale fusion

By combining traditional geometric fitting and deep learning feature modeling into a multi-scale fusion method, the problems of insufficient accuracy and robustness in 3D surface reconstruction are solved, and high-precision and interpretable 3D surface reconstruction results are achieved.

CN120976446AActive Publication Date: 2025-11-18SHANDONG INST OF BUSINESS & TECH +1

Patent Information

Application Number
CN202511500006.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2025-11-18
Estimated Expiration
2045-10-21

AI Technical Summary

Technical Problem

Existing 3D surface reconstruction methods struggle to guarantee high accuracy, robustness, and interpretability when dealing with issues such as noise, sparse sampling, complex topology, and local missing data. Traditional geometric methods rely on manually defined features, while deep learning methods lack explicit geometric constraints.

Method used

A robust 3D surface reconstruction method based on multi-scale fusion is adopted, which combines traditional geometric fitting and deep learning feature modeling. The local implicit surface is fitted by weighted least squares method, multi-scale convolution and attention mechanism are introduced to generate geometric and deep learning mesh model, and 3D surface reconstruction is achieved by dual-branch fusion and implicit function decoder.

Benefits of technology

It achieves high-precision, robust, and interpretable 3D surface reconstruction, can handle complex geometries and noisy point cloud data, improves the robustness and detail recovery capability of the model, and generates topologically continuous and smooth triangular mesh models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976446A_ABST
    Figure CN120976446A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of three-dimensional surface reconstruction, and particularly relates to a robust three-dimensional surface reconstruction method based on multi-scale fusion. Acquiring three-dimensional point cloud data, fitting local implicit curved surfaces of neighborhood points by adopting a weighted least square method, and generating a geometric reconstruction grid model; extracting geometric features in different receptive fields through multi-scale convolution, introducing channel attention, space attention and a gating mechanism to perform weighted fusion on the geometric features in different receptive fields, enhancing a global dependency relationship by adopting a feature mixing operation, and generating a deep learning grid model; performing double-branch fusion on the geometric reconstruction grid model and the deep learning grid model, highlighting a geometric sensitive area, and obtaining a fused enhanced feature; inputting the fused enhanced features into an implicit function decoder, and predicting a signed distance function value of each vertex; and based on the predicted signed distance function field, a contour surface extraction algorithm is adopted to generate a triangular mesh model, so that three-dimensional surface reconstruction is completed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of three-dimensional surface reconstruction, and particularly relates to a robust three-dimensional surface reconstruction method based on multi-scale fusion. BACKGROUND

[0002] Three-dimensional surface reconstruction is one of the core problems in the field of computer vision and graphics, and is widely used in many fields such as autonomous driving, virtual reality, cultural heritage digitization and industrial detection. With the development of three-dimensional scanning technology, point cloud data has gradually become the main input form of three-dimensional surface reconstruction. However, the actual collected point cloud data often has problems such as noise, sparse sampling, complex topological structure and local missing, which makes it difficult for traditional three-dimensional reconstruction methods to guarantee high precision, smoothness and robustness.

[0003] Existing three-dimensional surface reconstruction methods can be broadly divided into two categories: traditional geometric methods and deep learning-based methods.

[0004] Traditional geometric methods: traditional geometric reconstruction methods usually rely on explicit or implicit mathematical modeling, such as weighted least squares, polynomial fitting, radial basis functions, etc. These methods can generate accurate geometric surfaces through local fitting, especially in the case of less noise, sparse data and regular topology. However, the shortcomings of traditional methods are that they rely on manually set features and priori, and have limited processing capacity for noise, missing data and large-scale point clouds, especially when facing complex three-dimensional shapes, it is often difficult to guarantee precision and efficiency.

[0005] Deep learning-based methods: In recent years, deep learning-based three-dimensional reconstruction methods have made significant progress. Deep neural networks can better handle noisy data and irregular sampling by automatically learning complex geometric patterns from large-scale data. Many studies use convolutional neural networks, generative adversarial networks and other technologies for three-dimensional surface reconstruction, which can achieve stronger robustness and generalization ability. However, deep learning methods usually lack explicit geometric constraints, which limits their ability to model details and restore local geometric structures. In addition, the "black box" nature of deep networks also makes them less interpretable, which cannot provide intuitive understanding for engineering applications.

[0006] To overcome the above problems, some hybrid methods combining traditional geometric methods and deep learning techniques have been proposed in recent years, trying to take advantage of the high precision of traditional geometric methods and the powerful feature learning ability of deep learning methods. However, existing fusion methods can only take into account the advantages of both to some extent, and still have deficiencies in handling large-scale point clouds, complex geometric shapes and high-precision modeling.

[0007] Therefore, how to effectively combine the traditional geometric fitting method with the feature learning ability of deep learning, and propose a three-dimensional surface reconstruction method with high precision, robustness and explainability, has become an important direction of current research. SUMMARY

[0008] In order to overcome the problems in the prior art, the application provides a robust three-dimensional surface reconstruction method based on multi-scale fusion.

[0009] The technical scheme for solving the above technical problems is as follows: The application provides a robust three-dimensional surface reconstruction method based on multi-scale fusion, comprising the following steps: Obtain three-dimensional point cloud data, select target points and their neighborhood points, use weighted least squares method to fit the local implicit surface of the neighborhood points, extract local geometric priors, and generate a geometric reconstruction grid model; Obtain three-dimensional point cloud data, extract geometric features under different receptive fields through multi-scale convolution, introduce channel attention, spatial attention and gating mechanism to weight and fuse the geometric features under different receptive fields, and use feature mixing operation to enhance global dependency, generate a deep learning grid model; Fuse the geometric reconstruction grid model and the deep learning grid model in double branches, combine the comprehensive index of the Gaussian curvature and the adjacent edge length of each vertex in the geometric reconstruction grid model and the deep learning grid model, highlight the geometric sensitive area, and realize feature aggregation in the spatial neighborhood, to obtain the fused enhanced features; Input the fused enhanced features into the implicit function decoder to predict the signed distance function value of each vertex; based on the predicted signed distance function field, use the isosurface extraction algorithm to generate a triangular mesh model, thereby completing three-dimensional surface reconstruction.

[0010] Further, obtain three-dimensional point cloud data, select target points and their neighborhood points, use weighted least squares method to fit the local implicit surface of the neighborhood points, extract local geometric priors, and generate a geometric reconstruction grid model, comprising: Select target points in the three-dimensional point cloud data, define the local field point range, and construct the local implicit surface of the neighborhood points in the local field point range; Use weighted least squares method to fit the local implicit surface of the neighborhood points to form a global implicit surface; and sample on the global implicit surface to generate an initial triangular mesh model through gridding operation.

[0011] Furthermore, 3D point cloud data is acquired, and geometric features under different receptive fields are extracted through multi-scale convolution. Channel attention, spatial attention, and gating mechanisms are introduced to weightedly fuse the geometric features under different receptive fields. Feature fusion operations are then employed to enhance global dependencies, generating a deep learning mesh model, including: Acquire 3D point cloud data and extract geometric features under different receptive fields by setting multi-scale convolutional paths in parallel. The importance of geometric features under different receptive fields is weighted using the channel attention mechanism to obtain the geometric features after channel attention weighting. Based on the geometric features weighted by the channel attention mechanism, the fusion of channel attention and spatial attention is adaptively controlled by combining the spatial attention mechanism and introducing a gating factor, resulting in features fused by the channel attention and spatial attention mechanisms. The features fused by channel attention and spatial attention are stitched together with the acquired 3D point cloud data and passed as input to a multilayer perceptron. The corresponding directed distance function value is regressed point by point, and a deep learning grid model is generated using the MarchingCubes method.

[0012] Furthermore, by combining spatial attention mechanisms and introducing gating factors, the fusion of channel attention and spatial attention is adaptively controlled, resulting in features after the fusion of channel attention and spatial attention, including: ; In the above formula, This represents the features after the fusion of channel attention and spatial attention; Represents the learnable channel attention weights; Represents learnable spatial attention weights; Indicates the gating factor; This represents the geometric features after being weighted by the channel attention mechanism.

[0013] Furthermore, learnable spatial attention weights for: ; In the above formula, Indicates the first i The feature vector of each point is the original information input into the network; GAP represents global average pooling. This represents the LeakyReLU activation function. This represents the Sigmoid activation function; and The weight matrix of the fully connected layer helps the model capture and adjust the relationships between features, thereby improving the model's expressive power and performance.

[0014] Further, the learnable spatial attention weight is: ; In the above formula, represents a Sigmoid activation function; represents the input feature of the i th point; represents a depth separable convolution.

[0015] Further, the geometric reconstruction grid model and the deep learning grid model are fused in a double branch, the comprehensive indicators of the Gaussian curvature and the adjacent edge length of each vertex in the geometric reconstruction grid model and the deep learning grid model are combined, the geometric sensitive area is highlighted, and feature aggregation is realized in the spatial neighborhood to obtain the fused enhanced features, including: The Gaussian curvature and the edge saliency of the vertices in the deep learning grid model and the geometric reconstruction grid model are calculated respectively, the Gaussian curvature and the edge saliency are spliced, and then nonlinear mapping is performed to obtain nonlinear mapping features; Based on the deep learning grid model and the geometric reconstruction grid model, an attention subgraph based on near neighbor search is constructed in the spatial neighborhood range, and the correlation weight between the target point and the neighborhood point is calculated using the attention subgraph; based on the correlation weight, the features in the neighborhood are weighted and aggregated to obtain local perception enhanced features; The nonlinear mapping features and the local perception enhanced features are fused, and the importance of different features is adaptively adjusted through a gated channel-spatial attention mechanism to obtain enhanced local features, including multi-scale perception gated fusion enhanced features and geometric surface reconstruction enhanced features; The multi-scale perception gated fusion enhanced features and the geometric surface reconstruction enhanced features are respectively subjected to global average pooling to obtain multi-scale perception gated global representation vectors and geometric surface reconstruction global representation vectors; The multi-scale perception gated global representation vectors and the geometric surface reconstruction global representation vectors are spliced and then input into two layers of fully connected mapping and nonlinear activation function to obtain the fused enhanced features.

[0016] Further, based on the deep learning grid model and the geometric reconstruction grid model, an attention subgraph based on near neighbor search is constructed in the spatial neighborhood range, and the correlation weight between the target point and the neighborhood point is calculated using the attention subgraph, including: For each target point p , the near neighbor points are searched in the spatial neighborhood range to form a spatial neighborhood point set Within the spatial neighborhood set, an attention subgraph is constructed; the relevance weights between the target point and each neighboring point are calculated using the attention subgraph, and these relevance weights reflect the strength of the correlation between local features within the neighborhood; among them, the relevance weights... for: ; In the above formula, d The scaling factor for the feature dimension. For target point p The set of neighborhood points; Indicates the target point p A neighborhood index in the set of neighbors; Indicates the target point p The representation of the feature vector in the projection space, i.e., the query vector; Representing neighborhood points q The eigenvectors are represented in the projection space as key vectors.

[0017] Furthermore, the fused enhanced features are input into the implicit function decoder to predict the signed distance function value for each vertex, including: The fused enhanced features are then stitched together with the corresponding original 3D point cloud data to form stitched features. The concatenated features are input into an implicit function decoder, which employs a multilayer perceptron structure to perform nonlinear mapping and feature regression on each point, predicting the signed distance function value corresponding to that point point by point.

[0018] Furthermore, it also includes: introducing a joint loss function, which includes SDF regression loss, gradient consistency loss and marginal constraint loss.

[0019] Compared with the prior art, the present invention has the following technical effects: (1) Starting from the complex geometric features of point cloud data, this invention proposes a robust 3D surface reconstruction method based on multi-scale fusion. Combining the advantages of traditional geometric fitting and deep learning feature modeling, high-precision, end-to-end reconstruction of point cloud data is achieved through the collaborative design of a geometric surface reconstruction branch and a scale-aware gating fusion branch. In the geometric surface reconstruction branch, a weighted least squares method is used to perform a second implicit surface fitting on the local neighborhood of the point cloud, extracting local geometric priors, and generating a geometric reconstruction mesh model through sampling and meshing, thereby maintaining high fitting accuracy and interpretability in complex local regions. At the same time, a scale-aware gating fusion branch is proposed. The scale-aware gating fusion branch captures geometric features under different receptive fields through parallel multi-scale convolution, and combines channel attention, spatial attention and gating mechanism to perform weighted fusion of geometric features. At the same time, feature mixing operation is introduced to enhance global dependencies, so that the model can take into account both local structure and global consistency.

[0020] (2) In order to better handle the geometric sensitive area and local context relationship in the point cloud data, a double-branch fusion module is designed in the embodiment. The double-branch fusion module introduces a curvature and edge saliency fusion unit to highlight geometric details by combining the Gaussian curvature and edge intensity indicators, and combines a local perception attention unit to construct an attention subgraph in the neighborhood range, so as to realize dynamic weighted aggregation and consistency modeling of features, thereby effectively improving the robustness of the model to complex geometric details and noise points.

[0021] (3) On the basis of fusing features, the embodiment inputs the enhanced representation and the point coordinates into an implicit function decoder together, and uses a multi-layer perception to predict the signed distance function value point by point. Subsequently, the isosurface extraction algorithm is used at the zero level set to generate a triangular mesh model with topological continuity and smooth surface. This process can convert the implicit result predicted by the network into explicit geometric structure, taking into account both geometric precision and renderability.

[0022] (4) In order to ensure stable training and accurate reconstruction of the model, a joint loss function is introduced in the embodiment, including SDF regression loss, gradient consistency loss and edge constraint loss. Among them, the SDF regression loss is used to minimize the difference between the predicted value and the true value, the gradient consistency loss is used to improve the accuracy of the predicted normal vector, and the edge constraint loss assigns higher weights to high curvature and boundary regions to highlight geometric details. Through the above multiple constraints, the network can achieve a good balance between detail recovery and global consistency BRIEF DESCRIPTION OF DRAWINGS In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the drawings needed in the embodiment or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0023] Figure 1 The flowchart of the present application is shown in the figure. Figure 2 The reconstruction effect of the present application and one of the other methods is shown in the figure. Figure 3 The reconstruction effect of the present application and the other method is shown in the figure. DETAILED DESCRIPTION

[0024] To further clarify the technical means and effects taken by the present application to achieve the intended purpose, the specific implementation, structure, features and effects of the technical solutions proposed according to the present application are described in detail below in combination with the drawings and preferred embodiments. The specific features, structures or characteristics in one or more embodiments can be combined in any suitable form. Unless otherwise defined, all technical and scientific terms used in the present application have the same meaning as understood by those skilled in the art to which the present application belongs.

[0025] In the present embodiment, referring to Figure 1 , a robust three-dimensional surface reconstruction method based on multi-scale fusion is provided, comprising the following steps: Obtain three-dimensional point cloud data, select target points and their domain points, fit a quadratic implicit surface in the local neighborhood of the point cloud based on the weighted least squares method, extract local geometric priors, and generate a geometric reconstruction grid model; Obtain three-dimensional point cloud data, extract geometric features under different receptive fields through multi-scale convolution, introduce channel attention, spatial attention and gating mechanism to weight and fuse geometric features under different receptive fields, and use feature mixing operation to enhance global dependency, generate a deep learning grid model; Fuse the geometric reconstruction grid model and the deep learning grid model into a double-branch fusion, combine the comprehensive index of the Gaussian curvature of each vertex and the length of the adjacent edge, highlight the geometric sensitive area, and realize feature aggregation in the spatial neighborhood to obtain the fused enhanced features; Input the fused enhanced features into the implicit function decoder to predict the signed distance function value of each vertex; based on the predicted signed distance function field, use the isosurface extraction algorithm to generate a triangular mesh model, thereby completing the three-dimensional surface reconstruction.

[0026] The above steps are described in detail as follows: Step 100: Obtain three-dimensional point cloud data, in the geometric surface reconstruction branch, select target points and their domain points, use the weighted least squares method to fit the local implicit surface of the neighborhood points, extract local geometric priors, and generate a geometric reconstruction grid model, which is the output of the geometric surface reconstruction branch.

[0027] As an example, the present step 100 specifically includes: Step 110: Select a target point in the three-dimensional point cloud data, define its local domain point range, and construct a local implicit surface of the neighborhood points within the local domain point range.

[0028] For each target point in the three-dimensional point cloud , the first-order neighborhood points of the target point are defined as the local neighborhood points, wherein, represents the set of local neighborhood points related to the target point; Represents local neighborhood points; This represents the set of first-order local neighborhood points of the target point.

[0029] Based on the defined local neighborhood of the target point, a local implicit surface is constructed for the ternary quadratic neighborhood points: ; In the above formula, Indicates the first i A local implicit surface function constructed centered on a target point; , , , , , , , , , The coefficients to be fitted together determine the specific shape and position of the implicit surface; Indicates the target point The coordinates.

[0030] Step 120: Fit the local implicit surface of the neighborhood points using the weighted least squares method to form the global implicit surface; and sample the global implicit surface to generate the initial triangular mesh model through meshing operation.

[0031] The method of fitting the local implicit surface of neighboring points using weighted least squares to form a global implicit surface includes: ; In the above formula, Point Compared to the first i The weights of the local implicit surface functions reflect the degree of influence of the local surface on that location; S represents the total number of neighborhood points involved in the fitting. Represents a global implicit surface function.

[0032] The weighting function employs a spherical decay model, where only points falling within the effective influence range of the curved surface participate in the weighted fusion. The weighting function is defined as follows: ; in: Point With the target point The distance; The target point The distance from a first-order neighbor point to the farthest point in its own first-order neighborhood. A point can only participate in weighted fusion if it falls within the effective influence range of a certain surface.

[0033] On this basis, the multiple local implicit surfaces are fused by weighting to form a global consistent and continuous implicit surface, and sampling is performed on the global implicit surface to generate an initial triangular mesh model through meshing operation.

[0034] Step 200: Obtain three-dimensional point cloud data, in the multi-scale perception gate fusion branch, extract geometric features under different receptive fields through multi-scale convolution, introduce channel attention, spatial attention and gate mechanism to weight and fuse the geometric features under different receptive fields, and further use feature mixing operation to enhance global dependency, generate a deep learning grid model, that is, the output of the multi-scale perception gate fusion branch.

[0035] As an example, the step 200 specifically includes: Step 210: Obtain three-dimensional point cloud data After that, multiple-scale convolution paths are set in parallel to extract geometric features under different receptive fields: ; In the above formula, represents the geometric features under different receptive fields; represents splicing, splicing the output features of the three different scale convolution paths; is a parallel convolution path with three different convolution kernel sizes (1×1, 3×3, 5×5); represents the result of the convolution operation on the input point cloud x using a 1×1 convolution kernel; represents the result of the convolution operation on the input point cloud x using a 3×3 convolution kernel; represents the result of the convolution operation on the input point cloud x using a 5×5 convolution kernel; represents an activation function, which is used for nonlinear transformation or further feature extraction on ; represents an activation function, which is used for nonlinear transformation or further feature extraction on ; represents an activation function, which is used for nonlinear transformation or further feature extraction on ; N represents the number of points in the point cloud; D represents the dimension of the feature vector of each point.

[0036] A smaller convolution kernel (such as 1×1) can capture subtle geometric features in the point cloud, such as edges, corners, etc. A larger convolution kernel (such as 5×5) can capture more extensive context information, which helps the model understand the overall structure of the point cloud. By splicing features of different scales, the model can utilize both local details and global context information, thereby enhancing the understanding and representation ability of the point cloud geometric features.

[0037] Step 220: Weight the importance of geometric features under different receptive fields using a channel attention mechanism to obtain geometric features weighted by the channel attention mechanism: ; In the above formula, represents the geometric features weighted by the channel attention mechanism; is a learnable weight vector; represents an element-wise multiplication operation.

[0038] Step 230: Based on the geometric features weighted by the channel attention mechanism, the fusion of channel attention and spatial attention is adaptively regulated by combining the spatial attention mechanism and introducing a gating factor to obtain features fused by channel attention and spatial attention.

[0039] The weighted geometric features are combined with the spatial attention mechanism to highlight areas with significant local geometric changes, and a 1x1 convolution is used to generate a gating factor at each point position , which is used to adaptively regulate the fusion ratio of channel attention and spatial attention, enhance the complementarity and expression balance between the two, and adaptively regulate the fusion ratio of channel attention and spatial attention through the gating mechanism: ; In the above formula, represents the features fused by channel attention and spatial attention; represents a learnable channel attention weight; represents a learnable spatial attention weight.

[0040] Among them, through global average pooling and fully connected network, the expression ability of each channel feature in global semantics can be captured, thereby establishing global dependency in channel dimension and generating a set of learnable channel attention weights for each channel : ; In the above formula, represents the feature vector of the i th point, i.e. the original information input into the network.

[0041] GAP represents global average pooling; represents a LeakyReLU activation function, represents a Sigmoid activation function; and are weight matrices of the fully connected layer, which help the model capture and adjust the relationship between features, thereby improving the expression ability and performance of the model.

[0042] By applying depthwise separable convolution to the feature vector, the response intensity of each point in its local neighborhood is extracted, and a corresponding spatial attention weight is generated for each point. To highlight areas with significant geometric changes Represented as: ; In the above formula, This indicates that a depthwise separable convolution with a kernel size of 7×7 applies a feature mixing operation to the fused features to achieve non-local information interaction between different points, thereby enhancing global dependencies and geometric representation capabilities.

[0043] Step 240: Combine the features fused by channel attention and spatial attention with the acquired 3D point cloud data. The data is concatenated and passed as input to a multilayer perceptron (MLP). The corresponding SDF (Signed Distance Function) values ​​are regressed point by point, and the Marching Cubes method is used to generate a deep learning grid model.

[0044] Marching Cubes is a classic algorithm for extracting isosurfaces from three-dimensional scalar fields, such as volume data composed of SDF values.

[0045] Step 300: Input the output of the geometric surface reconstruction branch (geometric reconstruction mesh model) and the output of the multi-scale perception gating fusion branch (deep learning mesh model) into the dual-branch fusion module at the same time. In the dual-branch fusion module, the comprehensive index of Gaussian curvature of each vertex and adjacent edge length is combined to highlight the expressive ability of geometrically sensitive regions.

[0046] As an example, step 300 specifically includes: Step 310: Calculate the Gaussian curvature and edge saliency of the vertices, concatenate the Gaussian curvature and edge saliency, and then perform nonlinear mapping to obtain the nonlinear mapping features.

[0047] Gaussian curvature and edge saliency are two important indicators for measuring the geometric characteristics of a 3D model.

[0048] Gaussian curvature is a scalar value describing the degree of curvature of a surface at a point. For target points in deep learning mesh models and geometric reconstruction mesh models... p Its Gaussian curvature It can be calculated using the included angle between adjacent triangular faces: ; In the above formula, For this vertex on the triangle face t The included angle within; represents a set of triangular faces adjacent to the vertex.

[0049] Edge saliency is an index measuring the length variation of the adjacent edges of a vertex, which is used to capture the edge features in the model: ; In the above formula, is the set of adjacent edges of the vertex; represents the target point coordinates; represents the neighborhood point coordinates; p, q represents the vertex; represents the Euclidean distance.

[0050] After splicing the Gaussian curvature and the edge saliency, a nonlinear mapping is performed: ; In the above formula, represents the nonlinear mapping feature.

[0051] By splicing the Gaussian curvature and the edge saliency, the bending degree of the surface and the edge features can be comprehensively utilized to describe the geometric characteristics of the vertex. The nonlinear mapping can further extract and enhance these geometric features.

[0052] Step 320: Constructing an attention subgraph based on K-neighborhood search within the spatial neighborhood range, where K=16, and using the attention subgraph to calculate the correlation weight between the target point and the K-neighborhood points; based on the correlation weight, the features within the neighborhood are weighted and aggregated to obtain local perception enhanced features, i.e., multi-scale perception gated fusion enhanced features and geometric surface reconstruction enhanced features.

[0053] By constructing an attention subgraph based on neighbor search within the spatial neighborhood range through the local perception attention unit, and using the attention subgraph to calculate the correlation weight between the target point and the neighborhood points, including: For each target point p , retrieve the neighbor points within its spatial neighborhood range to form a set of spatial neighborhood points ; within the spatial neighborhood set range, construct an attention subgraph; use the attention subgraph to calculate the correlation weight between the target point and each neighborhood point, and the correlation weight reflects the strength of the correlation of local features within the neighborhood; wherein the correlation weight is: ; In the above formula, d is the scaling factor of the feature dimension, is the set of neighborhood points of the target point p ; and represents the target point pone of the neighborhood indexes in the neighbor set of the target point representation of the target point p representation of the feature vector of the target point in the projection space, i.e., the query vector representation of the feature vector of the neighborhood point q representation of the feature vector of the neighborhood point in the projection space, i.e., the key vector

[0054] The calculated correlation weight is used to weight and aggregate the features of the neighborhood points to obtain local perception enhanced features. The weighted aggregation can enhance the local consistency of the overall features and improve the expression ability of details. The local perception enhanced features are as follows: In the above formula, represents the feature vector of the i-th point. p

[0055] This mechanism can effectively introduce highly correlated neighborhood information in space while maintaining the features of the vertex itself, thereby improving the smoothness and structural consistency of the feature representation.

[0056] Step 330: Fuse the nonlinear mapping features and the local perception enhanced features , and adaptively adjust the importance of different features through the gated channel-space attention mechanism to obtain enhanced local features , which are represented as follows: In the above formula, Attention weights are assigned in the channel and spatial dimensions to highlight key geometric regions.

[0057] Step 340: Perform global average pooling on the multi-scale perception gated fusion enhanced features and the geometric surface reconstruction enhanced features respectively to obtain multi-scale perception gated global representation vectors and geometric surface reconstruction global representation vectors; concatenate the multi-scale perception gated global representation vectors and the geometric surface reconstruction global representation vectors, and then input them into two layers of fully connected mapping and nonlinear activation functions to obtain fused enhanced features.

[0058] For the multi-scale perception gated fusion branch and the geometric surface reconstruction branch, perform global average pooling on the multi-scale perception gated fusion enhanced features and the geometric surface reconstruction enhanced features to obtain multi-scale perception gated global representation vectors and global representation vectors . Concatenate the multi-scale perception gated global representation vectors and the global representation vectors ​​​​The concatenated features are input into two layers of fully connected mapping and a nonlinear activation function to learn the complementary semantics between the two reconstruction mechanisms and obtain fused enhanced features : ; In the above formula, represents a multi-scale perception gate fusion enhanced feature; represents a geometric surface reconstruction enhanced feature; represents a feature concatenation operation, , is a learnable mapping matrix, is the dimension of the fused features.

[0059] Step 400: input the fused enhanced features into an implicit function decoder to predict the signed distance function values of each vertex.

[0060] The fused enhanced features are concatenated with the corresponding original three-dimensional point cloud data to form concatenated features, The concatenated features are input into an implicit function decoder; wherein the implicit function decoder adopts a multi-layer perception structure, performs nonlinear mapping and feature regression on each point, and predicts the signed distance function value corresponding to the point.

[0061] By introducing a nonlinear activation function to transform the features of each layer, the expression ability of the network is enhanced, and at the same time, through an end-to-end regression training method, it is ensured that the predicted signed distance function value can accurately describe the spatial geometric distribution of the point cloud data, thereby realizing implicit modeling and continuous expression of three-dimensional surfaces.

[0062] Step 500: based on the predicted signed distance function field, an isosurface extraction algorithm is used to generate a continuous and smooth triangular mesh model, thereby completing three-dimensional surface reconstruction.

[0063] Based on the signed distance function field predicted by the implicit function decoder, an isosurface extraction algorithm is used at the zero level set to perform step-by-step traversal and judgment on the voxel space, calculate the intersection position of the voxel boundary and the isosurface, and generate a topologically continuous surface structure. Each intersection point is connected to form a triangular facet to construct a mesh, and finally a continuous, smooth and high-precision triangular mesh model is obtained, realizing complete three-dimensional surface reconstruction.

[0064] This embodiment introduces a joint loss function, including SDF regression loss, gradient consistency loss and edge constraint loss. Among them, the SDF regression loss is used to minimize the difference between the predicted value and the true value, the gradient consistency loss is used to improve the accuracy of the predicted normal vector, and the edge constraint loss is used to assign higher weights in high curvature and boundary areas to highlight geometric details.

[0065] In the training stage, a joint loss function is adopted The model is constrained in the form of: ; In the above formula, the first term is the SDF regression loss, which calculates the mean square error between the predicted SDF value and the true SDF value, and is used to minimize the difference between the predicted value and the true value; the second term is the gradient consistency constraint loss, which improves the surface normal accuracy by the similarity of the included angle between the predicted gradient and the normal vector; the third term is the edge constraint term loss, which allocates higher weights in high curvature or geometric detail areas to enhance the model's reconstruction ability for boundaries and details.

[0066] wherein, represents the true SDF value, represents the predicted SDF value, represents the number of training samples; represents the gradient consistency constraint loss weight; represents the included angle of the normal vector; represents the edge constraint term loss weight; represents the gradient of the predicted SDF value; is an edge weight calculated based on local curvature, edge saliency, etc.

[0067] In summary, the robust three-dimensional surface reconstruction method based on multi-scale fusion proposed in the embodiment can not only process large-scale, noisy point cloud data, but also accurately recover the geometric structure of high curvature and irregular boundary areas. Experimental results show that this method not only improves the reconstruction accuracy, but also significantly enhances the robustness and detail performance of the model, providing an efficient and feasible solution for high-fidelity reconstruction of point cloud data. The reconstruction effect visualization is shown in Figure 2 and Figure 3 .

[0068] Referring to Figures 2-3 , the distribution robust optimization method (SDRO) based on Sinkhorn distance, the HotSpot method, the OffsetOPT method, the divergence guided shape implicit neural representation method (DIGS), the implicit geometry regularization method (IGR), the predictable context prior method (PCP), the screened Poisson surface reconstruction method (SPSR), the unsigned learning method (SAL), the shape as point method (SAP), the point convolution method (POCO) for surface reconstruction, and the ground truth (GT), wherein the ground truth is the real data as the benchmark.

[0069] Figure 2We demonstrate the reconstruction results of our method and other methods. Most of the other methods have ghosting phenomenon, that is, false surfaces are generated at positions where they should not appear, or the reconstructed results are not connected. In contrast, our method largely alleviates these problems and maintains topological consistency. Although the offset optimization method (OffsetOPT) and the divergence-guided shape implicit neural representation method (DIGS) perform close to our method on some objects, their results significantly degrade on the third reconstruction task of a more difficult industrial part, while our method still maintains stable robustness and high accuracy.

[0070] Figure 3 We demonstrate the reconstruction results of our method and other methods. Most of the other methods have ghosting phenomenon, that is, false surfaces are generated at positions where they should not appear, or the reconstructed results are not connected. In contrast, our method largely alleviates these problems and maintains topological consistency. Although the offset optimization method (OffsetOPT) and the divergence-guided shape implicit neural representation method (DIGS) perform close to our method on some objects, their results significantly degrade on the third reconstruction task of a more difficult industrial part, while our method still maintains stable robustness and high accuracy. Figure 3 As shown in FIG. 6, the reconstruction results of other methods still have obvious defects in visual effects. The reconstruction results of the divergence-guided shape implicit neural representation method (DIGS), the point convolution method for surface reconstruction (POCO), the hot spot method (HotSpot), and the offset optimization method (OffsetOPT) have large areas of surface missing. The reconstruction results of the predictable context prior method (PCP), the unsigned learning method (SAL), the shape as a point method (SAP), and the Sinkhorn distance-based distribution robust optimization method (SDRO) are insufficient in terms of geometric detail description. The screened Poisson surface reconstruction method (SPSR) produces an error closure. In summary, the reconstruction results of our method are significantly better than those of other methods in terms of accuracy and stability.

[0071] The above examples are only used to illustrate the technical solutions of the present application, but not limit it; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing examples, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A robust three-dimensional surface reconstruction method based on multi-scale fusion, characterized in that, Includes the following steps: Acquire 3D point cloud data, select target points and their neighboring points, use weighted least squares method to fit the local implicit surface of the neighboring points, extract local geometric priors, and generate a geometric reconstruction mesh model; We acquire 3D point cloud data, extract geometric features under different receptive fields through multi-scale convolution, introduce channel attention, spatial attention and gating mechanism to perform weighted fusion of geometric features under different receptive fields, and use feature fusion operation to enhance global dependency to generate a deep learning grid model. By fusing the geometric reconstruction mesh model and the deep learning mesh model in a dual-branch manner, and combining the comprehensive indices of Gaussian curvature and adjacent edge length of each vertex in the geometric reconstruction mesh model and the deep learning mesh model, the geometrically sensitive areas are highlighted, and feature aggregation is achieved in the spatial neighborhood to obtain the enhanced features after fusion. The fused enhanced features are input into an implicit function decoder to predict the signed distance function value of each vertex; based on the predicted signed distance function field, an isosurface extraction algorithm is used to generate a triangular mesh model, thereby completing the 3D surface reconstruction.

2. The robust three-dimensional surface reconstruction method based on multi-scale fusion according to claim 1, characterized in that, Acquire 3D point cloud data, select target points and their neighborhood points, fit the local implicit surface of the neighborhood points using the weighted least squares method, extract local geometric priors, and generate a geometrically reconstructed mesh model, including: Select a target point in the 3D point cloud data, define its local neighborhood point range, and construct a local implicit surface for the neighboring points within the local neighborhood point range; We use the weighted least squares method to fit the local implicit surface of the neighborhood points to form the global implicit surface; and then sample the global implicit surface to generate the initial triangular mesh model through meshing.

3. The robust three-dimensional surface reconstruction method based on multi-scale fusion according to claim 1, characterized in that, 3D point cloud data is acquired, and geometric features under different receptive fields are extracted through multi-scale convolution. Channel attention, spatial attention, and gating mechanisms are introduced to weightedly fuse the geometric features under different receptive fields. Feature fusion operations are then used to enhance global dependencies, generating a deep learning mesh model, including: Acquire 3D point cloud data and extract geometric features under different receptive fields by setting multi-scale convolutional paths in parallel. The importance of geometric features under different receptive fields is weighted using the channel attention mechanism to obtain the geometric features after channel attention weighting. Based on the geometric features weighted by the channel attention mechanism, the fusion of channel attention and spatial attention is adaptively controlled by combining the spatial attention mechanism and introducing a gating factor, resulting in features fused by the channel attention and spatial attention mechanisms. The features fused by channel attention and spatial attention are stitched together with the acquired 3D point cloud data and passed as input to a multilayer perceptron. The corresponding directed distance function value is regressed point by point, and a deep learning grid model is generated using the Marching Cubes method.

4. The robust three-dimensional surface reconstruction method based on multi-scale fusion according to claim 3, characterized in that, By combining spatial attention mechanisms and introducing gating factors, the fusion of channel attention and spatial attention is adaptively controlled, resulting in features after the fusion of channel attention and spatial attention, including: ; In the above formula, This represents the features after the fusion of channel attention and spatial attention; Represents the learnable channel attention weights; Represents learnable spatial attention weights; Indicates the gating factor; This represents the geometric features after being weighted by the channel attention mechanism.

5. A robust three-dimensional surface reconstruction method based on multi-scale fusion according to claim 4, characterized in that, Learnable Spatial Attention Weights for: ; In the above formula, Indicates the first i The feature vector of a point; GAP represents global average pooling; This represents the LeakyReLU activation function. This represents the Sigmoid activation function; and This is the weight matrix of the fully connected layer.

6. A robust three-dimensional surface reconstruction method based on multi-scale fusion according to claim 5, characterized in that, Learnable Spatial Attention Weights for: ; In the above formula, This represents the Sigmoid activation function; Indicates the first i Input features of 1 point; This represents depthwise separable convolution.

7. A robust three-dimensional surface reconstruction method based on multi-scale fusion according to claim 1, characterized in that, A dual-branch fusion of the geometric reconstruction mesh model and the deep learning mesh model is performed. By combining the Gaussian curvature of each vertex and the length of adjacent edges in both models, geometrically sensitive regions are highlighted, and feature aggregation is achieved within the spatial neighborhood to obtain the enhanced fused features, including: The Gaussian curvature and edge saliency of vertices in the deep learning-based mesh model and the geometric reconstruction mesh model are calculated respectively. The Gaussian curvature and edge saliency are concatenated and then nonlinearly mapped to obtain the nonlinear mapping features. Based on deep learning grid models and geometric reconstruction grid models, attention subgraphs based on nearest neighbor retrieval are constructed in the spatial neighborhood, and the correlation weights between the target point and the neighborhood points are calculated using these attention subgraphs. Based on the correlation weights, the features in the neighborhood are weighted and aggregated to obtain local perception enhancement features. Nonlinear mapping features are fused with local perception enhancement features, and the importance of different features is adaptively adjusted through a gated channel-spatial attention mechanism to obtain enhanced local features, including multi-scale perception gated fusion enhancement features and geometric surface reconstruction enhancement features. The multi-scale perception gated fusion enhancement features and the geometric surface reconstruction enhancement features are respectively subjected to global average pooling to obtain the global representation vector of multi-scale perception gated fusion and the global representation vector of geometric surface reconstruction. The multi-scale perception gated global representation vector and the geometric surface reconstruction global representation vector are concatenated and then fed into two fully connected layers and a nonlinear activation function to obtain the fused enhanced features.

8. A robust three-dimensional surface reconstruction method based on multi-scale fusion according to claim 7, characterized in that, Based on deep learning grid models and geometric reconstruction grid models, attention subgraphs based on nearest neighbor retrieval are constructed within the spatial neighborhood. These attention subgraphs are then used to calculate the relevance weights between the target point and its neighboring points, including: For each target point p Search for nearest neighbors within its spatial neighborhood to form a spatial neighborhood point set. Within the spatial neighborhood set, construct an attention subgraph; use the attention subgraph to calculate the relevance weights between the target point and each neighboring point. for: ; In the above formula, d The scaling factor for the feature dimension. For target point p The set of neighborhood points; Indicates the target point p A neighborhood index in the set of neighbors; Indicates the target point p The representation of the feature vector in the projection space, i.e., the query vector; Representing neighborhood points q The eigenvectors are represented in the projection space as key vectors.

9. A robust three-dimensional surface reconstruction method based on multi-scale fusion according to claim 1, characterized in that, The fused enhanced features are input into the implicit function decoder to predict the signed distance function value for each vertex, including: The fused enhanced features are then stitched together with the corresponding original 3D point cloud data to form stitched features. The concatenated features are input into an implicit function decoder, which employs a multilayer perceptron structure to perform nonlinear mapping and feature regression on each point, predicting the signed distance function value corresponding to that point point by point.

10. A robust three-dimensional surface reconstruction method based on multi-scale fusion according to claim 1, characterized in that, Also includes: A joint loss function is introduced, which includes SDF regression loss, gradient consistency loss, and marginal constraint loss.

Citation Information

Patent Citations

  • Aircraft surface reconstruction method based on local geometric features and implicit distance field

    CN116468767A

  • Three-dimensional point cloud data segmentation method and system based on multi-scale point features

    CN119648536A

  • News scene three-dimensional reconstruction and visualization method based on multi-source remote sensing data

    CN119904592A

  • Robust three-dimensional reconstruction method based on weighted local curved surface approximation

    CN120014204A

  • Building digital twin three-dimensional reconstruction method and system based on large model

    CN120339540A

Cited By

  • Multi-scale scene reconstruction method based on multi-source data fusion

    CN121544812A

  • Surface shape measurement method based on implicit neural modeling and meta-learning

    CN121708085A