Three-dimensional reconstruction method and system, electronic equipment and storage medium

By introducing attention mechanisms and multi-head codebooks into the three-dimensional reconstruction method and constructing a triangle mesh model in combination with implicit fields, the challenges in the capture of complex shapes and geometric details in the prior art are solved, and high-precision three-dimensional reconstruction is achieved.

CN120182476APending Publication Date: 2025-06-20BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510123455.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Existing 3D reconstruction methods pose significant challenges in capturing complex shapes and geometric details, especially due to the sparsity, irregularity and noise interference of point cloud distributions.

Method used

A transformation network based on the attention mechanism is adopted that combines points and grids to extract the initial point features of point cloud data and combine the global-local relational attention mechanism to generate fusion point features. Then, a discrete embedded features are generated using the preset multi-head codebook, and finally a triangular mesh model is constructed through an implicit field to perform three-dimensional surface reconstruction.

Benefits of technology

It realizes high-precision three-dimensional reconstruction in complex shape and topological scenarios, and can simultaneously capture local geometric details and global structural information of point clouds, improving the accuracy and detail retention ability of three-dimensional reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182476A_ABST
    Figure CN120182476A_ABST
Patent Text Reader

Abstract

The invention provides a three-dimensional reconstruction method and system, and belongs to the technical field of computer vision, and the method comprises the steps: extracting an initial point feature of each space point from point cloud data, inputting the initial point feature into a transformation network, and obtaining a grid feature of a grid where each space point is located; based on a global-local relation attention mechanism and each grid feature, obtaining a fusion point feature of each spatial point; for any query point, obtaining an embedded feature according to the at least one fusion point feature, and generating a discretized embedded feature according to the embedded feature and a preset multi-head codebook; and obtaining an implicit field according to the embedded features and the discretized embedded features of all query points, and constructing a triangular mesh model through an edge-based contour surface extraction method to complete three-dimensional surface reconstruction. According to the method, implicit representation and a contour surface extraction method are combined through a global-local relation attention mechanism and discrete feature expression based on a multi-head codebook, and high-precision three-dimensional reconstruction is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a three-dimensional reconstruction method, system, electronic device and storage medium. Background Art

[0002] In the fields of computer graphics and vision, three-dimensional reconstruction is a key technology for constructing virtual environments and achieving geometric understanding, and is widely used in scenarios such as autonomous driving, virtual reality, and industrial design. As a common three-dimensional data representation method, point cloud has become an important basis for three-dimensional reconstruction due to its efficient spatial sampling ability. However, due to the sparsity, irregularity of point cloud distribution and noise interference, existing three-dimensional reconstruction methods still face significant challenges in capturing complex shapes and geometric details.

[0003] Under the current technical background, traditional explicit representation methods directly construct three-dimensional shapes through voxels, meshes or point clouds. Voxel methods have low storage and computational efficiency due to the large amount of data at high resolutions; although mesh methods can express complex topological structures, they usually rely on template deformation or predefined topological structures and are difficult to adapt to diverse scenarios; point cloud methods often require additional surface reconstruction steps due to the lack of continuous topological relationships, and the processing process is complex and prone to losing detail information. In contrast, implicit representation techniques have received attention in recent years due to their advantage of characterizing geometric shapes through continuous functions, but existing methods mostly rely on global features to describe shapes, and this method has limited ability in characterizing complex geometric shapes and local details and is difficult to capture the mutual relationship between global and local features. In addition, implicit representation techniques have insufficient ability to handle sharp edges in the mesh generation stage, and problems such as blurred boundaries or lost details often occur.

[0004] Therefore, how to simultaneously capture the local geometric details and global structure information of point clouds during the three-dimensional reconstruction process, so as to achieve high-precision three-dimensional reconstruction in complex shapes and topological scenarios, has become a technical problem that needs to be solved urgently. Summary of the Invention

[0005] The present invention provides a three-dimensional reconstruction method, system, electronic device and storage medium to solve the defects in the prior art, and simultaneously capture the local geometric details and global structure information of point clouds during the three-dimensional reconstruction process, so as to achieve high-precision three-dimensional reconstruction in complex shapes and topological scenarios.

[0006] The present invention provides a three-dimensional reconstruction method, including the following steps: Extract the initial point features of each spatial point from the point cloud data; Input each of the initial point features into a transformation network that combines points and meshes based on the attention mechanism to obtain the mesh features of the mesh where each spatial point is located; Based on the global-local relationship attention mechanism and each of the grid features, obtain the fused point feature of each spatial point; For any query point, based on at least one of the fused point features, obtain an embedded feature, and based on the embedded feature and a preset multi-head codebook, generate a discretized embedded feature; Based on the embedded features and the discretized embedded features of all the query points, obtain an implicit field for characterizing the three-dimensional shape; Based on the implicit field, construct a triangular mesh model through an edge-based isosurface extraction method to complete the three-dimensional surface reconstruction.

[0007] According to a three-dimensional reconstruction method provided by the present invention, the step of obtaining the fused point feature of each spatial point based on the global-local relationship attention mechanism and each of the grid features specifically includes: Based on the global-local relationship attention mechanism, construct a query-key-value generation network; Input the grid features of each face of the grid into the query-key-value generation network to correspondingly obtain the attention weights of each face; According to the attention weights of each face, perform weighted fusion on the grid features of each face to obtain the fused point feature.

[0008] According to a three-dimensional reconstruction method provided by the present invention, the preset multi-head codebook includes a first preset number of preset sub-codebooks; the step of obtaining an embedded feature based on at least one of the fused point features, and generating a discretized embedded feature based on the embedded feature and the preset multi-head codebook specifically includes: Determine a second preset number of neighboring points according to the Euclidean distance between the query point and each of the spatial points; Based on the fused point features of all the neighboring points, obtain the embedded feature; Divide the embedded feature into a first preset number of parts; For each part, input it into the corresponding preset sub-codebook, and by measuring the distance between the part and each code vector in the sub-codebook, select the code vector closest to the part to generate the discretized embedded sub-feature corresponding to the part; Connect the discretized embedded sub-features of each part into an overall discretized embedded feature.

[0009] According to a three-dimensional reconstruction method provided by the present invention, the step of obtaining the embedded feature based on the fused point features of all the neighboring points specifically includes: Obtain the embedded feature through the following formula: ; ; In the formula, is the embedded feature, is the connection operation, is the combined feature, is the fused point feature of the i-th neighboring point, q is the query point, MLP is the multi-layer perceptron, and k is the second preset quantity.

[0010] According to a 3D reconstruction method provided by the present invention, obtaining an implicit field for characterizing the 3D shape according to the embedded features and the discretized embedded features of all the query points specifically includes: Concatenate the fused point features and the discretized embedded features of all the query points and input them into a multi-layer perceptron to obtain the implicit field output by the multi-layer perceptron.

[0011] According to a 3D reconstruction method provided by the present invention, constructing an overall triangular mesh by an edge-based isosurface extraction method according to the implicit field to complete 3D surface reconstruction specifically includes: Calculate the implicit field values and normal directions of the eight vertices of each cubic unit of the implicit field respectively; Determine the intersecting edges of each cubic unit according to the respective implicit field values and the respective normal directions; Obtain the initial intersection position by interpolating the implicit field values of the endpoints of each intersecting edge, and correct the initial intersection position according to the normal directions of the endpoints of the intersecting edge to obtain the corrected intersection position; Perform triangulation on all the corrected intersection positions in combination with a preset lookup table to generate corresponding local triangular meshes; Repeat the above steps for all cubic units within the target area, and integrate all the local triangular meshes into the triangular mesh model to complete 3D surface reconstruction.

[0012] According to a 3D reconstruction method provided by the present invention, determining the intersecting edges of each cubic unit according to the respective implicit field values and the respective normal directions specifically includes: For each group of adjacent vertices of each cubic unit, when the signs of the implicit field values of the adjacent vertices are different and the normal directions of the adjacent vertices satisfy a preset angle threshold condition, determine that the edge between the current adjacent vertices is the intersecting edge.

[0013] The present invention also provides a 3D reconstruction system, including the following modules: The first processing module is used to extract the initial point features of each spatial point from the point cloud data; The first processing module is further configured to input each of the initial point features into a point-and-grid combination transformation network based on an attention mechanism to obtain the grid features of the grid where each spatial point is located; The first processing module is further configured to obtain the fused point features of each spatial point based on the global-local relationship attention mechanism and the respective grid features; The second processing module is configured to, for any query point, obtain an embedded feature according to at least one of the fused point features, and generate a discretized embedded feature according to the embedded feature and a preset multi-head codebook; The second processing module is further configured to obtain an implicit field for characterizing the three-dimensional shape according to the embedded features and the discretized embedded features of all the query points; The second processing module is further configured to construct an overall triangular mesh through an edge-based isosurface extraction method according to the implicit field to complete three-dimensional surface reconstruction.

[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, the three-dimensional reconstruction method as described in any one of the above is implemented.

[0015] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the three-dimensional reconstruction method as described in any one of the above is implemented.

[0016] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the three-dimensional reconstruction method as described in any one of the above is implemented.

[0017] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: By extracting the initial point features of each spatial point from the point cloud data, the local geometric information of the point cloud data can be effectively obtained, providing a basis for subsequent feature extraction and geometric analysis. By inputting each initial point feature into a transformation network that combines points and meshes based on the attention mechanism, the geometric relationship between the point cloud and the mesh can be utilized to dynamically adjust the weight distribution between features, thereby more accurately capturing the local and global geometric structure features of spatial points. By obtaining the fused point features of each spatial point based on the global-local relationship attention mechanism and each mesh feature, the local detail information of the points and the global structure information of the mesh can be effectively combined. For any query point, by obtaining the embedded feature according to at least one fused point feature and generating the discretized embedded feature according to the embedded feature and the preset multi-head codebook, the continuous feature of the spatial point can be mapped to an efficient discrete feature space, thereby utilizing the global geometric prior information provided by the multi-head codebook to further improve the compactness and generalization ability of the feature representation. By obtaining the implicit field for characterizing the three-dimensional shape according to the embedded features and discretized embedded features of all query points, the geometric characteristics of the three-dimensional shape can be represented as a continuous spatial function, thereby realizing the natural expression of complex topological structures. By constructing a triangular mesh model using the edge-based isosurface extraction method according to the implicit field, the implicitly represented three-dimensional shape can be efficiently converted into an explicit mesh representation, thereby simultaneously capturing the local geometric details and global structure information of the point cloud during the three-dimensional reconstruction process, and thus achieving high-precision three-dimensional reconstruction in complex shape and topological scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 FIG. is one of the flow diagrams of the three-dimensional reconstruction method provided by the present invention.

[0020] Figure 2 FIG. is another flow diagram of the three-dimensional reconstruction method provided by the present invention.

[0021] Figure 3 FIG. is yet another flow diagram of the three-dimensional reconstruction method provided by the present invention.

[0022] Figure 4 FIG. is still another flow diagram of the three-dimensional reconstruction method provided by the present invention.

[0023] Figure 5 FIG. is yet still another flow diagram of the three-dimensional reconstruction method provided by the present invention.

[0024] Figure 6 It is a schematic diagram of the visualization results of different categories on the ShapeNet dataset provided by the present invention.

[0025] Figure 7 It is a schematic diagram of the structure of the 3D reconstruction system provided by the present invention.

[0026] Figure 8 It is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed implementation manners

[0027] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0028] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitations, the element defined by the phrase "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element. The orientation or positional relationship indicated by the terms "upper", "lower", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the system or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation of the present invention. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0029] The terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type, and do not limit the number of objects. For example, the first object can be one or multiple. In addition, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0030] The following is combined withFigures 1 - 8 Describe the 3D reconstruction method, system, electronic device and storage medium provided by the present invention.

[0031] Figure 1 It is one of the schematic flowcharts of the 3D reconstruction method provided by the present invention. As Figure 1 shown, it includes but is not limited to the following steps: Step 101: Extract the initial point features of each spatial point from the point cloud data.

[0032] In the 3D surface reconstruction method of the present invention, the purpose of step 101 is to extract the initial point features of each spatial point from the point cloud data, so as to provide a basis for subsequent feature extraction and 3D shape implicit field modeling. The point cloud data is composed of discrete points in space, and each point carries its position coordinate information in the 3D space. Relying solely on the position coordinate information cannot fully describe the local geometric relationship and global distribution characteristics of the point cloud. Therefore, it is necessary to extract the initial point features that can express the local and global structures of the point cloud to lay a foundation for subsequent high-level feature calculations.

[0033] In a specific implementation, the point cloud data is input into a point-wise multi-layer perceptron (MLP) to learn the initial point feature fp for each point pi. The reason for choosing MLP as the initial feature extraction module is that MLP is an efficient and general feature transformation tool that can perform non-linear transformation on the input 3D coordinate data to extract the geometric and distribution features of the points. Specifically, MLP contains multiple fully connected layers, and each layer is followed by a ReLU activation function to enhance the non-linear expression ability of the network, and the batch normalization layer is used to improve the training stability and convergence speed.

[0034] When extracting the initial point features, the input of MLP is the 3D coordinates pi=(xi,yi,zi) of each point, and through layer-by-layer transformation, a feature vector fp with increased dimensions is obtained. This feature vector not only contains the position information of the point, but also integrates the local geometric relationship and global spatial distribution characteristics of the point. In terms of the specific configuration of the network, the first layer of MLP maps the 3D coordinates to a higher-dimensional feature space (such as 64 dimensions), and then gradually extracts deeper geometric features through multiple hidden layers, and finally outputs a point feature vector with a fixed dimension (such as 128 dimensions).

[0035] In this way, the initially extracted point features can effectively represent the local geometric properties of each point and its positional relationship in the overall structure, providing high-quality input for the subsequent point-mesh transformation network and global-local relationship attention mechanism. The experimental results show that the initially extracted point features using the above method can significantly improve the detailed reconstruction effect of 3D shapes, especially in the processing of details of complex geometries and sharp edges. This initial feature extraction process lays a solid foundation for the efficient global and local information fusion in 3D surface reconstruction, while effectively reducing the computational complexity of directly processing point cloud data.

[0036] Step 102: Input each of the initially extracted point features into the point-mesh combined transformation network based on the attention mechanism to obtain the mesh features of the grid where each spatial point is located.

[0037] In the 3D surface reconstruction method of the present invention, the purpose of step 102 is to transform the initially extracted point features of each spatial point into corresponding mesh features through the point-mesh combined transformation network, providing the necessary context information support for the feature fusion of the subsequent global-local relationship attention mechanism. The discreteness and sparsity of point cloud data limit the direct use of point features for 3D shape modeling. Introducing a mesh structure can provide more topologically continuous geometric information for point features, thereby achieving a more accurate 3D surface representation.

[0038] Specifically, input the initially extracted point feature fp of each spatial point into the point-mesh combined transformation network. Through the interaction between the point feature and the mesh feature, extract the mesh feature fg of the grid where the point is located. The core of this transformation network is to use the point-net Transformer module to process the initially extracted point features of the point cloud. The point-net Transformer is constructed with a multi-layer encoding structure, and captures the relationship between the point feature and the mesh feature through the query-key-value generation mechanism. Specifically, it can be expressed as: ; where is the point-net Transformer module.

[0039] After the above processing, the mesh feature fg of each point's location not only effectively compensates for the information loss caused by the sparsity of point cloud data, but also provides geometric information support with more continuity and global perception ability, laying a solid foundation for the global-local relationship attention mechanism processing and implicit field construction in the subsequent steps. The mesh features generated in this way show significant superiority in experiments, especially in the modeling of complex geometries and sharp edges, effectively improving the reconstruction accuracy and detail retention ability, while ensuring the improvement of computational efficiency.

[0040] Step 103: Based on the global-local relationship attention mechanism and each grid feature, obtain the fused point feature for each spatial point.

[0041] In the 3D surface reconstruction method of the present invention, the core objective of Step 103 is to combine each grid feature with its corresponding global and local geometric information through the global-local relationship attention mechanism, thereby generating richer and more accurate fused point features Fp. The local features of the point cloud data can well capture the geometric attributes of each point in space, but relying solely on local features cannot comprehensively describe the global structure of the 3D shape. Therefore, combining local geometric information with global context can enhance the detail expressiveness of the model while ensuring the overall accuracy of the reconstruction.

[0042] In a possible implementation manner, refer to Figure 2 , Figure 2 is the second flowchart of the 3D reconstruction method provided by the present invention. As Figure 2 shown, Step 103 specifically includes Steps 201 - 203: Step 201: Based on the global-local relationship attention mechanism, construct a query-key-value generation network.

[0043] Step 202: Input the grid feature of each face of the grid where the spatial point is located into the query-key-value generation network, and correspondingly obtain the attention weight for each face.

[0044] Step 203: According to the attention weight of each face, perform weighted fusion on the grid features of each face to obtain the fused point feature.

[0045] Specifically, the core objective of the global-local relationship attention module is to comprehensively consider the global structure information and local geometric details to dynamically adjust the contribution weight of each grid face feature to the point feature, thereby generating more accurate fused point features Fp. Specifically, by constructing a query-key-value (QKV) generation network, this module captures the complex relationship between point features and grid features, significantly improving the feature expression ability of the model.

[0046] The implementation of the module relies on the processing of the grid feature fg, and uses the Sequential container to unify the construction of the generation networks for Q, K, and V. Inside each container, the linear transformation of the features is achieved through the 1×1 convolutional layer. The role of the 1×1 convolution is that it only performs calculations in the channel dimension and does not change the spatial dimensions (such as height and width) of the features. Therefore, it is particularly suitable for efficiently processing the interactions between feature channels. Subsequently, the Batch Normalization layer is applied to the output of the convolution to stabilize the network training process and accelerate convergence. Finally, the ReLU activation function is used to perform a non-linear transformation on the normalized features to enhance the feature expression ability of the model.

[0047] After generating the query vector Q and the key vector K, the module calculates the correlation matrix through the dot product operation of the two. . This correlation matrix reflects the global relationships between different positions in the input features. For example, for a certain spatial point , the grid it belongs to is , and by calculating the global and local relationship terms , , , the attention weights for each face are generated , , . These relationship terms define the geometric relationships of the query point on each grid face and its influence within the global range.

[0048] Specifically, the calculation formula for the attention weights is: ; ; ; where and represent the embedding functions of the feature itself and the global relationship respectively, and are the weight mappings composed of the 1×1 convolution and the batch normalization layer, which are responsible for converting the feature channel size to the target dimension.

[0049] Through the above process, the module obtains the attention weights for each face, and uses these weights to perform a weighted sum on the corresponding grid features to generate the fused point feature of the query point: ; where t represents the three faces of the grid, namely the x, y, and z planes.

[0050] To ensure the smoothness of the fused point features in the geometric space, the module adopts the bilinear interpolation method to further optimize the continuity of the grid features in space. This design can effectively solve the information loss problem caused by the sparsity of point cloud data and demonstrate excellent performance in the reconstruction of complex shapes and sharp edges.

[0051] Step 104: For any query point, obtain the embedded feature according to at least one fused point feature, and generate the discretized embedded feature according to the embedded feature and the preset multi-head codebook.

[0052] In the 3D surface reconstruction method of the present invention, the core objective of step 104 is to perform embedded encoding on the fused point features of the query point and generate the discretized embedded feature in combination with the preset multi-head codebook. The purpose of this is to enhance the cross-object generalization ability of the model, enabling it to adapt to the feature expression requirements of different categories of objects, and at the same time optimizing the ability to capture details in the 3D reconstruction process. The introduction of the multi-head codebook can map the high-dimensional continuous feature representation to multiple compact prior subspaces by generating discretized features, thereby enhancing the robustness and expressiveness of the model when dealing with complex geometries.

[0053] In a possible implementation manner, the preset multi-head codebook includes a first preset number of preset sub-codebooks. Refer to Figure 3 , Figure 3 is the third flow diagram of the 3D reconstruction method provided by the present invention. As Figure 3 shown, step 104 specifically includes steps 301-305: Step 301: Determine a second preset number of neighboring points according to the Euclidean distance between the query point and each spatial point.

[0054] In the 3D surface reconstruction method of the present invention, the purpose of step 301 is to screen out a certain number of neighboring points through the relationship between the query point and other spatial points in the point cloud, so as to capture the local geometric features of the query point and provide basic support for the subsequent generation of embedded features. Point cloud data is inherently discrete and sparse. Directly processing the entire point cloud data will lead to too high computational complexity and is not conducive to effectively extracting local details. Therefore, by screening neighboring points, the computational burden can be reduced, and at the same time, focusing on the local geometric relationship of the query point provides key support for accurately describing the 3D shape.

[0055] In a specific implementation, with the query point q as the center, the neighboring point set is determined by calculating the Euclidean distance between it and all other points in the point cloud. To ensure that the selected neighboring points can fully cover the local spatial distribution of the query point while not introducing excessive computational redundancy, a fixed second preset quantity k is set. This quantity is set based on experimental experience and the need to balance computational efficiency and geometric information capture, usually taking a value between 10 and 30, and the specific value is dynamically adjusted according to the sparsity of the input point cloud and task requirements.

[0056] For each spatial point (where i is the point index in the point cloud), calculate its Euclidean distance from the query point q , and the calculation formula is: ; where represents the Euclidean distance. According to the calculation results, all points are sorted in ascending order of distance, and the k points with the smallest distance are selected as neighboring points. The indices and features of these neighboring points are recorded for use in the subsequent embedded feature generation process.

[0057] This screening process ensures that the query point can construct a neighborhood with a clear geometric relationship within its local space, so as to accurately capture the local geometric information through the features of these neighboring points in subsequent steps. In addition, to handle special cases (such as overly sparse point cloud data or the query point being located at the boundary), the algorithm also sets a minimum neighborhood search radius to ensure that the selected neighboring points can provide sufficient geometric coverage.

[0058] Step 302: Obtain the embedded feature according to the fused point features of all neighboring points.

[0059] In a possible implementation manner, the embedded feature in step 302 is specifically obtained through the following formula: ; ; In the formula, is the embedded feature, is the concatenation operation, is the joint feature, is the fused point feature of the i-th neighboring point, q is the query point, MLP is the multi-layer perceptron, and k is the second preset quantity.

[0060] In the specific implementation process, first, use the neighboring point set screened in step 301 and its corresponding fused point features , combine the position information of the query point q and the relative position relationship with the neighboring points, and generate the joint featureThis combined feature is represented by the following formula: ; where is the fused point feature of the i-th neighboring point, represents the relative position information between the query point and the neighboring points. Through the non-linear transformation of the MLP, these input information are mapped to a high-dimensional feature space, so that the generated combined feature can simultaneously characterize the local geometric relationship and spatial distribution characteristics of the query point.

[0061] Subsequently, all the combined features are integrated into the embedded feature of the query point through a concatenation operation , and the formula is: ; This integration method can fully express the contribution of each neighboring point to the geometric relationship of the query point, while maintaining the integrity of the query point features. The generation of the embedded feature fully considers the correlation and tightness of the local geometric features in design, so that it has strong description ability.

[0062] The embedded feature generated through the above operations not only contains the local geometric information of the query point, but also significantly enhances the overall expression ability of the point cloud features. Experimental results show that the introduction of this embedded feature significantly improves the 3D reconstruction accuracy of the model in complex geometric structures and sharp edge regions, while enhancing the generalization ability of the model in multi-category scenarios. In addition, the feature generation mechanism based on the MLP enables the embedded feature to flexibly adapt to the feature distribution requirements of different category objects, providing high-quality input for subsequent multi-head codebook operations and implicit field construction.

[0063] Step 303: Divide the embedded feature into a first preset number of parts.

[0064] In the 3D surface reconstruction method of the present invention, the core objective of step 303 is to divide the embedded feature of the query point into several parts to adapt to the sub-codebook operations of the subsequent multi-head codebook. The design of this step aims to enhance the diversity and adaptability of the embedded features, enabling them to better express the geometric characteristics of the query point in different dimensions, while making full use of the distributed coding ability of the multi-head codebook structure. Through feature division, the embedded features of the query point can be processed more finely to improve the generalization ability of the reconstruction model on multi-category shapes.

[0065] In the specific implementation process, the embedded feature is generated by the previous step 302 and has the shape of a high-dimensional vector. To fully utilize the distributed characteristics of the multi-head codebook, first, according to the first preset number h of sub-codebooks of the preset multi-head codebook, the embedded feature is equally divided into h parts, and the dimension of each part of the feature is . This equal division operation ensures that each part of the feature has the same dimension, facilitating subsequent matching with the sub-codebook.

[0066] Step 304: For each part, input it into the corresponding preset sub-codebook, and by measuring the distance between this part and each code vector in the sub-codebook, select the code vector closest to the part to generate the discretized embedded sub-feature corresponding to the part.

[0067] In the three-dimensional surface reconstruction method of the present invention, the core objective of step 304 is to match each feature part with the sub-codebook in the preset multi-head codebook. By measuring the distance between this feature part and each code vector in the sub-codebook, select the closest code vector as the discretized embedded sub-feature. The design of this process aims to map the continuous embedded feature to the discretized feature space, thereby reducing the storage and computational complexity. At the same time, prior information is introduced for the representation of the three-dimensional shape to improve the consistency and discriminability of feature expression.

[0068] In specific implementation, the embedded feature of the query point has been divided into h parts in step 303. The multi-head codebook contains h sub-codebooks , and each sub-codebook consists of M discrete code vectors, denoted as . For each feature part , by calculating the Euclidean distance between it and all code vectors in the corresponding sub-codebook, the closest code vector is determined. The specific calculation formula is: ; where is the index of the best code vector matched by the feature part in the sub-codebook, and represents the m-th code vector in the sub-codebook .

[0069] After the matching is completed, each feature part is mapped to a discrete code vector in the sub-codebook. This process effectively converts the high-dimensional continuous feature into a low-dimensional discrete feature, making subsequent feature processing and storage more efficient. In addition, to ensure that the sub-codebook can adapt to the diversity of different categories of objects, during the matching process, not only the distance metric between the feature and the code vector is utilized, but also the sub-codebook is optimized by combining a dynamic update mechanism, thereby improving the adaptability of the discretized feature.

[0070] Step 305: Connect the discretized embedded sub - features of each part into an overall discretized embedded feature.

[0071] In the 3D surface reconstruction method of the present invention, the core objective of Step 305 is to connect the discretized embedded sub - features corresponding to each feature part into an overall discretized embedded feature. The design intention of this process is to integrate the matching results of each sub - codebook, and uniformly represent the multi - dimensional discrete information generated by the multi - head codebook as an efficient and compact global feature, providing a unified input representation for subsequent implicit field modeling and 3D surface reconstruction.

[0072] In specific implementation, in Step 304, each feature part has obtained the discretized embedded sub - feature through matching with its corresponding sub - codebook . Step 305 sequentially splices these discretized embedded sub - features into an overall discretized embedded feature . The formula for this splicing operation can be expressed as: ; where represents the feature splicing operation.

[0073] The splicing operation connects all sub - features into a high - dimensional vector in the channel dimension, ensuring the integrity of the discretized features while retaining the independence of each sub - feature. This way can not only effectively integrate the feature information of different sub - codebooks, but also significantly improve the expression ability of the discretized features for global and local geometric relationships. Through the unified representation form, the model can more efficiently fuse information from different dimensions when dealing with complex geometric shapes, thereby improving the accuracy and detail retention ability of 3D reconstruction.

[0074] In addition, to ensure that the spliced discretized features have strong expression ability in subsequent processing, Step 305 also optimizes the splicing process. For example, by dynamically adjusting the number of sub - codebooks and the dimension of the code vectors, it is ensured that the finally spliced can adapt to the feature requirements of different categories of objects, thus showing higher robustness in complex scenes.

[0075] After integration, the overall discretized embedded feature is used for the construction of the implicit field in the subsequent steps.

[0076] Step 105: Obtain the implicit field for characterizing the 3D shape based on the embedded features and discretized embedded features of all query points.

[0077] In the 3D surface reconstruction method of the present invention, the core objective of step 105 is to construct an implicit field for characterizing the 3D shape based on the embedded features of all query points and the discretized embedded features An implicit field is a flexible and efficient representation in current 3D reconstruction techniques. By modeling the positions of spatial points with continuous functions, it can naturally describe complex geometric structures, including curved surfaces, boundaries, and topological changes. By combining embedded features and discretized features, this step aims to capture both global structural information and local geometric details simultaneously, providing high-quality data support for subsequent 3D surface extraction.

[0078] In a possible implementation, referring to Figure 4 Figure 4 is the fourth schematic diagram of the process of the 3D reconstruction method provided by the present invention. As shown in Figure 4 step 105 specifically includes step 401: Step 401: Concatenate the fused point features and discretized embedded features of all query points and input them into a multi-layer perceptron to obtain the implicit field output by the multi-layer perceptron.

[0079] In the specific implementation process, the fused point feature is generated in step 302, which captures the local geometric characteristics and relative spatial relationships of query points through the interaction between points and grid features; the discretized embedded feature is generated through multi-head codebook mapping in step 305, providing the representational ability of query points in the global geometric context. In step 401, first, through the concatenation operation, and are integrated into a unified feature representation: where is the feature concatenation operation, ensuring the integrity and independence of the two types of features when input.

[0080] The concatenated feature is used as the input and fed into a multi-layer perceptron (MLP) network. The MLP consists of several fully connected layers, each followed by a ReLU activation function to introduce non-linearity and enhance the feature representational ability. The MLP network performs layer-by-layer feature transformation on and finally outputs the implicit field of the query points, which is used to represent the spatial relationship of the query point q relative to the 3D shape surface. The output formula of the implicit field is: ​To improve the model's expressive power, the MLP network adopts a weight sharing mechanism, which can efficiently reuse network parameters when processing different query points. In addition, to avoid numerical instability caused by uneven feature distribution, a Batch Normalization layer is added after each fully connected operation to normalize the feature distribution and accelerate the network's convergence.

[0081] During the training process, by optimizing the error between the output value of the implicit field and the target value, the network can gradually learn the geometric characteristics and global topological structure of the point cloud. Common training losses include occupancy error or distance field error, and the specific choice depends on the design type of the implicit field. For example, in the implicit representation based on the distance field, the optimization goal is usually to minimize the difference between the implicit field value and the closest distance from the query point to the surface.

[0082] Through this implicit field generation process, the local geometric details and global structure information of the query points are effectively fused into a unified spatial field representation, which can naturally handle complex geometric shapes and sharp edge regions.

[0083] Step 106: According to the implicit field, construct a triangular mesh model through an edge-based isosurface extraction method to complete the three-dimensional surface reconstruction.

[0084] In the three-dimensional surface reconstruction method of the present invention, the core objective of step 106 is to convert the implicit field into a triangular mesh through an edge-based isosurface extraction method, thereby completing the three-dimensional surface reconstruction. The design of this process aims to efficiently generate a three-dimensional shape with a complete structure and rich details, while solving the ambiguity problem of traditional mesh extraction methods on sharp edges and complex geometric shapes. Through an innovative edge intersection detection mechanism, step 106 can accurately extract the intersection position of the surface and the implicit field, providing technical support for the generation of complex three-dimensional surfaces.

[0085] In a possible implementation manner, refer to Figure 5 , Figure 5 is the fifth schematic flow chart of the three-dimensional reconstruction method provided by the present invention. As Figure 5 shown, step 106 specifically includes steps 501-505: Step 501: Calculate the implicit field values and normal directions of the eight vertices of each cubic unit of the implicit field respectively.

[0086] In the three-dimensional surface reconstruction method of the present invention, the core objective of step 501 is to calculate the implicit field values and normal directions of the eight vertices of each cube unit of the implicit field. This step lays the foundation for subsequent edge intersection detection and triangle mesh generation. By accurately calculating the field values and normal directions of the vertices, the accuracy of mesh generation can be effectively improved, especially in complex geometric shapes and sharp edge regions. The necessity of doing so lies in that the implicit field value reflects the spatial position of the vertex relative to the three-dimensional surface, and the normal direction further provides local orientation information of the surface geometry. The combination of the two is the key to accurately extracting the surface intersection points.

[0087] In the specific implementation process, the implicit field is generated by step 401 and is represented as a continuous function that can return its field value for any query point and normal direction . In the cube unit, the position of each vertex is known. The implicit field queries are sequentially performed on these eight vertices to calculate their field values and normal directions: ; ; where is the implicit field value of vertex , and is its corresponding normal direction.

[0088] The sign of the implicit field value indicates the relationship between the vertex and the inside and outside of the surface (a positive value indicates that the vertex is outside the surface, a negative value indicates that it is inside the surface, and a zero value indicates that it is on the surface). The normal direction is obtained by calculating the gradient of the implicit field and is used to characterize the local geometric direction of the surface at this point.

[0089] To ensure the stability and accuracy of numerical calculations, the automatic differentiation technique is used for the calculation of the implicit field and its gradient, and the implicit field is analytically expressed by combining an efficient multi-layer perceptron (MLP) network structure. In the implementation process, batch calculation is used to calculate the field values and normal directions of all vertices of each cube unit simultaneously to improve the calculation efficiency.

[0090] Step 502: Determine the intersecting edges of each cube unit according to the respective implicit field values and respective normal directions.

[0091] In a possible implementation manner, step 502 specifically includes the following steps: For each group of adjacent vertices of each cube unit, when the signs of the implicit field values of the adjacent vertices are different and the normal directions of the adjacent vertices satisfy the preset angle threshold condition, it is determined that the edge between the current adjacent vertices is an intersecting edge.

[0092] In the three-dimensional surface reconstruction method of the present invention, the core objective of step 502 is to determine the intersecting edges of each cubic unit based on the vertex field values and normal directions of the implicit field. The design of this step aims to identify the intersection positions between the three-dimensional shape surface and the cube edges, thereby laying a foundation for the subsequent generation of the triangular mesh. Through high-precision edge intersection detection, not only can the geometric integrity of the generated mesh be ensured, but also the blurring phenomenon can be effectively avoided in sharp edges and complex structures.

[0093] In the specific implementation process, each cubic unit is composed of eight vertices, and the implicit field values and normal directions of each vertex have been calculated in step 501. For each edge within the cube (formed by connecting two adjacent vertices), it is determined whether it is an intersecting edge through the following conditions: 1. Implicit field value sign detection: The signs of the implicit field values and of adjacent vertices. If the signs are opposite, it indicates that the surface of the cube passes through this edge; if the signs are the same, it indicates that this edge is completely inside or outside the surface and is not an intersecting edge.

[0094] 2. Normal direction angle detection: Whether the normal directions and of adjacent vertices meet the preset angle threshold condition. Specifically, when the angle between the normal directions meets the preset angle threshold condition, that is, less than the angle threshold (usually taken between 30° - 60°), it can be considered that the vertex normal directions are consistent, which is conducive to the accurate positioning of the surface intersection points.

[0095] To improve the detection efficiency, parallel computing is used to simultaneously perform intersection detection on the 12 edges of each cubic unit. The detection results are stored in binary form, and the status (intersecting or non-intersecting) of each edge is recorded as a boolean value for quick lookup during subsequent intersection point calculation and triangular mesh generation.

[0096] Through this intersecting edge detection process, the surface penetration information of each cubic unit is accurately determined. Compared with the traditional strategy based on a single determination of the field value sign, this method significantly reduces the misjudgment rate by combining the geometric consistency constraint of the normal direction, especially performing excellently when dealing with sharp edges and complex shapes.

[0097] Step 503: Use interpolation to obtain the initial intersection position for the implicit field values of the endpoints of each intersecting edge, and correct the initial intersection position according to the normal directions of the endpoints of the intersecting edge to obtain the corrected intersection position.

[0098] In the 3D surface reconstruction method of the present invention, the core objective of step 503 is to perform interpolation calculations on the implicit field values of the endpoints of each intersecting edge to determine the initial intersection position, and to correct the initial intersection position in combination with the normal directions of the edge endpoints, thereby generating a more accurate intersection position. The design of this step aims to solve the possible numerical deviation problems in the implicit field value calculation process, and at the same time enhance the geometric accuracy of the edge intersection position, providing high-precision input for the subsequent triangular mesh generation.

[0099] During the specific implementation process, the endpoints of the intersecting edge and have been determined as an intersecting edge in step 502 by the implicit field value sign and the normal direction. To determine the intersection position of the surface and the intersecting edge, first, linear interpolation is performed through the implicit field values and to calculate the initial intersection position . The interpolation formula is: ; where and are the implicit field values of the two endpoints of the intersecting edge respectively, representing the distance direction and magnitude of the two points relative to the 3D surface.

[0100] Linear interpolation can quickly provide an initial estimate of the intersection position. However, due to possible numerical errors in the calculation of the implicit field values, the simple interpolation result may lead to an inaccurate intersection point position. To further improve the position accuracy, the initial intersection position is corrected in combination with the normal directions and of the intersecting edge endpoints. The correction process is achieved by optimizing the distribution consistency of the intersection points in the surface normal direction. The corrected intersection position is calculated by the formula: ; where is a coefficient for adjusting the correction amplitude, usually set to a small value (such as between 0.01 and 0.1) according to experimental experience to balance the accuracy and stability of the position adjustment.

[0101] The corrected intersection position can more accurately reflect the intersection position of the surface and the intersecting edge, and at the same time avoid the reconstruction result deviation caused by interpolation errors. In addition, to ensure the calculation efficiency, the correction process adopts a batch processing method, allowing the above calculations to be performed simultaneously on multiple intersecting edges within a cubic cell, thereby significantly improving the overall processing speed.

[0102] Step 504: Triangulate all the corrected intersection positions in combination with a preset lookup table to generate the corresponding local triangular mesh.

[0103] In the 3D surface reconstruction method of the present invention, the core objective of step 504 is to triangulate the corrected intersection positions through a lookup table, thereby generating the local triangular meshes of each cube unit. The necessity of this process lies in that the implicit field calculation generates the field values and intersection point positions of the cube vertices, but this information only describes the relationship between the surface and the cube edges and does not form a clear triangular mesh. Through the lookup table method, the triangulation scheme can be efficiently determined to ensure that the topological structure of the local mesh is consistent with the intersection state, providing an accurate geometric basis for 3D surface reconstruction.

[0104] In the specific implementation, first define the occupancy state for each cube unit, which represents the sign of the field value of the vertex (for example, a positive value represents outside the surface, and a negative value represents inside the surface). Based on the intersection edge information generated in step 502, the intersection point state of each edge has also been marked. For each pair of vertices i and j within the cube unit, a cost function L is defined by calculating the exclusive OR operation . This function is used to measure the suitability of the current triangulation scheme, and the specific expression is: ; where respectively represent the intersection point states of the edges connected by vertices i and j, and are the occupancy states of the vertices respectively.

[0105] By calculating the cost function L for all triangulation schemes in the lookup table, the scheme with the minimum cost is selected as the best triangulation scheme for the current cube unit. The lookup table is constructed based on the rules of the Marching Cubes algorithm and contains 256 possible vertex occupancy states, and each state corresponds to several vertex connection patterns of the triangular meshes. The use of the lookup table enables complex triangulation problems to be directly solved through efficient queries without having to calculate all possible geometric connections one by one.

[0106] After the triangulation scheme is selected, the corrected intersection positions are used as the vertices of the triangular mesh, and the local triangular mesh is generated according to the vertex connection relationship defined in the lookup table. During the generation process, through consistent vertex indexing and geometric verification, the continuity and geometric rationality of the triangular mesh are ensured.

[0107] Through this triangulation processing step, the implicit field information within each cube unit is successfully converted into a clear geometric representation, and the generated local triangular mesh can accurately describe the interaction relationship between the surface and the cube.

[0108] Step 505: Repeat the above steps for all cube units within the target area, and integrate all local triangular meshes into a triangular mesh model to complete the 3D surface reconstruction.

[0109] In the specific implementation process, the triangular meshes of each cube unit have been triangulated in Step 504 to form a local set of triangles. These local meshes are organized using vertex indices with the corrected intersection positions of the intersecting edges as vertices. To piece these local meshes together into an overall mesh, the following key issues need to be addressed: vertex sharing, boundary stitching, and topological structure consistency.

[0110] First, regarding the problem of mesh vertex sharing, to ensure the consistency of the common vertices of adjacent cube units, a global vertex index mapping mechanism is adopted. Specifically, a globally unique index is assigned to each vertex. When the vertex positions of two or more units are the same, their index representations are unified. This mechanism is implemented through a global lookup table. Each time a new vertex is added, it is checked whether it already exists in the global table. If it exists, its index is reused; otherwise, a new index is created. This process effectively avoids possible vertex redundancy or duplication problems during the mesh piecing process.

[0111] Second, to handle boundary stitching, the global index table is used to fuse the edges of adjacent cube units. By matching the intersecting edges and their vertex indices of adjacent units, the shared edges are merged into a single representation, thus ensuring the continuity of the pieced-together mesh at the boundary. In addition, to further enhance the smoothness of the boundary, a geometric correction algorithm is introduced to fine-tune the positions of the boundary vertices to eliminate possible cracks or minor misalignments.

[0112] Finally, to ensure the consistency of the topological structure, all generated triangular meshes are traversed to detect and optimize areas with overlap or intersection. The specific operations include removing duplicate triangles, correcting the vertex connection order of intersecting meshes, and simplifying locally over-dense mesh areas. These processing operations are optimized and calculated using geometric neighborhood information to ensure the structural stability and rendering quality of the mesh.

[0113] Through the above piecing steps, the local meshes of all cube units within the target area are successfully integrated into a complete triangular mesh model.

[0114] To verify the effectiveness of the proposed global-local relationship attention three-dimensional reconstruction method based on implicit representation, this solution trains and tests the method on the ShapeNet dataset. The ShapeNet dataset is a large-scale three-dimensional shape database containing over 50,000 high-quality three-dimensional models, covering multiple categories from furniture to transportation, from electronic devices to animals. The dataset is divided into a training set, a validation set, and a test set, providing rich resources for three-dimensional shape analysis, generation, and understanding. ShapeNet covers a variety of indoor layouts and object configurations, providing strong support for tasks such as indoor scene understanding, navigation, and three-dimensional reconstruction. In the ShapeNet dataset, the sizes and dimensions of target objects in different categories vary, covering small objects such as cups to large objects such as sofas and beds. During the training process, this solution performs corresponding preprocessing for each category's characteristics to ensure the consistency and stability of the input data, and reports the method's performance on the validation dataset.

[0115] Implement the model of this solution in Pytorch and use the Adam optimizer. During the training process, the learning rate in the first stage is 10 -4 , and the learning rate in the fine-tuning stage is 10 -6 . The depth of the U-NET-like encoder of this solution is 4, and this solution does not downsample or upsample the grid features of the two top layers. The radius r for searching relative points is set to 0.08, and the boundary m is set to 2.0. In the visualization stage, this solution uses an edge-based isosurface extraction algorithm to extract the grid for visualization. This solution uses Chamfer Distance (CD), Iou, and F-score as the evaluation metrics of this solution.

[0116] The functions and calculation formulas of the three metrics are as follows: CD: Used to calculate the similarity between two point clouds, measuring the average minimum distance between the predicted point cloud and the ground truth point cloud.

[0117] ; where P is the predicted point cloud and G is the ground truth point cloud, is the Euclidean distance between two points.

[0118] Iou: Measures the ratio of the intersection and union of two three-dimensional volumes, usually used to evaluate the overlapping degree of two three-dimensional objects.

[0119] ; where Vpre is the three-dimensional volume of the reconstruction prediction and Vgt is the ground truth three-dimensional volume.

[0120] FS: Considering both precision and recall, it is used to measure the matching situation between two point clouds in 3D reconstruction.

[0121] ; where is the recall rate of points, is the precision rate of points.

[0122] To verify the effectiveness and integrity of this application, the method of this solution was verified on the ShapeNet dataset. Under the condition of ensuring that other conditions are the same, the global-local relationship attention mechanism module and the edge-based isosurface extraction module were removed from the global-local relationship attention 3D reconstruction method based on implicit representation for ablation experiments. Table 1 shows the experimental results of the global-local relationship attention mechanism in the method proposed in the present invention and the experimental effects of the edge-based isosurface extraction algorithm. The experimental results show that the global-local attention mechanism proposed in this solution significantly improves the reconstruction accuracy by improving the sampling accuracy, and there are varying degrees of improvement in indicators such as Iou, CD, and FS. At the same time, the edge-based isosurface extraction algorithm of this solution can extract higher-quality triangular meshes, and the extracted meshes perform better in both specific details and sharp edges.

[0123] See Table 1 shown below. Table 1 is the ablation experiment analysis table under the ShapeNet dataset: Table 1

[0124] It can be fully seen from the experimental results in Table 1 the effectiveness of the present invention.

[0125] Refer to Figure 6 , Figure 6 is the schematic diagram of the visualization results of different categories on the ShapeNet dataset provided by the present invention. Table 2 and Figure 6 show the comparison between the proposed method and other existing methods.

[0126] See Table 2 shown below. Table 2 is the test result table on the ShapeNet dataset: Table 2

[0127] To evaluate the effectiveness of the reconstruction method of this solution, the baselines for comparison include ConvONet, POCO, and ALTO. Under the same experimental settings, the method of this solution is superior to the existing methods in terms of accuracy and visual effects. These improvements are not only reflected in the quantitative experimental results but also verified through qualitative visualization, such as Figure 6。The method of this solution can reconstruct a more detailed and smooth surface, thus achieving a better visual effect.

[0128] As can be seen from Table 2, the method of this solution has varying degrees of improvement compared to the baseline in different metrics. In terms of the IOU metric, it has increased by 6.6%, 3.6%, and 1.3% compared to the existing methods ConvONET, POCO, and ALTO respectively. In terms of the CD metric, it has decreased by 36%, 20%, and 9% compared to the existing methods ConvONET, POCO, and ALTO respectively. In terms of the FS metric, it has increased by 5.8%, 2.3%, and 1.1% compared to the existing methods ConvONET, POCO, and ALTO respectively.

[0129] In addition, from the visualization results, it can be seen that the reconstruction results of this application are more prominent in terms of reconstruction details. A smoother surface can be reconstructed for objects of different categories, maximizing the restoration of the original details of the objects. These advantages make the method of this solution have higher practical value and potential in practical applications.

[0130] Refer to Figure 7 , Figure 7 is a schematic structural diagram of the 3D reconstruction system provided by the present invention. The system includes: A first processing module, configured to extract the initial point features of each spatial point from the point cloud data; The first processing module is further configured to input the respective initial point features into a transformation network that combines points and meshes based on an attention mechanism to obtain the mesh features of the mesh where each spatial point is located; The first processing module is further configured to obtain the fused point features of each spatial point based on the global-local relationship attention mechanism and the respective mesh features; A second processing module, configured to, for any query point, obtain an embedded feature according to at least one fused point feature, and generate a discretized embedded feature according to the embedded feature and a preset multi-head codebook; The second processing module is further configured to obtain an implicit field for characterizing the 3D shape according to the embedded features and the discretized embedded features of all query points; The second processing module is further configured to construct a triangular mesh model by an edge-based isosurface extraction method according to the implicit field to complete the 3D surface reconstruction.

[0131] In a possible implementation manner, the first processing module is further configured to: Construct a query-key-value generation network based on the global-local relationship attention mechanism; Input the mesh features of each face of the mesh where the spatial point is located into the query-key-value generation network to correspondingly obtain the attention weights of each face; According to the attention weights of each face, the grid features of each face are weighted and fused to obtain fused point features.

[0132] In a possible implementation manner, the second processing module is further configured to: Determine a second preset number of neighboring points according to the Euclidean distance between the query point and each spatial point; Obtain embedded features according to the fused point features of all neighboring points; Divide the embedded features into a first preset number of parts; For each part, input it into the corresponding preset sub-codebook, and by measuring the distances between the part and each code vector in the sub-codebook, select the code vector closest to the part to generate a discretized embedded sub-feature corresponding to the part; Connect the discretized embedded sub-features of each part into an overall discretized embedded feature.

[0133] In a possible implementation manner, the second processing module is further configured to obtain embedded features through the following formula: ; ; In the formula, is the embedded feature, is the concatenation operation, is the joint feature, is the fused point feature of the i-th neighboring point, q is the query point, MLP is the multi-layer perceptron, and k is the second preset number.

[0134] In a possible implementation manner, the second processing module is further configured to splice the fused point features and the discretized embedded features of all query points and input them into the multi-layer perceptron to obtain an implicit field output by the multi-layer perceptron.

[0135] In a possible implementation manner, the second processing module is further configured to: Calculate the implicit field values and normal directions of the eight vertices of each cubic unit of the implicit field respectively; Determine the intersecting edges of each cubic unit according to the respective implicit field values and the respective normal directions; Use interpolation to obtain the initial intersection positions of the endpoints of each intersecting edge for the implicit field values, and correct the initial intersection positions according to the normal directions of the endpoints of the intersecting edges to obtain the corrected intersection positions; Combine with a preset lookup table to perform triangulation processing on all corrected intersection positions to generate corresponding local triangular meshes; Repeat the above steps for all cubic units within the target area, and integrate all local triangular meshes into a triangular mesh model to complete the three-dimensional surface reconstruction.

[0136] In a possible implementation, the second processing module is further configured to, for each group of adjacent vertices of each cube unit, determine that the edge between the current adjacent vertices is an intersecting edge when the signs of the implicit field values of the adjacent vertices are different and the normal directions of the adjacent vertices satisfy a preset included angle threshold condition.

[0137] It should be noted that the three-dimensional reconstruction system provided by the present invention can execute the three-dimensional reconstruction method of any of the above embodiments during specific operation, which will not be elaborated in this embodiment.

[0138] Figure 8 is a schematic structural diagram of an electronic device provided by the present invention, as Figure 8 shown, the electronic device may include: a processor 810 (processor), a communication interface 820 (Communications Interface), a memory 830 (memory), and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 complete mutual communication through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the three-dimensional reconstruction method, and the method includes: extracting the initial point features of each spatial point from the point cloud data; inputting each initial point feature into a transformation network that combines points and grids based on the attention mechanism to obtain the grid features of the grid where each spatial point is located; obtaining the fused point features of each spatial point based on the global-local relationship attention mechanism and each grid feature; for any query point, obtaining an embedded feature according to at least one fused point feature, and generating a discretized embedded feature according to the embedded feature and a preset multi-head codebook; obtaining an implicit field for characterizing the three-dimensional shape according to the embedded features and discretized embedded features of all query points; and constructing a triangular mesh model through an edge-based isosurface extraction method according to the implicit field to complete the three-dimensional surface reconstruction.

[0139] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0140] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the three-dimensional reconstruction method provided in each of the above embodiments. The method includes: extracting initial point features of each spatial point from point cloud data; inputting each initial point feature into a transformation network that combines points and grids based on an attention mechanism to obtain grid features of the grid where each spatial point is located; obtaining fused point features of each spatial point based on a global-local relationship attention mechanism and each grid feature; for any query point, obtaining an embedded feature according to at least one fused point feature, and generating a discretized embedded feature according to the embedded feature and a preset multi-head codebook; obtaining an implicit field for characterizing the three-dimensional shape according to the embedded features and discretized embedded features of all query points; and constructing a triangular mesh model through an edge-based isosurface extraction method according to the implicit field to complete three-dimensional surface reconstruction.

[0141] In another aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the three-dimensional reconstruction method provided in the above embodiments. The method includes: extracting initial point features of each spatial point from the point cloud data; inputting the respective initial point features into a transformation network that combines points and grids based on an attention mechanism to obtain grid features of the grid where each spatial point is located; obtaining fused point features of each spatial point based on a global-local relationship attention mechanism and the respective grid features; for any query point, obtaining an embedded feature according to at least one fused point feature, and generating a discretized embedded feature according to the embedded feature and a preset multi-head codebook; obtaining an implicit field for characterizing the three-dimensional shape according to the embedded features and discretized embedded features of all query points; and constructing a triangular mesh model through an edge-based isosurface extraction method according to the implicit field to complete the three-dimensional surface reconstruction.

[0142] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0143] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.

[0144] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A three-dimensional reconstruction method, characterized in that: include: Extract the initial point features of each spatial point from the point cloud data; Inputting each of the initial point features into a transformation network combining points and grids based on an attention mechanism to obtain grid features of the grid where each spatial point is located; Based on the global-local relationship attention mechanism and each of the grid features, a fusion point feature of each spatial point is obtained; For any query point, an embedded feature is obtained according to at least one of the fusion point features, and a discretized embedded feature is generated according to the embedded feature and a preset multi-head codebook; Obtaining an implicit field for characterizing a three-dimensional shape according to the embedded features and the discretized embedded features of all the query points; According to the implicit field, a triangular mesh model is constructed by an edge-based isosurface extraction method to complete three-dimensional surface reconstruction.

2. The three-dimensional reconstruction method according to claim 1, characterized in that: The fusion point feature of each spatial point is obtained based on the global-local relationship attention mechanism and each of the grid features, specifically including: Based on the global-local relationship attention mechanism, a query-key-value generation network is constructed; Inputting the grid features of each face of the grid where the spatial point is located into the query-key-value generation network, and obtaining the attention weight of each face accordingly; According to the attention weight of each face, the mesh features of each face are weighted fused to obtain the fusion point features.

3. The three-dimensional reconstruction method according to claim 1, characterized in that: The preset multi-head codebook includes a first preset number of preset sub-codebooks; the embedded feature is obtained according to at least one of the fusion point features, and a discretized embedded feature is generated according to the embedded feature and the preset multi-head codebook, specifically including: Determining a second preset number of neighboring points according to the Euclidean distance between the query point and each of the spatial points; Obtaining the embedded feature according to the fused point features of all the neighboring points; dividing the embedded feature into a first predetermined number of portions; For each part, input it into a corresponding preset sub-codebook, and select a code vector closest to the part by measuring the distance between the part and each code vector in the sub-codebook, so as to generate a discretized embedded sub-feature corresponding to the part; The discretized embedded sub-features of each part are connected into the discretized embedded features of the whole.

4. The three-dimensional reconstruction method according to claim 3, characterized in that: The step of obtaining the embedded feature according to the fusion point features of all the neighboring points specifically includes: The embedded features are obtained by the following formula: ; ; In the formula, For the embedded feature, For the connection operation, is the joint feature, is the fusion point feature of the i-th neighboring point, q is the query point, MLP is a multi-layer perceptron, and k is a second preset number.

5. The three-dimensional reconstruction method according to claim 1, characterized in that: The step of obtaining an implicit field for characterizing a three-dimensional shape according to the embedded features and the discretized embedded features of all the query points specifically includes: The fused point features and the discretized embedded features of all the query points are concatenated and input into a multi-layer perceptron to obtain an implicit field output by the multi-layer perceptron.

6. The three-dimensional reconstruction method according to claim 1, characterized in that: The method of constructing a whole triangle mesh according to the implicit field by an edge-based isosurface extraction method to complete the three-dimensional surface reconstruction specifically includes: Calculating implicit field values ​​and normal directions of eight vertices of each cubic unit of the implicit field respectively; Determine the intersection edge of each of the cube units according to each of the implicit field values ​​and each of the normal directions; The implicit field value of each endpoint of the intersecting edge is interpolated to obtain an initial intersection position, and the initial intersection position is corrected according to the normal direction of the endpoint of the intersecting edge to obtain a corrected intersection position; triangulate all the corrected intersection positions in combination with a preset lookup table to generate corresponding local triangular meshes; The above steps are repeated for all cubic units in the target area, and all the local triangular meshes are integrated into the triangular mesh model to complete the three-dimensional surface reconstruction.

7. The three-dimensional reconstruction method according to claim 6, characterized in that: Determining the intersection edge of each of the cube units according to each of the implicit field values ​​and each of the normal directions specifically includes: For each group of adjacent vertices of each of the cubic units, when the implicit field values ​​of the adjacent vertices have different signs and the normal directions of the adjacent vertices meet a preset angle threshold condition, the edge between the current adjacent vertices is determined to be the intersecting edge.

8. A three-dimensional reconstruction system, characterized in that: include: The first processing module is used to extract the initial point feature of each spatial point from the point cloud data; The first processing module is further used to input each of the initial point features into a transformation network combining points and grids based on an attention mechanism to obtain grid features of the grid where each spatial point is located; The first processing module is further used to obtain a fusion point feature of each spatial point based on a global-local relationship attention mechanism and each of the grid features; A second processing module is used to obtain an embedded feature for any query point according to at least one of the fusion point features, and generate a discretized embedded feature according to the embedded feature and a preset multi-head codebook; The second processing module is further used to obtain an implicit field for characterizing a three-dimensional shape according to the embedded features and the discretized embedded features of all the query points; The second processing module is also used to construct a triangular mesh model according to the implicit field through an edge-based isosurface extraction method to complete three-dimensional surface reconstruction.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the three-dimensional reconstruction method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the three-dimensional reconstruction method according to any one of claims 1 to 7 is implemented.