Method, device and equipment for three-dimensional reconstruction, medium and program product

Through depth value optimization and semantic information fusion, dense point clouds are generated, and surface reconstruction and triangular mesh model fitting are performed, which solves the problem of low reconstruction accuracy in complex scenarios and generates a more realistic three-dimensional model.

CN119963766APending Publication Date: 2025-05-09CHINA TELECOM CLOUD TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411794179.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In complex scenarios, complex surface objects are blocked by surrounding objects. In the incremental reconstruction of complex surface objects in related technologies, there are holes, noise and texture information missing, and the reconstruction accuracy is low.

Method used

By obtaining the scene image and depth image to be reconstructed, depth value optimization and semantic information extraction are performed, depth information and semantic information are integrated, dense point clouds are generated, patch reconstruction and triangular mesh model fitting are performed, and texture reconstruction and mapping are combined to improve reconstruction accuracy.

Benefits of technology

The accuracy of three-dimensional reconstruction is improved, the problem of missing holes, noise and texture information in complex volume increment reconstruction is solved, and a more realistic reconstruction model is generated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963766A_ABST
    Figure CN119963766A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, and discloses a method, a device, equipment, a medium and a program product for three-dimensional reconstruction, and the method comprises the steps: obtaining a to-be-reconstructed scene image and a depth image corresponding to the to-be-reconstructed scene image, carrying out the depth value optimization of the depth image, and obtaining the optimized depth information; semantic information in the scene image to be reconstructed is extracted, and the optimized depth information and semantic information are fused to obtain dense point clouds; based on the dense point cloud, performing patch reconstruction to obtain a reconstructed point cloud patch; performing segmentation and fitting based on the reconstructed point cloud surface patch to obtain a triangular mesh model; according to the method, semantic information is extracted from a large number of scene images to be reconstructed, semantic information and depth information are fused, complementary learning is carried out, the method is used for combined inference of scene geometry and semantics, and environment semantic information is fused into the depth reconstruction process, so that the reconstruction precision is improved. And the three-dimensional reconstruction precision can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a method, device, equipment, medium and program product for three-dimensional reconstruction. Background Art

[0002] Dense 3D reconstruction in complex scenes is crucial in game animation, 3D printing, augmented reality (AR), virtual reality (VR), drone applications, and intelligent control. The goal of 3D reconstruction is to generate digital representations of objects and environments, and to perform 3D accurate and real-time dense reconstruction of objects, which can be applied in many fields such as robotics, assisted living, monitoring, and industry.

[0003] However, in complex scenes, complex curved objects are occluded by surrounding objects. The problems of holes, noise and missing texture information in the incremental reconstruction of complex surface objects in related technologies need to be solved urgently, and the reconstruction accuracy is low. Summary of the invention

[0004] In view of this, the present invention provides a method, apparatus, device, medium and program product for three-dimensional reconstruction to solve the problem of incremental reconstruction of complex surface objects and low reconstruction accuracy.

[0005] In a first aspect, the present invention provides a method for three-dimensional reconstruction, the method comprising: obtaining a scene image to be reconstructed and a depth image corresponding to the scene image to be reconstructed, optimizing the depth value of the depth image to obtain optimized depth information; extracting semantic information in the scene image to be reconstructed, fusing the optimized depth information and the semantic information to obtain a dense point cloud; performing facet reconstruction based on the dense point cloud to obtain reconstructed point cloud faces; performing segmentation and fitting based on the reconstructed point cloud faces to obtain a triangular mesh model; performing texture reconstruction and mapping based on the triangular mesh model to obtain a reconstructed dense texture model.

[0006] In an optional embodiment, the depth value of the depth image is optimized to obtain optimized depth information, including: calculating the confidence of the depth value of the pixel point in the depth image; based on the confidence, determining a priority queue, the priority queue including multiple seed points with confidence levels ranging from high to low; and performing nonlinear depth optimization based on the seed points to obtain the optimized depth information.

[0007] In an optional embodiment, the optimized depth information and the semantic information are fused to obtain a dense point cloud, including: preprocessing the optimized depth information to obtain a preprocessing result; generating semantic classification labels for pixels in the depth image based on the semantic information; and fusing the preprocessing result and the semantic classification label based on a neural network to obtain the dense point cloud.

[0008] In an optional embodiment, the facet reconstruction is performed based on the dense point cloud to obtain the reconstructed point cloud facet, including: reconstructing the facet based on the dense point cloud and fusing multi-view depth information; screening matching point pairs based on constraint conditions, sorting the screened matching point pairs by distance, and generating triangular facets based on the sorted matching point pairs; filtering the triangular facets to obtain the reconstructed point cloud facet.

[0009] In an optional embodiment, the segmentation and fitting based on the reconstructed point cloud patches to obtain a triangular mesh model includes: obtaining a first patch and a second patch from a plurality of overlapping reconstructed point cloud patches, and determining a first intersection point and a second intersection point among the intersection points of the first patch and the second patch; determining a segmentation vector based on the first intersection point and the second intersection point; segmenting the first patch and the second patch based on the segmentation vector, and fitting a mesh based on the segmented patches to obtain the triangular mesh model.

[0010] In an optional embodiment, the texture reconstruction and mapping based on the triangular mesh model to obtain a reconstructed dense texture model includes: based on the target facets in the triangular mesh model, determining the target area from the scene image to be reconstructed, mapping the target area to the target facets, and obtaining the reconstructed dense texture model.

[0011] In a second aspect, the present invention provides a device for three-dimensional reconstruction, comprising: a first acquisition module, used to acquire a scene image to be reconstructed and a depth image corresponding to the scene image to be reconstructed, and optimize the depth value of the depth image to obtain optimized depth information; a fusion module, used to extract semantic information in the scene image to be reconstructed, and fuse the optimized depth information and the semantic information to obtain a dense point cloud; a surface reconstruction module, used to perform facet reconstruction based on the dense point cloud to obtain reconstructed point cloud faces; a mesh segmentation module, used to perform segmentation and fitting based on the reconstructed point cloud faces to obtain a triangular mesh model; a texture reconstruction module, used to perform texture reconstruction and mapping based on the triangular mesh model to obtain a reconstructed dense texture model.

[0012] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the method for three-dimensional reconstruction of the above-mentioned first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0013] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the method for three-dimensional reconstruction of the above-mentioned first aspect or any corresponding embodiment thereof.

[0014] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions for causing a computer to execute the method for three-dimensional reconstruction of the first aspect or any corresponding embodiment thereof.

[0015] The method, apparatus, device, medium and program product for 3D reconstruction provided in this embodiment extract semantic information from a large number of scene images to be reconstructed, fuse semantic and depth information, perform complementary learning, and use it for joint inference of scene geometry and semantics, and fuse environmental semantic information into the depth reconstruction process, thereby improving the 3D reconstruction accuracy; the segmentation and fitting of the reconstructed point cloud patches can accurately maintain the clear details of the object boundaries, and solve the problems of holes, noise and missing texture information in the incremental reconstruction of complex volumes; texture reconstruction and mapping can make the reconstructed model more realistic and improve the reconstruction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the related technologies, the drawings required for use in the specific embodiments or the related technical descriptions will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0017] Figure 1 A schematic flow chart of a method for three-dimensional reconstruction provided by an embodiment of the present invention is shown;

[0018] Figure 2 A schematic diagram of the process of generating a dense point cloud is shown;

[0019] Figure 3 A schematic diagram of semantic segmentation results is shown;

[0020] Figure 4 A schematic diagram of the process of triangular patch fitting is shown;

[0021] Figure 5 A schematic diagram of the process of texture reconstruction is shown;

[0022] Figure 6 A schematic flow chart of another method for three-dimensional reconstruction provided by an embodiment of the present invention is shown;

[0023] Figure 7 A schematic structural diagram of a device for three-dimensional reconstruction is shown;

[0024] Figure 8 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0025] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0026] Point cloud reconstruction based on depth images can realize real-time 3D reconstruction of static scenes indoors. Consumer-grade depth cameras can provide depth information based on the combination of structured light sensors and color images. Through this type of red, green and blue three-channel color image (RGB-D) sensor combined with depth information, algorithms for high-speed reconstruction of dense 3D objects can be implemented.

[0027] The surface reconstruction method in the related art uses smoothness as a priori to infer implicit surfaces globally or locally from the point cloud, and then uses isosurface extraction algorithms such as the Marching Cubes algorithm to obtain explicit meshes. Representative works include: moving least squares-based methods (MLS), radial basis functions (RBF) or Hermite variants of RBF (HRBF) and the most commonly used Poisson reconstruction. The reconstruction quality of these methods is limited by the resolution of triangles. They usually achieve high-fidelity and high-resolution triangular meshes by generating dense, accurate and distortion-free high-fidelity images, or require dense triangles to ensure high fidelity, especially for sharp edges, which may lead to unnecessary computational and storage pressures, as well as the computational burden brought by the presence of a large number of triangular meshes.

[0028] The shortcomings of the algorithms in the related technologies are mainly manifested in:

[0029] 1. Due to the low resolution of the point cloud image obtained by depth, the depth information matching accuracy contained in the camera is low, and the multi-view depth image point cloud registration and stitching have problems of low registration accuracy and reconstruction ambiguity.

[0030] 2. Reconstruct 3D mesh surface from point cloud. Currently there is no complete framework to obtain high-fidelity mesh model from point cloud.

[0031] 3. It is not easy to find the continuity of the patch-based representation of the object surface.

[0032] 4. The semantic information of the scene is not fully utilized to improve the reconstruction results that rely on geometry and depth information.

[0033] These limitations restrict the texture detail and clarity of 3D reconstructed models to some extent.

[0034] According to an embodiment of the present invention, a method embodiment for three-dimensional reconstruction is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0035] The concepts involved in the embodiments of the present invention are introduced below. The three coordinate systems of the visual system include the world coordinate system, the camera coordinate system and the pixel coordinate system. World coordinate system: the absolute coordinate system of the system, which is the coordinate system of the entire scene; camera coordinate system, the coordinate system of the camera at its own angle, with the origin at the optical center of the camera, and the Z axis parallel to the optical axis of the camera; pixel coordinate system, which is a continuous image coordinate system established with the intersection of the diagonals of the picture as the baseline origin. The three coordinate systems can be converted to each other according to rotation and translation parameters.

[0036] In this embodiment, a method for three-dimensional reconstruction is provided, which can be used in a terminal, such as a mobile phone, a tablet computer, a desktop computer, or a server, etc. Figure 1 A schematic diagram of a process for three-dimensional reconstruction provided by an embodiment of the present invention is shown. Figure 1 As shown, the process includes the following steps:

[0037] Step S101 : obtaining a scene image to be reconstructed and a depth image corresponding to the scene image to be reconstructed, optimizing the depth value of the depth image, and obtaining optimized depth information.

[0038] In this step, the region growing and dilation method can be used to perform nonlinear optimization on the depth values ​​in the depth image and establish an operation time series model. The depth image is fused on a dense surface, which is conducive to the reconstruction of the missing part.

[0039] Step S102: extracting semantic information from the scene image to be reconstructed, fusing the optimized depth information and semantic information, and obtaining a dense point cloud.

[0040] In this step, the Depth-Semantic Fusion Network (DSFNet), a model of embedded fusion front-end and back-end networks, can be used to extract rich object information from complex environmental information by aggregating multiple layers of semantics, and then extract semantic information based on the similarity of features between each layer of semantics.

[0041] In this way, in complex indoor environment scenes, both the depth information in the environment and the semantic information in the indoor environment are utilized, making full use of the rich information in the complex indoor scenes.

[0042] Step S103, performing face reconstruction based on the dense point cloud to obtain a reconstructed point cloud face.

[0043] In this step, a multi-view 3D stereo vision algorithm based on facets can be used to fuse multi-view depth information and perform point cloud reconstruction, thereby ensuring that each image block has at least one facet projection.

[0044] Among them, a patch is a small rectangle that reconstructs the surface of an object, one side of which is parallel to the x-axis of the reference camera. Each patch P has a center point and a unit normal vector (pointing to the optical center of the camera).

[0045] Step S104, segmenting and fitting are performed based on the reconstructed point cloud patches to obtain a triangular mesh model.

[0046] In this step, a coarse primitive detector (HPNet) based on supervised learning can be used to combine neural networks with geometric constraints using a two-stage hybrid model. After extracting and reconstructing point cloud patches based on the primitive detection module, a triangular mesh is fitted to obtain a triangular mesh model.

[0047] Among them, triangular meshes have the characteristics of simple representation and simple operation. Three-dimensional models are constructed based on the position and color information of these triangular meshes. Other polygonal meshes can be simplified into triangular meshes. Triangles need to represent the vertex, edge, and face information of the triangular mesh.

[0048] Primitive detection refers to the process of detecting and identifying basic shapes or specific local patterns in an image. Primitives can be simple geometric shapes or complex local graphic structures.

[0049] Step S105, performing texture reconstruction and mapping based on the triangular mesh model to obtain a reconstructed dense texture model.

[0050] In this step, texture reconstruction and mapping are the processes of reconstructing the three-dimensional scene from dense point cloud to outputting a renderable model, so that the reconstructed three-dimensional scene can present a more realistic appearance.

[0051] By mapping texture reconstruction, tens of thousands or even millions of vertices, faces, images used to reconstruct the mesh, and camera internal and external parameters can be processed into texture coordinates and texture maps corresponding to each vertex.

[0052] The method for 3D reconstruction provided in this embodiment extracts semantic information from a large number of scene images to be reconstructed, fuses semantic and depth information, performs complementary learning, and is used for joint inference of scene geometry and semantics. It fuses environmental semantic information into the depth reconstruction process, and can achieve better generalization performance and higher extraction accuracy in complex environments, thereby improving the accuracy of 3D reconstruction. The segmentation and fitting of the reconstructed point cloud patches can accurately maintain clear details of the object boundaries, and solve the problems of holes, noise, and missing texture information in the incremental reconstruction of complex volumes. Texture reconstruction and mapping can make the reconstructed model more realistic and improve the reconstruction accuracy.

[0053] In some optional embodiments, depth values ​​of a depth image are optimized to obtain optimized depth information, including: calculating the confidence of the depth values ​​of pixels in the depth image; determining a priority queue based on the confidence, the priority queue including multiple seed points with confidence levels ranging from high to low; and performing nonlinear depth optimization based on the seed points to obtain optimized depth information.

[0054] In this embodiment, the confidence can be calculated based on the accuracy of depth measurement, image clarity, or stability of feature points. Depth estimation can be performed based on initial sparse features. Sparse feature points, such as corner points or edge points, are extracted from the scene image to be reconstructed, and a preliminary depth estimation is performed on the area around the feature points based on the depth information of the feature points.

[0055] Perform nonlinear depth optimization on each seed point, and update the priority queue based on the nonlinear depth optimization results so that new high-confidence areas can be processed in the next iteration. For the optimized pixel points, check the pixels in the neighborhood of the pixel point. If there are unprocessed pixels in the neighborhood, that is, pixels with no depth value or pixels with higher confidence than the pixels in the adjacent range, add the aforementioned pixels to the priority queue.

[0056] In this way, the depth value is optimized, an optimization model is established, the depth information is fused into the dense surface, and the depth information can be used to supplement the reconstruction part.

[0057] In some optional embodiments, the optimized depth information and semantic information are fused to obtain a dense point cloud, including: preprocessing the optimized depth information to obtain a preprocessing result; generating semantic classification labels for pixels in the depth image based on the semantic information; and fusing the preprocessing result and the semantic classification label based on a neural network to obtain a dense point cloud.

[0058] In this embodiment, Figure 2 A schematic diagram of the process of generating a dense point cloud is shown in FIG. Figure 2 As shown, the scene image to be reconstructed includes an image of the current scene obtained by multiple rays, that is, a stereo image pair 201. The optimized depth information is corrected for noise, outliers and missing values, and the confidence value of each ray is further estimated. For each ray, a local voxel grid related to the depth and view is extracted. Each ray is sampled with the surface as the center.

[0059] Based on the semantic information, a semantic classification label of a pixel in a depth image is generated, including: extracting semantic information from a scene image to be reconstructed (RGB image) 202, providing a semantic classification label 203 for each pixel, storing the depth value and its semantic label as a separate channel, and updating the semantic label according to the confidence of the predicted depth map of each pixel. A network such as FuseNet-10, FuseNet-20, FuseNet-30, FuseNet-100, SSMA, SegNet, or LinkNet can be used to extract a two-dimensional (2D) semantic label of a depth image (RGB-D) as an initial feature map of the model input image. Figure 3 A schematic diagram of the semantic segmentation results is shown.

[0060] Based on the neural network, the preprocessing results and semantic classification labels are fused to obtain a dense point cloud, including: for a given training time, the noisy depth map, camera trajectory and 2D semantic labels, the depth of the frame is fused with the semantic labels in the scene space based on the neural network method. And for each ray in the point cloud mesh, the position is updated according to the back projection, and the predicted point cloud value is projected into the global point cloud mesh. The contextual information provided by the scene semantics helps the deep fusion network learn noise-resistant features, especially in complex environments, which can overcome the shortcomings of the current deep fusion method in dealing with thin object structures, thickening artifacts and false surfaces.

[0061] An embedded global inference attention module (GIAM) can be constructed to assign association weights between points in the model to reason about the global topology of the 3D model.

[0062] The specific implementation of the deep-semantic fusion model includes: sequentially passing the input 3D feature volume through two consecutive convolutional layers of encoding blocks to make the plane as close to the real scene as possible to reduce the noise on the mesh surface and sharpen the geometric features; the kernel size is 3*3, and the convolution operation can change the number of feature channels. Through leaky ReLUs for nonlinear activation, and a dropout layer, the output of each block is connected to its input and passed to the next block. With the increase of each block, the receptive field of the neural network increases. The output of the feature extraction result is a 100-dimensional feature vector on each ray. Then the feature quantity is obtained and the point cloud update along each ray is predicted. The features pass through a convolution block with two 1*1 convolutional layers and are crossed with leaky ReLUs, normalization and dropout layers to reduce the number of features. By comparing the fused features with the original feature map, a point cloud prediction value is generated, and finally the network can decide whether the point cloud value needs to be updated. For example, in the case of outliers, choose not to update, thereby reducing the impact of these values. It is updated step by step according to the input image, and the final point cloud model is calculated and improved step by step.

[0063] In this way, the deep-semantic fusion network (DSFNet) uses the predicted 2D semantic label prior, the depth sensor prior and the point cloud data of the current frame for efficient deep fusion and point cloud update. Through the semantic network and the deep fusion network, the output semantic label and its corresponding confidence, as well as the depth information output by the deep fusion network and its corresponding confidence, together with the previous frame point cloud data and corresponding weights, can be passed through the convolution layer to extract features.

[0064] The Deep-Semantic Fusion Network (DSFNet) improves depth information through complementary RGB information, fuses semantic features to compensate for lost depth information, and learns to reduce point cloud accuracy problems caused by noise. Aggregate multiple layers of semantics to extract object information from complex indoor environments, and then build a global topological structure based on the similarity of features between layers. Through the Deep-Semantic Fusion Network DSFNet, environmental semantic information is integrated into the depth reconstruction process to improve reconstruction accuracy. In addition, an embedded global inference attention module (GIAM) is added to assign association weights between points in the model.

[0065] In some optional embodiments, facet reconstruction is performed based on a dense point cloud to obtain a reconstructed point cloud facet, including: performing facet reconstruction based on the dense point cloud and fusing multi-view depth information; screening matching point pairs based on constraint conditions, sorting the screened matching point pairs by distance, and generating triangular faces based on the sorted matching point pairs; and filtering the triangular faces to obtain a reconstructed point cloud facet.

[0066] In this embodiment, a series of sparse patches are generated. The generation and screening of patches are performed multiple times to make the patches dense enough, and patches that do not meet the requirements are removed. The constraint condition can be a two-pixel error. By using the limit constraint of a two-pixel error in the target image, the same type of feature points are formed in other images. From these matching point pairs, the patches are generated one by one according to the distance sorted from small to large. Since the generated patches may have many errors, it is considered that in the target image I i The visible patch is an image where the angle between the normal vector of the patch and the line connecting the center of the patch to the optical center of the camera is less than a certain angle, satisfying the following relationship:

[0067]

[0068] Where V(p) represents the normal vector of the face p, I m is a collection of multi-view images, I i is the target image in the multi-view image, n(p) is the normal vector of the patch p, which is used to determine the local area or neighborhood of point p, and c(p) is the coordinate of the center point of the patch p, which is used to calculate the line connecting the patch to the optical center of the camera. c(p)o(I i ) represents the distance from the center point c(p) of patch p to image I i The camera optical center o(I i ) is used to calculate the angle between the patch normal and the line. i )| represents the modulus of the vector, that is, the distance from the center point of the patch p to the target image I i The distance from the camera optical center. cosτ represents a threshold used to determine the distance between the patch p and the target image I. i Specifically, when the normal vector n(p) of the patch p and the line c(p)o(I i ) is greater than cosτ, the patch p is considered to be in the image I i is visible.

[0069] The purpose of patch generation is to ensure that each image block corresponds to at least one patch. By giving a patch p, we first obtain a set of domain image blocks C(p) that meet the conditions, and then perform patch generation.

[0070] Image block set C(p):

[0071] C(p)={C i (x ′ ,y ′ )|p∈Q i (x,y),|xx ′ |+|yy ′ |=1}

[0072] The image block set C(p) includes all image blocks adjacent to the noise position of the patch p, that is, satisfying |xx ′ |+|yy ′ |=1 condition point p, C i (x ′ ,y ′ ) represents the i-th image block, and the center coordinate of the i-th image block is (x ′ ,y ′ ).

[0073] Patches p and p ′ When the following relationship is satisfied, the two are considered to be adjacent:

[0074] |(c(p)-c(p ′ ))·n(p)|+|(c(p)-c(p ′ ))·n(p ′ )|<2ρ1

[0075] When there is a patch p ′ The image block C to which it belongs i (x ′ ,y ′ ) satisfies C i (x ′ ,y ′ )∈C(p), and both p and p ′ When it is an adjacent relationship, C i (x ′ ,y ′ ) is deleted from C(p) and patch generation is not performed on it. For the remaining image blocks in C(p), patch generation will be performed to generate a new patch p ′ .

[0076] In the process of patch reconstruction, some patches with large errors may be generated and need to be filtered to ensure the accuracy of the patches. For a patch p in U(p), if the following formula is satisfied, the patch is filtered out:

[0077] |V(p)|(1-g(p))<∑p i ∈U(p)(1-g(p))

[0078] Among them, U(p) represents a set or neighborhood where patch p is located, which is used to evaluate the accuracy and reliability of patch p. g(p) is a function used to evaluate the quality or accuracy of patch p. g(p) can be determined based on the geometric shape of the patch, its consistency with other patches, or the degree of matching with the original data. The closer the value of g(p) is to 1, the higher the quality of patch p. ∑p i∈U(p)(1-g(p)) represents all patches p in the domain U(p) of patch p i The (1-g(p)) values ​​of the face are summed. The result of the summation can be regarded as the sum of the inaccuracies of all patches in the domain U(p). If the weight of the patch p multiplied by its uncertainty is less than the sum of the inaccuracies of all patches in its domain, then the patch is filtered out.

[0079] In this way, a texture reconstruction method of surface point cloud density reconstruction is proposed based on the surface dense reconstruction of the triangular patch multi-view stereo (PMVS) algorithm and the principle of multi-view selection of stereo images.

[0080] In some optional embodiments, segmentation and fitting are performed based on the reconstructed point cloud patches to obtain a triangular mesh model, including: obtaining a first patch and a second patch from overlapping reconstructed point cloud patches in a plurality of reconstructed point cloud patches, and determining a first intersection point and a second intersection point among the intersection points of the first patch and the second patch; determining a segmentation vector based on the first intersection point and the second intersection point; segmenting the first patch and the second patch based on the segmentation vector, and fitting a mesh based on the segmented patches to obtain a triangular mesh model.

[0081] In this embodiment, the input 3D point cloud P = {p i |1≤i≤N}, and every point p i Include Location and the normal direction at point p therefore

[0082] HPNet, a coarse primitive detector based on supervised learning, adopts a two-stage hybrid model to combine neural networks with geometric constraints. Although geometric constraints are used, the main limitation comes from the dynamic graph-based convolutional neural network (DGCNN) backbone, which will cause the throughput during training to be significantly lower than other learning-based point cloud processing methods, resulting in high training costs. First, eliminate the points smaller than N min The triangle of the point, calculate the convexity of the remaining triangles, N min represents a threshold of HPNet. Secondly, take two adjacent patches P p and P q , detect their adjacency and decide whether to merge them. Use the P-linkage clustering method to automatically cluster each patch into a collection of smaller slices, thus generating a cluster set C p ={C pi |1≤i≤N p} and C q ={C qi |1≤i≤N q}. Again, use principal component analysis to analyze each C p and C qPerform a local analysis, calculate the normals between them, represent each patch as multiple slices with normals. Finally, if more than N c The slice pairs satisfy the following equation, which means that the two triangles P p and P q With good adjacency, smooth transitions, they will be merged.

[0083]

[0084] Among them, θ t is the angle threshold that determines the curvature of the surfaces to be merged, Represents C p and C q The refinement clustering submodule in the embodiment of the present invention effectively eliminates the interference caused by trivial patches and merges over-segmented patches.

[0085] Figure 4 Figure 2 shows a schematic diagram of the process of triangular patch fitting. Figure 4 As shown, after extracting the point cloud patches from the primitive detection module, they are split into pairs. The triangle Δa and Δb are extended into a plane. Divide Δa into three small triangles (Δp1a3p2, Δp1p2a2, Δp1a2a1), according to the vector and Split them up, processing the triangles on the intersection lines.

[0086]

[0087] Where p1, p2 are the intersection points of planes Δa and Δb, is a normal vector known to point inside or outside the model, Perpendicular to vector If satisfied Then the triangle Δp1a3p2 is included in the set S a (A), on the contrary, is included in the set S a (B) Chinese.

[0088] For a triangle Δo in the non-intersecting area, locate the intersection point on the intersection line that is closest to Δo, and then classify Δo into S according to the following equation a (A) or S a (B).

[0089]

[0090] Where k is the set of two partitioned triangles S a (A) S a(B) k-domain, point p is the key reference point for segmenting Δo, locates the nearest intersection point p on the intersection line, and determines which segmented triangle set Δo should be included in based on point p. dis(a,b) represents the Euclidean distance from Δa to Δb. Applying these operations to all surfaces splits them into a set of candidate faces. This meshing and segmentation algorithm is crucial to producing high-quality boundaries.

[0091] In this way, by using the improved primitive extraction module and fusing the post-processing refinement module, the primitive blocks with the best performance can be extracted from the point cloud, which can retain the clear boundary features in complex scenes, realize the surface shape detail recovery of three-dimensional objects, and generate a high-fidelity mesh model. The efficient mesh fitting and segmentation module retains the sharp features in the point cloud, can accurately maintain the clear boundary features, and generate a high-fidelity reconstruction model. The original patch can be accurately and reasonably segmented, the mesh can be fitted in each patch, and the overlapping mesh can be segmented at the triangle level to ensure true clarity while obtaining a lightweight mesh model.

[0092] The primitive detection-based framework for reconstructing meshes from point clouds can accurately maintain the details of clear object boundaries, solve the problems of holes, noise and missing texture information in traditional incremental reconstruction of complex volumes, and generate high-fidelity models.

[0093] Mesh segmentation can separate overlapping triangular meshes, produce clear and continuous segmentation, reconstruct high-quality sharp edges, and realize lightweight mesh models of real scanned point clouds. Compared with related technologies, it uses fewer triangles to capture object boundary details and flexibility. In addition, this module can be combined with other modules, such as edge extraction or random sample consistency (RANSAC), as a plug-and-play post-processing module to generate clear, lightweight, high-quality mesh models.

[0094] In some optional embodiments, texture reconstruction and mapping are performed based on the triangular mesh model to obtain a reconstructed dense texture model, including: based on the target facets in the triangular mesh model, determining the target area from the scene image to be reconstructed, mapping the target area to the target facets, and obtaining a reconstructed dense texture model.

[0095] In this embodiment, for each patch, an optimal area is found from the input image to reflect its color, and this area is inserted into the texture map. The Markov random field (MRF) is introduced to de-delineate the texture details, and the texture details of the 3D reconstructed model are accurately calculated together with the texture coordinates and patches. Figure 5 A schematic diagram of the texture reconstruction process is shown in FIG. Figure 5As shown, based on the target patch in the triangular mesh model, the target area is determined from the scene image to be reconstructed, and the target area is mapped to the target patch to obtain a reconstructed dense texture model, including: a perspective map 502 determined based on the image and camera pose and a three-dimensional mesh 501 determined based on the patch and the vertex. The optimal perspective can be found for each patch, and the index number of the perspective is recorded with a label to obtain patches 503 with the same label, and then the patches using the same label are grouped into a block 504, the boundary of the block is found, the image of the area is cut out from the corresponding RGB image, and pasted into the texture map to obtain a texture reconstruction result 505, and the texture coordinates corresponding to each vertex are recorded.

[0096] In this way, the Markov random field (MRF) is introduced to describe the richness of image details, and combined with the precise calculation of regional coordinates and texture coordinates to reproduce the texture detail model of 3D reconstruction, by mapping the color and texture of the real object surface to the virtual model, the model can be made more realistic and the user's immersion and experience can be improved.

[0097] Figure 6 FIG. 2 is a flow chart of another method for three-dimensional reconstruction provided by an embodiment of the present invention. Figure 6 As shown, the method for three-dimensional reconstruction includes: step S601, obtaining a complex indoor scene image; step S602, based on the complex indoor scene image, collecting RGB-D data; step S603, based on the collected RGB-D data, using the regional growing method, optimizing the depth value to obtain the depth value optimization result; step S604, based on the collected RGB-D data, extracting semantic information to obtain semantic information; step S605, inputting the depth value optimization result and the semantic information into the deep-semantic fusion network DSFNet, performing multi-layer fusion training, and obtaining a dense point cloud; step S606, for the dense point cloud, using primitives to detect point cloud patches, extracting and segmenting the point cloud patches, and obtaining a triangular mesh model; step S607, performing texture reconstruction on the triangular mesh model to obtain a three-dimensional reconstruction result of a dense texture model.

[0098] In this embodiment, a device for three-dimensional reconstruction is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0099] This embodiment provides a device for three-dimensional reconstruction. Figure 7 A schematic diagram of the structure of a device for three-dimensional reconstruction is shown, Figure 7 As shown, including:

[0100] The first acquisition module 701 is used to acquire a scene image to be reconstructed and a depth image corresponding to the scene image to be reconstructed, and optimize the depth value of the depth image to obtain optimized depth information.

[0101] The fusion module 702 is used to extract semantic information from the scene image to be reconstructed, and fuse the optimized depth information and semantic information to obtain a dense point cloud.

[0102] The surface reconstruction module 703 is used to perform surface reconstruction based on the dense point cloud to obtain a reconstructed point cloud surface.

[0103] The mesh segmentation module 704 is used to perform segmentation and fitting based on the reconstructed point cloud patches to obtain a triangular mesh model.

[0104] The texture reconstruction module 705 is used to perform texture reconstruction and mapping based on the triangular mesh model to obtain a reconstructed dense texture model.

[0105] In some optional implementations, the first acquisition module includes:

[0106] The first acquisition unit is used to calculate the confidence of the depth value of the pixel point in the depth image; based on the confidence, determine the priority queue, the priority queue includes a plurality of seed points with confidence levels from high to low; perform nonlinear depth optimization based on the seed points to obtain optimized depth information.

[0107] In some optional embodiments, the fusion module includes:

[0108] The fusion unit is used to preprocess the optimized depth information to obtain a preprocessing result; generate semantic classification labels for pixels in the depth image based on the semantic information; and fuse the preprocessing result and the semantic classification label based on a neural network to obtain a dense point cloud.

[0109] In some optional embodiments, the surface reconstruction module includes:

[0110] The surface reconstruction unit is used to reconstruct the patch based on the dense point cloud and the multi-view depth information; filter the matching point pairs based on the constraint conditions, sort the distance of the filtered matching point pairs, and generate triangular patches based on the sorted matching point pairs; filter the triangular patches to obtain the reconstructed point cloud patches.

[0111] In some optional implementations, the grid segmentation module includes:

[0112] A mesh segmentation unit is used to obtain the first and second faces of overlapping reconstructed point cloud faces in multiple reconstructed point cloud faces, determine the first and second intersection points of the first and second faces; determine a segmentation vector based on the first and second intersection points; segment the first and second faces based on the segmentation vector, and fit a mesh based on the segmented faces to obtain a triangular mesh model.

[0113] In some optional implementations, the texture reconstruction module includes:

[0114] The texture reconstruction unit is used to determine the target area from the scene image to be reconstructed based on the target face in the triangular mesh model, map the target area to the target face, and obtain a reconstructed dense texture model.

[0115] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0116] The device for three-dimensional reconstruction in this embodiment is presented in the form of a functional unit, where the unit refers to an application specific integrated circuit (ASIC) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0117] The embodiment of the present invention also provides a computer device having the above Figure 7 The apparatus for three-dimensional reconstruction is shown.

[0118] See also Figure 8 , Figure 8 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 8 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the computer device, including instructions stored in or on the memory to display graphical information of a graphical user interface on an external input / output device (such as a display device coupled to an interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 8 A processor 10 is taken as an example.

[0119] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0120] The aforementioned memory 20 stores instructions executable by at least one processor 10, so that the aforementioned at least one processor 10 executes the method shown in the above embodiment.

[0121] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0122] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0123] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 8 The example of connecting through bus is taken in the following.

[0124] The input device 30 can receive input digital or character information, and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator rod, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (such as a light emitting diode) and a tactile feedback device (such as a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.

[0125] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0126] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.

[0127] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for three-dimensional reconstruction, characterized in that The method comprises: Acquire a scene image to be reconstructed and a depth image corresponding to the scene image to be reconstructed, and optimize the depth value of the depth image to obtain optimized depth information; Extracting semantic information from the scene image to be reconstructed, fusing the optimized depth information and the semantic information to obtain a dense point cloud; Based on the dense point cloud, performing face reconstruction to obtain a reconstructed point cloud face; Perform segmentation and fitting based on the reconstructed point cloud patches to obtain a triangular mesh model; Texture reconstruction and mapping are performed based on the triangular mesh model to obtain a reconstructed dense texture model.

2. The method according to claim 1, characterized in that Optimizing the depth value of the depth image to obtain optimized depth information includes: Calculating the confidence of the depth value of the pixel in the depth image; Based on the confidence, determining a priority queue, the priority queue comprising a plurality of seed points with confidence levels from high to low; Nonlinear depth optimization is performed based on the seed point to obtain the optimized depth information.

3. The method according to claim 1 or 2, characterized in that: The optimized depth information and the semantic information are integrated to obtain a dense point cloud, including: Preprocessing the optimized depth information to obtain a preprocessing result; Based on the semantic information, generating semantic classification labels for pixels in the depth image; The preprocessing result and the semantic classification label are fused based on a neural network to obtain the dense point cloud.

4. The method according to claim 1, characterized in that The step of performing facet reconstruction based on the dense point cloud to obtain a reconstructed point cloud facet comprises: Reconstructing the patch based on fusing multi-view depth information with the dense point cloud; Filter matching point pairs based on constraint conditions, sort the filtered matching point pairs by distance, and generate triangular facets based on the sorted matching point pairs; The triangular facets are filtered to obtain the reconstructed point cloud facets.

5. The method according to claim 1, characterized in that The segmentation and fitting based on the reconstructed point cloud patch to obtain a triangular mesh model includes: Acquire a first facet and a second facet among the overlapping reconstructed point cloud facets in the plurality of reconstructed point cloud facets, and determine a first intersection point and a second intersection point among the intersection points of the first facet and the second facet; Determine a segmentation vector based on the first intersection point and the second intersection point; The first face patch and the second face patch are segmented based on the segmentation vector, and meshes are fitted based on the segmented face patches to obtain the triangular mesh model.

6. The method according to claim 1, characterized in that The texture reconstruction and mapping based on the triangular mesh model to obtain a reconstructed dense texture model includes: Based on the target facet in the triangular mesh model, a target area is determined from the scene image to be reconstructed, and the target area is mapped to the target facet to obtain the reconstructed dense texture model.

7. A device for three-dimensional reconstruction, characterized in that The device comprises: A first acquisition module is used to acquire a scene image to be reconstructed and a depth image corresponding to the scene image to be reconstructed, and optimize the depth value of the depth image to obtain optimized depth information; A fusion module, used to extract semantic information from the scene image to be reconstructed, and fuse the optimized depth information and the semantic information to obtain a dense point cloud; A surface reconstruction module, used for performing surface reconstruction based on the dense point cloud to obtain a reconstructed point cloud surface; A mesh segmentation module, used for segmenting and fitting based on the reconstructed point cloud patches to obtain a triangular mesh model; The texture reconstruction module is used to perform texture reconstruction and mapping based on the triangular mesh model to obtain a reconstructed dense texture model.

8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method for three-dimensional reconstruction according to any one of claims 1 to 6 by executing the computer instructions.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method for three-dimensional reconstruction according to any one of claims 1 to 6.

10. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the method for three-dimensional reconstruction according to any one of claims 1 to 6.

Citation Information

Cited By

  • Building three-dimensional model intelligent reconstruction method and system based on deep learning

    CN120765863A