Scene surface reconstruction method and device, equipment and medium
By combining multi-view image sequence and rendering loss function to update the model in three-dimensional Gaussian splattering technology, the floating object phenomenon is solved and the quality and accuracy of scene reconstruction are improved.
Patent Information
- Application Number
- CN202510432933.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The existing three-dimensional Gaussian splattering technology when dealing with areas with smooth pixel value changes or poor lighting, resulting in floating objects appearing in the reconstruction scene, which has a lower reconstruction quality.
By introducing an initialized three-dimensional Gaussian model, combining multi-view image sequence to determine the attribute information and position information of two-dimensional Gaussians, using the Gaussian splattering method to generate the rendered image, and updating the three-dimensional Gaussian model by minimizing the rendering loss function, finally building the surface model of the target scene.
It significantly improves the quality and accuracy of scene reconstruction, avoids floating objects, and ensures accurate reflection of scene structure and details.
Smart Images

Figure CN120355848A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of computer graphics, and in particular, to a method, apparatus, device, and medium for scene surface reconstruction. Background Art
[0002] With the rapid development of Virtual Reality (VR) technology and Augmented Reality (AR) technology, the 3D Gaussian Splatting (3DGS) technology has emerged. The 3DGS technology can explicitly represent a three-dimensional scene by using three-dimensional Gaussian spheres, and synthesize new views of the three-dimensional scene from any perspective by means of Gaussian splatting, so that the rendered new views have high-quality visual effects and fast rendering speeds.
[0003] However, in practical applications, if there are weakly textured regions in the real image where the pixel values change very smoothly or are almost unchanged, or there are poorly lit regions, the finally reconstructed scene may not be accurately aligned with the real object. This may cause the three-dimensional Gaussian spheres to float in the air, forming incorrect "floating objects", resulting in a low quality of the reconstructed scene. Summary of the Invention
[0004] Based on the above problems, the present application provides a method, apparatus, device, and medium for scene surface reconstruction, which can more accurately reflect the structure and details of the target scene, avoid the appearance of the floating object phenomenon, and significantly improve the quality and accuracy of scene reconstruction.
[0005] The embodiments of the present application disclose the following technical solutions:
[0006] In a first aspect, the present application discloses a method for scene surface reconstruction, the method comprising:
[0007] When obtaining an initialized three-dimensional Gaussian model according to a multi-view image sequence of a target scene, determining the attribute information and position information of a two-dimensional Gaussian corresponding to a three-dimensional anchor Gaussian in the initialized three-dimensional Gaussian model;
[0008] Generating a rendered image by means of Gaussian splatting according to the attribute information and position information of the two-dimensional Gaussian, wherein the rendered image includes color information, semantic information, depth information, and normal information;
[0009] Based on the rendered image and the image in the multi-view image sequence that has the same view angle as the rendered image, update the initialized three-dimensional Gaussian model by minimizing a rendering loss function to obtain an updated three-dimensional Gaussian model, where the rendering loss function includes a color loss function, a semantic loss function, a depth loss function, and a normal loss function;
[0010] Based on the updated three-dimensional Gaussian model and the multi-view image sequence, determine the depth information of the target scene;
[0011] Based on the depth information and the rendered image, construct a surface model of the target scene through a three-dimensional reconstruction algorithm.
[0012] Optionally, determining the attribute information and position information of the two-dimensional Gaussian corresponding to the three-dimensional anchor Gaussian in the initialized three-dimensional Gaussian model includes:
[0013] Determine the projection radius of each three-dimensional anchor Gaussian in the initialized three-dimensional Gaussian model by projecting it onto the image plane;
[0014] Determine the attribute information and position information of the two-dimensional Gaussian corresponding to the three-dimensional anchor Gaussian with a projection radius greater than 0.
[0015] Optionally, determining the depth information of the target scene based on the updated three-dimensional Gaussian model and the multi-view image sequence includes:
[0016] Determine the average value of the alpha blending weights of the two-dimensional Gaussians corresponding to each three-dimensional anchor Gaussian in the updated three-dimensional Gaussian model;
[0017] Obtain a second-updated three-dimensional Gaussian model by deleting the three-dimensional anchor Gaussians corresponding to the two-dimensional Gaussians with an average value less than a first preset threshold;
[0018] Based on the second-updated three-dimensional Gaussian model and the multi-view image sequence, determine the depth information of the target scene.
[0019] Optionally, determining the depth information of the target scene based on the updated three-dimensional Gaussian model and the multi-view image sequence includes:
[0020] Generate an opacity map corresponding to the rendered image, where the opacity map characterizes the opacity of each pixel in the rendered image;
[0021] For pixels with an opacity less than a second preset threshold, determine the target pixel that is the closest to the pixel and has an opacity greater than the second preset threshold;
[0022] By performing back-projection on the target pixel in the updated three-dimensional Gaussian model, a three-dimensional Gaussian model updated three times is obtained;
[0023] Based on the three-dimensional Gaussian model updated three times and the multi-view image sequence, the depth information of the target scene is determined.
[0024] Optionally, the attribute information of the two-dimensional Gaussian includes color attribute information, geometric attribute information, semantic attribute information, and opacity attribute information; the determination method of the attribute information of the two-dimensional Gaussian is as follows:
[0025] Determine the color feature vector and semantic feature vector of the target scene;
[0026] Input the color feature vector into the color fully connected network and opacity fully connected network of the initialized three-dimensional Gaussian model respectively, and obtain the color attribute information and opacity attribute information respectively;
[0027] Input the semantic feature vector into the geometric fully connected network and semantic fully connected network of the initialized three-dimensional Gaussian model respectively, and obtain the geometric attribute information and opacity attribute information respectively.
[0028] Optionally, the determination method of the position information of the two-dimensional Gaussian is as follows:
[0029] Generate a sparse point cloud of the target scene according to the multi-view image sequence of the target scene;
[0030] According to the coordinates of the sparse point cloud and the preset distance between Gaussian anchor points, determine the anchor positions of the three-dimensional anchor Gaussians in the initialized three-dimensional Gaussian model;
[0031] According to the anchor positions and the position offsets, determine the position information of the two-dimensional Gaussian.
[0032] Optionally, the formulas of the depth loss function and the normal loss function are as follows respectively:
[0033]
[0034] where L d-smooth is the depth loss function, is the depth image, w h is the cosine similarity of the semantic features of adjacent pixels in the vertical direction, d(x i,j ) is the pixel value at position (i, j) in the depth image, w w is the cosine similarity of the semantic features of adjacent pixels in the horizontal direction, L n-smooth is the normal loss function, is the normal image, n(x i,j) is the pixel value at position (i, j) in the normal image.
[0035] In a second aspect, the present application discloses a scene surface reconstruction device, which includes: a first determination module, an image generation module, a model update module, a second determination module, and a model construction module;
[0036] The first determination module is configured to determine the attribute information and position information of the two-dimensional Gaussian corresponding to the three-dimensional anchor Gaussian in the initialized three-dimensional Gaussian model when obtaining the initialized three-dimensional Gaussian model according to the multi-view image sequence of the target scene;
[0037] The image generation module is configured to generate a rendered image by the Gaussian splashing method according to the attribute information and position information of the two-dimensional Gaussian, where the rendered image includes color information, semantic information, depth information, and normal information;
[0038] The model update module is configured to update the initialized three-dimensional Gaussian model by minimizing the rendering loss function according to the rendered image and the image in the multi-view image sequence with the same view as the rendered image, and obtain an updated three-dimensional Gaussian model, where the rendering loss function includes a color loss function, a semantic loss function, a depth loss function, and a normal loss function;
[0039] The second determination module is configured to determine the depth information of the target scene according to the updated three-dimensional Gaussian model and the multi-view image sequence;
[0040] The model construction module is configured to construct a surface model of the target scene by a three-dimensional reconstruction algorithm according to the depth information and the rendered image.
[0041] Optionally, the first determination module is specifically configured to: determine the projection radius of each three-dimensional anchor Gaussian in the initialized three-dimensional Gaussian model by projecting each three-dimensional anchor Gaussian onto the image plane;
[0042] Determine the attribute information and position information of the two-dimensional Gaussian corresponding to the three-dimensional anchor Gaussian with a projection radius greater than 0.
[0043] Optionally, the second determination module specifically includes: a first determination sub-module, a second determination sub-module, and a third determination sub-module;
[0044] The first determination sub-module is configured to determine the average value of the alpha blending weights of the two-dimensional Gaussian corresponding to each three-dimensional anchor Gaussian in the updated three-dimensional Gaussian model;
[0045] A second determination sub-module, configured to obtain a three-dimensional Gaussian model after secondary update by deleting three-dimensional anchor Gaussians corresponding to two-dimensional Gaussians whose average value is less than a first preset threshold;
[0046] A third determination sub-module, configured to determine depth information of the target scene according to the three-dimensional Gaussian model after secondary update and the multi-view image sequence.
[0047] Optionally, the second determination module specifically includes: a fourth determination sub-module, a fifth determination sub-module, a sixth determination sub-module, and a seventh determination sub-module;
[0048] The fourth determination sub-module is configured to generate an opacity map corresponding to the rendered image, where the opacity map represents the opacity of each pixel in the rendered image;
[0049] The fifth determination sub-module is configured to, for pixels with an opacity less than a second preset threshold, determine a target pixel that is closest to the pixel and has an opacity greater than the second preset threshold;
[0050] The sixth determination sub-module is configured to obtain a three-dimensional Gaussian model after tertiary update by back-projecting the target pixel in the updated three-dimensional Gaussian model;
[0051] The seventh determination sub-module is configured to determine depth information of the target scene according to the three-dimensional Gaussian model after tertiary update and the multi-view image sequence.
[0052] Optionally, the attribute information of the two-dimensional Gaussian includes color attribute information, geometric attribute information, semantic attribute information, and opacity attribute information; the determination units of the attribute information of the two-dimensional Gaussian are specifically as follows:
[0053] The first determination unit is configured to determine a color feature vector and a semantic feature vector of the target scene;
[0054] The second determination unit is configured to respectively input the color feature vector into a color fully-connected network and an opacity fully-connected network of the initialized three-dimensional Gaussian model to respectively obtain the color attribute information and the opacity attribute information;
[0055] The third determination unit is configured to respectively input the semantic feature vector into a geometric fully-connected network and a semantic fully-connected network of the initialized three-dimensional Gaussian model to respectively obtain the geometric attribute information and the opacity attribute information.
[0056] Optionally, the determination unit of the position information of the two-dimensional Gaussian is as follows:
[0057] The fourth determination unit is configured to generate a sparse point cloud of the target scene according to a multi-view image sequence of the target scene;
[0058] A fifth determination unit, configured to determine the anchor positions of the three-dimensional anchor Gaussians in the initialized three-dimensional Gaussian model according to the coordinates of the sparse point cloud and a preset distance between Gaussian anchors.
[0059] A sixth determination unit, configured to determine the position information of the two-dimensional Gaussian according to the anchor positions and the position offsets.
[0060] Optionally, the formulas of the depth loss function and the normal loss function are respectively as follows:
[0061]
[0062] where, L d-smooth is the depth loss function, is the depth image, w h is the cosine similarity of the semantic features of adjacent pixels in the vertical direction, d(x i,j ) is the pixel value at the position (i, j) in the depth image, w w is the cosine similarity of the semantic features of adjacent pixels in the horizontal direction, L n-smooth is the normal loss function, is the normal image, n(x i,j ) is the pixel value at the position (i, j) in the normal image.
[0063] In a third aspect, the present application discloses a scene surface reconstruction device, where the device includes: a memory and a processor;
[0064] The memory is configured to store a program;
[0065] The processor is configured to execute the program to implement each step of the scene surface reconstruction method as described in the first aspect.
[0066] In a fourth aspect, the present application discloses a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, each step of the scene surface reconstruction method as described in the first aspect is implemented.
[0067] Compared with the prior art, the present application has the following beneficial effects:
[0068] The embodiments of the present application provide a method, apparatus, device and medium for scene surface reconstruction. Firstly, an initialized three-dimensional Gaussian model is introduced, and the attribute information and position information of the two-dimensional Gaussian corresponding to the three-dimensional anchor Gaussian are determined in combination with the multi-view image sequence of the target scene. Secondly, the rendered image is generated by the Gaussian splashing method using the attribute information and position information. Subsequently, the initialized three-dimensional Gaussian model is updated by minimizing the rendering loss function according to the rendered image and the image in the multi-view image sequence that has the same view as the rendered image. Finally, the depth information is determined according to the updated three-dimensional Gaussian model and the multi-view image sequence, and the surface model of the target scene is constructed by a three-dimensional reconstruction algorithm according to the depth information and the rendered image, which can more accurately reflect the structure and details of the target scene, avoid the appearance of floating objects, and significantly improve the quality and accuracy of scene reconstruction. Description of the Drawings
[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0070] Figure 1 It is a flowchart of a method for scene surface reconstruction provided by an embodiment of the present application;
[0071] Figure 2 It is a flowchart of another method for scene surface reconstruction provided by an embodiment of the present application;
[0072] Figure 3 It is a schematic diagram of a device for scene surface reconstruction provided by an embodiment of the present application;
[0073] Figure 4 It is a schematic diagram of a computer-readable medium provided by an embodiment of the present application. Detailed Embodiments
[0074] First of all, the technical terms involved in the present application will be explained:
[0075] The 3D Gaussian Splatting (3DGS) technique is an advanced rendering method used in computer vision and graphics, especially suitable for reconstructing high-quality 3D scenes from a set of sparse 2D images. The 3DGS technique utilizes the Gaussian distribution to represent each local feature in the scene and renders by "splatting" these Gaussian functions onto the image plane, thereby generating realistic views. In the 3DGS technique, each 3D Gaussian sphere represents a local feature in the scene, such as color, shape, etc. It is defined by multiple parameters including position, size (standard deviation), orientation, etc. These Gaussian spheres are used to approximately describe the object surface in the scene.
[0076] As described above, in practical applications, if there are weakly textured regions in the real image where the pixel values change very smoothly or are almost constant, or there are poorly lit regions, the finally reconstructed scene may not be accurately aligned with the real object. This may cause the 3D Gaussian spheres to float in the air, forming incorrect "floating objects", resulting in a lower quality of the reconstructed scene.
[0077] After research, the inventors proposed a method, device, equipment, and medium for scene surface reconstruction. First, an initialized 3D Gaussian model is introduced, and combined with the multi-view image sequence of the target scene, the attribute information and position information of the 2D Gaussian corresponding to the 3D anchor Gaussian are determined. Second, using the attribute information and position information, a rendered image is generated through the Gaussian splatting method. Subsequently, according to the rendered image and the image in the multi-view image sequence with the same view as the rendered image, the initialized 3D Gaussian model is updated by minimizing the rendering loss function. Finally, the depth information is determined based on the updated 3D Gaussian model and the multi-view image sequence, and based on the depth information and the rendered image, the surface model of the target scene is constructed through a 3D reconstruction algorithm, which can more accurately reflect the structure and details of the target scene, avoid the appearance of floating object phenomena, and significantly improve the quality and accuracy of scene reconstruction. Further, based on the updated 3D Gaussian model, the method for scene surface reconstruction further determines the depth information of the target scene through various strategies (such as averaging the alpha blending weights and using the opacity map), which helps to eliminate inaccurate depth information and retain more reliable depth data, also significantly improving the quality and accuracy of scene reconstruction.
[0078] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0079] See Figure 1 , which is a flowchart of a scene surface reconstruction method provided by an embodiment of this application. The method includes:
[0080] S101: When obtaining an initialized three-dimensional Gaussian model according to a multi-view image sequence of a target scene, determine the attribute information and position information of the two-dimensional Gaussian corresponding to the three-dimensional anchor Gaussian in the initialized three-dimensional Gaussian model.
[0081] According to a multi-view image sequence of a target scene (the scene that needs to be surface-reconstructed), use the Structure From Motion (SFM) method to generate an initialized three-dimensional Gaussian model. Among them, the initialized three-dimensional Gaussian model includes several three-dimensional anchor Gaussians. A three-dimensional anchor Gaussian refers to a Gaussian distribution centered on a set of anchor points in three-dimensional space. Each anchor point represents a point in three-dimensional space, and the Gaussian distribution describes the uncertainty of this point in space.
[0082] Subsequently, for each three-dimensional anchor Gaussian in the initialized three-dimensional Gaussian model, find its projection on the two-dimensional image, that is, the two-dimensional Gaussian, through the camera parameters and the three-dimensional to two-dimensional projection relationship. Then, determine the attribute information and position information of the two-dimensional Gaussian. Among them, the attribute information of the two-dimensional Gaussian includes color attribute (describing the color information of the two-dimensional Gaussian), geometric attribute (describing the shape and size of the two-dimensional Gaussian), semantic attribute (describing the semantic information of the two-dimensional Gaussian, such as object category, material, etc.), and opacity attribute (describing the transparency of the two-dimensional Gaussian), etc.
[0083] S102: Generate a rendered image by the Gaussian splatting method according to the attribute information and position information of the two-dimensional Gaussian, where the rendered image includes color information, semantic information, depth information, and normal information.
[0084] The rendered image generated by the Gaussian splatting method includes color information (the color value of each pixel), semantic information (the semantic label of the object category of each pixel), depth information (the depth value of each pixel), and normal information (the surface normal direction of each pixel).
[0085] It can be understood that in S102, it is an example that the color channel, semantic label channel, depth channel, and normal channel are first generated respectively, and then the contributions of all two-dimensional Gaussians (that is, the color information in the color channel, the semantic information in the semantic label channel, the depth information in the depth channel, and the normal information in the normal channel) are accumulated to generate a final rendered image. In practical applications, it is also possible to not accumulate the contributions of all two-dimensional Gaussians respectively, but to generate separate rendered images by processing each channel separately Semantic feature image Depth image and normal image (i.e., four images). This application does not make any limitations on this.
[0086] S103: According to the rendered image and the image in the multi-view image sequence that has the same view angle as the rendered image, update the initialized three-dimensional Gaussian model by minimizing the rendering loss function to obtain an updated three-dimensional Gaussian model, where the rendering loss function includes a color loss function, a semantic loss function, a depth loss function, and a normal loss function.
[0087] In some specific implementation manners, the color loss function can be as shown in the following formula (1):
[0088]
[0089] Among them, is the color loss function, λ is the weight, the loss measures the absolute difference between the pixel values of the rendered image and the scene image in the multi-view image sequence that has the same view angle as the rendered image, while the loss takes into account the structural similarity between the rendered image and the scene image in the multi-view image sequence that has the same view angle as the rendered image.
[0090] It can be understood that λ is used to adjust the proportion of the L1 loss and the D-SSIM loss in the total color loss, and is usually set to 0.7 (this means that the L1 loss occupies a relatively large proportion in the total color loss, but the D-SSIM loss is also taken into account to balance the brightness / color difference and structural similarity of the image). For the specific value of λ, this application does not make any limitations.
[0091] In some specific implementation manners, the semantic loss function can be as shown in the following formula (2):
[0092]
[0093] Among them, is the semantic loss function, is the semantic information, v is the true category of the pixel, i is all categories of the pixel, and this formula (2) is the cross-entropy loss function.
[0094] In some specific implementation manners, the depth loss function and the normal loss function can be as shown in the following formulas (3) and (4) respectively:
[0095]
[0096] Among them, L d-smooth is the depth loss function, is the depth information, wh is the cosine similarity of the semantic features of adjacent pixels in the vertical direction, and the specific calculation method is shown in the following formula (5), where d(x i,j ) is the pixel value at position (i, j) in the depth image, and w w is the cosine similarity of the semantic features of adjacent pixels in the horizontal direction, and the specific calculation method is shown in the following formula (6). L n-smooth is the normal loss function, is the normal information, and n(x i,j ) is the pixel value at position (i, j) in the normal image.
[0097]
[0098] where S i,j is the semantic feature of pixel (i, j).
[0099] After obtaining the color loss value, semantic loss value, depth loss value, and normal loss value through the above formulas, the color loss value, semantic loss value, depth loss value, and normal loss value can be minimized through gradient descent or other optimization algorithms, thereby updating the initialized three-dimensional Gaussian model to obtain an updated three-dimensional Gaussian model. This updated three-dimensional Gaussian model can more accurately represent the three-dimensional structure, color information, semantic information, depth information, and normal information of the target scene.
[0100] It can be understood that by minimizing the depth loss value and normal loss value, the Gaussians of the same object can have similar depths and smooth normals through semantic constraints, thereby improving the rendering quality.
[0101] S104: Determine the depth information of the target scene according to the updated three-dimensional Gaussian model and the multi-view image sequence.
[0102] Depth information reveals the distance of each point in the target scene relative to the observer (or camera). The updated three-dimensional Gaussian model provides the three-dimensional structure information of the target scene, while the multi-view image sequence provides the observation data from different perspectives. By combining the two, the depth information can be estimated more accurately.
[0103] S105: Construct the surface model of the target scene through a three-dimensional reconstruction algorithm according to the depth information and the rendered image.
[0104] The surface model of the target scene is usually a three-dimensional mesh (such as a triangular mesh, a quadrilateral mesh, etc.), which approximately represents the surface of the objects in the target scene. Specifically, according to the depth value of each pixel in the depth information and the position of these pixels in the rendered image, their coordinates in the three-dimensional space can be calculated. Then, these points are connected to form a triangular mesh or a quadrilateral mesh to approximate the surface of the object, thereby constructing the surface model of the target scene.
[0105] It can be understood that since the rendered image includes semantic information, the semantic information enables the surface model of the target scene to understand and distinguish different objects and regions, which helps to more accurately identify and operate the objects in the scene in subsequent applications, thereby enhancing the overall understanding of the scene. Also, during the three-dimensional reconstruction process, due to factors such as occlusion and noise, ambiguities or errors may occur. The semantic information can help to more accurately judge and handle these situations, thereby improving the geometric accuracy and detail performance of the model.
[0106] In summary, the present application discloses a method for reconstructing a scene surface. First, an initialized three-dimensional Gaussian model is introduced, and in combination with a multi-view image sequence of the target scene, the attribute information and position information of the two-dimensional Gaussian corresponding to the three-dimensional anchor Gaussian are determined. Second, using the attribute information and position information, a rendered image is generated by the Gaussian splashing method. Subsequently, according to the rendered image and the image in the multi-view image sequence with the same view as the rendered image, the initialized three-dimensional Gaussian model is updated by minimizing the rendering loss function. Finally, the depth information is determined according to the updated three-dimensional Gaussian model and the multi-view image sequence, and according to the depth information and the rendered image, the surface model of the target scene is constructed by a three-dimensional reconstruction algorithm, which can more accurately reflect the structure and details of the target scene, avoid the appearance of floating objects, and significantly improve the quality and accuracy of scene reconstruction.
[0107] See Figure 2 , which is a flowchart of another method for reconstructing a scene surface provided by an embodiment of the present application. The method includes:
[0108] S201: Obtain a multi-view image sequence of the target scene.
[0109] In a specific implementation, electronic devices such as mobile phones and cameras can be used to collect a series of scene images of the target scene to obtain a multi-view image sequence I = {I k |k = 1, 2, … N}. It should be noted that the above scene images can be images obtained by a series of electronic devices at different orientations simultaneously shooting the target scene, or images obtained by a single electronic device shooting the target scene during the movement process. The present application does not limit the specific shooting method.
[0110] In another specific implementation, electronic devices such as mobile phones and cameras can be used to collect the scene video of the target scene, and through frame extraction processing on the scene video, a multi-view image sequence I = {I k | k = 1, 2, … N} is obtained. Exemplarily, after collecting the scene video of the target scene, the OpenCV software can be used to perform frame extraction processing on the scene video. For example, the images at an interval of 5 frames can be selected as key frame images, and the key frame image sequence is used as the multi-view image sequence. It should be noted that the images at intervals of 3 frames, 10 frames, and other numbers of frames can also be selected as key frame images, and the specific number of frames is not limited in this application.
[0111] It should also be noted that all the images included in the multi-view image sequence are images of the same target scene from different angles. And all the above images can be RGB images or images in other formats, and the specific image format is not limited in this application.
[0112] S202: Preprocess the multi-view image sequence to obtain camera parameters, the sparse point cloud of the target scene, and the semantic segmentation map of each scene image.
[0113] Camera parameters can be divided into camera internal parameters, camera external parameters, near-plane parameters, and far-plane parameters, etc. Among them, camera internal parameters represent the fixed parameters of the camera, usually a 3×3 matrix, including focal length, principal point, etc. Camera external parameters represent the rotation and position of the current camera relative to the world coordinate system, usually a 4×4 matrix, including rotation matrix and translation vector. Near-plane parameters and far-plane parameters define the depth range of the target scene, which helps to determine which parts should be included in the scene surface reconstruction process. The sparse point cloud represents the key feature points and their three-dimensional coordinates in the target scene. In one specific implementation, the COLMAP software can be used to perform sparse reconstruction on the multi-view image sequence to obtain camera parameters and the sparse point cloud.
[0114] The semantic segmentation map is to classify each scene image at the pixel level to obtain information on which pixel belongs to which object category (such as walls, furniture, etc.). In one specific implementation, the Grounded SAM model can be used to perform semantic segmentation on each scene image in the multi-view image sequence to obtain the semantic segmentation map of each scene image.
[0115] S203: Generate an initial three-dimensional Gaussian model based on the sparse point cloud of the target scene.
[0116] The initial three-dimensional Gaussian model includes several three-dimensional anchor Gaussians. Each three-dimensional anchor Gaussian includes the following information: anchor position P anchor(Indicating the position of the Gaussian sphere in three-dimensional space), color feature vector f c (Describing the color information of the Gaussian sphere), semantic feature vector f s (Describing the semantic information of the object category to which the Gaussian sphere belongs), and learnable position offset p offset (Each position offset is a three-dimensional vector. The position offset allows the position of the Gaussian sphere to be adjusted during training to better fit the scene).
[0117] The initialized three-dimensional Gaussian model also includes four fully connected networks (MultilayerPerceptron, MLP), including the color fully connected network MLP color 、geometric fully connected network MLP geometry 、semantic fully connected network MLP semantic , and the opacity fully connected network MLP opacity .
[0118] S204: Filter the three-dimensional anchor Gaussians in the initialized three-dimensional Gaussian model to obtain the filtered three-dimensional anchor Gaussians.
[0119] For each three-dimensional anchor Gaussian in the initialized scene model of the target scene, calculate its radius radii projected onto the image plane. Among them, the formula for calculating radii can be shown as the following formula (7):
[0120] radii = 3 * max(δ1, δ2) (7)
[0121] Among them, radii is the radius of the three-dimensional anchor Gaussian projected onto the image plane, and δ1 and δ2 are the eigenvalues of the covariance matrix Σ′ of the two-dimensional Gaussian projection on the image plane. Specifically, the formula for the covariance matrix Σ′ can be shown as the following formula (8):
[0122] Σ′ = JWRXX T R T W T J T (8)
[0123] Among them, Σ′ is the covariance matrix, J is the Jacobian matrix, W is the transformation matrix from world coordinates to image plane coordinates, R is the identity matrix (obtained from the quaternion r), and X is a 3×3 matrix representing the diagonal matrix diag(D anchor , D anchor , D anchor ). Specifically, J and W are calculated from the camera parameters (including the camera intrinsic parameters and the camera extrinsic parameters).
[0124] If radii ≤ 0, then the three-dimensional anchor Gaussian does not participate in the subsequent rendering process. Thus, the efficiency of the algorithm can be significantly improved because unnecessary calculations for invalid Gaussian spheres are not required.
[0125] S205: Determine the attribute information and position information of the two-dimensional Gaussian corresponding to the filtered three-dimensional anchor Gaussian according to the fully connected network, where the attribute information of the two-dimensional Gaussian includes color attribute information, geometric attribute information, semantic attribute information, and opacity attribute information.
[0126] In some specific implementation manners, the determination method of the position information μ of the two-dimensional Gaussian is specifically as follows:
[0127] First, initialize the anchor position P according to the preset distance D between Gaussian anchors anchor (for example, 0.5 unit length), and the specific formula is as shown in formula (9) below: anchor
[0128]
[0129] where P anchor is the anchor position, and the Unique function represents deleting the same anchor position P anchor , represents rounding down, and P colmap is the coordinate of the sparse point cloud. It can be seen from formula (9) that the anchor position P anchor is an integer multiple of the distance D between Gaussian anchors anchor . Thus, through the preset distance D between Gaussian anchors anchor and the COLMAP sparse point cloud coordinate P colmap , the unique and uniformly distributed anchor position P anchor can be calculated and de-duplicated.
[0130] Subsequently, according to the above anchor position P anchor and the learnable position offset p offset , the position information μ of the two-dimensional Gaussian can be obtained, and the specific formula is as shown in formula (10) below:
[0131] μ = p offset + P anchor (10)
[0132] In some specific implementation manners, the determination method of the attribute information of the two-dimensional Gaussian is specifically as follows:
[0133] First, input the color feature vector f c into the color fully connected network MLP color , and output a three-dimensional vector c (representing RGB color), and the specific formula is as shown in formula (11) below:
[0134] c = MLP color (f c ) (11)
[0135] Secondly, input the semantic feature vector f s into the geometric fully-connected network MLP geometry to output a four-dimensional vector r (representing the quaternion for rotation) and a two-dimensional vector s (representing the two scale factors s1, s2 of the elliptical patch). The specific formula is shown as formula (12) below:
[0136] r, s = MLP geometry (f s )(12)
[0137] Subsequently, input the color feature vector f c into the opacity fully-connected network MLP opacity to output a scalar o (representing opacity). The specific formula is shown as formula (13) below:
[0138] o = MLP opacity (f c ) (13)
[0139] Finally, input the semantic feature vector f s into the semantic fully-connected network MLP semantic to output a sixty-four-dimensional multi-dimensional vector S (representing semantic features). The specific formula is shown as formula (14) below:
[0140] S = MLP semantic (f s )(14)
[0141] Thus, color attributes, geometric attributes, semantic attributes, and opacity attributes (i.e., the attribute information of the two-dimensional Gaussian) can be generated through different MLP networks respectively.
[0142] S206: Generate a rendered image, a semantic feature image, a depth image, and a normal image according to the position information of the two-dimensional Gaussian and the attribute information of the two-dimensional Gaussian through the Gaussian splashing algorithm.
[0143] In some specific implementation manners, the formulas for generating the rendered image the semantic feature image the depth image and the normal image can be respectively shown as formulas (15) - (18) below:
[0144]
[0145] Among them, c iis the color of the i-th 2D Gaussian, and α i is the opacity factor, and T i is the cumulative transmittance, and S i is the semantic feature of the i-th 2D Gaussian, and d i is the intersection depth of the pixel-emitted ray and the 2D Gaussian, and n i is the normal of the 2D Gaussian. Among them, the opacity factor α i = o i G i (u), where u = [u, v] T is the intersection of the ray back-projected from pixel x on the corresponding 2D Gaussian, is the Gaussian distribution function.
[0146] S207: Determine the color loss value by calculating the L1 loss function and the D-SSIM structural dissimilarity loss function based on the rendered image and the scene image with the same view as the rendered image in the multi-view image sequence.
[0147] It can be understood that the steps of S207 - S210 are similar to the steps of S103, and will not be elaborated here.
[0148] S208: Determine the semantic loss value by calculating the per-pixel cross-entropy loss function based on the semantic feature image and the semantic segmentation map of the scene image with the same view as the rendered image in the multi-view image sequence.
[0149] S209: Determine the depth loss value and the normal loss value by calculating the depth loss function and the normal loss function respectively based on the depth image and the normal image.
[0150] It can be understood that the depth smoothness loss value and the normal smoothness loss value respectively make the regions with similar semantic features have similar values on the rendered depth image and normal image.
[0151] S210: Update the initialized 3D Gaussian model by minimizing the color loss value, semantic loss value, depth loss value, and normal loss value, and obtain the updated 3D Gaussian model.
[0152] S211: Determine the depth information of the target scene based on the updated 3D Gaussian model and the multi-view image sequence.
[0153] During the training process, it is necessary to control the density of the 3D anchors:
[0154] In some specific implementation manners, first, determine the average value of the alpha blending weights of the 2D Gaussians corresponding to each 3D anchor Gaussian in the updated 3D Gaussian model. Specifically, the alpha blending weights of the 2D Gaussians corresponding to the 3D anchor Gaussians can be shown as the following formula (19):
[0155]
[0156] Among them, w g is the alpha blending weight, N k is the number of pictures in which the Gaussian can be observed, α k is the opacity factor, T k is the cumulative transmittance.
[0157] Subsequently, by deleting the three-dimensional anchor Gaussian corresponding to the two-dimensional Gaussian whose average value is less than the first preset threshold, a three-dimensional Gaussian model after secondary update is obtained. Delete the three-dimensional anchors less than the first preset threshold τ d This means that this anchor is very likely to represent floating objects or noise in the air, so it will be deleted.
[0158] Finally, according to the three-dimensional Gaussian model after secondary update and the multi-view image sequence, the depth information of the target scene is determined. The depth information may be a depth image. For this, the present application does not make a limitation.
[0159] In some other specific implementation manners, in order to maintain the density and accuracy of the target scene, first, an opacity map corresponding to the rendered image is generated Among them, the opacity map represents the opacity of each pixel in the rendered image. Since there will be no infinite distance situation in the target scene, there should be no value close to 0 on the opacity map. Therefore, for the pixels with opacity less than the second preset threshold (that is, there is almost no Gaussian coverage in the area where the light passes through), determine the target pixel that is closest to the pixel and has an opacity greater than the second preset threshold. Subsequently, by performing back-projection on the target pixel in the updated three-dimensional Gaussian model, a three-dimensional Gaussian model after tertiary update is obtained. Finally, according to the three-dimensional Gaussian model after tertiary update and the multi-view image sequence, the depth information of the target scene is determined. The depth information may be a depth image. For this, the present application does not make a limitation.
[0160] It should be noted that for the above two methods for controlling the density of the three-dimensional anchors, both can be executed, or one can be selected for execution. For this, the present application does not make a limitation.
[0161] S212: According to the depth information and the rendered image, construct a surface model of the target scene through a three-dimensional reconstruction algorithm.
[0162] Specifically, based on the depth values of each pixel in the depth information and the positions of these pixels in the rendered image, their coordinates in the three-dimensional space can be calculated. Then, through Delaunay triangulation, Poisson surface reconstruction, or the Marching Cubes algorithm, these points are connected to form a triangular mesh or a quadrilateral mesh to approximate the surface of the object, thereby constructing a surface model of the target scene.
[0163] It should be noted that, to further improve the accuracy of the surface model of the target scene, the patches in the triangular mesh or the quadrilateral mesh can be screened, for example, selecting the top 500 clusters with the largest number of patches to eliminate inaccurate surfaces and floating objects in the air.
[0164] It can be understood that, similar to the first embodiment, since the rendered image includes semantic information, the semantic information enables the surface model of the target scene to understand and distinguish different objects and regions, which helps to more accurately identify and operate the objects in the scene in subsequent applications, thereby enhancing the overall understanding of the scene. Also, during the three-dimensional reconstruction process, due to factors such as occlusion and noise, ambiguities or errors may occur. The semantic information can help to more accurately judge and handle these situations, thereby improving the geometric accuracy and detail performance of the model.
[0165] In summary, the present application discloses a method for reconstructing the surface of a scene. First, an initialized three-dimensional Gaussian model is introduced, and in combination with a multi-view image sequence of the target scene, the attribute information and position information of the two-dimensional Gaussian corresponding to the three-dimensional anchor Gaussian are determined. Secondly, using the attribute information and position information, a rendered image is generated through the Gaussian splashing method. Subsequently, based on the rendered image and the image in the multi-view image sequence with the same view as the rendered image, the initialized three-dimensional Gaussian model is updated by minimizing the rendering loss function. Finally, based on the updated three-dimensional Gaussian model and the multi-view image sequence, the depth information is determined, and based on the depth information and the rendered image, a surface model of the target scene is constructed through a three-dimensional reconstruction algorithm, which can more accurately reflect the structure and details of the target scene, avoid the appearance of floating object phenomena, and significantly improve the quality and accuracy of scene reconstruction. Further, on the basis of the updated three-dimensional Gaussian model, the method for reconstructing the surface of the scene further determines the depth information of the target scene through various strategies (such as screening the average value of alpha blending weights, using the opacity map), which helps to eliminate inaccurate depth information and retain more reliable depth data, and also significantly improves the quality and accuracy of scene reconstruction.
[0166] See Figure 3, this figure shows a scene surface reconstruction device provided by an embodiment of the present application. The scene surface reconstruction device 300 includes: a first determination module 301, an image generation module 302, a model update module 303, a second determination module 304, and a model construction module 305.
[0167] The first determination module 301 is configured to determine the attribute information and position information of the two-dimensional Gaussian corresponding to the three-dimensional anchor Gaussian in the initialized three-dimensional Gaussian model when obtaining the initialized three-dimensional Gaussian model according to the multi-view image sequence of the target scene;
[0168] The image generation module 302 is configured to generate a rendered image by the Gaussian splashing method according to the attribute information and position information of the two-dimensional Gaussian, where the rendered image includes color information, semantic information, depth information, and normal information;
[0169] The model update module 303 is configured to update the initialized three-dimensional Gaussian model by minimizing the rendering loss function according to the rendered image and the image in the multi-view image sequence with the same view as the rendered image, and obtain the updated three-dimensional Gaussian model, where the rendering loss function includes a color loss function, a semantic loss function, a depth loss function, and a normal loss function;
[0170] The second determination module 304 is configured to determine the depth information of the target scene according to the updated three-dimensional Gaussian model and the multi-view image sequence;
[0171] The model construction module 305 is configured to construct a surface model of the target scene by a three-dimensional reconstruction algorithm according to the depth information and the rendered image.
[0172] In some specific implementation manners, the first determination module 301 is specifically configured to: determine the projection radius of each three-dimensional anchor Gaussian in the initialized three-dimensional Gaussian model by projecting each three-dimensional anchor Gaussian onto the image plane; determine the attribute information and position information of the two-dimensional Gaussian corresponding to the three-dimensional anchor Gaussian with a projection radius greater than 0.
[0173] In some specific implementation manners, the second determination module 304 specifically includes: a first determination sub-module, a second determination sub-module, and a third determination sub-module;
[0174] The first determination sub-module is configured to determine the average value of the alpha blending weights of the two-dimensional Gaussian corresponding to each three-dimensional anchor Gaussian in the updated three-dimensional Gaussian model;
[0175] The second determination sub-module is configured to obtain the three-dimensional Gaussian model after secondary update by deleting the three-dimensional anchor Gaussian corresponding to the two-dimensional Gaussian with an average value less than the first preset threshold;
[0176] A third determination sub-module, configured to determine the depth information of the target scene according to the three-dimensional Gaussian model after secondary update and the multi-view image sequence.
[0177] In some specific implementation manners, the second determination module 304 specifically includes: a fourth determination sub-module, a fifth determination sub-module, a sixth determination sub-module, and a seventh determination sub-module;
[0178] The fourth determination sub-module is configured to generate an opacity map corresponding to the rendered image, where the opacity map characterizes the opacity of each pixel in the rendered image;
[0179] The fifth determination sub-module is configured to, for pixels with an opacity less than a second preset threshold, determine a target pixel that is the closest to the pixel and has an opacity greater than the second preset threshold;
[0180] The sixth determination sub-module is configured to obtain a three-dimensional Gaussian model after tertiary update by performing back-projection on the target pixel in the updated three-dimensional Gaussian model;
[0181] The seventh determination sub-module is configured to determine the depth information of the target scene according to the three-dimensional Gaussian model after tertiary update and the multi-view image sequence.
[0182] In some specific implementation manners, the attribute information of the two-dimensional Gaussian includes color attribute information, geometric attribute information, semantic attribute information, and opacity attribute information; the determination units of the attribute information of the two-dimensional Gaussian are specifically as follows:
[0183] The first determination unit is configured to determine the color feature vector and the semantic feature vector of the target scene;
[0184] The second determination unit is configured to input the color feature vector into the color fully-connected network and the opacity fully-connected network of the initialized three-dimensional Gaussian model respectively, and obtain the color attribute information and the opacity attribute information respectively;
[0185] The third determination unit is configured to input the semantic feature vector into the geometric fully-connected network and the semantic fully-connected network of the initialized three-dimensional Gaussian model respectively, and obtain the geometric attribute information and the opacity attribute information respectively.
[0186] In some specific implementation manners, the determination unit of the position information of the two-dimensional Gaussian is as follows:
[0187] The fourth determination unit is configured to generate a sparse point cloud of the target scene according to the multi-view image sequence of the target scene;
[0188] The fifth determination unit is configured to determine the anchor position of the three-dimensional anchor Gaussian in the initialized three-dimensional Gaussian model according to the coordinates of the sparse point cloud and the preset distance between Gaussian anchor points;
[0189] A sixth determination unit, configured to determine the position information of the two-dimensional Gaussian according to the anchor position and the position offset.
[0190] In some specific implementation manners, the formulas of the depth loss function and the normal loss function are respectively shown in the following formulas (20) and (21):
[0191]
[0192]
[0193] Among them, L d-smooth is the depth loss function, is the depth image, w h is the cosine similarity of the semantic features of adjacent pixels in the vertical direction, d(x i,j ) is the pixel value at the position (i, j) in the depth image, w w is the cosine similarity of the semantic features of adjacent pixels in the horizontal direction, L n-smooth is the normal loss function, is the normal image, n(x i,j ) is the pixel value at the position (i, j) in the normal image.
[0194] In summary, the present application discloses a scene surface reconstruction device. First, an initialized three-dimensional Gaussian model is introduced, and in combination with a multi-view image sequence of the target scene, the attribute information and position information of the two-dimensional Gaussian corresponding to the three-dimensional anchor Gaussian are determined. Second, using the attribute information and position information, a rendering image is generated through a Gaussian splashing device. Subsequently, according to the rendering image and the image in the multi-view image sequence that has the same view as the rendering image, the initialized three-dimensional Gaussian model is updated by minimizing the rendering loss function. Finally, the depth information is determined according to the updated three-dimensional Gaussian model and the multi-view image sequence, and according to the depth information and the rendering image, a surface model of the target scene is constructed through a three-dimensional reconstruction algorithm, which can more accurately reflect the structure and details of the target scene, avoid the appearance of floating objects, and significantly improve the quality and accuracy of scene reconstruction. Further, based on the updated three-dimensional Gaussian model, the scene surface reconstruction device further determines the depth information of the target scene through various strategies (such as the average value screening of the alpha blending weight, the utilization of the opacity map), which helps to eliminate inaccurate depth information and retain more reliable depth data, and also significantly improves the quality and accuracy of scene reconstruction.
[0195] The embodiments of the present application also provide corresponding scene surface reconstruction devices and computer-readable media for implementing the scene surface reconstruction method provided by the embodiments of the present application.
[0196] Among them, the scene surface reconstruction device includes a memory and a processor. The memory is used to store instructions or codes, and the processor is used to execute the instructions or codes so that the device executes a scene surface reconstruction method according to any embodiment of the present application.
[0197] See Figure 4 , this figure is a schematic diagram of a computer-readable medium provided by an embodiment of the present application. A computer program 411 is stored on the computer-readable medium 400. When the computer program 411 is executed by a processor, it implements the steps of the above Figure 1 scene surface reconstruction method.
[0198] It should be noted that in the context of the present application, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0199] It should be noted that the machine-readable medium described above in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. And in this application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0200] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.
[0201] Although the subject matter has been described in language specific to structural features and / or method logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.
[0202] Although several specific implementation details are included in the above discussion, these should not be construed as limitations on the scope of this application. Certain features described in the context of separate embodiments can also be implemented combinatorially in a single embodiment. On the contrary, various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0203] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the present application that have similar functions.
Claims
1. A method for reconstructing a scene surface, characterized in that, The method includes: When obtaining an initialized three-dimensional Gaussian model based on a multi-view image sequence of a target scene, determining the attribute information and position information of the two-dimensional Gaussian corresponding to the three-dimensional anchor Gaussian in the initialized three-dimensional Gaussian model; Generating a rendered image by means of the Gaussian splashing method according to the attribute information and position information of the two-dimensional Gaussian, wherein the rendered image includes color information, semantic information, depth information, and normal information; Updating the initialized three-dimensional Gaussian model by minimizing a rendering loss function according to the rendered image and the image in the multi-view image sequence that has the same view angle as the rendered image, to obtain an updated three-dimensional Gaussian model, wherein the rendering loss function includes a color loss function, a semantic loss function, a depth loss function, and a normal loss function; Determining the depth information of the target scene according to the updated three-dimensional Gaussian model and the multi-view image sequence; Constructing a surface model of the target scene through a three-dimensional reconstruction algorithm according to the depth information and the rendered image.
2. The method according to claim 1, wherein The determining the attribute information and position information of the two-dimensional Gaussian corresponding to the three-dimensional anchor Gaussian in the initialized three-dimensional Gaussian model includes: Determining the projection radius of each three-dimensional anchor Gaussian in the initialized three-dimensional Gaussian model by projecting it onto the image plane; Determining the attribute information and position information of the two-dimensional Gaussian corresponding to the three-dimensional anchor Gaussian with a projection radius greater than 0.
3. The method according to claim 1, wherein The determining the depth information of the target scene according to the updated three-dimensional Gaussian model and the multi-view image sequence includes: Determining the average value of the alpha blending weights of the two-dimensional Gaussians corresponding to each three-dimensional anchor Gaussian in the updated three-dimensional Gaussian model; Obtaining a secondarily updated three-dimensional Gaussian model by deleting the three-dimensional anchor Gaussians corresponding to the two-dimensional Gaussians with an average value less than a first preset threshold; Determining the depth information of the target scene according to the secondarily updated three-dimensional Gaussian model and the multi-view image sequence.
4. The method according to claim 1, characterized in that, The determining the depth information of the target scene according to the updated three-dimensional Gaussian model and the multi-view image sequence includes: Generating an opacity map corresponding to the rendered image, wherein the opacity map represents the opacity of each pixel in the rendered image; For a pixel with an opacity less than a second preset threshold, determining a target pixel that is the closest to the pixel and has an opacity greater than the second preset threshold; Obtaining a tertially updated three-dimensional Gaussian model by performing back-projection on the target pixel in the updated three-dimensional Gaussian model; Determining the depth information of the target scene according to the tertially updated three-dimensional Gaussian model and the multi-view image sequence.
5. The method according to claim 1, wherein The attribute information of the two-dimensional Gaussian includes color attribute information, geometric attribute information, semantic attribute information, and opacity attribute information; the determination method of the attribute information of the two-dimensional Gaussian is as follows: Determining the color feature vector and semantic feature vector of the target scene; Input the color feature vector into the color fully-connected network and the opacity fully-connected network of the initialized three-dimensional Gaussian model respectively to obtain the color attribute information and the opacity attribute information respectively; Input the semantic feature vector into the geometry fully-connected network and the semantic fully-connected network of the initialized three-dimensional Gaussian model respectively to obtain the geometry attribute information and the opacity attribute information respectively.
6. The method according to claim 1, wherein The determination method of the position information of the two-dimensional Gaussian is as follows: Generate the sparse point cloud of the target scene according to the multi-view image sequence of the target scene; Determine the anchor position of the three-dimensional anchor Gaussian in the initialized three-dimensional Gaussian model according to the coordinates of the sparse point cloud and the preset distance between Gaussian anchors; Determine the position information of the two-dimensional Gaussian according to the anchor position and the position offset.
7. The method according to any one of claims 1-6, characterized in that, The formulas of the depth loss function and the normal loss function are respectively shown as follows: Among them, L d-smooth is the depth loss function, is the depth information, w h is the cosine similarity of the semantic features of adjacent pixels in the vertical direction, d(x i,j ) is the pixel value at position (i, j) in the depth image, w w is the cosine similarity of the semantic features of adjacent pixels in the horizontal direction, L n-smooth is the normal loss function, is the normal image, n(x i,j ) is the pixel value at position (i, j) in the normal image.
8. A scene surface reconstruction device, characterized in that The device includes: a first determination module, an image generation module, a model update module, a second determination module, and a model construction module; The first determination module is used to determine the attribute information and the position information of the two-dimensional Gaussian corresponding to the three-dimensional anchor Gaussian in the initialized three-dimensional Gaussian model when obtaining the initialized three-dimensional Gaussian model according to the multi-view image sequence of the target scene; The image generation module is used to generate a rendered image by the Gaussian splashing method according to the attribute information and the position information of the two-dimensional Gaussian, wherein the rendered image includes color information, semantic information, depth information, and normal information; The model update module is used to update the initialized three-dimensional Gaussian model by minimizing the rendering loss function according to the rendered image and the image in the multi-view image sequence with the same view angle as the rendered image to obtain an updated three-dimensional Gaussian model, wherein the rendering loss function includes a color loss function, a semantic loss function, a depth loss function, and a normal loss function; The second determination module is used to determine the depth information of the target scene according to the updated three-dimensional Gaussian model and the multi-view image sequence; The model construction module is used to construct the surface model of the target scene by a three-dimensional reconstruction algorithm according to the depth information and the rendered image.
9. A scene surface reconstruction device, characterized in that, The device includes: a memory and a processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the scene surface reconstruction method according to any one of claims 1 to 7.
10. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements each step of the scene surface reconstruction method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Scene three-dimensional reconstruction method based on prior depth and Gaussian sputtering model fusion
CN118351252A
Three-dimensional scene reconstruction method and electronic equipment
CN118365805A
Sparse visual angle three-dimensional reconstruction method based on depth prior information
CN118657888A
Three-dimensional scene reconstruction method and device, equipment, medium and program product
CN118823234A
Asteroid surface fine three-dimensional reconstruction method based on 3D Gaussian
CN118864767A
Cited By
3dgs reconstruction noise restoration method, apparatus and device, and medium
CN121147054A
Interactive three-dimensional teaching resource generation system and method based on 3D Gaussian splashing
CN121170164A
Remote sensing image three-dimensional reconstruction method based on semantic information
CN121259211A
A method for three-dimensional reconstruction of remote sensing images based on semantic information
CN121259211B
Urban scene-oriented multi-source remote sensing data fusion three-dimensional reconstruction method
CN121259212A