3D GS scene modeling method and device combining depth estimation and exposure correction

By combining depth estimation and exposure correction, the problems of ‘background collapse’ and ‘floating artifacts’ in complex scenes are solved, achieving high-quality visual effects and efficient rendering.

CN120198583APending Publication Date: 2025-06-24ACADEMY OF BROADCASTING SCI STATE ADMINISTATION OF PRESS PUBLICATION RADIO FILM & TELEVISION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510264118.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In complex scenes, the focal length error in camera pose estimation is large, resulting in ‘background collapse’ and ‘floating artifact’ problems, affecting the visual quality and rendering speed of scene modeling.

Method used

Combining the 3DGS scene modeling method of depth estimation and exposure correction, the viewing angle depth map is obtained through the monocular depth estimation model, and the loss constraint is performed with the rendered depth map. The exposure correction model is introduced to compensate for the brightness changes using two exposure coefficients, and a post-processing pruning scheme is used to clear floating artifacts.

Benefits of technology

It realizes high-fidelity scene details reconstruction, removes wrong floating objects in the scene, corrects the overexposure of viewing angles, and improves visual viewing effect and rendering speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198583A_ABST
    Figure CN120198583A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of scene modeling, in particular to a 3DGS scene modeling method combining depth estimation and exposure correction, which introduces depth constraint on the basis of the traditional 3DGS, that is, depth maps of various visual angles are obtained through a monocular depth estimation model, and the depth of the 3D GS is obtained. Making loss between the modeling scene and the rendering depth map of the corresponding camera pose to constrain scene optimization; meanwhile, aiming at the brightness deviation during shooting at different visual angles, an exposure correction model is introduced, and two luminosity coefficients are used for compensating the overall change of the image brightness; besides, in order to further eliminate floating artifacts in the scene, a post-processing pruning scheme is adopted, the floating artifacts are removed by utilizing explicit representation of three-dimensional Gaussian, and the model is encouraged to correctly learn the areas again; from a modeling result, the method provided by the invention can perform three-dimensional modeling on a shooting scene in a high-fidelity manner, and can perform unified adjustment on scene brightness and delete unfriendly floating artifacts in the scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of scene modeling, and in particular, to a 3DGS scene modeling method combining depth estimation and exposure correction. Background Art

[0002] Scene reconstruction is of great significance in many fields, including autonomous driving, aerial mapping, and virtual reality. These applications have extremely high requirements for realistic visual effects and real-time rendering. Although some methods attempt to extend Neural Radiance Fields (NeRF) to large-scale scenes, they still face problems such as insufficient details or slow rendering speed. In recent years, as an emerging technology, 3D Gaussian Splatter (3DGS) has received extensive attention for its excellent performance in visual quality and rendering speed, and can achieve realistic real-time rendering at 1080p resolution. In addition, this technology has also been applied to scene reconstruction and 3D content generation. However, these methods are usually applicable to object-centered or simple scenes, and there will be several scalability problems when facing complex 360° surround scenes.

[0003] Facing complex scenes, the focal length error in camera pose estimation is relatively large, and some areas may not be fully covered. During the scene modeling process, there may be a problem of "background collapse", that is, the background is wrongly represented as a position close to the camera by similar artifacts. At the same time, there is usually uneven illumination in the scene, which will cause "floating artifacts" in the scene, that is, low-transparency and high-density floating areas distributed in the scene. For example, bright spots are usually close to the camera with higher exposure, while dark spots are related to images with lower exposure. However, when viewed from a new perspective, these Gaussian spots will appear to float in the air in the form of "floating artifacts", bringing an unnatural visual effect.

[0004] In view of this, how to provide a 3DGS scene modeling method combining depth estimation and exposure correction, which can remove the wrong floating objects in the scene and correct the overexposure phenomenon of some perspectives on the basis of reconstructing high-fidelity scene details, bringing a high-quality visual viewing effect, has become an urgent technical problem to be solved currently. Summary of the Invention

[0005] Embodiments of this application provide a 3DGS scene modeling method combining depth estimation and exposure correction, a 3DGS scene modeling device combining depth estimation and exposure correction, a computer device, and a computer storage medium, which are used to remove the wrong floating objects in the scene and correct the overexposure phenomenon of some perspectives on the basis of reconstructing high-fidelity scene details, bringing a high-quality visual viewing effect.

[0006] In the first aspect of the embodiments of this application, a 3DGS scene modeling method combining depth estimation and exposure correction is provided, including:

[0007] Obtain scene pictures of different perspectives corresponding to the target reconstruction, perform camera pose estimation on the scene pictures based on an open-source algorithm, and determine the initial sparse point cloud corresponding to the target reconstruction scene and the camera poses and camera intrinsics corresponding to each of the scene pictures, wherein the overlapping regions between adjacent scene pictures satisfy a preset scene picture overlap threshold;

[0008] Adopt a pre-trained monocular depth estimation model to calculate the perspective depth maps corresponding to different perspectives of each scene picture, and convert the format of the perspective depth maps to the target format;

[0009] Initialize the initial sparse point cloud of the target reconstruction scene to generate three-dimensional Gaussian points, use a rasterizer, combine the camera pose and the camera intrinsics, project the three-dimensional Gaussian points onto a two-dimensional plane to obtain a rendered depth map, and calculate the loss value between the rendered depth map and the perspective depth map, and perform scene constraint based on the loss value;

[0010] Based on an exposure correction model, introduce two exposure coefficients into each of the scene pictures to perform exposure correction on each of the scene pictures;

[0011] Adopt a post-processing pruning scheme, utilize the explicit representation of three-dimensional Gaussian to remove floating artifacts in each of the scene pictures, and encourage the model to re-learn the regions for removing floating artifacts to perform three-dimensional Gaussian scene modeling.

[0012] In the second aspect of the embodiments of the present application, a 3DGS scene modeling device combining depth estimation and exposure correction is provided, including:

[0013] A determination module configured to obtain scene pictures of different perspectives corresponding to the target reconstruction, perform camera pose estimation on the scene pictures based on an open-source algorithm, and determine the initial sparse point cloud corresponding to the target reconstruction scene and the camera poses and camera intrinsics corresponding to each of the scene pictures, wherein the overlapping regions between adjacent scene pictures satisfy a preset scene picture overlap threshold;

[0014] A calculation module configured to adopt a pre-trained monocular depth estimation model to calculate the perspective depth maps corresponding to different perspectives of each scene picture, and convert the format of the perspective depth maps to the target format;

[0015] A constraint module configured to initialize the initial sparse point cloud of the target reconstruction scene to generate three-dimensional Gaussian points, use a rasterizer, combine the camera pose and the camera intrinsics, project the three-dimensional Gaussian points onto a two-dimensional plane to obtain a rendered depth map, and calculate the loss value between the rendered depth map and the perspective depth map, and perform scene constraint based on the loss value;

[0016] A correction module, configured to perform exposure correction on each of the scene pictures by introducing two exposure coefficients into each of the scene pictures based on an exposure correction model.

[0017] A clearing module, configured to adopt a post-processing pruning scheme and utilize an explicit representation of a three-dimensional Gaussian to clear floating artifacts in each of the scene pictures, and encourage the model to relearn the areas for clearing floating artifacts to perform three-dimensional Gaussian scene modeling.

[0018] In a third aspect of the embodiments of the present application, a computing device is provided, including:

[0019] A memory and a processor;

[0020] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above 3DGS scene modeling method combining depth estimation and exposure correction are implemented.

[0021] According to a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, which stores computer-executable instructions. When the instructions are executed by a processor, the steps of the above 3DGS scene modeling method combining depth estimation and exposure correction are implemented.

[0022] The present application provides a 3DGS scene modeling method combining depth estimation and exposure correction, including: obtaining scene pictures of different perspectives corresponding to a target reconstruction, performing camera pose estimation on the scene pictures based on an open-source algorithm, determining an initial sparse point cloud corresponding to the target reconstruction scene and the camera poses and camera intrinsics corresponding to each of the scene pictures, wherein the overlapping area between adjacent scene pictures satisfies a preset scene picture overlap threshold; using a pre-trained monocular depth estimation model to calculate the perspective depth maps corresponding to different perspectives of each scene picture, and converting the format of the perspective depth maps to a target format; initializing the initial sparse point cloud of the target reconstruction scene to generate three-dimensional Gaussian points, using a rasterizer, and combining the camera poses and the camera intrinsics to project the three-dimensional Gaussian points onto a two-dimensional plane to obtain a rendered depth map, and calculating the loss value between the rendered depth map and the perspective depth map, and performing scene constraints based on the loss value; performing exposure correction on each of the scene pictures by introducing two exposure coefficients into each of the scene pictures based on an exposure correction model; adopting a post-processing pruning scheme and utilizing an explicit representation of a three-dimensional Gaussian to clear floating artifacts in each of the scene pictures, and encouraging the model to relearn the areas for clearing floating artifacts to perform three-dimensional Gaussian scene modeling.

[0023] Applying the 3DGS scene modeling method combining depth estimation and exposure correction provided by the embodiments of the present application, depth constraints are introduced on the basis of traditional 3DGS, that is, monocular depth estimation models are used to obtain depth maps of each perspective, and losses are calculated between the depth maps and the rendered depth maps of the modeling scene at the corresponding camera poses to constrain scene optimization; at the same time, in view of the brightness deviation during shooting from different perspectives, the present invention introduces an exposure correction model, and two photometric coefficients are used to compensate for the overall change in image brightness; in addition, in order to further eliminate the "floating artifacts" in the scene, the present invention adopts a post-processing pruning scheme, using the explicit representation of three-dimensional Gaussian to remove the "floating artifacts" and encouraging the model to re-learn these areas correctly; from the modeling results, the method proposed by the present invention can perform three-dimensional modeling of the shooting scene with high fidelity, and can uniformly adjust the scene brightness and remove the unfriendly "floating artifacts" in the scene.

[0024] The above description is only an overview of the technical solutions of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented in accordance with the content of the description. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically described below. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0026] Figure 1 is a schematic flow chart of a 3DGS scene modeling method combining depth estimation and exposure correction provided by the embodiments of the present application;

[0027] Figure 2 is a schematic flow chart of another 3DGS scene modeling method combining depth estimation and exposure correction provided by the embodiments of the present application;

[0028] Figure 3 is a block diagram of a 3DGS scene modeling device combining depth estimation and exposure correction provided by the embodiments of the present application;

[0029] Figure 4 is a block diagram of a computing device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.

[0031] The present invention aims to reconstruct a high-quality three-dimensional scene through multi-view two-dimensional picture input. A recent part of the research work is based on the NeRF framework, which brings two main limitations: long training time and the black-box nature of the neural network, and this opacity limits the direct solution of the problem. Due to factors such as camera pose estimation error and uneven illumination in the traditional 3DGS algorithm, problems such as "background collapse" and "floating artifacts" will occur during the scene modeling process, which has a significant impact on the final scene modeling effect. The present invention introduces methods such as depth constraint, exposure correction, and post-processing pruning, aiming to remove the wrong floating objects in the scene and correct the overexposure phenomenon of some perspectives on the basis of reconstructing high-fidelity scene details, bringing a high-quality visual viewing effect.

[0032] See Figure 1 , Figure 1 is a schematic flowchart of a 3DGS scene modeling method combining depth estimation and exposure correction provided by an embodiment of the present application. Specifically, it includes the following steps.

[0033] Step S102: Obtain scene pictures of different perspectives corresponding to the target reconstruction, estimate the camera pose of the scene pictures based on an open-source algorithm, determine the initial sparse point cloud corresponding to the target reconstruction scene and the camera pose and camera intrinsics corresponding to each of the scene pictures, wherein the overlapping area between adjacent scene pictures satisfies a preset scene picture overlap threshold.

[0034] Step S104: Use a pre-trained monocular depth estimation model to calculate the perspective depth maps corresponding to different perspectives of each scene picture, and convert the format of the perspective depth maps to the target format.

[0035] Step S106: Initialize the initial sparse point cloud of the target reconstruction scene to generate three-dimensional Gaussian points, use a rasterizer, and combine the camera pose and the camera intrinsics to project the three-dimensional Gaussian points onto a two-dimensional plane to obtain a rendered depth map, and calculate the loss value between the rendered depth map and the perspective depth map, and perform scene constraint based on the loss value.

[0036] Step S108: Based on the exposure correction model, perform exposure correction on each of the scene pictures by introducing two exposure coefficients into each of the scene pictures.

[0037] Step S110: Adopt a post - processing pruning scheme, use the explicit representation of 3D Gaussian to remove floating artifacts in each of the scene pictures, and encourage the model to re - learn the areas where the floating artifacts are removed for 3D Gaussian scene modeling.

[0038] This application proposes a 3DGS scene modeling method that combines depth estimation and exposure correction. Given several multi - view pictures taken from different angles of a scene, the Pixsfm method is used to obtain the initial sparse point cloud of the scene, the camera poses and internal parameters corresponding to each view picture. A series of 3D Gaussian points are created based on the sparse point cloud, and the coefficients of the 3D Gaussian points are optimized using the same process as traditional 3DGS.

[0039] It is worth noting that to avoid the possible "background collapse" problem in scene modeling,

[0040] The present invention introduces depth constraints on the basis of traditional 3DGS, that is, monocular depth estimation models are used to obtain depth maps for each view, and losses are calculated between the rendered depth maps of the modeling scene at the corresponding camera poses to constrain scene optimization.

[0041] At the same time, for the brightness deviation when shooting from different views, the present invention introduces an exposure correction model, and two photometric coefficients are used to compensate for the overall change in image brightness.

[0042] In addition, to further eliminate the "floating artifacts" in the scene, the present invention adopts a post - processing pruning scheme, uses the explicit representation of 3D Gaussian to remove the "floating artifacts", and encourages the model to re - learn these areas correctly.

[0043] From the modeling results, the method proposed by the present invention can perform 3D modeling of the shooting scene with high fidelity, can uniformly adjust the scene brightness and remove the unfriendly "floating artifacts" in the scene.

[0044] Regarding camera pose estimation, for the scene to be reconstructed, several scene pictures with different views are taken around the scene, and it is required that the view deviation between adjacent pictures should not be too large, that is, there should be more than half of the overlapping area between adjacent view pictures. For the taken multi - view pictures, the open - source algorithm Pixsfm is used for camera pose estimation to obtain the initial sparse point cloud, camera poses and camera internal parameters.

[0045] Regarding monocular depth estimation, for the taken multi - view scene pictures, a pre - trained monocular depth estimation model is used to calculate the depth pictures corresponding to each view picture. In the embodiments of this application, the open - source Depth Anything V2 model is used for monocular depth estimation of each view picture. For the output results of the monocular depth estimation model, they are converted to the numpy format for storage to speed up reading and processing during training.

[0046] Regarding three-dimensional Gaussian modeling, for the sparse point cloud output by Pixsfm, it is initialized as a series of three-dimensional Gaussian points, that is, new parameters such as position, transparency, covariance, and spherical harmonic function coefficients are assigned on the original basis. For the three-dimensional Gaussian points, the fast rasterizer in the traditional 3DGS is used, and the three-dimensional Gaussian points are projected onto a two-dimensional plane in combination with the camera parameters to obtain a rendered image, and the loss is calculated between the rendered image and the real captured image of the corresponding view, and the position, transparency, covariance, spherical harmonic function coefficients, etc. are continuously optimized by backpropagation. In this process, the same strategy as the traditional 3DGS is adopted to adaptively control the number and density of Gaussians in the unit volume. Gaussian densification is performed once every 100 iterations during training, and the basically transparent Gaussians are removed.

[0047] Regarding depth constraints, the rasterizer in the traditional 3DGS cannot obtain a depth rendered image from the three-dimensional Gaussian point set, so this application designs a depth rendering scheme. The depth calculated in this application consists of two parts: mixed depth and modal depth.

[0048] The rendering method of the mixed depth is similar to the process of rendering a color image. Assume is the mixed depth at pixels x, y, and the calculation method is:

[0049] where T i is the cumulative transmittance of the i-th Gaussian, a i is the alpha blending weight, and d i is the depth of the Gaussian point.

[0050] The modal depth is the depth w i =T i a i of the Gaussian point with the largest contribution in the direction of this pixel, and the calculation formula is:

[0051] This application uses a combination of mixed depth and modal depth to represent the final rendered depth. The method for calculating the rendered depth of the rendered depth map includes: based on the first calculation formula, determining the rendered depth corresponding to the rendered depth map, where the first calculation formula includes:

[0052]

[0053] where d i is the depth of the Gaussian point, w i is the contribution in the corresponding direction of the pixel point, and β is the depth scale factor.

[0054] When β approaches 0, the rendering depth is the same as the blending depth; when β approaches infinity, the rendering depth is the same as the mode depth. Since the monocular depth estimation model obtains relative depth, while the rendering depth in the present invention is the anchored depth, the present application uses the Pearson correlation between image patches as the depth-related loss, and the calculation method is as follows:

[0055]

[0056] Wherein, and respectively represent the i-th image patch from the rendering depth map and the model estimated depth map.

[0057] Regarding exposure correction, in the actual scene, due to the change of lighting conditions, the camera will adjust the exposure time according to the shooting moment, resulting in the difference in the overall brightness of the image. The traditional 3DGS algorithm fails to effectively cope with these brightness changes, so problems such as "floating artifacts" are likely to occur in the rendering. To address this challenge, the present invention introduces two exposure coefficients for each image to accurately model the brightness change during shooting:

[0058] The exposure correction of each of the scene pictures by introducing two exposure coefficients into each of the scene pictures includes:

[0059] Based on the second calculation formula, perform exposure correction on each of the scene pictures, wherein the second calculation formula includes:

[0060]

[0061] Wherein, represents the rendered image, represents the image after exposure correction.

[0062] Therefore, the color loss of the present invention can be expressed as:

[0063]

[0064] Wherein, L1 represents the L1 loss constraint to ensure that the image after exposure adjustment is consistent with the real image in pixel values, and L SSIM represents the SSIM loss, which requires that the rendered image is similar to the real image in structure.

[0065] After compensating the exposure of the image using these coefficients, not only can the problem of inconsistent brightness be corrected, but also the performance of the algorithm under complex lighting conditions can be significantly improved, and the generation of artifacts can be reduced.

[0066] Regarding post - processing pruning, although the model introduces depth - related losses for training, optimizing only the depth is not sufficient to completely solve the problem of "floating artifacts". The present invention uses an explicit representation of 3D Gaussian to generate a mask for floating artifacts, remove the artifacts, and prompt the model to relearn the correct representation of these regions.

[0067] The post - processing pruning scheme is adopted, using an explicit representation of three - dimensional Gaussian to remove floating artifacts in each of the scene pictures, including:

[0068] Calculate the mixed depth map and the model depth map for each view corresponding to each camera pose, calculate the depth difference between the mixed depth map and the model depth map corresponding to each view, and determine the average deviation based on the depth difference;

[0069] Generate a threshold for floating - point detection for each view according to the preset curve parameters and the average deviation;

[0070] Based on the threshold and the depth difference, determine whether the target area is an area to be pruned.

[0071] The determining whether the target area is an area to be pruned based on the threshold and the depth difference includes:

[0072] In response to the threshold being less than the depth difference, determine that the target area is an area to be pruned and generate a mask for the floating - point area;

[0073] Use the mask to extract the Gaussian subset corresponding to the area to be pruned and remove it.

[0074] First, for each camera pose, calculate the mixed depth map d mix and the model depth map d mode for each view one by one, and then calculate the depth difference △ d = d mode - d mix . These depth differences are accumulated through a deviation test into a global deviation sum, and finally the average deviation

[0075] Next, the algorithm generates a threshold τ for floating - point detection for each view according to the preset curve parameters a, b and the average deviation d . Based on this threshold, compare the depth difference △ d for each view. If it is greater than the threshold τ d then mark it as an area to be pruned and generate a mask F d for the floating - point area. Use these masks to extract the Gaussian subset corresponding to the abnormal area and remove it.

[0076] SeeFigure 2 , Figure 2 It is a schematic flowchart of another 3DGS scene modeling method combining depth estimation and exposure correction provided by an embodiment of the present application.

[0077] Applying the 3DGS scene modeling method combining depth estimation and exposure correction provided by an embodiment of the present application, depth constraints are introduced on the basis of traditional 3DGS, that is, monocular depth estimation models are used to obtain depth maps of each view, and losses are calculated between the depth maps and the rendered depth maps of the modeling scene at the corresponding camera poses to constrain scene optimization; at the same time, in view of the brightness deviation when shooting from different views, an exposure correction model is introduced in the present invention, and two photometric coefficients are used to compensate for the overall change in image brightness; in addition, in order to further eliminate the "floating artifacts" in the scene, the present invention adopts a post-processing pruning scheme, using the explicit representation of three-dimensional Gaussian to remove the "floating artifacts" and encouraging the model to re-learn these areas correctly; from the modeling results, the method proposed in the present invention can perform three-dimensional modeling of the shooting scene with high fidelity, and can uniformly adjust the scene brightness and remove the unfriendly "floating artifacts" in the scene.

[0078] Corresponding to the above method embodiment, this specification also provides an embodiment of a 3DGS scene modeling device combining depth estimation and exposure correction. Figure 3 It is a block diagram of a 3DGS scene modeling device combining depth estimation and exposure correction provided by an embodiment of the present application. As Figure 3 shown, it specifically includes the following modules.

[0079] A determination module 302, configured to obtain scene pictures of different views corresponding to a target reconstruction, estimate the camera poses of the scene pictures based on an open-source algorithm, and determine the initial sparse point cloud corresponding to the target reconstruction scene, the camera poses and camera internal parameters corresponding to each of the scene pictures, wherein the overlapping area between adjacent scene pictures meets a preset scene picture overlap threshold;

[0080] A calculation module 304, configured to use a pre-trained monocular depth estimation model to calculate the view depth maps corresponding to different views of each scene picture, and convert the format of the view depth maps into a target format;

[0081] A constraint module 306, configured to initialize the initial sparse point cloud of the target reconstruction scene to generate three-dimensional Gaussian points, use a rasterizer, and combine the camera poses and the camera internal parameters to project the three-dimensional Gaussian points onto a two-dimensional plane to obtain a rendered depth map, and calculate the loss value between the rendered depth map and the view depth map, and perform scene constraint based on the loss value;

[0082] The correction module 308 is configured to perform exposure correction on each of the scene pictures by introducing two exposure coefficients into each of the scene pictures based on an exposure correction model;

[0083] The cleaning module 310 is configured to adopt a post-processing pruning scheme and utilize an explicit representation of a three-dimensional Gaussian to remove floating artifacts in each of the scene pictures and encourage the model to relearn the areas for removing floating artifacts for three-dimensional Gaussian scene modeling.

[0084] In an alternative embodiment, the constraint module 306 is further configured to:

[0085] Determine the rendering depth corresponding to the rendered depth map based on a first calculation formula, where the first calculation formula includes:

[0086]

[0087] where d i is the Gaussian point depth, w i is the contribution degree in the corresponding direction of the pixel point, and β is the depth scale factor.

[0088] In an alternative embodiment, the correction module 308 is further configured to:

[0089] Perform exposure correction on each of the scene pictures based on a second calculation formula, where the second calculation formula includes:

[0090]

[0091] where represents the rendered image, represents the image after exposure correction.

[0092] In an alternative embodiment, the cleaning module 310 is further configured to:

[0093] Calculate the mixed depth map and the mode depth map for each view corresponding to each camera pose, calculate the depth difference between the mixed depth map and the mode depth map corresponding to each view, and determine the average deviation based on the depth difference;

[0094] Generate a threshold for floating-point detection for each view according to preset curve parameters and the average deviation;

[0095] Determine whether the target area is an area to be pruned based on the threshold and the depth difference.

[0096] In an alternative embodiment, the cleaning module 310 is further configured to:

[0097] In response to the threshold being less than the depth difference, determine that the target area is the area to be trimmed, and generate a mask for the floating-point area;

[0098] Using the mask, extract the Gaussian subset corresponding to the area to be trimmed and clear it.

[0099] Applying the 3DGS scene modeling method combining depth estimation and exposure correction provided by the embodiments of the present application, depth constraints are introduced on the basis of traditional 3DGS, that is, depth maps of each view are obtained through a monocular depth estimation model, and the loss between the rendered depth map of the modeling scene in the corresponding camera pose is used to constrain the scene optimization; at the same time, in view of the brightness deviation when shooting from different perspectives, the present invention introduces an exposure correction model and uses two photometric coefficients to compensate for the overall change in image brightness; in addition, in order to further eliminate the "floating artifacts" in the scene, the present invention adopts a post-processing pruning scheme, uses the explicit representation of the three-dimensional Gaussian to remove the "floating artifacts", and encourages the model to re-learn these areas correctly; from the perspective of the modeling results, the method proposed by the present invention can perform three-dimensional modeling of the shooting scene with high fidelity, and can uniformly adjust the scene brightness and delete the unfriendly "floating artifacts" in the scene.

[0100] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the 3DGS scene modeling device combining depth estimation and exposure correction, since it is basically similar to the embodiment of the 3DGS scene modeling method combining depth estimation and exposure correction, the description is relatively simple, and the relevant parts can be referred to the partial description of the embodiment of the 3DGS scene modeling method combining depth estimation and exposure correction.

[0101] Figure 4 It is a structural block diagram of a computing device provided by the embodiments of the present application. The components of the computing device 400 include, but are not limited to, a memory 410 and a processor 420. The processor 420 is connected to the memory 410 through a bus 430, and a database 450 is used to store data.

[0102] The computing device 400 also includes an access device 440, which enables the computing device 400 to communicate via one or more networks 460. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 440 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.

[0103] In one embodiment of the present specification, the above components of the computing device 400, as well as Figure 4 other components not shown, may also be connected to each other, for example, via a bus. It should be understood that Figure 4 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present specification. Those skilled in the art may add or replace other components as needed.

[0104] The computing device 400 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 400 may also be a mobile or stationary server.

[0105] Wherein, the processor 420 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the 3DGS scene modeling method for depth estimation and exposure correction described above.

[0106] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiment of the computing device, since it is basically similar to the embodiment of the 3DGS scene modeling method combining depth estimation and exposure correction, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the embodiment of the 3DGS scene modeling method combining depth estimation and exposure correction.

[0107] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned 3DGS scene modeling method combining depth estimation and exposure correction.

[0108] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiment of the computer-readable storage medium, since it is basically similar to the embodiment of the 3DGS scene modeling method combining depth estimation and exposure correction, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the embodiment of the 3DGS scene modeling method combining depth estimation and exposure correction.

[0109] An embodiment of this specification also provides a computer program, which, when executed on a computer, causes the computer to execute the steps of the above-mentioned 3DGS scene modeling method combining depth estimation and exposure correction.

[0110] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiment of the computer program, since it is basically similar to the embodiment of the 3DGS scene modeling method combining depth estimation and exposure correction, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the embodiment of the 3DGS scene modeling method combining depth estimation and exposure correction.

[0111] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0112] The computer instructions include computer program code, which may be in the form of source code, object code, executable files, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0113] It should be noted that the specific embodiments of this specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be executed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.

[0114] In the above embodiments, each embodiment is described with its own emphasis. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0115] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not elaborate on all the details and do not limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.

[0116] The preferred embodiments of the present specification disclosed above are only used to help explain the present specification. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of the present specification. The present specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of the present specification, so that those skilled in the art can well understand and utilize the present specification. The present specification is only limited by the claims and their full scope and equivalents.

Claims

1. A 3DGS scene modeling method combining depth estimation and exposure correction, characterized in that: include: Obtain scene pictures of different perspectives corresponding to the target reconstruction, estimate the camera pose of the scene pictures based on an open source algorithm, determine the initial sparse point cloud corresponding to the target reconstruction scene and the camera pose and camera intrinsic parameters corresponding to each of the scene pictures, wherein the overlapping area between adjacent scene pictures satisfies a preset scene picture overlapping threshold; Using a pre-trained monocular depth estimation model, calculate the perspective depth map corresponding to different perspectives corresponding to each scene image, and convert the format of the perspective depth map to the target format; Initializing the initial sparse point cloud of the target reconstructed scene to generate three-dimensional Gaussian points, using a rasterizer, combining the camera pose and the camera intrinsic parameters, projecting the three-dimensional Gaussian points onto a two-dimensional plane to obtain a rendered depth map, and calculating a loss value between the rendered depth map and the viewing depth map, and performing scene constraints based on the loss value; Based on the exposure correction model, exposure correction is performed on each of the scene images by introducing two exposure coefficients into each of the scene images; A post-processing pruning scheme is adopted to remove floating artifacts in each scene picture using an explicit representation of a three-dimensional Gaussian, and the model is encouraged to relearn the area where the floating artifacts are removed to perform three-dimensional Gaussian scene modeling.

2. The method according to claim 1, characterized in that The rendering depth calculation method of the rendering depth map comprises: Based on a first calculation formula, a rendering depth corresponding to the rendering depth map is determined, wherein the first calculation formula includes: Among them, d i is the Gaussian point depth, w i is the contribution of the pixel in the corresponding direction, and β is the depth scale factor.

3. The method according to claim 1, characterized in that The exposure correction of each scene picture by introducing two exposure coefficients into each scene picture includes: Based on a second calculation formula, exposure correction is performed on each of the scene pictures, wherein the second calculation formula includes: in, Represents a rendered image, Represents the image after exposure correction.

4. The method according to claim 1, characterized in that: The post-processing pruning scheme is adopted to remove floating artifacts in each of the scene pictures by using an explicit representation of a three-dimensional Gaussian, including: Calculate the mixed depth map and the pattern depth map of each viewing angle corresponding to each camera pose, calculate the depth difference between the mixed depth map and the pattern depth map corresponding to each viewing angle, and determine the average deviation based on the depth difference; According to the preset curve parameters and the average deviation, a floating point detection threshold is generated for each viewing angle; Based on the threshold and the depth difference, it is determined whether the target area is an area to be cropped.

5. The method according to claim 4, characterized in that The determining whether the target area is an area to be trimmed based on the threshold and the depth difference includes: In response to the threshold being less than the depth difference, determining the target area as a to-be-pruned area, and generating a mask of the floating-point area; The mask is used to extract the Gaussian subset corresponding to the area to be pruned and clear it.

6. A 3DGS scene modeling device combining depth estimation and exposure correction, characterized in that: include: A determination module is configured to obtain scene pictures of different perspectives corresponding to the target reconstruction, estimate the camera pose of the scene pictures based on an open source algorithm, and determine an initial sparse point cloud corresponding to the target reconstruction scene and a camera pose and camera intrinsic parameters corresponding to each of the scene pictures, wherein an overlapping area between adjacent scene pictures satisfies a preset scene picture overlap threshold; A calculation module is configured to use a pre-trained monocular depth estimation model to calculate a perspective depth map corresponding to different perspectives corresponding to each scene image, and convert the format of the perspective depth map into a target format; A constraint module is configured to initialize the initial sparse point cloud of the target reconstructed scene, generate three-dimensional Gaussian points, use a rasterizer, combine the camera pose and the camera intrinsic parameters, project the three-dimensional Gaussian points onto a two-dimensional plane, obtain a rendered depth map, calculate a loss value between the rendered depth map and the viewing depth map, and perform scene constraints based on the loss value; A correction module is configured to perform exposure correction on each of the scene images by introducing two exposure coefficients into each of the scene images based on an exposure correction model; The cleaning module is configured to adopt a post-processing pruning scheme, use the explicit representation of the three-dimensional Gaussian to clean the floating artifacts in each of the scene pictures, and encourage the model to relearn the area of ​​clearing the floating artifacts to perform three-dimensional Gaussian scene modeling.

7. A computer device, characterized in that: The computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the method according to any one of claims 1 to 5 when executing the program.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores an implementation program for information transmission, and when the program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.