Illumination robustness 3DGS method for new view angle image synthesis

By capturing multiple images under extreme lighting conditions and introducing bilateral grid processing, the problem of inconsistent viewing angles during the generation of new viewing angle images in the prior art is solved, and high-quality and consistent new viewing angle image generation is achieved.

CN120047625AActive Publication Date: 2025-05-27QUFU NORMAL UNIV +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510233210.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-27
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

When the prior art generates new viewing angle images under extreme lighting conditions, the output quality is highly dependent on the input image quality, resulting in inconsistent viewing angles, resulting in artifacts, and significantly reducing the quality of the synthetic image.

Method used

Using the illumination robust 3DGS method, by capturing multiple images under extreme lighting conditions, using low-light enhancement or overexposure correction, an initial point cloud is generated, and a bilateral grid is introduced for pixel-level processing during the training phase, the loss function is optimized to solve color consistency and contrast loss, generating high-quality new perspective images.

Benefits of technology

Generate high-quality new viewing images under extreme or complex lighting conditions, eliminate the problem of view inconsistency and improve the quality and consistency of output images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047625A_ABST
    Figure CN120047625A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer image processing, in particular to an illumination robustness 3DGS method for new-view-angle image synthesis, which can generate a high-quality new-view-angle image under an extreme or complex illumination condition, and can generate a new-view-angle image in consideration of view inconsistency possibly caused by direct independent enhancement of an image in a two-dimensional image space. According to the method, a bilateral grid is introduced in a training stage to simulate an image processing flow, and the bilateral grid is removed in a reasoning stage, so that view inconsistency is effectively eliminated. In addition, # imgabs0 # and # imgabs1 # loss are introduced to generate a new visual angle image with higher quality under a normal illumination condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer image processing, and in particular to an illumination robust 3DGS method for new-viewing-angle image synthesis. Background Art

[0002] Extracting accurate 3D scene representations from a series of 2D images and generating new perspectives is a core challenge in the field of computer vision. This technology has a wide range of applications in many fields, including autonomous driving, medical imaging, robotics, and virtual reality. The emergence of NeRF (Neural Radiance Fields) has significantly accelerated the development of 3D reconstruction and new perspective synthesis. It can effectively interpret 3D scenes from 2D images and generate realistic new perspectives, injecting new vitality into these research fields. NeRF uses multi-layer perceptrons to implicitly encode volume density and color, thereby achieving the generation of high-quality new perspectives. This breakthrough technology has made significant progress in many fields and promoted the development of the intersection of computer vision and computer graphics.

[0003] Recently, 3DGS has shown comparable expressiveness to NeRF, while being hundreds of times faster in training and inference speed, showing that it is expected to be a better choice than NeRF; unlike NeRF's implicit surface representation, 3DGS uses an explicit 3D representation to efficiently construct a 3D space through a simplified and logically clear framework; 3DGS can not only achieve high-quality reconstruction and rendering, but also provide super-real-time rendering and inference speed, effectively solving the major problem of NeRF in time efficiency. Although 3DGS has made significant progress in new perspective synthesis, its output quality is still highly dependent on the quality of the input image. Under extreme lighting conditions such as low light or overexposure, the reduction of visible information and the loss of details in the input image pose a huge challenge to accurately recovering the geometric structure. A simple solution is to first perform low-light enhancement or overexposure correction in the 2D image space to enrich the image details. However, this approach introduces a key problem: existing 2D processing methods usually enhance each view independently, resulting in inconsistencies between views. Therefore, when observing the same position from different viewpoints, contradictory results may occur, resulting in artifacts in the reconstructed image, significantly reducing the quality of the synthesized new perspective image.

[0004] In this work, we propose a new approach to build an effective bridge between 2D and 3D representations by integrating 2D image processing techniques to enhance visibility while maintaining multi-view consistency. Summary of the invention

[0005] In order to make up for the deficiencies in the prior art and achieve the above functions, the present invention provides an illumination robust 3DGS method for new-viewing angle image synthesis.

[0006] The present invention is achieved through the following technical solutions: An illumination robust 3DGS method for new view image synthesis includes the following steps: Capture N images under extreme lighting conditions, and use low-light enhancement or overexposure correction to obtain a processed image set according to scene requirements. ,in represents a collection of individually enhanced multi-view images; Subsequently, these images were processed by the software COLMAP to generate an initial point cloud as the starting result; Then, initialize N bilateral grids , for processing at the pixel level Each training image in .

[0007] For a specific pixel (u, v) in the Lth training image, it uses the slice-based differentiable rasterization rendering algorithm in 3DGS to obtain the evaluated color after rendering , the slice operation expression is as follows: , , Among them, for the guiding function g(C), a logarithmic brightness formula is used to provide compression of the brightness dynamic range, aiming to handle scenes with high lighting intensity, thereby achieving a more natural visual effect; the c value is set to 6, and the guiding function g(C) is: , Apply an affine transformation After slicing, get the processed color .

[0008] Furthermore, in order to better implement the present invention, the bilateral grid is compatible with a neural network, and the bilateral grid maps the image data into a four-dimensional tensor , where W, H, and M represent the resolution of the grid, 12 corresponds to the length of the 3×4 affine color transformation matrix, and the color value of each pixel is , is used to locate the corresponding color change position in the bilateral grid by calculating the weighted sum based on color value and spatial position; Weighted weight By using the hat function The derived linear interpolation kernel is calculated to ensure that only neighboring pixels of similar color will affect the current pixel; Finally, the processed color The affine transformation matrix is ​​extracted by Generated by multiplying the pixel's color value.

[0009] Furthermore, in order to better implement the present invention, a total variation regularization term is introduced to control the smoothness of the bilateral grid, and the final objective function is expressed as: , , represents the finite difference operator, Set this parameter to 10 for best practical results.

[0010] Furthermore, in order to better implement the present invention, a color consistency loss function is used: , Restores the color balance and overall appearance saturation in an image when processed with and subsequently removed from an optimized bilateral grid.

[0011] Furthermore, in order to better implement the present invention, the predicted normal illumination image is obtained by using the contrast loss function To enhance the balance between local details and overall contrast, where: , C represents the number of channels, Represents the variance of channel C, which is set to 0.3 for overexposed scenes and 0.4 for low-light scenes.

[0012] Furthermore, in order to better implement the present invention, the color consistency loss and contrast loss are solved by optimizing the loss function, which consists of three parts: rendering loss and two unsupervised brightness correction losses and ; The overall training loss is expressed as: , .

[0013] The beneficial effects of the present invention are: The present invention can generate high-quality new-view images under extreme or complex lighting conditions. Considering that a single enhanced image directly in a two-dimensional image space may cause view inconsistency, the present invention introduces a bilateral grid in the training phase to simulate the image processing flow, and removes it in the inference phase to effectively eliminate view inconsistency. In addition, the present invention introduces and To generate new perspective images under normal lighting conditions with higher quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a principle framework diagram of the present invention. DETAILED DESCRIPTION

[0015] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Generally, the components of the embodiments of the present invention described and shown in the accompanying drawings here can be arranged and designed in various different configurations.

[0016] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.

[0017] (1) About 3D Gaussian sputtering: 3DGS is a prior art, and the following description is all background knowledge to help understand this technology.

[0018] When the density of the point cloud is high enough, the geometry and appearance of a three-dimensional object can be described from any perspective. However, when relatively sparse data is used to describe complex or smooth areas, the points in the point cloud need to be assigned a certain volume. If each point is regarded as a shape that follows a specific distribution, projecting these shapes onto a two-dimensional plane can generate high-quality images. 3DGS provides a solution, where each point is used as the center of a shape that conforms to a three-dimensional Gaussian distribution to describe a three-dimensional object. Real-time rendering is achieved through differentiable rasterization. Each 3DGS primitive {G} is defined by the following parameters: Gaussian center μ, opacity α, scaling described by the covariance matrix Σ, and color c modeled using spherical harmonics. The basis function of the Gaussian basis can be expressed by the following formula: , To represent a multivariate Gaussian distribution, only the spatial coordinates and the covariance matrix are needed. The covariance matrix can be expressed as , where R represents the rotation matrix and S is the scaling matrix. For a three-dimensional space point, only a 3D scaling vector s and a quaternion q need to be stored. In the rendering stage, a block-based differentiable rasterization algorithm is used. Before rendering, the 3D Gaussian points are pre-sorted according to their spatial depth. The final color is calculated by alpha blending, and the formula is as follows: , in, It is the color learned. Represents the opacity weighted by the probability density of the two-dimensional Gaussian distribution of the i-th projection to the target pixel position The loss function is a weighted combination of the L1 loss of color reconstruction and the D-SSIM loss of the reconstructed image, and is in the following form: .

[0019] (2) Color prior: Referring to the description in (1), 3DGS constructs a shape that conforms to a three-dimensional Gaussian distribution centered at each point, and requires processing images from multiple viewpoints to generate an initial point cloud. However, accurately estimating camera parameters and generating a complete initial point cloud for images taken under extreme lighting conditions is challenging. Some NeRF-based studies, such as AlethNeRF and LLNeRF, have improved neural radiance fields under poor lighting conditions and achieved good results. However, these methods assume that accurate camera poses can be obtained and usually require paired images to work effectively, which is often difficult to achieve in real scenes. Especially in low-light conditions, the extreme lack of visibility can lead to highly sparse point clouds. Using existing 2D image processing methods to perform preliminary image enhancement and then using COLMAP to generate a point cloud significantly improves the point cloud and its structure. Although there are deviations compared to the camera poses and number of point clouds obtained under normal lighting, the impact on the reconstruction results is small.

[0020] (3) Reference Figure 1 , the present invention comprises the following steps: Capture N images under extreme lighting conditions, and use low-light enhancement or overexposure correction to obtain a processed image set according to scene requirements. ,in represents a collection of individually enhanced multi-view images; Subsequently, these images were processed by the software COLMAP to generate an initial point cloud as the starting result; Then, initialize N bilateral grids , for processing at the pixel level Each training image in .

[0021] For a specific pixel (u, v) in the Lth training image, it uses the slice-based differentiable rasterization rendering algorithm in 3DGS to obtain the evaluated color after rendering , the slice operation expression is as follows: , , Among them, for the guiding function g(C), a logarithmic brightness formula is used to provide compression of the brightness dynamic range, aiming to handle scenes with high lighting intensity, thereby achieving a more natural visual effect; the c value is set to 6, and the guiding function g(C) is: , Apply an affine transformation After slicing, get the processed color .

[0022] The bilateral grid is compatible with neural networks and maps image data into a four-dimensional tensor. , where W, H, and M represent the resolution of the grid, 12 corresponds to the length of the 3×4 affine color transformation matrix, and the color value of each pixel is used to locate the corresponding color change position in the bilateral grid by calculating a weighted sum based on color value and spatial position; Weighted weight By using the hat function The derived linear interpolation kernel is calculated to ensure that only neighboring pixels of similar color will affect the current pixel; Finally, the processed color The affine transformation matrix is ​​extracted by Generated by multiplying the pixel's color value.

[0023] The total variation regularization term is introduced to control the smoothness of the bilateral grid, and the final objective function is expressed as: , , to train, represents the finite difference operator, Set to parameter 10 to obtain the best practical application results; Using color consistency loss function: , to optimize the color balance and overall appearance saturation in the image when processed and subsequently removed from the bilateral grid.

[0024] Predicted normal lighting image, through contrast loss function To enhance the balance between local details and overall contrast, where: , C represents the number of channels, Represents the variance of channel C, which is set to 0.3 for overexposed scenes and 0.4 for low-light scenes.

[0025] Color consistency loss and contrast loss are solved by optimizing the loss function, which consists of three parts: rendering loss and two unsupervised brightness correction losses and ; The overall training loss is expressed as: , .

[0026] The above is a simulation enhancement process of introducing bilateral grids during training. During testing, the bilateral grids are removed to eliminate inconsistent enhancements.

[0027] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Other modifications or equivalent substitutions made to the technical solution of the present invention by ordinary technicians in the field should be included in the scope of the claims of the present invention as long as they do not depart from the spirit and scope of the technical solution of the present invention.

Claims

1. An illumination robust 3DGS method for novel viewpoint image synthesis, characterized by: The following steps are involved: Capture N images under extreme lighting conditions, and use low-light enhancement or overexposure correction to obtain a processed image set according to scene requirements. ,in represents a collection of individually enhanced multi-view images; Subsequently, these images were processed by the software COLMAP to generate an initial point cloud as the starting result; Then, initialize N bilateral grids , for processing at the pixel level Each training image in ; For a specific pixel (u, v) in the Lth training image, it uses the slice-based differentiable rasterization rendering algorithm in 3DGS to obtain the evaluated color after rendering , the slice operation expression is as follows: , , Among them, for the guiding function g(C), a logarithmic brightness formula is used to provide compression of the brightness dynamic range, aiming to handle scenes with high lighting intensity, thereby achieving a more natural visual effect; the c value is set to 6, and the guiding function g(C) is: , Apply an affine transformation After slicing, get the processed color .

2. The illumination robust 3DGS method for new view image synthesis according to claim 1, characterized in that: The bilateral grid is compatible with neural networks and maps image data into a four-dimensional tensor. , where W, H, and M represent the resolution of the grid, 12 corresponds to the length of the 3×4 affine color transformation matrix, and the color value of each pixel , is used to locate the corresponding color change position in the bilateral grid by calculating the weighted sum based on color value and spatial position; Weighted weight By using the hat function The derived linear interpolation kernel is calculated to ensure that only neighboring pixels of similar color will affect the current pixel; Finally, the processed color The affine transformation matrix is ​​extracted by Generated by multiplying the pixel's color value.

3. The illumination robust 3DGS method for new view image synthesis according to claim 2, characterized in that: The total variation regularization term is introduced to control the smoothness of the bilateral grid, and the final objective function is expressed as: , , represents the finite difference operator, Set this parameter to 10 for best practical results.

4. The illumination robust 3DGS method for new view image synthesis according to claim 3, characterized in that: Using color consistency loss function: , Restores the color balance and overall appearance saturation in an image when processed with and subsequently removed from an optimized bilateral grid.

5. The illumination robust 3DGS method for new view image synthesis according to claim 4, characterized in that: Predicted normal lighting image, through contrast loss function To enhance the balance between local details and overall contrast, where: , C represents the number of channels, Represents the variance of channel C, which is set to 0.3 for overexposed scenes and 0.4 for low-light scenes.

6. The illumination robust 3DGS method for new view image synthesis according to claim 5, characterized in that: Color consistency loss and contrast loss are solved by optimizing the loss function, which consists of three parts: rendering loss and two unsupervised brightness correction losses and ; The overall training loss is expressed as: , 。

Citation Information

Patent Citations

  • 3D Gaussian sputtering scene reconstruction method based on view dependence difference decoupling

    CN118710792A

  • Nerve radiation field rendering method and device based on nerve base and tensor decomposition

    CN119180898A

  • Ray Image Modeling for Fast Catadioptric Light Field Rendering

    US20130038696A1

  • Generative model for 3D face synthesis with HDRI relighting

    US20240020915A1