3DGS outdoor scene rendering optimization method based on frequency domain separation
Through frequency domain separation and improved appearance learning network and feature extraction network, the problem of insufficient rendering details and appearance consistency in 3DGS in outdoor scenes is solved, improving the realistic rendering results and the accuracy of reconstruction model.
Patent Information
- Application Number
- CN202510382938.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-18
AI Technical Summary
3DGS under complex lighting conditions of outdoor scenes, insufficient rendering details and appearance consistency, resulting in reduced rendering results and reduced availability of reconstruction models.
The frequency domain separation method is adopted to convert the rendered image and ground reality diagram to the frequency domain space, and low-frequency appearance modeling and detailed enhancement modeling are performed respectively. Through improved appearance learning networks and feature extraction networks, the appearance and details of the rendered image are optimized, and iteratively optimized by combining the loss function of low-frequency appearance modeling and detail enhancement modeling.
It improves the detailed characterization and appearance consistency of outdoor scene renderings, enhances the realistic rendering results and the accuracy of the reconstruction model.
Smart Images

Figure CN120339480A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of three-dimensional reconstruction, and specifically relates to a method for optimizing 3DGS outdoor scene rendering based on frequency domain separation. Background Art
[0002] With the rapid development of computer vision and three-dimensional reconstruction technologies, image-based three-dimensional reconstruction methods are increasingly applied in various fields, such as virtual reality, augmented reality, geographic information system (GIS), and industrial inspection. However, traditional image-based three-dimensional reconstruction methods usually face the problem of difficult trade-off between reconstruction accuracy and speed. Especially when dealing with large-scale scenes and complex data, how to maintain high accuracy and high efficiency simultaneously becomes a challenge.
[0003] 3D Gaussian Splatting (3DGS), as a latest image-based three-dimensional reconstruction technology, represents a complex scene as a set of 3D Gaussian basis elements, and uses the EWA splatting method to project these basis elements onto a 2D imaging plane. On the 2D imaging plane, the pixel color is calculated using the alpha blending technique, thereby achieving efficient rendering. Finally, it is optimized through the pixel-level loss between the rendered image and the ground truth image. Compared with traditional multi-view stereo reconstruction methods, 3DGS has higher rendering efficiency and good geometric representation ability. However, 3DGS is only optimized through pure photometric loss. When facing outdoor scenes with complex lighting, due to shadows and reflections caused by uneven lighting under different viewpoints, 3DGS will have problems such as inconsistent appearance between the rendered image and the ground truth image, and there will be phenomena such as missing, blurring, and artifacts in the performance of complex details. These problems not only affect the realism of the rendering result, but also reduce the usability and accuracy of the final reconstructed model. Summary of the Invention
[0004] The problem to be solved by the present invention is the insufficient rendering detail representation and appearance consistency of 3DGS under complex lighting conditions in outdoor scenes, and a method for optimizing 3DGS outdoor scene rendering based on frequency domain separation is provided.
[0005] To solve the above problems, the present invention is implemented through the following technical solutions:
[0006] A method for optimizing 3DGS outdoor scene rendering based on frequency domain separation, including the following steps:
[0007] Step 1, through Fourier transform, convert the rendered image I r generated by 3DGS for a given viewpoint in each iteration and the ground truth image I g into the frequency domain space respectively, to obtain the rendered frequency domain image f r and the ground truth frequency domain image f g ;
[0008] Step 2: Move the DC components of the spectra of the rendered frequency-domain map f r and the ground truth frequency-domain map f g to the center of their spectra to obtain the rendered center frequency-domain map and the ground truth center frequency-domain map
[0009] Step 3: Based on the rendered center frequency-domain map and the ground truth center frequency-domain map construct the rendered mask M r and the ground truth mask M g respectively through the given radius r;
[0010] Step 4: Multiply the rendered mask M r with the rendered center frequency-domain map to obtain the rendered low-frequency frequency-domain map Then perform the inverse Fourier transform on the rendered low-frequency frequency-domain map to obtain the rendered low-frequency map
[0011] Step 5: Invert the rendered mask M r and multiply it with the rendered center frequency-domain map to obtain the rendered high-frequency frequency-domain map Then perform the inverse Fourier transform on the rendered high-frequency frequency-domain map to obtain the rendered high-frequency map Invert the ground truth mask M q and multiply it with the ground truth center frequency-domain map to obtain the ground truth high-frequency frequency-domain map Then perform the inverse Fourier transform on the ground truth high-frequency frequency-domain map to obtain the ground truth high-frequency map
[0012] Step 6: Initialize the appearance embedding that satisfies the normal distribution and concatenate the rendered low-frequency map and the appearance embedding to obtain the concatenated rendered map D;
[0013] Step 7: Feed the concatenated rendered map D into the improved appearance learning network for learning to obtain the adjusted map
[0014] Step 8: Perform per-pixel multiplication pixel transformation on the rendered map I through the adjusted map r to obtain the appearance-transformed rendered map
[0015] Step 9: The rendered high-frequency map and the ground truth high-frequency map are respectively fed into the feature extraction network for learning to obtain the rendered high-frequency feature map and the ground truth high-frequency feature map
[0016] Step 10: Based on the rendered image I r , the ground truth image I g , the appearance transformation rendered image the rendered high-frequency feature map and the ground truth high-frequency feature map calculate the total loss and optimize each iteration of the 3DGS by minimizing the total loss ;
[0017]
[0018] In the formula, represents the L1 loss, represents the structural similarity loss, represents the decoupled structural similarity loss, I r represents the rendered image, I g represents the ground truth image, represents the appearance transformation rendered image, represents the rendered high-frequency feature map, represents the ground truth high-frequency feature map, λ represents the photometric weight, α represents the low-frequency appearance weight, β represents the detail enhancement weight, and α + β = 1.
[0019] In the above step 6, the length of the appearance embedding is 96.
[0020] In the above step 7, the improved appearance learning network includes 2 convolutional layers, 1 LeakyReLU activation function layer, 1 Sigmoid activation function layer, and 5 appearance learning modules. Each appearance learning module consists of 1 convolutional layer, 1 LeakyReLU activation function layer, and 1 CBAM attention mechanism layer. The input of the convolutional layer forms the input of the appearance learning module. The output of the convolutional layer is connected to the input of the LeakyReLU activation function layer. The output of the LeakyReLU activation function layer is connected to the input of the CBAM attention mechanism layer. The output of the CBAM attention mechanism layer forms the output of the appearance learning module. The input of the first convolutional layer serves as the input of the improved appearance learning network. The output of the first convolutional layer is connected to the input of the LeakyReLU activation function layer. The output of the LeakyReLU activation function layer is connected to the input of the first appearance learning module. The output of the first appearance learning module is connected to the input of the second appearance learning module. The output of the second appearance learning module is connected to the input of the third appearance learning module. The output of the third appearance learning module is connected to the input of the fourth appearance learning module. The output of the fourth appearance learning module is connected to the input of the fifth appearance learning module. The output of the fifth appearance learning module is connected to the input of the second convolutional layer. The output of the second convolutional layer is connected to the input of the Sigmoid activation function layer. The output of the Sigmoid activation function layer serves as the output of the improved appearance learning network.
[0021] In the above step 9, the feature extraction network includes 5 feature extraction modules, 1 concatenation layer, 1 convolutional layer, 1 batch normalization layer, and 1 ReLU activation function layer. Each feature extraction module consists of 1 convolutional layer, 1 batch normalization layer, 1 ReLU activation function layer, and 1 ECA attention mechanism layer. The input of the convolutional layer forms the input of the feature extraction module. The output of the convolutional layer is connected to the input of the batch normalization layer. The output of the batch normalization layer is connected to the input of the ReLU activation function layer. The output of the ReLU activation function layer is connected to the input of the ECA attention mechanism layer. The output of the ECA attention mechanism layer forms the output of the feature extraction module. The input of the first feature extraction module serves as the input of the feature extraction network. The output of the first feature extraction module is connected to the input of the second feature extraction module. The output of the second feature extraction module is connected to the input of the third feature extraction module. The outputs of the first and third feature extraction modules are connected to the input of the concatenation layer. The output of the concatenation layer is connected to the input of the fourth feature extraction module. The output of the fourth feature extraction module is connected to the input of the fifth feature extraction module. The output of the fifth feature extraction module is connected to the input of the convolutional layer. The output of the convolutional layer is connected to the input of the batch normalization layer. The output of the batch normalization layer is connected to the input of the ReLU activation function layer. The output of the ReLU activation function layer serves as the output of the feature extraction network.
[0022] In the above step 10,
[0023] L1 loss is:
[0024]
[0025] Structural similarity loss is:
[0026]
[0027] Decoupled structural similarity loss is:
[0028]
[0029] In the formula, represents the pixel value of the i-th pixel of the rendered image, and I g (i) represents the pixel value of the i-th pixel of the ground truth image, and N represents the number of pixels in the image; represents the mean value of the rendered image, represents the mean value of the ground truth image, represents the covariance between the rendered image and the ground truth image, represents the variance of the rendered image, represents the variance of the ground truth image; represents the covariance between the rendered high-frequency feature map and the ground truth high-frequency feature map, represents the standard deviation of the rendered high-frequency feature map, represents the standard deviation of the ground truth high-frequency feature map; C1, C2, and C3 are all given non-zero constants.
[0030] Compared with the prior art, in view of the problems of insufficient detail representation and appearance adaptability in the outdoor scene rendering of the current 3DGS reconstruction method and the lack of rendered detail representation, the present invention proposes an optimization method combining low-frequency appearance modeling and detail enhancement modeling, which has the following characteristics:
[0031] 1. Before low-frequency appearance modeling and detail enhancement modeling, the image will be separated in the frequency domain. Before low-frequency appearance modeling, the low-frequency image of the rendered image is separated, and before detail enhancement modeling, the high-frequency images of the rendered image and the ground truth image are separated. Since the low-frequency image retains the brightness change and global information, and the high-frequency image retains the details such as edges and textures, low-frequency appearance modeling and detail enhancement modeling can perform dedicated optimizations for appearance changes and detail problems respectively;
[0032] 2. In the optimization of low-frequency appearance modeling, an improved appearance learning network is proposed. It does not perform downsampling but uses the original image size as the input, thus retaining more appearance details for appearance embedding to learn. At the same time, by increasing the length of the appearance embedding and introducing LReLU and CBAM attention, on the one hand, the appearance expression ability of the appearance embedding is improved, and on the other hand, the learning ability of the network is enhanced. For the characteristics of low-frequency images, the existing photometric loss is adopted in low-frequency appearance modeling, and the photometric loss is used to optimize the iterative process of 3DGS, making the loss focus on optimizing photometric information.
[0033] 3. In the optimization of detail enhancement modeling, a CNN-based feature extraction network is proposed. By introducing the ECA attention module, the focusing degree on important feature channels is improved. For the characteristics of high-frequency images, the structural similarity loss is modified in detail enhancement modeling, and the decoupled structural similarity loss is used to optimize the iterative process of 3DGS, making the loss focus on optimizing structural information. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a schematic diagram of a 3DGS outdoor scene rendering optimization method based on frequency domain separation.
[0035] Figure 2 It is a schematic diagram of the optimization of low-frequency appearance modeling.
[0036] Figure 3 It is a schematic diagram of the optimization of detail enhancement modeling. DETAILED DESCRIPTION OF THE INVENTION
[0037] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to specific examples and the accompanying drawings.
[0038] A 3DGS outdoor scene rendering optimization method based on frequency domain separation, as Figure 1 shown, includes the following steps:
[0039] (1) Separation of low-frequency and high-frequency features.
[0040] Step 1. Through Fourier transform, the rendering image I r generated by 3DGS at each iteration for a given viewing angle and the ground truth image I g are respectively transformed into the frequency domain space to obtain the rendering frequency domain image f r and the ground truth frequency domain image f g .
[0041] The formula for transforming the time domain image I(x,y) to the frequency domain image f(u,v) using Fourier transform is:
[0042]
[0043] Among them, (u, v) are frequency-domain coordinates, (x, y) are time-domain coordinates, W and H are the width and height of the image respectively, and i is the imaginary unit.
[0044] Step 2: Move the DC (Direct Current Component) component of the spectrum of the rendered frequency-domain image f r and the ground truth frequency-domain image f g to the center of their spectra to obtain the rendered center frequency-domain image and the ground truth center frequency-domain image
[0045] Step 3: Based on the rendered center frequency-domain image and the ground truth center frequency-domain image construct the rendered mask M r and the ground truth mask M g .
[0046] The formula for constructing the mask M(u c , v c ) is:
[0047]
[0048] Among them, is the center coordinate of the frequency-domain center image f c (u c , v c ), and r is the given radius. The area where the mask value is 1 is the low-frequency area, and the area where the mask value is 0 is the high-frequency area.
[0049] Step 4: Multiply the rendered mask M r with the rendered center frequency-domain image to obtain the rendered low-frequency frequency-domain image Then perform the inverse Fourier transform on the rendered low-frequency frequency-domain image to obtain the rendered low-frequency image
[0050] Step 5: Multiply the inverted rendered mask M r with the rendered center frequency-domain image to obtain the rendered high-frequency frequency-domain image Then perform the inverse Fourier transform on the rendered high-frequency frequency-domain image to obtain the rendered high-frequency image Multiply the inverted ground truth mask M g with the ground truth center frequency-domain image to obtain the ground truth high-frequency frequency-domain image Then perform the inverse Fourier transform on the ground truth high-frequency frequency-domain image to obtain the ground truth high-frequency image
[0051] (2) Low-frequency appearance modeling, such as Figure 2 as shown.
[0052] The existing Decoupled Appearance Modeling method stitches together an appearance embedding of length 64 and a rendered image downsampled 32 times as the input to the existing appearance learning network, and uses the existing appearance learning network to generate an adjustment map for adjusting the appearance of the rendered image. However, the purpose of downsampling is to reduce the network's learning of high-frequency details. The existing decoupled appearance modeling method reduces local low-frequency details and cannot adapt to scenarios with large appearance changes and local brightness and shadow changes. Therefore, the present invention improves the existing decoupled appearance modeling method by increasing the length of the appearance embedding and modifying the appearance learning network to adapt to appearance changes, so as to solve the problem of appearance changes in the rendered image of 3DGS under complex lighting conditions in outdoor scenes.
[0053] Step 6: Initialize the appearance embedding and stitch together the rendered low-frequency image of the original size and the appearance embedding to obtain the stitched rendered image D.
[0054] The present invention defines an appearance embedding with a length (depth channel) of 96 and initializes it as a learnable parameter that satisfies the following standard normal distribution:
[0055]
[0056] where, σ 2 represents the variance, and σ 2 = 1×10 -4 , indicating that the initialization of the appearance embedding satisfies a normal distribution with a mean of 0 and a variance of σ 2 .
[0057] In this embodiment, an appearance embedding with a size of 1047 (length) × 640 (width) × 96 (depth channel) is stitched together with a rendered low-frequency image with a size of 1047 (length) × 640 (width) × 96 (depth channel) to obtain a stitched rendered image D with a size of 1047 (length) × 640 (width) × 99 (depth channel).
[0058] Step 7: Feed the stitched rendered image D into the improved appearance learning network for learning to obtain an adjustment map
[0059] The existing appearance learning network gradually enlarges the size of the feature map through PixelShuffle and convolutional operations, and uses the ReLU activation function to increase non-linearity. Considering that the input of the appearance learning network is changed to the spliced image D with the same size as the rendered image, the upsampling module of the existing appearance learning network is removed in the present invention, and the LReLU and CBAM attention mechanism modules are introduced to redesigned the appearance learning network.
[0060] See Figure 2 , the improved appearance learning network includes 2 convolutional layers, 1 LeakyReLU activation function layer, 1 Sigmoid activation function layer and 5 appearance learning modules. Each appearance learning module consists of 1 convolutional layer, 1 LeakyReLU activation function layer and 1 CBAM attention mechanism layer. The input of the convolutional layer forms the input of the appearance learning module. The output of the convolutional layer is connected to the input of the LeakyReLU activation function layer. The output of the LeakyReLU activation function layer is connected to the input of the CBAM attention mechanism layer. The output of the CBAM attention mechanism layer forms the output of the appearance learning module. The input of the first convolutional layer serves as the input of the improved appearance learning network. The output of the first convolutional layer is connected to the input of the LeakyReLU activation function layer. The output of the LeakyReLU activation function layer is connected to the input of the first appearance learning module. The output of the first appearance learning module is connected to the input of the second appearance learning module. The output of the second appearance learning module is connected to the input of the third appearance learning module. The output of the third appearance learning module is connected to the input of the fourth appearance learning module. The output of the fourth appearance learning module is connected to the input of the fifth appearance learning module. The output of the fifth appearance learning module is connected to the input of the second convolutional layer. The output of the second convolutional layer serves as the input of the Sigmoid activation function layer. The output of the Sigmoid activation function layer serves as the output of the improved appearance learning network.
[0061] In this embodiment, a stitched rendering D with dimensions 1047 (length) × 640 (width) × 99 (depth channel) is used as the input image, and the three dimensions are the image height, image width, and number of channels respectively. The input image will first pass through a convolutional layer with a convolutional kernel size of 3×3 and a stride of 1 (the depth channel of the output feature matrix is 256), and through the LeakyReLU activation function layer to obtain an output feature map with dimensions 1047×640×256. The feature map with dimensions 1047×640×256 is input into 5 consecutive appearance learning modules. In each appearance learning module, the image first passes through a convolutional layer with a convolutional kernel size of 3×3 and a stride of 1, then through the LeakyReLU activation function, and then is input into a CBAM attention mechanism module. For the feature map with dimensions 1047×640×256, after passing through the first appearance learning module, the depth channel of the output feature matrix of the convolutional layer is 128, and the obtained feature map is 1047×640×128. For the feature map with dimensions 1047×640×128, after passing through the second appearance learning module, the depth channel of the output feature matrix of the convolutional layer is 64, and the obtained feature map is 1047×640×64. For the feature map with dimensions 1047×640×64, after passing through the third appearance learning module, the depth channel of the output feature matrix of the convolutional layer is 32, and the obtained feature map is 1047×640×32. For the feature map with dimensions 1047×640×32, after passing through the fourth appearance learning module, the depth channel of the output feature matrix of the convolutional layer is 16, and the obtained feature map is 1047×640×16. For the feature map with dimensions 1047×640×16, after passing through the fifth appearance learning module, the depth channel of the output feature matrix of the convolutional layer is 8, and the obtained feature map is 1047×640×8. The feature map with dimensions 1047×640×8 passes through a convolutional layer with a convolutional kernel size of 3×3 and a stride of 1 (the depth channel of the output feature matrix is 3), and through the Sigmoid activation function layer to obtain an adjusted map with dimensions 1047×640×3.
[0062] Step 8, through the adjusted map perform a pixel transformation of pixel-by-pixel multiplication on the rendering I r to obtain an appearance-transformed rendering
[0063] (3) Detail enhancement modeling, as Figure 3 shown.
[0064] Since high-frequency images contain detailed information such as the texture and edges of the image, the present invention designs a CNN-based feature extraction network, using the high-frequency maps as the input of the feature extraction network respectively to further mine the key features in the image. In addition, since the high-frequency feature maps extracted by the feature extraction network remove the low-frequency parts caused by illumination and contrast changes, and the original structural similarity loss includes three parts: brightness, contrast, and structure, which is not applicable to high-frequency feature maps, the present invention decouples the original structural similarity loss and only retains the structure part to optimize the structural details.
[0065] Step 9: Render the high-frequency map and the ground truth high-frequency map are respectively fed into the feature extraction network for learning to obtain the rendered high-frequency feature map and the ground truth high-frequency feature map
[0066] Refer to Figure 3 , the feature extraction network includes 5 feature extraction modules, 1 splicing layer, 1 convolutional layer, 1 batch normalization layer, and 1 ReLU activation function layer. Each feature extraction module consists of 1 convolutional layer, 1 batch normalization layer, 1 ReLU activation function layer, and 1 ECA attention mechanism layer. The input of the convolutional layer forms the input of the feature extraction module. The output of the convolutional layer is connected to the input of the batch normalization layer. The output of the batch normalization layer is connected to the input of the ReLU activation function layer. The output of the ReLU activation function layer is connected to the input of the ECA attention mechanism layer. The output of the ECA attention mechanism layer forms the output of the feature extraction module. The input of the first feature extraction module is used as the input of the feature extraction network. The output of the first feature extraction module is connected to the input of the second feature extraction module. The output of the second feature extraction module is connected to the input of the third feature extraction module. The output of the first feature extraction module and the output of the third feature extraction module are connected to the input of the splicing layer. The output of the splicing layer is connected to the input of the fourth feature extraction module. The output of the fourth feature extraction module is connected to the input of the fifth feature extraction module. The output of the fifth feature extraction module is connected to the input of the convolutional layer. The output of the convolutional layer is connected to the input of the batch normalization layer. The output of the batch normalization layer is connected to the input of the ReLU activation function layer. The output of the ReLU activation function layer is used as the output of the feature extraction network.
[0067] In this embodiment, a high-frequency map with a size of 1047 (length) × 640 (width) × 3 (depth channels) or is used as the input image.
[0068] First, input the high-frequency image with a size of 1047×640×3 into 3 consecutive feature extraction modules. In each feature extraction module, the image first passes through a convolutional layer with a convolutional kernel size of 3×3 and a stride of 1, a batch normalization layer, then passes through the activation function ReLU, and finally inputs an ECA attention mechanism module. For the feature map with a size of 1047×640×3, after passing through the first feature extraction module, the depth channel of the feature matrix output by the convolutional layer is 64, and the obtained feature map is 1047×640×64. For the feature map with a size of 1047×640×64, after passing through the second feature extraction module, the depth channel of the feature matrix output by the convolutional layer is 64, and the output feature map is 1047×640×64. For the feature map with a size of 1047×640×64, after passing through the third feature extraction module, the depth channel of the feature matrix output by the convolutional layer is 64, and the output feature map is 1047×640×64.
[0069] Then, the concatenation layer concatenates the output feature map of the first feature extraction module (with a size of 1047×640×64) and the output feature map of the third feature extraction module (with a size of 1047×640×64) to obtain a feature map with richer features, that is, a feature map with a size of 1047×640×128 as the input of the next layer.
[0070] Next, input the feature map with a size of 1047×640×128 into 2 consecutive feature extraction modules. In each feature extraction module, the image first passes through a convolutional layer with a convolutional kernel size of 3×3 and a stride of 1, a batch normalization layer, then passes through the activation function ReLU, and finally inputs an ECA attention mechanism module. For the feature map with a size of 1047×640×128, after passing through the fourth feature extraction module, the depth channel of the feature matrix output by the convolutional layer is 128, and the output feature map is 1047×640×128. For the feature map with a size of 1047×640×128, after passing through the fifth feature extraction module, the depth channel of the feature matrix output by the convolutional layer is 128, and the output feature map is 1047×640×128.
[0071] Finally, for the feature map with a size of 1047×640×128, first pass through a convolutional layer with a convolutional kernel size of 3×3 and a stride of 1 (the depth channel of the output feature matrix is 64), a batch normalization layer, and then pass through the activation function ReLU to obtain a feature map with a size of 1047×640×64.
[0072] (4) Optimization.
[0073] Step 10. Based on the rendering image I r and the ground truth image I g and the appearance transformation rendering image render the high-frequency feature map and the ground truth high-frequency feature map Calculate the total loss And by minimizing the total loss Optimize each iteration of 3DGS.
[0074] Total loss Consists of the photometric loss corresponding to the low-frequency appearance modeling And the decoupled structural similarity loss corresponding to the detail enhancement modeling That is:
[0075]
[0076] In the formula, Represents the photometric loss corresponding to the low-frequency appearance modeling Represents the decoupled structural similarity loss corresponding to the detail enhancement modeling, α represents the low-frequency appearance weight, β represents the detail enhancement weight, and α + β = 1.
[0077] The above photometric loss Is:
[0078]
[0079] In the formula, Represents the L1 loss Represents the structural similarity loss, and λ represents the photometric weight.
[0080] L1 loss Is:
[0081]
[0082] In the formula, Represents the pixel value of the i-th pixel of the rendered image, and I g (i) represents the pixel value of the i-th pixel of the ground truth image, and N represents the number of pixels in the image.
[0083] Structural similarity loss Is:
[0084]
[0085] In the formula, Represents the mean of the rendered image Represents the mean of the ground truth image Represents the covariance between the rendered image and the ground truth image Represents the variance of the rendered image Represents the variance of the ground truth image, and C1 and C2 are both given non-zero constants.
[0086] The above decoupled structural similarity loss is:
[0087]
[0088] In the formula, represents the covariance between the rendered high-frequency feature map and the ground truth high-frequency feature map, represents the standard deviation of the rendered high-frequency feature map, represents the standard deviation of the ground truth high-frequency feature map, and C3 is a given non-zero constant.
[0089] The present invention converts an image into the frequency domain space, obtains a low-frequency image and a high-frequency image through frequency domain separation and inverse Fourier transform. For the low-frequency image, it is learned through a low-frequency appearance modeling method to adapt to the appearance changes of the ground truth image during the training process. For the high-frequency image, a detail enhancement modeling method is adopted for learning to strengthen the detail expression of the rendered image, which can solve the problem of insufficient detail representation in the rendered image of the outdoor scene. The combination of the above two methods solves the problems of insufficient detail representation and insufficient adaptability to appearance changes of 3DGS in a large-scale outdoor scene.
[0090] It should be noted that although the embodiments described above of the present invention are illustrative, this is not a limitation of the present invention. Therefore, the present invention is not limited to the above specific embodiments. Without departing from the principle of the present invention, any other embodiments obtained by those skilled in the art under the inspiration of the present invention are regarded as within the protection scope of the present invention.
Claims
1. A 3DGS outdoor scene rendering optimization method based on frequency domain separation, characterized in that It includes the following steps: Step 1: Through Fourier transform, convert the rendered image I of a given perspective generated by 3DGS in each iteration and the ground truth image I r r into the frequency domain space respectively to obtain the rendered frequency domain image f g g and the ground truth frequency domain image f r r ; g ; Step 2: Move the DC components of the spectra of the rendered frequency-domain map f r and the ground truth frequency-domain map f g to the center of their spectra to obtain the rendered centered frequency-domain map and the ground truth centered frequency-domain map Step 3: Based on the rendered center frequency domain map and the ground truth center frequency domain map respectively construct a rendered mask M r and a ground truth mask M g ; Step 4: Multiply the rendering mask M r with the rendering center frequency domain diagram to obtain the rendered low-frequency domain diagram Then, perform an inverse Fourier transform on the rendered low-frequency domain diagram to obtain the rendered low-frequency diagram Step 5: Take the inverse of the rendering mask M r and multiply it with the rendering center frequency domain image to obtain the rendering high-frequency frequency domain image Then, perform an inverse Fourier transform on the rendering high-frequency frequency domain image to obtain the rendering high-frequency image Take the inverse of the ground truth mask M g and multiply it with the ground truth center frequency domain image to obtain the ground truth high-frequency frequency domain image Then, perform an inverse Fourier transform on the ground truth high-frequency frequency domain image to obtain the ground truth high-frequency image Step 6: Initialize the appearance embedding e that satisfies the normal distribution, and splice the rendered low-frequency map and the appearance embedding e to obtain the spliced rendering map D; Step 7: Send the spliced rendering image D into the improved appearance learning network for learning to obtain the adjusted image M; Step 8. By adjusting the figure Perform pixel multiplication pixel transformation on the rendered image I r to obtain the appearance transformation rendered image Step 9, send the rendered high-frequency map and the ground truth high-frequency map into the feature extraction network for learning respectively, to obtain the rendered high-frequency feature map and the ground truth high-frequency feature map Step 10: Based on the rendering graph I r , the ground truth graph I g , the appearance transformation rendering graph render the high-frequency feature graph and the ground truth high-frequency feature graph calculate the total loss and optimize each iteration of the 3DGS by minimizing the total loss ; In the formula, represents the L1 loss, represents the structural similarity loss, represents the decoupled structural similarity loss, I r represents the rendered image, I g represents the ground truth image, represents the appearance-transformed rendered image, represents the rendered high-frequency feature map, represents the ground truth high-frequency feature map, λ represents the photometric weight, α represents the low-frequency appearance weight, β represents the detail enhancement weight, and α + β = 1.
2. The 3DGS outdoor scene rendering optimization method based on frequency domain separation according to claim 1, characterized in that In step 6, the appearance embedding has a length of 96.
3. A 3DGS outdoor scene rendering optimization method based on frequency domain separation according to claim 1, characterized in that In Step 7, the improved appearance learning network includes 2 convolutional layers, 1 LeakyReLU activation function layer, 1 Sigmoid activation function layer, and 5 appearance learning modules; Each appearance learning module is composed of 1 convolutional layer, 1 LeakyReLU activation function layer, and 1 CBAM attention mechanism layer. The input of the convolutional layer forms the input of the appearance learning module. The output of the convolutional layer is connected to the input of the LeakyReLU activation function layer. The output of the LeakyReLU activation function layer is connected to the input of the CBAM attention mechanism layer. The output of the CBAM attention mechanism layer forms the output of the appearance learning module; The input of the first convolutional layer serves as the input of the improved appearance learning network. The output of the first convolutional layer is connected to the input of the LeakyReLU activation function layer. The output of the LeakyReLU activation function layer is connected to the input of the first appearance learning module. The output of the first appearance learning module is connected to the input of the second appearance learning module. The output of the second appearance learning module is connected to the input of the third appearance learning module. The output of the third appearance learning module is connected to the input of the fourth appearance learning module. The output of the fourth appearance learning module is connected to the input of the fifth appearance learning module. The output of the fifth appearance learning module is connected to the input of the second convolutional layer. The output of the second convolutional layer is connected to the input of the Sigmoid activation function layer. The output of the Sigmoid activation function layer serves as the output of the improved appearance learning network.
4. A 3DGS outdoor scene rendering optimization method based on frequency domain separation according to claim 1, characterized in that, In Step 9, the feature extraction network includes 5 feature extraction modules, 1 splicing layer, 1 convolutional layer, 1 batch normalization layer, and 1 ReLU activation function layer; Each feature extraction module is composed of 1 convolutional layer, 1 batch normalization layer, 1 ReLU activation function layer, and 1 ECA attention mechanism layer. The input of the convolutional layer forms the input of the feature extraction module. The output of the convolutional layer is connected to the input of the batch normalization layer. The output of the batch normalization layer is connected to the input of the ReLU activation function layer. The output of the ReLU activation function layer is connected to the input of the ECA attention mechanism layer. The output of the ECA attention mechanism layer forms the output of the feature extraction module; The input of the first feature extraction module serves as the input of the feature extraction network. The output of the first feature extraction module is connected to the input of the second feature extraction module. The output of the second feature extraction module is connected to the input of the third feature extraction module. The output of the first feature extraction module and the output of the third feature extraction module are connected to the input of the splicing layer. The output of the splicing layer is connected to the input of the fourth feature extraction module. The output of the fourth feature extraction module is connected to the input of the fifth feature extraction module. The output of the fifth feature extraction module is connected to the input of the convolutional layer. The output of the convolutional layer is connected to the input of the batch normalization layer. The output of the batch normalization layer is connected to the input of the ReLU activation function layer. The output of the ReLU activation function layer serves as the output of the feature extraction network.
5. A 3DGS outdoor scene rendering optimization method based on frequency domain separation according to claim 1, characterized in that, In Step 10, L1 loss is as follows: Structural similarity loss is as follows: Decoupled Structural Similarity Loss is as follows: Wherein, represents the pixel value of the i-th pixel of the rendered image, and I g (i) represents the pixel value of the i-th pixel of the ground truth image, and N represents the number of pixels in the image; represents the mean value of the rendered image, represents the mean value of the ground truth image, represents the covariance between the rendered image and the ground truth image, represents the variance of the rendered image, represents the variance of the ground truth image; represents the covariance between the rendered high-frequency feature map and the ground truth high-frequency feature map, represents the standard deviation of the rendered high-frequency feature map, represents the standard deviation of the ground truth high-frequency feature map; C1, C2, and C3 are all given non-zero constants.