New view synthesis method based on Uformer model and 3D Gaussian scattering technology
By combining the Uformer model and 3D Gaussian scattering technology and introducing cross-guided optimization strategies, the lack of performance of the new view synthesis technology when processing blurred input is solved, and efficient image defuzzing and new view synthesis are achieved, which significantly improves image quality and processing efficiency.
Patent Information
- Application Number
- CN202510274901.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
Smart Images

Figure CN120219598A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of novel view synthesis, and more particularly, to a novel view synthesis method based on the Uformer model and 3D Gaussian scattering technology. Background Art
[0002] Novel view synthesis (NVS) is a core technology in computer vision and computer graphics that can generate images from new perspectives from an existing set of images. The applications of this technology are extremely extensive, including from street view navigation to the enhancement of augmented reality (AR) and virtual reality (VR) experiences, and to providing environmental understanding and navigation support for autonomous vehicles and robotics. In recent years, the progress of multi-view geometry and deep learning has significantly improved the quality and efficiency of NVS, but these methods still have limitations when dealing with challenges common in real-world scenarios.
[0003] The emergence of Neural Radiance Field (NeRF) technology marks a major breakthrough in the field of NVS. It learns the continuous 5D radiance field of the scene (including spatial position and viewing direction) through a deep multi-layer perceptron (MLP), and successfully renders high-quality new perspective images. Although NeRF can produce extremely realistic rendering effects, it relies on dense voxel rendering technology, which requires a large number of samples of the light path during the rendering process to estimate pixel values, thus bringing a huge computational burden.
[0004] To overcome these limitations, researchers have explored alternative scene representation and rendering methods, including 3D Gaussian scattering (3D-GS). 3D-GS explicitly represents the scene through a set of colored 3D Gaussian functions, avoiding the dense sampling process in NeRF. These Gaussian functions optimize the parameters through a differentiable scattering process, significantly reducing the computational cost and training time, and improving the practicality and popularity of NVS technology in diverse environments.
[0005] However, in practical applications when dealing with blurred inputs (such as camera motion blur or depth of field blur), the performance of 3D-GS is still limited. Blurred inputs not only reduce the level of detail in the image but also increase the difficulty of estimating the scene geometry and appearance attributes, directly affecting the quality of novel view synthesis. Although 3D-GS has made progress in processing speed and efficiency, its limitations in blurred image inputs remain a major challenge in improving the performance of NVS.
[0006] In addition, although there are various image deblurring techniques in the NVS field, most methods only independently process the blurring problem at the 2D image level and fail to fully utilize 3D scene information. Although these traditional single-image deblurring methods can improve image clarity to a certain extent, they usually do not consider the 3D structure of the scene, limiting their effectiveness in dealing with complex scenes or severe blurring. In addition, these methods often rely on prior assumptions about the type of blurring (such as assuming that motion blur is linear or has a specific shape), which is often difficult to achieve in practical applications. Summary of the Invention
[0007] The object of the present invention is to overcome the deficiencies of the above-mentioned prior art and provide a novel view synthesis method based on the Uformer model and 3D Gaussian scattering technology.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] A novel view synthesis method based on the Uformer model and 3D Gaussian scattering technology, comprising the following steps:
[0010] Step S1, using a high-resolution camera to capture images of the target scene from multiple angles and preprocessing the captured images;
[0011] Step S2, applying the Transformer-based Uformer model to perform in-depth clarification processing on the preprocessed images and outputting clarified images;
[0012] Step S3, based on the clarified images, using 3D Gaussian scattering technology to simulate the light scattering effects under different perspectives;
[0013] Step S4, adopting a cross-guided optimization strategy to introduce 3D Gaussian scattering information into each processing layer of the Uformer model and strengthening the adaptability of the Uformer model to light scattering by setting cross layers;
[0014] Step S5, training the model constructed in Step S4 on an image data set containing different illuminations and different degrees of blurring;
[0015] Step S6, after the training of the constructed model is completed, evaluating the performance indicators of the constructed model and further optimizing the performance of the constructed model based on the evaluation results;
[0016] Step S7, using the trained and optimized model to process new scene images and synthesizing images of observing the scene from a new perspective.
[0017] Further, in step S2, the Uformer model has multiple processing layers for different blur levels, and the Uformer model dynamically adjusts the parameters and scales of each processing layer according to the blur level of the image.
[0018] Further, the adjustment mechanism of the Uformer model is as follows:
[0019]
[0020] In Equation (1), Q represents the query, K represents the key, V represents the value matrix, and d k represents the dimension of the key vector.
[0021] Further, in step S3, the 3D Gaussian scattering technology simulates the propagation of light in the scene based on the following scattering formula:
[0022]
[0023] In Equation (2), I0(x,y) represents the pixel intensity without scattering, σ represents the parameter related to the scattering radius, and characterizes the distribution of the scattering intensity.
[0024] Further, in step S4, the cross layer adopts dynamic adjustment, enabling the Uformer model to adaptively adjust the number and guiding strength of the cross layer according to different image features and scattering effects.
[0025] Further, the adjustment of the guiding strength of the cross layer is based on the feedback of the loss function, and the loss function is as follows:
[0026]
[0027] In Equation (3), is the set of rendered images, is the corresponding set of real images, and α and β are weight parameters for adjusting the contributions of the two losses.
[0028] Further, in step S5, the Adam optimizer is used for model training, the initial learning rate is set to 0.001, and it is automatically adjusted according to the change of the loss function. The learning rate adjustment strategy is:
[0029]
[0030] In Equation (3), t is the number of iterations, and LR0 is the initial learning rate.
[0031] The beneficial effects of the present invention are as follows: By combining the Uformer model and the 3D Gaussian scattering technology, and introducing a cross-guided optimization strategy, the present invention significantly improves the performance and efficiency of image processing. First, the Uformer model adopts a self-attention mechanism to perform feature decomposition and reconstruction on blurred images, which can precisely focus on local details and global context in the image, thereby effectively recovering clear details and textures from severely blurred images and enhancing the clarity and visual quality of the images. Second, the 3D Gaussian scattering technology creates more accurate depth maps and viewpoint images by simulating the propagation and scattering process of light in three-dimensional space and combining information collected from different viewpoints. This not only improves the accuracy of generating images from new viewpoints but also maintains the consistency of visual content, especially when maintaining the geometric and optical properties of the scene during viewpoint transformation. Finally, through the cross-guided optimization strategy, dynamic adjustment of model parameters is achieved, self-optimizing according to real-time performance feedback, significantly improving the processing efficiency, reducing the demand for computing resources, and lowering energy consumption and operation costs. The combination of these innovations makes the present invention have wide practicability and excellent performance in image deblurring, viewpoint generation, and other image processing applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 FIG. is a framework flowchart of a novel view synthesis method based on the Uformer model and 3D Gaussian scattering technology in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0034] Embodiment: A novel view synthesis system based on the Uformer model and 3D Gaussian scattering technology, as Figure 1 shown, includes two groups of 3D Gaussian scattering models and one group of Uformer models, and forms two branches, namely a leading branch and a subsequent branch. The leading branch uses the Uformer model to restore the image from multi-view blurred inputs before 3D Gaussian scattering, and the subsequent branch directly applies 3D Gaussian scattering to the blurred inputs and uses the same Uformer model to restore the synthesized views. Through the cross-guided optimization of the leading branch and the subsequent branch, the Uformer model is fine-tuned, so that the 3D Gaussian scattering model can achieve high-quality reconstruction under blurred inputs and maintain Figure 1 consistency.
[0035] A new view synthesis method based on the Uformer model and 3D Gaussian scattering technology, as Figure 1 shown, includes the following steps:
[0036] Step S1, use a high-resolution camera to capture images of the target scene from multiple angles. To ensure the high quality of the synthesized view, at least 20 shooting points are set within 360 degrees in an equiangular interval manner, and the resolution of the images captured at each shooting point is not lower than 2048x2048 pixels; preprocess the captured images to improve the quality and efficiency of subsequent processing. The preprocessing includes:
[0037] Denoising: Use a noise estimation algorithm, such as BM3D, to reduce random noise in the images;
[0038] Color correction: Adjust the color balance of the images to ensure the true restoration of colors;
[0039] Contrast enhancement: Use adaptive histogram equalization to enhance the image contrast.
[0040] Step S2, use the Uformer model to perform in-depth deblurring processing on the preprocessed images; the Uformer model is a deep learning model based on the Transformer structure, specifically designed for processing image deblurring problems; the input of the Uformer model is a set of preprocessed blurred images, and it outputs deblurred images by learning the detailed features and context information of the images.
[0041] To adapt to different degrees of blurring situations, the Uformer model has multiple processing layers to correspond to different blurring levels. The depth and width of the model are adjusted according to specific application scenarios. Generally, the model depth is 12 layers, and the dimension of each hidden layer is 384.
[0042] The Uformer model can dynamically adjust the parameters and scales of each processing layer according to the blurring level of the input images to achieve the optimal deblurring effect. Its adjustment mechanism is as follows:
[0043]
[0044] In Equation (1), Q represents the query, K represents the key, V represents the value matrix, and d k represents the dimension of the key vector.
[0045] The Uformer model includes a blur detector and a scale adjustment module. Among them, the blur detector (BlurDetector) is used to evaluate the blur level of the input image, which can be a simple image sharpness evaluation algorithm, such as the Laplacian algorithm, or a deep learning-based classifier. The scale adjustment module adjusts the parameters of the subsequent layers according to the detection results of the blur detector, such as the filter size and the number of layers. For example, for highly blurred images, increase the number of deep layers and adjust the filter size to capture broader context information.
[0046] For different blur levels, the Uformer model automatically adjusts the weights and biases of each layer through a dedicated parameter optimization network to ensure effective processing of different types of blur.
[0047] Through the above adjustment mechanism, the Uformer model not only improves the accuracy of image sharpness but also significantly enhances the adaptability and flexibility of the model to different blur states. This dynamic adjustment strategy ensures that the model can achieve the best performance when processing various blurred images, thus realizing high-quality image output during the view synthesis process.
[0048] Step S3, based on the obtained sharpened image, use the 3D Gaussian scattering technique to simulate the light scattering effect under different viewpoints, and simulate the propagation of light in the scene based on the following scattering formula:
[0049]
[0050] In Equation (2), I0(x,y) represents the pixel intensity without scattering, and σ represents the parameter related to the scattering radius, characterizing the distribution of the scattering intensity.
[0051] This technique is based on the principle of physical light propagation and can effectively simulate the reflection, refraction, and scattering processes of light in the real world.
[0052] At the same time, set the parameters of the 3D Gaussian scattering, including the scattering radius and the scattering intensity, to adapt to the visual effects under different materials and lighting conditions. The scattering radius is usually set between 0.5 and 1.5, and the scattering intensity is adjusted according to the lighting conditions of the scene to ensure the realism of the synthesized image.
[0053] By adopting the 3D Gaussian scattering technique, not only the visual quality of the image is enhanced, but also support is provided for scene rendering under complex lighting and texture conditions, greatly improving the naturalness and accuracy of the visual effect.
[0054] Step S4. Combining the advantages of the Uformer model and 3D Gaussian scattering technology, a cross-guided optimization strategy is adopted to further improve the effect of novel view synthesis. This strategy introduces 3D Gaussian scattering information into each processing layer of the Uformer model, enabling the Uformer model to consider the scattering effect of light in three-dimensional space while sharpening the image. By setting 3 to 5 cross layers, the adaptability of the Uformer model to light scattering is strengthened. At the same time, through the dynamic adjustment of the cross layers, the Uformer model can adaptively adjust the number of cross layers and the guiding strength according to different image features and scattering effects.
[0055] The implementation mechanism of the cross-guided optimization strategy is as follows:
[0056] (1) Design and function of the cross layers: Cross layers are embedded in each processing layer of the Uformer model. The role of these cross layers is to introduce the information of the 3D Gaussian scattering model during the blurring process, enhancing the naturalness and authenticity of view synthesis by simulating light scattering. The cross layers achieve the fusion and control of information through specific functional modules, such as Fusion Gates, or Modulation Layers.
[0057] (2) Dynamic adjustment mechanism of the cross layers:
[0058] (2.1) Adjustment of the number of cross layers: Dynamically adjust the number of cross layers according to the complexity of the input image and the scattering parameters. For example, in scenes with rich image details or significant scattering effects, increase the number of cross layers to better handle these features.
[0059] (2.2) Adjustment of the guiding strength: Adjust the guiding strength of the scattering information in the cross layers through a learning mechanism. This adjustment is based on the feedback of the loss function, automatically optimizing the proportion of the scattering information to ensure the balance between the scattering effect and image sharpness.
[0060] (3) Feedback adjustment during training: During the model training process, automatically adjust the parameters of the cross layers by analyzing the quality of the images output by the model in real time (for example, by evaluating the sharpness of the images and the realism of the scattering effect). This process is achieved using the backpropagation and gradient descent algorithms of deep learning to ensure that each iteration optimizes in the direction of improving the quality of the synthesized images.
[0061] (4) Implementation of the optimization strategy: When implementing this optimization strategy, an optimizer such as Adam can be used to adjust the parameters, and appropriate learning rates and other hyperparameters can be set to support the requirements of dynamic adjustment and fast convergence.
[0062] The core of the cross-guidance optimization strategy lies in leveraging the complementary advantages between the Uformer model and the 3D Gaussian scattering technique, and achieving the optimization goal through joint training. This includes two key steps: First, use the Uformer model to process the blurred image and restore its clarity; Second, take the restored clear image as input and apply the 3D Gaussian scattering technique for image rendering at a new viewpoint. The optimization goal is achieved through the following loss function, which takes into account pixel-level accuracy and visual quality:
[0063]
[0064] In Equation (3), is the set of rendered images, is the corresponding set of ground-truth images, and α and β are weight parameters that regulate the contributions of the two losses.
[0065] Through this strategy, the system can simultaneously optimize the clarity of the image and the authenticity of the viewpoint rendering, and ultimately achieve the generation of high-quality images at different viewpoints. This cross-guidance method not only improves the efficiency of the model but also enhances its practicality and robustness in their respective application scenarios.
[0066] In this embodiment, the cross-guidance optimization strategy not only includes the design of the loss function but also involves how to adjust the weights and parameters of the Uformer model and the 3D Gaussian scattering model during the training process. This adaptive adjustment mechanism is based on the performance of the model on the training set, and by dynamically adjusting the values of α and β, as well as other relevant hyperparameters, to find the optimal model configuration.
[0067] At each stage of the joint training, the performance of the model on the validation set is evaluated, including the clarity of the image, the viewpoint consistency, and the rendering efficiency. Based on these evaluation results, the model adaptively adjusts its internal parameters and training strategies to ensure optimal performance in the face of different inputs and challenges.
[0068] Step S5, train the model constructed in step S4 on an image dataset containing different illuminations and different degrees of blur. Use the Adam optimizer for model training, set the initial learning rate to 0.001, and automatically adjust it according to the change of the loss function. The learning rate adjustment strategy is:
[0069]
[0070] In Equation (3), t is the number of iterations, and LR0 is the initial learning rate.
[0071] During the training process, the learning rate is gradually decreased according to a preset strategy to finely adjust the model parameters and improve the stability and prediction accuracy of the model. Specifically, a learning rate warm-up and decay strategy is adopted. The learning rate is gradually increased to 0.001 in the first 10 epochs, and then halved every 100 epochs until the training ends.
[0072] Step S6, after the constructed model is trained, use metrics such as PSNR (Peak Signal-to-Noise Ratio), SSIM (Structural Similarity Index), and LPIPS (Perceptual Image Patch Similarity) to evaluate the performance of the model. These metrics can reflect the image quality and visual effects from different perspectives. Then, further optimize the model performance according to the evaluation results. The key points of optimization include adjusting the depth and width of the Uformer model, fine-tuning the 3D Gaussian scattering parameters, and the parameters in the cross-guidance optimization strategy. The optimization process aims to maximize the realism and visual comfort of the image while ensuring image clarity.
[0073] Step S7, based on the trained and optimized model, process new scene images to synthesize new views. This step mainly relies on the scene depth information and light scattering model learned by the model. By adjusting the viewing angle and viewpoint, an image of the scene observed from a new perspective is synthesized.
[0074] During the new view synthesis process, the viewing angle and focus of the synthesized view can be specified through user input to achieve the purpose of generating personalized views. In addition, by adjusting the parameters of the 3D Gaussian scattering, different lighting conditions and atmospheric effects can be simulated to enhance the realism of the synthesized view.
[0075] Furthermore, this embodiment also adopts dynamic viewpoint synthesis technology, which greatly improves the image quality and viewpoint coherence during the new view synthesis process. This technology uses 3D scene information and depth maps to dynamically adjust the ray tracing and scattering parameters, optimizing the rendering process, thus achieving the following purposes:
[0076] (1) Dynamic ray tracing: Using 3D scene information, the dynamic viewpoint synthesis technology adjusts the behavior of the ray tracing algorithm, enabling it to flexibly adjust the ray path according to the change of the viewpoint. This flexible ray tracing ensures that when observing from different viewpoints, the interaction between the rays and 3D objects can be realistically reflected, enhancing the three-dimensional sense and depth sense of the image.
[0077] (2) Parameterized scattering control: By analyzing the depth map, this technology can adjust the scattering effect in the image, such as the diffusion degree of halos or shadows. This not only improves the visual quality of the image but also makes the light and shadow effects in the image more natural and realistic.
[0078] (3)Enhanced realism: This technology enhances the realism of images by precisely controlling the details and lighting effects during the rendering process. This is particularly important for high-quality image rendering of complex scenes, especially in applications that require advanced visual effects, such as virtual reality and augmented reality.
[0079] (4)Improved rendering efficiency: Dynamically adjusting the ray tracing and scattering parameters not only improves the image quality but also optimizes the efficiency of the rendering process. This means that high-quality viewpoint images can be quickly generated even on devices with limited resources.
[0080] Through the application of these technologies, dynamic viewpoint synthesis not only improves the quality and coherence of images but also ensures consistent visual effects when the viewpoint changes, making the user experience smoother and more immersive. The comprehensive application of these technologies gives this embodiment significant advantages and application potential in the fields of image synthesis and virtual scene rendering.
[0081] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.
Claims
1. A new view synthesis method based on Uformer model and 3D Gaussian scattering technology, characterized in that: The following steps are involved: Step S1, using a high-resolution camera to capture images of a target scene from multiple angles, and preprocessing the captured images; Step S2, applying the Transformer-based Uformer model to perform in-depth sharpening processing on the preprocessed image and outputting the sharpened image; Step S3, based on the cleared image, using 3D Gaussian scattering technology to simulate light scattering effects at different viewing angles; Step S4, adopting a cross-guided optimization strategy, introducing 3D Gaussian scattering information in each processing layer of the Uformer model, and setting cross layers to enhance the adaptability of the Uformer model to light scattering; Step S5, training the model constructed in step S4 on a dataset of images with different illuminations and different blur levels; Step S6, after the constructed model training is completed, the performance indicators of the constructed model are evaluated, and the performance of the constructed model is further optimized according to the evaluation results; Step S7, using the trained and optimized model to process the new scene image, and synthesize an image of the scene observed from a new perspective.
2. According to claim 1, a new view synthesis method based on Uformer model and 3D Gaussian scattering technology is characterized in that: In step S2, the Uformer model is provided with a plurality of processing layers for different blur levels, and the Uformer model dynamically adjusts the parameters and scale of each processing layer according to the blur level of the image.
3. According to claim 2, a new view synthesis method based on Uformer model and 3D Gaussian scattering technology is characterized in that: The adjustment mechanism of the Uformer model is as follows: In formula (1), Q represents query, K represents key, V represents value matrix, and d k Represents the dimension of the key vector.
4. According to claim 1, a new view synthesis method based on Uformer model and 3D Gaussian scattering technology is characterized in that: In step S3, the 3D Gaussian scattering technique simulates the propagation of light in the scene based on the following scattering formula: In formula (2), I0(x, y) represents the pixel intensity without scattering, and σ represents a parameter related to the scattering radius, which characterizes the distribution of the scattering intensity.
5. According to claim 1, a new view synthesis method based on Uformer model and 3D Gaussian scattering technology is characterized in that: In step S4, the cross layer is dynamically adjusted so that the Uformer model can adaptively adjust the number and guidance strength of the cross layer according to different image features and scattering effects.
6. According to claim 5, a new view synthesis method based on Uformer model and 3D Gaussian scattering technology is characterized in that: The adjustment of the cross-layer guidance strength is based on the loss function feedback, and the loss function is as follows: In formula (3), is a collection of rendered images, is the corresponding set of real images, and α and β are weight parameters that adjust the contribution of the two losses.
7. According to claim 1, a new view synthesis method based on Uformer model and 3D Gaussian scattering technology is characterized in that: In step S5, the Adam optimizer is used for model training. The initial learning rate is set to 0.001 and automatically adjusted according to the change of the loss function. The learning rate adjustment strategy is: In formula (3), t is the number of iterations and LR0 is the initial learning rate.