Sparse dynamic 3D Gaussian splash method based on global-local feature extraction
Through the global-local feature extraction module, the inter-frame feature flow aggregation network and the deep prior combined with the 4D pseudo-pose supervision mechanism, the reconstruction incoherence and detail loss problems of dynamic scenes under sparse perspective are solved, and efficient dynamic scene rendering is achieved, which is suitable for virtual reality and augmented reality applications.
Patent Information
- Application Number
- CN202510720934.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-16
AI Technical Summary
Existing 3D GS methods have difficulty in accurately capturing the details of dynamic scenes under sparse viewing angles, have insufficient feature representation, and have poor temporal consistency, resulting in incoherent reconstruction results that are difficult to meet real-time rendering requirements.
A global-local feature extraction module, an inter-frame feature flow aggregation network, and a deep prior combined with a 4D pseudo-pose supervision mechanism are adopted, combined with a sparse dynamic 3D Gaussian splash method to fuse global and local features, solve the temporal consistency problem, and improve reconstruction accuracy and robustness.
High-quality dynamic scene reconstruction and rendering are achieved under sparse camera view, ensuring temporal consistency and real-time performance, reducing motion blur and flicker, and meeting the needs of applications such as virtual reality and augmented reality.
Smart Images

Figure CN120655799A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing technology, in particular to the field of computer vision and graphics technology, and relates to a dynamic view synthesis technology, in particular to a dynamic 3D Gaussian splash method for sparse camera perspectives. Background Art
[0002] Dynamic Novel View Synthesis (D-NVS) technology aims to generate dynamic scene renderings from arbitrary virtual perspectives from inputs from multiple camera viewpoints. While traditional Neural Radiance Field (NeRF)-based methods can achieve high-quality renderings, their training and rendering efficiency are low, making them difficult to meet the requirements of real-time applications.
[0003] In recent years, the 3D Gaussian splatter (3D GS) method, as an explicit representation method, has attracted widespread attention due to its efficient rendering capabilities. However, existing 3D GS methods still have the following problems when processing dynamic scenes: First, there is a performance bottleneck under sparse perspectives. In practical applications, the number of camera deployments is often limited, resulting in sparse input perspectives. Under sparse perspective conditions, traditional 3D GS methods have difficulty accurately capturing the dynamic changes of the scene and are prone to detail loss and geometric distortion problems. Secondly, there is insufficient feature representation: existing feature extraction methods cannot effectively fuse global scene features and local detail features, resulting in a lack of detail in the reconstructed scene, especially in fast-moving areas. Finally, there is poor temporal consistency: the motion of objects in dynamic scenes may cause incoherence in the reconstruction results between adjacent frames, affecting the visual experience. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of the existing technology and provide a sparse dynamic 3D Gaussian splash method based on global-local feature extraction. This method can achieve high-quality dynamic scene reconstruction and rendering under sparse camera viewing angles while ensuring temporal consistency and real-time performance. Under sparse camera viewing angles, the present invention accurately captures the geometric structure and texture details of dynamic scenes; effectively integrates global scene features and local detail features to improve reconstruction accuracy; ensures temporal consistency between adjacent frames, reducing motion blur and flicker; and supports real-time rendering to meet practical application requirements.
[0005] The method of the present invention includes three core components: a global-local feature extraction module, an inter-frame feature flow aggregation network, and a deep prior combined with a 4D pseudo-pose supervision mechanism.
[0006] The global-local feature extraction module extracts global and local features from the input sparse viewport image and fuses them. For global feature extraction, a six-plane encoder is used to encode the global structure of the scene. For local feature extraction, a stereo hash grid volume encoder is used to encode the local details of the scene. The global and local features are then further decomposed into static and dynamic features. This feature decomposition design enhances the representation capabilities of sparse point clouds.
[0007] The inter-frame feature flow aggregation network addresses temporal consistency in dynamic scenes. It achieves spatiotemporal alignment of dynamic features by predicting inter-frame feature flows. It first uses a stream MLP to predict the feature flows between adjacent frames. Based on the predicted feature flows, it then warps and aggregates the dynamic features of adjacent frames, addressing temporal consistency in sparse dynamic scenes.
[0008] A depth prior combined with a 4D pseudo-pose supervision mechanism is used to optimize the geometry of the sparse Gaussian splatter model, improving reconstruction accuracy. A pretrained monocular depth estimation model (such as DPT) is used to generate a depth map for each input image, and this depth information is incorporated into model training as prior knowledge. Furthermore, by generating pseudo camera poses and corresponding pseudo depth maps, the diversity of the training data is increased, improving the model's generalization capabilities.
[0009] The specific steps of the method of the present invention are: Step (1) for each first frame of the sparse input video, the structure-to-motion (SFM) algorithm is used to generate an initialized Gaussian sparse point cloud structure; Step (2) extracting global features and local features of the scene using a global-local feature extraction module on the generated Gaussian sparse point cloud; The global-local dynamic features extracted in step (3) are aggregated through inter-frame feature flows to enhance inter-frame feature consistency; Step (4) uses depth prior and 4D pseudo pose for supervision; Step (5) optimizes the Gaussian splash model using a loss function; Step (6) renders the output result through the Gaussian splash process.
[0010] Compared with the existing technology, the present invention has the following beneficial effects: 1) Higher reconstruction accuracy: By fusing global and local features, the geometric structure and texture details of the scene can be captured more accurately, especially under sparse viewing angle conditions; 2) Better temporal consistency: The inter-frame feature flow aggregation network effectively solves the temporal incoherence problem in dynamic scenes and reduces motion blur and flickering; 3) Stronger robustness: The depth prior and 4D pseudo-pose supervision mechanism improves the robustness of the model under sparse data conditions and reduces the risk of overfitting; 4) Real-time rendering capability: The rendering method based on 3D Gaussian splash supports efficient real-time rendering, meeting the needs of application scenarios such as virtual reality and augmented reality. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 It is the overall flow chart of the method of the present invention; Figure 2 Schematic diagram of the global-local feature extraction module in the method of the present invention; Figure 3 Schematic diagram of the inter-frame flow feature aggregation module in the method of the present invention; Figure 4 Schematic diagram of the supervision mechanism of deep prior combined with 4D pseudo pose in the method of the present invention. DETAILED DESCRIPTION
[0012] A sparse dynamic 3D Gaussian splash method based on global-local feature extraction achieves high-quality dynamic scene reconstruction and rendering under sparse camera view conditions through the collaborative work of a global-local feature extraction module, an inter-frame feature flow aggregation network, and a supervision mechanism combining deep priors with 4D pseudo-pose. This method compresses multi-view image information into a feature space for processing, reducing computational complexity while effectively preserving the scene's geometric structure and dynamic features. The fused deep prior information and 4D pseudo-pose provide accurate geometric constraints, while the inter-frame feature flow aggregation network is specifically optimized for temporal consistency in dynamic scenes. The overall design is robust and efficient, making it particularly suitable for processing complex dynamic scenes under sparse view conditions, significantly improving the realism and real-time performance of virtual view synthesis.
[0013] The global-local feature extraction module extracts global and local features from the input sparse viewport image and fuses them. For global feature extraction, a six-plane encoder is used to encode the global structure of the scene. For local feature extraction, a stereo hash grid volume encoder is used to encode the local details of the scene. The global and local features are then further decomposed into static and dynamic features. This feature decomposition design enhances the representation capabilities of sparse point clouds.
[0014] The inter-frame feature flow aggregation network addresses temporal consistency in dynamic scenes. It achieves spatiotemporal alignment of dynamic features by predicting inter-frame feature flows. It first uses a stream MLP to predict the feature flows between adjacent frames. Based on the predicted feature flows, it then warps and aggregates the dynamic features of adjacent frames, addressing temporal consistency in sparse dynamic scenes.
[0015] A depth prior combined with a 4D pseudo-pose supervision mechanism is used to optimize the geometry of the sparse Gaussian splatter model, improving reconstruction accuracy. A pretrained monocular depth estimation model (such as DPT) is used to generate a depth map for each input image, and this depth information is incorporated into model training as prior knowledge. Furthermore, by generating pseudo camera poses and corresponding pseudo depth maps, the diversity of the training data is increased, improving the model's generalization capabilities.
[0016] The overall process of this method is as follows Figure 1 The figure shows the complete processing process from the input of sparse camera frames to the generation of rendered images and depth maps, including the steps of generating an initialized Gaussian sparse point cloud using the structure-to-motion (SFM) algorithm, global-local feature extraction, inter-frame feature flow aggregation, supervision with depth priors and 4D pseudo poses, and various loss functions. The specific steps of this method are as follows: Step (1) For the first frame of each sparse input video, use the structure-to-motion (SFM) algorithm to generate an initialized Gaussian sparse point cloud structure. The details are as follows: (1-1) Acquisition of sparse video frames: Use several sparse cameras to capture image sequences of dynamic scenes from different perspectives.
[0017] (1-2) Camera calibration: Obtain the intrinsic and extrinsic parameters of each camera and establish the transformation relationship between the world coordinate system and the camera coordinate system.
[0018] (1-3) Initialize sparse point cloud: Use the structure-to-motion (SFM) algorithm to process the sparse view images and initialize the sparse point cloud.
[0019] Step (2) uses the global-local feature extraction module to extract the global and local features of the scene from the generated Gaussian sparse point cloud. Figure 2 As shown, the details are as follows: (2-1) Design a global six-plane feature encoder; the global six-plane feature encoder includes six multi-scale plane grid modules with the same low resolution: XY, XZ, YZ, XT, YT, and ZT. The Gaussian center position under time t The global feature encoding result is expressed as .
[0020] (2-2) Extract global static and dynamic features; input the sparsely initialized Gaussian center position into the global six-plane feature encoder to output the decomposable global static features of the dynamic scene and global dynamic features For global static features To extract the global dynamic features, we first perform bilinear interpolation on the XY, XZ, and YZ plane features without time information at the Gaussian center position, then perform pixel-by-pixel Hadamard product on the three static plane features, and finally connect the global static features at the two scales according to the channel dimension. To extract the global dynamic features, we first output the preliminary rough global dynamic features by bilinear interpolation of XT, YT, and ZT containing time information at the Gaussian center position, then perform the planar Hadamard product on the three dynamic planar features, and finally connect the global dynamic features of the two scales according to the channel dimension.
[0021] (2-3) Design a local hash grid feature encoder; the local hash grid feature encoder contains four multi-scale stereo hash grids with the same high resolution of XYZ, XYT, XZT, and YZT, and the Gaussian center position at time t The local feature encoding result is expressed as .
[0022] (2-4) Extract local static and dynamic features; input the sparsely initialized Gaussian center position into the local hash grid feature encoder and output the local static features and local dynamic features Local static features Directly perform linear interpolation of features from the static XYZ three-dimensional hash grid and connect the four scale features to obtain local dynamic features. The dynamic XYT, XZT, and YZT three-dimensional hash grid features are linearly interpolated and then Hadamard products are performed to reduce the number of parameters. The local dynamic features at four scales are then connected to obtain the feature.
[0023] The global-local dynamic features extracted in step (3) are aggregated through the inter-frame feature flow to enhance the consistency of inter-frame features. Figure 3 As shown, the details are as follows: (3-1) Predict the Gaussian center position flow; Combined with the multi-layer perceptron MLP for motion estimation, the 4D spatiotemporal coordinates after position encoding are input to predict the Gaussian center position flow between the previous and next frames: , and They represent the forward flow offset and backward flow offset of the current Gaussian 3D center position respectively, represents the position code, represents MLP encoding.
[0024] (3-2) Aggregate dynamic features; for the previous frame, calculate the current Gaussian center position and add the center position of the previous frame of the backward flow to obtain the center position of the Gaussian point of the previous frame after the backward flow is distorted ; For the next frame, calculate the center position of the current Gaussian point and add it to the center position of the next frame of the forward flow to get the center position of the Gaussian point of the next frame after the forward flow is distorted ; Then update the Gaussian center position at the current time t according to the weight The dynamic characteristics of , The final stream feature aggregation network outputs dynamic features after the aggregation of the previous and next frame features. Indicates the center position and time Dynamic global-local features of Gaussian points in the lower scene; (3-3) Final feature fusion: connect the global static features output by the previous step according to the channel dimension , local static features Dynamic characteristics of aggregation , output the final encoded features , put it into the multi-head Gaussian attribute decoder to decode the time domain deformation of each attribute, and learn the deformed Gaussian points ;in, represents a set of Gaussian center locations, represents a Gaussian ellipsoid point, represents the center position of this Gaussian point, A 4-element vector representing a Gaussian point, represents the scaling vector of the Gaussian point, represents the Gaussian point opacity, Represents the spherical harmonic coefficients of the Gaussian points.
[0025] Step (4) uses depth prior and 4D pseudo pose for supervision. Figure 4 As shown, the details are as follows: (4-1) Depth prior and Pearson loss term. First, a pre-trained 2D monocular depth estimator (such as the DPT model) is used to generate a monocular depth map at the current timestamp t during training. , which is then projected into 3D space with the depth map rendered by the Gaussian splash model Perform alignment and finally introduce the Pearson correlation loss to calculate the distribution difference between depth maps: , Cov represents covariance, and Var represents variance.
[0026] (4-2) Generate 4D pseudo camera pose. For any given timestamp t, first sample the synthetic viewpoint from the two closest training viewpoints cam0 and cam1 in Euclidean space, and then calculate the interpolated pseudo view deviation value of the camera Finally, we need to add 3 degrees of freedom random white noise To the current camera direction: .
[0027] (4-3) 4D pseudo depth supervision item. First, put its pseudo camera pose into the Gaussian depth map renderer to output the rendered depth map , which is then fed into a pre-trained depth estimator to output the pre-trained depth map , and finally get the 4D pseudo depth supervision item: .
[0028] Step (5) uses various loss functions to optimize the Gaussian splash model. The details are as follows: (5-1) Define the total loss function. First, use the L1 color loss and SSIM structural loss , then add Pearson depth loss , and finally introduce the 4D pseudo depth loss generated after interpolating the camera perspective : Total loss function ,in They represent the loss weights of color loss, SSIM structure loss, Pearson depth loss, and 4D pseudo depth loss respectively.
[0029] (5-2) Optimization process. The Adam optimizer is used to iteratively optimize the model parameters. The initial learning rate is set to 0.01 and gradually decays during the training process. They are set to 0.8, 0.2, 0.05, and 0.05 respectively. Pseudo view sampling is enabled after 1500 iterations, and the Gaussian deformation operation is not added in the first 3000 times to ensure that the Gaussian can roughly represent the scene at this time.
[0030] Step (6) renders the output result through the Gaussian splash process. The details are as follows: (6-1) Rendering from any perspective: For a given target camera pose, render using a 3D Gaussian splash model after deformation. (6-2) Differentiable rasterization: Through the differentiable rasterization process, the 3D Gaussian distribution is projected onto the 2D image plane to generate the final rendered image; (6-3) Real-time dynamic update: In dynamic scenes, as new frames are input, the model parameters are continuously updated to achieve real-time reconstruction and rendering of dynamic scenes.
[0031] In general, the present invention proposes a dynamic 3D Gaussian splash method for sparse perspectives by combining a global-local feature fusion architecture with spatiotemporal constraints of dynamic scenes. Compared with traditional dynamic view synthesis methods, the present invention shows significant advantages under sparse input conditions, and can generate dynamic views with more accurate geometric structures, stronger temporal continuity and richer details. In particular, the key components designed for sparse perspectives, such as the global-local feature extraction module, the inter-frame flow feature aggregation network and the deep prior combined with the 4D pseudo-pose supervision mechanism, enable this method to effectively solve the problem of feature incoherence, geometric instability and overfitting in dynamic scene reconstruction, and provide an efficient dynamic view synthesis solution for immersive applications such as virtual reality and augmented reality.
Claims
1. A sparse dynamic 3D Gaussian splash method based on global-local feature extraction, characterized by: This method achieves high-quality dynamic scene reconstruction and rendering under sparse camera view conditions by working together through a global-local feature extraction module, an inter-frame feature flow aggregation network, and a deep prior combined with a 4D pseudo-pose supervision mechanism. The global-local feature extraction module is used to extract global and local features of a scene from an input sparse view image and fuse the two. For global feature extraction, a six-plane encoder is used to encode the global structure of the scene; for local feature extraction, a stereo hash grid volume encoder is used to encode the local details of the scene. The global and local features are then decomposed into static and dynamic feature components, enhancing the representation capability of sparse point clouds. The inter-frame feature flow aggregation network is used to handle temporal consistency in dynamic scenes. It achieves spatiotemporal alignment of dynamic feature parts by predicting inter-frame feature flows. It first uses the stream MLP to predict the feature flows between adjacent frames, and then warps and aggregates the dynamic features of adjacent frames based on the predicted feature flows, thereby solving the temporal consistency problem in sparse dynamic scenes. The described depth prior combined with the 4D pseudo-pose supervision mechanism is used to optimize the geometric structure of the sparse Gaussian splash model and improve reconstruction accuracy; a pre-trained monocular depth estimation model is used to generate a depth map for each input image, and this depth information is incorporated into the model training as prior knowledge; by generating pseudo camera poses and corresponding pseudo depth maps, the diversity of training data is increased, thereby improving the generalization ability of the model.
2. The sparse dynamic 3D Gaussian splashing method based on global-local feature extraction according to claim 1, characterized in that: The steps of this method are as follows: Step (1) for each first frame of the sparse input video, the structure-to-motion (SFM) algorithm is used to generate an initialized Gaussian sparse point cloud structure; Step (2) extracting global features and local features of the scene using a global-local feature extraction module on the generated Gaussian sparse point cloud; The global-local dynamic features extracted in step (3) are aggregated through inter-frame feature flows to enhance inter-frame feature consistency; Step (4) uses depth prior and 4D pseudo pose for supervision; Step (5) optimizes the Gaussian splash model using a loss function; Step (6) renders the output result through the Gaussian splash process.
3. The sparse dynamic 3D Gaussian splashing method based on global-local feature extraction according to claim 2, characterized in that: Step (1) is as follows: (1-1) Use a sparse number of cameras to capture image sequences of dynamic scenes from different perspectives; (1-2) Obtain the intrinsic and extrinsic parameters of each camera and establish the conversion relationship between the world coordinate system and the camera coordinate system; (1-3) Use the structure-to-motion (SFM) algorithm to process the sparse view images and initialize the sparse point cloud.
4. The sparse dynamic 3D Gaussian splashing method based on global-local feature extraction according to claim 2, characterized in that: Step (2) is as follows: (2-1) Design a global six-plane feature encoder; the global six-plane feature encoder contains six multi-scale plane grid modules with the same low resolution: XY, XZ, YZ, XT, YT, and ZT. The Gaussian center position at time t The global feature encoding result is expressed as ; (2-2) Extract global static and dynamic features; The sparsely initialized Gaussian center position is input into the global six-plane feature encoder to output the global static features of the dynamic scene and global dynamic features ; For global static features To extract the time information, we first perform bilinear interpolation on the XY, XZ, and YZ plane features without time information at the Gaussian center position, then perform pixel-by-pixel Hadamard product on the three static plane features, and finally connect the global static features at the two scales according to the channel dimension. For global dynamic features To extract the initial rough global dynamic features, firstly, bilinear interpolation of XT, YT, and ZT containing time information is performed on the Gaussian center position, and then the plane Hadamard product is performed on the three dynamic plane features. Finally, the global dynamic features of the two scales are connected according to the channel dimension. (2-3) Design a local hash grid feature encoder; The local hash grid feature encoder contains four multi-scale stereo hash grids with the same high resolution: XYZ, XYT, XZT, and YZT. The Gaussian center position at time t The local feature encoding result is expressed as ; (2-4) Extract local static and dynamic features; input the sparsely initialized Gaussian center position into the local hash grid feature encoder and output the local static features and local dynamic features ; Local static features Directly perform linear interpolation of features from the static XYZ three-dimensional hash grid and connect the four scale features to obtain local dynamic features. It is obtained by linearly interpolating the dynamic XYT, XZT, and YZT three-dimensional hash grid features, performing Hadamard product, and then connecting the local dynamic features at four scales.
5. The sparse dynamic 3D Gaussian splashing method based on global-local feature extraction according to claim 2, characterized in that: Step (3) is as follows: (3-1) Predict the Gaussian center position flow; Combined with the multi-layer perceptron MLP for motion estimation, the 4D spatiotemporal coordinates after position encoding are input to predict the Gaussian center position flow between the previous and next frames: , and They represent the forward flow offset and backward flow offset of the current Gaussian 3D center position respectively, represents the position code, represents MLP encoding; (3-2) Aggregate dynamic features; for the previous frame, calculate the current Gaussian center position and add the center position of the previous frame of the backward flow to obtain the center position of the Gaussian point of the previous frame after the backward flow is distorted ; For the next frame, calculate the center position of the current Gaussian point and add it to the center position of the next frame of the forward flow to get the center position of the Gaussian point of the next frame after the forward flow is distorted. ; Then update the Gaussian center position at the current time t according to the weight The dynamic features of the final stream feature aggregation network output are the dynamic features after the aggregation of the previous and next frame features. , Indicates the center position and time Dynamic global-local features of Gaussian points in the lower scene; (3-3) Final feature fusion: connect the global static features output by the previous step according to the channel dimension , local static features Dynamic characteristics of aggregation , output the final encoded features , put it into the multi-head Gaussian attribute decoder to decode the time domain deformation of each attribute, and learn the deformed Gaussian points ;in, represents a set of Gaussian center locations, represents a Gaussian ellipsoid point, represents the center position of this Gaussian point, A 4-element vector representing a Gaussian point, represents the scaling vector of the Gaussian point, represents the Gaussian point opacity, Represents the spherical harmonic coefficients of the Gaussian points.
6. The sparse dynamic 3D Gaussian splashing method based on global-local feature extraction according to claim 2, characterized in that: Step (4) is as follows: (4-1) Depth prior and Pearson loss term; First, use the pre-trained 2D monocular depth estimator to generate a monocular depth map at the current timestamp t during training. , which is then projected into 3D space with the depth map rendered by the Gaussian splash model Perform alignment and finally introduce the Pearson correlation loss to calculate the distribution difference between depth maps; (4-2) Generate 4D pseudo camera pose; for any given timestamp t, first sample the synthetic viewpoint from the two closest training viewpoints cam0 and cam1 in Euclidean space, and then calculate the interpolated pseudo view deviation value of the camera , add 3 degrees of freedom random white noise To the current camera direction; (4-3) 4D pseudo depth supervision item; first put its pseudo camera pose into the Gaussian depth map renderer to output the rendered depth map , which is then fed into a pre-trained depth estimator to output the pre-trained depth map , and finally obtain the 4D pseudo depth supervision term.
7. The sparse dynamic 3D Gaussian splashing method based on global-local feature extraction according to claim 2, characterized in that: Step (5) is as follows: (5-1) Define the total loss function; first use L1 color loss and SSIM structure loss, then add Pearson depth loss, and finally introduce the 4D pseudo depth loss generated after interpolating the camera view; (5-2) Optimization process; The Adam optimizer is used to iteratively optimize the model parameters, which are gradually attenuated during the training process; pseudo view sampling is enabled after the set number of iterations is reached.
8. The sparse dynamic 3D Gaussian splashing method based on global-local feature extraction according to claim 2, characterized in that: Step (6) specifically includes: (6-1) Rendering from any perspective: For a given target camera pose, render using a 3D Gaussian splash model after deformation. (6-2) Differentiable rasterization: Through the differentiable rasterization process, the 3D Gaussian distribution is projected onto the 2D image plane to generate the final rendered image; (6-3) Real-time dynamic update: In dynamic scenes, as new frames are input, the model parameters are continuously updated to achieve real-time reconstruction and rendering of dynamic scenes.
Citation Information
Cited By
Binocular frame generation method and system based on central characteristic flow
CN121033344A
A binocular frame generation method and system based on central feature flow
CN121033344B
Flexible asset reconstruction method and device, electronic equipment, storage medium and product
CN121280734A