A Gaussian splashing method for dynamic scenes based on spatiotemporal motion distillation
By introducing learnable motion features and a spatiotemporal motion distillation mechanism, the computational complexity and resource consumption issues in dynamic scene rendering are resolved, improving rendering quality and efficiency and achieving high-quality dynamic rendering effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI UNIV
- Filing Date
- 2026-04-07
- Publication Date
- 2026-06-30
AI Technical Summary
Existing technologies suffer from high computational complexity, high resource consumption, weak generalization ability, and inconsistencies in occlusion and geometry in dynamic scene rendering, making it difficult to achieve high-quality, real-time dynamic rendering.
By introducing learnable motion features and a spatiotemporal motion distillation mechanism, a two-stage optimization strategy is adopted to enhance the spatiotemporal representation capability of Gaussian points, improve the consistency and stability of global motion, and constrain Gaussian points using the spatiotemporal motion distillation mechanism.
It effectively reduces model complexity, improves rendering quality and efficiency, reduces artifacts, enhances the ability to express local details in dynamic scenes, and supports real-time rendering.
Smart Images

Figure CN122312913A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of image synthesis, computer graphics and computer vision, and specifically to a dynamic scene Gaussian splashing method based on spatiotemporal motion distillation. Background Technology
[0002] Gaussian splashing-based neural rendering, a novel paradigm for 3D content generation that has emerged in recent years, offers an efficient solution for view composition. However, existing Gaussian splashing methods primarily target static scene modeling, facing numerous challenges in dynamic scenes. Due to object motion, occlusion variations, and viewpoint constraints, dynamic scenes are often underconstrained at any given moment, making it difficult to fully observe spatiotemporal geometry and appearance information. This not only affects the spatial distribution of Gaussian points but also weakens the model's ability to model complex motions and non-rigid deformations, resulting in phenomena such as ghosting, structural drift, and texture blurring in view composition. Therefore, achieving high-quality and temporally consistent dynamic rendering under limited observation conditions remains a crucial research challenge.
[0003] With the improvement of dynamic perception capabilities and the continuous evolution of visual acquisition devices, scene modeling and rendering for time-varying scenarios has gradually become a research hotspot in the fields of computer vision and computer graphics. Especially in application scenarios such as video analysis, augmented reality, intelligent robots, and digital human construction, the system not only needs to acquire accurate spatial geometric information, but also needs to model the motion process and deformation details of the target, so as to achieve the coordinated expression of spatial and temporal information.
[0004] Although existing methods have made some progress in dynamic scene reconstruction, there are still many key problems in how to efficiently and stably generate new perspective results from multi-frame image or video data in real complex environments. These problems are mainly reflected in the following aspects: (1) High computational complexity. Dynamic scene modeling usually relies on continuous frames or multi-view data to recover temporally consistent 3D structures. When processing high-resolution data, long-term sequences, or multi-camera inputs, the computational overhead increases significantly, and the efficiency of training and inference is limited, making it difficult to meet the needs of real-time or large-scale applications. (2) High resource consumption pressure. In the process of dynamic modeling, temporal data needs to be stored and updated. Especially in the four-dimensional representation framework, the consumption of memory and video memory is significantly higher than that of static scenes. For long sequences or complex dynamic scenes, problems such as insufficient resources, data overflow, or rendering failure are likely to occur. (3) Weak generalization ability. Different dynamic scenes have significant differences in motion amplitude and change form. For example, large-scale non-rigid deformation, high-speed motion, and subtle changes coexist. Existing methods often rely on specific assumptions or scene priors, making it difficult to maintain stable performance in diverse scenes. The adaptability and transferability of the model are still insufficient. (4) Occlusion and geometric inconsistency are prominent issues. During dynamic processes, complex occlusion can easily occur between objects or within objects themselves, leading to difficulties in cross-view matching and inaccurate depth estimation. At the same time, errors gradually accumulate over long time sequences, easily causing geometric structure shifts, thus affecting the accuracy and spatiotemporal consistency of the overall reconstruction results. To address these challenges, it is urgent to construct an efficient, robust, and scalable dynamic scene rendering method that can not only adapt to the processing of large-scale image sequences but also possess strong geometric modeling and motion representation capabilities to support complete, continuous, and detailed rendering of dynamic objects.
[0005] In recent years, Gaussian sputtering technology has achieved great success in the field of real-time rendering of dynamic scenes. Inspired by this, some researchers have applied Gaussian sputtering technology to the problem of dynamic scene rendering and have made certain research progress. The relevant research papers are: [1] "4D Gaussian Splatting for Real-Time Dynamic SceneRendering", [2] "Real-time Photorealistic Dynamic Scene Representation and Rendering with 4D Gaussian Splatting", [3] "Per-Gaussian Embedding-Based Deformation for Deformable 3D Gaussian Splatting", [4] "SpatioGS: Spatiotemporal-Aware Density Control for Dynamic Scene Rendering with Gaussian Splatting". Among them, [1] was published at CVPR2024, [2] at ICLR2024, [3] at ECCV2024, and [4] at TVC2026. The key point of these dynamic scene rendering methods is to extend the representation of static scenes from the spatial dimension to the spatiotemporal dimension, so as to represent dynamic scenes. Although these methods have made some progress, they are limited to Gaussian points themselves and ignore the influence of the underconstrained characteristics of discrete Gaussian points in the scene during global motion.
[0006] Chinese patent application CN119991973A discloses a dynamic scene reconstruction method and apparatus based on a multi-scale Gaussian sphere. This technical solution only models Gaussian points in the spatiotemporal dimension without distillation or compression, which may lead to problems such as low efficiency.
[0007] Chinese patent application CN120298581A discloses a four-dimensional Gaussian video method and device that combines motion layering. Although the technical solution groups Gaussian points in different states, it still does not have a corresponding distillation operation. It has efficiency improvement at the training level, but the final output still contains a large number of useless Gaussian points.
[0008] Chinese patent application CN120355851A discloses a dynamic scene reconstruction method and system based on spatial decomposition and Gaussian splashing. The technical solution extracts the initial point cloud from multi-view images and converts it into Gaussian units. If the input image has poor quality, lacks texture, or has insufficient view coverage, the generated initial point cloud will be sparse or contain noise, which directly affects the representation quality of subsequent Gaussian units.
[0009] Therefore, when applied to dynamic scenes with complex changes, existing technologies still face the following challenges: (1) artifacts are easily generated in areas where objects move quickly or are partially occluded, affecting the rendering quality; (2) the geometry and appearance modeling of dynamic areas is not refined enough, resulting in blurred local structures; (3) some methods have high computational burden and storage overhead, making it difficult to meet the needs of real-time applications. Summary of the Invention
[0010] Purpose of the Invention: The purpose of this invention is to address the shortcomings of existing technologies and provide a Gaussian splashing method for dynamic scenes based on spatiotemporal motion distillation. First, this invention introduces learnable motion features to enhance the spatiotemporal representation of Gaussian points in dynamic scenes. Second, it utilizes a spatiotemporal motion distillation mechanism to enhance the consistency and stability of global motion. Finally, through a two-stage optimization strategy, it further improves the training efficiency of the model while ensuring rendering quality, making the local details of the dynamic scene more refined and the structure clearer, thereby improving both rendering quality and computational efficiency. This effectively overcomes the shortcomings of existing dynamic scene rendering methods and can promote the development of fields such as autonomous driving, virtual reality, and augmented reality.
[0011] Technical solution: This invention provides a dynamic scene Gaussian splashing method based on spatiotemporal motion distillation, comprising the following steps:
[0012] Step S1: Input multi-view image data and camera parameters Combine multi-view image data, camera parameters, and timestamp information into tuples. Each tuple corresponds to a training sample, including image data, camera pose, and time.
[0013] in, Indicates time The next shot Zhang Image Number of camera angles This represents the number of timestamps, corresponding to the number of video frames. For the camera intrinsic parameter matrix, Let be a rotation matrix. It is a translation vector;
[0014] Step S2: For the initial image data tuple Using motion inference structure methods to generate scenes in time The initial point cloud below Then, through the initial point cloud Perform Gaussian property initialization to obtain a set of Gaussian points. ;
[0015] in, Indicates the first The three-dimensional coordinates of the points For the number of point clouds, The Gaussian point scaling matrix, For color coefficients, For opacity, Let be a rotation matrix. Indicates the first One Gaussian point;
[0016] Step S3: Combine the Gaussian point set obtained in step S2. The input is a differentiable rasterization pipeline to generate the rendered image. Update the set using backpropagation. The various attributes are controlled using an adaptive density mechanism to determine the number of Gaussian points. ;
[0017] This process is iterated several times to finally generate a standard space Gaussian point set for the dynamic scene. ;
[0018] Step S4: Based on the standard spatial Gaussian point set generated in Step S3, the dynamic scene is reconstructed through a deformation field network. At the same time, learnable motion feature vectors are introduced into the Gaussian point attributes, and the global spatiotemporal consistency of Gaussian points in the neighborhood is constrained by a spatiotemporal motion distillation mechanism. Finally, the PLY file and model weights are output.
[0019] Step S5: Construct the overall loss term as follows:
[0020] ;
[0021] ;
[0022] in, For image reconstruction loss, For the total variational loss, For entropy loss; Let the opacity be the value at the i-th Gaussian point;
[0023] Step S6: The overall loss term obtained through step S5. Perform back gradient propagation and update the set. Related attributes, spatiotemporal structure encoder Weights and multi-head deformable decoders The weights;
[0024] Step S7: After several iterations of training, output the final set of Gaussian points. and spatiotemporal structure encoder Weights and multi-head deformable decoders The weights;
[0025] The final rendered image is then generated through a differentiable raster rendering pipeline.
[0026] Furthermore, the detailed method of step 4 is as follows:
[0027] Step S4.1: The deformation field network includes a spatiotemporal structure encoder and a multi-head deformation decoder. The spatiotemporal structure encoder and the multi-head deformation decoder sequentially perform deformation processing on the attribute information of Gaussian points in the standard spatial Gaussian point set at different times.
[0028] First, the spatiotemporal structure encoder Contains a multi-resolution plane and a compact multilayer perceptron The expression is as follows:
[0029] ;
[0030] in, Indicates the first Multi-resolution plane of layers, This represents a multilayer perceptron. It is a combination of two-dimensional planes;
[0031] Spatiotemporal structure encoder pass Gaussian attributes are distributed and encoded across different two-dimensional planes, while temporal information is incorporated into the corresponding two-dimensional planes. Features are then extracted from the multiple two-dimensional planes using bilinear interpolation and multiplied. The results from each layer are then fused, and finally processed using a multilayer perceptron. Mapped to the final spatiotemporal features The expression is:
[0032] ;
[0033] in, This represents the bilinear interpolation used to query the voxel features located at the four vertices of the grid;
[0034] Next, after all the 3D Gaussian features have been encoded by the spatiotemporal structure encoder, a multi-head deformation decoder is used. To calculate any time The Gaussian property deformation value is calculated using independent multilayer perceptrons; Rotational deformation Scale deformation Opacity distortion and color distortion ;
[0035] Gaussian is obtained from the above deformation value. The corresponding time is the first The properties of each Gaussian point are:
[0036] ;
[0037] in, Indicates the first The coordinates of a Gaussian point Indicates the first A rotation of Gaussian points, Indicates the first A scale of Gaussian points Indicates the first Opacity at Gaussian points Indicates the first The color of a Gaussian point;
[0038] Step S4.2: Model the dynamic behavior of Gaussian points using learnable motion feature vectors, introducing a dimension of for each Gaussian point. Learnable motion feature vectors By combining the spatial location of Gaussian points With motion feature vector This enables the model to distinguish Gaussian points that are spatially adjacent but have different motion behaviors; after introducing motion feature vectors, the deformation field network in step S4.1 is expanded at the input end to make it time-dependent. The following uses the standard spatial coordinates of the Gaussian point Time variables and motion characteristics As a joint input, its mapping relationship is defined as follows:
[0039] ;
[0040] in, The learnable parameters of the deformation field network are represented. Represents the dimension of motion features. This represents the dimension of the Gaussian property deformation value, which is composed of position, rotation, scale, opacity, and color.
[0041] Step S4.3, for the first A Gaussian point, in time The property deformation values are determined by the deformation field network. The prediction is as follows:
[0042] ;
[0043] in, Indicates the first The spatial location of a Gaussian point Indicates the first Motion feature vectors of Gaussian points;
[0044] Construct a set of anchor points based on spatial location and motion feature vectors. And by combining the weighted residual propagation mechanism, the spatiotemporal consistency constraint of Gaussian points is achieved:
[0045] ;
[0046] in, Indicates the first Anchor points, Indicates the number of anchor points;
[0047] In time Below, for non-anchor Gaussian points, their attribute deformation values are obtained by weighting the deformation values of their neighboring anchor points, where for the th A Gaussian point, which in time The transformation values of each attribute are represented as follows:
[0048] ;
[0049] ;
[0050] ;
[0051] ;
[0052] ;
[0053] in, Indicates the anchor index within the neighborhood. Indicates the relationship with the first The set of nearest neighbor anchor points associated with a Gaussian point; , , , , These represent the weighting coefficients for each attribute dimension;
[0054] The weighting coefficients are learnable parameters and are constrained by a normalization function; their unified expression is as follows:
[0055] ;
[0056] in, Indicates the first The Gaussian point and the first Weighting coefficients between anchor points;
[0057] Constructing the nearest neighbor anchor set At the same time, motion feature vectors are introduced. As a constraint, a joint screening is conducted from two dimensions: spatial proximity and motion consistency.
[0058] The above method achieves spatiotemporal co-constraint between Gaussian points, thereby improving the stability of motion modeling in complex dynamic scenes, reducing artifacts, and enhancing the spatiotemporal consistency and detail representation of rendering results.
[0059] Beneficial Effects: This invention employs a spatiotemporal motion distillation mechanism to constrain the global motion of Gaussian points, thereby improving the modeling accuracy and rendering efficiency of dynamic regions. This method introduces learnable motion feature representations into the attributes of Gaussian points, explicitly modeling their spatiotemporal motion information, thus enhancing the spatiotemporal representation capability of Gaussian points in dynamic scenes. A spatiotemporal motion distillation mechanism is proposed, which effectively reduces model complexity while significantly enhancing the consistency and stability of global motion. A two-stage optimization strategy is adopted: the first stage focuses on learning spatiotemporal motion features at the Gaussian point level to fully fit the spatiotemporal changes of the dynamic scene; the second stage introduces the motion distillation mechanism to explicitly apply global spatiotemporal consistency constraints. Finally, a PLY file containing Gaussian point attribute information and model weights are generated, which can then be rendered in real-time through a rasterization rendering pipeline to generate high-quality rendered images.
[0060] Compared with the prior art, the advantages of the present invention are as follows:
[0061] (1) This invention solves the problems of poor spatiotemporal consistency in previous methods by using a dynamic scene Gaussian splashing method and system based on spatiotemporal motion distillation, and achieves high rendering quality and low storage overhead in different dynamic scenes.
[0062] (2) This invention proposes a learnable motion feature, providing a motion representation of Gaussian points that is decoupled from their spatial location, enabling the model to characterize multimodal motion behaviors existing within the same local region. This residual deformation modeling strategy based on motion feature embedding lays a stable and more discriminative dynamic representation foundation for further compression and prediction of subsequent motion information.
[0063] (3) This invention proposes a spatiotemporal motion distillation mechanism, which can effectively constrain the motion consistency of Gaussian points at different times, reduce artifacts caused by the motion of objects in the scene, and improve image rendering quality and rendering efficiency.
[0064] (4) This invention can solve the underfitting problem caused by existing dynamic scene rendering methods when processing sparse view image data, improve the real-time rendering speed after dynamic scene reconstruction, and reduce the memory overhead generated during training, so that the method can run on different devices. Attached Figure Description
[0065] Figure 1 This is an overall flowchart of the method of the present invention.
[0066] Figure 2 These are samples of image data from different viewpoints at the same time in the embodiments.
[0067] Figure 3 The image data is rendered in the example.
[0068] Figure 4 This is a flowchart of the spatiotemporal motion distillation mechanism in the embodiment.
[0069] Figure 5 The rendering results of the technical solution of the present invention in different scenarios are shown.
[0070] Figure 6 This is a comparison of the technical solution of the present invention with other existing technologies. Detailed Implementation
[0071] The technical solution of the present invention will be described in detail below, but the scope of protection of the present invention is not limited to the embodiments described.
[0072] This invention presents a Gaussian splashing method for dynamic scenes based on spatiotemporal motion distillation. It introduces learnable motion feature representations into the attributes of Gaussian points, explicitly modeling their spatiotemporal motion to enhance the model's spatiotemporal representation capability. Subsequently, by distilling the motion information of the Gaussian points, motion anchor points capable of characterizing the global motion pattern of the scene are extracted, reducing model complexity while improving the consistency and stability of global motion. This invention fully utilizes the spatiotemporal characteristics of dynamic objects to efficiently reconstruct and render dynamic scenes, effectively solving the artifact and noise problems that occur during dynamic scene rendering and outputting high-quality rendered images.
[0073] like Figure 1As shown, this invention first inputs multi-view image data and initializes Gaussian attributes with point cloud data; secondly, it obtains a standard space Gaussian point set through pre-training, trains the standard space through dynamic Gaussian sputtering, and enhances the spatiotemporal consistency of Gaussian point motion in the scene through learnable motion feature vectors and spatiotemporal motion distillation mechanism; then, it calculates the overall loss term, and updates the model weights by backpropagation combined with position gradient and density control; finally, it outputs the Gaussian model and deformation field network weights, and outputs multi-view images through a differentiable raster rendering pipeline. The spatiotemporal distillation mechanism and motion feature vectors of this invention can effectively solve the problems of rendering artifacts, low efficiency, and structural blur caused by fast-moving objects in the scene.
[0074] The dynamic scene Gaussian splashing method based on spatiotemporal motion distillation in this embodiment includes the following steps:
[0075] Step S1: Input multi-view image data and camera parameters Combine multi-view image data, camera parameters, and timestamp information into tuples. Each tuple corresponds to a training sample, including image data, camera pose, and time.
[0076] in, Indicates time The next shot Zhang Image Number of camera angles This represents the number of timestamps, corresponding to the number of video frames. For the camera intrinsic parameter matrix, Let be a rotation matrix. It is a translation vector;
[0077] Step S2: For the initial image data tuple Using motion inference structure methods to generate scenes in time The initial point cloud below Then, through the initial point cloud Perform Gaussian property initialization to obtain a set of Gaussian points. ;
[0078] in, Indicates the first The three-dimensional coordinates of the points For the number of point clouds, The Gaussian point scaling matrix, For color coefficients, For opacity, Let be a rotation matrix. Indicates the first One Gaussian point;
[0079] Step S3: Combine the Gaussian point set obtained in step S2. The input is a differentiable rasterization pipeline to generate the rendered image. Update the set using backpropagation. The various attributes are controlled using an adaptive density mechanism to determine the number of Gaussian points. ;
[0080] This process is iterated several times to finally generate a standard space Gaussian point set for the dynamic scene;
[0081] Step S4: Based on the standard spatial Gaussian point set generated in step S3, the dynamic scene is reconstructed through a deformation field network. At the same time, learnable motion feature vectors are introduced into the Gaussian point attributes, and the global spatiotemporal consistency of Gaussian points in the neighborhood is constrained by a spatiotemporal motion distillation mechanism.
[0082] Step S5: Construct the overall loss term as follows:
[0083] ;
[0084] in, For image reconstruction loss, For the total variational loss, For entropy loss;
[0085] Step S6: The overall loss term obtained through step S5. Perform back gradient propagation and update the set. Related attributes, spatiotemporal structure encoder Weights and multi-head deformable decoders The weights;
[0086] Step S7: After several iterations of training, output the final set of Gaussian points. and spatiotemporal structure encoder Weights and multi-head deformable decoders The weights;
[0087] The final rendered image is then generated through a differentiable raster rendering pipeline.
[0088] The detailed method for step 4 of this embodiment is as follows:
[0089] Step S4.1: The deformation field network includes a spatiotemporal structure encoder and a multi-head deformation decoder. The spatiotemporal structure encoder and the multi-head deformation decoder sequentially perform deformation processing on the attribute information of Gaussian points in the standard spatial Gaussian point set at different times.
[0090] First, the spatiotemporal structure encoder Contains a multi-resolution plane and a compact multilayer perceptron The expression is as follows:
[0091] ;
[0092] in, Indicates the first Multi-resolution plane of layers, This represents a multilayer perceptron. It is a combination of two-dimensional planes;
[0093] Spatiotemporal structure encoder pass Gaussian attributes are distributed and encoded across different two-dimensional planes, while temporal information is incorporated into the corresponding two-dimensional planes. Features are then extracted from the multiple two-dimensional planes using bilinear interpolation and multiplied. The results from each layer are then fused, and finally processed using a multilayer perceptron. Mapped to the final spatiotemporal features The expression is:
[0094] ;
[0095] in, This represents the bilinear interpolation used to query the voxel features located at the four vertices of the grid;
[0096] Next, after all the 3D Gaussian features have been encoded by the spatiotemporal structure encoder, a multi-head deformation decoder is used. To calculate any time The Gaussian property deformation value is calculated using independent multilayer perceptrons; Rotational deformation Scale deformation Opacity distortion and color distortion ;
[0097] Gaussian is obtained from the above deformation value. The corresponding time is the first The properties of each Gaussian point are:
[0098] ;
[0099] in, Indicates the first The coordinates of a Gaussian point Indicates the first A rotation of Gaussian points, Indicates the first A scale of Gaussian points Indicates the first Opacity at Gaussian points Indicates the first The color of a Gaussian point;
[0100] Step S4.2: Model the dynamic behavior of Gaussian points using learnable motion feature vectors, introducing a dimension of for each Gaussian point. Learnable motion feature vectors By combining the spatial location of Gaussian points With motion feature vector This enables the model to distinguish Gaussian points that are spatially adjacent but have different motion behaviors; after introducing motion feature vectors, the deformation field network in step S4.1 is expanded at the input end to make it time-dependent. The following uses the standard spatial coordinates of the Gaussian point Time variables and motion characteristics As a joint input, its mapping relationship is defined as follows:
[0101] ;
[0102] in, The learnable parameters of the deformation field network are represented. Represents the dimension of motion features. This represents the dimension of the Gaussian property deformation value, which is composed of position, rotation, scale, opacity, and color.
[0103] Step S4.3, for the first A Gaussian point, in time The property deformation values are determined by the deformation field network. The prediction is as follows:
[0104] ;
[0105] in, Indicates the first The spatial location of a Gaussian point Indicates the first Motion feature vectors of Gaussian points;
[0106] Construct a set of anchor points based on spatial location and motion feature vectors. And by combining the weighted residual propagation mechanism, the spatiotemporal consistency constraint of Gaussian points is achieved:
[0107] ;
[0108] in, Indicates the first Anchor points, Indicates the number of anchor points;
[0109] In time Below, for non-anchor Gaussian points, their attribute deformation values are obtained by weighting the deformation values of their neighboring anchor points, where for the th A Gaussian point, which in time The transformation values of each attribute are represented as follows:
[0110] ;
[0111] ;
[0112] ;
[0113] ;
[0114] ;
[0115] in, Indicates the anchor index within the neighborhood. Indicates the relationship with the first The set of nearest neighbor anchor points associated with a Gaussian point; , , , , These represent the weighting coefficients for each attribute dimension;
[0116] The weighting coefficients are learnable parameters and are constrained by a normalization function; their unified expression is as follows:
[0117] ;
[0118] in, Indicates the first The Gaussian point and the first Weighting coefficients between anchor points;
[0119] Constructing the nearest neighbor anchor set At the same time, motion feature vectors are introduced. As a constraint, a joint screening is conducted from two dimensions: spatial proximity and motion consistency.
[0120] Example 1:
[0121] In this embodiment, the input image data sample is as follows: Figure 2 As shown, the rendered image in this embodiment is as follows: Figure 3 As shown, the rendered image data has a high degree of pixel consistency with the real scene.
[0122] In this embodiment, the spatiotemporal motion of discrete Gaussian points in the scene is constrained by a spatiotemporal motion distillation mechanism, as follows: Figure 4 As shown. The final rendering results in different scenarios are as follows. Figure 5 As shown. The comparison results with other existing methods are as follows. Figure 6 As shown, the method provided by this invention can output high-quality rendered images.
[0123] As can be seen from the above embodiments, this invention first takes a set of images from different perspectives and at different times as input, then outputs a set of sparse point clouds through a motion inference structure. The sparse point clouds are used to initialize Gaussian point attributes, and the initialized Gaussian point set undergoes a fixed number of warm-up training iterations to obtain a standard spatial set that can effectively represent the geometric structure of the scene. In subsequent training, learnable motion feature representations are introduced into the Gaussian point attributes to explicitly model the spatiotemporal motion of the Gaussian points, thereby enhancing the model's spatiotemporal representation capability. Furthermore, by distilling the motion information of the Gaussian points, motion anchor points that can characterize the global motion pattern of the scene are extracted, reducing model complexity while improving the consistency and stability of global motion. During iteration, the density distribution of Gaussian points is continuously updated through an adaptive density control mechanism. Finally, the corresponding Gaussian model and deformation field weights are output for subsequent rendering evaluation. This invention fully utilizes the spatiotemporal characteristics of dynamic objects to efficiently reconstruct and render dynamic scenes, effectively solving the artifact and noise problems that occur during dynamic scene rendering, and thus outputting high-quality rendered images.
Claims
1. A dynamic scene Gaussian splashing method based on spatiotemporal motion distillation, characterized in that, Includes the following steps: Step S1: Input multi-view image data and camera parameters ; Multi-view image data, camera parameters, and timestamp information are combined into tuples. Each tuple corresponds to a training sample, including image data, camera pose, and time. in, Indicates time The next shot Zhang Image Number of camera angles This represents the number of timestamps, corresponding to the number of video frames. For the camera intrinsic parameter matrix, Let be a rotation matrix. It is a translation vector; Step S2: For the initial image data tuple Using motion inference structure methods to generate scenes in time The initial point cloud below Then, through the initial point cloud Perform Gaussian property initialization to obtain a set of Gaussian points. ; in, Indicates the first The three-dimensional coordinates of the points For the number of point clouds, The Gaussian point scaling matrix, For color coefficients, For opacity, Let be a rotation matrix. Indicates the first One Gaussian point; Step S3: Combine the Gaussian point set obtained in step S2. The input is a differentiable rasterization pipeline to generate the rendered image. Update the set using backpropagation. The various attributes are controlled using an adaptive density mechanism to determine the number of Gaussian points. ; This process is iterated several times to finally generate a standard space Gaussian point set for the dynamic scene. ; Step S4: Based on the standard spatial Gaussian point set generated in step S3, the dynamic scene is reconstructed through a deformation field network. At the same time, learnable motion feature vectors are introduced into the Gaussian point attributes, and the global spatiotemporal consistency of Gaussian points in the neighborhood is constrained by a spatiotemporal motion distillation mechanism. Step S5: Construct the overall loss term as follows: ; ; in, For image reconstruction loss, For the total variational loss, For entropy loss, This represents the opacity of the i-th Gaussian point; Step S6: The overall loss term obtained through step S5. Perform back gradient propagation and update the set. Related attributes, spatiotemporal structure encoder Weights and multi-head deformable decoders The weights; Step S7: After several iterations of training, output the final set of Gaussian points. and spatiotemporal structure encoder Weights and multi-head deformable decoders The weights; The final rendered image is then generated through a differentiable raster rendering pipeline.
2. The dynamic scene Gaussian splashing method based on spatiotemporal motion distillation according to claim 1, characterized in that, The detailed method for step 4 is as follows: Step S4.1: The deformation field network includes a spatiotemporal structure encoder and a multi-head deformation decoder. The spatiotemporal structure encoder and the multi-head deformation decoder sequentially perform deformation processing on the attribute information of Gaussian points in the standard spatial Gaussian point set at different times. First, the spatiotemporal structure encoder Contains a multi-resolution plane and a compact multilayer perceptron The expression is as follows: ; in, Indicates the first Multi-resolution plane of layers, This represents a multilayer perceptron. It is a combination of two-dimensional planes; i and j represent coordinates on different two-dimensional planes; Spatiotemporal structure encoder pass Gaussian attributes are distributed and encoded across different two-dimensional planes, while temporal information is incorporated into the corresponding two-dimensional planes. Features are then extracted from the multiple two-dimensional planes using bilinear interpolation and multiplied. The results from each layer are then fused, and finally processed using a multilayer perceptron. Mapped to the final spatiotemporal features The expression is: ; in, This represents the bilinear interpolation used to query the voxel features located at the four vertices of the grid; Next, after all the 3D Gaussian features have been encoded by the spatiotemporal structure encoder, a multi-head deformation decoder is used. To calculate any time The Gaussian property deformation value is calculated using independent multilayer perceptrons; Rotational deformation Scale deformation Opacity distortion and color distortion ; Gaussian is obtained from the above deformation value. The corresponding time is the first The properties of each Gaussian point are: ; in, Indicates the first The coordinates of a Gaussian point Indicates the first A rotation of Gaussian points, Indicates the first A scale of Gaussian points Indicates the first Opacity at Gaussian points Indicates the first The color of a Gaussian point; , , The property transformation value of the Gaussian point set; Step S4.2: Model the dynamic behavior of Gaussian points using learnable motion feature vectors, introducing a dimension of for each Gaussian point. Learnable motion feature vectors By combining the spatial location of Gaussian points With motion feature vector This enables the model to distinguish Gaussian points that are spatially adjacent but have different motion behaviors; after introducing motion feature vectors, the deformation field network in step S4.1 is expanded at the input end to make it time-dependent. The following uses the standard spatial coordinates of the Gaussian point Time variables and motion characteristics As a joint input, its mapping relationship is defined as follows: ; in, The learnable parameters of the deformation field network are represented. Represents the dimension of motion features. This represents the dimension of the Gaussian property deformation value, which is composed of position, rotation, scale, opacity, and color. Step S4.3, for the first A Gaussian point, in time The property deformation values are determined by the deformation field network. The prediction is as follows: ; in, Indicates the first The spatial location of a Gaussian point Indicates the first Motion feature vectors of Gaussian points; Construct a set of anchor points based on spatial location and motion feature vectors. And by combining the weighted residual propagation mechanism, the spatiotemporal consistency constraint of Gaussian points is achieved: ; in, Indicates the first Anchor points, Indicates the number of anchor points; In time Below, for non-anchor Gaussian points, their attribute deformation values are obtained by weighting the deformation values of their neighboring anchor points, where for the th A Gaussian point, which in time The transformation values of each attribute are represented as follows: ; ; ; ; ; in, Indicates the anchor index within the neighborhood. Indicates the relationship with the first The set of nearest neighbor anchor points associated with a Gaussian point; , , , , These represent the weighting coefficients for each attribute dimension; The weighting coefficients are learnable parameters and are constrained by a normalization function; their unified expression is as follows: ; in, Indicates the first The Gaussian point and the first Weighting coefficients between anchor points; Constructing the nearest neighbor anchor set At the same time, motion feature vectors are introduced. As a constraint, a joint screening is conducted from two dimensions: spatial proximity and motion consistency.
Citation Information
Patent Citations
Dynamic scene reconstruction method and device based on multi-scale Gaussian sphere
CN119991973A
Four-dimensional Gaussian video method and device combined with motion layering
CN120298581A
Dynamic scene reconstruction method and system based on spatial decomposition and Gaussian splashing
CN120355851A