A method for surging Gaussian rendering structure for dynamic scene modeling

By using the surging Gaussian rendering structure method, a dynamic Gaussian scene model is constructed using single-frame point clouds and multi-frame images. This solves the problems of insufficient dynamic modeling capabilities and high cost dependence in existing technologies, and realizes low-cost and efficient dynamic 3D modeling and rendering, which is suitable for scenarios such as autonomous driving and virtual reality.

CN121095411BActive Publication Date: 2026-04-03CHINA UNIV OF MINING & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing Gaussian rendering methods are mainly designed for static scenes and lack the ability to model dynamic changes over continuous time. They rely on global LiDAR point clouds or multi-frame images, resulting in high system costs and weak generalization ability, making it difficult to apply on low-cost platforms. Furthermore, insufficient temporal modeling leads to inconsistent inter-frame rendering, motion blur, and loss of details.

Method used

We adopt the surge-flow Gaussian rendering structure method, construct an initial Gaussian field model by using a single-frame point cloud and multiple frames of images, and combine a multi-source joint encoder and a dual-path cooperative decoding mechanism to introduce inter-frame constraint loss and incremental motion consistency loss, thereby realizing incremental dynamic modeling and image reconstruction of Gaussian elements.

Benefits of technology

It enables high-precision dynamic 3D modeling at low cost, reduces reliance on high-cost hardware, improves modeling efficiency and rendering quality, and is suitable for dynamic 3D modeling scenarios such as autonomous driving and virtual reality. It also has good online inference capabilities and hardware compatibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121095411B_ABST
    Figure CN121095411B_ABST
Patent Text Reader

Abstract

This invention discloses a surge-flow Gaussian rendering structure method for dynamic scene modeling. The method first establishes a Gaussian primitive set based on the initial 3D global point cloud image, outputting a 3D static Gaussian scene model. Then, combining the input consecutive adjacent frame images, a multi-source joint encoder is established and run to obtain Gaussian primitive feature encodings. Next, a dual-path collaborative decoding mechanism of feature enhancement and time-frequency domain decomposition, along with a training strategy for dynamic scene reconstruction, is constructed. The trained model is then used for forward inference modeling and image rendering of the dynamic scene, outputting the final result. This invention introduces an adaptive Gaussian modeling mechanism based on spatiotemporal location, combined with a multimodal encoder and a time-frequency joint prediction structure, to achieve self-evolutionary modeling of dynamic Gaussian primitives in continuous time. Dynamic modeling and image reconstruction are completed using only a single frame point cloud and multiple frames of images, improving modeling efficiency, scene adaptability, and rendering quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and 3D computer vision technology, and in particular to a surge-flow Gaussian rendering structure method for dynamic scene modeling. Background Technology

[0002] In recent years, with the rapid development of technologies such as artificial intelligence, virtual reality, intelligent manufacturing, and autonomous driving, the demand for 3D visual modeling and image rendering technologies in industries such as industry, transportation, healthcare, and entertainment has been increasing. Especially in applications such as digital twins, immersive simulations, and virtual scene reconstruction, systems not only need to perform high-fidelity modeling of static environments but also need to have the ability to accurately depict dynamic objects and changing environments. Therefore, constructing dynamic 3D modeling methods with spatiotemporal consistency and high rendering efficiency has become an important research direction in the field of computer vision and graphics.

[0003] Currently, significant progress has been made in 3D reconstruction methods based on neural rendering, among which 3D Gaussian rendering has attracted widespread attention due to its advantages such as strong differentiability and high rendering efficiency. This method achieves joint modeling of scene geometry and appearance by establishing a Gaussian distributed particle field in 3D space. However, existing Gaussian rendering methods are mainly designed for static scenes and lack the ability to model dynamic changes over continuous time, making it difficult to meet the comprehensive needs of dynamic modeling, temporal representation, and real-time rendering in real-world scenes. In addition, some methods rely on global LiDAR point clouds or multi-frame images for initialization and supervision, resulting in high system costs, weak generalization ability, and difficult deployment.

[0004] While existing methods have achieved some success in static scene synthesis and a small number of dynamic modeling tasks, they still face the following challenges when dealing with multi-objective and multi-scale changes over continuous time:

[0005] 1. Insufficient dynamic modeling capability: Most existing methods are based on static modeling mechanisms, which are difficult to adapt to the continuous changes of dynamic objects. They lack the ability to independently model and incrementally update targets within the scene, resulting in insufficient modeling accuracy and serious computational redundancy.

[0006] 2. Strong Initialization Dependency: Some neural rendering methods rely on global point clouds or multi-frame labels as initialization conditions, which limits their versatility and deployment capability under low sensor dependence and weak supervision data. Especially in real-world applications, many methods rely heavily on multi-frame dense LiDAR point clouds or high-quality global geometric priors to obtain high-precision 3D modeling results. This not only increases the dependence on sensor hardware but also significantly increases the system deployment cost, limiting the promotion and application of the algorithm on low- and medium-cost platforms. For example, in urban-level autonomous driving or lightweight intelligent terminal scenarios, it is usually impossible to equip multi-line LiDAR or high-frequency data acquisition equipment, and only limited perception resources such as sparse point clouds or a small number of images can be obtained.

[0007] 3. Insufficient temporal modeling: Traditional unified modeling methods cannot effectively capture the heterogeneous temporal evolution of Gaussian meta-attributes, resulting in problems such as inconsistent inter-frame rendering, motion blur, and loss of details.

[0008] Therefore, how to achieve high-precision, low-cost dynamic 3D modeling using only initial sparse point cloud or city-level mapping data has become a key issue in promoting the "hardware-software decoupling" of this technology and facilitating its widespread application. A novel method that can balance dynamic modeling accuracy, temporal consistency, and computational efficiency is urgently needed to reduce dependence on high-cost hardware and improve applicability in resource-constrained scenarios. Summary of the Invention

[0009] This invention aims to address at least one of the technical problems existing in the prior art. To this end, this invention proposes a surge-flow Gaussian rendering structure method for dynamic scene modeling, which can complete dynamic modeling and image reconstruction using only a single-frame point cloud and multiple frames of images, thereby improving modeling efficiency, scene adaptability, and rendering quality.

[0010] A surge-type Gaussian rendering structure method for dynamic scene modeling according to an embodiment of the present invention includes the following steps:

[0011] S1. Construct the initial Gaussian field model: Input the 3D global point cloud map at the initial moment, establish a Gaussian element set based on each sampling point in the 3D global point cloud map data, and output a 3D static Gaussian scene model;

[0012] S2. Establish a multi-source joint encoder to extract Gaussian field spatiotemporal features: Based on the three-dimensional static Gaussian scene model constructed in step S1 and the input consecutive adjacent frame images, establish and run a multi-source joint encoder to obtain Gaussian meta-feature encoding.

[0013] S3. Constructing a dual-path cooperative decoding mechanism: A dual-path cooperative decoding mechanism is constructed based on feature enhancement and time-frequency domain decomposition;

[0014] S4. Constructing a training strategy for dynamic scene reconstruction: Utilizing a dual loss constraint mechanism of inter-frame constraint loss and incremental motion consistency loss, high-fidelity dynamic reconstruction is achieved through multi-objective collaborative training.

[0015] S5. Dynamic Inference Rendering and Output: Utilize the trained model to perform forward inference modeling and image rendering of dynamic scenes, and output the final result.

[0016] According to embodiments of the present invention, a surge-flow Gaussian rendering structure method for dynamic scene modeling is proposed. This method constructs an incremental dynamic modeling framework for Gaussian elements based on an initial 3D global point cloud map, enabling continuous evolution modeling and autonomous updating of Gaussian elements over time. Simultaneously, a multi-source joint encoder is introduced to represent multi-source information in a high-dimensional manner. A dual-path collaborative decoding mechanism, utilizing feature enhancement and time-frequency joint prediction, is employed to adaptively model the changing trends of different particle attributes during the modeling process, thereby improving the model's adaptability to moving targets, complex lighting, and multimodal inputs. Unlike traditional one-time modeling strategies, this invention adopts a particle-driven modeling approach with on-demand updates, updating only key Gaussian elements in each frame, significantly reducing redundant computation and improving modeling efficiency. Furthermore, to ensure the continuity and physical consistency of the dynamic modeling process, this invention introduces inter-frame consistency loss and incremental motion constraint mechanisms. Through cross-frame prediction and error backpropagation, the model is globally adjusted, further enhancing the structural stability and detail fidelity of the rendered image. In summary, the present invention can be widely applied to dynamic 3D modeling scenarios such as autonomous driving simulation, digital twins, and virtual reality. It effectively alleviates the problems of high computing power consumption and data redundancy in traditional methods, reduces the system's dependence on continuous high-precision point clouds or high-frequency data, and has good online inference capabilities and hardware adaptability. It is particularly suitable for resource-constrained platforms or actual deployment scenarios with high real-time requirements.

[0017] In some embodiments of the present invention, the multi-source joint encoder in step S2 includes a spatiotemporal adaptive resolution encoder and an axial compression cross-modal encoder. The spatiotemporal adaptive resolution encoder is used to dynamically adjust the feature resolution, while the axial compression cross-modal encoder fuses point clouds, images, and temporal motion features through axial decomposition and dimensionality reduction.

[0018] In some embodiments of the present invention, step S2 specifically includes the following steps:

[0019] S21. Construct a spatiotemporal adaptive resolution encoder; assign a normalized timestamp to each Gaussian primitive of the 3D static Gaussian scene model output in step S1, decompose it into orthogonal planes, extract features through bilinear interpolation, and output the encoded cross-plane spatiotemporal fusion features.

[0020] S22. Construct an axial compression cross-modal encoder, taking the initial three-dimensional global point cloud map and consecutive adjacent frame images as input, extracting features through an independent CNN backbone network, and outputting the processed data using an axial compression strategy.

[0021] S23. The feature vectors output by the spatiotemporal adaptive resolution encoder and the axial compression cross-modal encoder are fused after pooling to generate a unified Gaussian meta-feature code.

[0022] In some embodiments of the present invention, step S3 specifically includes the following steps:

[0023] S31. Construct a multimodal feature-driven incremental prediction decoder, perform incremental prediction on the Gaussian meta-feature code output in step S2, and output a deformed four-dimensional Gaussian model.

[0024] S32. Construct a dual-domain adaptive dynamic incremental decoder, which processes the data through time-domain and frequency-domain branches, and outputs updated motion incremental parameters through an adaptive weighting mechanism.

[0025] In some embodiments of the present invention, the increment output of the multimodal feature-driven incremental prediction decoder is used for parameter fitting of the dual-domain adaptive dynamic incremental decoder, and the output of the dual-domain adaptive dynamic incremental decoder is used as a priori correction of the output of the multimodal feature-driven incremental prediction decoder.

[0026] In some embodiments of the present invention, step S4 specifically includes the following steps:

[0027] S41. Construct inter-frame constraint loss by comparing the differences between directly rendered frames and predicted rendered frames to establish temporal consistency constraints.

[0028] S42. Construct incremental motion consistency loss and apply balance displacement constraints and direction consistency constraints to the motion increment parameters output in step S3.

[0029] S43. Jointly optimize the dual loss constraint mechanism, combining the loss functions in steps S41 and S42 to optimize the scene representation.

[0030] In some embodiments of the present invention, the dynamic three-dimensional reconstruction result output in step S5 includes continuous image frames and a dynamic three-dimensional model. Attached Figure Description

[0031] Figure 1 This is a flowchart illustrating the surge-type Gaussian rendering structure method for dynamic scene modeling according to the present invention.

[0032] Figure 2 This is the overall operational framework of the present invention;

[0033] Figure 3 This is an overall flowchart of the axial compression cross-modal feature encoder in this invention;

[0034] Figure 4 This is an overall flowchart of the present invention;

[0035] Figure 5 This is a graph showing the qualitative comparison results of this invention on the Waymo dataset;

[0036] Figure 6 This is a time visualization experiment diagram of the present invention on the KITTI dataset. Detailed Implementation

[0037] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0038] To facilitate understanding, before introducing the embodiments of this disclosure, several terms involved in the embodiments of this disclosure will be explained as follows:

[0039] STAR: Spatio Temporal Adaptive Resolution Encoder, is a feature extraction module that dynamically adjusts computational resources. It can adaptively adjust the feature resolution according to the spatiotemporal complexity of the scene, optimizing computational efficiency while maintaining high accuracy.

[0040] AXMEC: Axial Modality-Embedded Coder, is a lightweight cross-modal feature fusion module that reduces computational complexity through axial decomposition while effectively integrating multi-source data such as point clouds, images, and temporal motion.

[0041] MLP Decoder: Multi-layer Perceptron Decoder is a feature decoding architecture based on a fully connected neural network. It maps high-dimensional latent space features to target outputs (such as 3D Gaussian parameters, RGB color, depth, etc.) by stacking multiple linear layers and non-linear activation functions.

[0042] T-FAD, Time-Frequency Adaptive Decoder for Dynamic Gaussian Units, is a dual-domain adaptive dynamic incremental decoder.

[0043] To achieve efficient modeling and continuous rendering of complex dynamic 3D scenes, this invention discloses a dynamic visual modeling framework based on the Gaussian particle evolution mechanism, the overall operational framework of which is as follows: Figure 1 As shown. The entire framework includes a Gaussian modeling module, a feature encoding module, an attribute prediction module, and a rendering module, which are connected sequentially.

[0044] The functions and roles of each module are as follows:

[0045] The Gaussian modeling module can generate a set of Gaussian elements based on the initial 3D global point cloud map and output a 3D static Gaussian scene model to represent the static structures and potential dynamic targets in the scene.

[0046] The feature encoding module includes a spatiotemporal adaptive resolution feature encoder STAR and an axial compression cross-modal feature encoder AXMEC, which are responsible for extracting the spatiotemporal location features and cross-modal appearance features of each Gaussian element, respectively, and fusing them to generate a unified Gaussian element feature code.

[0047] The attribute prediction module receives the fused encoded features and uses an incremental prediction decoder and a dual-domain adaptive dynamic incremental decoder to perform incremental predictions of the position, shape, velocity and color of Gaussian elements over time, thereby achieving continuous evolution at the particle level.

[0048] The rendering module is used for rendering. It takes the attribute parameters predicted by the attribute prediction module as input, and uses a differentiable Gaussian rasterization method based on transparency mixing to complete high-fidelity image synthesis and output a continuous frame image sequence.

[0049] Meanwhile, the entire framework also includes a training supervision module, which enhances the model's capabilities in temporal modeling, structural representation, and detail rendering by introducing inter-frame consistency loss, motion consistency loss, deep supervision, and structural similarity loss.

[0050] Specifically, in practical applications, the input data includes a 3D global point cloud map at the initial moment and a sequence of images from several consecutive frames. First, an initial static Gaussian scene model is constructed, and particle features are generated guided by spatiotemporal coordinates and image semantics. In each subsequent frame, the model automatically updates particle attributes based on the Gaussian unit state of the previous frame and the prediction increment, achieving particle-level self-evolution. When new targets appear in the scene or the structure changes drastically, the system generates new Gaussian units through strategic point sampling to supplement the modeling density, thereby adapting to rapidly changing dynamic scenes.

[0051] Through this architecture, the present invention achieves the goal of completing continuous dynamic scene modeling and high-quality rendering without global point cloud at all times and without frame-by-frame reconstruction, which greatly reduces modeling costs, improves modeling speed, and has good temporal generalization ability and multi-task adaptability.

[0052] The following is for reference. Figures 1-6 A method for modeling dynamic scenes using a surge-flow Gaussian rendering structure according to embodiments of the present invention is described, comprising the following steps:

[0053] S1. Construct the initial Gaussian field model: Input the 3D global point cloud map at the initial moment, establish a Gaussian element set based on each sampling point in the 3D global point cloud map data, and output a 3D static Gaussian scene model;

[0054] S2. Establish a multi-source joint encoder to extract Gaussian field spatiotemporal features: Based on the three-dimensional static Gaussian scene model constructed in step S1 and the input consecutive adjacent frame images, establish and run a multi-source joint encoder to obtain Gaussian meta-feature encoding.

[0055] S3. Constructing a dual-path cooperative decoding mechanism: A dual-path cooperative decoding mechanism is constructed based on feature enhancement and time-frequency domain decomposition;

[0056] S4. Constructing a training strategy for dynamic scene reconstruction: Utilizing a dual loss constraint mechanism of inter-frame constraint loss and incremental motion consistency loss, high-fidelity dynamic reconstruction is achieved through multi-objective collaborative training.

[0057] S5. Dynamic Inference Rendering and Output of 3D Dynamic Reconstruction Results: Utilizing the trained model, forward inference modeling and image rendering are performed on the dynamic scene, outputting the final result. The output may include consecutive image frames and the dynamic 3D reconstruction results.

[0058] It is understandable that this invention, by introducing an incremental evolution mechanism of Gaussian particles into dynamic 3D modeling tasks, uses the initial 3D global point cloud map as a basis and combines multi-source information from a multi-source joint encoder to achieve dynamic perception and attribute updates of Gaussian elements on a continuous time axis from multiple perspectives. This effectively improves inter-frame consistency and rendering fidelity in the 3D modeling process and is widely applicable to scenarios such as autonomous driving simulation, virtual reality, and robot perception. Simultaneously, this invention employs a dual-path collaborative decoding mechanism, introducing feature enhancement and time-frequency domain decomposition in particle attribute modeling. This allows for a balance between smooth changes and high-frequency details in modeling the evolution of particle attributes over time, effectively improving the model's ability to fit changes such as target motion trajectories. Compared to a single modeling strategy, this method significantly improves modeling efficiency and temporal generalization ability while maintaining rendering quality. Furthermore, this invention introduces inter-frame image consistency loss and incremental motion consistency loss during training. By jointly supervising the predicted images and actual rendered images between consecutive frames, it strengthens the physical rationality and temporal stability of the model. The proposed loss function does not depend on additional labels and is designed solely based on the structural information of the image and the model itself, thereby enhancing its transferability and adaptability to weak supervision.

[0059] Overall, through the coordination of each step, this invention has the following advantages in dynamic scene reconstruction tasks: it can achieve high-quality dynamic modeling without relying on global dense point clouds; it can achieve incremental updates of Gaussian particles frame by frame, reducing redundant calculations; it has the characteristics of good structural continuity, high rendering quality, and strong spatiotemporal generalization ability, and is suitable for a variety of complex dynamic 3D modeling tasks, and has good scalability and engineering deployment prospects.

[0060] It should be noted that the above-mentioned 3D global point cloud map can be represented using world coordinates. In practice, the initial 3D global point cloud map can be obtained through a LiDAR system, and each sampling point in the map data can be converted into a Gaussian unit with geometric and appearance attributes.

[0061] Each Gaussian element is represented as {μ,Σ,α,c}, where μ is the coordinate of the Gaussian center point; Σ is the covariance matrix, Σ=L·L T L is a matrix composed of a rotation matrix and a scaling matrix, L = R·S, where R is the rotation matrix describing the direction of the Gaussian distribution, and S is the scaling matrix defining the range of the Gaussian distribution along each principal axis; α is the opacity, α∈R; and c is the color, c∈R. 3 (k+1) 2 The color c is defined by the coefficients of the spherical harmonic function (SH), and k represents the order of the SH function.

[0062] Gaussian units can be embedded into a 3D scene, and the Gaussian function G(x) can be used to define the coordinates of any point x∈R in the 3D global point cloud map. 3Modeling of illumination density:

[0063]

[0064] In some embodiments of the present invention, the multi-source joint encoder in step S2 includes a spatiotemporal adaptive resolution encoder STAR and an axial compression transmodal encoder AXMEC. The spatiotemporal adaptive resolution encoder STAR is used to dynamically adjust the feature resolution, while the axial compression transmodal encoder AXMEC fuses point clouds, images, and temporal motion features through axial decomposition and dimensionality reduction.

[0065] In some embodiments of the present invention, step S2 may include the following steps:

[0066] S21. Construct STAR (short for Spatiotemporal Adaptive Resolution Encoder, hereinafter the same); assign a normalized timestamp to each Gaussian primitive of the 3D static Gaussian scene model output in step S1, decompose it into orthogonal planes, extract features through bilinear interpolation, and output the encoded cross-plane spatiotemporal fusion features.

[0067] Specifically, the initial 3D static Gaussian scene model is first used as the static background. Simultaneously, potential moving point clouds are randomly sampled and supplemented in the dynamic sensitive region to form a hybrid initialization field with the static background. Each Gaussian element is assigned a normalized timestamp t∈[0,1]. The four-dimensional spatiotemporal coordinates of each Gaussian element in the four-dimensional spatiotemporal field are (x,y,z,t), where (x,y,z) are the corresponding 3D coordinates, and t is mapped to high-dimensional features through Fourier position encoding to improve temporal discriminability.

[0068] γ(t)=[sin(2 0 πt),cos(2 0 πt),...,sin(2 L-1 πt),cos(2 L-1 πt)] (2);

[0069] In the formula, L is the encoding frequency, which captures the high / low frequency components of motion.

[0070] Next, the obtained four-dimensional spatiotemporal field is decomposed into six sets of orthogonal point feature encoding planes, and efficient modeling is achieved through low-rank constraints. Spatial static plane S = {P} xy ,P xz ,P yz} is used to encode spatially geometrically invariant structures (such as rigid constraints in building facades); while the dynamic plane T = {P} xt ,P yt ,P zt} is used to encode the motion trajectory of a dynamic target along each axis, applicable to any four-dimensional point (x, y, z, t). Feature vectors {f} of the six planes are extracted using bilinear interpolation.xy ,f xz ,f yz ,f xt ,f yt ,f zt}, and apply a gated attention mechanism for cross-plane feature fusion:

[0071]

[0072] Where, α i ,β j For learnable weight matrix, This is an element-wise multiplication.

[0073] The above F 4D This refers to the cross-plane spatiotemporal fusion feature encoded by the spatiotemporal adaptive resolution encoder output.

[0074] S22. Construct AXMEC (short for Axial Compression Cross-Modal Encoder, hereinafter the same), taking the initial three-dimensional global point cloud map and consecutive adjacent frame images as input, extracting features through an independent CNN backbone network, and outputting them after processing with an axial compression strategy.

[0075] Specifically, such as Figure 3 As shown, the initial 3D global point cloud map and any consecutive adjacent frame images I t-1 ,I t As input, feature maps F are obtained by extracting features through an independent CNN backbone network. lidar ∈R H×W and consecutive adjacent frame image feature maps Using an axial compression strategy, the feature map is first compressed into three groups of feature strips along both the horizontal and vertical directions. The horizontal feature strips of the point cloud are as follows: The vertical feature bands of consecutive adjacent frames are Then, axial convolution and 1×1 convolution are used to restore the compressed strips into complete enhanced feature maps. and

[0076] Enhance feature maps and First, features are learned independently through a self-attention mechanism to enhance the spatiotemporal consistency within each modality. Then, a cross-attention mechanism is used to realize feature interaction between different modalities, capturing the correlation between dynamic objects in point clouds and images, ultimately forming a coupled feature representation F. coupled .

[0077] To accurately characterize the progressive evolution of motion properties, a lightweight UNeT branch is introduced to couple the feature representation F. coupledBy refining dynamic features through an encoder-decoder structure, preserving global contextual information, and suppressing high-frequency noise, the temporal smoothness of incremental prediction is improved, ultimately outputting F. AXMEC .

[0078] S23. The feature vectors output by STAR and AXMEC are pooled and then fused to generate a unified Gaussian meta-feature code.

[0079] Specifically, the cross-plane spatiotemporal fusion feature F encoded by STAR will be used. 4D and AXMEC output F AXMEC Pooling and concatenation are performed to generate a unified Gaussian meta-feature code F:

[0080]

[0081] It is understood that the multi-source joint encoder constructed in this invention consists of STAR and AXMEC. The former generates particle embedding representations with dynamic semantics based on the spatiotemporal location input of Gaussian units, while the latter utilizes multiple frames of images and the initial point cloud to construct structural detail-aware features, thereby comprehensively capturing spatiotemporal changes and texture details in dynamic scenes. This feature fusion approach effectively enhances the model's ability to model complex motion structures and multi-view inputs.

[0082] In some embodiments of the present invention, the construction of the dual-path cooperative decoding mechanism in step S3 may specifically include:

[0083] S31. Construct a multimodal feature-driven incremental prediction decoder, perform incremental prediction on the Gaussian meta-feature code output in step S2, and output a deformed four-dimensional Gaussian model.

[0084] Specifically, using the Gaussian meta-encoder F as input, incremental prediction is performed using three parallel MLP branches in the MLP decoder (short for Multimodal Feature-Driven Incremental Prediction Decoder, hereinafter the same). The three branches are merged to output a four-dimensional Gaussian model representing the incremental representation. The MLP decoder architecture employs three parallel MLP branches, namely: geometric decoder D... g Speed ​​Decoder D v and texture decoder D c Geometric Decoder D g Used to predict position increment Δμ and deformation increment Δ∑ (covariance matrix update); velocity decoder D v Used to predict velocity increment Δv, driving temporal motion; texture decoder D c It is used to predict the color increment Δc, focusing on the features of the current frame image; the three branches are merged to finally output a deformed four-dimensional Gaussian model.

[0085] S32. Construct a dual-domain adaptive dynamic incremental decoder, which processes the data through time-domain and frequency-domain branches, and outputs updated motion incremental parameters through an adaptive weighting mechanism.

[0086] Specifically, the Gaussian feature code F is used as input, and the constructed T-FAD (short for dual-domain adaptive dynamic incremental decoder, the same below) is used to perform dual-branch processing in the time domain and frequency domain, and the fitting is combined to form a unified model; then an adaptive weighting mechanism is introduced, and the updated motion increment parameters are output after the dual-branch weighted summation.

[0087] For example, in the time-domain branch, polynomial functions can be used to describe the motion of Gaussian units; in the frequency-domain branch, Fourier series can be used to describe the motion; then, the polynomial fitting in the time-domain branch and the Fourier series fitting in the frequency-domain branch are combined to form a unified model to better describe complex motion trajectories. Here, the time-varying attribute residual D(t) at each time t is expressed as:

[0088] D(t)=P N (t)+F L (t) (5);

[0089] Among them, P N (t) is an Nth-order polynomial whose coefficients are given by Composition; F L (t) is an L-order Fourier series, whose coefficients are given by Composition; Corresponding to:

[0090]

[0091] Next, an adaptive weighting mechanism is introduced, which dynamically adjusts the weights of the polynomial fitting and Fourier series fitting based on the smoothness of the motion. For smooth motion, the weight of the polynomial fitting is increased; for fast motion, the weight of the Fourier series fitting is increased. Therefore, a weight function w(t) is defined to dynamically adjust the polynomial fitting P. N (t) and Fourier series fitting F L The contribution of (t):

[0092]

[0093] The weighting function w(t) can be calculated based on the smoothness of the velocity:

[0094]

[0095] It should be noted that the increment of the multimodal feature-driven incremental predictive decoder output is used for parameter fitting of the dual-domain adaptive dynamic incremental decoder, and the output of the dual-domain adaptive dynamic incremental decoder is used as a priori to correct the output of the multimodal feature-driven incremental predictive decoder.

[0096] Understandably, this invention introduces a dual-branch structure of multinomial time modeling and Fourier frequency modeling in particle attribute modeling, combined with an adaptive weighting mechanism. This balances smooth changes and high-frequency details in modeling the evolution of particle attributes over time, effectively improving the model's ability to fit target motion trajectories and color changes. Compared to a single modeling strategy, this method significantly improves modeling efficiency and temporal generalization ability while maintaining rendering quality.

[0097] To ensure the continuity and physical consistency of the dynamic modeling process, in some embodiments of the present invention, step S4, constructing the training strategy for dynamic scene reconstruction, may include:

[0098] S41. Construct inter-frame constraint loss. By comparing the differences between directly rendered frames and predicted rendered frames, establish temporal consistency constraints.

[0099] Specifically, I of consecutive adjacent frames at any given time t-1 ,I t As input, the inter-frame constraint loss establishes temporal consistency constraints by comparing the differences between directly rendered frames and predicted rendered frames. It can include three parts, which respectively constrain color consistency, depth consistency, and normal vector consistency.

[0100] At any time t k There are two methods to render the image: the first is to directly use timing t k Gaussian model rendering image The second type is based on the previous time t. k-1 Given a Gaussian model, velocity increment Δv, and time interval Δt, predict the Gaussian model at time t. k The state at any given moment is used to generate the final rendered image. The inter-frame constraint loss formula is defined as follows:

[0101]

[0102] Among them, L1 1 Loss manifestation and Differences, L1 2 Loss manifestation With real images Differences, L1 3 Loss manifestation and The differences; λ1, λ2, and λ3 are weighting coefficients used to balance the contributions of the three parts of the absolute error loss function.

[0103] S42. Construct incremental motion consistency loss and apply balance displacement constraints and direction consistency constraints to the motion increment parameters output in step S3.

[0104] Specifically, the motion increment parameters output from the S3 decoder are used as input to apply equilibrium displacement constraints and orientation consistency constraints. The incremental motion consistency loss L... motion as follows:

[0105] L motion =λ4·|Δμ-(v+Δv)·Δt|+λ5·(1-cos(Δμ,(v+Δv)·Δt)) (11);

[0106] Where λ4 is the weighting coefficient for the equilibrium displacement loss, λ5 is the weighting coefficient for the directional consistency loss, Δμ is the displacement increment, Δv is the velocity increment, and Δt is the time interval.

[0107] S43. Jointly optimize the dual loss constraint mechanism, combining the loss functions in steps S41 and S42 to optimize the scene representation.

[0108] It is understandable that there are five loss functions in steps S41 and S42. These loss functions are used to optimize the scene, and the overall loss function is expressed as:

[0109] L=λ i L inter +λ m L motion +λ d L depth +λ s L sky +λ ssim L ssim (12);

[0110] Among them, L inter For inter-frame constraint loss; L motion For incremental motion consistency loss; L depth For depth loss, it can be calculated by comparing the absolute error between the depth rendered by the Gaussian model and the sparse depth generated by projecting LiDAR points onto the camera plane; L sky For the binary cross-entropy loss in the sky region; L ssim λ is a structural similarity loss function used to evaluate the similarity between rendered and real images. i , λ m , λ d , λ s , λ ssimThese are the weighting coefficients for the contributions of each loss function.

[0111] Understandably, this invention introduces inter-frame consistency loss and incremental motion constraint mechanisms during training, and globally adjusts the model through cross-frame prediction and error backpropagation, further improving the structural stability and detail fidelity of the rendered images. The proposed loss function does not depend on additional labels, but is designed solely based on the structural information of the image and the model itself, enhancing transferability and weak supervision adaptation capabilities.

[0112] After obtaining the particle attributes at the current moment, a Gaussian rasterization method based on a transparency blending mechanism can be used to project all Gaussian primitives onto the image plane for rendering, resulting in the rendered image of the current frame. This process is based on the Gaussian primitive attributes predicted in step S3, including parameters such as position, shape, color, and opacity. To ensure the continuity and consistency of the rendering results, the network parameters optimized by the training strategy constructed in step S4 are used to regulate the temporal evolution of particle attributes.

[0113] The entire inference and rendering process can be completed automatically with each frame of image input. It can update attributes based on the particle state of the previous frame and the features of the current image, achieving particle-level incremental modeling. Finally, it can continuously output high-fidelity images of each frame, along with the corresponding Gaussian particle set for each frame, for the reconstruction and visualization updates of dynamic 3D scenes, achieving particle-level incremental modeling.

[0114] Example 1:

[0115] In this embodiment, to verify the effectiveness of the surge-type Gaussian rendering structure method for dynamic scene modeling proposed in this invention, a complete experimental system was constructed to perform functional verification and performance evaluation of this invention.

[0116] In this embodiment, the system implements an incremental modeling structure based on Gaussian particles, with the overall framework built upon 3D Gaussian rendering (3DGS). Model training employs the Adam optimizer, performing 30,000 iterations on a single RTX 4090 GPU.

[0117] The initial input is the first frame of global 3D point cloud data acquired by the LiDAR system, used to initialize the static Gaussian particle set. Subsequent time frames of Gaussian particles are initialized through random sampling, eliminating the need for dense laser data across multiple consecutive frames and significantly reducing reliance on the sensor. The encoding stage utilizes the spatiotemporal adaptive resolution encoder (STAR) and the axial compression cross-modal encoder (AXMEC) proposed in this invention. The former captures the local evolutionary features of particles in time and space, while the latter fuses information from the point cloud and image modalities to achieve efficient perception of dynamic semantics. The features output by both are then uniformly fused and input into the dual-path cooperative decoder module for attribute prediction.

[0118] In the decoder section, an MLP decoder is first used to make preliminary predictions on the position, shape, velocity, and color attribute increments of Gaussian particles. Then, a T-FAD decoder with a time-frequency joint structure is used to further fit the temporal evolution trend of particle attributes. T-FAD adopts a combination of time-domain polynomial and frequency-domain Fourier series modeling and introduces an adaptive weight control mechanism to enhance the fitting effect, which has good adaptability to smooth motion and rapidly changing scenes.

[0119] During the training phase, two core constraints, inter-frame consistency loss and incremental motion consistency loss, were introduced. In addition, a multi-objective collaborative optimization mechanism was constructed by combining deep supervision loss, sky region loss and structural similarity loss to supervise the model’s accurate modeling of dynamic structures.

[0120] Reference Figure 5 and Figure 6 As shown, the experimental section selected two publicly available datasets, KITTI and Waymo, for validation. Both datasets have a frame rate of 10Hz and contain numerous dynamic targets and complex lighting environments. The proposed method achieved excellent performance in both novel perspective synthesis and image reconstruction tasks, especially when compared to state-of-the-art methods such as EmerNeRF, SUDS, and PVG, demonstrating comprehensive superiority in PSNR, SSIM, and LPIPS metrics. Specifically, in the fast-moving scenes of the KITTI dataset, PSNR was improved by 9.4%, SSIM by 10%, and LPIPS by 56.1%.

[0121] Furthermore, in ablation experiments, the STAR encoder, AXMEC encoder, T-FAM decoder, MLP decoder, adaptive weighting mechanism, and inter-frame loss constraints designed in this invention were all proven to be key factors in performance improvement. In particular, using only the initial LiDAR point cloud as a geometric prior, high-precision modeling capability was maintained in subsequent time frames, further verifying the effectiveness and generalization ability of the "initial-driven + incremental evolution" strategy proposed in this invention.

[0122] This embodiment verifies the feasibility and advancement of the present invention in dynamic scene 3D modeling and image rendering. It can not only achieve continuous time-level particle-level updates and rendering, but also significantly reduce the dependence on dense priors of multiple frames while maintaining high-fidelity visual quality, and has good potential for practical application.

[0123] In summary, this invention combines the concept of dynamic particle modeling with the attribute prediction capabilities of neural networks. By introducing an adaptive Gaussian modeling mechanism based on spatiotemporal location, it constructs a fluid-like Gaussian modeling framework. Using Gaussian atoms as particles, through time-driven continuous evolution and the dynamic injection of multimodal features, the entire Gaussian field naturally flows and evolves in spacetime like a fluid, achieving unified modeling of dynamic structures and changing details. Compared to traditional neural rendering methods, this invention completes dynamic modeling and image reconstruction using only a single-frame point cloud and multiple frames of images, without relying on globally dense point clouds or additional label information, significantly improving modeling efficiency, scene adaptability, and rendering quality. Furthermore, this invention not only demonstrates superior modeling quality in dynamic visual scenes such as autonomous driving and virtual reality, but also possesses good versatility and resource adaptability, providing an effective path for promoting the large-scale application of 3D vision technology in low- to medium-cost systems.

[0124] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

Claims

1. A surge-flow Gaussian rendering structure method for dynamic scene modeling, characterized in that, Includes the following steps: Step S1: Construct the initial Gaussian field model: Input the 3D global point cloud map at the initial moment, establish a Gaussian element set based on each sampling point in the 3D global point cloud map data, and output the 3D static Gaussian scene model; Step S2: Establish a multi-source joint encoder to extract Gaussian field spatiotemporal features: Based on the three-dimensional static Gaussian scene model constructed in step S1 and the input consecutive adjacent frame images, establish and run a multi-source joint encoder to obtain Gaussian meta-feature encoding. Step S3: Construct a dual-path cooperative decoding mechanism: Construct a dual-path cooperative decoding mechanism based on feature enhancement and time-frequency domain decomposition; Step S4: Construct a training strategy for dynamic scene reconstruction: Utilize the dual loss constraint mechanism of inter-frame constraint loss and incremental motion consistency loss to achieve high-fidelity dynamic reconstruction through multi-objective collaborative training; Step S5: Dynamic inference rendering and output: Use the trained model to perform forward inference modeling and image rendering of the dynamic scene, and output the final result. Step S2 specifically includes the following steps: Step S21: Construct a spatiotemporal adaptive resolution encoder; assign a normalized timestamp to each Gaussian primitive of the 3D static Gaussian scene model output in step S1, decompose it into orthogonal planes, extract features through bilinear interpolation, and output the encoded cross-plane spatiotemporal fusion features. Step S22: Construct an axial compression cross-modal encoder. The initial three-dimensional global point cloud map and consecutive adjacent frame images are used as input. Features are extracted through an independent CNN backbone network and processed using an axial compression strategy before output. Step S23: The feature vectors output by the spatiotemporal adaptive resolution encoder and the axial compression cross-modal encoder are fused after pooling to generate a unified Gaussian meta-feature code. Step S3 specifically includes the following steps: Step S31: Construct a multimodal feature-driven incremental prediction decoder, perform incremental prediction on the Gaussian meta-feature code output in step S2, and output a deformed four-dimensional Gaussian model. Step S32: Construct a dual-domain adaptive dynamic incremental decoder, which processes the data through time-domain and frequency-domain branches, and outputs updated motion incremental parameters through an adaptive weighting mechanism; Step S4 specifically includes the following steps: Step S41: Construct inter-frame constraint loss by comparing the differences between directly rendered frames and predicted rendered frames to establish temporal consistency constraints. Step S42: Construct incremental motion consistency loss and apply balance displacement constraints and direction consistency constraints to the motion increment parameters output in step S3; Step S43: Jointly optimize the dual loss constraint mechanism, combining the loss functions in steps S41 and S42 to optimize the scene representation.

2. The surge-flow Gaussian rendering structure method for dynamic scene modeling according to claim 1, characterized in that, The multi-source joint encoder in step S2 includes a spatiotemporal adaptive resolution encoder and an axial compression cross-modal encoder. The spatiotemporal adaptive resolution encoder is used to dynamically adjust the feature resolution, while the axial compression cross-modal encoder fuses point clouds, images, and temporal motion features through axial decomposition and dimensionality reduction.

3. The surge-flow Gaussian rendering structure method for dynamic scene modeling according to claim 1, characterized in that, The increment of the multimodal feature-driven incremental prediction decoder is used for parameter fitting of the dual-domain adaptive dynamic incremental decoder, and the output of the dual-domain adaptive dynamic incremental decoder is used as a priori correction of the output of the multimodal feature-driven incremental prediction decoder.

4. The surge-flow Gaussian rendering structure method for dynamic scene modeling according to claim 1, characterized in that, The output of step S5 includes consecutive image frames and dynamic 3D reconstruction results.

Citation Information

Patent Citations

  • AI visual special effect dynamic generation system fused with multi-modal perception

    CN120318379A

  • Three-Dimensional Model Reconstruction and Rendering Device

    US20250200861A1