Four-dimensional Gaussian video method and device combined with motion layering
Through grouped Gaussian reconstruction and motion hierarchical strategies, combining optical flow and semantic segmentation models, the dynamic and static separation of Gaussian point clouds is optimized, which solves the rendering and storage challenges of long-term videos and complex motion scenes, and realizes high-quality and low-storage real-time rendering and streaming.
Patent Information
- Application Number
- CN202510347720.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
AI Technical Summary
When handling long-term video and complex motion scenarios, the prior art has problems such as limited scalability and surge in storage requirements, making it difficult to achieve high-quality and low-storage real-time rendering and streaming.
Using a reconstruction method based on grouping Gaussian (GOG), the video is divided into multiple frame groups, combined with optical flow and semantic segmentation models for motion hierarchy, and using adaptive static point conversion and progressive frame sampling strategies, the dynamic and static separation of Gaussian point clouds are optimized, and rendered through a lightweight four-dimensional hash grid.
It improves the reconstruction quality of complex motion scenarios, reduces the computing and storage overhead of static points, improves system efficiency, and supports real-time rendering and streaming of long-term videos.
Smart Images

Figure CN120298581A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer graphics and 3D vision, and particularly to a four-dimensional Gaussian video method and apparatus combining motion stratification. Background Art
[0002] With the progress of computer vision and computer graphics, volumetric video, as an emerging media form, has shown broad prospects in multiple application fields such as remote conferencing and holographic classrooms. The core feature of volumetric video lies in its ability to encode complex four-dimensional spatio-temporal scene information, thereby providing users with an immersive experience. Therefore, in the development process of online volumetric video systems, achieving high-quality, real-time rendering while adopting a compact representation form to improve transmission efficiency has become a crucial research direction.
[0003] Among them, the three-dimensional reconstruction technology for dynamic scenes has become an important research direction in the graphics and vision communities. In recent years, neural radiance field technology has been applied, which can render high-quality new visual images, but its training and rendering computational costs are very high, and it is difficult to achieve real-time rendering even with the help of existing compression and acceleration technologies. Three-dimensional Gaussian Splatting (3DGS), as an efficient three-dimensional scene reconstruction solution, combines an implicit motion field and uses a neural network to predict the change of Gaussian point attributes over time, enabling the reconstruction and real-time rendering of dynamic scenes.
[0004] Although existing methods can support real-time rendering and provide relatively high visual quality, they still face the following challenges when dealing with long videos or complex motion scenes: (1) Limited scalability: Existing methods based on implicit neural deformation fields have limited representation capabilities due to the fixed model size, making it difficult to effectively handle long videos and complex motions. (2) Surge in storage requirements: The 3DGS method has high storage requirements because it stores a large amount of explicit attributes for each point. When extending 3DGS to four dimensions, the required storage will further increase, thus bringing more serious storage challenges to the representation of dynamic scenes. Therefore, achieving high-quality, low-storage, and streamable volumetric video remains a major challenge. Summary of the Invention
[0005] The objective of the present invention is to address the deficiencies in the prior art and provide a four-dimensional Gaussian video method and apparatus combining motion stratification. This method has scalability and compactness, can achieve high-quality dynamic scene reconstruction, real-time free-viewpoint rendering, and supports streaming.
[0006] The objective of the present invention is achieved through the following technical solutions: In the first aspect, the present invention also provides a four-dimensional Gaussian video method combining motion stratification, which includes the following steps:
[0007] (1) Reconstruction based on Grouped Gaussian (GOG): The input video is divided into multiple groups of video frames in chronological order, and each group is reconstructed using a GOG. This helps reduce storage and computational complexity, provides scalability, and supports the processing of long-duration volumetric videos. The GOG structure helps decompose complex scene motions into simple piecewise functions, thereby reducing the complexity of motion modeling.
[0008] (2) Motion stratification of Gaussian points: Two-dimensional motion labels are generated through an optical flow model and a semantic segmentation model, and the two-dimensional motion labels are lifted to a Gaussian point cloud for static and dynamic division. At the same time, an adaptive static point conversion technique is used to convert some static points into dynamic points during the GOG optimization process to model the appearance changes of static objects.
[0009] Furthermore, the deformation of static points can be omitted, and static points are shared among multiple GOGs, thereby reducing computational and storage costs. In the present invention, two-dimensional motion labels corresponding to the input video frames are obtained through an optical flow and a semantic segmentation model, the two-dimensional motion labels are lifted to three-dimensional Gaussian points using differentiable rendering, and an adaptive static point conversion method is combined to convert some static points into dynamic points for optimization during GOG reconstruction.
[0010] Furthermore, the specific steps of the GOG-based reconstruction are as follows:
[0011] (1.1) Based on the results of the motion stratification of Gaussian points, fix the static point parameters during the GOG reconstruction process;
[0012] (1.2) Use a lightweight four-dimensional hash grid to perform multiple estimations of the motion offset of the Gaussian point cloud for each GOG respectively;
[0013] (1.3) Design an as-rigid-as-possible (ARAP) loss function and a temporal smoothing loss function to integrate the physical characteristics of real-world motions;
[0014] (1.4) Use a progressive frame sampling strategy to sample video frames from near to far within the time range covered by one GOG.
[0015] Furthermore, the specific steps of the motion stratification of Gaussian points are as follows:
[0016] (2.1) Use an optical flow model to identify the motion pixels in the input video frames, and fuse semantic information to identify the moving objects to obtain a two-dimensional motion label map;
[0017] (2.2) Lift the two-dimensional motion label map to the Gaussian point cloud by adding motion label parameters to the Gaussian point cloud;
[0018] (2.3) By monitoring the position gradient of static points in screen space, some static points are converted into dynamic points to participate in the optimization of Gaussian point parameters and motion, and the appearance changes of static objects are modeled.
[0019] Furthermore, the motion labels of pixels in the input video frame are initialized through an optical flow model. Dynamic labels are set for pixels with optical flow vectors greater than a set threshold, and static labels are set for other pixels. Since the optical flow model usually marks the slowly moving parts in a moving object as static, the present invention enhances the consistency of motion labels of different pixels on the same object through a semantic segmentation model, which can improve the quality of subsequent motion modeling. Then, the present invention binds the motion labels to Gaussian points, and uses the Gaussian splash differentiable renderer to optimize the Gaussian point motion labels with the two-dimensional video frame motion labels as supervision, thereby efficiently using the two-dimensional motion labels to perform dynamic and static layering of Gaussian points in three-dimensional space.
[0020] Furthermore, in the static reconstruction of the Gaussian point cloud, for the Gaussian point cloud, a series of attributes representing the position, rotation, scale, opacity, and color of the Gaussian point stored in each Gaussian point are optimized; in the dynamic reconstruction of the Gaussian point cloud, the dynamic Gaussian point attributes in the standard space and the implicit motion field are jointly optimized; the implicit motion field predicts the motion offsets of the Gaussian point position and rotation over time through a neural network and applies them to the dynamic Gaussian points in the standard space to obtain the corresponding Gaussian point cloud at each moment for optimization and rendering.
[0021] Furthermore, in the GOG reconstruction stage, a progressive frame sampling strategy is adopted. During the motion offset optimization process, video frames close to the key frames are sampled in the early training stage to pre-train the motion field, and then video frames far from the key frames are gradually sampled. This strategy effectively avoids the difficulty of predicting long-distance motion in the initial stage of optimization and improves the reconstruction and rendering quality of large moving objects by gradually expanding the distance of the predicted motion.
[0022] Furthermore, the present invention proposes an adaptive static point conversion method to further optimize motion layering during the GOG reconstruction process. Since the attributes of static points are locked during the GOG reconstruction optimization process to save computational overhead, the appearance changes (such as shadows and reflections, etc.) on static objects in the scene cannot be well explained. Therefore, the present invention monitors the magnitude of the position gradient of static points in screen space and converts static points with large gradients into dynamic points. The converted points optimize their own attributes as dynamic points and participate in the motion optimization over time, improving the learning ability of the changing appearance of static objects.
[0023] In a second aspect, the present invention also provides a four-dimensional Gaussian video device combined with motion stratification, including a memory and one or more processors. Executable code is stored in the memory. When the processor executes the executable code, the four-dimensional Gaussian video method combined with motion stratification as described above is implemented.
[0024] In a third aspect, the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, the four-dimensional Gaussian video method combined with motion stratification as described above is implemented.
[0025] In a fourth aspect, the present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the four-dimensional Gaussian video method combined with motion stratification as described above is implemented.
[0026] The beneficial effects of the present invention are as follows: By adopting the strategy of Gaussian group (GOG) reconstruction, the reconstruction quality of complex motion is improved, and the reconstruction of long-duration volumetric videos is supported; Through the motion stratification strategy of Gaussian points, the calculation and storage overhead of static points are reduced, and the system efficiency is improved; An adaptive static point conversion method is proposed to improve the appearance representation ability of static objects; A progressive frame sampling strategy is proposed to improve the reconstruction quality of large-amplitude motion.
[0027] In summary, the present invention provides a four-dimensional Gaussian video method and device combined with motion stratification, which can effectively improve the rendering quality and storage efficiency of streaming volumetric videos, and is applicable to various application scenarios such as remote immersive communication and holographic classrooms, and has broad application prospects and commercial value. Description of the Drawings
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art:
[0029] Figure 1 It is the overall algorithm flowchart of the present invention.
[0030] Figure 2 It is the flowchart for generating two-dimensional motion labels in the present invention.
[0031] Figure 3 It is the GOG reconstruction flowchart in the present invention.
[0032] Figure 4 It is the structural schematic diagram of a four-dimensional Gaussian video device combined with motion stratification provided by the present invention. Detailed Embodiments
[0033] The following will further describe the specific embodiments of the present invention in detail with reference to the drawings.
[0034] Figure 1 shows the overall framework of the present invention, namely a four-dimensional Gaussian video method combining motion stratification, which reconstructs a free-viewpoint volume video from multi-view video inputs. First, the input multi-view videos are divided into multiple groups of video frames with a fixed length (usually 30 frames) in chronological order, and each group of video frames is used for multi-stage reconstruction using a Grouped Gaussian (GOG): Gaussian point cloud initialization stage, motion stratification initialization stage, and GOG reconstruction stage. Among them, the first GOG is initialized from the sparse point cloud output by Structure-from-Motion (SfM), and subsequent GOGs are initialized from the point cloud of the previous GOG moving to the current key frame, and fast improvement training is performed to repair the error accumulation brought by the previous GOG. Each contains static Gaussian points dynamic Gaussian points and their motion offset D k . During the reconstruction process, Gaussian point cloud initialization and motion stratification initialization are performed at key frames to obtain the stratification of static Gaussian points and dynamic Gaussian points , and then the parameters of the dynamic Gaussian points in the standard space and the motion offset D k are jointly optimized in the GOG reconstruction stage.
[0035] The following further details each key step in the present invention:
[0036] (1) Adaptive motion stratification: The present invention first obtains two-dimensional motion labels corresponding to the key frames of the multi-view input video through an optical flow model and a semantic segmentation model. The specific process is as Figure 2 shown. Taking adjacent key frames in the corresponding view ( and ) as input to query the optical flow model F raft to obtain multi-view RAFT labels: where v ∈ V, V represents all views, and k i ∈ K, K represents all key frame numbers. Pixels with the magnitude of the optical flow vector greater than the threshold λ raft represent significant motion and are thus assigned dynamic labels, while other pixels are assigned static labels.
[0037] However, it is found in practical applications that usually misses the slower-moving parts of dynamic objects, resulting in inconsistent motion labels assigned to different pixels occupied by the same object. This inconsistency may cause object tearing problems in the subsequent GOG reconstruction process. Therefore, the present invention uses a semantic segmentation model to obtain consistent motion labels at the object level to complete the Figure 2As shown, this step selects a certain perspective v i , and samples the dynamic label area from . The sampled points are used as the query input of the Segment Anything Model 2 (SAM2) segmentation model F sam to obtain the object-level SAM labels corresponding to the sampled points: Then, the memory attention function of SAM2 is used to expand to each perspective Finally, the judgment of motion based on optical flow and the object-level completion are added together at the corresponding perspective to obtain the multi-perspective two-dimensional motion labels at the key frame: In the following description, the subscript k is omitted i for concise expression.
[0038] The present invention adds a motion label variable m to each Gaussian point, and renders m through the Gaussian splatter rasterization differentiable rendering pipeline to obtain at the key frame with the multi-perspective two-dimensional motion label M v as the supervision. By minimizing the loss function L m optimize the motion labels m of all Gaussian points: where V is all perspectives and P is all Gaussian points. The first term requires that the rendered is similar to M v and promotes the generated two-dimensional motion label information to the Gaussian points in the three-dimensional space. In the second term, the subscript nn represents the serial number of the clustering center point corresponding to the i-th point in the KNN clustering. This term requires that the motion labels m of the local Gaussian points in the three-dimensional space are similar.
[0039] Through the above method, the Gaussian points in the three-dimensional space can be divided into dynamic and static according to their motion labels m, so as to accurately identify the dynamic objects in the scene. However, it is found in the experiment that complex lighting effects such as dynamic shadows and reflections usually cause the appearance of static objects to change as well. Since the parameters of static points are locked during the GOG reconstruction process to reduce the computational and storage overhead, the appearance changes of static objects cannot be effectively modeled. Therefore, the present invention proposes an adaptive static point transformation method during the GOG reconstruction process, which transforms some static points into dynamic points, so that their parameters can continue to be optimized and can move with time, enhancing the flexibility and robustness of the motion separation module and improving the overall rendering quality. The adaptive static point transformation is based on the monitoring of the position gradient of static points in the screen space. Since the pixels with larger loss function values in the final rendered image usually backpropagate larger position gradients to the corresponding Gaussian points, the present invention sets the position gradient in the screen space greater than the threshold λ cThe static points with a value of 4e-3 are transformed into dynamic ones to adaptively interpret the appearance changes of static objects.
[0040] (2) GOG reconstruction: During the reconstruction process, jointly optimize the dynamic point cloud in the standard space and its motion offset D k , while the parameters of the static point cloud are fixed and do not participate in the optimization, but monitor the gradient of the static point cloud for adaptive static point transformation. In the following descriptions, the subscript k is omitted to simplify the expression. As Figure 3 shown, the dynamic point P d in the standard space undergoes a deformation network to obtain the motion offset D, including the offsets of the Gaussian point position and rotation Δx i,t , Δq i,t . After passing through the motion offset, the dynamic point cloud at time t is obtained, combined with the static point cloud to form the complete point cloud at time t, and after rasterization, the rendered image is obtained. The entire system is trained under the supervision of a loss function. For the parameterization method of the motion deformation field, the present invention comprehensively considers the expression ability and computational efficiency and selects to use a four-dimensional hash grid (4D Hashgrid), which includes four three-dimensional spatio-temporal hash grids (xyz, xyt, xzt, and yzt). The features extracted from the four hash grids are decoded through a lightweight multi-layer perceptron (MLP) network to obtain the motion offset of the dynamic Gaussian point cloud.
[0041] The reconstruction process of GOG is optimized through the following loss functions: the photometric loss function L photo , the as-rigid-as-possible (ARAP) loss function L arap , and the temporal smoothness loss function L temp . The photometric loss function calculates the difference between the rendered image and the video frame supervision image: where the first term is the per-pixel L1 loss term, is the input video frame image, and the second term is the local similarity loss based on image patches, where λ = 0.2 is set. Since physical motions in the real world usually have similarity in local space, the present invention is inspired by this and applies the physics-based ARAP loss function:
[0042] L arap = |Δq i,t - Δq nn,t | + |Δτ i,t - Δτ nn,t | + |∥x i - x nn ∥ - ∥x i + Δx i,t - x nn - Δxnn,t ∥|
[0043] Among them, represents the translation offset of the Gaussian point, ΔR is the matrix form of Δq, and the subscript nn is the serial number of the KNN nearest neighbor points. L arap The first and second terms in L encourage the offsets of point rotation and translation to be similar to those of neighboring points, and the third term encourages maintaining the distances between local points after deformation. To make the movement of the Gaussian point smooth in time and conform to the physical laws of the objective world, the present invention further designs a temporal smoothing loss function:
[0044] L temp = |(1 - δ)Δx i,t + δΔx i,t′ - Δx i,t+δ(t′-t) | + |(1 - δ)Δq i,t + δΔq i,t′ - Δq i,t+δ(t′-t) |
[0045] Among them, δ ∈ (0, 1) is a random perturbation quantity. L temp encourages the position and rotation of the Gaussian point to change uniformly within the local time window [t, t′], where t′ = t + 2 is set.
[0046] To improve the quality of the implicit motion field in modeling large-scale motions, the present invention designs a progressive frame sampling strategy. During the optimization process of GOG reconstruction, first sample the frames closer to the current starting key frame on the time axis for training, as a warm-up for the motion field neural network, and then gradually expand the time range of frame sampling from near to far. In the first N iterations of the GOG reconstruction training, at the nth iteration, the probability that the video frame corresponding to the time point t is sampled is:
[0047]
[0048] Among them, T is the length of the video frame grouping. In the present invention, N is set to half of the total number of training iterations, and progressive frame sampling is performed in the first half of the iterations, and all frames are sampled uniformly in the remaining iterations.
[0049] The comparative experiments have demonstrated the advantages of the present invention over the prior art solutions. Table 1 shows the comparison of metrics between the present invention and other prior art solutions on the widely used Neu3DV dataset for testing. The metrics for comparison include: Peak Signal-to-Noise Ratio (PSNR) which measures the rendering quality, and a higher value indicates a higher rendering quality; storage overhead, measured in MBtyes; reconstruction time, measured in minutes; and rendering speed, measured in Frames per second (FPS). The data in the table are all from the data published by each technical solution, where "N / A" means that the corresponding data metric is not published for that technical solution. As can be seen from Table 1, compared with the prior art solutions, the present invention simultaneously achieves the highest rendering quality and a smaller storage overhead, and the reconstruction time and rendering speed are also at a relatively leading level.
[0050] Table 1. Quantitative comparison between the present invention and prior art solutions on the Neu3DV dataset
[0051]
[0052] Meanwhile, the ablation experiments have demonstrated the necessity of the core steps in the present invention. Table 2 lists the performance of the system on the Neu3DV dataset after removing the motion stratification step in the present invention, which is used to illustrate the role played by the motion stratification step in the whole system. It can be seen that the point cloud motion stratification can significantly reduce the model size, improve the rendering quality, and shorten the reconstruction time.
[0053] Table 2. Influence of the main steps in the present invention on the whole system
[0054] Test scheme Steps to remove motion layering Complete method of the present invention Storage overhead (MB) 1142 340 Rendering quality (PSNR) 32.28 32.56 Reconstruction time (minutes) 194.5 62.7
[0055] Corresponding to the foregoing embodiment of a four-dimensional Gaussian video method combining motion stratification, the present invention also provides an embodiment of a four-dimensional Gaussian video device combining motion stratification.
[0056] See Figure 4 , a four-dimensional Gaussian video device combining motion stratification provided by an embodiment of the present invention includes a memory and one or more processors. An executable code is stored in the memory. When the processor executes the executable code, it is used to implement a four-dimensional Gaussian video method combining motion stratification in the foregoing embodiment.
[0057] An embodiment of a four-dimensional Gaussian video device combining motion layering provided by the present invention can be applied to any device with data processing capabilities, and such a device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory and running them. At the hardware level, as Figure 4 shown, it is a hardware structure diagram of any device with data processing capabilities where a four-dimensional Gaussian video device combining motion layering provided by the present invention is located. In addition to Figure 4 the processor, memory, network interface, and non-volatile memory shown, generally, according to the actual functions of any device with data processing capabilities where the device in the embodiment is located, other hardware may also be included, which will not be elaborated here.
[0058] For the implementation processes of the functions and roles of each unit in the above device, please refer to the implementation processes of the corresponding steps in the above method for details, which will not be elaborated here.
[0059] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0060] The embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements a four-dimensional Gaussian video method combining motion layering in the above embodiment.
[0061] The computer-readable storage medium may be an internal storage unit of any data processing capable device described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any data processing capable device, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any data processing capable device. The computer-readable storage medium is used to store the computer program and other programs and data required by any data processing capable device, and may also be used to temporarily store the data that has been output or is to be output.
[0062] The present invention also provides a computer program product, including a computer program, which when executed by a processor, implements the four-dimensional Gaussian video method combined with motion layering described above.
[0063] In summary, the present invention provides a four-dimensional Gaussian video method and apparatus combined with motion layering. The above embodiments are used to explain and illustrate the present invention, rather than to limit the present invention. Any modification and change made to the present invention within the spirit and scope of the claims of the present invention fall within the protection scope of the present invention.
Claims
1. A four-dimensional Gaussian video method combined with motion stratification, characterized in that, The method includes the following steps: (1) Reconstruction based on grouped Gaussian (GOG): The input video is divided into multiple groups of video frames in chronological order, and each group is reconstructed using a GOG; (2) Motion stratification of Gaussian points: Two-dimensional motion labels are generated through an optical flow model and a semantic segmentation model, and the two-dimensional motion labels are lifted to the Gaussian point cloud for static and dynamic division. At the same time, an adaptive static point conversion technique is used to convert some static points into dynamics during the GOG optimization process to model the appearance changes of static objects.
2. The four-dimensional Gaussian video method combined with motion stratification according to claim 1, characterized in that: The specific steps of the reconstruction based on GOG are as follows: (1.1) Based on the results of the motion stratification of Gaussian points, the static point parameters are fixed during the GOG reconstruction process; (1.2) A lightweight four-dimensional hash grid is used to perform multiple estimations of the motion offset of the Gaussian point cloud for each GOG respectively; (1.3) Design an as-rigid-as-possible (ARAP) loss function and a temporal smoothing loss function to integrate the physical characteristics of real-world motion; (1.4) Use a progressive frame sampling strategy to sample video frames from near to far within the time range covered by a GOG.
3. The four-dimensional Gaussian video method combined with motion stratification according to claim 1, wherein: The specific steps of the motion stratification of Gaussian points are as follows: (2.1) Use an optical flow model to identify the moving pixels in the input video frame, and fuse semantic information to identify the moving objects to obtain a two-dimensional motion label map; (2.2) Lift the two-dimensional motion label map to the Gaussian point cloud by adding motion label parameters to the Gaussian point cloud; (2.3) By monitoring the position gradient of the static points in the screen space, some static points are converted into dynamic points to participate in the optimization of the Gaussian point parameters and motion, and the appearance changes of static objects are modeled.
4. The four-dimensional Gaussian video method combined with motion stratification according to claim 3, characterized in that: Initialize the motion labels of the pixels in the input video frame through the optical flow model, set dynamic labels for the pixels with optical flow vectors greater than the set threshold, set static labels for other pixels, and bind the motion labels to the Gaussian points. Optimize based on the Gaussian splash differentiable renderer, and use the two-dimensional motion labels to perform static and dynamic stratification of the Gaussian points in three-dimensional space.
5. The four-dimensional Gaussian video method combined with motion stratification according to claim 1, characterized in that: In the static reconstruction of the Gaussian point cloud, for the Gaussian point cloud, a series of attributes representing the position, rotation, scale, opacity, and color of the Gaussian points stored in each Gaussian point are optimized; in the dynamic reconstruction of the Gaussian point cloud, the dynamic Gaussian point attributes in the standard space and the implicit motion field are jointly optimized; the implicit motion field predicts the motion offset of the Gaussian point position and rotation over time through a neural network and applies it to the dynamic Gaussian points in the standard space to obtain the corresponding Gaussian point cloud at each moment for optimization and rendering.
6. The four-dimensional Gaussian video method combined with motion stratification according to claim 1, characterized in that: In the GOG reconstruction stage, a progressive frame sampling strategy is adopted. During the motion offset optimization process, in the early training stage, video frames close to the key frames are sampled first to perform warm-up training on the motion field, and then video frames far from the key frames are gradually sampled.
7. A four-dimensional Gaussian video device combined with motion layering, comprising a memory and one or more processors, wherein executable code is stored in the memory, characterized in that, When the processor executes the executable code, it implements a four-dimensional Gaussian video method combined with motion stratification as described in any one of claims 1-6.
8. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements a four-dimensional Gaussian video method combined with motion stratification as described in any one of claims 1-6.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements a four-dimensional Gaussian video method combined with motion layering as described in any one of claims 1-6.
Citation Information
Cited By
Dynamic three-dimensional scene reconstruction method and system based on optical flow and multi-layer Gaussian representation
CN120953520A
Method and system for dynamic 3d scene reconstruction based on optical flow and multi-layered gaussian representation
CN120953520B
Dynamic scene Gaussian splashing method based on space-time motion distillation
CN122312913A