Dynamic scene image domain adaptive system and method based on four-dimensional gaussian sputtering, computer storage medium
Patent Information
- Application Number
- CN202511755406.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2045-11-26
AI Technical Summary
[0003]然而,现有基于神经辐射场或高斯溅射的方法在处理动态场景图像域适应任务时,存在两个局限性
本申请实施例提供的基于四维高斯溅射的动态场景图像域适应方法,基于四维高斯溅射表示的动态场景图像域适应框架,与传统的基于单张图像特征的迁移方法不同,本发明所提出的方法利用多张图像来捕获更丰富和多样的风格、纹理与光照条件。本发明首先利用扩展将四维高斯分解为条件三维高斯和基于时间分布的一维高斯,此外,设计域嵌入模块使模型能够学习基于目标域图像的条件嵌入,设计域协调机制以确保多视角的视觉一致性。
Smart Images

Figure CN121582073B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and computer graphics, specifically to a dynamic scene image domain adaptation system and method based on four-dimensional Gaussian sputtering, and a computer storage medium. Background Technology
[0002] Image domain adaptation aims to transfer information from a source domain to a target domain while addressing inter-domain discrepancies, and has been extensively studied in 2D vision tasks. Recent advances in Neural Radiance Field (NeRF) and Gaussian Splatting (GS) in 3D scene representation and rendering have provided new opportunities for exploring image domain adaptation in 3D or 4D space. NeRF-based methods render scenes by encoding the volume density and radiance values of the scene using a multi-layer perceptron. Gaussian sputtering uses a combination of 3D Gaussians to model the scene to reconstruct spatial information and achieves new perspective rendering through point sputtering-based rasterization.
[0003] However, existing methods based on neural radiation fields or Gaussian sputtering have two limitations when handling image domain adaptation tasks in dynamic scenes. First, existing methods are mainly designed for static scenes and struggle to effectively handle dynamic objects, thus limiting their applicability in real-world scenarios. Second, existing methods primarily rely on single target images for transformation, leading to overfitting of the model to specific features of that image, resulting in poor generalization ability to other images or scenes. Currently, there is still a lack of technologies that simultaneously address these issues, thus affecting the quality of dynamic scene reconstruction and image domain adaptation. Summary of the Invention
[0004] This invention provides a dynamic scene image domain adaptation system and method based on four-dimensional Gaussian sputtering, as well as a computer storage medium, which can effectively improve the quality of dynamic scene reconstruction and image domain adaptation.
[0005] This invention is achieved through the following technical solution: On one hand, embodiments of this application provide a dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering, including: Acquire dynamic images captured from multiple perspectives at different points in time; Generate a four-dimensional Gaussian model based on dynamic images; The four-dimensional Gaussian model is decomposed into a conditional three-dimensional Gaussian model and a marginal one-dimensional time component. Extract the target domain embedding vector; Based on the extracted target domain embedding vector and the embedding vector of the 3D Gaussian model, an affine transformation is performed on each Gaussian point to map the Gaussian representation onto the distribution of the target domain. Maintain multi-view consistency by predicting the correspondence between different training perspectives.
[0006] In some embodiments, the target domain embedding vector is extracted using a domain multilayer perceptron in the step of extracting the target domain embedding vector.
[0007] In some embodiments, performing an affine transformation on each Gaussian point to map the Gaussian representation onto the distribution of the target domain includes: inputting the target domain embedding vector and the embedding vector of each 3D Gaussian model into a domain embedding multilayer perceptron, performing an affine transformation on each Gaussian point to map the Gaussian representation onto the distribution of the target domain.
[0008] In some embodiments, maintaining multi-view consistency by predicting the correspondence between different training viewpoints includes: projecting the pixels of the first image into a three-dimensional space to obtain three-dimensional space points, and mapping the three-dimensional space points into image B to constrain the mapping depth of image B to be consistent with the rendering depth.
[0009] In some embodiments, maintaining multi-view consistency by predicting the correspondence between different training viewpoints includes: dividing the image region into reliable regions and dynamically changing regions, and for reliable regions, using low-uncertainty regions as the main constraint objects. For dynamically changing regions, dynamic masks are used to identify moving objects or changing regions in the scene, and flexible constraints are applied to the moving objects or changing regions.
[0010] In some embodiments, the loss function for Gaussian model reconstruction is: in, Representative reference image, This represents the affine image; as can be seen, the reconstruction loss function includes two terms: Constraints and constraint, Used to balance the proportions of the two items.
[0011] In some embodiments, the domain coordination loss function is: in, This is a region of low uncertainty. This represents the overlapping area between image A and image B. and These represent the pixels in the overlapping region of image A and image B, respectively. To render depth, The mapping depth.
[0012] On the other hand, embodiments of this application provide a dynamic scene image domain adaptation system based on four-dimensional Gaussian sputtering, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering according to any of the above embodiments.
[0013] This application also provides a computer storage medium storing a computer program, which is executed to implement the dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering in any of the above embodiments.
[0014] Compared with the prior art, the present invention has the following advantages and beneficial effects: The dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering provided in this application, and the dynamic scene image domain adaptation framework based on four-dimensional Gaussian sputtering representation, differs from traditional transfer methods based on single image features. The method proposed in this invention utilizes multiple images to capture richer and more diverse styles, textures, and lighting conditions. This invention first uses an extension to decompose the four-dimensional Gaussian into a conditional three-dimensional Gaussian and a one-dimensional Gaussian based on temporal distribution. Furthermore, a domain embedding module is designed to enable the model to learn conditional embeddings based on the target domain image, and a domain coordination mechanism is designed to ensure visual consistency across multiple viewpoints. Attached Figure Description
[0015] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a flowchart illustrating a dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering provided in some embodiments of the present invention. Figure 2 This is a schematic diagram of different domain adaptation results on an object dataset provided in some embodiments of the present invention; Figure 3 This is a schematic diagram illustrating the adaptation results of different domains on a scene dataset provided in some embodiments of the present invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0018] The terms “comprising” and “having”, and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not limited to the steps or modules listed, but may optionally include steps or modules not listed, or may optionally include other steps or modules inherent to such process, method, product, or device.
[0019] On the one hand, please refer to Figure 1 This application provides a dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering, including the following steps: S10. Acquire dynamic images captured from multiple perspectives at different points in time.
[0020] S20. Generate a four-dimensional Gaussian model based on dynamic images.
[0021] S30. Decompose the four-dimensional Gaussian model into a conditional three-dimensional Gaussian model and a marginal one-dimensional time component.
[0022] In step S30, given dynamic input captured from multiple perspectives at different time points, the four-dimensional Gaussian is first decomposed into a conditional three-dimensional Gaussian and an edge one-dimensional temporal component. Since spatial coordinates and time are independent, this decomposition not only simplifies the spatiotemporal modeling process formally but also provides more flexible operational space for domain adaptation. Subsequently, the relevant processing for domain adaptation is focused on the three-dimensional Gaussian, aligning it with the target domain in appearance, style, and geometry. After domain adaptation is completed, it is combined with the temporal distribution to restore the overall representation of the dynamic scene in four-dimensional space.
[0023] S40. Extract the target domain embedding vector.
[0024] S50. Based on the extracted target domain embedding vector and the embedding vector of the three-dimensional Gaussian model, perform an affine transformation on each Gaussian point to map the Gaussian representation onto the distribution of the target domain.
[0025] S60. Maintain multi-view consistency by predicting the correspondence between different training perspectives.
[0026] In step S60, the core idea of the domain coordination mechanism is to maintain multi-view consistency by predicting the depth correspondence between different training views. Unlike traditional methods that mainly calculate pixel position differences, the depth correspondence-based method emphasizes the consistency of depth values in 3D space. By fully utilizing the spatial relationships between multiple views, the domain coordination mechanism can provide more robust and accurate optimization results than traditional pixel difference constraints, thereby improving the geometric and appearance consistency of scene reconstruction.
[0027] There is no necessary order between steps S40 and S10-S30. Steps S40 and S10-S30 can be performed sequentially or simultaneously.
[0028] In some embodiments, in step S40, extracting the target domain embedding vector, the target domain embedding vector is extracted by a domain multilayer perceptron.
[0029] In the above embodiments, a domain multilayer perceptron can be designed through a domain embedding module to extract the target domain embedding vector, which can capture the overall scene features, style and appearance information of the target domain.
[0030] In some embodiments, step S50, performing an affine transformation on each Gaussian point to map the Gaussian representation onto the distribution of the target domain, includes: inputting the target domain embedding vector and the embedding vector of each 3D Gaussian model into the DE-MLP, performing an affine transformation on each Gaussian point to map the Gaussian representation onto the distribution of the target domain.
[0031] In the above embodiments, the target domain embedding vector and the embedding vector of each 3D Gaussian are simultaneously used as inputs to the domain embedding multilayer perceptron. The domain embedding multilayer perceptron utilizes the joint information of the two types of embeddings to perform an affine transformation on each Gaussian point, mapping its Gaussian representation, which originally belonged to the input domain, to the distribution of the target domain, thereby achieving cross-domain style adjustment and appearance transfer.
[0032] In some embodiments, step S60, which maintains multi-view consistency by predicting the correspondence between different training viewpoints, includes: projecting the pixels of the first image into a three-dimensional space to obtain three-dimensional space points, and mapping the three-dimensional space points into image B to constrain the mapping depth of image B to be consistent with the rendering depth.
[0033] In the above embodiments, the domain coordination mechanism first projects the pixels of image A into a three-dimensional space, and then uses the camera's intrinsic and extrinsic parameter matrices to map these three-dimensional space points onto image B, thereby constraining the mapping depth of image B to maintain consistency with the rendering depth. This method can effectively ensure the geometric consistency of multiple views in static scenes.
[0034] In some other embodiments, S60, the step of maintaining multi-view consistency by predicting the correspondence between different training viewpoints includes: dividing the image region into a reliable region and a dynamically changing region, and using the low uncertainty region as the main constraint object for the reliable region. For dynamically changing regions, dynamic masks are used to identify moving objects or changing regions in the scene, and flexible constraints are applied to the moving objects or changing regions.
[0035] In contrast, the approach of projecting pixels from the first image into 3D space to obtain 3D spatial points, and then mapping these points onto image B to constrain the mapping depth of image B to maintain consistency with the rendering depth, has limitations when handling dynamic regions. To address this, the domain coordination mechanism introduces uncertainty maps and dynamic masks, employing different optimization strategies for reliable and dynamically changing regions. For reliable regions, low-uncertainty regions are used as the primary constraint objects to ensure accurate reconstruction of the geometry and appearance of static parts. For dynamic regions, moving objects or changing areas in the scene are identified through dynamic masks, and flexible constraints are applied to these regions to avoid over-penalizing dynamic objects. This strategy enables the domain coordination mechanism to effectively enhance scene reconstruction accuracy under multi-view and temporally varying conditions.
[0036] In some embodiments, the loss function used in training the above model is mainly mentioned in the following embodiments: The loss function for Gaussian model reconstruction is: in, Representative reference image, This represents the affine image; as can be seen, the reconstruction loss function includes two terms: Constraints and constraint, Used to balance the proportions of the two items.
[0037] The domain coordination loss function is: in, This is a region of low uncertainty. This represents the overlapping area between image A and image B. and These represent the pixels in the overlapping region of image A and image B, respectively. To render depth, The mapping depth.
[0038] In some examples, the domain adaptation loss function includes content loss and style loss, with the content loss defined as follows: in, This is the ReLU4_2 layer of the VGG19 network. Features extracted by VGG19 are used here to measure high-level semantic differences, thus forcing the preservation of content similarity. Similarly, a style loss function is designed based on the mean and standard deviation of the source and target domain features extracted by VGG19. in, For the VGG19 network, the relu1_1, relu2_1, relu3_1, and relu4_1 layers are... and The mean and standard deviation represent the characteristics of the source domain. and The mean and standard deviation represent the characteristics of the source domain.
[0039] On the other hand, embodiments of this application provide a dynamic scene image domain adaptation system based on four-dimensional Gaussian sputtering, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering according to any of the above embodiments.
[0040] This application also provides a computer storage medium storing a computer program, which is executed to implement the dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering in any of the above embodiments.
[0041] Effect demonstration: Figure 2 and Figure 3 The domain adaptation results of this invention on two different datasets are shown. Figure 2 The first row shows the domain adaptation results on the object dataset. The second and third rows show the domain adaptation results guided by different target domains. Figure 3 To achieve domain adaptation results on scene datasets, this invention can effectively learn different target styles, such as painting styles and animation styles, while realizing high-quality dynamic scene reduction.
[0042] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of the invention in any way. Any simple modifications or equivalent changes made to the above embodiments based on the technical essence of the present invention shall fall within the protection scope of the present invention.
Claims
1. A dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering, characterized in that, include: Acquire dynamic images captured from multiple perspectives at different points in time; Generate a four-dimensional Gaussian model based on dynamic images; The four-dimensional Gaussian model is decomposed into a conditional three-dimensional Gaussian model and a marginal one-dimensional time component. Extract the target domain embedding vector; Based on the extracted target domain embedding vector and the embedding vector of the 3D Gaussian model, an affine transformation is performed on each Gaussian point to map the Gaussian representation onto the distribution of the target domain. Maintain multi-view consistency by predicting the correspondence between different training perspectives; Performing an affine transformation on each Gaussian point to map the Gaussian representation onto the distribution of the target domain includes: inputting the target domain embedding vector and the embedding vector of each 3D Gaussian model into a domain embedding multilayer perceptron, performing an affine transformation on each Gaussian point to map the Gaussian representation onto the distribution of the target domain; Maintaining multi-view consistency by predicting the correspondence between different training viewpoints includes: projecting the pixels of the first image into three-dimensional space to obtain three-dimensional space points, and mapping the three-dimensional space points into image B to constrain the mapping depth of image B to be consistent with the rendering depth. Maintaining multi-view consistency by predicting the correspondence between different training perspectives also includes dividing the image region into reliable regions and dynamically changing regions. For reliable regions, low-uncertainty regions are used as the main constraint objects. For dynamically changing regions, dynamic masks are used to identify moving objects or changing regions in the scene, and flexible constraints are applied to the moving objects or changing regions.
2. The dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering as described in claim 1, characterized in that, In the step of extracting the target domain embedding vector, the target domain embedding vector is extracted using a domain multilayer perceptron.
3. The dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering as described in claim 1, characterized in that, The loss function for Gaussian model reconstruction is: in, Representative reference image, Represents the reconstructed image. This represents the affine image; as can be seen, the reconstruction loss function includes two terms: Constraints and constraint, Used to balance the proportions of the two items.
4. The dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering as described in claim 1, characterized in that, The domain coordination loss function is: in, This is a region of low uncertainty. This represents the overlapping area between image A and image B. and These represent the pixels in the overlapping region of image A and image B, respectively. To render depth, The mapping depth.
5. A dynamic scene image domain adaptation system based on four-dimensional Gaussian sputtering, characterized in that, The system includes a memory and a processor, wherein the memory stores a computer program and the processor executes the computer program to implement the dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering as described in any one of claims 1-4.
6. A computer storage medium, characterized in that, It stores a computer program, which is executed to implement the dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering as described in any one of claims 1-4.
Citation Information
Patent Citations
Reconstruction method of three-dimensional reconstruction model based on two-dimensional Gaussian splashing
CN120374867A
Monocular video scene dynamic three-dimensional reconstruction method based on optical flow
CN120747366A