Dynamic scene image domain adaptation system and method based on four-dimensional Gaussian sputtering, and computer storage medium
By generating a four-dimensional Gaussian model and performing decomposition and affine transformation, combined with a domain coordination mechanism, the problems of dynamic object processing and model generalization in dynamic scene image domain adaptation are solved, achieving high-quality dynamic scene reconstruction and image domain adaptation.
Patent Information
- Application Number
- CN202511755406.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2045-11-26
AI Technical Summary
Existing methods based on neural radiation fields or Gaussian sputtering struggle to effectively handle dynamic objects in dynamic scene image domain adaptation tasks, and the models overfit to the features of a single target image, resulting in poor generalization ability in other images or scenes.
By acquiring dynamic images from multiple perspectives at different time points, a four-dimensional Gaussian model is generated and decomposed into a conditional three-dimensional Gaussian model and a one-dimensional time component. The target domain embedding vector is extracted, and affine transformation and domain coordination are performed. Multi-view consistency prediction is used, and a domain embedding module and domain coordination mechanism are designed to ensure multi-view consistency and the quality of dynamic scene reconstruction.
It improves the quality of dynamic scene reconstruction and image domain adaptation, enhances the model's adaptability under different viewpoints and time conditions, and improves the geometric and appearance consistency of scene reconstruction.
Smart Images

Figure CN121582073A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and computer graphics, specifically to a dynamic scene image domain adaptation system and method based on four-dimensional Gaussian sputtering, and a computer storage medium. Background Technology
[0002] Image domain adaptation aims to transfer information from a source domain to a target domain while addressing inter-domain discrepancies, and has been extensively studied in 2D vision tasks. Recent advances in Neural Radiance Field (NeRF) and Gaussian Splatting (GS) in 3D scene representation and rendering have provided new opportunities for exploring image domain adaptation in 3D or 4D space. NeRF-based methods render scenes by encoding the volume density and radiance values of the scene using a multi-layer perceptron. Gaussian sputtering uses a combination of 3D Gaussians to model the scene to reconstruct spatial information and achieves new perspective rendering through point sputtering-based rasterization.
[0003] However, existing methods based on neural radiation fields or Gaussian sputtering have two limitations when handling image domain adaptation tasks in dynamic scenes. First, existing methods are mainly designed for static scenes and struggle to effectively handle dynamic objects, thus limiting their applicability in real-world scenarios. Second, existing methods primarily rely on single target images for transformation, leading to overfitting of the model to specific features of that image, resulting in poor generalization ability to other images or scenes. Currently, there is still a lack of technologies that simultaneously address these issues, thus affecting the quality of dynamic scene reconstruction and image domain adaptation. Summary of the Invention
[0004] This invention provides a dynamic scene image domain adaptation system and method based on four-dimensional Gaussian sputtering, as well as a computer storage medium, which can effectively improve the quality of dynamic scene reconstruction and image domain adaptation.
[0005] This invention is achieved through the following technical solution: On one hand, embodiments of this application provide a dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering, including: Acquire dynamic images captured from multiple perspectives at different points in time; Generate a four-dimensional Gaussian model based on dynamic images; The four-dimensional Gaussian model is decomposed into a conditional three-dimensional Gaussian model and a marginal one-dimensional time component. Extract the target domain embedding vector; Based on the extracted target domain embedding vector and the embedding vector of the 3D Gaussian model, an affine transformation is performed on each Gaussian point to map the Gaussian representation onto the distribution of the target domain. Maintain multi-view consistency by predicting the correspondence between different training perspectives.
[0006] In some embodiments, the target domain embedding vector is extracted using a domain multilayer perceptron in the step of extracting the target domain embedding vector.
[0007] In some embodiments, performing an affine transformation on each Gaussian point to map the Gaussian representation onto the distribution of the target domain includes: inputting the target domain embedding vector and the embedding vector of each 3D Gaussian model into a domain embedding multilayer perceptron, performing an affine transformation on each Gaussian point to map the Gaussian representation onto the distribution of the target domain.
[0008] In some embodiments, maintaining multi-view consistency by predicting the correspondence between different training viewpoints includes: projecting the pixels of the first image into a three-dimensional space to obtain three-dimensional space points, and mapping the three-dimensional space points into image B to constrain the mapping depth of image B to be consistent with the rendering depth.
[0009] In some embodiments, maintaining multi-view consistency by predicting the correspondence between different training viewpoints includes: dividing the image region into reliable regions and dynamically changing regions, and for reliable regions, using low-uncertainty regions as the main constraint objects. For dynamically changing regions, dynamic masks are used to identify moving objects or changing regions in the scene, and flexible constraints are applied to the moving objects or changing regions.
[0010] In some embodiments, the loss function for Gaussian model reconstruction is: in, Representative reference image, Represents the reconstructed image. This represents the affine image. As can be seen, the reconstruction loss function consists of two terms: Constraints and constraint, Used to balance the proportions of the two items.
[0011] In some embodiments, the domain coordination loss function is: in, This is a region of low uncertainty. This represents the overlapping area between image A and image B. and These represent the pixels in the overlapping region of image A and image B, respectively. To render depth, The mapping depth.
[0012] On the other hand, embodiments of this application provide a dynamic scene image domain adaptation system based on four-dimensional Gaussian sputtering, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering according to any of the above embodiments.
[0013] This application also provides a computer storage medium storing a computer program, which is executed to implement the dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering in any of the above embodiments.
[0014] Compared with the prior art, the present invention has the following advantages and beneficial effects: The dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering provided in this application, and the dynamic scene image domain adaptation framework based on four-dimensional Gaussian sputtering representation, differs from traditional transfer methods based on single image features. The method proposed in this invention utilizes multiple images to capture richer and more diverse styles, textures, and lighting conditions. This invention first uses an extension to decompose the four-dimensional Gaussian into a conditional three-dimensional Gaussian and a one-dimensional Gaussian based on temporal distribution. Furthermore, a domain embedding module is designed to enable the model to learn conditional embeddings based on the target domain image, and a domain coordination mechanism is designed to ensure visual consistency across multiple viewpoints. Attached Figure Description
[0015] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a flowchart illustrating a dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering provided in some embodiments of the present invention. Figure 2 This is a schematic diagram of different domain adaptation results on an object dataset provided in some embodiments of the present invention; Figure 3 This is a schematic diagram illustrating the adaptation results of different domains on a scene dataset provided in some embodiments of the present invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0018] The terms “comprising” and “having”, and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not limited to the steps or modules listed, but may optionally include steps or modules not listed, or may optionally include other steps or modules inherent to such process, method, product, or device.
[0019] On the one hand, please refer to Figure 1 This application provides a dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering, including the following steps: S10. Acquire dynamic images captured from multiple perspectives at different points in time.
[0020] S20. Generate a four-dimensional Gaussian model based on dynamic images.
[0021] S30. Decompose the four-dimensional Gaussian model into a conditional three-dimensional Gaussian model and a marginal one-dimensional time component.
[0022] In step S30, given dynamic input captured from multiple perspectives at different time points, the four-dimensional Gaussian is first decomposed into a conditional three-dimensional Gaussian and an edge one-dimensional temporal component. Since spatial coordinates and time are independent, this decomposition not only simplifies the spatiotemporal modeling process formally but also provides more flexible operational space for domain adaptation. Subsequently, the relevant processing for domain adaptation is focused on the three-dimensional Gaussian, aligning it with the target domain in appearance, style, and geometry. After domain adaptation is completed, it is combined with the temporal distribution to restore the overall representation of the dynamic scene in four-dimensional space.
[0023] S40. Extract the target domain embedding vector.
[0024] S50. Based on the extracted target domain embedding vector and the embedding vector of the three-dimensional Gaussian model, perform an affine transformation on each Gaussian point to map the Gaussian representation onto the distribution of the target domain.
[0025] S60. Maintain multi-view consistency by predicting the correspondence between different training perspectives.
[0026] In step S60, the core idea of the domain coordination mechanism is to maintain multi-view consistency by predicting the depth correspondence between different training views. Unlike traditional methods that mainly calculate pixel position differences, the depth correspondence-based method emphasizes the consistency of depth values in 3D space. By fully utilizing the spatial relationships between multiple views, the domain coordination mechanism can provide more robust and accurate optimization results than traditional pixel difference constraints, thereby improving the geometric and appearance consistency of scene reconstruction.
[0027] There is no necessary order between steps S40 and S10-S30. Steps S40 and S10-S30 can be performed sequentially or simultaneously.
[0028] In some embodiments, in step S40, extracting the target domain embedding vector, the target domain embedding vector is extracted by a domain multilayer perceptron.
[0029] In the above embodiments, a domain multilayer perceptron can be designed through a domain embedding module to extract the target domain embedding vector, which can capture the overall scene features, style and appearance information of the target domain.
[0030] In some embodiments, step S50, performing an affine transformation on each Gaussian point to map the Gaussian representation onto the distribution of the target domain, includes: inputting the target domain embedding vector and the embedding vector of each 3D Gaussian model into the DE-MLP, performing an affine transformation on each Gaussian point to map the Gaussian representation onto the distribution of the target domain.
[0031] In the above embodiments, the target domain embedding vector and the embedding vector of each 3D Gaussian are simultaneously used as inputs to the domain embedding multilayer perceptron. The domain embedding multilayer perceptron utilizes the joint information of the two types of embeddings to perform an affine transformation on each Gaussian point, mapping its Gaussian representation, which originally belonged to the input domain, to the distribution of the target domain, thereby achieving cross-domain style adjustment and appearance transfer.
[0032] In some embodiments, step S60, which maintains multi-view consistency by predicting the correspondence between different training viewpoints, includes: projecting the pixels of the first image into a three-dimensional space to obtain three-dimensional space points, and mapping the three-dimensional space points into image B to constrain the mapping depth of image B to be consistent with the rendering depth.
[0033] In the above embodiments, the domain coordination mechanism first projects the pixels of image A into a three-dimensional space, and then uses the camera's intrinsic and extrinsic parameter matrices to map these three-dimensional space points onto image B, thereby constraining the mapping depth of image B to maintain consistency with the rendering depth. This method can effectively ensure the geometric consistency of multiple views in static scenes.
[0034] In some other embodiments, S60, the step of maintaining multi-view consistency by predicting the correspondence between different training viewpoints includes: dividing the image region into a reliable region and a dynamically changing region, and using the low uncertainty region as the main constraint object for the reliable region. For dynamically changing regions, dynamic masks are used to identify moving objects or changing regions in the scene, and flexible constraints are applied to the moving objects or changing regions.
[0035] In contrast, the approach of projecting pixels from the first image into 3D space to obtain 3D spatial points, and then mapping these points onto image B to constrain the mapping depth of image B to maintain consistency with the rendering depth, has limitations when handling dynamic regions. To address this, the domain coordination mechanism introduces uncertainty maps and dynamic masks, employing different optimization strategies for reliable and dynamically changing regions. For reliable regions, low-uncertainty regions are used as the primary constraint objects to ensure accurate reconstruction of the geometry and appearance of static parts. For dynamic regions, moving objects or changing areas in the scene are identified through dynamic masks, and flexible constraints are applied to these regions to avoid over-penalizing dynamic objects. This strategy enables the domain coordination mechanism to effectively enhance scene reconstruction accuracy under multi-view and temporally varying conditions.
[0036] In some embodiments, the loss function used in training the above model is mainly mentioned in the following embodiments: The loss function for Gaussian model reconstruction is: in, Representative reference image, Represents the reconstructed image. This represents the affine image. As can be seen, the reconstruction loss function consists of two terms: Constraints and constraint, Used to balance the proportions of the two items.
[0037] The domain coordination loss function is: in, This is a region of low uncertainty. This represents the overlapping area between image A and image B. and These represent the pixels in the overlapping region of image A and image B, respectively. To render depth, The mapping depth.
[0038] In some examples, the domain adaptation loss function includes content loss and style loss, with the content loss defined as follows: in, This is the ReLU4_2 layer of the VGG19 network. Features extracted by VGG19 are used here to measure high-level semantic differences, thus forcing the preservation of content similarity. Similarly, a style loss function is designed based on the mean and standard deviation of the source and target domain features extracted by VGG19. in, For the VGG19 network, the relu1_1, relu2_1, relu3_1, and relu4_1 layers are... and The mean and standard deviation represent the characteristics of the source domain. and The mean and standard deviation represent the characteristics of the source domain.
[0039] On the other hand, embodiments of this application provide a dynamic scene image domain adaptation system based on four-dimensional Gaussian sputtering, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering according to any of the above embodiments.
[0040] This application also provides a computer storage medium storing a computer program, which is executed to implement the dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering in any of the above embodiments.
[0041] Effect demonstration: Figure 2 and Figure 3 The domain adaptation results of this invention on two different datasets are shown. Figure 2 The first row shows the domain adaptation results on the object dataset. The second and third rows show the domain adaptation results guided by different target domains. Figure 3 To achieve domain adaptation results on scene datasets, this invention can effectively learn different target styles, such as painting styles and animation styles, while realizing high-quality dynamic scene reduction.
[0042] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of the invention in any way. Any simple modifications or equivalent changes made to the above embodiments based on the technical essence of the present invention shall fall within the protection scope of the present invention.
Claims
1. A dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering, characterized in that, The method comprises: acquiring dynamic images captured at different time points from multiple perspectives; generating a four-dimensional Gaussian model based on the dynamic images; decomposing the four-dimensional Gaussian model into conditional three-dimensional Gaussian models and edge one-dimensional time components; extracting target domain embedding vectors; performing affine transformation on each Gaussian point based on the extracted target domain embedding vectors and embedding vectors of the three-dimensional Gaussian models, and mapping the Gaussian representation onto the distribution of the target domain; maintaining multi-perspective consistency by predicting the correspondence between different training perspectives.
2. The dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering according to claim 1, wherein, In the step of extracting target domain embedding vectors, the target domain embedding vectors are extracted by a domain multi-layer perception.
3. The dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering of claim 1, wherein, The step of performing affine transformation on each Gaussian point and mapping the Gaussian representation onto the distribution of the target domain comprises: inputting the target domain embedding vectors and the embedding vectors of each three-dimensional Gaussian model into a domain embedding multi-layer perception, performing affine transformation on each Gaussian point, and mapping the Gaussian representation onto the distribution of the target domain.
4. The dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering of claim 1, wherein, The step of maintaining multi-perspective consistency by predicting the correspondence between different training perspectives comprises: projecting pixels of the first image into a three-dimensional space to obtain three-dimensional space points, and mapping the three-dimensional space points into the image B to constrain the mapping depth of the image B to be consistent with the rendering depth.
5. The dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering of claim 1, wherein, The step of maintaining multi-perspective consistency by predicting the correspondence between different training perspectives comprises: dividing the image region into a reliable region and a dynamic change region, for the reliable region, using the low-uncertainty region as the main constraint object; for the dynamic change region, identifying moving objects or change regions in the scene through a dynamic mask, and using flexible constraints on the moving objects or change regions.
6. The dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering according to claim 5, wherein, The loss function of the Gaussian model reconstruction is: wherein, represents a reference image, represents a reconstructed image, represents an affine post image. It can be seen that the reconstruction loss function includes two terms: constraints and constraints, for balancing the proportion of the two terms.
7. The dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering according to claim 5, wherein, The domain coordination loss function is: wherein, is a low uncertainty region, represents the overlapping region of image A and image B, and are pixels of the overlapping region of image A and image B, respectively, is a rendered depth, is a mapped depth.
8. A dynamic scene image domain adaptation system based on four-dimensional Gaussian sputtering, characterized in that, The method comprises a memory and a processor, the memory stores a computer program, and the processor executes the computer program to implement the four-dimensional Gaussian sputtering based dynamic scene image domain adaptation method of any one of claims 1-7.
9. A computer storage medium, characterized in that A computer program is stored thereon, and the computer program is executed to implement the four-dimensional Gaussian sputtering based dynamic scene image domain adaptation method of any one of claims 1-7.