Dynamic four-dimensional content generation method and device based on single image, equipment and medium

By generating multi-view image sets based on a single image, a dynamic four-dimensional scene representation model is constructed and optimized. This solves the problems of discontinuity between viewpoints and instability of dynamic details in the time dimension during multi-view image generation, enabling the generation of high-quality dynamic content and improving the user experience.

CN120912758APending Publication Date: 2025-11-07GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510813505.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing technologies suffer from insufficient consistency between viewpoints and unstable dynamic details in the time dimension in multi-view image generation, resulting in discontinuous generated images that affect panoramic stitching and user experience.

Method used

By generating a multi-view image set based on a single image, a static 3D scene representation model is constructed and converted into a dynamic 4D scene representation model. Temporal consistency optimization and controllable background lighting editing are then performed to generate a coherent dynamic frame sequence in the time dimension.

Benefits of technology

It enhances the realism of dynamic content and user experience, generating high-quality, time-coherent dynamic four-dimensional content suitable for fields such as virtual reality, augmented reality, game development, and film and television production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120912758A_ABST
    Figure CN120912758A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, and discloses a dynamic four-dimensional content generation method and device based on a single image, equipment and a medium, and the method comprises the steps: generating a corresponding multi-view image set based on an input single static image; constructing a static three-dimensional scene representation model based on the multi-view image set; converting the static three-dimensional scene representation model into a dynamic four-dimensional scene representation model; performing time consistency optimization processing on the dynamic four-dimensional scene representation model to generate a dynamic frame sequence coherent in time dimension; and background illumination controllable editing processing is carried out on the dynamic frame sequence after time consistency optimization processing, and final editable dynamic four-dimensional content is generated. According to the method and the device, the three-dimensional model is constructed by utilizing the multi-view image set and is further converted into the dynamic four-dimensional model, so that the problems of insufficient multi-view image continuity and unstable dynamic details in time dimension in the prior art are solved, and the reality sense of dynamic contents and the user experience are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision, in particular to a single-image-based dynamic four-dimensional content generation method, device, equipment and medium. BACKGROUND

[0002] In the field of computer vision and graphics, with the growing demand for virtual scene construction and dynamic content generation, dynamic content generation technology has become a research hotspot. The current mainstream technology mainly relies on large-scale data to learn geometric and visual priors. However, the existing technology has some defects in multi-view image generation. For example, the existing method is difficult to ensure consistency between different views, especially when dealing with a large view span, the generated image often has the problem of incoherence, which brings great challenges to panoramic stitching and other applications. In addition, in the process of dynamic four-dimensional content generation, the existing technology is difficult to maintain the coherence of dynamic details in the time dimension, resulting in flickering or distortion of the generated video content, which greatly affects the realism of the generated results and the immersive experience of the user. It can be seen that the existing technology has the defects of insufficient coherence of multi-view images and instability of dynamic details in the time dimension, which seriously affects the practicality of the generated content and the user experience.

[0003] The foregoing narrative is to provide general background information and does not necessarily constitute the prior art. SUMMARY

[0004] The embodiments of the present application provide a single-image-based dynamic four-dimensional content generation method, device, equipment and medium, which constructs a three-dimensional model by using a multi-view image set, and further converts it into a dynamic four-dimensional model, solves the problems of insufficient coherence of multi-view images and instability of dynamic details in the time dimension in the existing technology, and improves the realism of dynamic content and user experience.

[0005] In a first aspect, the embodiments of the present application provide a single-image-based dynamic four-dimensional content generation method, comprising:

[0006] Generating a corresponding multi-view image set based on an input single static image;

[0007] Constructing a static three-dimensional scene representation model based on the multi-view image set;

[0008] Converting the static three-dimensional scene representation model into a dynamic four-dimensional scene representation model;

[0009] Performing time consistency optimization processing on the dynamic four-dimensional scene representation model to generate a dynamic frame sequence that is coherent in the time dimension;

[0010] The dynamic frame sequence after the time consistency optimization processing is subjected to background light controllable editing processing, and a final editable dynamic four-dimensional content is generated.

[0011] Optionally, in some embodiments of the present application, the single static image is input, and a corresponding multi-view image set is generated based on the input single static image, and the method comprises the following steps:

[0012] The input single static image and target view parameter are input into a view condition diffusion model to generate a preliminary multi-view image, wherein the target view parameter represents a camera pose change amount relative to the input image view angle;

[0013] The camera pose parameter of the preliminary multi-view image is adjusted, and the smoothness of the visual transition of the preliminary multi-view image is optimized by a control variable method to obtain an optimized multi-view image set.

[0014] Optionally, in some embodiments of the present application, the static three-dimensional scene representation model is constructed based on the multi-view image set, and the method comprises the following steps:

[0015] A multi-view fusion technology is adopted to fuse the three-dimensional structure information of different view images in the multi-view image set based on geometric and light consistency;

[0016] The three-dimensional point cloud in space is initialized, and the parameters of the three-dimensional point cloud are optimized through a diffusion model loss function to construct a static three-dimensional scene representation model.

[0017] Optionally, in some embodiments of the present application, the parameters of the three-dimensional point cloud are optimized through a diffusion model loss function, and the method comprises the following steps:

[0018] A unit scale non-rotating three-dimensional Gaussian point cloud is initialized at a random position in space;

[0019] A two-dimensional diffusion prior is used as supervision information, and the parameters of the three-dimensional point cloud are optimized through a fractional distillation sampling loss function.

[0020] Optionally, in some embodiments of the present application, the static three-dimensional scene representation model is converted into a dynamic four-dimensional scene representation model, and the method comprises the following steps:

[0021] A dynamic deformation field in the time dimension is determined based on the multi-view image set, and the dynamic deformation field is used to represent the dynamic characteristics in the time dimension;

[0022] The static three-dimensional scene representation model is subjected to time sequence optimization based on the dynamic deformation field, and a dynamic four-dimensional scene representation model is generated by minimizing the error between the model rendering image at each timestamp and the multi-view reference image.

[0023] Optionally, in some embodiments of the present application, the time consistency optimization processing on the dynamic four-dimensional scene representation model generates a dynamic frame sequence that is coherent in the time dimension, including:

[0024] selecting an image of any view from the multi-view image set as a target single-view image;

[0025] generating a time sequence dynamic frame of a single view based on the target single-view image through a stable 4D diffusion model;

[0026] inputting each frame of the time sequence dynamic frame into a multi-view image generation module to generate a multi-view dynamic frame under corresponding view parameters;

[0027] constructing the multi-view dynamic frame into an image matrix according to view and time step, wherein the columns of the image matrix correspond to different views, and the rows correspond to different time steps;

[0028] optimizing the dynamic four-dimensional scene representation model based on the image matrix to generate a dynamic frame sequence that is coherent in the time dimension.

[0029] Optionally, in some embodiments of the present application, the background light controllable editing processing on the dynamic frame sequence after the time consistency optimization processing generates a final editable dynamic four-dimensional content, including:

[0030] based on light transmission consistency constraints, constraint processing is performed on the multi-view and multi-time step dynamic frame sequence to retain the inherent properties of the image;

[0031] light consistency optimization is performed on the dynamic frame sequence after the constraint processing to adaptively adjust the background of the dynamic frame sequence under different lighting conditions, thereby generating the final editable dynamic four-dimensional content.

[0032] In a second aspect, the embodiments of the present application provide a dynamic four-dimensional content generation device based on a single image, including:

[0033] a multi-view image module configured to generate a corresponding multi-view image set based on an input single static image;

[0034] a three-dimensional model construction module configured to construct a static three-dimensional scene representation model based on the multi-view image set;

[0035] a four-dimensional model construction module configured to convert the static three-dimensional scene representation model into a dynamic four-dimensional scene representation model;

[0036] an optimization module configured to perform time consistency optimization processing on the dynamic four-dimensional scene representation model to generate a dynamic frame sequence that is coherent in the time dimension;

[0037] An editing module is configured to perform background light controllable editing processing on the dynamic frame sequence after the time consistency optimization processing, and generate the final editable dynamic four-dimensional content.

[0038] In a third aspect, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the single-image-based dynamic four-dimensional content generation method according to the first aspect when executing the computer program.

[0039] In a fourth aspect, a storage medium is provided, which stores a computer program capable of being loaded and executed by a processor to implement the single-image-based dynamic four-dimensional content generation method according to the first aspect.

[0040] The present application provides a single-image-based dynamic four-dimensional content generation method, device, equipment and medium. First, based on an input single static image, a corresponding multi-view image set is generated to solve the problem of insufficient continuity of multi-view images in the prior art. Second, the multi-view image set is used to construct a static three-dimensional scene representation model to provide a basis for subsequent dynamic content generation. Third, the static three-dimensional scene representation model is converted into a dynamic four-dimensional scene representation model to introduce a time dimension and solve the problem of time continuity in dynamic content generation. Fourth, the dynamic four-dimensional scene representation model is subjected to time consistency optimization processing to generate a dynamic frame sequence that is continuous in the time dimension and improve the realism of dynamic content. Finally, the dynamic frame sequence after the time consistency optimization processing is subjected to background light controllable editing processing to generate the final editable dynamic four-dimensional content, effectively improving user experience. It can be seen that the present application solves the problems of insufficient continuity of multi-view images and unstable dynamic details in the time dimension in the prior art by using a multi-view image set to construct a static three-dimensional scene representation model and further converting it into a dynamic four-dimensional scene representation model, realizing the conversion from a single image to high-quality dynamic four-dimensional content, and thus improving the realism of dynamic content and user experience. BRIEF DESCRIPTION OF DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0042] Figure 1 is an application environment diagram of the single-image-based dynamic four-dimensional content generation method provided by the embodiments of the present application;

[0043] Figure 2FIG. 1 is a flow diagram of a method for generating dynamic four-dimensional content based on a single image according to an example embodiment of the present application;

[0044] Figure 3 FIG. 2 is a schematic diagram of a 4D content generation model based on a single image according to an example embodiment of the present application;

[0045] Figure 4 FIG. 3 is a flow diagram of a learning process of a multi-view image generation module according to an example embodiment of the present application;

[0046] Figure 5 FIG. 4 is a flow diagram of a working principle of an image matrix module according to an example embodiment of the present application;

[0047] Figure 6 FIG. 5 is a flow diagram of a working principle of a background controllable editing module according to an example embodiment of the present application;

[0048] Figure 7 FIG. 6 is a data flow diagram of a 4D content generation model based on a single image according to an example embodiment of the present application;

[0049] Figure 8 FIG. 7 is a structural diagram of a device for generating dynamic four-dimensional content based on a single image according to an example embodiment of the present application;

[0050] Figure 9 FIG. 8 is a structural diagram of an electronic device according to an example embodiment of the present application. DETAILED DESCRIPTION

[0051] The example embodiments will be described in detail herein with reference to the attached drawings. The following description is made with reference to the accompanying drawings in which like reference numerals refer to like elements in the several figures. The following description of example embodiments does not represent all of the implementations consistent with the present application. Instead, they are merely examples of implementations consistent with some aspects of the present application as detailed in the appended claims.

[0052] It should be noted that, in this document, by the term "comprising" or "including" or any other variant thereof, it is intended to encompass non-exclusive inclusion, such that processes, methods, articles, or apparatuses that comprise a list of elements are not necessarily limited to those elements, but can include other elements not expressly listed or inherent to such processes, methods, articles, or apparatuses. Without further limitation, an element preceded by "comprises... a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or apparatus that includes the element. Also, like terms have like meanings, unless stated otherwise.

[0053] It should be understood that the specific embodiments described herein merely exemplify the application and should not be considered limiting.

[0054] In the following description, the suffix used for an element such as "module", "part", or "unit" is merely for convenience in describing the present application, and does not have a specific meaning by itself. Thus, "module", "part", or "unit" can be used interchangeably.

[0055] In the field of multi-view generation, existing methods rely on large-scale data to learn geometric and visual priors. However, when generating multi-view images from single-view input, the consistency between views is particularly insufficient. Especially in scenes with large view span, the discontinuity of generated images makes it difficult to perform panoramic stitching, limiting the application of complex scenes. In addition, existing technologies are often limited by low output resolution, insufficient sampling efficiency, difficulty in supporting real-time interaction, and high demand for computing resources during training. The image processing process after generation is complicated, increasing the cost of practical application.

[0056] In the field of dynamic scene (4D content) generation, although rendering efficiency and storage cost are gradually optimized, the robustness of dynamic representation methods still faces challenges. Existing solutions struggle to balance training speed, storage overhead, and rendering quality. Some methods rely on high-quality reference videos, limiting their generalization ability. The realization of real-time high-resolution output still needs to sacrifice some dynamic details, and the contradiction between generation speed and computing resource consumption has not been effectively solved. During the reconstruction of dynamic scenes, complex motion modeling and the preservation of spatiotemporal continuity remain technical bottlenecks.

[0057] For the editing of background and light in dynamic scenes, existing researches focus on static 2D or 3D objects, and realize the adaptation of surface characteristics by adjusting the light conditions. However, the real-time adaptive adjustment of the gloss of object surfaces in 4D dynamic scenes with changing backgrounds is still in the exploratory stage. Dynamic light perception and rendering technology is not mature, and it is difficult to maintain the consistency of texture and light under complex spatiotemporal changes, resulting in insufficient realism of generated content. In addition, variance shift and weak inter-frame correlation problems often occur in multi-view texture aggregation, which need to be solved.

[0058] To solve the above technical problems and overcome the defects of the prior art, the embodiments of the present application provide a single-image-based dynamic four-dimensional content generation method, device, equipment and medium, which can solve the problems of insufficient continuity of multi-view images and unstable dynamic details in the time dimension in the prior art, thereby improving the realism of dynamic content and user experience.

[0059] Figure 1 An application environment diagram for a single-image-based dynamic four-dimensional content generation method in an embodiment. Referring to Figure 1 The single-image-based dynamic four-dimensional content generation method is based on a single-image-based dynamic four-dimensional content generation system. The single-image-based dynamic four-dimensional content generation system includes a terminal 110 and a server 120. The terminal 110 and the server 120 are connected through a network, and the terminal 110 can be a desktop terminal or a mobile terminal, and the mobile terminal can be at least one of a mobile phone, a tablet computer, a notebook computer, etc. The server 120 can be implemented by an independent server or a server cluster composed of multiple servers. The terminal 110 can be used to generate a corresponding multi-view image set based on an input single static image, construct a static three-dimensional scene representation model based on the multi-view image set, convert the static three-dimensional scene representation model into a dynamic four-dimensional scene representation model, perform time consistency optimization processing on the dynamic four-dimensional scene representation model to generate a dynamic frame sequence that is coherent in the time dimension, and perform background and light controllable editing processing on the dynamic frame sequence after the time consistency optimization processing to generate a final editable dynamic four-dimensional content.

[0060] Please refer to Figure 2 , Figure 2 is a flowchart of a single-image-based dynamic four-dimensional content generation method provided by an embodiment of the present application. The embodiment mainly takes the application of the single-image-based dynamic four-dimensional content generation method in a computer device as an example to illustrate, and the single-image-based dynamic four-dimensional content generation method provided by an embodiment of the present application can specifically include the following steps:

[0061] S1. Based on an input single static image, a corresponding multi-view image set is generated.

[0062] Specifically, for step S1, the step generates a preliminary multi-view image using a view condition diffusion model, taking the input single static image and target view parameters as inputs. The target view parameters are used to represent the camera pose change relative to the input image view. The camera pose parameters of the preliminary multi-view image are adjusted through the control variable method to optimize the smoothness of the visual transition, and finally an optimized multi-view image set is obtained.

[0063] S2. Construct a static three-dimensional scene representation model based on the multi-view image set;

[0064] Specifically, for step S2, a multi-view fusion technology is used to fuse the three-dimensional structure information of different view images in the multi-view image set based on geometric and lighting consistency. At the same time, the three-dimensional point cloud in space is initialized, and the parameters of the three-dimensional point cloud are optimized through a diffusion model loss function, thereby constructing a static three-dimensional scene representation model. The multi-view fusion technology can more accurately recover the three-dimensional structure of the scene by matching and fusing the features of multi-view images.

[0065] S3. Convert the static three-dimensional scene representation model into a dynamic four-dimensional scene representation model;

[0066] Specifically, for step S3, a dynamic deformation field in the time dimension is determined based on the multi-view image set, which is used to represent the dynamic characteristics in the time dimension. Then, the static three-dimensional scene representation model is time-optimized based on the dynamic deformation field, and a dynamic four-dimensional scene representation model is generated by minimizing the error between the model rendering image at each timestamp and the multi-view reference image.

[0067] S4. Perform time consistency optimization processing on the dynamic four-dimensional scene representation model to generate a dynamic frame sequence that is continuous in the time dimension;

[0068] Specifically, for step S4, an image of any view is selected from the multi-view image set as a target single-view image, and a time-series dynamic frame of a single view is generated based on the target single-view image through a stable 4D diffusion model. Then, each frame of the time-series dynamic frame is input into a multi-view image generation module to generate a multi-view dynamic frame under the corresponding view parameters. Finally, the multi-view dynamic frames are constructed into an image matrix according to the view and time step, and the dynamic four-dimensional scene representation model is optimized based on the image matrix to generate a dynamic frame sequence that is continuous in the time dimension.

[0069] S5. Perform background lighting controllable editing processing on the dynamic frame sequence after time consistency optimization processing to generate a final editable dynamic four-dimensional content;

[0070] Specifically, for step S5, based on the light transmission consistency constraint, the multi-view, multi-time step dynamic frame sequence is constrained and processed to retain the inherent properties of the image. Then, the light consistency optimization is performed on the dynamic frame sequence after the constraint processing, so as to adaptively adjust the background of the dynamic frame sequence under different lighting conditions, and generate the final editable dynamic four-dimensional content.

[0071] The embodiment effectively solves the problems of insufficient continuity of multi-view images and instability of dynamic details in the time dimension in the prior art by sequentially going through a series of innovative steps of multi-view image generation, static three-dimensional scene construction, dynamic four-dimensional scene conversion, time consistency optimization processing, and background lighting editing from a single static image. It can generate high-quality dynamic four-dimensional content with time continuity, significantly improve the realism and user experience of dynamic content, and provide more realistic and immersive dynamic 3D experience for virtual reality (VR), augmented reality (AR), game development, film production and other fields.

[0072] Optionally, in some embodiments, step S1 "generating a corresponding multi-view image set based on the input single static image" can specifically include:

[0073] S11. input the input single static image and the target view parameter into the view condition diffusion model to generate a preliminary multi-view image, wherein the target view parameter represents the camera pose change amount relative to the input image view angle;

[0074] Specifically, for step S11, the input single static image and the target view parameter are input as input to generate a preliminary multi-view image through the view condition diffusion model. The target view parameter is used to represent the camera pose change amount relative to the input image view angle. The view condition diffusion model is a deep learning-based generative model that can generate corresponding multi-view images according to the input image and specified view parameters. In the training process, the model learns the feature changes of the image under different view angles, so as to generate images with similar content but different view angles as the input image when given new view parameters. In this way, a basis can be provided for subsequent view adjustment.

[0075] S12. adjust the camera pose parameters of the preliminary multi-view image, and optimize the smoothness of the visual transition of the preliminary multi-view image through the control variable method to obtain an optimized multi-view image set;

[0076] Specifically, for step S12, the preliminary generated multi-view images are further optimized to improve the smoothness of visual transition by adjusting the camera pose parameters. The control variable method is used to fine-tune these parameters to ensure that the images between different views can transition naturally. The control variable method is an experimental design method that studies the effect of a variable on the result by controlling other variables to be constant. Here, by controlling other image features to be constant and only adjusting the camera pose parameters, the visual effect during view transition can be accurately optimized. In this way, the geometric relationship between different view images can be ensured to be reasonable, thereby improving the smoothness of the visual effect.

[0077] The embodiment introduces a view condition diffusion model and a control variable method to generate a high-quality multi-view image set from a single static image, solves the problem of discontinuity between views in the prior art multi-view image generation, and provides high-quality multi-view data support for subsequent three-dimensional scene construction. In addition, by optimizing the camera pose parameters, the visual effect of the multi-view image is improved, and the smooth transition between different views is ensured, thereby laying a solid foundation for generating dynamic four-dimensional content with temporal coherence.

[0078] Optionally, in some embodiments, step S2 "constructing a static three-dimensional scene representation model based on the multi-view image set" can specifically include:

[0079] S21. Using multi-view fusion technology, the three-dimensional structure information of different view images in the multi-view image set is fused based on geometric and lighting consistency;

[0080] Specifically, for step S21, the multi-view fusion technology is used to integrate the three-dimensional structure information of different view images in the multi-view image set. Geometric consistency ensures that the three-dimensional coordinates of corresponding points in different view images are consistent, while lighting consistency ensures that the color and brightness of the object surface under different views reasonably reflect its true physical properties. Multi-view fusion technology can more accurately recover the three-dimensional structure of the scene by matching and fusing the features of multi-view images. In actual operation, geometric consistency can be achieved through feature point matching, optical flow calculation, etc., while lighting consistency can be ensured through color correction, lighting compensation, etc. For example, by calculating the optical flow field between different view images, the motion trajectory of the pixel points under different views can be estimated, thereby achieving geometric alignment; through color histogram matching or lighting estimation and compensation based on physical models, the lighting difference between different views can be eliminated, and the lighting consistency can be improved.

[0081] S22. The three-dimensional point cloud in space is initialized, and the parameters of the three-dimensional point cloud are optimized through a diffusion model loss function to construct a static three-dimensional scene representation model;

[0082] Specifically, in step S22, the 3D point cloud in space is first initialized, and then the parameters of the 3D point cloud are optimized using the diffusion model loss function to construct a static 3D scene representation model. 3D point cloud initialization typically involves randomly distributing a certain number of points in space, or making a preliminary estimate based on depth information from multi-view images. The diffusion model loss function guides the optimization process of the 3D point cloud parameters by measuring the difference between the model-generated image and the real image, enabling the optimized 3D point cloud to more accurately represent the geometric shape and appearance characteristics of the scene. Specifically, the diffusion model optimizes the point cloud parameters by progressively adding noise and learning a denoising process, thereby capturing detailed information in the image during the optimization process, allowing the final 3D point cloud model to more realistically reflect the geometric structure of the scene.

[0083] This embodiment achieves effective integration and optimization of the three-dimensional structural information of a multi-view image set by employing multi-view fusion technology and diffusion model loss function, solving the problem of geometric and lighting inconsistencies in multi-view images, and also constructing an accurate static three-dimensional scene representation model, providing a high-quality three-dimensional data foundation for subsequent dynamic four-dimensional scene generation.

[0084] Optionally, in some embodiments, the step S22 of "optimizing the parameters of the 3D point cloud using the diffusion model loss function" may specifically include:

[0085] Initialize a unit-scale, unrotated 3D Gaussian point cloud at a random location in space;

[0086] Specifically, a unit-scale, unrotated 3D Gaussian point cloud is initialized at random locations in 3D space to provide an initial state for subsequent 3D point cloud optimization. A 3D Gaussian point cloud is a mathematical model used to represent a 3D scene; each Gaussian point cloud is described by parameters such as position, scale, rotation, and color. During initialization, the Gaussian point clouds are uniformly distributed in 3D space, and the initial state is unrotated, ensuring that each Gaussian point cloud has the same initial orientation. This initialization method can cover a general area of ​​the scene, providing a reasonable starting point for the subsequent optimization process.

[0087] Using two-dimensional diffusion priors as supervisory information, the parameters of the three-dimensional point cloud are optimized through fractional distillation sampling loss function;

[0088] Specifically, the two-dimensional diffusion prior is a prior knowledge based on a diffusion model, which can capture the detailed information in the image. The fractional distillation sampling loss function is a loss function used to optimize the generative model, which guides the update of the model parameters by comparing the differences between the generated image and the real image. During the optimization process, the three-dimensional point cloud is projected onto the two-dimensional image plane and compared with the two-dimensional diffusion prior, and the fractional distillation sampling loss function is used to adjust the parameters of the three-dimensional point cloud, so that the projection of the three-dimensional point cloud on the two-dimensional image plane is closer to the real image.

[0089] The embodiment realizes effective conversion from the initial point cloud to the high-quality static three-dimensional scene representation model by randomly initializing the unit scale non-rotating three-dimensional Gaussian point cloud in space and optimizing the point cloud parameters by using the two-dimensional diffusion prior and the fractional distillation sampling loss function, improves the accuracy and detail richness of the three-dimensional scene representation, and also ensures the efficiency of the optimization process and the reliability of the results.

[0090] Optionally, in some embodiments, the step S3 of converting the static three-dimensional scene representation model into a dynamic four-dimensional scene representation model can specifically include:

[0091] S31. determining a dynamic deformation field in the time dimension based on the multi-view image set, the dynamic deformation field being used to represent the dynamic characteristics in the time dimension;

[0092] Specifically, for step S31, the multi-view image set is mainly used to determine the dynamic deformation field in the time dimension, which is used to describe the motion and deformation characteristics of the object in the time dimension. The determination of the dynamic deformation field usually involves the analysis of the multi-view image sequence to estimate the position, pose and shape change of the object at different time points. Optical flow estimation, feature point tracking or deep learning methods can be used to analyze the motion information in the image sequence. For example, by calculating the optical flow field between consecutive frames, the motion trajectory of the pixel points can be obtained, and then the motion field of the object can be estimated. In addition, the time convolution network (TCN) or recurrent neural network (RNN) in deep learning can also be used to learn the time sequence features of the object motion, so as to more accurately construct the dynamic deformation field.

[0093] S32. performing time sequence optimization on the static three-dimensional scene representation model based on the dynamic deformation field, generating a dynamic four-dimensional scene representation model by minimizing the error between the model rendering image at each time stamp and the multi-view reference image;

[0094] Specifically, for step S32, the static three-dimensional scene representation model is time-optimized using the dynamic deformation field determined in the foregoing. By adjusting the state of the model at different timestamps, the error between the rendered image of the model at each time point and the corresponding multi-view reference image is minimized, thereby generating a dynamic four-dimensional scene representation model. The time-optimization process can be implemented by various algorithms, such as gradient descent-based optimization algorithms, physics simulation-based optimization methods, etc. In the optimization process, the multi-view reference images can be used to provide constraint conditions to ensure that the optimized model is consistent with the reference images at different angles. At the same time, a regularization term can be introduced to prevent overfitting and ensure the generalization ability of the model. For example, a regularization term about the smoothness of the model can be added in the optimization process, so that the generated dynamic four-dimensional scene is more smooth and natural in time.

[0095] The embodiment determines a dynamic deformation field in the time dimension, and time-optimizes a static three-dimensional scene representation model based on the dynamic deformation field, thereby achieving effective conversion from a static three-dimensional scene to a dynamic four-dimensional scene, generating a dynamic four-dimensional scene representation model highly consistent with multi-view reference images, improving the time coherence and realism of the dynamic scene, and effectively solving the problem of unstable dynamic details in the time dimension in the prior art.

[0096] Optionally, in some embodiments, step S4 "time-consistency optimization processing is performed on the dynamic four-dimensional scene representation model to generate a dynamic frame sequence coherent in the time dimension" can specifically include:

[0097] S41. Select an image of any view from the multi-view image set as a target single-view image;

[0098] Specifically, for step S41, an image of a view is selected from the multi-view image set as a target single-view image to provide a basis for subsequent generation of a single-view time sequence dynamic frame. When selecting the view, it can be determined according to the specific application scenario or user demand. For example, when generating a dynamic human model, the front view can be selected as the target single-view image to better capture the facial expressions and actions of the human. In addition, the most representative view can also be automatically selected as the target view by analyzing the content of the multi-view images.

[0099] S42. Generate a single-view time sequence dynamic frame based on the target single-view image by stabilizing the 4D diffusion model;

[0100] Specifically, for step S42, a stable 4D diffusion model is used to generate a series of time sequence dynamic frames under the target single perspective image, capturing the dynamic changes of the object in the time dimension. The stable 4D diffusion model is a generative model specially designed for processing dynamic scenes, which can generate a dynamic frame sequence that meets the time coherence according to the input single perspective image. By learning a large amount of dynamic scene data, the model captures the rules and characteristics of object motion, and can maintain the temporal consistency between dynamic frames during generation. For example, the model can learn the motion trajectory of the limbs when a person is walking, ensuring the naturalness and coherence of limb movement when generating dynamic frames.

[0101] S43. Input each frame of the time sequence dynamic frame into the multi-perspective image generation module to generate a multi-perspective dynamic frame under the corresponding perspective parameter;

[0102] Specifically, for step S43, the generated single-perspective time sequence dynamic frames are input into the multi-perspective image generation module frame by frame to generate multi-perspective dynamic frames under the corresponding perspective parameter for each frame, enriching the perspective information of the dynamic scene. The multi-perspective image generation module can be implemented based on a deep learning generative model or a geometric transformation algorithm. For example, a view transformer network (View Transformer Network) can be used to predict image content under different perspectives, or a feature matching and interpolation-based method can be used to generate multi-perspective images. The module can generate images of other perspectives corresponding to the input single-perspective image and perspective parameter, thereby realizing perspective expansion.

[0103] S44. Construct the multi-perspective dynamic frames into an image matrix according to the perspective and time step, wherein the columns of the image matrix correspond to different perspectives and the rows correspond to different time steps;

[0104] Specifically, for step S44, the generated multi-perspective dynamic frames are organized according to the perspective and time step to form an image matrix, where the columns of the matrix represent different perspectives and the rows represent different time steps, facilitating subsequent overall optimization processing. The construction of the image matrix can be done in various ways, such as arranging in the order of perspective and time step. During the construction process, the images can be preprocessed, such as size adjustment, normalization, etc., to improve the efficiency and effectiveness of subsequent processing.

[0105] S45. Based on the image matrix, optimize the dynamic four-dimensional scene representation model to generate a time-dimensionally coherent dynamic frame sequence;

[0106] Specifically, for step S45, the constructed image matrix is used to optimize the dynamic four-dimensional scene representation model, so that the generated dynamic frame sequence remains coherent in the time dimension, ensuring the naturalness and realism of the dynamic scene. Various methods can be used in the optimization process, such as gradient descent-based optimization algorithms, energy minimization-based optimization methods, etc. By analyzing the image content at different viewing angles and time steps in the image matrix, the parameters of the dynamic four-dimensional scene representation model are adjusted so that the images generated by the model are as consistent as possible with the images in the image matrix in terms of vision. At the same time, a regularization term can be introduced to maintain the smoothness and stability of the model and avoid overfitting.

[0107] The embodiment realizes the time consistency optimization of the dynamic four-dimensional scene representation model by selecting target single-view images, generating single-view time sequence dynamic frames, expanding to multi-view dynamic frames, constructing an image matrix, and optimizing based on the image matrix, thereby generating a dynamic frame sequence that is coherent in the time dimension and rich in viewing angles.

[0108] Optionally, in some embodiments, step S5 "performing background lighting controllable editing processing on the dynamic frame sequence after time consistency optimization processing to generate the final editable dynamic four-dimensional content" can specifically include:

[0109] S51. Based on the light transmission consistency constraint, the dynamic frame sequence of multiple viewing angles and multiple time steps is processed to preserve the inherent properties of the image;

[0110] Specifically, for step S51, the dynamic frame sequence is processed using the light transmission consistency constraint to ensure that the inherent properties of the image, such as albedo and surface details, are preserved at different viewing angles and time steps. The light transmission consistency constraint is based on the physical principle that the appearance of an object under different lighting conditions can be represented as a linear combination of the appearance under different lighting conditions. Specifically, for each pixel, its color change under different lighting conditions should comply with the principle of linear superposition. By constructing a light transmission matrix and using the matrix to constrain the dynamic frame sequence, it can be ensured that the inherent properties of the image are not affected during the lighting editing process. For example, by estimating the response of each pixel under different lighting conditions, a light transmission matrix is constructed, and then the characteristics of the matrix are ensured to be maintained during optimization.

[0111] S52. Perform lighting consistency optimization on the dynamic frame sequence after constraint processing to adaptively adjust the background of the dynamic frame sequence under different lighting conditions to generate the final editable dynamic four-dimensional content;

[0112] Specifically, for step S52, the dynamic frame sequence after the light transport consistency constraint processing is further optimized for lighting consistency. By adjusting the background lighting, the dynamic frame sequence under different lighting conditions presents a natural visual effect. Lighting consistency optimization can be achieved by various methods, such as using a physically-based rendering model or a data-driven method. Specifically, the IC-Light method can be used, which trains the model by introducing a light transport consistency loss function, so that the model can focus on lighting modification without changing other intrinsic properties of the image. In addition, a deep learning-based diffusion model can also be used, which learns a large amount of image data under different lighting conditions, so that the model can adaptively adjust the background lighting. For example, by training a conditional diffusion model, it can generate corresponding image content according to the input lighting conditions, thereby achieving precise control of the background lighting.

[0113] By introducing light transport consistency constraints and lighting consistency optimization, this embodiment achieves precise control of the background lighting of dynamic four-dimensional content, ensures that the inherent properties of the image are preserved during lighting editing, and can adaptively adjust the background lighting to adapt to different lighting conditions. The generated dynamic four-dimensional content not only has a high degree of realism, but also maintains the consistency of visual effects under various lighting conditions.

[0114] In specific embodiments, the single-image-based dynamic four-dimensional content generation method provided by the embodiment mainly includes high-definition and boundary sharpness dynamic 4D model construction, 4D model time consistency optimization, background lighting controllable 4D content generation and editing. Through the three research contents, single image to image matrix containing motion information is generated, and then background controllable 4D content generation is realized.

[0115] I. High-definition and boundary sharpness dynamic 4D model construction

[0116] The technology of generating 4D content from a single image provides an innovative method for 4D content generation. This method combines a multi-view image generation module with 4D dynamic content optimization technology. This method optimizes the Gaussian deformation field based on 3DGaussianSplatting(3DGS) technology, not only enhancing the continuity and clarity of the generated four-dimensional content, but also significantly speeding up the four-dimensional content generation process. For example, Figure 3As shown, the specific process is as follows: inputting an image, which can be a single static image provided by a user, and all subsequent operations are based on this image. The multi-view image generation module is responsible for generating multiple view images from the input single image. Fine-tuning is performed on the original basis to improve the continuity and clarity of multi-view image generation, including two sub-stages of preliminary multi-view image generation and fine adjustment. A 2D diffusion model is fine-tuned to generate consistent multi-view images, and further fine-tuning of camera pose parameters optimizes the transition between viewpoints to provide a data basis for subsequent 3D model construction. The 3DGS construction module uses multi-view fusion and volume reconstruction technology to convert multi-view images into a 3D Gaussian Splatting model. Multi-view fusion uses geometric and lighting consistency between images to accurately reconstruct the three-dimensional structure; volume reconstruction initializes Gaussian point clouds from random spatial positions, optimizes the fractional distillation sampling (SDS) loss, and recovers three-dimensional geometry and texture information from two-dimensional images to improve the geometric accuracy and visual realism of the three-dimensional model, ensuring consistency when viewing the model from different angles. The picture stack of multiple perspectives represents a series of image sets of different perspectives obtained after multi-view image generation and 3DGS construction. These images are important intermediate results for constructing 3D models and generating 4D content, and they show the appearance and structure information of the object from different perspectives. The 4D content synthesis module converts the static 3DGS model into a dynamic 4D model, introduces a time-dependent dynamic deformation field, and optimizes its parameters to enable the 3D model to exhibit dynamic characteristics that evolve over time while minimizing flicker during motion, ensuring that the generated 4D model is visually consistent with the original video frames and generating an image matrix containing motion information to provide materials for subsequent background light controllable 4D content generation and editing. 4DGS is a dynamic scene representation method used in this embodiment, which is based on 3D Gaussian Splatting and is used to synthesize and represent the final dynamic 4D content. It has higher rendering efficiency, computational stability, and reduces dependence on training data, enabling the generation of high-quality dynamic 4D content from a single image, improving the clarity, dynamic diversity, and visual coherence of the generated content.

[0117] 1) Multi-view image generation

[0118] The enhanced multi-view image generation module is used to generate multi-view images. This model is extensively fine-tuned based on the original version to improve the continuity and clarity of multi-view image generation. This process includes two different sub-stages.

[0119] ① Preliminary multi-view image generation

[0120] The core of this research is to fine-tune a 2D diffusion model conditioned on a view, generating consistent multi-view images for each video frame. During inference, the model takes the original frame image and the required camera-related parameters (such as angular offset and depth variation) as input, and outputs the new view image accordingly. During training, the object is placed at the origin of a standard 3D coordinate system, and a spherical camera setup is simulated. The cameras are placed on a sphere centered on the object and are constrained to always face the origin. Two camera viewpoints are defined using spherical coordinates (θ1, r1) and (θ2, r2), where θ i , and r i represent the polar angle, azimuthal angle, and radius, respectively. Their relative transformation is parameterized as (θ1-θ2, r1-r2). The training objective of the diffusion model is to learn a conditional diffusion model f such that given an input image x1 and a viewpoint transformation (Δθ, Δr), the model can generate a new view image that is very similar to the ground-truth image x2 captured from the target viewpoint. This process is shown in Figure 4 . The model learns the general mechanism of camera viewpoint control and can infer the target image x2 from x1 under arbitrary relative viewpoint changes. The mathematical model is represented as:

[0121]

[0122] where x represents the input image, c(x, R, T) represents the conditional embedding containing the input image and target viewpoint information, t is the diffusion time step, ε is the image encoder, ∈ θ is the u-net-based denoiser, and z t is the latent representation of x at time step t.

[0123] ②Fine-tuning

[0124] After generating the preliminary images, the pose parameters of the cameras are further fine-tuned to optimize the transition between viewpoints. This process is crucial for ensuring a visually smooth transition between the generated images. In dynamic models, fine-tuning the accuracy of the generated viewpoints is key to avoiding visual discontinuities.

[0125] 2) 3D Gaussian Splatting construction

[0126] After obtaining continuous, high-quality multi-view images, advanced three-dimensional reconstruction techniques are used to convert these two-dimensional images into a 3D Gaussian Splatting model. This process includes two key steps.

[0127] ① Multi-view fusion

[0128] Multi-view fusion techniques are used to fuse image information under different views. This technique uses the geometric and lighting consistency between images to accurately reconstruct the three-dimensional structure of the image.

[0129] ② Volume reconstruction

[0130] Initialize unit-scale, rotation-free 3D Gaussian point clouds at random positions in space, and periodically densify them during optimization. Unlike the reconstruction pipeline, start with fewer Gaussian point clouds, but densify them more frequently to align with the generation process. Optimize 3D Gaussian point clouds using fractional distillation sampling (SDS) loss. In each step, randomly sample camera view parameters centered on the object and render RGB images from the current view. During training, linearly reduce the time step t, which weights the random noise added to the rendered RGB images. Then use the input picture as a 2D diffusion prior to optimize the underlying 3D Gaussian point cloud with SDS loss:

[0131]

[0132] where ω(t) is the weight function, ∈ φ is the noise predicted by the two-dimensional diffusion prior φ, Δp is the relative change in camera view parameters relative to the reference camera view parameters r. x is the input image, x r is the new view image obtained by the two-dimensional diffusion model. Through this method, three-dimensional geometry and texture information can be effectively recovered from two-dimensional images, providing a solid foundation for subsequent generation of four-dimensional dynamic models. The optimization goal of this stage is to improve the geometric accuracy and visual realism of the three-dimensional model. In addition, it aims to ensure the consistency of the model from various angles.

[0133] 3) Synthesize 4D content using 4DGaussianSplatting

[0134] Finally, convert the static 3D GS model into a dynamic 4D model. This stage is crucial for introducing time-dependent dynamic deformation fields, which enable the 3D model to exhibit dynamic characteristics that evolve over time. By optimizing the parameters of the dynamic deformation field, ensure that the generated 4D model is visually consistent with the original video frames and minimizes flicker during motion. Based on the multi-view images generated in the first stage, optimize the dynamic deformation field to minimize the mean square error between the rendered images at each timestamp and the multi-view images:

[0135]

[0136] where τ is the time variable, o Ref represents the view information corresponding to the multi-view images generated in the first stage, and the associated images are represented as xτ Ref The image of this viewpoint is rendered using the 4D warping field model f, and the error between the rendered image and x τ Ref is calculated to optimize the model f.

[0137] Through the study of the MVG4G method, not only can an accurate 3D model be reconstructed from a single image, but also dynamic 4D content can be generated. Compared with traditional methods, the technology of this study significantly improves the generation efficiency, naturalness and visual continuity of the generated content. The optimized model improves the clarity in dynamic performance and reduces motion flicker. These improvements are crucial for the practical application of 4D content in fields such as augmented reality and virtual reality, especially for creating more realistic and immersive experiences.

[0138] II. Time-consistent optimization of dynamic 4D models

[0139] After obtaining the multi-view views, the SV4D model is further utilized to generate dynamic frames that are time-consistent while maintaining the original image semantics and style. Specifically, each generated new view is individually input into the SV4D model to generate a continuous dynamic frame sequence under that view.

[0140] After completing the multi-view dynamic frame generation, these frames are organized into an image matrix in a specific manner. The columns of the matrix represent different views, while the rows represent different time steps. The image sequence within the same column reflects the temporal evolution under a fixed view, while the images in the same row exhibit multi-view synchronous frames at the same time step. This matrix representation method not only ensures consistency in the spatial dimension, but also accurately captures the dynamic change characteristics of objects in the time dimension.

[0141] As shown in Figure 5 , this image matrix module, as a key component of the model, generates multi-view images containing dynamic changes in batches, providing rich supervision information for subsequent 4DGS optimization. With this matrix module, more refined dynamic content optimization can be achieved, effectively improving the geometric and texture consistency of 4D scenes in the multi-view and time dimensions, providing higher realism and coherence for the final generated dynamic 4D content.

[0142] III. 4D content generation and editing with controllable background lighting

[0143] In the third phase of this research, the goal is to achieve 4D content generation with controllable background lighting. Currently, the IC-Light method is referenced, which is based on the physical principle that the linear mixture of an object's appearance under different lighting conditions is consistent with its appearance under the mixed lighting. This principle provides a strong constraint mechanism for lighting editing in 4D content generation, ensuring that the model only changes the lighting part of the image during editing, while preserving other inherent properties such as albedo and image details.

[0144] Specifically, the IC-Light method trains the model by introducing a light transport consistency loss function. The purpose of this loss function is to minimize the difference between the mixed lighting and the individual lighting of the images under different lighting conditions, thereby forcing the model to learn a mapping that focuses on lighting modification without changing other intrinsic properties of the image. This method allows for the uniform processing of batch images obtained from the image matrix module, achieving precise lighting editing.

[0145] To achieve this goal, the embodiment designs a background editing module that can process batch images from the image matrix module and perform lighting consistency optimization. As shown in Figure 6 , the background editing module will receive the multi-view, multi-time step image sequence generated by the image matrix module and apply the light transport consistency loss function to ensure that the 4D content generated under different views and time points maintains geometric and texture consistency. In this way, this research can generate dynamic 4D content with high realism and visual coherence, while allowing precise control and editing of background lighting.

[0146] In a specific embodiment, as shown in Figure 7 , the background editing module is implemented as a neural network that takes the input image sequence and outputs the optimized 4D content with controllable background lighting. Figure 7 The complete architecture and data flow of the single-image-based 4D content generation model provided for the embodiment are as follows:

[0147] (1) Diffusion model and 3D GS reconstruction

[0148] Diffusion model optimization: Use the fractional distillation sampling (SDS) loss to combine the diffusion model to gradually generate high-quality 3D GS models from random noise. During training, linearly reduce the time step t, add random noise and use the two-dimensional diffusion prior to optimize, thereby obtaining accurate three-dimensional geometric structure and texture information, ensuring consistency when observing the model from different angles.

[0149] 3D GS construction: Through densification and other operations, the generated Gaussian point cloud is converted into an accurate 3D GS model, laying the foundation for subsequent 4D dynamic content generation.

[0150] (2) Multi-view image generation and matrix construction

[0151] Multi-view image generation module: generates a sequence of images from a single input image, each corresponding to a different timestamp and camera perspective parameter, for subsequent 4D dynamic content generation and background lighting editing.

[0152] Image matrix module: organizes the generated multi-view images into an image matrix, where columns represent different perspectives and rows represent different time steps. This matrix representation ensures consistency in both spatial and temporal dimensions, providing rich supervision information for 4D dynamic content optimization.

[0153] (3) Background editing and 4D content synthesis

[0154] Background editing module: based on the IC-Light method, introduces a light transmission consistency loss function to optimize the batch of images obtained from the image matrix module for lighting consistency, ensuring that the generated 4D content maintains geometric and texture consistency at different perspectives and time points, while allowing precise control and editing of background lighting.

[0155] 4D content synthesis module: converts the multi-view image sequence after background editing into a dynamic 4D GS model. By optimizing the parameters of the dynamic deformation field, the 3D model can exhibit dynamic characteristics over time, while minimizing flicker during motion, ensuring that the generated 4D model is visually consistent with the original video frames.

[0156] (4) 4D GS representation and optimization

[0157] 4D GS representation: uses 4D Gaussian Splatting as a dynamic scene representation method, making the generated 4D content have higher rendering efficiency, computational stability and dynamic characteristics, reducing dependence on training data, and improving the clarity, dynamic diversity and visual coherence of the generated content.

[0158] SDS loss and MSE error optimization: in the entire model, SDS loss and mean square error (MSE) loss functions are used to optimize the generated 4D content. SDS loss is used to guide the model to generate high-quality 3D and 4D content consistent with the input image from the diffusion model, while MSE error is used to ensure the consistency and accuracy of the generated dynamic content in multiple perspectives and time dimensions.

[0159] In summary, the single-image-based dynamic four-dimensional content generation method provided in the embodiment greatly reduces the data acquisition and storage costs compared with the traditional 4D generation method which relies on multi-view image or video data. By optimizing the multi-view image generation module (MVIG) and the 4DGS representation method, the quality of the generated 4D content is improved, making the edges of dynamic objects clearer and the details richer. The image matrix module is designed to introduce time constraints, making the generated 4D content have more dynamic changes and avoiding repetitive and rigid motion patterns. The SV4D is used for timing optimization to reduce flickering, geometric distortion, and texture drift in dynamic scenes, ensuring consistency in different time steps. By introducing a lighting constraint mechanism, the background lighting changes are optimized, making the dynamic scene consistent under different lighting conditions and improving visual fidelity. By optimizing the 4DGS training strategy, the generation process is more efficient, significantly reducing the computational resource requirements compared to methods such as NeRF, and improving the feasibility of practical applications. It is suitable for virtual reality (VR), augmented reality (AR), game development, film production, and other fields, providing a more realistic and immersive dynamic 3D experience.

[0160] It should be understood that, although Figure 2 the steps in the flowchart of the method are displayed in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, Figure 2 at least part of the steps in the method can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be alternately or alternately executed with other steps or sub-steps or stages of other steps.

[0161] To better implement the single-image-based dynamic four-dimensional content generation method of the present application embodiment, the present application embodiment further provides a single-image-based dynamic four-dimensional content generation device based on the above-mentioned single-image-based dynamic four-dimensional content generation method. The meanings of the terms are the same as in the above-mentioned single-image-based dynamic four-dimensional content generation method, and the specific implementation details can be referred to the description in the method embodiment.

[0162] Please refer to Figure 8 , Figure 8A structural schematic diagram of a single-image-based dynamic four-dimensional content generation apparatus provided by an embodiment of the present application is shown in FIG. 2. The single-image-based dynamic four-dimensional content generation apparatus can specifically include a multi-view image module 201, a three-dimensional model construction module 202, a four-dimensional model construction module 203, an optimization module 204, and an editing module 205, and can be specifically as follows:

[0163] The multi-view image module 201 is configured to generate a corresponding multi-view image set based on an input single static image.

[0164] The three-dimensional model construction module 202 is configured to construct a static three-dimensional scene representation model based on the multi-view image set.

[0165] The four-dimensional model construction module 203 is configured to convert the static three-dimensional scene representation model into a dynamic four-dimensional scene representation model.

[0166] The optimization module 204 is configured to perform time consistency optimization processing on the dynamic four-dimensional scene representation model to generate a dynamic frame sequence that is coherent in the time dimension.

[0167] The editing module 205 is configured to perform background light controllable editing processing on the dynamic frame sequence after the time consistency optimization processing to generate a final editable dynamic four-dimensional content.

[0168] The specific limitations of the single-image-based dynamic four-dimensional content generation apparatus can be seen from the limitations of the single-image-based dynamic four-dimensional content generation method described above, and will not be repeated here. The various modules of the single-image-based dynamic four-dimensional content generation apparatus described above can be realized by software, hardware, and combinations thereof, in whole or in part. The various modules described above can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to the various modules.

[0169] The single-image-based dynamic four-dimensional content generation device provided in the embodiment generates a corresponding multi-view image set based on an input single static image through the multi-view image module 201, solves the problem of insufficient continuity of multi-view images in the prior art, secondly, constructs a static three-dimensional scene representation model by using the multi-view image set through the three-dimensional model construction module 202, and provides a basis for subsequent dynamic content generation, then, converts the static three-dimensional scene representation model into a dynamic four-dimensional scene representation model through the four-dimensional model construction module 203, introduces a time dimension, and solves the problem of time continuity in dynamic content generation, then, performs time consistency optimization processing on the dynamic four-dimensional scene representation model through the optimization module 204, generates a dynamic frame sequence that is continuous in the time dimension, and improves the realism of dynamic content, finally, performs background light controllable editing processing on the dynamic frame sequence that has been subjected to the time consistency optimization processing through the editing module 205, and generates final editable dynamic four-dimensional content, thereby effectively improving user experience. It can be seen that, by using the multi-view image set to construct a static three-dimensional scene representation model and further converting the static three-dimensional scene representation model into a dynamic four-dimensional scene representation model, the problems of insufficient continuity of multi-view images and unstable dynamic details in the time dimension in the prior art are solved, the conversion from a single image to high-quality dynamic four-dimensional content is realized, and the realism of dynamic content and user experience are improved.

[0170] In addition, the embodiment of the present application further provides an electronic device, as shown in the figure, which shows a structural schematic diagram of the electronic device related to the embodiment of the present application, in particular: Figure 9

[0171] The electronic device can include a processor 301 with one or more processing cores, a memory 302 with one or more computer readable storage media, a power supply 303, and an input unit 304, and the like. Those skilled in the art can understand that the structure of the electronic device shown in the figure does not constitute a limitation on the electronic device, and can include more or fewer components than the figure, or combine certain components, or different component arrangements. Among them: Figure 9

[0172] The processor 301 is the control center of the electronic device, connects all parts of the electronic device through various interfaces and lines, executes various functions of the electronic device and processes data by running or executing software programs and / or modules stored in the memory 302 and calling data stored in the memory 302, thereby overall monitoring the electronic device. Optionally, the processor 301 can include one or more processing cores; preferably, the processor 301 can integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 301.​​

[0173] The memory 302 can be used to store software programs and modules, and the processor 301 executes various functions and the single-image-based dynamic four-dimensional content generation method by running the software programs and modules stored in the memory 302. The memory 302 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, application programs required by at least one function (such as a sound playing function, an image playing function, etc.), and the like; and the data storage area can store data created according to the use of the electronic device, etc. In addition, the memory 302 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device. Accordingly, the memory 302 can also include a memory controller to provide the processor 301 with access to the memory 302.

[0174] The electronic device also includes a power supply 303 for powering various components. Preferably, the power supply 303 can be logically connected to the processor 301 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 303 can also include one or more direct current or alternating current power supplies, a recharging system, a power supply fault detection circuit, a power supply converter or inverter, a power supply status indicator, and the like.

[0175] The electronic device can also include an input unit 304, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0176] Although not shown, the electronic device can also include a display unit and the like, which will not be described here. Specifically, in the present embodiment, the processor 301 in the electronic device will load the executable file corresponding to the process of one or more application programs into the memory 302 according to the following instructions, and run the application programs stored in the memory 302 by the processor 301, so as to realize various functions, as follows:

[0177] Based on the input single static image, a corresponding multi-view image set is generated; based on the multi-view image set, a static three-dimensional scene representation model is constructed; the static three-dimensional scene representation model is converted into a dynamic four-dimensional scene representation model; the dynamic four-dimensional scene representation model is subjected to time consistency optimization processing to generate a dynamic frame sequence that is coherent in the time dimension; and the dynamic frame sequence subjected to the time consistency optimization processing is subjected to background light controllable editing processing to generate a final editable dynamic four-dimensional content.

[0178] The specific implementation of each operation can refer to the previous embodiments, which will not be described here.

[0179] The embodiment of the present application solves the problems of insufficient continuity of multi-view images and instability of dynamic details in the time dimension in the prior art by constructing a static three-dimensional scene representation model using a multi-view image set and further converting the static three-dimensional scene representation model into a dynamic four-dimensional scene representation model, and realizes conversion from a single image to high-quality dynamic four-dimensional content, thereby improving the realism and user experience of dynamic content.

[0180] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling relevant hardware by the instructions, which can be stored in a computer readable storage medium and loaded and executed by a processor.

[0181] To this end, the embodiment of the present application provides a storage medium, which stores a plurality of instructions capable of being loaded by a processor to execute the steps in any of the dynamic four-dimensional content generation methods based on a single image provided by the embodiment of the present application. For example, the instructions can execute the following steps:

[0182] Based on the input single static image, a corresponding multi-view image set is generated; a static three-dimensional scene representation model is constructed based on the multi-view image set; the static three-dimensional scene representation model is converted into a dynamic four-dimensional scene representation model; time consistency optimization processing is performed on the dynamic four-dimensional scene representation model to generate a dynamic frame sequence that is continuous in the time dimension; and background light controllable editing processing is performed on the dynamic frame sequence after the time consistency optimization processing to generate a final editable dynamic four-dimensional content.

[0183] The specific implementation of each operation can be referred to the foregoing embodiments, which will not be described here.

[0184] The storage medium can include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0185] Since the instructions stored in the storage medium can execute the steps in any of the dynamic four-dimensional content generation methods based on a single image provided by the embodiment of the present application, the beneficial effects that can be achieved by any of the dynamic four-dimensional content generation methods based on a single image provided by the embodiment of the present application can be achieved, which will not be described here for details, and can be referred to the foregoing embodiments.

[0186] The above describes in detail the method, device, equipment and medium provided by the embodiment of the application for generating dynamic four-dimensional content based on a single image. The principles and implementation manners of the application are described by using specific examples. The above description of the embodiments is only used to help understand the method of the application and its core idea. Meanwhile, for those skilled in the art, the specific implementation manners and application ranges can be changed according to the idea of the application. In conclusion, the content of the specification should not be understood as a limitation of the application.

Claims

1. A method for dynamic four-dimensional content generation based on a single image, characterized in that, The method comprises the following steps: Based on the input single static image, a corresponding multi-view image set is generated; Based on the multi-view image set, a static three-dimensional scene representation model is constructed; The static three-dimensional scene representation model is converted into a dynamic four-dimensional scene representation model; The dynamic four-dimensional scene representation model is subjected to time consistency optimization processing to generate a dynamic frame sequence that is continuous in the time dimension; The dynamic frame sequence subjected to the time consistency optimization processing is subjected to background light controllable editing processing to generate a final editable dynamic four-dimensional content.

2. The single-image-based dynamic four-dimensional content generation method of claim 1, wherein, The method comprises the following steps: The input single static image and target view angle parameters are input into a view condition diffusion model to generate preliminary multi-view images, wherein the target view angle parameters represent the camera pose change amount relative to the input image view angle; The camera pose parameters of the preliminary multi-view images are adjusted, and the smoothness of the visual transition of the preliminary multi-view images is optimized by a control variable method to obtain an optimized multi-view image set.

3. The single-image-based dynamic four-dimensional content generation method of claim 1, wherein, The method comprises the following steps: A multi-view fusion technology is used to fuse the three-dimensional structure information of different view images in the multi-view image set based on geometric and light consistency; Three-dimensional point clouds in space are initialized, and the parameters of the three-dimensional point clouds are optimized through a diffusion model loss function to construct a static three-dimensional scene representation model.

4. The single-image-based dynamic four-dimensional content generation method of claim 3, wherein, The method comprises the following steps: Unit scale non-rotating three-dimensional Gaussian point clouds are initialized at random positions in space; A two-dimensional diffusion prior is used as supervision information, and the parameters of the three-dimensional point clouds are optimized through a fractional distillation sampling loss function.

5. The single image based dynamic four-dimensional content generation method of claim 1, wherein, The method comprises the following steps: A dynamic deformation field in the time dimension is determined based on the multi-view image set, and the dynamic deformation field is used to represent the dynamic characteristics in the time dimension; The static three-dimensional scene representation model is subjected to time sequence optimization based on the dynamic deformation field, and a dynamic four-dimensional scene representation model is generated by minimizing the error between the model rendering image at each timestamp and the multi-view reference image.

6. The single image based dynamic four-dimensional content generation method of claim 1, wherein, The method comprises the following steps: An image of any view is selected from the multi-view image set as a target single-view image; A single-view time sequence dynamic frame is generated based on the target single-view image through a stable 4D diffusion model; Each frame of the time sequence dynamic frame is input into a multi-view image generation module to generate a multi-view dynamic frame under corresponding view angle parameters; The multi-view dynamic frame is constructed into an image matrix according to the view angle and the time step, wherein the columns of the image matrix correspond to different view angles, and the rows correspond to different time steps; The dynamic four-dimensional scene representation model is optimized based on the image matrix to generate a dynamic frame sequence that is continuous in the time dimension.

7. The single-image-based dynamic four-dimensional content generation method of claim 1, wherein, The background light controllable editing processing is performed on the dynamic frame sequence after the time consistency optimization processing, to generate the final editable dynamic four-dimensional content, including: Based on the light transmission consistency constraint, the dynamic frame sequence of multiple views and multiple time steps is processed to retain the inherent properties of the image; The dynamic frame sequence after the constraint processing is subjected to light consistency optimization, to adaptively adjust the background of the dynamic frame sequence under different light conditions, to generate the final editable dynamic four-dimensional content.

8. A single image based dynamic four-dimensional content generation apparatus, characterized by, It includes: A multi-view image module is configured to generate a corresponding multi-view image set based on an input single static image; A three-dimensional model construction module is configured to construct a static three-dimensional scene representation model based on the multi-view image set; A four-dimensional model construction module is configured to convert the static three-dimensional scene representation model into a dynamic four-dimensional scene representation model; An optimization module is configured to perform time consistency optimization processing on the dynamic four-dimensional scene representation model, to generate a dynamic frame sequence that is coherent in the time dimension; An editing module is configured to perform background light controllable editing processing on the dynamic frame sequence after the time consistency optimization processing, to generate the final editable dynamic four-dimensional content.

9. An electronic device, comprising: It includes: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the single-image-based dynamic four-dimensional content generation method according to any one of claims 1-7.

10. A storage medium, characterized by A computer program is stored, which can be loaded and executed by the processor to perform the single-image-based dynamic four-dimensional content generation method according to any one of claims 1-7.

Citation Information

Cited By

  • A single-image vehicle four-dimensional reconstruction method, device, electronic equipment and storage medium

    CN122368349A