Geometry and motion joint optimization four-dimensional scene generation method and device
By employing deep-guided motion normalization and motion-aware adaptive normalization strategies, combined with a view synthesis module, the problem of decoupling geometric reconstruction and motion generation in existing technologies is solved. This enables the generation of high-quality, multi-view consistent four-dimensional dynamic scenes from a single static image, thereby enhancing the application effects of virtual reality and simulation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-02
- Publication Date
- 2026-03-17
AI Technical Summary
Existing technologies suffer from the problem of decoupling geometric reconstruction and motion generation when generating high-quality four-dimensional dynamic scenes from a single static image. This results in geometric inconsistencies or limited dynamic performance of the generated results across multiple viewpoints, and high-quality four-dimensional scene data is scarce.
A depth-guided motion normalization strategy is adopted to eliminate scale bias in the motion of objects at different depths. Potential motion features are extracted through the motion perception module and injected into the diffusion model using an adaptive normalization mechanism. Combined with the view synthesis module, dynamic video is rendered and occluded areas are repaired, achieving tight coupling optimization of geometry and motion.
It has achieved the generation of high-quality four-dimensional dynamic scenes from a single static image to multiple perspectives with spatiotemporal consistency, improving the physical rationality and visual coherence of dynamic scenes, and enhancing the realism and adaptability of virtual reality and simulation applications.
Smart Images

Figure CN121685773A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and artificial intelligence, and in particular to a method, apparatus, device and storage medium for generating four-dimensional scenes by joint optimization of geometry and motion. Background Technology
[0002] With the rapid development of computer vision and machine learning technologies, especially the rise of applications such as virtual reality (VR), augmented reality (AR), and immersive content creation, generating high-quality four-dimensional dynamic scenes from single static images has become a core technical challenge in computer graphics and computer vision. Four-dimensional scene generation aims to reconstruct a comprehensive spatiotemporal representation containing explicit three-dimensional geometry and complex temporal dynamics. This technological capability has transformative significance for driving cutting-edge applications such as VR / AR, filmmaking, game development, and digital twins. However, recovering complete spatiotemporal information from inherently limited 2D observation remains a fundamental technical challenge. Although video generation models have been able to produce realistic dynamic content in recent years, they often lack an explicit understanding of three-dimensional structure, leading to inconsistencies across viewpoints and difficulty in capturing physically plausible motion patterns. Currently, existing four-dimensional generation technologies mainly follow two separate technical paradigms, but both have significant limitations: The "post-generative reconstruction" paradigm first utilizes powerful video generation models to synthesize multi-view video sequences, then reconstructs a four-dimensional representation from these videos. The advantage of this approach lies in its ability to fully leverage the high fidelity and rich dynamic characteristics offered by existing video generation models to generate visually realistic dynamic content. However, because the video generation models themselves lack explicit understanding and constraints on three-dimensional geometry, the generated multi-view videos struggle to maintain strict geometric consistency. This geometric inconsistency leads to serious structural problems in the subsequent 3D reconstruction stage, including geometric collapse, spatial distortion, depth inconsistencies, and various visual artifacts, severely impacting the quality and usability of the final four-dimensional scene.
[0003] The "reconstruction-to-generation" paradigm, responding to the limitations of the previous approach, first reconstructs a static 3D geometric model from a single input image, then performs animation generation or motion synthesis based on this pre-established 3D structure. The core advantage of this method lies in providing a solid geometric foundation for subsequent motion generation through the pre-established static 3D structure, effectively avoiding geometric inconsistencies. However, because this paradigm completely decouples static geometric reconstruction from dynamic motion generation, this separate processing method discards the potentially rich dynamic information and motion cues in the source image. Therefore, the motion generated by this type of method is often limited to simple, predictable externally driven actions under physical constraints (such as the swaying or rotation of objects), making it difficult to generate large-scale, spontaneous, and complex motion patterns driven by elements within the scene or the characteristics of the objects themselves. This greatly limits the dynamic expressiveness and realism of the generated 4D scenes.
[0004] Furthermore, existing technologies face the problem of scarce high-quality four-dimensional scene data. Acquiring large-scale, high-quality four-dimensional scene data containing complex dynamics is extremely challenging both technically and costly, further limiting the training effectiveness and generalization ability of related models. Simultaneously, effectively fusing geometric and motion priors from images while maintaining geometric consistency is also a significant challenge for current technologies.
[0005] Existing technologies generally suffer from the fundamental problem of decoupling geometric reconstruction and motion generation, resulting in either geometric inconsistencies across multiple viewpoints or severe limitations in dynamic performance. This technological bottleneck hinders the further development of generating high-quality four-dimensional scenes from single static images. Therefore, there is an urgent need for an innovative technical solution that can tightly couple geometric reconstruction and motion generation to achieve the generation of high-quality four-dimensional scenes from single static images that possess both geometric consistency and rich dynamic details, providing strong technical support for VR / AR and other application fields. Summary of the Invention
[0006] The present invention aims to at least partially solve one of the technical problems in the related art.
[0007] To address this, this invention proposes a four-dimensional scene generation method that jointly optimizes geometry and motion. Based on the depth information of the input image, a depth-guided motion normalization strategy is used to process the motion trajectory of the three-dimensional point cloud, eliminating scale deviations in the motion of objects at different depths. The motion perception module extracts potential motion features from the static image and uses an adaptive normalization mechanism to inject the feature parameters into the diffusion model network, realizing dynamic guidance of motion priors for trajectory generation. Based on the jointly optimized four-dimensional point cloud representation, the view synthesis module renders dynamic video along arbitrary camera trajectories and repairs occluded areas, thereby generating a spatiotemporally consistent four-dimensional dynamic scene with multiple perspectives.
[0008] Another objective of this invention is to propose a four-dimensional scene generation device that combines geometry and motion optimization.
[0009] The third objective of this invention is to provide a computer device.
[0010] A fourth objective of this invention is to provide a non-transitory computer-readable storage medium.
[0011] To achieve the above objectives, this invention proposes a four-dimensional scene generation method that combines geometry and motion optimization, comprising: S1, based on the depth information of the input image, adopts a depth-guided motion normalization strategy to perform scale-invariant representation processing on the motion trajectory of the 3D point cloud, thereby eliminating the scale bias of motion perception of objects at different depths; S2 extracts potential motion features from static images through the motion perception module and injects the feature parameters into each layer of the diffusion model network using the motion perception adaptive normalization mechanism, so as to realize the dynamic guidance of motion prior knowledge on trajectory generation. S3, based on the jointly optimized four-dimensional point cloud representation, renders dynamic video along arbitrary camera trajectories through the view synthesis module, and repairs and completes the occluded areas generated by the rendering, generating a four-dimensional dynamic scene with consistent spatiotemporal perspectives.
[0012] The four-dimensional scene generation method of the present invention, which combines geometry and motion optimization, may also have the following additional technical features: In one embodiment of the present invention, the step of performing scale-invariant representation processing on the motion trajectory of the 3D point cloud based on the depth information of the input image and employing a depth-guided motion normalization strategy to eliminate scale bias in motion perception of objects at different depths includes: S11, based on the focal length of the input image , and image size Calculate the scaling factor and ; S12, representing the relative motion of the 3D point cloud. Convert to scale-invariant representation , , ;in, The initial depth of the point .
[0013] In one embodiment of the present invention, the step of extracting latent motion features from a static image through a motion sensing module and injecting feature parameters into each layer of the diffusion model network using a motion sensing adaptive normalization mechanism to achieve dynamic guidance of trajectory generation based on prior motion knowledge includes: S21 employs a feature extractor pre-trained on joint image-video data as a motion perception module to extract block-level motion features from static images. ; S22, generates token-level adaptive parameters through a linear layer. , , , and according to the formula , The intermediate features of the diffusion model are modulated; whereby, , These are learnable global gating coefficients.
[0014] In one embodiment of the present invention, the step of rendering dynamic video along arbitrary camera trajectories using a view compositing module based on the jointly optimized four-dimensional point cloud representation, and repairing and completing the occluded areas generated by the rendering to generate a multi-view spatiotemporally consistent four-dimensional dynamic scene includes: S31, set the value to 0.5 for the area in the rendered video not covered by the projection point; S32 repairs occluded areas by fine-tuning the video generation model, ensuring visual coherence in both time and space dimensions after repair.
[0015] In one embodiment of the present invention, it further includes: S4 represents the generated four-dimensional point cloud. By jointly optimizing with dynamic semantic constraints input by the user, and adjusting the physical rationality parameters of the motion trajectory, the dynamic behavior of the generated scene is made consistent with the user's needs.
[0016] To achieve the above objectives, another aspect of the present invention proposes a four-dimensional scene generation device that jointly optimizes geometry and motion, comprising: The depth-guided motion normalization module is used to perform scale-invariant representation processing on the motion trajectory of 3D point clouds based on the depth information of the input image and adopt the depth-guided motion normalization strategy to eliminate the scale bias of motion perception of objects at different depths. The motion-aware feature injection module is used to extract potential motion features from static images through the motion-aware module, and inject the feature parameters into each layer of the diffusion model network using the motion-aware adaptive normalization mechanism, so as to realize the dynamic guidance of motion prior knowledge for trajectory generation. The view composition and occlusion repair module is used to render dynamic video along arbitrary camera trajectories based on the jointly optimized four-dimensional point cloud representation, and repair and complete the occluded areas generated by the rendering to generate a four-dimensional dynamic scene with spatiotemporal consistency from multiple perspectives.
[0017] In one embodiment of the present invention, it further includes: The joint optimization module is used to represent the generated four-dimensional point cloud. By jointly optimizing with dynamic semantic constraints input by the user, and adjusting the physical rationality parameters of the motion trajectory, the dynamic behavior of the generated scene is made consistent with the user's needs.
[0018] This invention discloses a geometrically and kinematically optimized four-dimensional scene generation method and apparatus. Through a depth-guided motion normalization strategy, it effectively eliminates scale deviations in the motion of objects at different depths. Furthermore, it utilizes a motion-aware adaptive normalization mechanism to dynamically guide trajectory generation based on motion priors, thus solving the core problems of inaccurate motion representation and insufficient spatiotemporal consistency of dynamic scenes in existing technologies. It achieves full automation from static image input to four-dimensional dynamic scene generation, significantly improving the physical plausibility of motion trajectories and the visual coherence of rendered videos. While generating consistent dynamic scenes from multiple perspectives, it optimizes computational efficiency and user interaction experience, enhancing the realism and adaptability of four-dimensional scene generation in complex applications such as virtual reality and simulation.
[0019] To achieve the above objectives, a third aspect of this application provides a computer device, including a processor and a memory; wherein the processor reads executable program code stored in the memory to run a program corresponding to the executable program code, for implementing a geometry and motion joint optimization four-dimensional scene generation method as described in the first aspect embodiment.
[0020] To achieve the above objectives, a fourth aspect of this application provides a non-transitory computer-readable storage medium storing a computer program that, when executed by a processor, implements a geometry and motion joint optimization method for generating a four-dimensional scene as described in the first aspect.
[0021] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0022] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart of a four-dimensional scene generation method based on the joint optimization of geometry and motion according to an embodiment of the present invention; Figure 2 This is a model architecture diagram illustrating the specific steps of a four-dimensional scene generation method based on geometry and motion joint optimization according to an embodiment of the present invention. Figure 3This is a comparative schematic diagram of a four-dimensional scene generation method based on a geometry and motion joint optimization method according to an embodiment of the present invention. Figure 4 This is a schematic diagram of a four-dimensional scene generation device that combines geometry and motion optimization according to an embodiment of the present invention. Figure 5 It is a computer device according to an embodiment of the present invention. Detailed Implementation
[0023] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0024] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0025] The following description, with reference to the accompanying drawings, describes a method, apparatus, device, and storage medium for generating four-dimensional scenes using joint optimization of geometry and motion, according to embodiments of the present invention.
[0026] The core idea of this invention is to integrate depth information, motion priors, and dynamic rendering by constructing a four-dimensional scene generation framework that jointly optimizes geometry and motion. First, based on the depth information of the input image, a depth-guided motion normalization strategy is used to perform scale-invariant representation processing on the motion trajectory of the 3D point cloud, fundamentally eliminating scale bias in motion perception of objects at different depths. Then, a motion perception module extracts latent motion features from the static image and uses a motion perception adaptive normalization mechanism to inject feature parameters into each layer of the diffusion model network, achieving dynamic guidance of trajectory generation based on motion prior knowledge. Finally, based on the jointly optimized four-dimensional point cloud representation, a view synthesis module renders dynamic video along arbitrary camera trajectories and repairs and completes occluded areas generated during rendering. This unifies the traditionally separate scene reconstruction and motion generation processes into a tightly coupled optimization system, achieving end-to-end generation from a single static image to a multi-view, spatiotemporally consistent four-dimensional dynamic scene. This significantly improves the physical rationality, visual coherence, and viewpoint consistency of dynamic scene generation, providing reliable technical support for applications such as virtual reality and simulation.
[0027] Example 1 To achieve the above invention, embodiments of the present invention provide a four-dimensional scene generation method that combines geometry and motion optimization, such as... Figure 1 As shown, it includes: S1, based on the depth information of the input image, uses a depth-guided motion normalization strategy to perform scale-invariant representation processing on the motion trajectory of the 3D point cloud, eliminating the scale bias in motion perception of objects at different depths.
[0028] Specifically, this strategy is based on the depth information of the input image, and measures the three-dimensional motion of the point cloud. Normalization is performed to ensure that objects of different depths have a consistent perception scale during motion generation, thereby improving the stability of the generative model and the accuracy of motion modeling.
[0029] Specifically, this normalization process depends on the camera's focal length parameter. , and image resolution Introducing a scaling factor and This is used to map the motion of a point to the relative scale of the view frustum. For each point, the initial depth... Its lateral and longitudinal scales in the visual cone are respectively and Therefore, the formula for normalizing the motion of a point is: ; in, , , This represents the normalized motion magnitude. This processing method inversely proportionalizes the motion magnitude of a point to its depth, thereby maintaining consistency in motion perception across different depths. For example, even if a distant point undergoes a large 3D displacement, its projection on the image plane may change only slightly. However, this normalization strategy amplifies its motion magnitude, giving it a perceptual scale similar to that of nearby points in the model.
[0030] Furthermore, this step plays a crucial role in the four-dimensional scene trajectory generator. By converting motion trajectories into scale-invariant representations, the model can more effectively learn the motion patterns of objects at different depths, avoiding training instability caused by depth differences. This strategy is particularly suitable for generating four-dimensional scenes containing multi-scale object motion from a single image, such as in virtual reality or augmented reality applications, ensuring that foreground and background objects have a unified perceptual logic and visual coherence in their dynamic performance.
[0031] Furthermore, S1 includes: S11, based on the focal length of the input image , and image size Calculate the scaling factor and .
[0032] Specifically, this step aims to map the motion of points in the image from pixel space to a normalized three-dimensional space, thereby eliminating scale bias caused by differences in camera parameters and image resolution, and providing a consistent input representation for subsequent motion generation.
[0033] Specifically, this step first reads the image's intrinsic parameter information, including the horizontal focal length. and vertical focal length and the width of the image and height In practical applications, these parameters are typically derived from camera calibration results or image metadata (such as EXIF information). Then, using a simple division operation, the scaling factors in the horizontal and vertical directions are calculated respectively. and These two scaling factors reflect the projection scale of each pixel in the image in the view frustum space, and are the key bridge for converting 2D pixel displacement into 3D spatial displacement.
[0034] Furthermore, and It is usually expressed in pixels, representing the camera's projected focal length on the image plane. Its value ranges from hundreds to thousands of pixels, depending on the camera's optical characteristics and image resolution. and The width and height of the image, usually integers, such as... or Scaling factor and The value range is usually in arrive The exact value depends on the degree of matching between the image resolution and the focal length. In this invention, this step requires high computational precision, and floating-point operations are typically used to ensure numerical stability.
[0035] Specifically, this step is widely used in the preprocessing stage of four-dimensional scene generation, especially when reconstructing dynamic point cloud trajectories from a single image. By normalizing pixel displacements to relative motion at the view frustum scale, the model can more accurately perceive the magnitude of motion in different depth regions, thereby avoiding motion perception distortion caused by depth differences. For example, when dealing with complex scenes containing a mixture of foreground and background, this normalization strategy ensures that small movements of foreground objects are not misjudged as violent movements in 2D images, while large movements of background objects are not ignored due to small pixel displacements.
[0036] Specifically, scale invariance of motion representation is achieved by introducing a depth-related normalization factor. In particular, the normalized motion quantity... and Able to reach different depths Maintaining consistent motion perception across the entire model improves the training stability and generation quality of the four-dimensional trajectory generator. This strategy provides a more uniform input distribution for subsequent diffusion models and is a crucial foundation for achieving joint geometry and motion modeling.
[0037] S12, representing the relative motion of the 3D point cloud. Convert to scale-invariant representation , , ;in, The initial depth of the point .
[0038] Specifically, this step is one of the core implementations of the depth-guided motion normalization strategy of this invention, which aims to solve the problem of scale inconsistency caused by depth differences during motion generation, thereby improving the learning efficiency and generation stability of the diffusion model for motion patterns.
[0039] Furthermore, in the 3D point cloud representation, the motion trajectory of each point... Due to time coordinate offset on Composition. Since the projection scales of points at different depths on the image plane vary significantly, directly using the original motion data for modeling will make it difficult for the model to learn a unified motion pattern. Therefore, this invention introduces a normalization strategy based on the view frustum scale. Specifically, it utilizes the image's focal length... , With image size Calculate the scaling factor , This reflects the projection scale of the view frustum in the horizontal and vertical directions. The initial depth of each point... Used to normalize its motion in each direction, thereby , , Mapped to a depth-independent scale-invariant representation , , .
[0040] Furthermore, during the normalization process, key parameters include focal length. , Image width and height and the initial depth of the point Scaling factor , This value is typically determined during the camera calibration phase, and its range depends on the specific imaging device's parameter configuration. For example, in a standard camera setting, , ,but Normalized exercise volume , , It is usually limited to the range of [-1, 1] to fit the input requirements of the diffusion model.
[0041] Specifically, this step is widely used in the motion generation stage of four-dimensional scene generation, especially in point cloud trajectory prediction tasks based on diffusion models. By normalizing the motion of points into a depth-independent representation, the model can more effectively learn the motion patterns of objects at different distances, avoiding motion perception biases caused by depth differences. For example, in virtual reality scenes, small movements of near-field objects may appear as significant displacements in the image, while large movements of distant objects may be almost invisible. Through the normalization process in this step, the model can uniformly model these motions, thereby generating geometrically consistent and dynamically reasonable four-dimensional scenes from multiple perspectives.
[0042] Specifically, in actual testing, this strategy effectively mitigated the motion prediction distortion caused by depth variations, ensuring the stability of the generated point cloud trajectory across different viewpoints and avoiding the geometric collapse or discontinuous motion issues common in traditional methods. Furthermore, this normalization method provides a standardized input representation for the subsequent Motion Perception Module (MPM) and Motion Perception Adaptive Normalization (MAdaNorm), enhancing the overall system's collaborative optimization capabilities and serving as a key technological support for achieving high-quality four-dimensional scene generation.
[0043] S2 extracts potential motion features from static images through a motion perception module and injects the feature parameters into each layer of the diffusion model network using a motion perception adaptive normalization mechanism, thereby realizing the dynamic guidance of motion prior knowledge for trajectory generation.
[0044] Specifically, this step is crucial for achieving spatiotemporal consistency and dynamic fidelity in the entire four-dimensional generation system. At the technical implementation level, the MPM module employs a feature extractor pre-trained on an image-video joint dataset to extract motion-potential regions from the image. Specifically, this module extracts block-level motion-aware features from the input image using a convolutional neural network (CNN) or visual transformer (ViT) structure. Its dimensions are usually ,in , This represents the spatial resolution of the feature map. This represents the number of feature channels. To ensure compatibility with diffusion models, motion features... Feature alignment is achieved by adjusting the length of the word sequence in the diffusion transformer to match the length of the word sequence in the spatial interpolation or upsampling operation.
[0045] Furthermore, the MAdaNorm mechanism performs fine-grained feature modulation in each transformer layer of the diffusion model. Specifically, aligned motion features are modulated through linear layers. Mapped to token-level adaptive parameters ,in This represents the number of eigenvectors in the diffusion model. These parameters are used for the feature dimensions. The modulation process involves normalization and scaling operations, and the modulation procedure is as follows: ; in, For learnable global gating coefficients, Presentation layer normalization operation, This represents element-wise multiplication. Through this mechanism, the model can dynamically adjust its internal feature representation based on the underlying motion semantics in the image, thereby achieving adaptive guidance of the motion generation process.
[0046] Specifically, this step is widely applicable to generating spatiotemporally consistent four-dimensional dynamic scenes from a single image, and is particularly valuable in virtual reality (VR), augmented reality (AR), and digital content creation. By injecting motion priors, the model can generate more natural and complex motion trajectories, such as object deformation and material changes, rather than being limited to rigid motion.
[0047] Specifically, experiments show that the combination of MPM and MAdaNorm can effectively enhance the model's ability to perceive dynamic details, reduce the discontinuity and irrationality of motion trajectories, and thus generate higher quality four-dimensional point cloud trajectory representations.
[0048] Furthermore, S2 includes: S21 employs a feature extractor pre-trained on joint image-video data as a motion perception module to extract block-level motion features from static images. .
[0049] Specifically, the module is pre-trained on a large-scale image-video joint dataset and is able to capture potential motion cues in images, such as the motion trends of objects, the distribution of dynamic regions, and temporal semantic change patterns.
[0050] Specifically, MPM is implemented as follows: First, a feature extraction network (such as OmniMAE) trained on joint image-video data is used to encode features of the input image. This network learns the semantic relationships between the image and the video, enabling it to identify regions in the image that may change over time. During feature extraction, the model outputs block-level features. ,in This indicates the number of blocks into which the image is divided. The feature dimensions for each block. Block division typically employs a sliding window or image segmentation strategy; the block size can be set to... or Pixels, to balance spatial resolution and computational efficiency.
[0051] Furthermore, in order to incorporate block-level motion characteristics Injected into each layer of the diffusion transformer, this invention introduces a motion-aware adaptive normalization (MAdaNorm) mechanism. This mechanism uses a linear transformation from... Generate token-level scaling parameters and bias parameters And apply it to the intermediate features of the diffusion converter. The normalization and modulation process. The specific modulation formula is as follows: ; in, For learnable global gating coefficients, Presentation layer normalization operation, This represents element-wise multiplication. Through this mechanism, the model can dynamically adjust the feature distribution based on the motion potential of different regions in the image, thereby enhancing the spatiotemporal consistency and detail fidelity of motion generation.
[0052] Specifically, MPMs are typically deployed in the image preprocessing stage to extract motion semantic features and serve as conditional inputs to the diffusion model. Their output features... This can be used to guide point cloud trajectory generators in motion prediction at different time steps, and is especially suitable for scenarios that require generating complex, spontaneous motion, such as changes in human facial expressions, liquid flow, or deformation of flexible objects. This step significantly improves the model's ability to understand dynamic content, providing crucial motion perception support for four-dimensional scene generation.
[0053] S22, generates token-level adaptive parameters through a linear layer. , , , and according to the formula , The intermediate features of the diffusion model are modulated; whereby, , These are learnable global gating coefficients.
[0054] Specifically, this step is a core component of the motion-aware adaptive normalization mechanism (MAdaNorm), which enhances the model's ability to perceive potential dynamic regions in static images and improves the spatiotemporal consistency of the generated four-dimensional scene.
[0055] Specifically, the motion sensing module (MPM) first extracts block-level motion features from the input image. This feature is typically generated by a feature extractor pre-trained on an image-video joint dataset. Then, Mapped to token-level scaling parameters via a lightweight linear layer. , and bias parameters , Its dimension and intermediate features of the diffusion converter To remain consistent, among which Indicates the number of tokens. These parameters represent the feature dimensions. They are used for token-by-token adaptive modulation of intermediate features.
[0056] Specifically, in the diffusion converter's... Within each block, the intermediate features Perform the following modulation operation:
[0057]
[0058] in, Presentation layer normalization operation, This represents element-wise multiplication. and These are the attention mechanism and the multilayer perceptron module, respectively. , The gating coefficients are globally learnable, and their dimensions are... This is used to adjust the intensity of the influence of adaptive parameters on features.
[0059] Furthermore, linear layers typically employ a fully connected structure, with the number of input channels varying depending on the motion characteristics. The dimensions are consistent, and the number of output channels is , respectively corresponding , , , Global gating coefficient , Automatic learning is achieved through model training, with the initial value typically set to 1 to ensure a smooth transition in the modulation process.
[0060] Specifically, in practical applications, this step is mainly used to enhance the diffusion model's ability to perceive potential dynamic regions in static images. For example, when generating dynamic scenes containing people or animals, the MPM can identify regions with motion potential and adjust their feature representation through an adaptive normalization mechanism, thereby guiding the model to generate motion trajectories that are more semantically consistent. Furthermore, this mechanism can also be used to control the intensity and direction of motion generation, achieving fine-grained control over dynamic content.
[0061] Specifically, this step significantly improves the expressive power of the diffusion model in motion generation tasks by introducing token-level adaptive parameters and a global gating mechanism. Experiments show that this method can effectively enhance the model's accuracy in modeling dynamic details, improve the spatiotemporal consistency and visual realism of the generated four-dimensional content, and provide key support for generating high-quality four-dimensional dynamic scenes from single images.
[0062] S3, based on the jointly optimized four-dimensional point cloud representation, renders dynamic video along arbitrary camera trajectories through the view synthesis module, and repairs and completes the occluded areas generated by the rendering, generating a four-dimensional dynamic scene with consistent spatiotemporal perspectives.
[0063] Specifically, this step is a key step in the present invention to generate high-quality four-dimensional dynamic content from a single static image. Its core lies in transforming the abstract point cloud trajectory into a dynamic video sequence with visual coherence and geometric consistency.
[0064] Specifically, the view composition module is based on four-dimensional point cloud representation. ,in Indicates the number of time frames. This indicates the number of points in the point cloud. and These represent the height and width of the image, respectively. This module first projects the point cloud frame-by-frame using a 3D rendering engine (such as Gaussian Splatting) based on any camera trajectory specified by the user, generating a preliminary dynamic video sequence. Since the point cloud projection may not cover all pixel areas from different viewpoints, an occlusion mask is generated during the rendering process to identify missing or invisible areas in the video frames.
[0065] Furthermore, the generation of the occlusion mask relies on the visibility determination of the point cloud projection, typically employing depth buffering or frustum culling techniques. During training, the mask value for the occlusion mask is set to 0.5, indicating that the region needs to be filled in by the generative model. The video generation model learns during fine-tuning how to generate reasonable pixel content in occluded regions while maintaining temporal motion coherence. The resolution of the rendered output is usually consistent with the input image, i.e. And time frame count The frame rate can be set to 16, 32, or 64 frames depending on the length of the target video.
[0066] Specifically, this step is widely applicable to fields such as virtual reality (VR), augmented reality (AR), digital twins, and film and television special effects production. For example, in VR scenes, users can freely set the camera path, and the system will render and fill in occluded areas in real time based on the four-dimensional point cloud trajectory, generating continuous, artifact-free dynamic video, thereby achieving an immersive interactive experience. In film and television production, this technology can be used to generate multi-view dynamic scenes with complex motion from a single background image, reducing the reliance on multi-view shooting.
[0067] Specifically, this step significantly improves the visual quality and spatiotemporal consistency of the generated video through joint optimization of view synthesis and occlusion repair. Specifically, the rendering process ensures the consistency of the geometric projection of the four-dimensional point cloud under different viewpoints, while the occlusion repair mechanism effectively fills in missing areas by introducing a pre-trained video generation model, avoiding the holes and artifacts caused by viewpoint switching in traditional methods. The final output video sequence possesses reasonable structure, natural motion, and visual coherence under any novel viewpoint, achieving high-quality generation from a single image to an interactive four-dimensional dynamic scene.
[0068] Furthermore, S3 includes: S31 sets the value to 0.5 for areas in the rendered video not covered by projection points.
[0069] Specifically, the technical implementation of this step is based on the geometric relationship between point cloud projection and image rendering, combined with the masking mechanism of the video generation model, to ensure that missing areas can be reasonably repaired under the new perspective.
[0070] Specifically, this step first obtains the depth map of the input image using a depth estimation method, and then performs a four-dimensional point cloud analysis based on the camera trajectory. The system performs frame-by-frame projection to generate the rendered image for each frame. Since the projection coverage of the point cloud is limited at different viewpoints, some pixel areas may not be covered by any point cloud projection during rendering, resulting in holes or missing areas. In this case, the system determines whether a pixel is covered by the point cloud and generates a binary or floating-point occlusion mask. The mask value of 1 indicates that the pixel is covered by the point cloud, and 0 indicates that it is not covered at all.
[0071] Optionally, the setting of the value 0.5 for regions not covered by projection points is not random, but based on the training strategy of the video generation model. During the training phase, the four-dimensional view synthesis module uses a fine-tuned video generation model, whose inputs include rendered images, occlusion masks, real video, and text descriptions. By learning how to generate reasonable content in regions with a value of 0.5, the model achieves high-quality completion of missing regions under new perspectives during the inference phase. This mask value setting helps the model distinguish between completely occluded and partially occluded regions, avoiding the generation of unreasonable backgrounds or noise in completely uncovered areas.
[0072] Furthermore, the occlusion mask settings must meet the following technical specifications: the mask resolution must be consistent with the rendered video, typically [resolution value missing]. The mask value ranges from [0, 1], and the mask update frequency is consistent with the video frame rate to ensure temporal continuity. In practical applications, this step is widely used for free-viewpoint video generation in VR / AR scenarios, especially when dealing with complex geometric structures or dynamic occlusion relationships, which can significantly improve the visual integrity and immersion of the video.
[0073] Specifically, by setting the occlusion mask value appropriately, the video generation model is guided to generate dynamic content consistent with the context in the missing area, thereby effectively eliminating the hollow artifacts caused by changes in viewpoint and improving the visual quality and temporal consistency of the four-dimensional scene.
[0074] S32 repairs occluded areas by fine-tuning the video generation model, ensuring visual coherence in both time and space dimensions after repair.
[0075] Specifically, the core objective of this step is to use a video generation model to fill in any gaps that may appear in the new perspective video generated by point cloud rendering, thereby ensuring the visual coherence of the repaired video in both the temporal and spatial dimensions.
[0076] Specifically, this step first involves the generated four-dimensional point cloud representation. A preliminary video sequence is generated by a renderer along arbitrary camera trajectories. Since point clouds cannot cover the entire scene from certain viewpoints, occluded areas appear in the video. To address this issue, this invention employs a fine-tuned video generation model to repair these areas. This model is pre-trained on an image-video joint dataset and further fine-tuned in a four-dimensional view synthesis task to adapt to the video structure generated from point cloud rendering. During fine-tuning, the model input includes rendered video, occlusion masks, real video, and text descriptions, where the occlusion mask indicates areas not covered by the point cloud in each frame. Specifically, for areas without projection points in each frame, the mask value is set to 0.5 to guide the model to focus on and repair these areas.
[0077] Furthermore, the occlusion mask is set using a strategy combining binarization and grayscale; a mask value of 0.5 indicates that the area is a transitional region to be repaired, rather than a complete loss. The input resolution of the video generation model is typically [resolution missing]. or To adapt to the accuracy requirements of different application scenarios, the model outputs videos at a frame rate of typically 24 FPS or 30 FPS to ensure smooth transitions over time. Furthermore, the model employs a joint optimization strategy of adversarial loss and reconstruction loss during training. The adversarial loss enhances the visual realism of the generated video, while the reconstruction loss maintains structural consistency with the real video.
[0078] Specifically, this step is widely applicable to fields such as virtual reality (VR), augmented reality (AR), and digital content creation. For example, in a VR scene, users can move freely along any path, and the system needs to generate unobstructed, visually coherent dynamic video in real time. Through the repair mechanism in this step, visual breaks or artifacts caused by perspective switching can be effectively avoided, improving the continuity and realism of the immersive experience.
[0079] Specifically, by introducing a fine-tuned video generation model, high-quality completion of occluded areas is achieved, significantly improving the temporal and spatial coherence of the video. Simultaneously, this method avoids the structural distortion or motion inconsistency problems commonly found in traditional restoration methods, providing a reliable guarantee for generating interactive four-dimensional dynamic scenes from single images.
[0080] S4 represents the generated four-dimensional point cloud. By jointly optimizing with dynamic semantic constraints input by the user, and adjusting the physical rationality parameters of the motion trajectory, the dynamic behavior of the generated scene is made consistent with the user's needs.
[0081] Specifically, this step is widely applicable to the fields of virtual reality (VR), augmented reality (AR), and digital content creation. It ensures that the generated point cloud trajectory is smooth in time and conforms to physical laws in space, thereby generating high-quality, dynamic videos that match the user's intent.
[0082] Specifically, this step significantly improves the spatiotemporal consistency and user controllability of four-dimensional point cloud trajectories by introducing a joint optimization mechanism of physically reasonable parameters and semantic constraints. Experiments show that this method can improve performance in motion video generation tasks. The motor consistency score, and improve The aesthetic score provides a key guarantee for generating high-quality four-dimensional dynamic scenes from a single image.
[0083] This invention discloses a geometrically and motion-optimized four-dimensional scene generation method. Through a depth-guided motion normalization strategy, it effectively eliminates scale perception bias in the motion of objects at different depths. Furthermore, it utilizes a motion-aware adaptive normalization mechanism to dynamically guide trajectory generation based on motion priors, thus solving the core problems of motion representation distortion and insufficient spatiotemporal consistency in existing technologies. This method automates the entire process from static image input to four-dimensional dynamic scene generation, significantly improving the physical plausibility of motion trajectories and the visual coherence of rendered videos. While generating consistent dynamic scenes from multiple perspectives, it also enhances user interactivity, providing a highly realistic four-dimensional scene generation solution for complex applications such as virtual reality and simulation.
[0084] Example 2 To achieve the above invention, embodiments of the present invention also provide specific steps for a four-dimensional scene generation method based on joint optimization of geometry and motion, such as... Figure 2 As shown, it includes: Given a single input image I∈R (H×W×3) The goal is to reconstruct a physically plausible and temporally consistent four-dimensional scene, described by its textual description C. The scene is represented as a point cloud sequence P∈R. (T×N×3) This captures the 3D geometry and motion trajectories of N = H × W points on T frames. From this representation, dynamic scene video V∈R is rendered from any new perspective. (T×H'×W'×3) This system enables the transformation from a single static image to complete multi-view four-dimensional dynamic generation. It comprises two core components: a four-dimensional scene trajectory generator and a four-dimensional view synthesis module. Unlike existing decoupled paradigms, this method tightly integrates motion generation and geometric reconstruction, ensuring the inherent consistency between the inferred dynamics and the evolving 3D structure. By co-optimizing structure and motion, it generates a more coherent and stable four-dimensional representation, avoiding the problem of geometric reconstruction errors propagating to motion estimation in traditional methods. Specifically, it includes: S101: Deep-guided motion normalization strategy.
[0085] Specifically, to enhance training stability and ensure compatibility with the scale of generative models, the four-dimensional scene trajectory generator only predicts the relative motion ΔP. t = P t - P0 = {[Δx t , Δy t , Δz t ]}, where t∈[0,T], and P0 represents the coordinates of the first frame. To avoid an unlimited range of data values, ΔP is further optimized. t Normalization is performed.
[0086] Furthermore, given that small 3D movements of nearby objects produce large displacements on the 2D image plane, while the same movements of distant objects appear minuscule, we propose a depth-guided motion normalization strategy. This strategy normalizes the absolute motion of each point based on the view frustum at the initial depth, converting the original motion into a scale-invariant representation. This ensures perceptual consistency across different distances and provides a more uniform data distribution for the generative model. Specifically, given a focal length f... x f y And image size W×H, introduce a scaling factor = f x / W and = f y / H. Geometrically, z / and z / These correspond to the width and height of the view frustum at depth z, respectively. (Through depth z = ...) The size of the visual cone at a given point is normalized to the motion quantity: This depth-dependent normalization achieves scale invariance across different depth ranges, enabling the diffusion model to effectively learn motion patterns without being biased by the absolute spatial location of points.
[0087] S102: Architecture design of a four-dimensional scene trajectory generator.
[0088] Specifically, based on the existing image-to-video diffusion model architecture, its spatiotemporal variational autoencoder and diffusion transformer modules are fine-tuned respectively. During training, the model accepts image-capture pairs as input and learns the diffusion model to predict pixel-level point trajectories.
[0089] Specifically, the motion-sensitive variational autoencoder first needs to be fine-tuned to enable it to process trajectory signals. The relative point displacement ΔP tA lightweight trajectory encoder converts the image to an RGB motion map, where spatial motion is represented by color changes, while shape and appearance remain static as in the first frame. Accordingly, a trajectory decoder is appended after the variational autoencoder decoder to ensure accurate recovery of point trajectories from the RGB motion map. After fine-tuning the variational autoencoder, an adaptive diffusion transformer processes the latent representation encoded from the RGB motion map. To explicitly inject strong geometric priors into the model, depth information from the initial frame is encoded into a latent representation using a variational autoencoder, providing the model with robust structural priors and geometric cues, significantly improving its ability to infer scene layout and object relationships. The image, noise, and depth latent representation are concatenated along the feature dimension to form the final input to the diffusion transformer: z. 合并 = splicing (z 图像 , z 噪声 , z 深度 We employ flow matching to train the Diffusion Transformer model. By minimizing the error between the predicted and actual flow fields, we learn a deterministic flow from noise to data distribution, achieving accurate pixel-level motion modeling. The objective function is defined as: L = E t,x0,x1 [||v θ (t,x t ) - (x1-x0)||²], where x0 and x1 represent the initial and target states respectively, x t v represents the interpolated state at time t. θ It is the flow vector predicted by the model.
[0090] S103: Motion sensing module and adaptive normalization mechanism.
[0091] Specifically, adapting the robust motion priors of existing image-to-video diffusion models presents a significant challenge. These models excel at generating dynamic moving scenes from static images, while our approach aims to leverage this pre-trained temporal knowledge to handle a different task: generating videos that are structurally static but exhibit dynamic color changes. This creates a core conflict, as the model must learn to translate its inherent understanding of physical displacement into a new domain of temporal color evolution.
[0092] Furthermore, to effectively guide this adaptation and enable the model to identify candidate regions containing color dynamic semantics within static scenes, we introduce a motion-aware module (MPM). A feature extractor pre-trained on joint image-video data is used as the motion feature extractor to obtain motion-aware block-level features S from static images. To embed motion information into the diffusion process, a motion-aware adaptive normalization (MAdaNorm) is introduced, performing fine-grained feature modulation in each diffusion transformer layer. Specifically, after spatially aligning the motion features S by resizing to match the Diffusion Transformer token sequence length, token-level adaptive parameters are generated through linear layers. For intermediate features in the i-th diffusion transformer block... ∈R (N×d) The modulation process is as follows: First, the adaptive parameters are calculated: , , , = Linear(S), where , , , ∈R (N×d) These are token-level scaling and bias parameters. Then feature modulation is performed: ; in, , ∈R d These are learnable global gating coefficients, and ⊙ represents word-level multiplication. Experimental results show that MPM significantly improves the fidelity of motion trajectory modeling and enhances the spatiotemporal consistency of generated four-dimensional content.
[0093] S104: Design and Implementation of a Four-Dimensional View Synthesis Module
[0094] Specifically, after obtaining a dense four-dimensional point cloud representation, a four-dimensional view synthesis module is proposed to achieve new perspective video synthesis along arbitrary camera trajectories. The original point cloud may not completely cover the image region in the new perspective, resulting in holes in the rendered output. Therefore, a generative model is used to fill in these missing regions, ensuring visual coherence and plausibility.
[0095] Furthermore, considering the superior performance of existing video generation models in video generation tasks, a four-dimensional view synthesis module was constructed by fine-tuning the model. Training data included rendered videos, corresponding occlusion masks, real videos, and titles. During training, a masking strategy was followed, setting the value to 0.5 for regions without projection points in each frame. Through fine-tuning, the model achieved high-quality and visually consistent new-view video synthesis, effectively handling occlusion regions while maintaining temporal coherence. In the inference phase, depth estimation methods were first used to obtain depth information from the input images, ensuring consistency with the trajectory tracking model settings. The relative motion map was de-normalized and fused with the initial point cloud to form a complete four-dimensional scene representation. This representation can be used to render dynamic scene videos from any viewpoint and along any camera path, bridging the gap from static images to fully multi-view four-dimensional dynamic generation.
[0096] The above steps, through the tight integration of motion generation and geometric reconstruction, ensure that the inferred dynamics are not only reasonable but also intrinsically consistent with the evolving 3D structure. By co-optimizing structure and motion, a more coherent and stable 4D representation is generated, effectively avoiding the error accumulation problem of traditional decoupling methods. The depth-guided motion normalization strategy achieves scale-invariant motion representation, significantly improving the learning efficiency and stability of the generative model. The motion perception module and adaptive normalization mechanism effectively utilize the motion priors of the pre-trained model, significantly improving the fidelity of motion trajectory modeling and the spatiotemporal consistency of the generated 4D content. The entire method represents a technological breakthrough from single static images to complete multi-view 4D dynamic generation.
[0097] Furthermore, existing technologies generally suffer from the fundamental problem of decoupling geometric reconstruction and motion generation, resulting in either geometric inconsistencies across multiple viewpoints or severe limitations in dynamic performance. This technological bottleneck hinders the further development of generating high-quality four-dimensional scenes from single static images. Therefore, there is an urgent need for an innovative technical solution that can tightly couple geometric reconstruction and motion generation to achieve the generation of high-quality four-dimensional scenes from single static images that possess both geometric consistency and rich dynamic details, providing strong technical support for applications such as VR / AR. A comparison of this invention with existing technologies is provided below. Figure 3 As shown.
[0098] This invention discloses a method for generating four-dimensional scenes through joint optimization of geometry and motion. By leveraging the synergistic effect of a depth-guided motion normalization strategy and a motion-aware adaptive normalization mechanism, it effectively solves the spatiotemporal inconsistencies and motion distortion problems caused by the decoupling of geometric reconstruction and motion generation in existing technologies. It achieves end-to-end generation from a single static image to a multi-view four-dimensional dynamic scene. Through the joint optimization of tightly coupled geometric structures and motion trajectories, it significantly improves the physical plausibility and visual coherence of the generated scene. While enhancing the consistency of multi-view rendering, it ensures the realism and detail fidelity of dynamic content, providing a high-quality and highly adaptable four-dimensional scene generation solution for complex applications such as virtual reality and simulation.
[0099] Example 3 To achieve the above invention, such as Figure 4 As shown, this embodiment also provides a four-dimensional scene generation device 10 with joint optimization of geometry and motion, the device 10 including: The depth-guided motion normalization module 100 is used to perform scale-invariant representation processing on the motion trajectory of the 3D point cloud based on the depth information of the input image and adopt the depth-guided motion normalization strategy to eliminate the scale deviation of motion perception of objects at different depths. The motion-aware feature injection module 200 is used to extract potential motion features from static images through the motion-aware module, and inject the feature parameters into each layer of the diffusion model network using the motion-aware adaptive normalization mechanism, so as to realize the dynamic guidance of motion prior knowledge for trajectory generation. The view composition and occlusion repair module 300 is used to render dynamic video along arbitrary camera trajectories based on the jointly optimized four-dimensional point cloud representation, and repair and complete the occluded areas generated by the rendering to generate a four-dimensional dynamic scene with spatiotemporal consistency from multiple perspectives.
[0100] In one embodiment of the present invention, it further includes: a joint optimization module, used to represent the generated four-dimensional point cloud. By jointly optimizing with dynamic semantic constraints input by the user, and adjusting the physical rationality parameters of the motion trajectory, the dynamic behavior of the generated scene is made consistent with the user's needs.
[0101] This invention discloses a four-dimensional scene generation device that jointly optimizes geometry and motion. Through the coordinated operation of its various functional modules, it effectively solves the problems of scale perception deviation and spatiotemporal consistency caused by the decoupling of geometric reconstruction and motion generation in existing technologies. The device eliminates motion perception differences between objects at different depths based on a depth-guided motion normalization mechanism, utilizes motion perception feature injection to dynamically guide trajectory generation based on motion priors, and ensures visual coherence in multi-view rendering through view synthesis and occlusion repair. The entire device achieves fully automated processing from static image input to four-dimensional dynamic scene generation, significantly improving the physical rationality and visual realism of the generated scene, and providing an efficient and reliable four-dimensional scene generation solution for applications such as virtual reality and digital twins.
[0102] To implement the methods of the above embodiments, the present invention also provides a computer device, such as... Figure 5 As shown, the computer device 600 includes a memory 601 and a processor 602; wherein, the processor 602 reads the executable program code stored in the memory 601 to run a program corresponding to the executable program code, so as to implement the various steps of the geometry and motion joint optimization four-dimensional scene generation method described above.
[0103] To implement the above embodiments, this application also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a four-dimensional scene generation method with joint optimization of geometry and motion as described in the foregoing embodiments.
[0104] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0105] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
Claims
1. A method for four-dimensional scene generation with joint geometry and motion optimization, characterized in that, Comprise: S1, based on the depth information of the input image, using depth guided motion normalization strategy to process the scale invariant representation of three-dimensional point cloud motion trajectory, eliminate the scale deviation of different depth object motion perception; S2, through the motion perception module to extract potential motion features from static images, and use motion perception adaptive normalization mechanism to inject feature parameters into each layer network of diffusion model, realize the dynamic guidance of motion prior knowledge to trajectory generation; S3, according to the joint optimization of four-dimensional point cloud representation, through view synthesis module along any camera trajectory to render dynamic video, and repair the occlusion area generated by rendering, generate multi-view spatiotemporal consistent four-dimensional dynamic scene.
2. The method of claim 1, wherein, Based on the depth information of the input image, using depth guided motion normalization strategy to process the scale invariant representation of three-dimensional point cloud motion trajectory, eliminate the scale deviation of different depth object motion perception, comprising: S11, calculating a scaling factor from the focal length of the input image , and the image size and ; S12, converting the relative motion of the three-dimensional point cloud to a scale-invariant representation , , ; wherein, is the initial depth of the point .
3. The method of claim 1, wherein, Through the motion perception module to extract potential motion features from static images, and use motion perception adaptive normalization mechanism to inject feature parameters into each layer network of diffusion model, realize the dynamic guidance of motion prior knowledge to trajectory generation, comprising: S21, using a feature extractor pre-trained on image-video joint data as a motion-aware module to extract block-level motion features from the static image ; S22, generating token-level adaptive parameters by linear layer 、 、 、 and modulating the intermediate features of the diffusion model according to the formula 、 wherein 、 are learnable global gating coefficients.
4. The method of claim 1, wherein, According to the joint optimization of four-dimensional point cloud representation, through view synthesis module along any camera trajectory to render dynamic video, and repair the occlusion area generated by rendering, generate multi-view spatiotemporal consistent four-dimensional dynamic scene, comprising: S31, set the area not covered by the projected point in the rendered video to 0.5; S32, repair the occlusion area through the fine-tuned video generation model, ensure the visual coherence of the repaired video in time and space dimensions.
5. The method of claim 1, wherein, Also include: S4, generating a four-dimensional point cloud representation Joint optimization with dynamic semantic constraints input by users, through adjusting the physical rationality parameters of the motion trajectory, the consistency of the dynamic behavior of the generated scene and the user's demand is realized.
6. A device for generating a four-dimensional scene with joint optimization of geometry and motion, characterized in that Comprise: Depth guided motion normalization module, for based on the depth information of the input image, using depth guided motion normalization strategy to process the scale invariant representation of three-dimensional point cloud motion trajectory, eliminate the scale deviation of different depth object motion perception; Motion perception feature injection module, for through the motion perception module to extract potential motion features from static images, and use motion perception adaptive normalization mechanism to inject feature parameters into each layer network of diffusion model, realize the dynamic guidance of motion prior knowledge to trajectory generation; View synthesis and occlusion repair module, for according to the joint optimization of four-dimensional point cloud representation, through view synthesis module along any camera trajectory to render dynamic video, and repair the occlusion area generated by rendering, generate multi-view spatiotemporal consistent four-dimensional dynamic scene.
7. The apparatus of claim 6, wherein, Also include: A joint optimization module is configured to jointly optimize the generated four-dimensional point cloud representation The dynamic semantic constraints input by the user are jointly optimized, and the consistency between the dynamic behavior of the generated scene and the user demand is realized by adjusting the physical reasonableness parameters of the motion trajectory.
8. An electronic device comprising: a processor; a memory storing executable instructions; the processor executes the instructions to implement a geometric and motion joint optimization four-dimensional scene generation method according to any one of claims 1-5.
9. A computer readable storage medium storing a computer program, the computer program being executed by a processor to implement a geometric and motion joint optimization four-dimensional scene generation method according to any one of claims 1-5.