Multi-modal 4D content generation method and system based on alignment

A multimodal 4D content generation method designed with focal length search and alignment loss solves the problems of insufficient generation efficiency and quality in existing technologies, and realizes efficient and diversified 4D asset generation, which is suitable for virtual reality, augmented reality and other fields.

CN120672972AActive Publication Date: 2025-09-19ZHEJIANG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511179418.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-09-19
Estimated Expiration
2045-08-22

AI Technical Summary

Technical Problem

Existing multimodal 4D content generation technologies have significant deficiencies in data acquisition, input modality processing, generation diversity, multimodal information fusion, and generation efficiency, making it difficult to generate diverse and dynamic 4D assets efficiently and with high quality.

Method used

Accurate video alignment focal length and multi-view alignment focal length are obtained through focal length search. Combined with motion alignment loss and geometric alignment loss, a total loss function is designed for asynchronous optimization to generate high-quality 4D asset models.

Benefits of technology

It achieves efficient generation of 4D assets faithful to the input conditions from any single modal input, reduces data acquisition costs, improves view-motion alignment accuracy, and improves generation efficiency and quality, making it suitable for industrial-grade applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672972A_ABST
    Figure CN120672972A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal 4D content generation method and system based on alignment, and belongs to the field of computer vision processing. According to the method, firstly, single-mode input is converted into a video-3D model pair through a diffusion model, and a 4D model is initialized according to the video-3D model pair; a video alignment focal length and a multi-view alignment focal length are obtained, and precise space-time registration is realized through two-stage focal length alignment; minimizing a total loss function formed by known / unknown view angle-moment alignment loss by adopting an asynchronous optimization strategy of alternately optimizing a 3D Gaussian model and a deformation network through odd and even steps; the finally generated 4D asset model can output rendered images or continuous videos of any view angle-moment, and high-quality multi-mode 4D content generation is achieved. According to the method, various inputs can be flexibly processed, and the 4D asset model faithful to the inputs is efficiently generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision processing, and in particular relates to an alignment-based multimodal 4D content generation method and system. Background Art

[0002] With the rapid development of artificial intelligence and computer graphics technologies, the generation of three-dimensional (3D) content has become a research hotspot. However, static 3D models cannot fully express the dynamic changes of objects and scenes in the real world. Therefore, four-dimensional (4D) content (i.e., dynamic 3D content) that can capture and generate temporal information has become increasingly important. 4D content has broad application prospects in a wide range of fields, including virtual reality (VR), augmented reality (AR), digital humans, film production, game development, metaverse construction, and robotic simulation.

[0003] Currently, 3D / 4D content generation faces the following major challenges and limitations: 1) High complexity in data acquisition and modeling: Traditional 4D content generation typically relies on multi-sensor capture (e.g., multi-camera arrays, depth sensors) and complex reconstruction algorithms. These methods are costly, time-consuming, and labor-intensive, with stringent requirements for the environment and equipment, making them difficult to popularize. 2) Text-based 3D / 4D generation: While research on generating 3D content from textual descriptions (e.g., text-to-3D) has made progress, directly generating high-fidelity, dynamic, and controllable 4D assets from text remains a significant challenge, primarily due to the limited semantic richness of textual information, which makes it difficult to fully define complex dynamic scenes. 3) Image-based 3D / 4D generation: 3D reconstruction from a single or small number of images (e.g., NeRF and 3D Gaussian Splatting) has become increasingly mature, but extending these approaches to dynamic 4D generation often requires multi-view or time-series image input and faces difficulties in maintaining temporal consistency and dynamic details. Furthermore, inferring dynamic information from static images is inherently an underdetermined problem. 4) Video-Based 3D / 4D Generation: Generating high-quality, editable 4D content directly from a single video remains challenging. Although videos contain temporal information, the dimensionality of the information they provide (2D projection) limits the ability to directly reconstruct complex 3D geometry and dynamics, especially when dealing with occlusion, perspective loss, and motion blur. 5) 4D Generation from Existing 3D Models: While 4D can be generated by deforming or animating existing 3D models, this requires specialized modeling and animation knowledge, and the diversity of the generated content is limited by the quality and type of the initial 3D model. 6) Insufficient Diversity and Controllability of Generated Content: Existing methods often struggle to achieve high diversity and fine-grained control when generating 4D content. The generated results may lack variation in motion, pose, appearance, or scene dynamics, making them difficult to meet the personalized needs of different application scenarios. 7) Inefficient Utilization of Multimodal Information: Many existing methods tend to focus on single-modal input and fail to effectively leverage the complementary information between different modal data (such as the semantics of text, the details of images, and the dynamics of videos). How to effectively fuse information from multiple modalities to drive more realistic and richer 4D generation is a problem that current technology urgently needs to solve. 8) The trade-off between quality and efficiency in generating 4D assets: While pursuing high-fidelity 4D generation, computational efficiency often becomes a bottleneck. For example, although 4D reconstruction based on neural radiance fields (NeRF) can generate high-quality views, the rendering speed and training time are usually long, and it is difficult to edit directly. 9) In recent years, diffusion models have demonstrated powerful capabilities in the fields of image, video, and 3D content generation, and can generate high-fidelity and diverse samples. However, fully applying the advantages of diffusion models to generate dynamic 4D content from arbitrary multimodal inputs and addressing the above limitations remains a challenging frontier research direction.

[0004] In summary, current multimodal 4D generation technologies still have significant shortcomings in data acquisition, input modality processing, generation diversity, multimodal information fusion, and generation efficiency. Therefore, the industry urgently needs a new method that can overcome these shortcomings and achieve efficient, high-quality generation of diverse and dynamic 4D assets from any modality input. Summary of the Invention

[0005] The purpose of the present invention is to solve the problems existing in the prior art and provide a method and system for generating multimodal 4D content based on alignment.

[0006] The inventive concept of this invention is to generate and convert input data from any modality into standard data pairs (video, 3D model). Based on these standard data pairs, the invention generates 4D asset models that are faithful to the input conditions. Specifically, the invention obtains accurate focal length parameters for training through focal length search, proposes motion alignment loss and geometric alignment loss, and designs an unknown moment-view alignment loss based on this. This loss is weighted and summed with the known moment-view alignment loss as the total loss to jointly optimize the 4D model, ultimately generating vivid and diverse 4D asset models.

[0007] In order to achieve the above-mentioned object of the invention, the present invention specifically adopts the following technical solutions:

[0008] In a first aspect, the present invention provides a method for generating multimodal 4D content based on alignment, which comprises the following steps:

[0009] S1: Input data of any single modality is fed into the diffusion model to generate target data pairs consisting of video and 3D model;

[0010] S2: Initialize the 4D model from the 3D model in the target data pair, and use the first frame of the video in the target data pair as the alignment reference image. Render the 4D model within a preset focal length range to obtain the main perspective rendering image corresponding to each focal length at the initial moment. Calculate the mean square error between each main perspective rendering image and the alignment reference image, and use the focal length corresponding to the minimum mean square error as the video alignment focal length.

[0011] S3: Within the focal length range of S2, the focal length that is smaller than the video alignment focal length is retained to form a new focal length range. The 4D model at the initial moment is rendered at different perspectives within the new focal length range. The fractional distillation sampling loss value corresponding to the rendering image of each focal length is calculated at each perspective. The loss values ​​of all perspectives at the same focal length are averaged, and the focal length corresponding to the minimum loss average value is used as the multi-perspective alignment focal length. ;

[0012] S4: Asynchronously optimize the 4D model with the goal of minimizing the total loss function. In the odd-numbered steps of the optimization process, only the 3D Gaussian model in the 4D model is optimized, while in the even-numbered steps, only the deformable network in the 4D model is optimized. The optimized 4D model is used as the final high-quality 4D asset model. The total loss function is formed by the weighted sum of the known moment-view alignment loss and the unknown moment-view alignment loss: for known perspectives and moments, the 4D model is rendered in the main perspective based on the video alignment focal length to generate multi-moment main perspective renderings, and the known moment-view alignment loss is calculated from the multi-moment main perspective renderings and the videos in the target data. For unknown perspectives and moments, the 4D model is rendered in non-main perspectives based on the multi-view alignment focal length to generate multi-moment multi-view renderings. By aligning the multi-moment multi-view renderings with the multi-view diffusion model prior, the motion information and geometric information are transferred to the 4D model to construct the motion alignment loss and the geometric alignment loss. The motion alignment loss and the geometric alignment loss are each weightedly fused with a time-related fusion parameter to obtain the unknown moment-view alignment loss.

[0013] S5: Input the observation angle and time parameters into the high-quality 4D asset model to obtain the corresponding rendered image; input the continuous observation angle and time parameters into the high-quality 4D asset model to obtain multi-time and multi-view rendered videos, completing the alignment-based multimodal 4D content generation.

[0014] Based on the above solution, each step can be implemented in the following preferred specific manner.

[0015] As a preferred embodiment of the first aspect, in step S1, the input data is text, image, video or 3D model.

[0016] As a preferred embodiment of the first aspect above, in step S1, the specific process of generating the video and 3D model in the target data pair is: obtaining input data, when the input data is text, inputting it into the text-video diffusion model to generate a video corresponding to the text, and inputting the first frame of the generated video into the image-3D diffusion model as a control condition to obtain a 3D model corresponding to the generated video; when the input data is an image, inputting it into the image-video diffusion model to generate a video corresponding to the image, and using the input image as a control condition to input it into the image-3D diffusion model to obtain a 3D model corresponding to the input image; when the input data is a video, inputting its first frame as a control condition into the image-3D diffusion model to obtain a 3D model corresponding to the input video; when the input data is a 3D model, using its frontal perspective rendered image as a control condition to input it into the image-video diffusion model to obtain a video corresponding to the input 3D model.

[0017] As a preference of the first aspect above, in step S2, the specific process of obtaining the video alignment focal length is: initializing the initial moment of the 4D model from the 3D model in the target data pair, using the video frame at the initial moment as the alignment reference image for the T-frame video in the target data pair, rendering the 4D model at the same focal length interval within the preset focal length range, obtaining the main perspective rendering image of each focal length at the initial moment, calculating the mean square error between each main perspective rendering image and the alignment reference image, and taking the focal length corresponding to the minimum mean square error as the video alignment focal length .

[0018] As a preferred embodiment of the above-mentioned first aspect, in step S3, the specific process of calculating the fractional distillation sampling loss corresponding to the rendering image of each focal length is: at each viewing angle, the rendering image corresponding to each focal length is input into the encoder respectively to obtain the feature map corresponding to each rendering image, and the randomly sampled Gaussian noise is superimposed on each feature map respectively. With the alignment reference image as the control condition, the fractional distillation sampling loss value is calculated based on the feature map after superimposing the noise to obtain the loss value corresponding to the rendering image at the focal length.

[0019] As a preferred embodiment of the first aspect, in step S4, a multi-moment main-view rendering image is generated by rendering the 4D model at the video alignment focal length and the main view. The multi-moment main-view rendering image is composed of multiple single-moment main-view rendering images, and each single-moment main-view rendering image corresponds one-to-one to each frame of the video in the target data pair. The mask of each frame and the single-moment main-view rendering image is obtained, and the mean square error loss between each frame and the single-moment main-view rendering image is calculated. The mean square error loss between the mask of each frame and the mask of the single-moment main-view rendering image is calculated, and the two parts of the mean square error loss are added as the known moment-view alignment loss.

[0020] Furthermore, given the moment-view alignment loss It is expressed as follows:

[0021]

[0022] in, Indicates the total number of frames in the target data pair, which is equal to the total number of moments; Indicates the main perspective; express Main perspective rendering at all times; express Mask; Indicates the first frame; express Mask; Represents the square of the L2 norm.

[0023] As a preferred embodiment of the first aspect, in step S4, a multi-time multi-view rendering image is generated by rendering the 4D model in multiple perspectives with aligned focal lengths and non-primary perspectives. The multi-time multi-view rendering image is composed of single-time single-view rendering images at different perspectives at different times. The single-view rendering at the moment is multiplied by a first weight coefficient related to the timestamp, and the randomly sampled Gaussian noise is multiplied by a second weight coefficient related to the timestamp. The single-view rendering image and the weighted Gaussian noise are added to form Add noise to the rendering at any time, The noise rendering image and the target data are aligned with the video frame, The five parts of the perspective corresponding to the single-view rendering at the moment, the multi-view alignment focal length, and the timestamp randomly sampled in the multi-view diffusion model are input into the U-Net network of the multi-view diffusion model, and the noise prediction result is output; The mean square error between the noise prediction results and the randomly sampled Gaussian noise at all views at the moment is used as the action alignment loss.

[0024] Furthermore, the action alignment loss It is expressed as follows:

[0025]

[0026] in, Indicates the total number of randomly sampled view angles; U-Net network representing the multi-view diffusion model; Represents the output noise prediction result of the U-Net network; represents the first weight coefficient related to the timestamp; Indicates the indivual Single-view rendering at all times; express corresponding perspective; represents a second weight coefficient related to the timestamp; represents randomly sampled Gaussian noise; the variable Represents the timestamp of random sampling in the multi-view diffusion model.

[0027] As a preferred embodiment of the first aspect, in step S4, The perspective corresponding to the single-view rendering at that moment is used as the reference perspective, and the 3D model of the target data pair is rendered under the reference perspective to generate 3D rendering of the moment, Constantly add noise to the rendering, Moment 3D rendering, , multi-view alignment focal length, and timestamps randomly sampled in the multi-view diffusion model are input into the U-Net network of the multi-view diffusion model to output a new noise prediction result; The mean square error between the new noise prediction results and the randomly sampled Gaussian noise at all views at the moment is used as the geometric alignment loss.

[0028] Furthermore, the geometric alignment loss It is expressed as follows:

[0029]

[0030] in, From the perspective Rendering 3D model generated 3D rendering of the moment.

[0031] As a preferred embodiment of the above-mentioned first aspect, in step S4, the unknown moment-viewpoint alignment loss is formed by adding and averaging the action-geometry alignment losses of all moments, and the action-geometry alignment loss of each moment is composed of the addition of two items, the first item is the first fusion parameter multiplied by the action alignment loss, and the second item is the second fusion parameter multiplied by the geometric alignment loss. The first fusion parameter is the first ratio multiplied by the preset hyperparameter, and the second fusion parameter is the second ratio multiplied by the hyperparameter. The first ratio is the ratio of the current moment to the total number of moments, and the second ratio is the ratio of the moment difference to the total number of moments. The moment difference is the difference between the current moment and the total number of moments.

[0032] Furthermore, the unknown moment-view alignment loss It is expressed as follows:

[0033]

[0034] in, Represents the preset hyperparameters; represents the first fusion parameter; represents the second fusion parameter; represents the first ratio; Represents the second ratio.

[0035] In a second aspect, the present invention provides an alignment-based multimodal 4D content generation system, comprising:

[0036] A data acquisition module is used to feed the input data of any single modality into the diffusion model to generate target data pairs consisting of video and 3D model;

[0037] A first focal length acquisition module is configured to initialize a 4D model from the 3D model in the target data pair, use the first frame of the video in the target data pair as an alignment reference image, render the 4D model within a preset focal length range, obtain a primary perspective rendering image corresponding to each focal length at the initial moment, calculate the mean square error between each primary perspective rendering image and the alignment reference image, and use the focal length corresponding to the minimum mean square error as the video alignment focal length;

[0038] The second focal length acquisition module is used to retain the focal lengths smaller than the video alignment focal length within the focal length range of the first focal length acquisition module to form a new focal length range. The 4D model at the initial moment is rendered at different perspectives within the new focal length range. The fractional distillation sampling loss value corresponding to the rendered image at each focal length is calculated at each perspective. The loss values ​​of all perspectives at the same focal length are averaged, and the focal length corresponding to the minimum average loss value is used as the multi-perspective alignment focal length.

[0039] A model optimization module is configured to asynchronously optimize the 4D model with the goal of minimizing a total loss function. In the odd-numbered steps of the optimization process, only the 3D Gaussian model in the 4D model is optimized, while in the even-numbered steps, only the deformable mesh in the 4D model is optimized. The optimized 4D model is then used as the final high-quality 4D asset model. The total loss function is formed by the weighted summation of a known-time-perspective alignment loss and an unknown-time-perspective alignment loss. For known perspectives and times, the 4D model is rendered from the primary perspective based on the video alignment focal length to generate multi-time primary-perspective renderings. The known-time-perspective alignment loss is calculated from the multi-time primary-perspective renderings and the videos in the target data. For unknown perspectives and times, the 4D model is rendered from non-primary perspectives based on the multi-perspective alignment focal lengths to generate multi-time multi-perspective renderings. By aligning the multi-time multi-perspective renderings with the multi-perspective diffusion model prior, motion information and geometric information are transferred to the 4D model to construct motion alignment loss and geometric alignment loss. The motion alignment loss and geometric alignment loss are then weightedly fused with a time-related fusion parameter to obtain the unknown-time-perspective alignment loss.

[0040] The result acquisition module is used to input the observation perspective and time parameters into the high-quality 4D asset model to obtain the corresponding rendered image; continuous observation perspective and time parameters are input into the high-quality 4D asset model to obtain multi-time and multi-perspective rendered videos, completing the alignment-based multimodal 4D content generation.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] The present invention proposes a multimodal 4D content generation method based on alignment for input data of any single modality, which can efficiently generate a 4D asset model that is faithful to the input data. First, the present invention inputs data of any single modality (image / video / 3D model) into the diffusion model to generate video-3D model data pairs, breaking through the traditional 4D generation's dependence on multimodal joint input. Compared with traditional methods that require multi-view shooting or 3D scanning, the present invention has low data acquisition costs and wider applicability; secondly, the present invention uses the first frame of the video as a reference to calculate the MSE loss, determine the video alignment focal length, and achieve main perspective alignment. Then, within the reduced focal length range, the multi-perspective consistency is optimized based on the SDS loss, the multi-perspective alignment focal length is determined, and multi-perspective alignment is achieved. This two-stage focal length alignment optimization mechanism is used to improve The perspective-action alignment accuracy improves the geometric consistency of the 4D model and avoids the geometric distortion caused by focal length mismatching in traditional methods. Then, in the 4D model optimization process, the present invention uses the true value of the video for supervision at known moments, and transfers the action / geometry information by diffusion prior at unknown moments, and designs the fusion weights of the action / geometry loss that adjusts over time. In addition, the present invention also designs an alternating training strategy of optimizing the 3D Gaussian model in odd steps and optimizing the deformable network in even steps to accelerate the training convergence speed and improve the ability to retain dynamic details. Finally, the trained 4D asset model supports parametric rendering and is suitable for industrial-grade applications. From the perspective of practical application, the method of the present invention has achieved a breakthrough effect in the field of 4D content generation through multimodal input adaptation, dual focal length alignment optimization, asynchronous training strategy and spatiotemporal decoupling loss design, realizing efficient and high-quality 4D content generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a flow chart of the steps of the present invention;

[0044] Figure 2 This is a schematic diagram of the first type of input data in the X4D dataset according to an embodiment of the present invention;

[0045] Figure 3 Schematic diagram of rendering results of different methods on the first type of input data in an embodiment of the present invention;

[0046] Figure 4 This is a schematic diagram of the second type of input data in the X4D dataset according to an embodiment of the present invention;

[0047] Figure 5 Schematic diagram of rendering results of different methods on the second input data in an embodiment of the present invention;

[0048] Figure 6 Schematic diagram of input data in the Consistent4D dataset of an embodiment of the present invention;

[0049] Figure 7 Schematic diagram of rendering results of different methods on the Consistent4D dataset in an embodiment of the present invention;

[0050] Figure 8 This is a system block diagram of the present invention. DETAILED DESCRIPTION

[0051] In order to make the above-mentioned objects, features and advantages of the present invention more clearly understood, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways than those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. The technical features in the various embodiments of the present invention can be combined accordingly without conflicting with each other.

[0052] In the description of the present invention, it should be understood that the terms "first" and "second" are used solely for descriptive purposes and are not to be construed as indicating or implying relative importance or implicitly specifying the number of technical features being described. Therefore, features defined as "first" or "second" may explicitly or implicitly include at least one of such features.

[0053] like Figure 1 As shown, in a preferred implementation of the present invention, the above-mentioned alignment-based multimodal 4D content generation method includes the following steps S1 to S5. The specific implementation process is described below.

[0054] S1: Generation of multimodal data pairs and task transformation: Input data of any single modality (e.g., text, image, video, or 3D model) is fed into the diffusion model to generate target data pairs consisting of video and 3D model.

[0055] It should be noted that, in step S1 of the present invention, the input data is text, image, video or 3D model.

[0056] It should be noted that in step S1 of the present invention, the specific process of generating the video and 3D model in the target data pair is as follows: obtaining input data, and when the input data is text, inputting it into the text-video diffusion model to generate a video corresponding to the text, and inputting the first frame of the generated video into the image-3D diffusion model as a control condition to obtain a 3D model corresponding to the generated video; when the input data is an image, inputting it into the image-video diffusion model to generate a video corresponding to the image, and using the input image as a control condition to input it into the image-3D diffusion model to obtain a 3D model corresponding to the input image; when the input data is a video, inputting its first frame as a control condition into the image-3D diffusion model to obtain a 3D model corresponding to the input video; when the input data is a 3D model, using its frontal perspective rendered image as a control condition to input it into the image-video diffusion model to obtain a video corresponding to the input 3D model.

[0057] In this embodiment, step S1 effectively transforms the multimodal 4D generation task into a (video, 3D model) to 4D generation task. In this step, after generating the target data pair of video and 3D model according to the above process, the generated products are unified into a (video, 3D model) data pair for subsequent 4D generation.

[0058] S2: Obtain the video alignment focal length. Initialize the 4D model from the 3D model in the target data pair. Use the first frame of the target data pair's video as the alignment reference image. Render the 4D model within a preset focal length range to obtain the primary perspective rendering image corresponding to each focal length at the initial moment. Calculate the mean square error (MSE) between each primary perspective rendering image and the alignment reference image. Use the focal length corresponding to the minimum MSE as the video alignment focal length.

[0059] It should be noted that in step S2 of the present invention, the specific process of obtaining the video alignment focal length is as follows: the initial moment (first moment) of the 4D model is initialized by the 3D model in the target data pair, Frame video, use the video frame at the initial moment as the alignment reference image, render the 4D model at the same focal length interval within the preset focal length range, obtain the main perspective rendering image of each focal length at the initial moment, calculate the mean square error between each main perspective rendering image and the alignment reference image, and use the focal length corresponding to the minimum mean square error as the video alignment focal length .

[0060] In step S2 of this embodiment, a focal length value for rendering the 4D model is generated within a focal length range of 0 to 5 with an interval of 0.001, and a primary perspective rendering image of each focal length value at the initial moment is recorded.

[0061] S3: Get the multi-view alignment focal length. Within the focal length range of S2, the focal length that is smaller than the video alignment focal length is retained to form a new focal length range. The 4D model at the initial moment is rendered at different perspectives within the new focal length range. The Score Distillation Sampling (SDS) loss value corresponding to the rendering image of each focal length is calculated at each perspective. The loss values ​​of all perspectives at the same focal length are averaged, and the focal length corresponding to the minimum loss average value is used as the multi-view alignment focal length. .

[0062] It should be noted that in step S3 of the present invention, the specific process of calculating the fractional distillation sampling loss corresponding to the rendering image of each focal length is as follows: at each viewing angle, the rendering image corresponding to each focal length is input into the encoder respectively to obtain the feature map corresponding to each rendering image, and the randomly sampled Gaussian noise is superimposed on each feature map respectively. With the alignment reference image as the control condition, the fractional distillation sampling loss value is calculated based on the feature map after superimposing the noise to obtain the loss value corresponding to the rendering image at the focal length.

[0063] In step S3 of this embodiment, four viewing angles [0, 90, 180, 270] are preset, and the 4D model at the initial moment is rendered at each viewing angle. Then, a rendering image will be obtained at each viewing angle and focal length, that is, the same focal length will correspond to rendering images at four viewing angles. , Each represents a viewing angle. Then, the rendering images of each focal length under the four viewing angles are input into the encoder to obtain the feature maps of each focal length under the four viewing angles. Randomly sample a Gaussian noise and superimpose it on the feature map and calculate the SDS loss value.

[0064] S4: Optimize the 4D model. Asynchronously optimize the 4D model with the goal of minimizing the total loss function. During odd-numbered steps, only the 3D Gaussian model (3DGS) in the 4D model is optimized, while during even-numbered steps, only the deformable mesh in the 4D model is optimized. The optimized 4D model is used as the final high-quality 4D asset model.

[0065] The total loss function is formed by the weighted summation of the known moment-view alignment loss and the unknown moment-view alignment loss: for known perspectives and moments, the 4D model is rendered in the main perspective based on the video alignment focal length to generate multi-moment main perspective renderings, and the known moment-view alignment loss is calculated from the multi-moment main perspective renderings and the video in the target data; for unknown perspectives and moments, the 4D model is rendered in non-main perspectives based on the multi-perspective alignment focal length to generate multi-moment multi-perspective renderings, and the multi-moment multi-perspective renderings are aligned with the multi-perspective diffusion model prior, and the action information and geometric information are migrated to the 4D model to construct the action alignment loss and the geometric alignment loss, and the action alignment loss and the geometric alignment loss are each weightedly fused with a time-related fusion parameter to obtain the unknown moment-view alignment loss.

[0066] It should be noted that in step S4 of the present invention, a multi-time main perspective rendering image is generated by rendering the 4D model under the video alignment focal length and the main perspective. The multi-time main perspective rendering image is composed of multiple single-time main perspective rendering images, and each single-time main perspective rendering image corresponds one-to-one to each frame of the video in the target data pair; the mask of each frame and the single-time main perspective rendering image is obtained, and the mean square error loss between each frame and the single-time main perspective rendering image is calculated. The mean square error loss between the mask of each frame and the mask of the single-time main perspective rendering image is calculated, and the two parts of the mean square error loss are added as the known moment-perspective alignment loss.

[0067] In this embodiment, for known viewpoints and moments, the known moment-viewpoint alignment loss is obtained by calculating the mean square error loss between the single-moment main view rendering and the corresponding video frame and the mean square error loss between the masks. :

[0068]

[0069] in, Indicates the total number of frames in the target data pair, which is also the total number of moments; Indicates the main perspective (frontal perspective); express Main perspective rendering at all times; express The value on the alpha channel, that is Mask; Indicates the first frame; express Mask; Represents the square of the L2 norm.

[0070] It should be noted that in step S4 of the present invention, a multi-time multi-perspective rendering image is generated by rendering the 4D model in multiple perspectives with aligned focal lengths and in a non-primary perspective. The multi-time multi-perspective rendering image is composed of single-time single-perspective rendering images at different perspectives at different times.

[0071] It should be noted that, in step S4 of the present invention, A single-view rendering at a given moment and a first weight coefficient related to the timestamp Multiply the randomly sampled Gaussian noise and a second weight coefficient related to the timestamp Multiply the weighted The single-view rendering image and the weighted Gaussian noise are added to form Add noise to the rendering at any time, The noise rendering image and the target data are aligned with the video frame, The five parts of the perspective corresponding to the single-view rendering at the moment, the multi-view alignment focal length, and the timestamp randomly sampled in the multi-view diffusion model are input into the U-Net network of the multi-view diffusion model, and the noise prediction result is output; The mean square error between the noise prediction results and the randomly sampled Gaussian noise at all views at the moment is used as the action alignment loss.

[0072] In this embodiment, for At a certain moment, randomly sample a perspective in the non-main perspective range, render the 4D model under the multi-perspective alignment focal length and the randomly sampled perspective, and generate a rendering image at that moment and perspective, which is a single-time single-perspective rendering image. At the perspective At this moment, the corresponding indivual These renderings belong to the same moment, but not the same view. Furthermore, after generating the corresponding noise prediction results based on the single-view renderings at a single moment and a view at a single view, the noise prediction results of all views at the same moment are counted. The mean square error between these noise prediction results and the randomly sampled Gaussian noise is calculated to obtain the action alignment loss. :

[0073]

[0074] in, Indicates the total number of randomly sampled view angles; U-Net network representing the multi-view diffusion model; Represents the output noise prediction result of the U-Net network; represents the first weight coefficient related to the timestamp; Indicates the indivual Single-view rendering at all times; express corresponding perspective; represents a second weight coefficient related to the timestamp; represents randomly sampled Gaussian noise; the variable Represents the timestamp of random sampling in the multi-view diffusion model.

[0075] It should be noted that, in step S4 of the present invention, The perspective corresponding to the single-view rendering at that moment is used as the reference perspective, and the 3D model of the target data pair is rendered under the reference perspective to generate 3D rendering of the moment, Constantly add noise to the rendering, Moment 3D rendering, , multi-view alignment focal length, and timestamps randomly sampled in the multi-view diffusion model are input into the U-Net network of the multi-view diffusion model to output a new noise prediction result; The mean square error between the new noise prediction results and the randomly sampled Gaussian noise at all views at the moment is used as the geometric alignment loss.

[0076] In this embodiment, for At this moment, by changing the two inputs in the multi-view diffusion model, we can obtain The noise prediction result corresponding to a certain perspective at the moment. Then the noise prediction results of all perspectives at the same moment are counted. The mean square error between these noise prediction results and the randomly sampled Gaussian noise is calculated to obtain the geometric alignment loss. :

[0077]

[0078] in, From the perspective Rendering 3D model generated 3D rendering of the moment.

[0079] It should be noted that in step S4 of the present invention, the unknown moment-viewpoint alignment loss is formed by adding and averaging the action-geometry alignment losses of all moments. The action-geometry alignment loss at each moment is composed of the addition of two items. The first item is the multiplication of the first fusion parameter and the action alignment loss, and the second item is the multiplication of the second fusion parameter and the geometric alignment loss. The first fusion parameter is the multiplication of the first ratio and the preset hyperparameter, and the second fusion parameter is the multiplication of the second ratio and the hyperparameter. The first ratio is the ratio of the current moment to the total number of moments, and the second ratio is the ratio of the moment difference to the total number of moments. The moment difference is the difference between the current moment and the total number of moments.

[0080] In this embodiment, due to the movement time tend to When the geometric representation is far away from the initial state, the present invention designs two time-related fusion parameters to obtain the motion-geometry alignment loss at each moment. moment, the corresponding action-geometry alignment loss is:

[0081]

[0082] in, Represents the preset hyperparameters; represents the first fusion parameter; represents the second fusion parameter; represents the first ratio; Represents the second ratio.

[0083] Furthermore, the unknown moment-view alignment loss can be formed by adding and averaging the action-geometry alignment losses at all moments :

[0084]

[0085] Therefore, the total loss function formed by the weighted summation of the known moment-view alignment loss and the unknown moment-view alignment loss is as follows:

[0086]

[0087] In the total loss function, the action alignment loss ensures that the position motion posture of the 4D object at each moment is consistent with the motion posture of the front view; by optimizing the known moment-view alignment loss, the multi-moment frontal view of the 4D model is aligned with the video in the target data.

[0088] In this embodiment, gradient backpropagation is performed based on the obtained total loss function. When the number of training steps is odd, the deformable network weights are frozen and only the 3D Gaussian model is optimized. When the number of training steps is even, the 3D Gaussian model is frozen and only the deformable network is optimized. This asynchronous optimization method has the following two advantages: First, using multi-frame videos and random perspective parameters as control conditions, the 4D renderings under the corresponding time-perspective are optimized to enhance the generalization ability of the model under random perspectives; second, using the multi-perspective 4D renderings at time 0 and the main perspective parameters as control conditions, the 4D renderings under the same perspective at other times are optimized to ensure the perspective consistency of the model at different times.

[0089] S5: Rendered image or rendered video generation. Input the viewing angle and time parameters into the high-quality 4D asset model to generate the corresponding rendered image. Input continuous viewing angles and time parameters into the high-quality 4D asset model to generate rendered videos at multiple times and perspectives, completing alignment-based multimodal 4D content generation.

[0090] Therefore, the high-quality 4D asset model trained in step S4 can drive the generation of diverse 4D content using the input of any modality as a control condition.

[0091] In order to better demonstrate the specific implementation and technical effects of the present invention, the alignment-based multimodal 4D content generation method shown in steps S1 to S5 in the above preferred implementation is applied to a specific example.

[0092] Example

[0093] The steps of this embodiment are the same as the alignment-based multimodal 4D content generation method shown in the aforementioned steps S1 to S5, and will not be repeated here. The specific data set, some specific parameter settings and implementation results of this embodiment are mainly demonstrated.

[0094] As a quantitative indicator, this embodiment conducts multiple quantitative tests on the Consistent4D and X4D datasets. The test results of the method of the present invention and the four comparative methods of L4GM, SC4D, STAG4D, and DG4D on the Consistent4D dataset are shown in Table 1, and the test results on the X4D dataset are shown in Table 2. In the X4D dataset, one input data is as follows: Figure 2 As shown, the rendering results of different methods on the input data are as follows Figure 3 As shown, another input data in the X4D dataset is Figure 4 As shown, the rendering results of different methods on the input data are as follows Figure 5 As shown, one input data in the Consistent4D dataset is as follows Figure 6As shown, the rendering results of different methods on the input data are as follows Figure 7 As shown. It should also be noted that L4GM, SC4D, STAG4D, and DG4D are all existing technologies, so the implementation process of each method will not be detailed here. In Table 1, PSNR (Peak Signal-to-Noise Ratio) is the peak signal-to-noise ratio, SSIM (Structural Similarity Index Measure) is the structural similarity index, LPIPS (Learned Perceptual Image Patch Similarity) is the learned perceptual image patch similarity, and FVD (Fréchet Video Distance) is the Fréchet video distance. CLIP (Contrastive Language–Image Pretraining) indicates the degree of match between the rendering results generated by the model and the text description or semantic intent. In Table 2, Appearance indicates the realism and temporal coherence of the rendering results in terms of visual appearance (such as color, texture, and lighting consistency); Structure indicates the geometric accuracy, temporal consistency, and physical plausibility of the rendering results; Motion is used to measure the motion accuracy, naturalness, and physical plausibility of the dynamic model in the temporal dimension; and Fidelity is a multi-dimensional fidelity metric.

[0095] Table 1 Consistent4D dataset evaluation table

[0096] Table 2 X4D dataset evaluation table

[0097] It should also be noted that the alignment-based multimodal 4D content generation method in the above embodiment can essentially be executed by a computer program or module. Therefore, similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides an alignment-based multimodal 4D content generation system corresponding to the alignment-based multimodal 4D content generation method provided in the above embodiment, such as Figure 8 As shown, it includes:

[0098] A data acquisition module is used to feed the input data of any single modality into the diffusion model to generate target data pairs consisting of video and 3D model;

[0099] A first focal length acquisition module is configured to initialize a 4D model from the 3D model in the target data pair, use the first frame of the video in the target data pair as an alignment reference image, render the 4D model within a preset focal length range, obtain a primary perspective rendering image corresponding to each focal length at the initial moment, calculate the mean square error between each primary perspective rendering image and the alignment reference image, and use the focal length corresponding to the minimum mean square error as the video alignment focal length;

[0100] The second focal length acquisition module is used to retain the focal lengths smaller than the video alignment focal length within the focal length range of the first focal length acquisition module to form a new focal length range. The 4D model at the initial moment is rendered at different perspectives within the new focal length range. The fractional distillation sampling loss value corresponding to the rendered image at each focal length is calculated at each perspective. The loss values ​​of all perspectives at the same focal length are averaged, and the focal length corresponding to the minimum average loss value is used as the multi-perspective alignment focal length.

[0101] A model optimization module is configured to asynchronously optimize the 4D model with the goal of minimizing a total loss function. In the odd-numbered steps of the optimization process, only the 3D Gaussian model in the 4D model is optimized, while in the even-numbered steps, only the deformable mesh in the 4D model is optimized. The optimized 4D model is then used as the final high-quality 4D asset model. The total loss function is formed by the weighted summation of a known-time-perspective alignment loss and an unknown-time-perspective alignment loss. For known perspectives and times, the 4D model is rendered from the primary perspective based on the video alignment focal length to generate multi-time primary-perspective renderings. The known-time-perspective alignment loss is calculated from the multi-time primary-perspective renderings and the videos in the target data. For unknown perspectives and times, the 4D model is rendered from non-primary perspectives based on the multi-perspective alignment focal lengths to generate multi-time multi-perspective renderings. By aligning the multi-time multi-perspective renderings with the multi-perspective diffusion model prior, motion information and geometric information are transferred to the 4D model to construct motion alignment loss and geometric alignment loss. The motion alignment loss and geometric alignment loss are then weightedly fused with a time-related fusion parameter to obtain the unknown-time-perspective alignment loss.

[0102] The result acquisition module is used to input the observation perspective and time parameters into the high-quality 4D asset model to obtain the corresponding rendered image; continuous observation perspective and time parameters are input into the high-quality 4D asset model to obtain multi-time and multi-perspective rendered videos, completing the alignment-based multimodal 4D content generation.

[0103] It should also be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the system described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here. In the various embodiments provided in this application, the division of steps or modules in the system and method is only a logical function division. In actual implementation, there may be other division methods, for example, multiple modules or steps can be combined or integrated together, and a module or step can also be split.

[0104] The embodiment described above is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Persons skilled in the art may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution obtained by equivalent substitution or equivalent transformation falls within the scope of protection of the present invention.

Claims

1. A method for generating multimodal 4D content based on alignment, characterized in that: The following steps are involved: S1: Input data of any single modality is fed into the diffusion model to generate target data pairs consisting of video and 3D model; S2: Initialize the 4D model from the 3D model in the target data pair, and use the first frame of the video in the target data pair as the alignment reference image. Render the 4D model within a preset focal length range to obtain the main perspective rendering image corresponding to each focal length at the initial moment. Calculate the mean square error between each main perspective rendering image and the alignment reference image, and use the focal length corresponding to the minimum mean square error as the video alignment focal length. S3: Within the focal length range of S2, the focal lengths smaller than the video alignment focal length are retained to form a new focal length range. The 4D model at the initial moment is rendered at different perspectives within the new focal length range. The fractional distillation sampling loss value corresponding to the rendered image at each focal length is calculated at each perspective. The loss values ​​of all perspectives at the same focal length are averaged, and the focal length corresponding to the minimum average loss value is used as the multi-perspective alignment focal length. S4: Asynchronously optimize the 4D model with the goal of minimizing the total loss function. In odd-numbered steps of the optimization process, only the 3D Gaussian model in the 4D model is optimized, while in even-numbered steps, only the deformable mesh in the 4D model is optimized. The optimized 4D model is used as the final high-quality 4D asset model. The total loss function is formed by the weighted sum of the known moment-view alignment loss and the unknown moment-view alignment loss. For known viewpoints and time, the 4D model is rendered in the main view based on the video alignment focal length to generate multi-moment main view renderings. The known moment-view alignment loss is calculated from the multi-moment main view renderings and the video in the target data pair. For unknown viewpoints and moments, the 4D model is rendered in non-primary viewpoints based on the multi-view alignment focal length to generate multi-time and multi-view renderings. By aligning the multi-time and multi-view renderings with the multi-view diffusion model prior, motion and geometric information are transferred to the 4D model to construct motion alignment loss and geometric alignment loss. The motion alignment loss and geometric alignment loss are then weightedly fused with a time-related fusion parameter to obtain the unknown time-view alignment loss. S5: Input the observation angle and time parameters into the high-quality 4D asset model to obtain the corresponding rendered image; input the continuous observation angle and time parameters into the high-quality 4D asset model to obtain multi-time and multi-view rendered videos, completing the alignment-based multimodal 4D content generation.

2. The alignment-based multimodal 4D content generation method according to claim 1, wherein: In step S1 , the input data is text, image, video or 3D model.

3. The alignment-based multimodal 4D content generation method according to claim 2, wherein: In step S1, the specific process of generating the video and 3D model in the target data pair is as follows: obtaining input data, when the input data is text, inputting it into the text-video diffusion model to generate a video corresponding to the text, and inputting the first frame of the generated video as a control condition into the image-3D diffusion model to obtain a 3D model corresponding to the generated video; when the input data is an image, inputting it into the image-video diffusion model to generate a video corresponding to the image, and using the input image as a control condition to input it into the image-3D diffusion model to obtain a 3D model corresponding to the input image; when the input data is a video, inputting its first frame as a control condition into the image-3D diffusion model to obtain a 3D model corresponding to the input video; when the input data is a 3D model, using its frontal perspective rendered image as a control condition to input it into the image-video diffusion model to obtain a video corresponding to the input 3D model.

4. The alignment-based multimodal 4D content generation method according to claim 1, wherein: In step S2, the specific process of obtaining the video alignment focal length is as follows: initialize the initial moment of the 4D model from the 3D model in the target data pair, use the video frame at the initial moment as the alignment reference image for the T-frame video in the target data pair, render the 4D model at the same focal length interval within the preset focal length range, obtain the main perspective rendering image of each focal length at the initial moment, calculate the mean square error between each main perspective rendering image and the alignment reference image, and use the focal length corresponding to the minimum mean square error as the video alignment focal length.

5. The alignment-based multimodal 4D content generation method according to claim 1, wherein: In step S3, the specific process of calculating the fractional distillation sampling loss corresponding to the rendering image of each focal length is as follows: at each viewing angle, the rendering image corresponding to each focal length is input into the encoder respectively to obtain the feature map corresponding to each rendering image, and the randomly sampled Gaussian noise is superimposed on each feature map respectively. With the alignment reference image as the control condition, the fractional distillation sampling loss value is calculated based on the feature map after superimposing the noise to obtain the loss value corresponding to the rendering image at this focal length.

6. The alignment-based multimodal 4D content generation method according to claim 1, wherein: In step S4, a multi-time main-view rendering is generated by rendering the 4D model under the video alignment focal length and the main-view. The multi-time main-view rendering is composed of multiple single-time main-view renderings, and each single-time main-view rendering corresponds to each frame of the target data in the video. The mask of each frame and the single-time main-view rendering is obtained, and the mean square error loss between each frame and the single-time main-view rendering is calculated. The mean square error loss between the mask of each frame and the mask of the single-time main-view rendering is calculated, and the two parts of the mean square error loss are added as the known moment-view alignment loss.

7. The alignment-based multimodal 4D content generation method according to claim 1, wherein: In step S4, a multi-time multi-perspective rendering image is generated by rendering the 4D model under multi-perspective alignment focal length and non-main perspective. The multi-time multi-perspective rendering image is composed of single-time single-perspective rendering images of different perspectives at different times; the single-perspective rendering image at time t is multiplied by a first weight coefficient related to the timestamp, the randomly sampled Gaussian noise is multiplied by a second weight coefficient related to the timestamp, the weighted single-perspective rendering image at time t and the weighted Gaussian noise are added to form a noisy rendering image at time t, the noisy rendering image at time t, the tth frame of the video in the target data, the perspective corresponding to the single-perspective rendering image at time t, the multi-perspective alignment focal length, and the randomly sampled timestamp in the multi-perspective diffusion model are input into the U-Net network of the multi-perspective diffusion model, and the noise prediction result is output; the mean square error between the noise prediction results at all perspectives at time t and the randomly sampled Gaussian noise is used as the action alignment loss.

8. The alignment-based multimodal 4D content generation method according to claim 7, wherein: In step S4, the perspective corresponding to the single-view rendering image at time t is used as the reference perspective, and the 3D model in the target data pair is rendered under the reference perspective to generate the 3D rendering image at time t. The five parts, namely the noisy rendering image at time t, the 3D rendering image at time t, 0, the multi-view alignment focal length, and the randomly sampled timestamp in the multi-view diffusion model, are input into the U-Net network of the multi-view diffusion model, and a new noise prediction result is output. The mean square error between the new noise prediction result and the randomly sampled Gaussian noise under all perspectives at time t is used as the geometric alignment loss.

9. The alignment-based multimodal 4D content generation method according to claim 8, wherein: In step S4, the unknown moment-viewpoint alignment loss is formed by adding and averaging the action-geometry alignment losses of all moments. The action-geometry alignment loss at each moment is composed of two items. The first item is the multiplication of the first fusion parameter and the action alignment loss, and the second item is the multiplication of the second fusion parameter and the geometric alignment loss. The first fusion parameter is the multiplication of the first ratio and the preset hyperparameter, and the second fusion parameter is the multiplication of the second ratio and the hyperparameter. The first ratio is the ratio of the current moment to the total number of moments, and the second ratio is the ratio of the moment difference to the total number of moments. The moment difference is the difference between the current moment and the total number of moments.

10. A multimodal 4D content generation system based on alignment, characterized in that: include: A data acquisition module is used to feed the input data of any single modality into the diffusion model to generate target data pairs consisting of video and 3D model; A first focal length acquisition module is configured to initialize a 4D model from the 3D model in the target data pair, use the first frame of the video in the target data pair as an alignment reference image, render the 4D model within a preset focal length range, obtain a primary perspective rendering image corresponding to each focal length at the initial moment, calculate the mean square error between each primary perspective rendering image and the alignment reference image, and use the focal length corresponding to the minimum mean square error as the video alignment focal length; The second focal length acquisition module is used to retain the focal lengths smaller than the video alignment focal length within the focal length range of the first focal length acquisition module to form a new focal length range. The 4D model at the initial moment is rendered at different perspectives within the new focal length range. The fractional distillation sampling loss value corresponding to the rendered image at each focal length is calculated at each perspective. The loss values ​​of all perspectives at the same focal length are averaged, and the focal length corresponding to the minimum average loss value is used as the multi-perspective alignment focal length. A model optimization module is used to asynchronously optimize the 4D model with the goal of minimizing a total loss function. During odd-numbered steps of the optimization process, only the 3D Gaussian model in the 4D model is optimized, while during even-numbered steps, only the deformable mesh in the 4D model is optimized. The optimized 4D model is used as the final high-quality 4D asset model. The total loss function is formed by the weighted sum of a known moment-view alignment loss and an unknown moment-view alignment loss. For known viewpoints and moments, the 4D model is rendered from the primary perspective based on the video alignment focal length to generate multi-moment primary perspective renderings. The known moment-view alignment loss is calculated from the multi-moment primary perspective renderings and the video in the target data pair. For unknown viewpoints and moments, the 4D model is rendered in non-primary viewpoints based on the multi-view alignment focal length to generate multi-time and multi-view renderings. By aligning the multi-time and multi-view renderings with the multi-view diffusion model prior, motion and geometric information are transferred to the 4D model to construct motion alignment loss and geometric alignment loss. The motion alignment loss and geometric alignment loss are then weightedly fused with a time-related fusion parameter to obtain the unknown time-view alignment loss. The result acquisition module is used to input the observation perspective and time parameters into the high-quality 4D asset model to obtain the corresponding rendered image; continuous observation perspective and time parameters are input into the high-quality 4D asset model to obtain multi-time and multi-perspective rendered videos, completing the alignment-based multimodal 4D content generation.

Citation Information

Patent Citations

  • Multi-modal image registration method based on parallax estimation

    CN115471397A

  • Four-dimensional Gaussian model generation method, system and equipment based on Gaussian sputtering

    CN119338966A

  • Unmanned aerial vehicle auxiliary photographing correction method based on improved RT-DETR model

    CN120494053A

  • Dynamic human modeling method based on three-dimensional gaussians

    WO2025118224A1