Single-image four-dimensional dynamic human body video generation method based on diffusion converter

The diffusion transformer-based method decomposes the generation process into spatial, angular, and temporal dimensions to efficiently generate high-quality, realistic four-dimensional dynamic human videos, addressing the high resource demands of existing methods.

CN120318383APending Publication Date: 2025-07-15TSINGHUA UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510243470.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing technology requires high computing resources and video memory costs when establishing video diffusion models and dynamic three-dimensional human body correlation, making it difficult to achieve high-quality and high-reality four-dimensional dynamic human body generation.

Method used

A single-image four-dimensional dynamic human video generation method based on diffusion converter is adopted. Through a layered four-dimensional diffusion converter, including space, viewing angle and time dimension transformer, combined with self-attention decomposition operations and feature injection, the calculation complexity is reduced and the space-time consistency is optimized.

Benefits of technology

It effectively reduces the computing cost and video memory space, and realizes high-quality and high-reality four-dimensional dynamic human body generation, suitable for the generation of various dynamic human body scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318383A_ABST
    Figure CN120318383A_ABST
Patent Text Reader

Abstract

The invention provides a single-image four-dimensional dynamic human body video generation method based on a diffusion converter, and relates to the field of human body dynamic three-dimensional modeling and video generation in computer vision and computer graphics, and the method comprises the steps: obtaining any character action video; the character action video is input into a layered four-dimensional diffusion converter, the diffusion converter is updated, and the diffusion converter comprises a spatial dimension converter, a visual angle dimension converter and a time dimension converter; executing self-attention decomposition operation in the diffusion converter; obtaining a single figure image, and extracting figure identity features, camera posture features and human body action features in the image; and respectively injecting the extracted features into corresponding converters, generating a video sequence through a diffusion converter, and carrying out space-time consistency optimization on the generated video sequence to obtain a target video sequence. The dynamic human body scene generation method and device are suitable for generation of various dynamic human body scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of dynamic four-dimensional generation and video generation in computer vision and computer graphics, and particularly relates to a method and device for generating a single-image four-dimensional dynamic human body video based on a diffusion transformer. Background Art

[0002] In recent years, with the proposal and development of video diffusion models, the generation of dynamic human bodies through video generation has achieved a high-fidelity digital human generation effect in general scenarios. However, extending video generation to three-dimensional scenarios remains a huge challenge, and how to establish the association between video diffusion models and dynamic three-dimensional human bodies is the core issue. A direct solution is to use a spatio-temporal full attention mechanism to build the association. However, establishing connections in four-dimensional space in this way requires high computing resources and video memory costs. Summary of the Invention

[0003] This application aims to solve at least one of the technical problems in the related art to some extent.

[0004] To this end, the first object of this application is to propose a method for generating a single-image four-dimensional dynamic human body video based on a diffusion transformer, which solves the technical problem of the high computing resources and video memory costs required by existing methods, effectively reduces the computing cost and video memory space occupancy, and realizes high-quality and high-fidelity four-dimensional dynamic human body generation.

[0005] The second object of this application is to propose a computer device.

[0006] The third object of this application is to propose a non-transitory computer-readable storage medium.

[0007] To achieve the above object, the first aspect embodiment of this application proposes a method for generating a single-image four-dimensional dynamic human body video based on a diffusion transformer, including:

[0008] Obtain any human action video;

[0009] Input the human action video into a hierarchical four-dimensional diffusion transformer and update the diffusion transformer, where the diffusion transformer includes a spatial dimension transformer, a perspective dimension transformer, and a time dimension transformer;

[0010] Perform self-attention decomposition operations in the diffusion transformer;

[0011] Obtain a single human image and extract the human identity feature, camera pose feature, and human body action feature in the image;

[0012] Inject the extracted features into the corresponding transformers respectively, generate a video sequence through a diffusion transformer, and optimize the spatio-temporal consistency of the generated video sequence to obtain the target video sequence.

[0013] Optionally, in an embodiment of the present application, input a human action video into a hierarchical four-dimensional diffusion transformer and update the diffusion transformer, including:

[0014] Through the spatial dimension transformer, divide each input video frame spatially, and calculate the correlation between spatial positions;

[0015] Through the perspective dimension transformer, calculate the features between different perspectives, and construct the corresponding relationship between any two perspectives within the range of 0 degrees to 360 degrees;

[0016] Through the perspective dimension transformer, establish a correlation between any two video frames at different times in the video sequence;

[0017] Stack the multi-dimensional transformers with each other through residual connections and layer normalization to form a hierarchical four-dimensional diffusion transformer.

[0018] Optionally, in an embodiment of the present application, perform self-attention decomposition operations in the diffusion transformer, including:

[0019] In the spatial dimension transformer, divide the input image into feature grids of a preset size, each feature grid position contains a feature vector of a preset size, and capture the long-range spatial dependence within the image by calculating the self-attention between spatial positions;

[0020] In the perspective dimension transformer, adopt the multi-head self-attention mechanism, and establish the correlation mapping between perspectives by calculating the similarity of query-key-value pairs between different perspectives, and construct the feature corresponding relationship between any two perspectives within the range of 0 degrees to 360 degrees;

[0021] In the time dimension transformer, adopt the multi-head self-attention mechanism, divide the time window by sampling window in time series, and perform time series modeling between adjacent and non-adjacent frames within the time window of the video sequence.

[0022] Optionally, in an embodiment of the present application, extract the human identity features, camera pose information, and human body action features in the image, including:

[0023] Use a pre-trained CLIP visual encoder to process the input image and extract the human identity feature vector, where the human identity feature vector includes the facial features, body shape features, and clothing features of the person;

[0024] Encode the input image using a variational autoencoder to obtain a human action feature vector, where the human action feature vector includes the pose information, joint position information, and motion state information of the human body;

[0025] Use a multi-layer perceptron combined with positional encoding to process the camera parameters, and convert the intrinsic matrix and extrinsic matrix of the camera into a camera pose feature vector, where the camera pose feature vector includes the focal length, principal point coordinates, rotation matrix, and translation vector of the camera.

[0026] Optionally, in an embodiment of the present application, inject the extracted features into the corresponding transformers respectively, including:

[0027] Convert the person identity feature through a linear projection layer to change the dimension, and use an additive attention mechanism to inject it into each feature grid position of the spatial dimension transformer to ensure the consistency of the person identity feature of the generated image with the input image;

[0028] Inject the camera pose feature into the view dimension transformer through a cross-attention mechanism to guide the image generation process under different views;

[0029] Copy the human action feature to obtain a first feature and a second feature, inject the first feature into the spatial dimension transformer after linear transformation to guide the pose generation of each frame, and inject the second feature into the time dimension transformer after processing through a temporal cross-attention mechanism to control the motion change between consecutive frames;

[0030] Inject the extracted features into the corresponding transformers respectively, and further include:

[0031] All feature injection processes adopt a residual connection structure, and use LayerNorm for feature normalization before and after injection.

[0032] Optionally, in an embodiment of the present application, optimize the spatio-temporal consistency of the generated video sequence to obtain a target video sequence, including:

[0033] Initialize Gaussian random noise in the spatio-temporal dimension, and the noise dimension is consistent with the target video sequence;

[0034] During the diffusion denoising process, adopt a sliding window strategy for sampling, including:

[0035] Set a sliding window of a preset size in the time dimension, and for the current frame to be generated, consider the noise states of several frames before and after it within the preset size at the same time;

[0036] Set a sliding window of a preset number of views in the view dimension, and for the current view, consider the noise states of several adjacent views within the preset number at the same time;

[0037] For the noise signals within each spatio-temporal sampling window, weight coefficients are calculated through an attention mechanism, and the weight values are adaptively adjusted according to the spatio-temporal distance, such that samples closer to the current moment or the current perspective obtain higher weights;

[0038] The weighted multiple noise signals are fused to obtain an updated denoised signal.

[0039] Optionally, in an embodiment of the present application, before inputting the human action video into the hierarchical four-dimensional diffusion transformer, the diffusion transformer is trained using a multi-dimensional human body dataset, including:

[0040] Training the spatial dimension transformer using static image data to optimize the reconstruction loss and the perceptual loss;

[0041] Training the perspective dimension transformer using multi-perspective data to optimize the perspective consistency loss and the feature matching loss;

[0042] Training the time dimension transformer using dynamic video data and 3D reconstruction data to optimize the temporal smoothing loss and the 3D consistency loss;

[0043] Performing end-to-end joint optimization of the diffusion transformer using four-dimensional sequence data;

[0044] Training the diffusion transformer using a multi-dimensional human body dataset further includes:

[0045] In each training stage, the Adam optimizer is adopted.

[0046] To achieve the above object, an embodiment of the second aspect of the present invention proposes a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above method for generating a single-image four-dimensional dynamic human body video based on a diffusion transformer is implemented.

[0047] To achieve the above object, an embodiment of the third aspect of the present invention proposes a non-transitory computer-readable storage medium, which can execute the method for generating a single-image four-dimensional dynamic human body video based on a diffusion transformer when the instructions in the storage medium are executed by a processor.

[0048] The method for generating a single-image four-dimensional dynamic human body video based on a diffusion transformer according to the embodiments of the present application decomposes the generation process of the four-dimensional dynamic human body into three dimensions: space, perspective, and time domain. It constructs the associations between different dimensions hierarchically in each dimension, reducing the complexity of the spatio-temporal full attention mechanism from four dimensions to two dimensions, thereby greatly reducing the computational cost and the occupancy of video memory space, and finally achieving a high-quality and high-fidelity four-dimensional dynamic human body generation effect. This embodiment can be applied and extended to the generation of various dynamic human body scenarios, including using only a single-person image to freely control the generation of the character's actions and the generation of the camera in the scene. The achieved generation effect can be widely applied to various applications, including film production, holographic video generation, high-precision digital human generation, AR / VR interaction, remote conferencing, and other fields.

[0049] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, in which:

[0051] Figure 1 is a schematic flowchart of a method for generating a single-image four-dimensional dynamic human body video based on a diffusion transformer according to Embodiment 1 of the present application;

[0052] Figure 2 is a hierarchical four-dimensional transformer architecture diagram of the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present application and should not be construed as limiting the present application.

[0054] The method and device for generating a single-image four-dimensional dynamic human body video based on a diffusion transformer according to the embodiments of the present application will be described below with reference to the drawings.

[0055] Figure 1 is a schematic flowchart of a method for generating a single-image four-dimensional dynamic human body video based on a diffusion transformer according to Embodiment 1 of the present application.

[0056] As Figure 1 shown, the method for generating a single-image four-dimensional dynamic human body video based on a diffusion transformer includes the following steps:

[0057] Step 101, obtain an arbitrary human action video;

[0058] In this embodiment, a video of a person's action is captured using a single RGB camera or a mobile phone camera, and the whole body of the person should be included.

[0059] Step 102: Input the person's action video into a hierarchical four-dimensional diffusion transformer and update the diffusion transformer. The diffusion transformer includes a spatial dimension transformer, a perspective dimension transformer, and a time dimension transformer.

[0060] Step 103: Perform self-attention decomposition operation in the diffusion transformer.

[0061] Step 104: Obtain a single person image and extract the person's identity features, camera pose features, and human body action features in the image.

[0062] In this embodiment, a high-definition image of a person is captured using a single RGB camera or a mobile phone camera, and the whole body of the person should be included.

[0063] Step 105: Inject the extracted features into the corresponding transformers respectively, generate a video sequence through the diffusion transformer, and perform spatio-temporal consistency optimization on the generated video sequence to obtain the target video sequence.

[0064] Optionally, in an embodiment of the present application, inputting the person's action video into a hierarchical four-dimensional diffusion transformer and updating the diffusion transformer includes:

[0065] First, implement the spatial dimension transformer to spatially divide each input video frame and improve the quality of human body generation by calculating the correlation between spatial positions. Then establish the perspective dimension transformer to build the corresponding relationship between any two perspectives within the range of 0 degrees to 360 degrees by calculating the features between different perspectives; then construct the time dimension transformer to establish the correlation between any two video frames at different moments in the video sequence; these three-dimensional transformers are stacked with each other through residual connections and layer normalization, with a total of 24 layers, forming a complete hierarchical four-dimensional transformer architecture, as Figure 2 shown.

[0066] Optionally, in an embodiment of the present application, performing the self-attention decomposition operation in the diffusion transformer includes:

[0067] Perform different attention decomposition mechanisms in the three multi-dimensional transformers of the diffusion transformer. In the perspective transformer, mainly use the multi-head self-attention mechanism, the number of heads is 16, and the dimension of each attention head is 64. Establish the correlation mapping between perspectives by calculating the query-key-value pair similarity between different perspectives, so that effective feature correspondence can be established between any two perspectives within the range of 0 degrees to 360 degrees;

[0068]

[0069] In the temporal dimension transformer, a 16 - head self - attention mechanism is also adopted. In the time series, a 24 - frame sampling window is used to divide the time window, and temporal modeling is performed between adjacent and non - adjacent frames within a time window of the video sequence to ensure that the generated sequence maintains temporal coherence within a duration range of 2 - 4 seconds;

[0070]

[0071] In the spatial dimension transformer, the input image is divided into a 32×32 feature grid, and each grid position contains a 1024 - dimensional feature vector. By calculating the self - attention between spatial positions, long - range spatial dependencies within the image are captured to achieve precise modeling of human body local details;

[0072]

[0073] Optionally, in an embodiment of the present application, the human identity features, camera pose information, and human body motion features in the image are extracted, including:

[0074] The pre - trained CLIP visual encoder is used to process the input image to extract a 512 - dimensional human identity feature vector, which contains the facial features, body type features, and clothing features of the person; the variational auto - encoder (VAE) is used to encode the input image to obtain a 1024 - dimensional human body motion feature vector, which encodes the pose information, joint position information, and motion state information of the human body; a multi - layer perceptron (MLP) combined with position encoding is used to process the camera parameters to convert the internal and external parameter matrices of the camera into a 1024 - dimensional camera pose feature vector, which contains information such as the focal length, principal point coordinates, rotation matrix, and translation vector of the camera. The resolution requirement of the input image is 768×768 or higher, the image needs to contain a complete human body, and the human body area occupies the main part of the image (at least 70% of the image area). The extracted feature vectors will be used as conditional inputs for the subsequent diffusion transformer.

[0075] Optionally, in an embodiment of the present application, a cross - attention mechanism is adopted to inject the extracted features into the corresponding transformers respectively, which is expressed as:

[0076]

[0077] This process includes:

[0078] First, the 512-dimensional person identity features are transformed into 1024 dimensions through a linear projection layer, and an additive attention mechanism is used to inject them into each feature grid position of the spatial dimension transformer to ensure that the generated image is consistent with the input image in terms of facial features, clothing details, etc. Then, the 1024-dimensional camera pose features are injected into the perspective dimension transformer through a cross-attention mechanism to guide the image generation process under different perspectives and achieve precise perspective control. Finally, the 1024-dimensional human motion features are copied twice. The first copy is injected into the spatial dimension transformer after linear transformation to guide the pose generation of each frame, and the second copy is injected into the temporal dimension transformer after being processed by the temporal cross-attention mechanism to control the motion changes between consecutive frames. All feature injection processes adopt a residual connection structure, and LayerNorm is used for feature normalization before and after injection.

[0079] Optionally, in an embodiment of the present application, spatio-temporal consistency optimization is performed on the generated video sequence to obtain a target video sequence, including:

[0080] First, Gaussian random noise is initialized in the spatio-temporal dimension, and the noise dimension is consistent with the target video sequence (the temporal dimension is 25fps × video duration, the perspective dimension is 12 uniformly sampled perspectives, and the spatial dimension is 768×768 pixel resolution);

[0081] During the diffusion denoising process, a sliding window strategy is adopted for sampling, specifically including: setting a sliding window of 24 frames in the temporal dimension, and for the current frame to be generated, considering the noise states of its 12 frames before and after at the same time; setting a sliding window of 4 perspectives in the perspective dimension, and for the current perspective, considering the noise states of its three adjacent perspectives (±90 degrees) at the same time; for the noise signals within each spatio-temporal sampling window, weight coefficients are calculated through an attention mechanism, and the weight values are adaptively adjusted according to the spatio-temporal distance. Samples closer to the current moment or the current perspective obtain higher weights; the weighted multiple noise signals are fused to obtain an updated denoising signal, which has strong consistency in both the temporal and perspective dimensions.

[0082] Optionally, in an embodiment of the present application, when inputting a human motion video into a hierarchical four-dimensional diffusion transformer, training the diffusion transformer using a multi-dimensional human dataset, including:

[0083] Construct a multi-dimensional human body dataset, which includes the following subsets: a static image dataset containing 50,000 high-definition human images with a resolution of 1080 or above, covering human samples of different genders, ages, body types, and clothing; a dynamic video dataset containing 10,000 human body movement video clips, each clip with a duration of 2 - 10 seconds and a frame rate of 25fps; a perspective dataset containing 5,000 groups of multi-perspective human body data, each group containing synchronously captured images of 12 different perspectives (spaced 30 degrees apart); a dataset containing 5,000 groups of 3D human body reconstruction data, each group containing a complete 3D human body mesh model and corresponding texture information; a dataset containing 1,000 groups of four-dimensional dynamic human body data sequences, each group containing a complete temporal 3D reconstruction result. Adopt a multi-stage training strategy: First, use static image data to train the spatial dimension transformer, optimize the reconstruction loss and perceptual loss, with a learning duration of 500,000 iterations; then use multi-perspective data to train the perspective dimension transformer, adopt the perspective consistency loss and feature matching loss, and train for 200,000 iterations; then use dynamic video data and 3D reconstruction data to train the temporal dimension transformer, optimize the temporal smoothness loss and 3D consistency loss, and train for 300,000 iterations; finally, use four-dimensional sequence data to perform end-to-end joint optimization on the entire model, and train for 100,000 iterations. In each training stage, use the Adam optimizer with an initial learning rate set to 0.0001.

[0084] In the single-image four-dimensional dynamic human body video generation method based on the diffusion transformer according to the embodiments of the present application, the generation process of the four-dimensional dynamic human body is decomposed into three dimensions: space, perspective, and time domain. Associations between different dimensions are hierarchically constructed in each dimension, reducing the complexity of the spatio-temporal full attention mechanism from four dimensions to two dimensions, thereby greatly reducing the computational cost and the occupancy of video memory space, and finally achieving a high-quality and highly realistic four-dimensional dynamic human body generation effect. This embodiment can be applied and extended to the generation of various dynamic human body scenarios, including using only a single human image, freely controlling the generation of human actions and the generation of the camera in the scene. The achieved generation effect can be widely applied to various applications, including film production, holographic video generation, high-precision digital human generation, AR / VR interaction, remote conferencing, and other fields.

[0085] To implement the above embodiments, the present invention also proposes a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method described in the above embodiments is implemented.

[0086] To implement the above embodiments, the present invention also proposes a non-temporary computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor, the method described in the above embodiments is implemented.

[0087] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0088] In addition, the terms "first" and "second" are used only for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0089] Any process or method description shown in a flowchart or described in other ways herein can be understood as representing a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process. And the scope of the preferred embodiments of the present application includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present application belong.

[0090] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.

[0091] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0092] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0093] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0094] The above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present application.

Claims

1. A method for generating a four-dimensional dynamic human body video from a single image based on a diffusion transformer, characterized in that, Including: Obtain any human action video; Input the human action video into a hierarchical four-dimensional diffusion transformer and update the diffusion transformer, where the diffusion transformer includes a spatial dimension transformer, a perspective dimension transformer, and a temporal dimension transformer; Perform self-attention decomposition operation in the diffusion transformer; Obtain a single human image and extract the human identity feature, camera pose feature, and human body action feature in the image; Inject the extracted features into the corresponding transformers respectively, generate a video sequence through the diffusion transformer, and perform spatio-temporal consistency optimization on the generated video sequence to obtain the target video sequence.

2. The method according to claim 1, characterized in that, The step of inputting the human action video into a hierarchical four-dimensional diffusion transformer and updating the diffusion transformer includes: Through the spatial dimension transformer, divide each input video frame spatially, and calculate the correlation between spatial positions; Through the perspective dimension transformer, calculate the features between different perspectives, and construct the corresponding relationship between any two perspectives within the range of 0 degrees to 360 degrees; Through the perspective dimension transformer, establish the correlation between any two video frames at different times in the video sequence; Stack the multi-dimensional transformers through residual connections and layer normalization to form a hierarchical four-dimensional diffusion transformer.

3. The method according to claim 1, characterized in that, The step of performing self-attention decomposition operation in the diffusion transformer includes: In the spatial dimension transformer, divide the input image into feature grids of a preset size, each feature grid position contains a feature vector of a preset size, and capture the long-range spatial dependence relationship in the image by calculating the self-attention between spatial positions; In the perspective dimension transformer, adopt the multi-head self-attention mechanism, establish the correlation mapping between perspectives by calculating the similarity of query-key-value pairs between different perspectives, and construct the feature corresponding relationship between any two perspectives within the range of 0 degrees to 360 degrees; In the temporal dimension transformer, adopt the multi-head self-attention mechanism, divide the time window by sampling window in time series, and perform temporal modeling between adjacent and non-adjacent frames within the time window in the video sequence.

4. The method according to claim 1, characterized in that Extracting the human identity feature, camera pose information, and human body action feature in the image includes: Use the pre-trained CLIP visual encoder to process the input image and extract the human identity feature vector, where the human identity feature vector includes the facial feature, body type feature, and clothing feature of the person; Use the variational autoencoder to encode the input image to obtain the human body action feature vector, where the human body action feature vector includes the pose information, joint position information, and motion state information of the human body; Use a multi-layer perceptron combined with position encoding to process the camera parameters, and convert the internal parameter matrix and external parameter matrix of the camera into a camera pose feature vector, where the camera pose feature vector includes the focal length, principal point coordinates, rotation matrix, and translation vector of the camera.

5. The method according to claim 1, characterized in that, The step of injecting the extracted features into the corresponding transformers respectively includes: Convert the dimension of the human identity feature through a linear projection layer, and use the additive attention mechanism to inject it into each feature grid position of the spatial dimension transformer to ensure that the human identity feature of the generated image is consistent with the input image; Inject the camera pose features into the view dimension transformer through the cross-attention mechanism to guide the image generation process under different views; Copy the human action features to obtain the first feature and the second feature. Inject the first feature into the spatial dimension transformer after linear transformation to guide the pose generation of each frame, and process the second feature through the temporal cross-attention mechanism and then inject it into the temporal dimension transformer to control the motion changes between consecutive frames; The injecting the extracted features into the corresponding transformers respectively further includes: All feature injection processes adopt a residual connection structure, and LayerNorm is used for feature normalization before and after injection.

6. The method according to claim 1, wherein The optimizing the spatio-temporal consistency of the generated video sequence to obtain the target video sequence includes: Initialize Gaussian random noise in the spatio-temporal dimension, and the noise dimension is consistent with the target video sequence; During the diffusion denoising process, adopt a sliding window strategy for sampling, including: Set a sliding window of a preset size in the temporal dimension. For the current frame to be generated, consider the noise states of several frames before and after it within the preset size at the same time; Set a sliding window of a preset number of views in the view dimension. For the current view, consider the noise states of several adjacent views within the preset number at the same time; For the noise signals within each spatio-temporal sampling window, calculate the weight coefficients through the attention mechanism, and adaptively adjust the weight values according to the spatio-temporal distance, so that the samples closer to the current moment or the current view obtain higher weights; Fuse the weighted multiple noise signals to obtain the updated denoised signal.

7. The method according to claim 1, characterized in that, Before inputting the human action video into the hierarchical four-dimensional diffusion transformer, train the diffusion transformer using a multi-dimensional human dataset, including: Train the spatial dimension transformer using static image data to optimize the reconstruction loss and the perceptual loss; Train the view dimension transformer using multi-view data to optimize the view consistency loss and the feature matching loss; Train the temporal dimension transformer using dynamic video data and 3D reconstruction data to optimize the temporal smoothness loss and the 3D consistency loss; Perform end-to-end joint optimization on the diffusion transformer using four-dimensional sequence data; Training the diffusion transformer using a multi-dimensional human dataset further includes: In each training stage, adopt the Adam optimizer.

8. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the method according to any one of claims 1-7.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1-7.

Citation Information

Cited By

  • Personalized video generation model training and reasoning method based on spatio-temporal representation alignment

    CN121074561A

  • Personalized video generation model training and inference method based on spatio-temporal representation alignment

    CN121074561B