Personalized video generation model training and inference method based on spatio-temporal representation alignment

By combining spatial and temporal transformers in the diffusion video model and training with intermediate and optical flow features, the problem of video generation discrepancies caused by reconstruction loss is solved, and higher quality video generation is achieved.

CN121074561BActive Publication Date: 2026-03-17PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511616564.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2026-03-17
Estimated Expiration
2045-11-06

AI Technical Summary

Technical Problem

In existing technologies, reconstruction loss results in differences between the generated video and the input data in terms of subject identity and motion patterns, which cannot meet the requirements of customized videos.

Method used

By inputting image data into the spatial transformer and self-supervised encoder of the diffusion video model, intermediate and global coding features are extracted, and the spatial transformer is trained; by inputting video data into the optical flow encoder, intermediate and standard optical flow features are extracted, and the temporal transformer is trained. The combined components are used to obtain the trained diffusion video model.

Benefits of technology

The model enhances the fidelity of the model in reproducing the subject and motion patterns in the generated video, thereby improving the accuracy and naturalness of the video generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121074561B_ABST
    Figure CN121074561B_ABST
Patent Text Reader

Abstract

This application proposes a training and inference method for a personalized video generation model based on spatiotemporal representation alignment. The method includes: inputting first image data into the spatial transformer of a diffusion video model to obtain intermediate features corresponding to the first image data; training the spatial transformer based on the intermediate features and global encoding features; inputting first video data into the diffusion video model to obtain reconstructed video data corresponding to the first video data output by the diffusion video model; and training a temporal transformer based on the intermediate optical flow features and standard optical flow features corresponding to each frame of video data. In this embodiment, the spatial transformer is trained using global encoding features to enhance the model's global supervision and interaction capabilities, and the temporal transformer is trained using standard optical flow features to enhance the model's ability to capture motion at the object level, thereby improving the fidelity of the subject and motion patterns in the model-generated video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, specifically relating to a training and inference method for a personalized video generation model based on spatiotemporal representation alignment. Background Technology

[0002] With the development of multimedia technology, customized video generation technology plays a crucial role in the fields of computer vision and graphics. Customized video generation technology aims to generate videos that faithfully preserve the appearance of the subject in the reference image while maintaining the temporal continuity of the reference video.

[0003] In related technologies, significant progress has been made in generating videos with consistent subject identity and motion patterns by relying on reconstruction loss to learn the latent space representation of input data. However, due to its inherent limitations, the reconstruction loss results in some differences between the generated video and the input data in terms of subject identity and motion patterns, which cannot meet the requirements of customized videos. Summary of the Invention

[0004] This application proposes a training and inference method for a personalized video generation model based on spatiotemporal representation alignment. This method can solve the technical problem that current reconstruction loss methods, due to their inherent limitations, result in differences between the generated videos and the input data in terms of subject identity and motion patterns, thus failing to meet the requirements of customized videos.

[0005] The first aspect of this application proposes a method for training a personalized video generation model based on spatiotemporal representation alignment, including:

[0006] The first image data is input into the spatial transformer of the diffusion video model to obtain the intermediate features corresponding to the first image data. The first image data is any image data in the input dataset, and the intermediate features are the feature maps of the largest scale among the feature maps of different scales output by the spatial transformer.

[0007] The first image data is input into a trained self-supervised encoder to obtain global coding features corresponding to the first image data. The global coding features include the context information of the first image data. The intermediate features and the global coding features are used to extract the subject information in the first image data.

[0008] The spatial transformer is trained based on the intermediate features and the global encoded features;

[0009] Input the first video data into the diffusion video model, and obtain the reconstructed video data corresponding to the first video data output by the diffusion video model. The first video data is any video data in the input dataset.

[0010] The first video data and the reconstructed video data are respectively input into a pre-trained optical flow encoder to obtain intermediate optical flow features and standard optical flow features corresponding to each frame of video data in the first video data. The intermediate optical flow features and the standard optical flow features are used to capture the motion information of the first video data.

[0011] The time transformer of the diffusion video model is trained based on the intermediate optical flow features and standard optical flow features corresponding to each frame of video data.

[0012] A well-trained diffusion video model is obtained by combining the trained spatial transformer and the trained temporal transformer.

[0013] An embodiment of the second aspect of this application provides a personalized video generation model inference method based on spatiotemporal representation alignment, including:

[0014] The target dataset is input into the trained diffusion video model, which is obtained by the personalized video generation model training method based on spatiotemporal representation alignment described in the first aspect. The target dataset includes target image data and target video data.

[0015] The target image features of the target image data are extracted by the spatial transformer in the trained diffusion video model, and the target optical flow features of each frame of the target image data in the target video data are extracted by the temporal transformer in the trained diffusion video model. The target image features are used to extract the subject information in the target image data, and the target optical flow features are used to capture the motion information of the target image data.

[0016] The trained diffusion video model outputs the target video based on the target image features and the target optical flow features of each frame of target image data.

[0017] An embodiment of the third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in the first aspect above.

[0018] An embodiment of the fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the method described in the first aspect above.

[0019] The technical solutions provided in this application embodiment have at least the following technical effects or advantages:

[0020] This application proposes a training and inference method for a personalized video generation model based on spatiotemporal representation alignment. The method includes: inputting first image data into a spatial transformer of a diffusion video model to obtain intermediate features corresponding to the first image data. The first image data is any image data in the input dataset, and the intermediate features are the largest-scale feature map among the feature maps of different scales output by the spatial transformer; inputting the first image data into a trained self-supervised encoder to obtain global encoding features corresponding to the first image data. The global encoding features include contextual information of the first image data. The intermediate features and global encoding features are used to extract subject information from the first image data; and training a spatial transformer based on the intermediate features and global encoding features. The process involves several steps: inputting first video data into a diffusion video model to obtain reconstructed video data corresponding to the first video data output by the diffusion video model (the first video data being any video data from the input dataset); inputting the first video data and the reconstructed video data into a pre-trained optical flow encoder to obtain intermediate optical flow features and standard optical flow features corresponding to each frame of the first video data (the intermediate and standard optical flow features are used to capture motion information of the first video data); training the temporal transformer of the diffusion video model based on the intermediate and standard optical flow features corresponding to each frame of the video data; and combining the trained spatial transformer and the trained temporal transformer to obtain the trained diffusion video model. In this embodiment, the spatial transformer is trained using global encoding features to enhance the model's global supervision and interaction capabilities, and the temporal transformer is trained using standard optical flow features to enhance the model's ability to capture motion at the object level, thereby improving the fidelity of the subject and motion patterns in the model-generated video.

[0021] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0022] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0023] Figure 1 A flowchart of a personalized video generation model training method based on spatiotemporal representation alignment provided in an embodiment of this application is shown.

[0024] Figure 2 A flowchart of a space transformer training method according to an embodiment of this application is shown;

[0025] Figure 3A flowchart of another space transformer training method provided in one embodiment of this application is shown;

[0026] Figure 4 A flowchart of a time transformer training method according to an embodiment of this application is shown;

[0027] Figure 5 A flowchart of another space transformer training method provided in one embodiment of this application is shown;

[0028] Figure 6 A flowchart of a personalized video generation model inference method based on spatiotemporal representation alignment provided in an embodiment of this application is shown;

[0029] Figure 7 This illustration shows a schematic diagram of a model structure provided in one embodiment of this application;

[0030] Figure 8 This illustration shows a schematic diagram of the structure of a personalized video generation model training device based on spatiotemporal representation alignment according to an embodiment of this application;

[0031] Figure 9 This illustration shows a schematic diagram of the structure of a personalized video generation model training device based on spatiotemporal representation alignment according to an embodiment of this application;

[0032] Figure 10 This illustration shows a schematic diagram of the structure of an electronic device according to an embodiment of this application;

[0033] Figure 11 A schematic diagram of a storage medium provided in one embodiment of this application is shown. Detailed Implementation

[0034] Exemplary embodiments of this application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of this application are shown in the drawings, it should be understood that this application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of this application and to fully convey the scope of this application to those skilled in the art.

[0035] It should be noted that, unless otherwise stated, the technical or scientific terms used in this application shall have the ordinary meaning as understood by one of ordinary skill in the art to which this application pertains.

[0036] The personalized video generation model training method based on spatiotemporal representation alignment proposed in this application can be executed by a computing device. The computing device can be a server, such as a single server, multiple servers, a server cluster, a cloud computing platform, etc. Optionally, the computing device can also be a terminal device, such as a mobile phone, tablet computer, game console, portable computer, desktop computer, advertising machine, all-in-one machine, etc. This application does not limit the type or number of computing devices.

[0037] Related techniques, which rely on reconstruction loss to learn latent space representations, have made significant progress in generating videos with consistent subject identity and motion patterns. However, for subject learning, this approach neglects global spatial information, such as overall structure and semantic coherence, because the reconstruction loss focuses on local images and lacks global supervision and interaction.

[0038] Furthermore, while existing methods utilize reconstruction loss and techniques such as interpolation-based methods or identity adapters to decouple the appearance and temporal motion of videos, the reconstruction loss only focuses on local frame consistency, limiting the model's ability to capture object-level motion. Moreover, during video generation based on reconstruction loss, the temporal and appearance features of the video data remain intertwined, making motion generation highly dependent on appearance and resulting in incomplete decoupling. Therefore, a personalized video generation method based on spatiotemporal representation is urgently needed to address these technical problems.

[0039] To address the aforementioned issues, this application proposes a training and inference method for a personalized video generation model based on spatiotemporal representation alignment. This method includes: inputting first image data into a spatial transformer of a diffusion video model to obtain intermediate features corresponding to the first image data. The first image data is any image data from the input dataset, and the intermediate features are the largest-scale feature map among the feature maps of different scales output by the spatial transformer; inputting the first image data into a trained self-supervised encoder to obtain global encoded features corresponding to the first image data. The global encoded features include contextual information of the first image data. The intermediate features and global encoded features are used to extract subject information from the first image data; based on the intermediate features and global encoded features... The process involves training a spatial transformer; inputting first video data into a diffusion video model to obtain reconstructed video data corresponding to the first video data output by the diffusion video model, where the first video data is any video data from the input dataset; inputting the first video data and the reconstructed video data into a pre-trained optical flow encoder to obtain intermediate optical flow features and standard optical flow features corresponding to each frame of the first video data, which are used to capture motion information of the first video data; training a temporal transformer for the diffusion video model based on the intermediate and standard optical flow features corresponding to each frame of the video data; and combining the trained spatial transformer and the trained temporal transformer to obtain a trained diffusion video model. This embodiment trains the spatial transformer using global encoding features to enhance the model's global supervision and interaction capabilities, and trains the temporal transformer using standard optical flow features to enhance the model's ability to capture motion at the object level, thereby improving the fidelity of the subject and motion patterns in the model-generated video.

[0040] The following describes, with reference to the accompanying drawings, a training method for a personalized video generation model based on spatiotemporal representation alignment proposed according to an embodiment of this application.

[0041] See Figure 1 The method specifically includes the following steps:

[0042] S101. Input the first image data into the spatial transformer of the diffusion video model to obtain the intermediate features corresponding to the first image data.

[0043] The first image data is any image data in the input dataset, and the intermediate features are the largest scale feature map among the feature maps of different scales output by the spatial transformer.

[0044] The input dataset typically includes image data and video data. To quickly obtain the image and video data from the input dataset, the corresponding image and video data can be filtered from existing mature datasets.

[0045] In some embodiments, the method further includes:

[0046] Extract target image data from an image dataset, where the subject is classified as motion.

[0047] Extract target video data with one subject from the video dataset;

[0048] The target image data and target video data are preprocessed to obtain the input dataset.

[0049] In some embodiments, images of all moving categories, such as plush toys, cars, and pets, are selected from the existing DreamBooth image customization generation task dataset. These categories possess dynamic characteristics, better supporting the customization needs of the generation task. Simultaneously, images of static categories, such as vases, sneakers, and beverages, are removed to avoid static features in the dataset interfering with feature extraction during the subject learning phase. Furthermore, single-object motion videos, such as car driving, animal activity, and human movement, are selected from the TGVE video dataset. These videos provide clear motion pattern references, facilitating temporal representation modeling during the motion learning phase. In addition, conceptual materials related to movement are collected from the internet, specifically including sports (such as golf and weightlifting) and musical instrument performances (such as playing guitar and flute). This data further enriches the diversity of motion patterns, providing a wider range of application scenarios for the generation task.

[0050] The preprocessing process can be implemented as follows: after the data screening is completed, an automated script is used to clean the image and video data, remove low-quality or noisy data, and ensure the purity of the final dataset.

[0051] The diffusion video model consists of two Markov chains: a forward noise reduction process and a reverse denoising process. The forward process gradually adds noise to the clean video data over multiple time steps, while the reverse process uses a neural network to predict the noise injected at the forward time step and gradually recover the original video.

[0052] In this embodiment of the application, the generation capability of the diffusion model can be used to replace or modify specific parts of the video, thereby achieving more flexible video editing effects.

[0053] Spatial Transformers focus on spatial features within a single frame, such as object shape, texture, and color distribution. They process each frame independently, effectively extracting and processing intra-frame details, thus providing a more accurate feature representation for subsequent temporal modeling.

[0054] In general, spatial transformers are typically U-Net architectures, which can output intermediate features at different scales. In order to maintain the same size as the global encoded features output by the subsequent self-supervised encoder, since the features output by the trained self-supervised encoder have not undergone a downsampling process, the size of the output feature map is the same as the size of the first image data. Therefore, the largest scale feature map among the different scales of the spatial transformer output can be obtained as an intermediate feature, so that the spatial transformer can be trained using the intermediate features and global encoded features under the condition of the same size.

[0055] S102. Input the first image data into the trained self-supervised encoder to obtain the global encoding features corresponding to the first image data.

[0056] The global encoding features include contextual information from the first image data, while intermediate features and global encoding features are used to extract subject information from the first image data.

[0057] Understandably, the features output by the spatial transformer are primarily used to extract the main subject information from the first image data.

[0058] Generally, compared to traditional local feature extraction methods, self-supervised encoders can learn richer subject semantic information, that is, the global encoded features output by the self-supervised encoder include the contextual information of the first image data.

[0059] S103. Train the spatial transformer based on intermediate features and global encoded features.

[0060] Furthermore, a corresponding loss function can be constructed based on the difference between intermediate features and global encoded features, and a spatial transformer can be trained based on this loss function, enabling the spatial transformer to learn richer subject semantic information.

[0061] S104. Input the first video data into the diffusion video model and obtain the reconstructed video data corresponding to the first video data output by the diffusion video model.

[0062] The first video data is any video data in the input dataset.

[0063] S105. Input the first video data and the reconstructed video data into the pre-trained optical flow encoder respectively to obtain the standard optical flow features and intermediate optical flow features corresponding to each frame of video data in the first video data.

[0064] Intermediate optical flow features and standard optical flow features are used to capture motion information from the first video data.

[0065] Optical flow features are vector fields describing pixel motion in an image sequence, capable of capturing motion information in video. By training a temporal transform based on optical flow features, the model can focus on learning motion patterns in the video. This approach is particularly suitable for scenarios requiring accurate motion modeling, such as action recognition or behavior analysis. By leveraging optical flow features, the model can better understand and predict motion changes in video, thereby improving the naturalness and smoothness of the generated video.

[0066] The first video data is input into the diffusion video model. At each time step of the diffusion model, by reversing the noise addition process of the diffusion model, the model can restore the video features from the latent space to the pixel space. That is, the diffusion video model can encode and decode the first video data to obtain the reconstructed video data corresponding to the first video data.

[0067] Furthermore, the reconstructed video data is input into the pre-trained optical flow encoder to obtain the intermediate optical flow features corresponding to each frame of video data, and the first video data is input into the pre-trained optical flow encoder to obtain the standard optical flow features corresponding to each frame of video data.

[0068] S106. A time transformer for training a diffusion video model based on the intermediate optical flow features and standard optical flow features corresponding to each frame of video data.

[0069] Furthermore, a loss function can be constructed based on the difference between the intermediate optical flow features and the standard optical flow features corresponding to each frame of video data to train the time transformer of the diffusion video model, learn the dynamic change law of the temporal pattern, and ensure that the generated video is consistent with the reference in motion mode.

[0070] In the diffusion video model, the Temporal Transformer is responsible for modeling the temporal dimension of the feature sequence processed by the Spatial Transformer. The Temporal Transformer uses a self-attention mechanism to capture long-range dependencies between frames, thereby generating temporally coherent video content.

[0071] S107. Based on the trained spatial transformer and the trained temporal transformer, a trained diffusion video model is obtained by combining them.

[0072] Based on the trained spatial transformer and the trained temporal transformer, the two are stitched together to obtain the trained diffusion video model.

[0073] This application proposes a training and inference method for a personalized video generation model based on spatiotemporal representation alignment. The method includes: inputting first image data into a spatial transformer of a diffusion video model to obtain intermediate features corresponding to the first image data. The first image data is any image data in the input dataset, and the intermediate features are the largest-scale feature map among the feature maps of different scales output by the spatial transformer; inputting the first image data into a trained self-supervised encoder to obtain global encoding features corresponding to the first image data. The global encoding features include contextual information of the first image data. The intermediate features and global encoding features are used to extract subject information from the first image data; and training a spatial transformer based on the intermediate features and global encoding features. The process involves several steps: inputting first video data into a diffusion video model to obtain reconstructed video data corresponding to the first video data output by the diffusion video model (the first video data being any video data from the input dataset); inputting the first video data and the reconstructed video data into a pre-trained optical flow encoder to obtain intermediate optical flow features and standard optical flow features corresponding to each frame of the first video data (the intermediate and standard optical flow features are used to capture motion information of the first video data); training the temporal transformer of the diffusion video model based on the intermediate and standard optical flow features corresponding to each frame of the video data; and combining the trained spatial transformer and the trained temporal transformer to obtain the trained diffusion video model. In this embodiment, the spatial transformer is trained using global encoding features to enhance the model's global supervision and interaction capabilities, and the temporal transformer is trained using standard optical flow features to enhance the model's ability to capture motion at the object level, thereby improving the fidelity of the subject and motion patterns in the model-generated video.

[0074] In some embodiments, training a spatial transformer based on intermediate features and global encoded features includes: calculating a relationship graph between the intermediate features and global encoded features; fusing the relationship graph and global encoded features to obtain fused features; generating a first loss function value based on the fused features and global encoded features; and training the spatial transformer based on the first loss function value.

[0075] Since the encoder of the diffusion video model is generally U-Net architecture, while the self-supervised encoder is generally ViT architecture, in order to solve the problem of differences in receptive field between encoders of different architectures and ensure that the model accurately captures the subject features at different scales, in the process of training the spatial transformer based on intermediate features and global coding features, in addition to directly calculating the loss between the two, we can also calculate the relationship graph between intermediate features and target features and fuse these two features in a block-by-block manner.

[0076] In some embodiments, the process of training the spatial transformer based on intermediate features and global encoded features is as follows: Figure 2 As shown:

[0077] First, intermediate features and global encoded features are obtained. The intermediate features and global encoded features are then fused and input into a normalized unit to obtain a relational graph.

[0078] In some embodiments, the relationship graph between intermediate features and global encoded features is calculated as shown in equation (1):

[0079] (1)

[0080] in, This is a graph showing the relationship between intermediate features and global encoded features. For normalization function, This represents a learnable linear layer. This represents the intermediate features generated. For global encoding features, This is the transpose of the globally encoded features. is the scaling factor, where is the dimension of the feature vector. The scaling factor is used to prevent the dot product result from becoming too large, which could lead to gradient explosion.

[0081] Furthermore, the relationship graph and global coding features are fused together.

[0082] The process of fusing the relationship graph and global encoding features to obtain the fused features is shown in equation (2):

[0083] (2)

[0084] This is a feature of fusion.

[0085] Based on the fusion features and global encoding features, the process of generating the first loss function value can be shown as strengthening the model's ability to capture the main features through the spatial representation alignment module. The first loss function value can be obtained through the similarity loss function. The similarity loss function is used to shorten the distance between the target feature and the fusion feature. The loss function is shown in Equation (3):

[0086] (3)

[0087] in, The first loss function value, It is the set of parameters of the model. Represents all possible The set, where N is the number of image data. It is a similarity function used to measure the similarity between generated samples and target samples. Samples for globally encoded features For samples with fused features, The initial encoding feature typically represents the original latent representation of the data in a diffusion model. s is a condition variable, which can be a text description. Generally, a text description exists to guide the model in generating specific content. The time step represents the current step in the diffusion process within the diffusion model.

[0088] In some embodiments, the method further includes: acquiring a reconstructed image obtained by a diffusion video model based on first image data, the reconstructed image including a predicted mask; inputting the first image data into a trained mask generation model to obtain a target mask; generating a second loss function value based on the predicted mask and the target mask; and training a spatial transformer based on the first loss function value, including: adjusting the spatial low-rank matrix of the spatial transformer based on the first loss function value and the second loss function value, continuing training until a first training completion condition is met, and obtaining a trained spatial transformer.

[0089] The first training completion condition may be that the number of training sessions reaches a first preset threshold, or that either the first loss function value or the second loss function value reaches the first preset threshold. The first preset number of training sessions and the first preset threshold can be flexibly set based on the actual situation.

[0090] In some embodiments, in addition to training the spatial transformer using a pre-trained image self-supervised encoder, the spatial transformer can also be trained using a corresponding mask generation model.

[0091] Specifically, a reconstructed image obtained by the diffusion video model based on the first image data is acquired. The reconstructed image includes a prediction mask. The first image data is then input into a trained mask generation model to obtain a target mask. The prediction mask represents the predicted subject information corresponding to the first image data, and the target mask represents the target subject information corresponding to the first image data. Training the spatial transformer based on the target mask and the prediction mask makes the prediction mask output by the trained spatial transformer more accurate.

[0092] In some embodiments, the value of the second loss function can be obtained by equation (4):

[0093] (4)

[0094] in, Let the mean square error loss function be the mask. The expectation operator represents averaging over all possible input data. Let r and t be derived from the standard normal distribution. Random variables sampled in the middle, where It is the identity matrix. This is typically used to simulate noise or uncertainty in data. The time step represents the current step in the diffusion process. These are conditional variables, typically textual descriptions, used to guide the model in generating specific content. The true latent space representation represents the true representation of the input data in the model's latent space. The latent space representation of the model's predictions represents the model's performance under given conditions. and time step The latent variables to be predicted, This is a mask matrix used to specify which specific regions of the input data the loss function should focus on. Elements in the mask matrix are typically 0 or 1, where 1 represents the region of interest and 0 represents the region to be ignored. It is the L2 norm, representing the square root of the sum of squares of the elements of a vector or matrix, used to calculate the distance between two vectors.

[0095] Correspondingly, the definition of the overall loss function of the above space transformer is shown in equation (5):

[0096] (5)

[0097] Where L1 is the overall loss function value, This is the value of the second loss function. The first loss function value, It is a hyperparameter used to balance the alignment loss of the subject representation.

[0098] In some embodiments, low-rank decomposition techniques can be used to update the weight matrix of the space transformer. This yields a trained spatial transformer.

[0099] The weight matrix can be obtained through equation (6):

[0100] (6)

[0101] in, This represents the original weights of the pre-trained model. These are the weights of the low-rank matrix, used to adjust the output of the space transformer. and Denotes the low-rank factor, where Much smaller than the original dimension and Low-rank matrix fine-tuning requires fewer computational resources and can be easily deployed as a plugin for pre-trained models.

[0102] Furthermore, the low-rank matrix BA of the trained spatial transformer can be obtained by minimizing the overall loss function value.

[0103] In some embodiments, the training process of the spatial transformer is as follows: Figure 3 As shown:

[0104] The diffusion video model includes residual network blocks and a spatial transformer. The residual network block is a commonly used building block in deep learning, addressing the vanishing gradient problem in deep network training by introducing residual connections. Residual connections allow the network to learn residual mappings instead of directly learning unmapped features, enabling more efficient training of deeper architectures.

[0105] The first image data is input into the residual network block and spatial transformer to obtain feature maps at different scales. The first image data is then input into a trained self-supervised encoder to obtain the global encoded features corresponding to the first image data. Figure 2 The feature alignment process shown yields the first loss function value.

[0106] Furthermore, the predicted mask of the first image data is obtained through the diffusion video model, the first image data is input into the trained mask generation model to obtain the target mask, and the second loss function value is obtained through equation (4). Thus, the spatial transformer is trained using the first loss function value and the second loss function value.

[0107] In some embodiments, training a time transformer for a diffusion video model based on intermediate optical flow features and standard optical flow features corresponding to each frame of video data includes: calculating a third loss function value for a first intermediate optical flow feature and a first standard optical flow feature of any frame of video data; and training the time transformer based on the third loss function value.

[0108] In some embodiments, during the motion learning phase, the noise-free video is obtained through one-step inference by reversing the noise-adding process of the diffusion model. During this process, the diffusion video model encodes and decodes the first video data to obtain the reconstructed video. Furthermore, the intermediate optical flow features of each frame of the reconstructed video data and the standard optical flow features of each frame of the first video data can be used to train the time transformer to improve the time transformer's ability to extract global temporal motion features.

[0109] For video data frames via optical flow encoder The process of extracting standard optical flow features from adjacent frames is shown in equation (7):

[0110] (7)

[0111] in, It represents the absolute displacement of pixels in the horizontal and vertical directions between two adjacent frames of a video.

[0112] Furthermore, at each time step of the diffusion video model, by reversing the noise addition process of the diffusion model, the model can reconstruct video features from the latent space to the pixel space and further extract intermediate optical flow features, ensuring the temporal consistency of the generated video. Specifically, for the video features in the latent space... Using VAE decoder The decoding process is shown in equation (8):

[0113] (8)

[0114] in, Let N be the set of reconstructed data samples, representing the N samples reconstructed from the latent space back into the data space. The decoder, the neural network part of the VAE, is responsible for transforming the representation in the latent space back into the data space. Let be the latent variable of the i-th sample at time step t, representing the noisy latent representation after the diffusion process. The noise sampled from the standard normal distribution N(0,I) is used to introduce randomness into the latent space. Let be the cumulative variance parameter at time step t, which is the cumulative product of the variance parameters αt at each step during the diffusion process. To represent the ratio of the standard deviation between the raw data and the noise at time step t, To represent the standard deviation of the noise at time step t, This is a latent variable. The purpose of the denoising process is to recover the original latent representation from the noisy latent representation.

[0115] Furthermore, an optical flow encoder is used. To extract intermediate optical flow features from adjacent frames during the diffusion process, as shown in Equation (9):

[0116] (9)

[0117] in, This represents the absolute displacement of pixels in the horizontal and vertical directions between two adjacent frames of the decoded video.

[0118] By utilizing the temporal representation alignment module, the model's ability to capture temporal motion is enhanced through the alignment loss of intermediate optical flow features and standard optical flow features. Specifically, this is achieved by calculating the value of the third loss function and training the aforementioned time transformer based on the value of the third loss function.

[0119] The value of the third loss function can be obtained through equation (10):

[0120]

[0121] in, This is the value of the third loss function.

[0122] In some embodiments, the method further includes: acquiring first anchor video frame data of the first video data and second anchor video frame data of the reconstructed image data corresponding to the first video data, wherein the first anchor video frame data is any video frame data in the first video data, and the second anchor video frame data is video frame data in the reconstructed image data with the same number of frames as the first anchor video frame data; inputting the first video data and the reconstructed image data into an encoder respectively to acquire the target latent space representation of each frame of video data in the first video data, the first anchor feature of the first anchor video frame data, the predicted latent space representation of each frame of video data in the reconstructed image data, and the second anchor feature corresponding to the second anchor video frame data; The target motion features of each frame of video data in the first video data are determined based on the target latent space representation and the first anchor feature. The predicted motion features of each frame of video data in the reconstructed image data are determined based on the predicted latent space representation and the second anchor feature. The fourth loss function value is calculated based on the predicted motion features and the target motion features. The time transformer of the diffusion video model is trained based on the third loss function value, including: adjusting the time low-rank matrix of the time transformer based on the third loss function value and the fourth loss function value, and continuing training until the second training completion condition is met to obtain the trained time transformer.

[0123] The second training completion condition can be that the number of training sessions reaches a second preset threshold, or that any one of the third or fourth loss function values ​​reaches the second preset threshold. The second preset threshold and the second preset threshold can be flexibly set based on the actual situation.

[0124] First, it is necessary to obtain the reconstructed video data corresponding to the first video data, and obtain the first anchor video frame data of the first video data and the second anchor video frame data of the reconstructed image data corresponding to the first video data. The second anchor video frame data is the video frame data in the reconstructed image data that has the same frame number as the first anchor video frame data. For example, the first anchor video frame data is the 10th frame video data in the first video data, and the second anchor video frame data is the 10th frame video data in the reconstructed image data.

[0125] The first video data is input into a pre-trained encoder to obtain the target latent space representation of each frame of video data in the first video data. The first anchor feature corresponding to the first anchor video frame data in the first video data is obtained. The target motion feature of each frame of video data in the first video data is determined based on the target latent space representation and the first anchor feature.

[0126] The reconstructed video data is input into a pre-trained encoder to obtain the predicted latent space representation of each frame of video data in the reconstructed data. The second anchor feature, corresponding to the second anchor video frame data (which is the video frame data in the reconstructed image data with the same frame number as the first anchor video frame data), is then obtained. Based on the predicted latent space representation and the second anchor feature of each frame of video data in the reconstructed video data, the predicted motion features of each frame of video data are determined.

[0127] Furthermore, the fourth loss function value is calculated based on the predicted motion features and the target motion features.

[0128] In some embodiments, the process of training the time transformer based on the intermediate optical flow features and standard optical flow features corresponding to each frame of video data is as follows: Figure 4 As shown.

[0129] First, the first video data is input into the diffusion video model for denoising, that is, the diffusion video model is encoded. The denoised first video data is input into the decoder to obtain the reconstructed video data. The reconstructed video data is input into the pre-trained optical flow encoder to obtain the intermediate optical flow features of each frame of video data in the first video data.

[0130] Furthermore, the first video data is input into the pre-trained optical flow encoder to obtain the standard optical flow features of each frame of video data in the first video data. Then, the time transformer is trained by using Equation (10) through the intermediate optical flow features and standard optical flow features corresponding to each frame of video data.

[0131] In some embodiments, optical flow features include appearance features and motion features. However, the temporal features and appearance features are intertwined, making motion generation highly dependent on appearance and resulting in incomplete decoupling. Therefore, it is necessary to decouple the appearance features and motion features of the video in the latent space.

[0132] The anchor feature serves as a baseline, helping the model understand the changes in each frame relative to the anchor frame. These changes can be variations in spatial location, pose, or other motion features.

[0133] The first anchor feature can be obtained as follows: select any frame of video data from the first video data, input the frame of video data into the encoder, and obtain the first anchor feature. Correspondingly, select a frame of video data from the reconstructed image data with the same number of frames as the first anchor video frame data, input the frame of video data into the encoder, and obtain the second anchor feature.

[0134] After obtaining any anchor feature, the decoupled motion features can be determined based on the anchor features, as shown in Equation (11):

[0135] (11)

[0136] in, The motion features decoupled from the i-th frame of video data. This is a hyperparameter used to control the degree of influence of anchor features during the decoupling process. Anchor features Let be the latent space representation of the i-th frame of video data.

[0137] That is, the motion characteristics of any frame of video data can be determined by the latent space representation and anchor features of that frame.

[0138] The first video data is input into a pre-trained encoder to obtain the target latent space representation of each frame of video data in the first video data. The first anchor feature corresponding to the first anchor video frame data in the first video data is obtained. The target motion feature of each frame of video data in the first video data is determined based on the target latent space representation and the first anchor feature.

[0139] The reconstructed video data is input into a pre-trained encoder to obtain the predicted latent space representation of each frame of video data in the reconstructed data. The second anchor feature corresponding to the second anchor video frame data in the reconstructed video data is obtained. Based on the predicted latent space representation and the second anchor feature of each frame of video data in the reconstructed video data, the predicted motion feature of each frame of video data in the reconstructed video data is determined.

[0140] Furthermore, the time transformer can be trained based on the predicted motion features and target motion features of each frame of video data to improve its decoupling capability.

[0141] The training of the time transformer based on the predicted motion features and target motion features of each frame of video data can be achieved by: calculating the fourth loss function value based on the predicted motion features and target motion features, wherein the fourth loss function value can be obtained by equation (12):

[0142] (12)

[0143] in, This is the fourth loss function value, i.e., the anchor loss function value. The expectation operator represents averaging over all possible input data. These are the initial latent variables, typically representing the original latent representation of the data in a diffusion model. These are conditional variables, which can be textual descriptions used to guide the model in generating specific content. The latent space representation represents the input data in the model's latent space. The time step, in the diffusion model, represents the current step in the diffusion process. This is a transformation function used to convert latent space representations into decoupled motion features. The latent space representation of the model predictions represents the latent variables predicted by the model given the conditions y and time step t. Let be a latent variable at time step t, representing an intermediate state in the diffusion process. The condition encoder is a function that transforms the condition variable y into a form that the model can understand for use during the generation process. It is the L2 norm, representing the square root of the sum of squares of the elements of a vector or matrix, used to calculate the distance between two vectors.

[0144] in, The target motion features of each frame of video data in the first video data. To reconstruct the predicted motion features of each frame of video data.

[0145] The purpose of this loss function is to train the model to better generate video content that meets given conditions by minimizing the difference between the model's predicted latent space representation and the target latent space representation. In this way, the model can learn how to generate video frames with specific motion features based on conditional information.

[0146] Furthermore, the time transformer can be trained using the third and fourth loss function values ​​as the total loss function value:

[0147] The time low-rank matrix of the time transformer, which is adjusted based on the third and fourth loss function values, is shown in equation (13):

[0148] (13)

[0149] in, The total loss function value of the time transformer. The value of the third loss function. This is the value of the fourth loss function. and All of these are hyperparameters.

[0150] Furthermore, the low-rank matrix BA of the trained time transformer can be obtained by minimizing the overall loss function value.

[0151] In some embodiments, the training process of the time transformer is as follows: Figure 5 As shown:

[0152] A diffusion video model can include residual network blocks and a time transformer.

[0153] The first video data is input into a residual network block and a time transformer. The residual network block and time transformer can encode and decode the first video data to obtain the reconstructed video data corresponding to the first video data. The reconstructed video data is input into a pre-trained optical flow encoder to obtain the intermediate optical flow features of each frame of the first video data. The first video data is then input into the pre-trained optical flow encoder to obtain the standard optical flow features of each frame of the first video data.

[0154] The value of the third loss function is further calculated based on equation (10).

[0155] The first video data is input into a pre-trained encoder to obtain the target latent space representation of each frame of video data in the first video data. The first anchor feature corresponding to the first anchor video frame data in the first video data is obtained. The target motion feature of each frame of video data in the first video data is determined based on the target latent space representation and the first anchor feature.

[0156] Furthermore, the reconstructed video data is input into a pre-trained encoder to obtain the predicted latent space representation of each frame of video data in the reconstructed data. The second anchor feature corresponding to the second anchor video frame in the reconstructed video data is then obtained. The second anchor video frame is a video frame in the reconstructed image data with the same frame number as the first anchor video frame; for example, the first anchor video frame is the 10th frame in the first video data, and the second anchor video frame is the 10th frame in the reconstructed image data. Based on the predicted latent space representation and the second anchor feature of each frame of video data in the reconstructed video data, the predicted motion features of each frame are determined.

[0157] The fourth loss function value is further calculated based on equations (11) and (12).

[0158] The time transformer is trained by combining the third and fourth loss function values ​​using equation (13) to obtain a trained time transformer. Then, the trained spatial transformer and the trained time transformer are combined to obtain a trained diffusion video model.

[0159] The following describes, with reference to the accompanying drawings, a personalized video generation model inference method based on spatiotemporal representation alignment proposed according to an embodiment of this application.

[0160] See Figure 6 The method specifically includes the following steps:

[0161] S601. Input the target dataset into the trained diffusion video model.

[0162] The trained diffusion video model is Figure 1The target dataset, which is obtained by training a personalized video generation model based on spatiotemporal representation alignment, includes target image data and target video data.

[0163] S602. Extract target image features from the target image data using the spatial transformer in the trained diffusion video model, and extract target optical flow features from each frame of target image data in the target video data using the temporal transformer in the trained diffusion video model.

[0164] Target image features are used to extract subject information from target image data, while target optical flow features are used to capture motion information from target image data.

[0165] S603. Based on the target image features and the target optical flow features of each frame of target video data, the trained diffusion video model outputs the target video.

[0166] In some embodiments, subject information and target optical flow features of each frame of target video data are extracted from the features of the target image to combine the motion features of the subject information and the target optical flow features of each frame of target video data, thereby generating a target video that includes subject information in the target image and motion features in the target video data.

[0167] This application proposes a personalized video generation model inference method based on spatiotemporal representation alignment. It extracts target image features from target image data by using a spatial transformer trained based on global coding features and extracts target optical flow features from each frame of target image data in target video data by using a temporal transformer trained based on standard optical flow features. The output is a target video that includes subject information in the target image data and motion features in the target video.

[0168] In some embodiments, extracting target image features from target image data using a spatial transformer in a trained diffusion video model, and extracting target optical flow features from each frame of target image data in the target video data using a temporal transformer in a trained diffusion video model, includes: obtaining a first original weight and a second original weight, wherein the first original weight is the original weight corresponding to the target spatial low-rank matrix of the spatial transformer in the trained diffusion video model, and the second original weight is the original weight corresponding to the target spatial low-rank matrix of the temporal transformer in the trained diffusion video model; when the current time step is less than or equal to a preset time step, adjusting the first running weight of the target spatial low-rank matrix to be lower than the first original weight; extracting target image features from the target image data using a spatial transformer with the adjusted first running weight; when the current time step is greater than a preset time step, gradually increasing the first running weight of the target spatial low-rank matrix to the first original weight; extracting target image features from the target image data using a spatial transformer with the adjusted first running weight; and at any time step, extracting target optical flow features from each frame of target image data in the target video data using a temporal transformer with the running weight set to the second running weight.

[0169] In some embodiments, the diffusion video model outputs a video at each time step, and the video output at each time step includes some noise.

[0170] In the diffusion video model, the motion pattern corresponding to the motion feature is recovered first in the earlier time steps. That is, the subject information is not recovered in the earlier time steps, or the need to recover the subject is relatively low. Therefore, the output of the spatial transformer is suppressed in the earlier time steps.

[0171] In some embodiments, the preset time step can be obtained through experiments or can be flexibly set based on actual conditions.

[0172] In the early stages of diffusion, when the time step is less than the preset time step, the model injects a low-rank spatial matrix with smaller weights. This helps the model to freely explore the solution space and generate diverse latent representations. The smaller weights mean that the model has greater freedom to explore different combinations of spatial features during the generation process, which helps to capture the dynamic changes of video content.

[0173] Specifically, the first original weight can be obtained, and if the current time step is less than or equal to the preset time step, the first running weight of the target space low-rank matrix can be adjusted to be lower than the first original weight.

[0174] It should be noted that when the current time step is less than or equal to the preset time step, the first running weight can be increased linearly, or the same running weight can be maintained, or a function (such as an exponential function, logarithmic function, or S-curve) can be used to control the rate of weight increase. However, when the previous time step is less than or equal to the preset time step, the first running weight of the target space low-rank matrix needs to be controlled to be lower than the first original weight.

[0175] When the current time step is longer than the preset time step, gradually increasing the weight of the first run can make the content generated by the model gradually stabilize and focus on the main features, achieving a smooth transition from exploration to refinement. Increasing the weight helps the model to more accurately depict the appearance and details of the main subject, improving the quality of the generated content.

[0176] The gradual increase of the first running weight can be achieved by linearly increasing the first running weight, maintaining the same running weight, or using a function (such as an exponential function, a logarithmic function, or an S-curve) to control the rate of weight increase. Furthermore, if the current time step is greater than a preset time step, the first running weight can be made equal to the first original weight.

[0177] It should be noted that in the output stage of all time steps of the diffusion video model, the synchronous customization of motion features is always emphasized. Therefore, at any time step, the target optical flow features of each frame of target image data in the target video data are extracted by running a time changer with the second running weight.

[0178] In some embodiments, combined with Figure 7 The diffusion video model structure shown illustrates the reasoning method of the aforementioned diffusion-based personalized video generation model aligned with spatiotemporal representations. Figure 7 As shown, the diffusion video model includes: residual network blocks, spatial transformers, and temporal transformers.

[0179] The target dataset, which includes target image data and target video data, is input into the trained diffusion video model.

[0180] The target image features of the target image data are extracted by the spatial transformer in the trained diffusion video model, and the target optical flow features of each frame of the target image data in the target video data are extracted by the temporal transformer in the trained diffusion video model.

[0181] Specifically, when the current time step is less than or equal to a preset time step, the first running weight of the target spatial low-rank matrix is ​​adjusted to be lower than the first original weight; the target image features of the target image data are extracted by the spatial transformer after adjusting the first running weight. When the current time step is greater than the preset time step, the first running weight of the target spatial low-rank matrix is ​​gradually increased to the first original weight; the target image features of the target image data are extracted by the spatial transformer after adjusting the first running weight; at any time step, the target optical flow features of each frame of target image data in the target video data are extracted by the time transformer with the running weight being the second running weight.

[0182] Furthermore, at each time step, subject information is extracted through target image features and motion features are extracted through target optical flow features. The target video, which includes subject information from the target image data and motion features from the target video, is then output by combining the subject information and motion features.

[0183] This application also provides a training device for a personalized video generation model based on spatiotemporal representation alignment, which is used to perform the above-described... Figure 1 The embodiment provides a training method for a personalized video generation model based on spatiotemporal representation alignment. For example... Figure 8 As shown, the device includes an input module 801, a training module 802, and a combination module 803.

[0184] Input module 801 is used to input the first image data into the spatial transformer of the diffusion video model to obtain the intermediate features corresponding to the first image data. The first image data is any image data in the input dataset, and the intermediate features are the feature map of the largest scale among the feature maps of different scales output by the spatial transformer.

[0185] The input module 801 is used to input the first image data into the trained self-supervised encoder to obtain the global coding features corresponding to the first image data. The global coding features include the context information of the first image data. The intermediate features and the global coding features are used to extract the main information in the first image data.

[0186] Training module 802 is used to train the spatial transformer based on the intermediate features and the global encoded features;

[0187] The input module 801 is further configured to input the first video data into the diffusion video model and obtain the reconstructed video data corresponding to the first video data output by the diffusion video model, wherein the first video data is any video data in the input dataset;

[0188] The input module 801 is further configured to input the first video data and the reconstructed video data into the pre-trained optical flow encoder respectively to obtain the intermediate optical flow features and standard optical flow features corresponding to each frame of video data in the first video data. The intermediate optical flow features and the standard optical flow features are used to capture the motion information of the first video data.

[0189] The training module 802 is also used to train the time transformer of the diffusion video model based on the intermediate optical flow features and standard optical flow features corresponding to each frame of video data;

[0190] The combination module 803 is used to combine the trained spatial transformer and the trained temporal transformer to obtain a trained diffusion video model.

[0191] This application proposes a training device for a personalized video generation model based on spatiotemporal representation alignment. The method includes: inputting first image data into a spatial transformer of a diffusion video model to obtain intermediate features corresponding to the first image data. The first image data is any image data in the input dataset, and the intermediate features are the largest-scale feature map among the feature maps of different scales output by the spatial transformer; inputting the first image data into a trained self-supervised encoder to obtain global encoding features corresponding to the first image data. The global encoding features include contextual information of the first image data. The intermediate features and global encoding features are used to extract subject information from the first image data; and training the spatial transformer based on the intermediate features and global encoding features. The method involves inputting first video data into a diffusion video model to obtain reconstructed video data corresponding to the first video data output by the diffusion video model. The first video data can be any video data in the input dataset. The first video data and the reconstructed video data are then input into a pre-trained optical flow encoder to obtain intermediate optical flow features and standard optical flow features corresponding to each frame of the first video data. These intermediate and standard optical flow features are used to capture motion information of the first video data. A temporal transformer of the diffusion video model is trained based on the intermediate and standard optical flow features corresponding to each frame of the video data. The trained spatial transformer and the trained temporal transformer are then combined to obtain the trained diffusion video model. This embodiment trains the spatial transformer using global encoding features to enhance the model's global supervision and interaction capabilities, and trains the temporal transformer using standard optical flow features to enhance the model's ability to capture motion at the object level, thereby improving the fidelity of the subject and motion patterns in the model-generated video.

[0192] In some embodiments, the training module 802 is specifically used for:

[0193] Calculate the relationship graph between the intermediate features and the global encoded features;

[0194] The relationship graph and the global encoding features are fused to obtain fused features;

[0195] Based on the fused features and the global encoding features, a first loss function value is generated;

[0196] The spatial transformer is trained based on the first loss function value.

[0197] In some embodiments, the training module 802 is further specifically used for:

[0198] Obtain the reconstructed image obtained by the diffusion video model based on the first image data, the reconstructed image including a prediction mask;

[0199] The first image data is input into the trained mask generation model to obtain the target mask;

[0200] A second loss function value is generated based on the prediction mask and the target mask;

[0201] Training the spatial transformer based on the first loss function value includes:

[0202] Based on the first loss function value and the second loss function value, adjust the spatial low-rank matrix of the spatial transformer, and continue training until the first training completion condition is met to obtain a trained spatial transformer.

[0203] In some embodiments, the training module 802 is further specifically used for:

[0204] For any frame of video data, calculate the value of the third loss function based on the first intermediate optical flow feature and the first standard optical flow feature;

[0205] The time transformer is trained based on the value of the third loss function.

[0206] In some embodiments, the training module 802 is further specifically used for:

[0207] Acquire the first anchor video frame data of the first video data and the second anchor video frame data of the reconstructed image data, wherein the first anchor video frame data is any frame of video data in the first video data, and the second anchor video frame data is video frame data in the reconstructed image data with the same number of frames as the first anchor video frame data;

[0208] The first video data and the reconstructed image data are respectively input into the encoder to obtain the target latent space representation of each frame of video data in the first video data, the first anchor feature of the first anchor video frame data, the predicted latent space representation of each frame of video data in the reconstructed image data, and the second anchor feature corresponding to the second anchor video frame data.

[0209] The target motion features of each frame of the first video data are determined based on the target latent space representation of each frame of the first video data and the first anchor feature.

[0210] The predicted motion features of each frame of video data in the reconstructed image data are determined based on the predicted latent space representation of each frame of video data in the reconstructed image data and the second anchor feature;

[0211] Calculate the value of the fourth loss function based on the predicted motion features and the target motion features;

[0212] Training the time transformer based on the third loss function value includes:

[0213] Based on the third and fourth loss function values, the time low-rank matrix of the time transformer is adjusted, and training continues until the second training completion condition is met, resulting in a well-trained time transformer.

[0214] In some embodiments, the above-described training apparatus for a personalized video generation model based on spatiotemporal representation alignment further includes a preprocessing module, used for:

[0215] Extract target image data from an image dataset, where the subject is classified as motion.

[0216] Extract target video data with one subject from the video dataset;

[0217] The target image data and the target video data are preprocessed to obtain the input dataset.

[0218] This application also provides an inference device for a personalized video generation model based on spatiotemporal representation alignment, which is used to perform the above-mentioned... Figure 6 The embodiment provides a personalized video generation model inference method based on spatiotemporal representation alignment. For example... Figure 9 As shown, the device includes an input module 901, an extraction module 902, and an output module 903.

[0219] Input module 901 is used to input the target dataset into the trained diffusion video model, wherein the trained diffusion video model is based on... Figure 1 The target dataset, obtained by the training method of the personalized video generation model based on spatiotemporal representation alignment, includes target image data and target video data.

[0220] The extraction module 902 is used to extract target image features of the target image data through the spatial transformer in the trained diffusion video model, and to extract target optical flow features of each frame of target image data in the target video data through the temporal transformer in the trained diffusion video model. The target image features are used to extract the subject information in the target image data, and the target optical flow features are used to capture the motion information of the target image data.

[0221] The output module 903 is used to output the target video based on the target image features and the target optical flow features of each frame of target image data using the trained diffusion video model.

[0222] This application proposes a personalized video generation model inference device based on spatiotemporal representation alignment. It extracts target image features from the target image data by using a spatial transformer trained based on global coding features and extracts target optical flow features from each frame of the target image data by using a temporal transformer trained based on standard optical flow features, so as to output a target video including subject information in the target image data and motion features in the target video.

[0223] In some embodiments, the extraction module 902 is specifically used for:

[0224] Obtain the first original weight and the second original weight. The first original weight is the original weight corresponding to the target space low-rank matrix of the spatial transformer in the trained diffusion video model, and the second original weight is the original weight corresponding to the target space low-rank matrix of the time transformer in the trained diffusion video model.

[0225] If the current time step is less than or equal to the preset time step, adjust the first running weight of the target space low-rank matrix to be lower than the first original weight.

[0226] The target image features of the target image data are extracted by a spatial transformer after adjusting the first running weight;

[0227] If the current time step is greater than the preset time step, the first running weight of the target space low-rank matrix is ​​gradually increased to the first original weight.

[0228] The target image features of the target image data are extracted by a spatial transformer after adjusting the first running weight;

[0229] At any given time step, the target optical flow features of each frame of target image data in the target video data are extracted by running a time changer with a weight of the second running weight.

[0230] This application also provides an electronic device for executing the above-described method for training a personalized video generation model based on spatiotemporal representation alignment. Please refer to... Figure 10 It illustrates a schematic diagram of an electronic device provided by some embodiments of this application. For example... Figure 10 As shown, the electronic device 7 includes: a processor 700, a memory 701, a bus 702, and a communication interface 703. The processor 700, the communication interface 703, and the memory 701 are connected via the bus 702. The memory 701 stores a computer program that can run on the processor 700. When the processor 700 runs the computer program, it executes the training method or inference method of the personalized video generation model based on spatiotemporal representation alignment provided in any of the foregoing embodiments of this application.

[0231] The memory 701 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between the device network element and at least one other network element is achieved through at least one communication interface 703 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc.

[0232] Bus 702 can be an ISA bus, PCI bus, or EISA bus, etc. Buses can be divided into address buses, data buses, control buses, etc. Memory 701 is used to store programs. After receiving execution instructions, processor 700 executes the program. The personalized video generation model training method based on spatiotemporal representation alignment disclosed in any of the foregoing embodiments of this application can be applied to processor 700, or implemented by processor 700.

[0233] The processor 700 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 700 or by instructions in software form. The processor 700 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software units in the decoding processor. The software units may reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 701. Processor 700 reads the information in memory 701 and, in conjunction with its hardware, completes the steps of the above method.

[0234] The electronic device provided in this application embodiment and the personalized video generation model training method or the personalized video generation model inference method based on spatiotemporal representation alignment provided in this application embodiment are based on the same inventive concept and have the same beneficial effects as the methods they adopt, operate or implement.

[0235] This application also provides a computer-readable storage medium corresponding to the personalized video generation model training method based on spatiotemporal representation alignment provided in the foregoing embodiments. Please refer to [reference needed]. Figure 11 The computer-readable storage medium shown is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it executes the training method or inference method of the personalized video generation model based on spatiotemporal representation alignment provided in any of the foregoing embodiments.

[0236] It should be noted that examples of computer-readable storage media may also include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated here.

[0237] The computer-readable storage medium provided in the above embodiments of this application and the personalized video generation model training method based on spatiotemporal representation alignment provided in the embodiments of this application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the applications stored therein.

[0238] It should be noted that:

[0239] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this application may be practiced without these specific details. In some instances, well-known structures and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0240] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of this application and form different embodiments. For example, in the following claims, any of the claimed embodiments can be used in any combination.

[0241] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for training a personalized video generation model based on spatio-temporal representation alignment, characterized in that, The method comprises: inputting first image data into a spatial transformer of a diffusion video model to obtain intermediate features corresponding to the first image data, the first image data being any image data in an input data set, the intermediate features being a feature map of the largest scale in different scale feature maps output by the spatial transformer; inputting the first image data into a trained self-supervised encoder to obtain global encoding features corresponding to the first image data, the global encoding features including context information of the first image data, the intermediate features and the global encoding features being used to extract subject information in the first image data; training the spatial transformer of the diffusion video model based on the intermediate features and the global encoding features; inputting first video data into the diffusion video model to obtain reconstructed video data corresponding to the first video data output by the diffusion video model, the first video data being any video data in the input data set; inputting the first video data and the reconstructed video data into a pre-trained optical flow encoder respectively to obtain intermediate optical flow features and standard optical flow features corresponding to each frame of video data in the first video data, the intermediate optical flow features and the standard optical flow features being used to capture motion information of the first video data; training a temporal transformer of the diffusion video model based on the intermediate optical flow features and the standard optical flow features corresponding to each frame of video data; combining the trained spatial transformer and the trained temporal transformer to obtain a trained diffusion video model.

2. The method of claim 1, wherein, The training of the spatial transformer based on the intermediate features and the global encoding features comprises: calculating a relationship graph of the intermediate features and the global encoding features; fusing the relationship graph and the global encoding features to obtain fused features; generating a first loss function value based on the fused features and the global encoding features; training the spatial transformer based on the first loss function value.

3. The method of claim 2, wherein, The method further comprises: obtaining a reconstructed image based on the first image data by the diffusion video model, the reconstructed image including a predicted mask; inputting the first image data into a trained mask generation model to obtain a target mask; generating a second loss function value based on the predicted mask and the target mask; The training of the spatial transformer based on the first loss function value comprises: adjusting a spatial low-rank matrix of the spatial transformer based on the first loss function value and the second loss function value, continuing training until a first training completion condition is met, and obtaining a trained spatial transformer.

4. The method of claim 1, wherein, The training of the temporal transformer of the diffusion video model based on the intermediate optical flow features and the standard optical flow features corresponding to each frame of video data comprises: calculating a third loss function value for the first intermediate optical flow features and the first standard optical flow features of any frame of video data; training the temporal transformer of the diffusion video model based on the third loss function value.

5. The method of claim 4, wherein, The method further comprises: obtaining first anchor video frame data of first video data and second anchor video frame data corresponding to reconstructed video data of the first video data, the first anchor video frame data being any video frame data in the first video data, and the second anchor video frame data being video frame data in the reconstructed video data corresponding to the frame number of the first anchor video frame data; inputting the first video data and the reconstructed video data into an encoder respectively to obtain target latent space representation of each frame of video data in the first video data, first anchor feature of the first anchor video frame data, predicted latent space representation of each frame of video data in the reconstructed video data, and second anchor feature corresponding to the second anchor video frame data; determining target motion feature of each frame of video data in the first video data based on the target latent space representation of each frame of video data in the first video data and the first anchor feature; determining predicted motion feature of each frame of video data in the reconstructed video data based on the predicted latent space representation of each frame of video data in the reconstructed video data and the second anchor feature; calculating a fourth loss function value based on the predicted motion feature and the target motion feature; the training of the time transformer of the diffusion video model based on the third loss function value comprises: adjusting the time low-rank matrix of the time transformer based on the third loss function value and the fourth loss function value, continuing the training until a second training completion condition is met, and obtaining the trained time transformer.

6. The method according to any one of claims 1 to 5, characterized in that, The method further comprises: extracting target image data of a motion category from an image data set; extracting target video data with a subject quantity of one from a video data set; preprocessing the target image data and the target video data to obtain the input data set.

7. A personalized video generation model inference method based on spatio-temporal representation alignment, characterized in that, comprises: inputting a target data set into a trained diffusion video model, the trained diffusion video model being obtained based on the personalized video generation model training method based on spatio-temporal representation alignment of any one of claims 1-6, the target data set comprising target image data and target video data; extracting target image features of the target image data by a spatial transformer in the trained diffusion video model, and extracting target optical flow features of each frame of target image data in the target video data by a time transformer in the trained diffusion video model, the target image features being used to extract subject information in the target image data, and the target optical flow features being used to capture motion information of the target image data; outputting a target video based on the target image features and the target optical flow features of each frame of target image data by the trained diffusion video model.

8. The method of claim 7, wherein, the extraction of the target image features of the target image data by the spatial transformer in the trained diffusion video model, and the extraction of the target optical flow features of each frame of target image data in the target video data by the time transformer in the trained diffusion video model, comprises: obtaining a first original weight and a second original weight, the first original weight being an original weight corresponding to a target spatial low-rank matrix of a spatial transformer in the trained diffusion video model, and the second original weight being an original weight corresponding to a target spatial low-rank matrix of a temporal transformer in the trained diffusion video model; in a case where the current time step is less than or equal to a preset time step, adjusting the first running weight of the target spatial low-rank matrix to be lower than the first original weight; extracting a target image feature of the target image data by the spatial transformer with the adjusted first running weight; in a case where the current time step is greater than the preset time step, adjusting the first running weight of the target spatial low-rank matrix to gradually increase to the first original weight; extracting a target image feature of the target image data by the spatial transformer with the adjusted first running weight; in any time step, extracting a target optical flow feature of each frame of target image data in the target video data by the temporal transformer with a second running weight.

9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor runs the computer program to implement the method of any one of claims 1-8.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the method of any one of claims 1-8.

Citation Information

Patent Citations

  • Single-image four-dimensional dynamic human body video generation method based on diffusion converter

    CN120318383A

  • Training method of variational auto-encoder, video generation method and corresponding device

    CN120471107A