Video generation method and apparatus, computing device cluster, and storage medium
Patent Information
- Application Number
- CN202510392711.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2026-09-29
AI Technical Summary
[0030]此外,本申请实施例采用分布式推理技术,可通过将AI模型部署在不同设备上,来实现对多个视频片段的分布式的并行推理。例如本申请实施例可将目标视频切分后的多个视频片段分配到不同GPU上进行并行的风格化推理,每个GPU可负责处理一个或多个视频片段的风格化生成任务。本申请实施例的这种分布式推理方式大幅提升了视频风格化的速度,显著提升了视频风格化的效率,并使AI模型在目标视频的画面场景比较复杂的情况下的鲁棒性更强,避免了单一设备风格化视频时可能出现的内存瓶颈和计算延迟问题。
Smart Images

Figure CN122845882A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a video generation method, apparatus, computer equipment cluster, and storage medium. Background Technology
[0002] Currently, video stylization technology has received widespread attention in recent years, especially in fields that have high-quality requirements for video stylization technology, such as film and television production, animation creation, virtual reality (VR), augmented reality (AR), and advertising production.
[0003] Video stylization transforms a video from one style to another, such as a surreal, anime, or Van Gogh style. The video to be processed can be of any style or the original, unstylized video.
[0004] With the increasing demand for video content creation, especially the growing demand for personalized videos, how to achieve high-quality video stylization is an urgent problem to be solved. Summary of the Invention
[0005] This application provides a video generation method, apparatus, computer equipment cluster, and storage medium, which can stylize target videos with complex scenes to ensure that the stylized videos have natural and smooth transition effects.
[0006] In a first aspect, embodiments of this application provide a video generation method. The method includes: obtaining an inference request, the inference request indicating a target video to be processed and a target style to which the target video is to be processed; obtaining a first video segment and a second video segment from the target video; obtaining global style features of global frames obtained from preprocessing the target video, wherein the global frames are the top k video frames of the target video arranged in descending order based on at least one of pixel variation degree and scene variation degree, and the global style features indicate the style features of the global frames, where k is a positive integer; then, inputting the first video segment, the global style features, and the information indicating the target style into an artificial intelligence (AI) model for stylizing the video segments to output a first video segment with the target style; and inputting the second video segment, the global style features, and the information indicating the target style into the AI model to output a second video segment with the target style; finally, obtaining a target video with the target style based on the first video segment and the second video segment with the target style.
[0007] In one possible implementation, the method can take the top k video frames (i.e. the k video frames with the greatest degree of pixel change) from multiple video frames of the target video, arranged in descending order according to the degree of pixel change, as the k global frames.
[0008] In one possible implementation, the method can take the top k video frames (i.e. the k video frames with the greatest scene change) from the multiple video frames of the target video, arranged in descending order according to the degree of scene change, as the k global frames.
[0009] In one possible implementation, the method can take the top k video frames (i.e. the k video frames with the greatest overall change in pixel and scene changes) from multiple video frames of the target video, sorted in descending order according to the degree of change in pixel and scene changes, as the k global frames.
[0010] In some embodiments, multiple video frames in the target video can be sorted according to the degree of change (at least one of pixel change degree and scene change degree) in descending order of degree of change, so that the top k video frames are used as global frames.
[0011] In one possible implementation, the degree of pixel change of a video frame vi can indicate the magnitude of the difference in pixel values between the video frame vi and an adjacent video frame in the target video (e.g., video frame vi+1, i.e., the next video frame).
[0012] In one possible implementation, the degree of scene change in a video frame vi can indicate the magnitude of the difference in color distribution between that video frame vi and its adjacent video frames in the target video (e.g., video frame vi+1, i.e., the next video frame). Therefore, the k video frames in the target video with the greatest degree of scene change indicate the k video frames with the greatest differences in color distribution. A large difference in color distribution between two adjacent frames indicates that there is a high probability that a scene switch occurred between those two frames.
[0013] In one possible implementation, the global style feature of the global frame can be a key-value (KV) pair of the global frame.
[0014] In this application, when stylizing a target video using an AI model, multiple video segments within the target video can be stylized separately to obtain multiple video segments with the target style, thereby improving the stylization efficiency of the target video. Furthermore, this method can obtain the global style features of global frames obtained from the preprocessing of the target video. These global frames are the top k video frames in descending order based on at least one of pixel variation and scene variation among the multiple video frames of the target video, and the global style features indicate the style characteristics of these global frames within the target video. Thus, the global frames in the target video represent the k video frames with the most significant variations in at least one of pixel variation and scene variation. Therefore, the style features of the k video frames with the most drastic inter-frame fluctuations (i.e., the global style features) can reflect the style features of some video frames with the most significant scene variations and / or the style features of some video frames with the most significant pixel variations, enabling the style features of the global frames to reflect the overall style characteristics of the target video.
[0015] When stylizing a single video segment using an AI model, the input includes not only the video segment to be stylized (which can be the video segment itself or a feature sequence derived from it, without restriction), but also the global style features of global frames in the target video. This allows the AI model to reference both temporal and global information (e.g., global style features of global frames) during inference, ensuring that each stylized video segment not only possesses the target style but also incorporates the style of the global frames. This guarantees that the stylized video segments maintain a consistent style with the global frames, improving global consistency and the coherence of video actions. Furthermore, the input global style features can be style features of video frames with significant scene changes in the target video. This enhances the adaptability of the video generation method to complex scenes, enabling better stylization of target videos with complex scenes. The resulting video, after stylization, exhibits natural and smooth transitions, consistent style across different frames, and smooth changes in object movements. During the inference process, the global style features of each of the k global frames can be used for the inference of the AI model (e.g., the calculation of temporal attention). This can ensure the global consistency of the stylized target video generated by this method, and can filter redundant information (such as consecutive similar frames), reduce noise interference, and improve the adaptability of this method to complex scenes (such as fast motion and sudden changes in lighting).
[0016] In one possible implementation, during the process of obtaining the global style features of the global frames obtained from the preprocessing of the target video, the global style features are the style features of the global frames obtained from the preprocessing of the target video. Therefore, the preprocessing of the target video can be achieved as follows: The first k video frames from multiple video frames of the target video are obtained, sorted in descending order based on at least one of the degree of pixel change and the degree of scene change, to obtain k global frames, where k is a positive integer; the global style features of each of the k global frames are then obtained.
[0017] In this application, global frames can be adaptively selected based on the inter-frame fluctuation distribution, thereby effectively ensuring global consistency during the stylization process. Specifically, embodiments of this application can calculate the fluctuation information between target video frames and select video frames with strong representativeness (drastic fluctuations) as global frames. For example, the selected global frames can be the top k video frames arranged in descending order based on at least one of the pixel change degree and scene change degree among multiple video frames of the target video. The global frames obtained in this way are usually video frames with drastic motion (corresponding to large pixel change degree) or video frames with scene switching (corresponding to large scene change degree) in the target video. In the subsequent inference process, the global style features of each of the k global frames can be used for the inference of the AI model (e.g., the calculation of temporal attention), which can ensure the global consistency of the stylized target video generated by this method, and can filter redundant information (such as consecutive similar frames), reduce noise interference, and improve the adaptability of this method to complex scenes (such as fast motion, sudden changes in lighting).
[0018] In one possible implementation, obtaining the global style features of global frames obtained from preprocessing the target video includes: inputting the above-mentioned information indicating the target style, the above-mentioned k global frames, and at least one frame sequence obtained based on a third video segment into an AI model to obtain the global style features of each of the k global frames. The global style features of each global frame indicate the style features of the corresponding global frame in the target video and indicate the style features of the video frames in the third video segment after being stylized with the target style. The third video segment is either a first video segment or a second video segment.
[0019] In this application, when obtaining the global style features of each of the k global frames, not only can these k global frames be input into the AI model used for stylizing video segments, but information indicating the target style and at least one frame sequence obtained based on a third video segment can also be input into the AI model. Thus, the global style features of each global frame can indicate the style features of the corresponding global frame in the target video to be processed, enabling the global style features to indicate the style features of various scenes in the target video and the style features of video frames with significant pixel variations. Furthermore, the global style features of each global frame can also indicate the style features of video frames in the third video segment after stylization with the target style. The third video segment is any video segment in the target video. During inference, the AI model can utilize the global style features of each of the k global frames for inference (e.g., calculation of temporal attention). In this way, the global style feature of each global frame indicates the style feature (original style feature) of the corresponding global frame in the target video. This ensures that the content structure (such as object outlines and motion trajectories) in the target video is not destroyed by the stylization process, guaranteeing the global consistency of each video frame in the stylized target video. In addition, the global style feature of each global frame indicates the style feature of the video frame in the third video segment after being stylized with the target style. This can provide the required target style information, thereby ensuring the style consistency of the stylized target video on the time axis, preventing style drift between frames, and making the stylized target video maintain a consistent target style visual effect across different frames.
[0020] In one possible implementation, obtaining the global style features of global frames obtained from preprocessing the target video includes: obtaining key-value (KV) pairs for each of the k global frames obtained from preprocessing the target video, wherein the KV pairs indicate the style features of the corresponding global frames.
[0021] In this way, the AI model can obtain key-value pairs (KV pairs) for each of the k global frames, with each KV pair indicating the style features of the corresponding global frame. These KV pairs can then be used to indicate the style features of the global frame. During the inference phase, the AI model's input can include these KV pairs from the k global frames. The AI model can then use these KV pairs from each of the k global frames for inference (e.g., calculating temporal attention). Since each global frame's KV pair indicates the style features (original style features) of the corresponding global frame in the target video, this ensures that the content structure (such as object outlines and motion trajectories) in the target video is not destroyed by the stylization process, guaranteeing global consistency across all frames in the stylized target video.
[0022] In one possible implementation, the key-value pairs of the global frame indicate the style features of the corresponding global frame in the target video and indicate the style features of the video frames in the third video segment after being stylized in the target style, wherein the third video segment is either the first video segment or the second video segment.
[0023] In this way, the AI model can obtain key-value pairs for each of the k global frames. Each key-value pair indicates the style features of the corresponding global frame in the target video and the style features of the video frames in the third video segment after being stylized in the target style. Therefore, the key-value pairs can be used to indicate the style features of the global frame and the style features of the video frames in the third video segment after being stylized in the target style. During the inference phase, the input to the AI model can include key-value pairs (KV pairs) of the k global frames. The AI model can then use the KV pairs of each of these k global frames for inference (e.g., calculation of temporal attention). Since each global frame's KV pair indicates the style features (original style features) of the corresponding global frame in the target video, this ensures that the content structure (such as object outlines and motion trajectories) in the target video is not destroyed by the stylization process, guaranteeing global consistency among all video frames in the stylized target video. Furthermore, the global style features of each global frame indicate the style features of the video frames in the third video segment after stylization with the target style, providing the necessary target style information. This ensures style consistency across the timeline of the stylized target video, preventing inter-frame style drift and maintaining a consistent visual effect of the target style across different frames.
[0024] In one possible implementation, the AI model includes a temporal action module that inputs the first video segment, the global style features, and the information indicating the target style into the AI model for stylizing the video segment, to output a first video segment with the target style. This includes: the temporal action module obtaining global attention based on the global style features and the style features of the first video segment stylized according to the target style. The style features of the first video segment stylized according to the target style are intermediate results obtained by the AI model stylizing each video frame in the first video segment according to the target style; the temporal action module obtaining local attention based on the intermediate results corresponding to the first video segment; the temporal action module obtaining temporal attention based on the global attention and local attention; and the temporal action module outputting a first video segment with the target style based on the temporal attention.
[0025] In this application, the temporal action module can calculate the global attention (calculated by combining the global style features of the preprocessed global frame) and the local attention for each frame in the first video segment, thereby obtaining the temporal attention for each frame in the first video segment. This allows the temporal action module to not only focus on short-term dependencies through local attention but also on long-term dependencies through global attention, thus ensuring the consistency between local details and global style in the generated stylized target video. This improves the naturalness and realism of the stylized target video, thereby enhancing the adaptability of this method to complex scenes in long-term video stylization tasks and avoiding the matching problem between the training and inference phases.
[0026] In one possible implementation, the temporal action module obtains temporal attention based on global attention and local attention, including: the temporal action module determines a first weight corresponding to global attention and a second weight corresponding to local attention based on the probability distribution of global attention and local attention; and weights the global attention and local attention based on the first weight and the second weight to obtain temporal attention.
[0027] In this application, the temporal action module can determine the first weight for global attention and the second weight for local attention based on the probability distribution of global and local attention corresponding to video frames in the first video segment. Then, it calculates the weighted result of global and local attention for each frame to obtain the temporal attention. This allows for adaptive adjustment of the weights of the two attention types according to the specific scene of the video segment. When the target video is a long-form video, which may involve complex and varied scene sequences, this method can meet the different needs of different video segments in terms of coordination between local details and global style.
[0028] In one possible implementation, there are video frames with overlapping display times between the first video segment and the second video segment. An AI model for outputting the first video segment with a target style runs on a first device, and an AI model for outputting the second video segment with a target style runs on a second device. The method further includes: the first device sending first hidden layer features of the video frames with overlapping display times calculated by the running AI model to the second device; the first device receiving second hidden layer features of the video frames with overlapping display times calculated by the AI model running on the second device sent by the second device; and the first device performing feature fusion on the first hidden layer features and the second hidden layer features to output the first video segment with the target style.
[0029] In this application, AI models can be distributed across multiple devices. Each device's AI model can then be used to stylize one or more video segments within a target video, enabling parallel stylization of multiple video segments and improving efficiency. Furthermore, when video segments processed by different devices have overlapping display times, the AI models on both devices can exchange the hidden features of those frames during inference within adjacent video segments. These exchanged hidden features are then fused, for example, as input to the next hidden layer in the AI model. This achieves feature fusion for the same video frame across different devices, ensuring stylistic consistency among the stylized video segments obtained from different devices. It also facilitates subsequent stitching of the stylized video segments, resulting in a smoother, more natural final video with the target style.
[0030] Furthermore, this application employs distributed inference technology, which enables distributed parallel inference of multiple video segments by deploying the AI model on different devices. For example, this application can distribute multiple video segments after the target video is segmented onto different GPUs for parallel stylization inference, with each GPU responsible for processing the stylization generation task of one or more video segments. This distributed inference method significantly improves the speed and efficiency of video stylization, and makes the AI model more robust when the target video has complex scenes, avoiding memory bottlenecks and computational latency issues that may occur when stylizing videos on a single device.
[0031] Secondly, embodiments of this application provide a video generation apparatus. The apparatus includes: an acquisition module, configured to acquire an inference request, the inference request indicating a target video to be processed and a target style to which the target video is to be processed; the acquisition module is further configured to acquire a first video segment and a second video segment from the target video; the acquisition module is further configured to acquire global style features of global frames obtained by preprocessing the target video, the global frames being the top k video frames arranged in descending order based on at least one of pixel change degree and scene change degree among multiple video frames of the target video, the global style features indicating the style features of the global frames, where k is a positive integer; an inference module, configured to input the first video segment, the global style features, and the information indicating the target style into an artificial intelligence (AI) model for stylizing the video segments to output a first video segment with the target style, and to input the second video segment, the global style features, and the information indicating the target style into the AI model to output a second video segment with the target style; and a generation module, configured to obtain a target video with the target style based on the first video segment with the target style and the second video segment with the target style.
[0032] In one possible implementation, the acquisition module is specifically used to: acquire the top k video frames from multiple video frames of the target video, sorted in descending order according to at least one of the degree of pixel change and the degree of scene change, to obtain k global frames, where k is a positive integer; and acquire the global style features of each of the k global frames.
[0033] In one possible implementation, the acquisition module is specifically used to: input the information indicating the target style, the k global frames, and at least one frame sequence obtained based on the third video segment into the AI model to obtain the global style features of each of the k global frames. The global style features of each global frame indicate the style features of the corresponding global frame in the target video and indicate the style features of the video frames in the third video segment after being stylized with the target style. The third video segment is either the first video segment or the second video segment.
[0034] In one possible implementation, the acquisition module is specifically used to: acquire key-value (KV) pairs for each of the k global frames obtained from the preprocessing of the target video, wherein the KV pairs indicate the style features of the corresponding global frames.
[0035] In one possible implementation, the inference module includes an AI model, which includes a temporal action module. This temporal action module is configured to: obtain global attention based on global style features and style features of a first video segment stylized to the target style; the style features of the first video segment stylized to the target style are intermediate results obtained by the AI model stylizing each video frame in the first video segment according to the target style; obtain local attention based on the intermediate results corresponding to the first video segment; obtain temporal attention based on global attention and local attention; and output a first video segment with the target style based on the temporal attention.
[0036] In one possible implementation, the aforementioned temporal action module is specifically used to: determine a first weight corresponding to global attention and a second weight corresponding to local attention based on the probability distributions of global attention and local attention; and weight the global attention and local attention based on the first weight and the second weight to obtain temporal attention.
[0037] The effects of the video generation apparatus in the above embodiments are similar to those of the video generation methods in the above embodiments, and will not be repeated here.
[0038] Thirdly, embodiments of this application provide a video generation system. The system includes a first device deploying a video generation apparatus as described in the second aspect or any possible implementation of the second aspect, and a second device deploying a video generation apparatus as described in the second aspect or any possible implementation of the second aspect. The first device is configured to input a first video segment, global style features, and information indicating a target style into an AI model in the video generation apparatus deployed on the first device, to output a first video segment with the target style; the second device is configured to input a second video segment, global style features, and information indicating a target style into an AI model in the video generation apparatus deployed on the second device, to output a second video segment with the target style; the first device or the second device is configured to obtain a target video with the target style based on the first video segment with the target style and the second video segment with the target style.
[0039] In one possible implementation, there are video frames with overlapping display times between the first video segment and the second video segment; the first device is configured to send the first hidden layer features of the video frames with overlapping display times, calculated by the AI model in the video generation device deployed on the first device, to the second device; the first device receives the second hidden layer features of the video frames with overlapping display times, calculated by the AI model running on the second device, sent by the second device; the first device performs feature fusion on the first hidden layer features and the second hidden layer features to output the first video segment with the target style.
[0040] The effects of the video generation systems described in the above embodiments are similar to those of the video generation methods described in the above embodiments, and will not be repeated here.
[0041] Fourthly, embodiments of this application provide a computing device cluster, including at least one computing device, each computing device including a processor and a memory, wherein the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, such that the computing device cluster performs the video generation method of the first aspect or any possible implementation thereof.
[0042] The effect of the computing device cluster in this embodiment is similar to that of the video generation methods in the above embodiments, and will not be described again here.
[0043] Fifthly, embodiments of this application provide a computer program product containing instructions that, when executed by a cluster of computing devices, cause the cluster of computing devices to perform the video generation method of the first aspect or any possible implementation thereof.
[0044] The effect of the computer program product in this embodiment is similar to that of the video generation methods in the above embodiments, and will not be described again here.
[0045] In a sixth aspect, embodiments of this application provide a computer-readable storage medium including computer program instructions, which, when executed by a cluster of computing devices, enable the cluster of computing devices to perform a video generation method in the first aspect or any possible implementation thereof.
[0046] The effect of the computer-readable storage medium in this embodiment is similar to that of the video generation methods in the above embodiments, and will not be described again here. Attached Figure Description
[0047] Figure 1 This application provides a schematic diagram of the structure of a cloud system according to an embodiment of the present application.
[0048] Figure 2 A flowchart illustrating a video generation method provided in an embodiment of this application;
[0049] Figure 3 A schematic diagram illustrating a process for acquiring video clips, provided as an embodiment of this application;
[0050] Figure 4a This application provides a schematic diagram of a video preprocessing workflow.
[0051] Figure 4b A schematic diagram of a process for obtaining key-value pairs of a global frame, provided in an embodiment of this application;
[0052] Figure 5a This application provides a schematic diagram of a process for stylizing video clips according to an embodiment of the present application.
[0053] Figure 5b A schematic diagram illustrating a distributed parallel processing of video segments provided in an embodiment of this application;
[0054] Figure 6 A schematic diagram illustrating the processing procedure of a temporal attention module provided in an embodiment of this application;
[0055] Figure 7 This is a schematic diagram of the structure of a video generation device provided in an embodiment of this application;
[0056] Figure 8 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application;
[0057] Figure 9 This is a schematic diagram of the structure of a computing device cluster provided in an embodiment of this application;
[0058] Figure 10 This is a schematic diagram of another computing device cluster provided in an embodiment of this application. Detailed Implementation
[0059] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0060] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0061] The terms "first" and "second," etc., used in the specification and claims of this application are used to distinguish different objects, not to describe a specific order of objects. For example, "first target object" and "second target object," etc., are used to distinguish different target objects, not to describe a specific order of target objects.
[0062] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0063] In the description of the embodiments in this application, unless otherwise stated, "multiple" means two or more. For example, multiple processing units means two or more processing units; multiple systems means two or more systems.
[0064] Before describing the technical solutions of the embodiments of this application, a brief introduction to the background technology and technical terms involved in the embodiments of this application will be given first:
[0065] 1) Public cloud is a cloud platform provided by a third-party public cloud provider to a wide range of individuals or businesses. In a public cloud, the hardware, software, and other infrastructure are owned and managed by the third-party public cloud provider.
[0066] 2) A private cloud is a dedicated cloud platform provided for a single enterprise or organization. Private clouds can be operated internally by the respective enterprise or organization. Private clouds are primarily geared towards enterprise users and are also known as enterprise clouds.
[0067] 3) Hybrid cloud refers to a cloud platform formed by different cloud platforms. A hybrid cloud includes at least two cloud platforms, also known as a multi-cloud platform or multi-cloud. Optionally, a hybrid cloud integrates public and private clouds. For security reasons, some enterprise users prefer to store data in a private cloud, but at the same time want to obtain the computing resources of a public cloud. In this case, hybrid clouds, which include both public and private clouds, are increasingly being adopted. Hybrid clouds combine and match public and private clouds to achieve optimal performance.
[0068] 4) Diffusion model: This model generates new samples by simulating the "diffusion" process of data samples. Its core idea is to gradually remove noise to generate high-quality samples by simulating the diffusion and reverse diffusion processes of data over time. The basic principles of the diffusion model can include forward diffusion and reverse diffusion processes.
[0069] The forward diffusion process is a crucial part of the diffusion model training process. Its goal is to gradually add noise to a real sample (such as the original image or video) until it eventually becomes pure noise. Through a series of noise-adding steps, the original image is gradually transformed into random noise.
[0070] The backdiffusion process starts with pure noise and gradually removes it to recover a new data sample. This process is the core of training. The diffusion model receives pure noise obtained from the forward diffusion process and learns how to generate a clear sample from a noisy image. During training, the goal is to learn the inverse of the forward diffusion process, that is, to recover the true distribution of the samples by optimizing the loss function. Thus, during training, the diffusion model learns how to recover the original data from noise.
[0071] In this way, the diffusion model can generate high-quality samples during the inference phase by back-diffusion of noise (a denoising process), that is, gradually restoring the noise into a clear image or data.
[0072] 5) ControlNet, which can serve as an enhancement module for the diffusion model, is a network used to control the generation process, aiming to improve the ability of the diffusion model to generate images under specific conditions. ControlNet accepts additional control information as input, which can be specific features of the image, class labels, bounding boxes, pose information, etc. ControlNet can enhance the forward and backward diffusion processes of the diffusion model, ensuring that the generated images conform to the control information while maintaining the realism and quality of the images.
[0073] 6) Stable Diffusion is a type of diffusion model based on deep learning for image generation. Its core idea is to view the image generation process as a diffusion process. In this process, the Stable Diffusion model progressively denoises the random noise (i.e., the aforementioned back-diffusion) based on the input random noise and a series of control conditions (also called denoising conditions) to generate the final image. These control conditions can be provided by ControlNet.
[0074] For example, a Stable Diffusion model may include a mask-based neural network (e.g., a U-Net) or an attention-based neural network (e.g., a Transformer).
[0075] 7) Animate Diffusion (AnimateDiff) Model: This is a deep learning-based image generation model, a short-time video generation model specifically designed for generating animated frames or transforming and compositing dynamic images. It typically combines image and video generation techniques, capable of transforming static images into dynamic animations or creating smooth animation effects through frame-by-frame generation. When generating animations, AnimateDiff ensures that the details and style of the image remain consistent across different frames, while providing natural transitions, resulting in a smooth and visually appealing final animation.
[0076] 8) Long-duration videos and short-duration videos: Long-duration videos are videos of a longer duration, while short-duration videos are videos of a shorter duration. The terms "long" and "short" are relative and not limited to a specific range. In some scenarios, long-duration videos are those with a duration (i.e., length) exceeding one minute; for example, the frame rate of long-duration videos is typically 20 frames per second. Short-duration videos are those with a duration in the tens of seconds, such as videos less than 10 seconds long.
[0077] 9) Video stylization refers to using a model to regenerate the video to be processed, so that the video to be processed becomes a video of another style, such as surreal style, anime style, Van Gogh style, etc.
[0078] 10) Short-time video generation models are deep learning models used to generate short-time videos. Unlike traditional static image generation models (such as Generative Adversarial Networks (GANs) and Variational Autoencoders (VAEs), short-time video generation models need to consider the temporal dimension, i.e., the coherence and dynamic changes between image frames. It not only generates single-frame images but also continuous sequences of multiple frames, ensuring consistency between these frames to guarantee the smoothness and watchability of the video. Short-time video generation models can be used as plug-in frameworks in conjunction with existing diffusion models. For example, a short-time video generation model could be an AnimateDiffusion (AnimateDiff) model.
[0079] 11) Low-rank adaptation (LoRA) is an efficient parameter fine-tuning technique for large models (such as diffusion models and language models). It reduces the number of parameters that need to be updated during the fine-tuning process by decomposing the model's weight matrix into a low-rank form, thereby achieving efficient model fine-tuning.
[0080] Long-duration video stylization technology has received widespread attention in recent years, especially in fields such as film and television production, animation creation, virtual reality (VR), augmented reality (AR), and advertising production, where there is a high demand for high-quality video stylization. With the continuous growth in demand for video content creation, particularly the increasing need for personalized custom videos, how to achieve high-quality long-duration video stylization is a pressing issue that needs to be addressed.
[0081] In related technologies, long-term video stylization implementation schemes can be mainly divided into two categories: schemes based on video base models and schemes based on image base models.
[0082] Video-based model-based approaches typically rely on the temporal information and structured features of video content to achieve long-term video stylization through pre-trained video generation models. However, these approaches are costly in terms of pre-training the video generation model. For example, pre-training requires a large number of high-quality videos, and collecting and labeling such a large amount of high-quality video data is also a significant undertaking. Furthermore, the computational complexity of these approaches is generally high, which in turn leads to low generation efficiency when generating stylized long-term videos. In addition, current video-based models are prone to temporal instability and missing local details when processing long-term videos, resulting in a lack of temporal coherence and the potential blurring or loss of details in specific regions, thus reducing the quality of the generated videos.
[0083] Image-based model-based approaches achieve long-term video stylization by individually stylizing each frame. However, these approaches suffer from poor global consistency during stylization due to the lack of temporal and global information about the video content during inference, making it difficult to guarantee the continuity of object movements in the generated video. Furthermore, when faced with videos featuring complex scenes, these approaches cannot effectively handle diverse background changes and scene transitions in long-term videos, resulting in a lack of natural and smooth transitions. This leads to inconsistent styles between frames and abrupt changes in object movements, making these approaches less adaptable to complex scenes.
[0084] Based on this, this application provides a video generation method. This method acquires multiple video segments (e.g., a first video segment and a second video segment) of a target video and processes each segment using an AI model for stylizing the video segments. Compared to directly processing a complete long video, this method produces a higher quality video. Furthermore, parallel processing of multiple video segments significantly improves video generation efficiency. When stylizing each video segment using the AI model, this method inputs not only the video segment but also global style features, ensuring that all segments in the stylized target video maintain a consistent style with the global frames, improving global consistency and the coherence of video actions. Moreover, the input global style features can be style features of frames in the target video with significant scene changes, thus this method exhibits good adaptability to complex scenes.
[0085] The method described in this application can be applied to cloud scenarios or scenarios outside the cloud, without limitation.
[0086] The following describes the cloud system according to an embodiment of this application. In this embodiment, the cloud system is a public cloud as an example. In other embodiments, the cloud system can also be a private cloud and / or a hybrid cloud. This application does not limit the scope of the cloud system.
[0087] The following is combined with Figure 1 The cloud system 100 of this application embodiment will be introduced.
[0088] like Figure 1 As shown, cloud system 100 may include public cloud 101 and one or more tenants (here, one tenant 102 is taken as an example).
[0089] The public cloud 101 may include a cloud management platform 103 and cloud infrastructure 104 with communication connections.
[0090] The cloud management platform 103, also known as the cloud platform or simply the cloud management platform, is a software system used by cloud providers to provide cloud technology (also known as cloud computing) services and can be used to manage cloud infrastructure 104.
[0091] Cloud infrastructure 104 refers to the hardware devices that provide cloud services. Cloud infrastructure 104 may include multiple data centers (DCs) located in different regions, with at least one data center in each region. Each data center may contain multiple physical servers, and each physical server can be used to support various cloud services. For example, physical servers can be bare-metal servers; there are no restrictions on this.
[0092] Cloud services may include computing services, storage services, virtual machine services, container services, network services, etc. Devices or functions accessible to tenant 102 upon logging into the cloud management platform 103 can all be considered cloud services provided by the cloud infrastructure 104.
[0093] In this embodiment, the cloud service may include a service for implementing the video generation method of this application, so that tenants can purchase the service to stylize the target video.
[0094] like Figure 1 As shown, the cloud management platform 103 provides an interface related to cloud services for tenant 102 (or a client) to remotely access cloud services. Tenant 102 can log in to the cloud management platform 103 through a pre-registered account and password on the cloud service access page, and after successful login, purchase and use the corresponding cloud services on the cloud service access page. Since the cloud management platform 103 is communicatively connected to the cloud infrastructure 104, the cloud management platform 103 can provide various cloud services supported by the cloud infrastructure 104 and purchased by tenant 102 to tenant 102 for use.
[0095] like Figure 1 The client shown refers to the terminal or browser on the terminal used by tenant 102 that can access cloud services. This terminal may include, but is not limited to, mobile phones, tablets, computers, personal computers (PCs), and devices in internet systems.
[0096] For example, such as Figure 1 The cloud management platform 103 shown can receive inference requests sent by tenant 102, which specify the target video to be processed and the target style to be applied to the target video. In response to this inference request, cloud management platform 103 can interact with cloud infrastructure 104 to use an AI model running on cloud infrastructure 104 to perform inference on the target video to obtain a target video with the target style. Then, cloud management platform 103 can output the target video with the target style to tenant 102 in response to the inference request. In this way, tenant 102 achieves stylization processing of the video through purchased cloud services.
[0097] The video generation method of this application is described below to achieve stylization of the target video to be processed.
[0098] The entity executing this video generation method can be the cloud management platform mentioned above, or cloud infrastructure, or a client and / or server in a scenario outside the cloud.
[0099] The client can be software, applications, browsers, in-vehicle systems, terminal devices, etc. When the client is implemented as a terminal device, it can include, but is not limited to, any of the following: mobile phone, personal computer (PC), virtual reality (VR) device, augmented reality (AR) device, tablet computer, laptop computer, etc. The client can also be a smart TV, mobile internet device (MID), wearable device (such as a smartwatch, smart glasses, or smart helmet), smart car, wireless terminal device in industrial control, wireless terminal device in self-driving, wireless terminal device in remote medical surgery, wireless terminal device in smart grid, wireless terminal device in transportation safety, wireless terminal device in smart city, wireless terminal device in smart home, etc. The following embodiments do not impose special limitations on the specific form of the client.
[0100] This server can be implemented through software or hardware.
[0101] In the first example, when the functionality of the server is implemented through software, the server may be, for example, an application running on a computing instance (e.g., an application that implements the methods of this application), which may be, for example, a virtual machine, a container, or a host.
[0102] In the second example, when the server functionality is implemented through hardware, the server can be implemented through at least one physical device including a processor. This physical device can be a server (e.g., a compute node (also called a compute server), a central server, or an edge server), a base station, a relay device, satellite equipment, etc., without limitation.
[0103] The processor can be a central processing unit (CPU) or a graphics processing unit (GPU), or it can be any type of processor or any combination thereof, such as an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), a system-on-chip (SoC), a software-defined infrastructure (SDI) chip, an AI chip, or a data processing unit (DPU).
[0104] Furthermore, the number of processors included in the server can be arbitrary, and the types of processors included can be one or more. The specific number and types of processors can be set according to the actual business needs of the application, and this application does not impose any restrictions on this.
[0105] In the third example, when the server-side functionality is implemented through hardware, the server can also be a computing cluster comprising multiple computing nodes. Furthermore, these multiple computing nodes can communicate through the at least one switching node. Exemplarily, a computing node can be a computing server including an accelerator card. This accelerator card can be, for example, a deep-learning processing unit (DPU), a GPU, a neural-network processing unit (NPU), or a tensor processing unit (TPU), or other types of accelerator cards. Alternatively, a computing node can be a computing server including a general-purpose processor (such as a CPU).
[0106] Figure 2 A flowchart illustrating a video generation method according to an embodiment of this application. Figure 2 As shown, the video generation method may include, but is not limited to, the following steps: S101, S111, S121, S102 and S103.
[0107] S101: Obtain inference request.
[0108] In some embodiments, the inference request indicates the target video to be processed and the target style to which the target video is to be processed.
[0109] For example, the target video can be a short video or a long video, and this application does not impose a specific limitation on the duration of the target video. The target video can be a video of any style, or an unstylized original video shot by the client, and this application does not impose any limitations on this either.
[0110] For example, the target style to be processed for the target video can be, but is not limited to, any of the following: surrealism, animation, Van Gogh, retro, futurism, black and white film, watercolor, pixel art, etc.
[0111] As an example, a user can upload a local video (e.g., Video 1) via a mobile application and select a target style (e.g., black and white film style) from a style list provided within the application. This style list represents the video stylization styles supported by the application. The user can then trigger an inference request within the application, allowing the mobile application to process the target video's style to the target style using the method described in this embodiment, outputting a target video with the target style. Optionally, the mobile application can interact with a server to achieve video stylization; this is not limited here.
[0112] After S101, the method may also include S111 and S121. The execution order of S111 may be before S121, after S121, or in parallel with S121. This application does not impose any specific restrictions on this.
[0113] S111: Obtain the first and second video segments from the target video.
[0114] The target video may include multiple video segments, which may include a first video segment and a second video segment.
[0115] The method in this embodiment can obtain the first video segment and the second video segment by dividing the target video into segments. Alternatively, the method in this embodiment can directly obtain the first video segment and the second video segment obtained by dividing the target video, without performing segmentation processing. No limitation is imposed here.
[0116] For example, the first video segment and the second video segment may be all or part of the video segments in the target video.
[0117] For example, the duration of the first video segment and the duration of the second video segment may be the same or different.
[0118] For example, the video segments obtained from the target video can be video segments with a duration of seconds.
[0119] For example, there may be at least one video frame with overlapping display times between the first video segment and the second video segment. For instance, the first video segment is a video segment of the target video with a display time of 0 to 10 seconds, and the second video segment is a video segment of the target video with a display time of 8 to 18 seconds. In this way, there may be multiple video frames with overlapping display times from 8 to 10 seconds between the first video segment and the second video segment.
[0120] The number of video segments obtained from the target video can be two or more. In this embodiment and subsequent embodiments, we take the processing of two video segments as an example to illustrate how to stylize the target video. The processing of other video segments in the target video is the same as the processing of the first video segment (or the second video segment) described in this embodiment, so it will not be described in detail.
[0121] S121: Obtain the global style features of the global frame obtained from the preprocessing of the target video.
[0122] In some embodiments, the global frame can be the top k video frames from a plurality of video frames of the target video, sorted in descending order based on at least one of the degree of pixel change and the degree of scene change, where k is a positive integer.
[0123] Thus, the number of global frames is k, and each global frame can be one or more video frames. This application does not impose a specific limit on the number of global frames (i.e., the size of k).
[0124] As an example 1:
[0125] For example, the k global frames can be the top k video frames in descending order of pixel change from multiple video frames in the target video, that is, the k video frames with the greatest pixel change in the target video.
[0126] As an example 2:
[0127] For example, the k global frames can be the top k video frames in descending order of the scene change from multiple video frames in the target video, that is, the k video frames in the target video with the greatest scene change.
[0128] As an example 3:
[0129] For example, the k global frames can be the top k video frames in the target video, sorted in descending order based on the overall degree of change of pixel changes and scene changes (e.g., the weighted result of the two degrees of change), that is, the k video frames in the target video with the greatest overall degree of change of pixel changes and scene changes.
[0130] The degree of pixel change in a video frame vi can indicate the magnitude of the difference in pixel values between that video frame vi and adjacent video frames in the target video (e.g., video frame vi+1, i.e., the next video frame).
[0131] The degree of scene change in a video frame vi indicates the magnitude of the difference in color distribution between that video frame vi and its adjacent video frames in the target video (e.g., video frame vi+1, the next video frame). Therefore, the k video frames with the greatest degree of scene change in the target video indicate the k video frames with the greatest differences in color distribution. A significant difference in color distribution between two adjacent frames indicates that there is a high probability that a scene switch occurred between those two frames.
[0132] The global style feature can indicate the style features of the aforementioned global frames in the target video. As an example, the style feature can be a key-value (KV) pair of the global frame.
[0133] Following S101, S111, and S121, the method may further include S102.
[0134] S102: Input the first video segment obtained in S111, the global style features of the global frame obtained in S121, and the information indicating the target style into the Artificial Intelligence (AI) model to output a first video segment with the target style; and input the second video segment obtained in S111, the global style features of the aforementioned global frame obtained in S121, and the information indicating the target style into the AI model to output a second video segment with the target style.
[0135] The information regarding the indicated target style can be obtained, for example, based on the inference request acquired in S101.
[0136] The aforementioned AI model is used to stylize video clips.
[0137] For example, the AI model for stylizing video clips described above may include an image generation module, a temporal action module, and optionally a control network module. The specific structure of this AI model will be discussed later. Figure 4b A detailed introduction will be provided in the following section, which will not be elaborated upon here.
[0138] In some embodiments, the first video segment and the second video segment use the same AI model. This ensures good consistency between the generated first and second video segments with the target style in terms of color, style, and video clarity, resulting in better coherence between the two video segments in the final target video and improving the overall consistency of the target video.
[0139] In some embodiments, the AI model that stylizes the first video segment and the AI model that stylizes the second video segment can run on the same graphics processing unit (GPU), so that different video segments in the target video can be stylized sequentially.
[0140] In some embodiments, the AI model for stylizing the first video segment and the AI model for stylizing the second video segment can run on different GPUs. This allows for distributed and parallel stylization of different video segments in the target video, improving the efficiency of generating the stylized video. Furthermore, compared to related technologies that pre-train video generation models, the AI model in this embodiment does not require training, thereby reducing the cost of video stylization and eliminating the need to collect and label large amounts of high-quality video data, thus reducing workload.
[0141] Furthermore, AI models on the same GPU can be used to infer one or more video segments in a target video to sequentially output one or more stylized video segments.
[0142] For example, Figure 1 The cloud infrastructure 104 shown may include multiple GPUs, each of which can run an AI model for stylizing video clips. The cloud infrastructure 104 can process the first and second video clips in parallel using these multiple GPUs based on inference requests. For example, the cloud infrastructure 104 can input the first video clip and global style features into the AI model based on an inference request, and simultaneously input the second video clip and global style features into the same AI model based on the same inference request. These two AI models run on different GPUs, thereby achieving parallel processing of the first and second video clips.
[0143] Following S102, the method may further include S103.
[0144] S103: Based on the first video segment with the target style and the second video segment with the target style, obtain the target video with the target style.
[0145] In some embodiments, the method can splice a first video segment with a target style and a second video segment with a target style in chronological order to obtain a target video with the target style.
[0146] Continuing with the example mentioned above, where the first video segment is a video segment whose display time is from 0 to 10 seconds and the second video segment is a video segment whose display time is from 8 to 18 seconds, we will illustrate two possible splicing schemes when splicing the first video segment with the target style and the second video segment with the target style in chronological order.
[0147] The main difference between the two splicing schemes lies in how they handle multiple video frames from the 8th to the 10th second that overlap in display time in the first video clip with the target style and the second video clip with the target style.
[0148] The video clips mentioned in the following two splicing schemes are all video clips with the target style obtained after stylization processing.
[0149] The first splicing scheme involves discarding the stylized video frames in the second video segment whose target video is displayed between 8 and 10 seconds. The first video segment (the video segment whose target video is displayed between 0 and 10 seconds) and the video segment in the second video segment whose target video is displayed between 11 and 18 seconds are spliced together in chronological order to obtain a target video with the target style.
[0150] The second splicing scheme involves discarding the stylized video frames in the first video segment whose target video is displayed between 8 and 10 seconds. The video segments in the first video segment whose target video is displayed between 0 and 7 seconds are spliced with the second video segment (the video segments whose target video is displayed between 8 and 18 seconds) in chronological order to obtain a target video with the target style.
[0151] Each stylized video frame exhibiting temporal overlap between the first and second video clips with the target style is identical (e.g., the style and content of the temporally overlapping video frames are consistent). This is achieved through latent feature fusion of the temporally overlapping video frames during parallel processing of the two video clips. For details, please refer to the following text. Figure 5b Introduction.
[0152] Therefore, during the splicing process, any video frame with overlapping time in a video segment can be discarded to obtain the spliced result of the two video segments.
[0153] The above description is based on the case where the first and second video segments cover the entire target video. In other embodiments, other video segments besides the first and second video segments can be obtained based on the target video. The principle of processing these other video segments is similar to the principle of processing the first and second video segments described above, and will not be repeated here.
[0154] In this embodiment, when stylizing a target video using an AI model, multiple video segments within the target video can be stylized separately to obtain multiple video segments with the target style, thereby improving the stylization efficiency of the target video. Furthermore, this method can obtain the global style features of global frames obtained from the preprocessing of the target video. These global frames are the top k video frames arranged in descending order based on at least one of pixel change and scene change, and the global style features indicate the style characteristics of these global frames within the target video. Thus, the global frames in the target video represent the k video frames with the most significant changes in at least one of pixel change and scene change. Therefore, the style features of the k video frames with the most drastic inter-frame fluctuations (i.e., the global style features) can reflect the style features of some video frames with the most significant scene changes and / or the style features of some video frames with the most significant pixel changes, enabling the style features of the global frames to reflect the overall style characteristics of the target video.
[0155] When stylizing a single video segment using an AI model, the input includes not only the video segment to be stylized (which can be the video segment itself or a feature sequence derived from it, without restriction), but also the global style features of global frames in the target video. This allows the AI model to reference both temporal and global information (e.g., global style features of global frames) during inference, ensuring that each stylized video segment not only possesses the target style but also incorporates the style of the global frames. This guarantees that the stylized video segments maintain a consistent style with the global frames, improving global consistency and the coherence of video actions. Furthermore, the input global style features can be style features of video frames with significant scene changes in the target video. This enhances the adaptability of the video generation method to complex scenes, enabling better stylization of target videos with complex scenes. The resulting video, after stylization, exhibits natural and smooth transitions, consistent style across different frames, and smooth changes in object movements.
[0156] In this embodiment, multiple video segments (e.g., a first video segment and a second video segment) of the target video can be obtained, and these multiple video segments are processed based on an AI model used for stylizing the video segments. Thus, even if the target video is a long video, the video processed by the AI model is actually a short video, resulting in a higher quality generated video. Furthermore, processing multiple video segments in parallel can significantly improve video generation efficiency.
[0157] Combined with Figure 2 In any embodiment, for the above S111, in one possible implementation, when the method acquires the first video segment and the second video segment in the target video, it can acquire multiple video segments in the target video through a sliding window mechanism, wherein the multiple video segments include the first video segment and the second video segment.
[0158] In some embodiments, any two adjacent video segments among the plurality of video segments may have video frames with overlapping display times.
[0159] For example, such as Figure 3 As shown, the target video includes M video frames, and the video frames in the target video constitute a frame sequence {v1,v2,……,vM}. This method can obtain P video segments in the target video through a sliding window mechanism. The P video segments include video segment 1 and video segment 2, where P is an integer greater than or equal to 2.
[0160] At least two adjacent video segments among the P video segments may have the same video frames, but the video frames of the two adjacent video segments are not exactly the same. For example, video segment 1 includes the video frame sequence {v1,v2,...,v10}, and video segment 2 includes the video frame sequence {v8,v9,...,v18}. Video segments 1 and 2 have overlapping video frames v8, v9, and v10. In other words, among two adjacent video segments, at least one video frame at the end of the video segment that appears earlier in the display time sequence (e.g., v8, v9, v10 above) is the same video frame as at least one video frame at the beginning of the video segment that appears later in the display time sequence (e.g., v8, v9, v10 above).
[0161] In some embodiments, the method can segment the target video using overlapping sliding windows to obtain multiple video segments with the same video frames between adjacent video segments.
[0162] An overlapping sliding window is a sliding window in which, after each sliding step, the current position and the previous position in the target video's video frame sequence contain the same video frame. The step size of an overlapping sliding window is smaller than the window size.
[0163] For example, the size (width) of the sliding window can be d (frames), and the step size can be s (frames), where both s and d are positive integers and s<d. The position of the sliding window W(d,s) on the video frame sequence of the target video can be obtained by the following formula (1), specifically:
[0164] W(d,s)={[i,i+d-1]:1≤i≤M,i=1(mods)}; (1)
[0165] Wherein, M is the total number of frames of the target video, i is the starting video frame at each position of the sliding window on the video frame sequence of the target video, and [i,i+d-1] indicates that the position of the sliding window on the video frame sequence of the target video is from the i-th frame to the (i+d-1)-th frame of the target video. i≡1(mod s) means that the remainder after dividing the integer i by s is 1, which is equivalent to i=a*s+1, a=0,1,2,……
[0166] For example, if the total number of frames of the target video is 50, the size of the sliding window is 10 (frames), and the step size is 8 (frames), then when the sliding window is at the initial position, i=0*8+1=1, and the initial position of the sliding window on the video frame sequence of the target video is [1,1+10-1], that is [1,10], which means the initial position of the sliding window on the video frame sequence of the target video is from the 1st frame to the 10th frame of the target video. When the sliding window is at the next position, i=1*8+1=9, and the next position of the sliding window on the video frame sequence of the target video is [9,9+10-1], that is [9,18], which means the next position of the sliding window on the video frame sequence of the target video is from the 9th frame to the 18th frame of the target video. By analogy, the sliding window can split the 50-frame target video into 6 video segments, which sequentially include the 1st frame to the 10th frame, the 9th frame to the 18th frame, the 17th frame to the 26th frame, the 25th frame to the 34th frame, the 33rd frame to the 42nd frame, and the 41st frame to the 50th frame of the target video.
[0167] In this embodiment, since the generation of long-duration videos involves a large amount of computation, directly processing the entire video would lead to high computational resource load and low generation efficiency. Therefore, this embodiment divides the target video (e.g., a long-duration video) into multiple video segments and uses a sliding window mechanism to segment the video, ensuring consistency and smooth transitions between segments. Furthermore, this embodiment can use overlapping sliding windows to segment the target video, ensuring that the end of one video segment and the beginning of the next in adjacent video segments have the same video frames. This allows multiple GPUs to communicate and fuse the hidden features of the same video segment during subsequent parallel processing of multiple video segments, ensuring consistency between video segments and smooth transitions between adjacent video segments, thereby guaranteeing the global consistency and coherence of the stylized target video.
[0168] In the above Figure 2 In S121 shown, the global style features of the global frame obtained by this method can be obtained by preprocessing the target video. In conjunction with any of the above embodiments, the following describes... Figure 4a This illustrates one possible implementation of preprocessing the target video to obtain global style features of the global frame.
[0169] In one possible implementation, the method can be achieved through, for example... Figure 4a The process shown is used to preprocess the target video to obtain global style features of the global frame.
[0170] like Figure 4a As shown, the process may include, but is not limited to, S201 and S202.
[0171] S201: Obtain the first k video frames from multiple video frames of the target video, sorted in descending order based on at least one of the degree of pixel change and the degree of scene change, to obtain k global frames, where k is a positive integer.
[0172] In some embodiments, the specific value of k can be preset or user-defined.
[0173] For example, a mobile application can use the method of the embodiments of this application to process the style of a target video into a target style, so as to output a target video with the target style, and the specific value of k can be preset in the mobile application.
[0174] For example, when the application scenario is a cloud scenario, such as Figure 1The cloud infrastructure 104 shown can obtain the information indicating the number of global frames k sent by the tenant (client) 102 through the cloud management platform 103 to obtain the specific value of k. For example, the information indicating the number of global frames k can be carried in the inference request sent by the tenant (client).
[0175] In some embodiments, multiple video frames in the target video can be sorted according to the degree of change (at least one of pixel change degree and scene change degree) in descending order of degree of change, so that the top k video frames are used as global frames.
[0176] The degree of change can be the degree of pixel change. In this way, the method can take the top k video frames in descending order of the degree of pixel change from multiple video frames of the target video as global frames. That is, the k video frames with the largest degree of pixel change in the target video can be taken as global frames.
[0177] The degree of change can also be the degree of scene change. In this way, the method can take the top k video frames in descending order of the degree of scene change from multiple video frames of the target video as global frames. That is, the k video frames with the greatest degree of scene change in the target video can be taken as global frames.
[0178] The degree of change can also be the overall degree of change of pixel changes and scene changes. In this way, the method can take the top k video frames in descending order of the overall degree of change of pixel changes and scene changes in the target video as global frames. That is, the k video frames with the largest overall degree of change of pixel changes and scene changes in the target video can be taken as global frames.
[0179] In some embodiments, the degree of pixel variation between adjacent frames in the method can indicate the magnitude of the mean squared error (MSE) of pixels between adjacent video frames (referred to as "adjacent frames") in the target video.
[0180] In some embodiments, for any two adjacent video frames in the target video, the method can calculate the mean square error (MSE) of the pixels between these two adjacent video frames using the following formula (2). i ,v i+1 ):
[0181]
[0182] Where H is the height of each video frame (i.e., each image frame) in the target video, W is the width of each video frame (i.e., each image frame) in the target video, and v i (x,y) represents video frame v in the target video. i The pixel value at coordinate (x, y) in the middle, vi+1 (x,y) represents video frame v in the target video. i+1 The pixel value at the coordinate (x, y) position in the middle.
[0183] In some embodiments, video frame v in the target video i The degree of pixel variation can indicate the video frame v i With video frame v in the target video i+1 Mean square error (MSE) of pixels between i ,v i+1 The method can obtain the k video frames with the largest mean square error in the target video as global frames, that is, obtain the k video frames with the largest pixel change in the target video as global frames.
[0184] In some embodiments, in order to compare the MES of different video frames in the target video at the same scale, the video frame v calculated according to the above formula (1) can be compared. i Mean square error MSE(v) i ,v i+1 Normalization is performed to obtain video frame v i normalized mean square error Then, the normalized mean square error of the M video frames in the target video is... The comparisons are performed, and the k video frames with the largest normalized mean square error are selected as global frames. Thus, the video frames v in the target video... i The degree of pixel variation can indicate the video frame v i normalized mean square error
[0185] In some embodiments, the method can calculate the video frame v in the target video using the following formula (3). i normalized mean square error
[0186]
[0187] Wherein, min(MSE) represents the video v in the target video. i To video frame v M-1 The minimum value among the MSE values (calculated according to Formula 1 above), max(MSE) represents the value of video frame v in the target video. i To video frame v M-1 The largest value among the MSEs (calculated according to Formula 1 above).
[0188] In some embodiments, the degree of scene change between adjacent frames in the method can indicate the degree of change in the color histogram between adjacent frames in the target video.
[0189] In some embodiments, for any two adjacent video frames in the target video, the degree of change in the color histogram between the two adjacent video frames can be determined by calculating the distance between their color histograms (e.g., the Bhattacharyya distance). Specifically, the video frame v can be calculated using the following formula (4). i and video frame v i+1 The distance D between the color histograms of Bhattacharyya BH (h i ,h i+1 ):
[0190]
[0191] Among them, h i Indicates video frame v i The color histogram, h i+1 Indicates video frame v i+1 The color histogram, h i (b) is video frame v i The histogram value of the b-th color channel, h i+1 (b) is video frame v i+1 The histogram value of the b-th color channel, where B is the number of bins in the color histogram, which refers to the number of intervals into which the pixel value range of the video frame (image) is divided.
[0192] In this embodiment, b and B have different meanings.
[0193] For example, when a video frame is an RGB or YUV image, the pixel value range for each color channel is 0 to 255. Therefore, B can be 256. This means the pixel value range of the video frame can be divided into 256 intervals, each interval corresponding to a pixel value (from 0 to 255). Thus, the video frame v... i The histogram values for different pixel value ranges refer to the number of times each pixel value appears.
[0194] In some embodiments, if the video frame in the target video is an image with multiple color channels, then the method can calculate the Bhattacharyya distance for each of the multiple color channels (e.g., each color channel) of the video frame.
[0195] In this way, the changes in the color histograms of multiple color channels in adjacent frames can be considered to obtain the global frame, which can capture the overall fluctuations of all color components in the target video and avoid misjudgments caused by local fluctuations or noise in a single color channel.
[0196] For example, if the target video frame is an RGB image, then this method can calculate the Bhattacharyya distance for the color histograms corresponding to the R, G, and B channels of the video frame to obtain the global frame.
[0197] Alternatively, this method can also calculate the Bhattacharyya distance for the color histogram corresponding to any color channel of the video frame. In this way, if the target style is related to specific color characteristics (such as a retro film style that enhances the red channel or a technological style that enhances the blue channel), then only the changes in the color histogram corresponding to a single color channel can be considered when acquiring the global frame, and the key information can be directly focused.
[0198] For example, if the target video frame is an RGB image and the target style is a retro film style with enhanced red channels, then this method can also calculate the Bhattacharyya distance for the color histogram corresponding to the R channel of the video frame to obtain the global frame.
[0199] In some embodiments, video frame v in the target video i The degree of scene change can indicate the video frame v i With video frame v in the target video i+1 The method determines the Bhattacharyya distance between the color histograms. It can obtain the k video frames corresponding to the k largest Bhattacharyya distances in the target video as global frames, that is, obtain the k video frames with the greatest scene change in the target video as global frames.
[0200] In some embodiments, in order to compare the Bhattacharyya distances of different video frames in the target video at the same scale, the video frame v calculated according to the above formula (4) can be compared. i Bhattacharyya is far from D BH (h i ,h i+1 Normalization is performed to obtain video frame v i Normalized Bhattacharyya distance Then, the normalized Bhattacharyya distance of the M video frames in the target video is calculated. The comparisons are performed, and the k video frames with the largest normalized Bhattacharyya distance are selected as global frames. Thus, video frames v in the target video... i The degree of scene change can indicate the video frame v i Normalized Bhattacharyya distance Size.
[0201] In some embodiments, the method can calculate the video frame v in the target video using the following formula (5). i Normalized Bhattacharyya distance
[0202]
[0203] Wherein, min(D) BH ) represents the video v in the target video. i To video frame v M-1 The minimum value among the Bhattacharyya distances (calculated according to Formula 3 above), max(D) BH ) represents the video v in the target video. i To video frame v M-1 The largest value among the Bhattacharyya distances (calculated according to Formula 4 above).
[0204] In some embodiments, the method can obtain the k video frames with the largest weighted result of the mean square error of the pixels of the corresponding adjacent video frames and the Bhattacharyya distance between the color histograms of the adjacent video frames, and use these as the k global frames.
[0205] In some embodiments, the method can indicate video frames v in the target video. i Data on the degree of pixel variation (e.g., normalized mean square error), and video frame v i The color histograms are weighted and summed based on their degree of variation (e.g., the normalized Bhattacharyya distance), and the top k video frames with the largest weighted sum in the target video are selected as global frames.
[0206] In some embodiments, the method can determine the normalized MSE(v) corresponding to video frame vi using the following formula (6). i ,v i+1 ) and D BH (h i ,h i+1 The weighted result w i :
[0207]
[0208] Here, α is a weighting parameter used to control the impact of pixel and scene changes on the selection of global frames. For example, α can be a hyperparameter. When α is large, this method tends to select video frames with greater pixel changes as global frames; when α is small, this method tends to select video frames with greater scene changes as global frames.
[0209] In some embodiments, the specific value of α can be preset or user-defined.
[0210] In some embodiments, the method can use the normalized MES and Bhattacharyya distance of each frame in the target video to perform a weighted calculation using the above formula (6) to obtain the weighted result w corresponding to each frame. The weighted results corresponding to each frame are then sorted in ascending order, and the top k video frames with the most significant changes are determined as global frames based on the sorting results. In this way, the top k video frames with the largest combined changes in pixel change and scene change (i.e., overall change) can be obtained as global frames, denoted as {g1, g2, ..., gk}. These global frames can be used in the temporal action module of the above AI model (e.g., ...). Figure 5a In the timing action module shown.
[0211] like Figure 4a As shown, after S201, the method may further include the following S202.
[0212] S202: Obtain the global style features of each of the k global frames obtained in S201.
[0213] In this way, the global style features of each of the acquired k global frames can be cached. (As mentioned above...) Figure 2 In S121 shown, the global style features of each of the k global frames can be extracted from the cached data without having to calculate the global style features in real time during inference, thereby improving inference efficiency and video stylization efficiency.
[0214] Optionally, there are two possible implementations of S202, which will be described below. In other embodiments, there may be more implementations, without limitation.
[0215] In the first possible implementation, when obtaining the KV pairs of each of the k global frames obtained in S201 above, the method can input the k global frames into the temporal action module in the AI model used to stylize video segments to obtain the global style features of each of the k global frames.
[0216] In this way, the global style features of each global frame obtained indicate the style features of the corresponding global frame in the target video, that is, the original style features of the corresponding global frame in the target video to be processed.
[0217] The temporal action module can be AnimateDiff. In some embodiments, the AI model may include an image generation module, which can be used to stylize single-frame images in a video clip. However, in image stylization tasks, the success of single-frame stylization does not guarantee the continuity of actions in the video. To address the issues of inconsistency between frames and video flicker, the AI model in this application embodiment also includes a temporal action module, such as AnimateDiff. This temporal action module can solve the temporal consistency problem in the stylization of short videos (e.g., video clips). Specifically, AnimateDiff performs fine-tuning on video clips through Stable Diffusion, achieving high-quality stylization generation within a video clip of up to 32 frames (e.g., video clip 1 above includes 32 video frames), thereby ensuring that the generated video has continuous and unabrupt actions.
[0218] AnimateDiff relies on relative temporal information between video frames during the temporal generation process. It utilizes short-term dependencies (e.g., in the example video clip 1 above, which includes 32 video frames, the short-term dependency is the contextual information of those 32 frames) to enhance motion consistency during generation, avoiding significant inter-frame differences and unnatural flickering that may occur in traditional methods. In this way, AnimateDiff ensures that the stylization effect of each frame remains consistent with the overall temporal behavior of the video, further improving the quality of video stylization.
[0219] In the second possible implementation, when obtaining the KV pairs of each of the k global frames obtained in S201 above, the method can input the information indicating the target style, the k global frames, and at least one frame sequence obtained based on the third video segment into the AI model used to stylize the video segment to obtain the global style features of each of the k global frames.
[0220] The information indicating the target style could be, for example, the information carried in the aforementioned reasoning request.
[0221] The third video segment can be either the first or second video segment from the target video, or it can be any other video segment obtained from the target video.
[0222] The at least one frame sequence obtained based on the third video segment can be at least one of the outline video frame sequence and the color video frame sequence corresponding to the third video segment.
[0223] Here, we take the third video segment, which contains N video frames, as an example.
[0224] For example, the method can obtain a contour video sequence by extracting the contour image (i.e., contour video frame) of each video frame in the third video segment, for example, denoted as {o1,o2,……,oN}, where contour video frame o1 is the contour image of video frame v1.
[0225] For example, the color video frame sequence corresponding to the third video segment can be the frame sequence {v1,v2,……,vN} of the third video segment itself, which is represented here as {c1,c2,……,cN}, where color video frame c1 = video frame v1.
[0226] Thus, this method obtains global style features for each global frame based on the latent space features of the k global frames and the third video segment. The latent space features of the third video segment indicate the stylized latent space features of each frame in the third video segment (the latent space features of the target style for each frame in the third video segment). Therefore, the global style features of each global frame can fuse the original style features of the corresponding global frame in the target video to be processed, as well as the stylized latent space features of each frame in the third video segment. In other words, the global style features of each global frame obtained by this method not only indicate the style features of the corresponding global frame in the target video but also indicate the style features of the video frames in the third video segment after stylization with the target style.
[0227] In some embodiments, the method may also cache the global style features of each of the k global frames obtained from preprocessing for use in subsequent inference processes.
[0228] In some embodiments, the method described above Figure 2 In step S121, when obtaining the global style features of the global frames obtained from the preprocessing of the target video, the key-value (KV) pairs of each of the k global frames obtained from the preprocessing of the target video can be obtained.
[0229] In some embodiments, the KV pairs of each of the k global frames can indicate the global style features of the corresponding global frame.
[0230] For example, it can be done through the above Figure 4a S201 shows that k global frames are obtained. Then, global style features are obtained through the first possible implementation in S202 shown in Figure 4. The global style features are, for example, KV pairs of global frames. Then, the KV pairs of each of the k global frames indicate the global style features of the corresponding global frame.
[0231] In some embodiments, the KV pair of each of the k global frames can indicate the style features of the corresponding global frame in the target video and the style features of the video frames in the third video segment after being stylized in the target style.
[0232] For example, it can be done through the above Figure 4a As shown in S201, k global frames are obtained. Then, global style features are obtained through the second possible implementation in S202 shown in Figure 4. These global style features are, for example, KV pairs of global frames. Then, the KV pairs of each of the k global frames indicate the style features of the corresponding global frame in the target video and indicate the style features of the video frames in the third video segment after being stylized with the target style.
[0233] In some embodiments, the key-value pairs of each global frame are intermediate results obtained by the timing action module processing the corresponding global frame.
[0234] In this embodiment, global frames can be adaptively selected based on the inter-frame fluctuation distribution, thereby effectively ensuring global consistency during the stylization process. Specifically, this embodiment can calculate the fluctuation information between target video frames and select video frames with strong representativeness (drastic fluctuations) as global frames. For example, the selected global frames can be the k video frames with the largest changes in at least one of pixel change and scene change in the target video. The global frames obtained in this way are usually video frames with drastic motion (corresponding to large pixel changes) or video frames with scene transitions (corresponding to large scene changes) in the target video. In the subsequent inference process, the global style features of each of the k global frames can be used for the inference of the AI model (e.g., the calculation of temporal attention), which can ensure the global consistency of the stylized target video generated by this method, and can filter redundant information (such as consecutive similar frames), reduce noise interference, and improve the adaptability of this method to complex scenes (such as rapid motion, sudden changes in lighting).
[0235] In this embodiment, when obtaining the global style features of each of the k global frames, not only can the k global frames be input into the AI model used for stylizing video segments, but also information indicating the target style and at least one frame sequence obtained based on a third video segment can be input into the AI model. Thus, the global style features of each global frame can indicate the style features of the corresponding global frame in the target video to be processed, enabling the global style features to indicate the style features of each scene in the target video and the style features of video frames with significant pixel variations. Furthermore, the global style features of the global frames can also indicate the style features of video frames in the third video segment after stylization with the target style. The third video segment is any video segment in the target video. During inference, the AI model can utilize the global style features of each of the k global frames (e.g., the key-value pairs of each global frame) for inference (e.g., the calculation of temporal attention). In this way, the global style feature of each global frame indicates the style feature (original style feature) of the corresponding global frame in the target video. This ensures that the content structure (such as object outlines and motion trajectories) in the target video is not destroyed by the stylization process, guaranteeing the global consistency of each video frame in the stylized target video. In addition, the global style feature of each global frame indicates the style feature of the video frame in the third video segment after being stylized with the target style. This can provide the required target style information, thereby ensuring the style consistency of the stylized target video on the time axis, preventing style drift between frames, and making the stylized target video maintain a consistent target style visual effect across different frames.
[0236] In some embodiments, such as Figure 4b As shown, the AI model used in this method may include an image generation module, a timing action module, and a control network module (e.g., control network module 1 and control network module 2).
[0237] The image generation module can be, for example, a diffusion model such as Stable Diffusion. This image generation model can perform individual stylization processing on each frame of the target video, obtaining the stylized latent space features of each frame.
[0238] The control network module can provide control conditions for the image generation module, ensuring that each stylized frame obtained by the image generation module matches the original frame of the target video in detail. For example, if the image generation module is a diffusion model such as Stable Diffusion, this model can progressively denoise the random noise based on the input random noise and a series of control conditions (also called denoising conditions) to generate the final image. In this case, the control network module can act as a conditional encoder for the diffusion model, providing the aforementioned denoising conditions.
[0239] Control network module 1 and control network module 2 can be, for example, ControlNet models. Control network module 1 can provide contour conditions, for example, and control network module 2 can provide color conditions, for example.
[0240] For example, a timing action module could be AnimateDiff, as detailed above. Figure 4a The relevant information will not be repeated here.
[0241] In addition, for example, the image generation module described above can also be Flux, Playground, etc.
[0242] For example, the control network module described above can also be a T2I-Adapter, Composable Diffusion, etc. Control network module 1 and control network module 2 can be the same or different, and this application does not impose any restrictions.
[0243] The following is combined with Figure 4b To illustrate the second possible implementation of obtaining the global style features of the global frame in S202 above, a specific example will be used. Figure 4b The example below uses global style features as key-value pairs for global frames.
[0244] like Figure 4b As shown, the third video segment mentioned above is video segment 1 in the target video, and the video frame sequence of video segment 1 is {v1,v2,……,vN}.
[0245] First, this method can input the video frame sequence {v1,v2,……,vN} of the target video into the contour detection module to obtain the frame sequence {o1,o2,……,oN} of the contour video frames corresponding to video segment 1. Furthermore, this method can obtain the frame sequence {c1,c2,……,cN} of the color video frames based on the video frame sequence {v1,v2,……,vN}. For example, this method can add image details to each frame of the video frame sequence {v1,v2,……,vN} to improve image quality, and then use the improved video frame sequence {v1,v2,……,vN} as the frame sequence {c1,c2,……,cN} of the color video frame.
[0246] Subsequently, the method can input the frame sequence of the contour video frame into the control network module 1 of the AI model, input the frame sequence of the color video frame into the control network module 2 of the AI model, and input the frame sequence of the video segment 1 and the information indicating the target style into the image generation module of the AI model.
[0247] Thus, control network module 1 can output a contour feature sequence including N contour features to the image generation module, and control network module 2 can output a color feature sequence including N color features to the image generation module. Then, based on the contour feature sequence, color feature sequence, video frame sequence of video segment 1, and information indicating the target style, the image generation module can output the stylized latent space features of each frame (a total of N frames) in the frame sequence of video segment 1 to the temporal action module. Here, the stylized latent space features of each frame input to the temporal action module refer to the sequence of stylized latent space features (also known as latent space features) of each frame output by the image generation module.
[0248] In other words, the frame sequences of the outline video frames and the color video frames mentioned above can be used as inputs to control network module 1 and control network module 2, respectively, to generate conditional features (denoising conditions). These features can be combined with the frame sequence of video segment 1 and input into the image generation module, for example, into the Stable Diffusion UNet structure. Thus, the image generation module can stylize each frame of video segment 1 according to the denoising conditions provided by control network module 1 and control network module 2 and the frame sequence of video segment 1, based on the target style, to output the stylized latent space features of each frame.
[0249] Next, the timing action module can obtain the key-value pairs for each of the k global frames based on the stylized latent space features of the N frames and the k global frames. These key-value pairs are the intermediate results of the timing action module. The k global frames used by the timing action module are pre-input into the timing action module and can be the k global frames obtained in S201 above.
[0250] Finally, the method can cache the KV pairs of each of the obtained k global frames for use in subsequent inference processes.
[0251] In conjunction with any of the above embodiments, the following is a description of... Figure 5a To explain the above Figure 2 The S102 shown is one possible specific implementation process for stylizing video clips to obtain video clips with the target style. However, the process of stylizing video clips is not limited to this. Figure 5a Examples. Figure 5a The example below uses global style features as key-value pairs for global frames.
[0252] Please refer to Figure 5aFor example, the first video segment obtained in S111 above can be video segment 1, and the second video segment can be video segment 2. The frame sequence of video segment 1 is {v1,v2,……,vN}, and the frame sequence of video segment 2 is {vN-1,vN,……,vM}. The frame sequences of these two video segments have the same video frames vN-1 and vN.
[0253] Figure 5a This example demonstrates how the method stylizes video clips 1 and 2 in a distributed, parallel manner when the AI model is running on different GPUs. The following detailed explanation uses video clip 1 as an example.
[0254] like Figure 5a As shown, this method can input the frame sequence {v1,v2,……,vN} of video segment 1 into the AI model. Optionally, the frame sequence of the outline video frame and the frame sequence of the color video frame can also be input into the AI model. Furthermore, this method can also input information indicating the target style into the AI model.
[0255] Compared to Figure 4b , Figure 5a The processing procedures of control network module 1, control network module 2, and image generation module within the AI model, and... Figure 4b The processes shown are all the same; please refer to the following for details. Figure 4b The relevant information will not be repeated here.
[0256] In some embodiments, Figure 5a In the process of obtaining the frame sequences of the aforementioned contour video frames and the aforementioned color video frames, the process is performed before inference using the AI model. If, after preprocessing the target video, the frame sequences of the contour video frames and the frame sequences of the color video frames corresponding to video segment 1 have already been obtained, then in Figure 5a In the embodiment shown, the frame sequence of the contour video frame corresponding to video segment 1 and the frame sequence of the color video frame corresponding to video segment 1 can be used directly without contour detection.
[0257] Reference Figure 5a , combined Figure 4b As can be seen from the introduction, the image generation module can output the stylized latent space features of each frame (a total of N frames) in the frame sequence of video clip 1 to the temporal action module.
[0258] The stylized latent space features input to the timing action module already contain the timing information of video clip 1, so there is no need to input video clip 1 separately to the timing action module.
[0259] Different from Figure 4b The process, such as Figure 5aAs shown, this method can extract key-value pairs (KV pairs) of k global frames from cached data and input them into the temporal action module of the AI model. In this way, the temporal action module can smoothly transition consecutive video frames based on the KV pairs of each of the k global frames and the stylized latent space features of N frames from the image generation module, to output a video segment 1 stylized according to the target style.
[0260] The key-value pairs for each of the k global frames can be as described above. Figure 2 The KV pairs are obtained in S202. For example, after obtaining the KV pairs of each of the k global frames in S202, the KV pairs of each of the k global frames are cached. Then, this method can input the cached KV pairs of each of the k global frames into the timing action module during model inference.
[0261] In some embodiments, this method can fine-tune model parameters using Low-Rank Adaptation (LoRA) to make the AI model of this application compatible with various video styles and avoid generating low-quality stylized videos. For example, LoRA can be applied to... Figure 5a The temporal attention calculation process of the temporal action module of the AI model shown is achieved by injecting a low-rank decomposition matrix into the weight matrix of the temporal action module (such as the Q / V matrix of the attention layer). This method only requires training 1%-10% of the parameters of the temporal action module of the original AI model, thereby reducing the computational cost.
[0262] In some embodiments, the AI model used in this method includes a text encoder (e.g., Figure 5a The image generation module shown includes a text encoder. This method applies textual inversion by training a learnable placeholder token embedding vector. This embedding vector is then mapped to the latent space features of the target style in the AI model's text encoder. For example, the image generation module of this AI model can be triggered to generate images in the target style via text prompts. This method requires training only a very small number of embedding vectors (typically <1MB), reducing computational costs. Furthermore, this method is compatible with various video styles and avoids generating low-quality video results.
[0263] In some embodiments, the method may apply LoRA and text inversion techniques simultaneously, thereby reducing computational costs and improving the overall computational efficiency of the method.
[0264] exist Figure 5aThe process described above may only involve processing the video frames (i.e., images in the video) of the target video, without involving the processing of the audio of the target video.
[0265] exist Figure 5a In the illustrated embodiment, the method not only inputs the acquired k global frames into the temporal action module of the AI model for stylizing video clips, but also inputs the frame sequences of the contour video frames and the frame sequences of the color video frames of video clip 1 (the example of the third video clip mentioned above) into the control network modules (control network module 1 and control network module 2) of the AI model. Thus, the frame sequences of the contour video frames and the frame sequences of the color video frames can be used to control the stylization process, making it more stable and preventing distortion of the content in the generated stylized video clip. Furthermore, the KV pairs of each of the acquired k global frames not only contain the global style features of the global frames, but also the stylized latent space features of each frame in video clip 1. These KV pairs of the k global frames can be used for temporal attention calculation in the temporal action module during inference to obtain the stylized video clip 1. The KV pairs of the k global frames contain global style features that ensure the content of video segment 1 is not destroyed by the stylization process, ensuring the global consistency of the stylized video segment 1. The stylized latent space features of each frame in the video segment 1 contained by the KV pairs of the k global frames can ensure the stylized target video has stylistic consistency on the time axis.
[0266] In conjunction with any of the above embodiments, the following description continues. Figure 2 Some possible embodiments of the specific implementation process of S102 shown.
[0267] In some embodiments, the above Figure 2 In step S111, there are video frames with overlapping display times between the first and second video segments obtained. Figure 2 In S102, the AI model that outputs a first video clip with the target style runs on the first device, and the AI model that outputs a second video clip with the target style runs on the second device.
[0268] The first device and the second device are two different devices, which can be GPUs, clients, servers, etc., without restriction. For an introduction to clients and servers, please refer to the above text.
[0269] Thus, the embodiments of this application can deploy AI models in a distributed manner to achieve parallel processing of stylization tasks for different video segments in the target video, thereby improving the stylization efficiency of the target video.
[0270] In some embodiments, when the method stylizes a first video segment using an AI model running on a first device, the first device may send the first hidden layer features of the aforementioned temporally overlapping video frames calculated by the running AI model to a second device; and the first device may receive the second hidden layer features of the temporally overlapping video frames calculated by the AI model running on the second device, sent by the second device; subsequently, the first device may perform feature fusion on the calculated first hidden layer features and the received second hidden layer features to output a first video segment with a target style. The temporally overlapping video frames may be one frame or multiple frames.
[0271] Similarly, when the method stylizes the second video segment using an AI model running on a second device, the second device can send the second hidden layer features of the video frames with overlapping display times calculated by the running AI model to the first device; and the second device can receive the first hidden layer features of the video frames with overlapping display times calculated by the AI model running on the first device sent by the first device; then, the second device can perform feature fusion on the calculated second hidden layer features and the received first hidden layer features to output a second video segment with the target style.
[0272] For any two video segments in the target video that have overlapping display time frames, when the two video segments are processed separately by running AI models on different devices, the two devices can communicate the hidden features of the overlapping display time video frames and perform feature fusion on the hidden features of the same video frame obtained by each device to output the corresponding video segment with the target style.
[0273] The hidden layer features mentioned above will be explained below.
[0274] First, it's important to understand that a neural network typically consists of multiple layers, usually including an input layer, hidden layers, and an output layer.
[0275] Output layers typically accept raw data input (such as pixel values of an image or word vectors of text). Hidden layers are intermediate layers in a neural network, usually containing multiple neurons. Each hidden layer transforms the input, extracts features, and gradually maps the input to a higher-dimensional, more abstract feature space. Output layers are typically used to generate the final prediction or classification result based on the output of the hidden layers.
[0276] The aforementioned hidden features, also known as hidden layer representations, hidden layer outputs, latent space features, or latent space features, refer to the features generated in the hidden layers of a neural network. These features are typically abstract information extracted from the input data, used to help the model make predictions or decisions. Hidden features are computed in the intermediate layers of the neural network (i.e., the layers between the input and output layers), and they are usually more abstract and informative than the input data.
[0277] In some embodiments, the AI model running by the first device may include a timing action module, which includes a neural network. When processing video frames with overlapping display times, the timing action module can perform the aforementioned communication and feature fusion on the hidden features output by each hidden layer of the neural network, and then output the result of feature fusion to the next hidden layer. In this way, except for the first hidden layer connected to the input layer, the input of each hidden layer of the neural network included in the timing action module is the result of feature fusion. The second device can perform a similar process for feature fusion, so as to ensure that the video frames with overlapping display times are consistent in the stylized first video segment and the stylized second video segment.
[0278] In some embodiments, when fusing the first hidden layer features and the second hidden layer features, the first device may perform a weighted calculation on the first hidden layer features and the second hidden layer features to obtain the feature fusion result. The specific feature fusion method is not limited. The principle of feature fusion in other devices is similar.
[0279] Please refer to Figure 5b For example, the first device's GPU1 and the second device's GPU2 each run AI models for stylizing video clips. Both AI models include a temporal action module. The AI model running on the first device is used to stylize video clip 1, and the AI model running on the second device is used to stylize video clip 2. The specific principles can be found above. Figure 5a Examples of implementations. Figure 5b The example below uses global style features as key-value pairs for global frames.
[0280] For example Figure 5b The video clip 1 (e.g., video frames v1 to vN) and video clip 2 (e.g., video frames vN-1 to vM) shown have video frames vN-1 and vN with overlapping display times. Taking the video frame vN-1 as an example for feature fusion, the process of feature fusion of video frames with overlapping display times is explained.
[0281] The AI model (referred to as AI model 1) running on GPU1 of the first device may include a timing action module, combined with Figure 5aIt can be understood that the input of the timing action module can be the stylized latent space features of video frame vN-1. The timing action module can include multiple hidden layers. The first hidden layer among these multiple hidden layers processes the stylized latent space features of video frame vN-1 based on the KV pairs of k global frames to output hidden layer feature 1.
[0282] like Figure 5b As shown, the AI model running on GPU2 of the second device (referred to as AI model 2) may include a timing action module, combined with Figure 5a The principle can be understood that the input of the timing action module can be the stylized latent space features of video frame vN-1. The timing action module can include multiple hidden layers. The first hidden layer among these multiple hidden layers is based on the KV pairs of k global frames to process the stylized latent space features of video frame vN-1 to output hidden layer feature 1'.
[0283] GPU1 can send the hidden layer feature 1 to GPU2, and GPU2 can send the hidden layer feature 1' to GPU1. In this way, the timing action module on GPU1 can perform feature fusion of hidden layer feature 1 and hidden layer feature 1' to obtain feature fusion result 1, and then output the feature fusion result 1 to the second hidden layer in the timing action module. After that, the second hidden layer in the timing action module on GPU1 can output hidden layer feature 2.
[0284] Similarly, the timing action module on GPU2 can fuse hidden layer feature 1 and hidden layer feature 1' to obtain feature fusion result 1', and then output the feature fusion result 1' to the second hidden layer in the timing action module on GPU2; afterwards, the second hidden layer in the timing action module on GPU2 can output hidden layer feature 2'.
[0285] In this manner, the two GPUs continuously exchange the hidden layer features corresponding to the video frame vN-1 they obtain, in order to fuse the hidden layer features, and use the fused hidden layer features as the input of their respective next hidden layer. This process is repeated until the output of the last hidden layer in the timing action module of the two GPUs is fused.
[0286] In this way, AI model 1 on GPU1 can output stylized video clip 1, and AI model 2 on GPU2 can output stylized video clip 2. In stylized video clip 1 and stylized video clip 2, the video frames vN-1 and vN with overlapping display times have a consistent style.
[0287] The above text Figure 3The corresponding embodiment mentions that the method can segment the target video through overlapping sliding windows, so that the target video can be divided into multiple video segments, and each video segment includes a portion of video frames with overlapping display times with adjacent video segments.
[0288] In some embodiments, for video frames i (also called the i-th video frame) that exhibit temporal overlap between two adjacent video segments (e.g., video segment 1 and video segment 2), the method can denote the hidden layer features corresponding to video frame i as e. t,+ (l,i,r), can be obtained through the following formula (7):
[0289]
[0290] Where i represents the sequence number of the video frame in the target video (different from the meaning of i in formulas (1) to (6) above), l represents the starting video frame in the video frames with overlapping display times in two adjacent video segments, and r represents the ending video frame in the video frames with overlapping display times in two adjacent video segments. For example, l = 8 and r = 10, which means that the video frames with overlapping display times in two adjacent video segments are: the 8th video frame in the target video (e.g., v8 above) to the 10th video frame in the target video (e.g., v10 above). Let represent the i-th video frame, the outline video frame corresponding to the i-th video frame, and the color video frame corresponding to the i-th video frame, respectively. t is the time step, and τ is the prompt. This indicates the case where i is in the interval from l to r. The range of values for i is from l to r, for example, i = 8, 9, 10.
[0291] by Figure 5b Taking video clip 1 as an example, Video frame vi represents the video frame in the frame sequence {v1,v2,……,vN} of video segment 1. The video frame oi represents the outline of video frame {o1,o2,...,oN} in the video frame sequence. The video frame ci represents the video frame ci in the frame sequence {c1,c2,……,cN} of the color video frame of video segment 1.
[0292] The time step t mentioned above refers to a unit of step in the sequence data processing or generation process during AI model inference. For example, when the AI model processes the video frame sequence {v1,v2,……,vN}, each time step corresponds to the prediction of one element in the input (such as the generation of a stylized video frame vi).
[0293] The aforementioned cue τ refers to the input text provided by the user to the model, used to guide the model to generate specific content or perform tasks. For example, cue τ may be derived based on the information indicating the target style mentioned above.
[0294] To enhance consistency between video segments and smooth transitions between different windows, this application employs a weighted fusion mechanism. The features of each video segment are weighted and smoothly fused using the overlapping portions with other windows.
[0295] For example, the hidden features of the i-th video frame in the target video obtained by the first device and the hidden features of the i-th video frame in the target video obtained by the second device can be weighted and fused using the following formula (8) to obtain the feature fusion result. Specifically, it could be:
[0296]
[0297] Here, (l,r) represents the range of video frames i that overlap in time between two adjacent video segments corresponding to the sliding window W(d,s). In other words, (l,r) represents the overlapping video frames between two adjacent video segments. Specifically, the video frames i that overlap in time between two adjacent video segments can be any one of the l-th to r-th video frames, and f(l,i,r) represents the hidden layer feature e of the i-th video frame in the target video. t,+ The weights of (l,i,r) and the description of the sliding window W(d,s) can be found in the explanation of formula (1) above.
[0298] Where, ∑ (l,r)∈w(d,s) f(l,i,r) represents the sum of the weights of the hidden layer features of the i-th video frame obtained by the first device and the weights of the hidden layer features of the i-th video frame obtained by the second device.
[0299] This represents the normalized result of the weight f(l,i,r) of the hidden feature obtained from one of the two devices on the device side of the i-th video frame with temporal overlap between two adjacent video segments, obtained through the sliding window W(d,s). This avoids the problem of excessively large differences in the weights of the hidden features obtained from different devices for the same video frame in temporally overlapping video frames of an adjacent video segment.
[0300] Formula (8) above can use the hidden layer features e between two video segments with respect to the i-th video frame. t,+ (l,i,r) is weighted and summed with its normalized weights to obtain the feature fusion result of the hidden layer features of the i-th video frame between the two video segments.
[0301] For example, the hidden layer feature e of the i-th video frame in the target video can be calculated using the following formula (9). t,+ The weights f(l,i,r) of (l,i,r) are as follows:
[0302]
[0303] Based on the above formula (9), the weights of the hidden layer features of the video frames that show temporal overlap in two adjacent video segments (e.g., the l-th to r-th video frames in the target video) can be obtained through... The calculation shows that the weight of the hidden features of the video frames that do not overlap in display time between the two video segments is 0. That is, when i is not in the range [l,r], the weight of the hidden features of the i-th video frame is 0.
[0304] In some embodiments, the method can store the hidden layer features corresponding to the video frames whose display times overlap in the hidden layer output in the random access memory (RAM) of the device (e.g., the first device), and move the hidden layer features to the GPU memory of the device (e.g., the memory of the GPU running the AI model in the first device) for processing when the sliding window is overwritten. In this way, memory resources can be managed flexibly while improving computational efficiency.
[0305] In this embodiment, the AI model can be distributed across multiple devices. Each device's AI model can then be used to stylize one or more video segments within the target video, enabling parallel stylization of multiple video segments and improving efficiency. Furthermore, when video segments processed by different devices have overlapping display times, the AI models on both devices can exchange the hidden features of those frames during inference and fuse them. For example, these fused hidden features can be used as input to the next hidden layer in the AI model, thus achieving feature fusion for the same video frame across different devices. This ensures stylistic consistency among the stylized video segments obtained from different devices. It also facilitates subsequent stitching of the stylized video segments, resulting in a smoother, more natural final video with the target style.
[0306] Furthermore, this application employs distributed inference technology, which enables distributed parallel inference of multiple video segments by deploying the AI model on different devices. For example, this application can distribute multiple video segments after the target video is segmented onto different GPUs for parallel stylization inference, with each GPU responsible for processing the stylization generation task of one or more video segments. This distributed inference method significantly improves the speed and efficiency of video stylization, and makes the AI model more robust when the target video has complex scenes, avoiding memory bottlenecks and computational latency issues that may occur when stylizing videos on a single device.
[0307] In conjunction with any of the above embodiments, the AI model may include a time-series action module. Figure 4a In S202 shown, when the method inputs a first video segment and global style features to an AI model for stylizing the video segment based on an inference request, in order to output a first video segment with a target style, the method may include Figure 6 The calculation process for temporal attention is shown here. This explanation uses the first video clip as an example; the processing of the second and other video clips can include a similar process.
[0308] like Figure 6 As shown, the method may include the following steps: S301, S302, S303 and S304.
[0309] S301: The temporal action module in this AI model obtains global attention based on global style features and the style features of the first video segment after being stylized with the target style.
[0310] Among them, the global style features can be as described above. Figure 2 The global style features of the global frames obtained in S121 are shown. For example, the global style features can be the KV pairs of the aforementioned k global frames to indicate the style features of the k global frames, which can be, for example, the global frames obtained in S201 above.
[0311] The style features of the first video segment after being stylized according to the target style are the intermediate results obtained by the AI model stylizing each video frame in the first video segment according to the target style.
[0312] For example, the AI model is as follows: Figure 5a Given the AI model shown, the style features of the first video segment after being stylized to the target style can be as follows: Figure 5aThe image generation module in the AI model shown outputs stylized latent space features of N video frames of the first video segment. Here, the first video segment comprises N video frames, and for each video frame in the N frames, the image generation module can output the stylized latent space features of that video frame.
[0313] Reference Figure 6 Combining with Figure 5a As can be seen, the input of this timing action module may include the KV pairs of each of the k global frames (specifically the K vector and V vector of the global frame) and the style features of the first video segment (specifically the stylized latent space features of the N video frames).
[0314] In some embodiments, the first video segment includes N video frames. Then, the timing action module can calculate the Q vector, K vector, and V vector of each video frame based on the latent space features of the input N video frames.
[0315] In some embodiments, the timing action module can calculate N global attention scores based on the Q vector corresponding to each video frame in the first video segment and the K vector of each of the k global frames. Since there are N Q vectors, the calculated attention scores (called global attention scores because K is for global frames) are also N. Then, the timing action module can calculate N global attentions based on the obtained N global attention scores and the V vector of each of the k global frames. These N global attentions correspond to each frame of the first video segment.
[0316] In a specific embodiment, when calculating the global attention score, the temporal action module can calculate the similarity between the Q vector and the K vector based on the Q vector of each video frame in the first video segment and the K vector corresponding to each global frame in the k global frames, by means of dot product, scaled dot product, etc., to obtain N global attention scores.
[0317] The N global attention scores calculated by this timing action module are derived from the K vector of the global frame and the Q vector of the video frame within the first video segment. Since the K and Q vectors are not from the same video frame, this global attention score is also called a cross-attention score. Correspondingly, the global attention subsequently calculated using the global attention score and the V vector of the global frame is also called cross-attention.
[0318] Global attention relies on the global style features of global frames (such as the KV pairs of the k global frames mentioned above) to reflect global information. Global attention focuses on long-term dependencies and can ensure the consistency of each video segment in the target video during the stylization process. Even if the scene changes in different frames of the video, the style of each frame remains consistent.
[0319] like Figure 6 As shown, the method may further include the following S302. S302 may be executed before S301, after S301, or in parallel with S301. This application does not impose specific restrictions on the order of S301 and S302.
[0320] S302: This timing action module is based on the style features of the first video segment after being stylized to the target style (e.g., ...). Figure 5a The latent space features of the N video frames shown are used to obtain local attention.
[0321] As mentioned above, the timing action module can calculate the Q vector, K vector, and V vector of each video frame based on the latent space features of the input N video frames.
[0322] The timing action module can then calculate the similarity between the Q-vector and the K-vector corresponding to each video frame in the first video segment using methods such as dot product and scaled dot product to obtain an attention score (here called a local attention score). The number of local attention scores is N, which is the same as the number of video frames in the first video segment. Subsequently, the timing action module can calculate N local attention points based on the obtained N local attention scores and the V-vector corresponding to each video frame in the first video segment. These N local attention points correspond to each frame of the first video segment.
[0323] The N local attention scores calculated by this timing action module are based on the Q and K vectors of the same frame within the first video segment; these local attention scores are also called internal attention scores. Correspondingly, the local attention subsequently calculated using the local attention scores and the V vectors of the video frames within the first video segment is also called internal attention.
[0324] Local attention focuses on adjacent video frames within the processed video segment (e.g., the first video segment), primarily focusing on short-term dependencies, i.e. the continuity of action between adjacent frames, ensuring a smooth transition between adjacent frames after stylization and avoiding abrupt action jumps.
[0325] Following S301 and S302, the method may further include S303.
[0326] S303: This timing action module obtains timing attention based on the global attention obtained in S301 and the local attention obtained in S302.
[0327] The temporal attention (also N) in this application embodiment can be calculated based on the N global attention and the N local attention.
[0328] Following S303, the method may further include S304.
[0329] S304: This timing action module can output a first video clip with the target style based on the timing attention obtained in S303.
[0330] For example, the temporal action module can obtain the temporal attention through a cross-attention mechanism (e.g., global attention obtained by cross-calculating the K and V vectors of the global frame and the Q vectors of N video frames) and a self-attention mechanism (e.g., calculated local attention). Cross-attention differs from self-attention in that it not only focuses on the information of the input sequence itself (e.g., the latent space features of N frames in the first video segment) but also on the information of another related input sequence (e.g., the KV pairs of k global frames). In this embodiment, the cross-attention mechanism allows the temporal action module to focus not only on the information of the first video segment itself but also on the global information provided by the global style features when processing the first video segment, thereby improving the global consistency of the stylized target video.
[0331] Since the temporal action module only processes video segments of no more than N frames during training, and attention calculation is difficult to directly adapt to longer key-value (KV) lengths, effectively handling these long-term dependencies becomes a key issue when generating long-term videos. Therefore, to ensure the consistency between local details and global style, this embodiment decouples the calculation of temporal attention into two parts: local temporal attention and global temporal attention. For example, the temporal action module can calculate the global attention (calculated by combining the global style features of the preprocessed global frame) and the local attention of each frame in the first video segment, thereby obtaining the temporal attention of each frame in the first video segment. This allows the temporal action module to not only focus on short-term dependencies through local attention but also on long-term dependencies through global attention, thus ensuring the consistency between local details and global style in the generated stylized target video, improving the naturalness and realism of the stylized target video, and thereby improving the adaptability of this method to complex scenes in long-term video stylization tasks, avoiding the matching problem between the training and inference phases.
[0332] In some embodiments, the temporal attention can be obtained through a temporal attention mechanism. Temporal attention is an attention mechanism used when processing sequential data, particularly suitable for video, audio, or time-series data. In temporal tasks, AI models need to focus on the relationships between different time steps in the input sequence. The temporal attention mechanism determines which parts the AI model should focus on at each time step by calculating temporal attention, thereby improving the model's understanding of temporal data. Therefore, this embodiment of the application, by calculating temporal attention, can determine the parts that the AI model needs to focus on at each time step, thus improving the AI model's understanding of video frame sequence data.
[0333] In some embodiments, the temporal action module may determine a first weight corresponding to the global attention and a second weight corresponding to the local attention based on the probability distribution of global attention and local attention; then, based on the first weight and the second weight, the global attention and local attention are weighted to obtain temporal attention.
[0334] In some embodiments, both the first weight and the second weight can be preset values.
[0335] In some embodiments, since long videos involve complex and varied scene sequences, the need for coordination between local details and global style may differ at different times. To adaptively adjust the weights of the two attention parameters according to the specific scene, the first and second weights in this embodiment can be dynamically allocated based on attention scores.
[0336] When the first weight is higher, it means that more attention needs to be paid to global features (such as the global style features of global frames) when processing video frames. When the second weight is higher, it means that more attention needs to be paid to local features (such as the stylized latent space features of the first video segment) when processing video frames. In this way, the method can adjust the size of the first and second weights to adapt to different first video segments. When inferring the first video segment, it can reasonably pay attention to global and local features, thereby avoiding the problems of poor stability and content disorder in the stylized first video segment.
[0337] In a specific embodiment, the first weight corresponding to the global attention of each frame of the first video segment can be the same, and the second weight corresponding to the local attention of each frame of the first video segment can also be the same. When calculating the temporal attention of each frame of the first video segment, the temporal action module can weight the global attention and local attention corresponding to each frame based on the first weight and the second weight to obtain the temporal attention corresponding to each frame in the first video segment.
[0338] In some embodiments, the method can determine the second weight β of local attention and the first weight 1-β of global attention using the following formula (10):
[0339]
[0340] Where max(Attn) g ) represents the maximum value among the N global attention values in the N frames of the first video segment, max(Attn) l ) represents the largest value among the N local attention values of the N frames in the first video segment.
[0341] In this way, the method can adaptively adjust the relative weights of local and global attention based on the probability distribution of global and local attention corresponding to the first video segment, so as to adapt to the needs of different video scenarios.
[0342] In some embodiments, the temporal attention Attn corresponding to each frame in the first video segment time It can be obtained through the following formula (11):
[0343] Attn time =β·Attn l +(1-β)·Attn g (11)
[0344] Among them, Attn l Attn represents local attention. g This represents global attention, where β is the second weight mentioned above, and 1-β is the first weight mentioned above.
[0345] In some embodiments, β may be a preset value or a user-defined value.
[0346] In this embodiment, the temporal action module can determine the first weight for global attention and the second weight for local attention based on the probability distribution of global and local attention corresponding to video frames in the first video segment. Then, it calculates the weighted result of global and local attention for each frame to obtain the temporal attention. This allows for adaptive adjustment of the weights of the two attention types according to the specific scene of the video segment. When the target video is a long-form video, which may involve complex and varied scene sequences, this method can meet the different needs of different video segments in terms of coordination between local details and global style.
[0347] The video generation method provided in this application can solve the problems of poor global consistency, low generation efficiency, and poor adaptability to complex scenes in the long-term video stylization process in related technologies.
[0348] In some embodiments, the method utilizes an AI model that stylizes video segments to process each video segment of the target video separately, and can decouple temporal attention into two parts: global attention and local attention. The global attention can be obtained based on global style features, which can be shared by different video segments. Furthermore, when performing distributed parallel processing on video segments, the method can fuse the hidden features of video frames with overlapping display times. Based on these features and the other contents mentioned above, the video generation method provided in this application embodiment can adopt an efficient and accurate video stylization generation framework to achieve video stylization.
[0349] Specifically, this method can utilize an AI model that includes a control network module and a timing action module to stylize video clips, ensuring the continuity of actions within the video clips and avoiding abrupt changes in action during the stylization process.
[0350] Furthermore, the method can adaptively select multiple global frames (e.g., the k global frames mentioned above) based on the inter-frame fluctuation distribution of the target video, such as the degree of pixel change and / or scene change of the video frames, and obtain the global style features of each global frame containing global information. These global style features (e.g., KV pairs of global frames) can be used for the calculation of temporal attention in the subsequent temporal action module, thereby effectively ensuring global consistency in the stylization process.
[0351] Furthermore, this method can adaptively adjust the weight distribution of local and global attention for each video segment based on the probability distribution of global and local attention corresponding to different video segments. This effectively balances the coordination between local details and global style for different video segments with different scenes, thereby improving the naturalness and realism of the stylized target video.
[0352] Furthermore, this method can also divide the target video into multiple video segments, and these multiple video segments can be distributed and processed in parallel on different devices. When processing video frames with overlapping display times, different devices can communicate and weighted fuse the hidden features of the video frames with overlapping display times, thereby significantly improving the video generation efficiency and enhancing the robustness of the method in complex scenarios.
[0353] Furthermore, compared to the pre-training approach for video generation models in related technologies, the AI model in this application embodiment does not require training and can be directly used for inference, thereby reducing the cost of video stylization and eliminating the need to collect and label large amounts of high-quality video data, thus reducing workload.
[0354] Based on this, this application provides a video generation method that requires no training, is compatible with multiple styles, has high generation efficiency, and strong global consistency. It is suitable for stylized generation of various long videos and has significant application advantages. This method can be deeply integrated into various fields, providing personalized services and intelligent generation, making users' lives more convenient and efficient.
[0355] The video generation method provided by the embodiments of this application has been described above with reference to various embodiments. Next, the structure of the video generation apparatus and computing device provided by the embodiments of this application will be described with reference to the accompanying drawings.
[0356] This application provides a video generation device 700. Figure 7 For a schematic diagram of the structure of the video generation apparatus 700 as an example, please refer to... Figure 7 The video generation apparatus 700 includes: an acquisition module 701, configured to acquire an inference request, the inference request indicating a target video to be processed and a target style to be processed into the target video; the acquisition module 701 is further configured to acquire a first video segment and a second video segment from the target video; the acquisition module 701 is further configured to acquire global style features of global frames obtained by preprocessing the target video, the global frames being the top k video frames arranged in descending order based on at least one of pixel change degree and scene change degree among multiple video frames of the target video, the global style features indicating the style features of the global frames, and k being a positive integer; an inference module 702, configured to input the first video segment, the global style features, and the information indicating the target style into an AI model for stylizing the video segments to output a first video segment with the target style, and to input the second video segment, the global style features, and the information indicating the target style into the AI model to output a second video segment with the target style; and a generation module 703, configured to obtain a target video with the target style based on the first video segment with the target style and the second video segment with the target style.
[0357] In one possible implementation, the acquisition module 701 is specifically used to: acquire the first k video frames of the target video in descending order based on at least one of pixel change degree and scene change degree, to obtain k global frames, where k is a positive integer; and acquire the global style features of each of the k global frames.
[0358] In one possible implementation, the acquisition module 701 is specifically used to: input the above-mentioned information indicating the target style, the above-mentioned k global frames, and at least one frame sequence obtained based on the third video segment into the AI model to obtain the global style features of each global frame in the k global frames. The global style features of each global frame indicate the style features of the corresponding global frame in the target video and indicate the style features of the video frames in the third video segment after being stylized with the target style. The third video segment is either the first video segment or the second video segment.
[0359] In one possible implementation of the above method, the acquisition module 701 is specifically configured to: acquire a KV pair for each of the k global frames obtained from the preprocessing of the target video, wherein the KV pair indicates the style features of the corresponding global frame.
[0360] In one possible implementation, the inference module 702 includes an AI model, which includes a temporal action module. The temporal action module is configured to: obtain global attention based on global style features and style features of a first video segment stylized with the target style, wherein the style features of the first video segment stylized with the target style are intermediate results obtained by the AI model stylizing each video frame in the first video segment according to the target style; obtain local attention based on the intermediate results corresponding to the first video segment; obtain temporal attention based on global attention and local attention; and output the first video segment with the target style based on the temporal attention.
[0361] In one possible implementation, the aforementioned temporal action module is specifically used to: determine a first weight corresponding to global attention and a second weight corresponding to local attention based on the probability distributions of global attention and local attention; and weight the global attention and local attention based on the first weight and the second weight to obtain temporal attention.
[0362] Figure 7 The video generation apparatus 700 shown corresponds to the video generation method in the above embodiments, therefore Figure 7 The specific implementation of the video generation device 700 and its technical effects can be found in the relevant descriptions of the embodiments of the aforementioned video generation methods, and will not be repeated here.
[0363] This application provides a video generation system, which includes a system deployed as described above. Figure 7 The first device of the video generation apparatus 700 shown above, and, are deployed with the above-described... Figure 7The video generation apparatus 700 shown includes a second device; a first device for inputting a first video segment, global style features, and information indicating the target style into an AI model in the video generation apparatus deployed on the first device to output a first video segment with the target style; a second device for inputting a second video segment, global style features, and information indicating the target style into the AI model in the video generation apparatus deployed on the second device to output a second video segment with the target style; and either the first or second device for obtaining a target video with the target style based on the first video segment with the target style and the second video segment with the target style.
[0364] In one possible implementation, there are video frames with overlapping display times between the first video segment and the second video segment; the first device is configured to send the first hidden layer features of the video frames with overlapping display times, calculated by the AI model in the video generation device deployed on the first device, to the second device; the first device receives the second hidden layer features of the video frames with overlapping display times, calculated by the AI model running on the second device, sent by the second device; the first device performs feature fusion on the first hidden layer features and the second hidden layer features to output the first video segment with the target style.
[0365] This video generation system corresponds to the video generation methods in the above embodiments. Therefore, the specific implementation of the video generation system and its technical effects can be found in the relevant descriptions of the embodiments of the aforementioned video generation methods, and will not be repeated here.
[0366] All of the above modules can be implemented in software or hardware. As an example of a software functional unit, a module may include code running on a computing instance. A computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Furthermore, there may be one or more computing instances. For example, module 701 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.
[0367] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0368] As an example of a hardware functional unit, the module mentioned above may include at least one computing device, such as a server. Alternatively, the module may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0369] The multiple computing devices included in the aforementioned video generation apparatus 700 can be distributed within the same region or in different regions. Similarly, the multiple computing devices included in the aforementioned video generation apparatus 700 can be distributed within the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices included in the aforementioned video generation apparatus 700 can be distributed within the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0370] It should be noted that, in other embodiments, the above-described modules can be used to perform corresponding steps in the video generation method to realize all the functions of the video generation device 700.
[0371] This application also provides a computing device 800. For example... Figure 8 As shown, the computing device 800 includes a bus 802, a processor 804, a memory 806, and a communication interface 809. The processor 804, the memory 806, and the communication interface 809 communicate with each other via the bus 802. The computing device 800 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 800.
[0372] The 802 bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 8 The bus 802 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 802 may include a path for transmitting information between various components of the computing device 800 (e.g., memory 806, processor 804, communication interface 809).
[0373] Processor 804 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0374] Memory 806 may include volatile memory, such as random access memory (RAM). Memory 806 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0375] The memory 806 stores executable program code, which the processor 804 executes to implement the functions of the aforementioned acquisition module 701, inference module 702, and generation module 703, thereby realizing the video generation method. In other words, the memory 806 stores instructions for executing the video generation method.
[0376] The communication interface 809 uses transceiver modules, such as, but not limited to, network interface cards and transceivers, to enable communication between the computing device 800 and other devices or communication networks.
[0377] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0378] like Figure 9 As shown, the computing device cluster includes at least one computing device 900. The memory 906 of one or more computing devices 900 in the computing device cluster may store the same instructions for executing the video generation method.
[0379] The computing device 900 includes a bus 902, a processor 904, a memory 906, and a communication interface 908. The processor 904, the memory 906, and the communication interface 908 communicate with each other via the bus 902.
[0380] In some possible implementations, the memory 906 of one or more computing devices 900 in the computing device cluster may also store partial instructions for executing the video generation method. In other words, a combination of one or more computing devices 900 can jointly execute the instructions for executing the video generation method.
[0381] It should be noted that the memory 906 in different computing devices 900 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the video generation device 700. That is, the instructions stored in the memory 906 of different computing devices 900 can implement the functions of one or more modules among the acquisition module 701, the inference module 702, and the generation module 703.
[0382] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 10 One possible implementation is shown. For example... Figure 10 As shown, the two computing devices 1000A and 1000B are connected via a network. Specifically, they are connected to the network through the communication interface 1008 in each computing device.
[0383] The computing device 1000A includes a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. The processor 1004, the memory 1006, and the communication interface 1008 communicate with each other via the bus 1002.
[0384] The computing device 1000B includes a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. The processor 1004, the memory 1006, and the communication interface 1008 communicate with each other via the bus 1002.
[0385] In this type of possible implementation, computing device 1000A may be, for example, the first device described above, and computing device 1000B may be, for example, the second device described above, such as... Figure 10 As shown, the memory 1006 in computing device 1000A may store instructions for executing the functions of the acquisition module 701 and the inference module 702. Simultaneously, the memory 1006 in computing device 1000B may store instructions for executing the functions of the acquisition module 701, the inference module 702, and the generation module 703.
[0386] For example, the acquisition module 701 in computing device 1000A can acquire the aforementioned inference request, the aforementioned first video segment, the global style features of the aforementioned global frame, and the aforementioned information indicating the target style. The inference module 702 in computing device 1000A can input the first video segment, the global style features, and the information indicating the target style into the AI model in the inference module 702 to output a first video segment with the target style. Computing device 1000A can send the first video segment with the target style to computing device 1000B via a network.
[0387] The acquisition module 701 in computing device 1000B can acquire the aforementioned inference request, the aforementioned second video segment, the aforementioned global style features, and the aforementioned information indicating the target style. The inference module 702 in computing device 1000B can input the second video segment, the global style features, and the information indicating the target style into the AI model in the inference module 702 to output a second video segment with the target style. Subsequently, the generation module 703 in computing device 1000B can obtain a target video with the target style based on the second video segment with the target style and the first video segment with the target style sent by computing device 1000A.
[0388] In some embodiments, computing device 1000A may further include a generation module 703. After the inference module 702 in computing device 1000B obtains a second video segment with the target style, computing device 1000B can send the second video segment with the target style to computing device 1000A via a network. Then, the generation module 703 in computing device 1000A can obtain a target video with the target style based on the first video segment with the target style and the second video segment with the target style sent by computing device 1000B.
[0389] In some embodiments, there are video frames with overlapping display times between the first video segment and the second video segment; the computing device 1000A can send the first hidden layer features of the video frames with overlapping display times calculated by the AI model included in the inference module 702 to the second device via the network; the computing device 1000A can also receive the second hidden layer features of the video frames with overlapping display times calculated by the AI model included in the inference module 702 in the computing device 1000B sent by the computing device 1000B; the computing device 1000A can perform feature fusion on the first hidden layer features and the second hidden layer features to output the first video segment with the target style.
[0390] It should be understood that Figure 10 The functions of computing device 1000A shown can also be performed by multiple computing devices 1000. Similarly, the functions of computing device 1000B can also be performed by multiple computing devices 1000.
[0391] This application also provides another computing device cluster. The connection relationships between the computing devices in this computing device cluster can be similarly referred to... Figure 9 and Figure 10 The connection method of the computing device cluster is different in that the memory 1006 of one or more computing devices 1000 in the computing device cluster can store the same instructions for executing the video generation method.
[0392] In some possible implementations, the memory 1006 of one or more computing devices 1000 in the computing device cluster may also store partial instructions for executing the video generation method. In other words, a combination of one or more computing devices 1000 can jointly execute the instructions for executing the video generation method.
[0393] It should be noted that the memory 1006 in different computing devices 1000 within the computing device cluster can store different instructions for executing some functions of the video generation apparatus 700. That is, the instructions stored in the memory 1006 of different computing devices 1000 can implement the functions of one or more modules in the video generation apparatus 700.
[0394] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform the video generation method described above.
[0395] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct a computing device to perform the video generation method described in the above embodiments.
[0396] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of this application.
Claims
1. A video generation method, characterized in that, The method includes: Obtain an inference request, the inference request indicating the target video to be processed and the target style to which the target video is to be processed; Obtain the first and second video segments from the target video; Obtain the global style features of the global frames obtained by preprocessing the target video. The global frames are the first k video frames in descending order of at least one of the pixel change degree and the scene change degree among multiple video frames of the target video. The global style features indicate the style features of the global frames, where k is a positive integer. The first video segment, the global style feature, and information indicating the target style are input into an artificial intelligence (AI) model for stylizing the video segment to output a first video segment with the target style; and the second video segment, the global style feature, and information indicating the target style are input into the AI model to output a second video segment with the target style. Based on a first video segment having the target style and a second video segment having the target style, a target video having the target style is obtained.
2. The method according to claim 1, characterized in that, The step of obtaining the global style features of the global frame obtained from the preprocessing of the target video includes: The information indicating the target style, the k global frames, and at least one frame sequence obtained based on the third video segment are input into the AI model to obtain the global style features of each of the k global frames. The global style features of each global frame indicate the style features of the corresponding global frame in the target video and indicate the style features of the video frames in the third video segment after being stylized with the target style. The third video segment is either the first video segment or the second video segment.
3. The method according to claim 1 or 2, wherein obtaining the global style features of the global frame obtained by preprocessing the target video includes: Obtain key-value (KV) pairs for each of the k global frames obtained from the preprocessing of the target video, wherein the KV pairs indicate the style features of the corresponding global frames.
4. The method according to any one of claims 1 to 3, characterized in that, The AI model includes a temporal action module. The step of inputting the first video segment, the global style features, and information indicating the target style into the AI model for stylizing the video segment, to output a first video segment with the target style, includes: The timing action module obtains global attention based on the global style features and the style features of the first video segment after being stylized with the target style. The style features of the first video segment after being stylized with the target style are the intermediate results obtained by the AI model stylizing each video frame in the first video segment according to the target style. The timing action module obtains local attention based on the intermediate results corresponding to the first video segment; The temporal action module obtains temporal attention based on the global attention and the local attention; The temporal action module outputs a first video clip with the target style based on the temporal attention.
5. The method according to claim 4, characterized in that, The temporal action module obtains temporal attention based on the global attention and the local attention, including: The timing action module determines the first weight corresponding to the global attention and the second weight corresponding to the local attention based on the probability distribution of the global attention and the local attention; Based on the first weight and the second weight, the global attention and the local attention are weighted to obtain temporal attention.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: There are video frames with overlapping display times between the first video segment and the second video segment; an AI model of the first video segment with the target style is run on a first device; and an AI model of the second video segment with the target style is run on a second device. The first device sends the first hidden layer features of the video frames with overlapping display times, calculated by the running AI model, to the second device; The first device receives the second hidden layer features of the video frames with overlapping display times, which are calculated by the AI model running on the second device and sent by the second device. The first device performs feature fusion on the first hidden layer features and the second hidden layer features to output a first video segment with the target style.
7. A video generation apparatus, characterized in that, The device includes: The acquisition module is used to acquire an inference request, wherein the inference request indicates the target video to be processed and the target style to which the target video is to be processed; The acquisition module is further configured to acquire a first video segment and a second video segment from the target video; The acquisition module is further configured to acquire global style features of global frames obtained by preprocessing the target video. The global frames are the first k video frames in descending order of at least one of pixel change degree and scene change degree among multiple video frames of the target video. The global style features indicate the style features of the global frames, where k is a positive integer. The inference module is used to input the first video segment, the global style features, and information indicating the target style into an artificial intelligence (AI) model for stylizing the video segment, so as to output a first video segment with the target style; and based on the inference request, to input the second video segment, the global style features, and information indicating the target style into the AI model, so as to output a second video segment with the target style. A generation module is used to obtain a target video with the target style based on a first video segment with the target style and a second video segment with the target style.
8. The apparatus according to claim 7, characterized in that, The acquisition module is specifically used for: The information indicating the target style, the k global frames, and at least one frame sequence obtained based on the third video segment are input into the AI model to obtain the global style features of each of the k global frames. The global style features of each global frame indicate the style features of the corresponding global frame in the target video and indicate the style features of the video frames in the third video segment after being stylized with the target style. The third video segment is either the first video segment or the second video segment.
9. The apparatus according to claim 7 or 8, wherein the acquisition module is specifically used for: Obtain key-value (KV) pairs for each of the k global frames obtained from the preprocessing of the target video, wherein the KV pairs indicate the style features of the corresponding global frames.
10. The apparatus according to any one of claims 7 to 9, characterized in that, The reasoning module includes the AI model, and the AI model includes a time-series action module; The timing action module is used for: Based on the global style features and the style features of the first video segment after being stylized with the target style, global attention is obtained. The style features of the first video segment after being stylized with the target style are the intermediate results obtained by the AI model stylizing each video frame in the first video segment according to the target style. Based on the intermediate results corresponding to the first video segment, local attention is obtained; Based on the global attention and the local attention, temporal attention is obtained; Based on the temporal attention, a first video segment with the target style is output.
11. The apparatus according to claim 10, characterized in that, The timing action module is specifically used for: Based on the probability distributions of the global attention and the local attention, determine the first weight corresponding to the global attention and the second weight corresponding to the local attention; Based on the first weight and the second weight, the global attention and the local attention are weighted to obtain temporal attention.
12. A video generation system, characterized in that, The system includes a first device having a video generation apparatus as described in any one of claims 7 to 11, and a second device having a video generation apparatus as described in any one of claims 7 to 11; The first device is configured to input the first video segment, the global style features, and information indicating the target style into the AI model in the video generation device deployed by the first device, so as to output a first video segment having the target style; The second device is configured to input the second video segment, the global style features, and information indicating the target style into the AI model in the video generation apparatus deployed by the second device, so as to output a second video segment having the target style; The first device or the second device is used to obtain a target video with the target style based on a first video segment with the target style and a second video segment with the target style.
13. The system according to claim 12, characterized in that, There are video frames with overlapping display times between the first video segment and the second video segment; The first device is configured to send the first hidden layer features of the video frames with overlapping display times, calculated by the AI model in the video generation device deployed on the first device, to the second device; The first device receives the second hidden layer features of the video frames with overlapping display times, which are calculated by the AI model running on the second device and sent by the second device. The first device performs feature fusion on the first hidden layer features and the second hidden layer features to output a first video segment with the target style.
14. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1 to 6.
15. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster performs the method as described in any one of claims 1 to 6.
16. A computer-readable storage medium, characterized in that, It includes computer program instructions, which, when executed by a cluster of computing devices, perform the method as described in any one of claims 1 to 6.