Video generation method based on diffusion model optical flow trajectory control

By introducing optical flow trajectory control and dynamic optical flow attention mechanism based on diffusion model in the video generation method, the problems of insufficient motion information extraction and poor video coherence in the prior art are solved, and personalized and refined video generation is realized, and the quality and stability of video generation are improved.

CN120201260APending Publication Date: 2025-06-24GUANGDONG BOHUA UHD INNOVATION CENT CO LTD

Patent Information

Application Number
CN202510402278.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing few-sample video generation methods have shortcomings in motion information extraction and video coherence, and cannot effectively capture complex motion patterns and personalized video generation requirements.

Method used

The optical flow trajectory control method based on the diffusion model is adopted to extract the optical flow characteristic map of the video frame through the optical flow estimation network, calculate the pixel-level motion trajectory, and optimize it with the dynamic optical flow attention mechanism to accurately extract and model the motion information in the video.

Benefits of technology

It realizes personalized and refined video generation, improves the action consistency and detailed performance of video generation, and can quickly adapt to different personalized needs in a small sample data environment to generate high-quality and stable videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201260A_ABST
    Figure CN120201260A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method based on diffusion model optical flow trajectory control. The video generation method comprises the following steps: S1, processing a training video and modeling an optical flow trajectory; s2, optical flow trajectory calculation and dynamic attention optimization; s3, carrying out diffusion modeling; and S4, model reasoning and implicit vector reconstruction. Through three core technologies of optical flow trajectory modeling, dynamic attention optimization and implicit vector reconstruction, the precision of motion information capture and the detail quality of video generation are improved, personalized style control is realized, and compared with a traditional method, the method has remarkable advantages in motion accuracy, video coherence and personalized customization ability, and the method is suitable for popularization and application. The problems that in the prior art, the precision in few-sample video generation is insufficient, and motion information in a training video cannot be effectively extracted are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and particularly to a video generation method based on diffusion model optical flow trajectory control. Background Art

[0002] Video generation is a cutting-edge technology in the fields of computer vision and artificial intelligence, aiming to automatically generate continuous and realistic video sequences through algorithms, and has shown broad application prospects in multiple scenarios such as game development and cultural creation. In recent years, large text-to-image models represented by Stable Diffusion have achieved remarkable results in the field of image generation and have gradually been introduced into video generation tasks. However, current models (such as the Sora model, Conch Video, etc.) usually consume a large amount of computing resources, which to a certain extent limits the application of ordinary scientific researchers and self-media creators.

[0003] Currently, video generation methods mainly rely on diffusion models. Diffusion models are a type of generative model, and their core idea is to simulate the gradual generation process of data from noise to clear samples. Specifically, the model first gradually adds noise to the data through the forward process, and then combines neural networks to learn the noise distribution through the reverse process, ultimately converting random noise into the target data. In this process, each step is optimized by combining text semantic information with the result of the previous step, thereby generating images that highly match the text description. Existing few-shot video generation methods include patents: CN118053090A, CN119094788A, and CN118984414B. For example, in patent CN118053090A: By inserting convolutional layers and temporal attention layers, it learns the temporal information and correlation between video frames. At the same time, this technology uses the generated front and back video frames for interpolation processing to generate additional frames, achieving ultra-high-resolution video generation. However, this method still has the following deficiencies: 1. Loss of motion details: This method mainly relies on the key frames of the training video for training, which may lead to insufficient extraction of temporal information, thereby losing motion details and affecting the dynamic performance of the video. 2. Limited video coherence: In the video generation stage, this method first generates key frames and then generates transitional frames through interpolation processing. Essentially, this method still relies on interpolation to achieve smooth transitions between frames and cannot directly generate a coherent video sequence from an overall level. 3. Lack of reference image-driven ability: This method cannot generate the target video based on a specific reference image, limiting its application in video synthesis and generation under specific conditions. In patent CN118984414B: The problems of this patent are similar to those of the above patent. Their common point is that they cannot effectively fit the motion information in the training video and can only generate corresponding video actions based on text descriptions for the generation of personalized videos. Patent CN119094788A: Proposed a video action generation method based on an enhanced video diffusion model. This method uses a reference motion video collection and trains the enhanced video diffusion model with multiple source videos, making an innovation from the perspective of motion information fitting. However, this method still has the following limitations: 1. Limited expression of motion information: This method enhances motion information by constructing different attention modules, but these attention mechanisms are essentially still optimized fittings of the input data and do not introduce additional motion information representation methods (such as optical flow, motion trajectories, etc.), which may lead to limited modeling ability for complex motion patterns. 2. Lack of video generation ability based on reference images: This method cannot directly generate the corresponding motion video according to a single-frame reference image, limiting its application in motion synthesis and personalized video generation under specific conditions. In summary, these methods still have some problems and difficulties in solving them: interference from various background objects in video training data, lack of effective feature extraction means, difficulty in accurately capturing motion information in videos, and inability to generate corresponding action videos according to the expected content.

[0004] The significance of solving the above problems and deficiencies is as follows: It solves the problem of insufficient accuracy in few-shot video generation in the prior art and provides a new technical path for personalized customized video generation. Summary of the Invention

[0005] The present invention provides a video generation method based on diffusion model optical flow trajectory control, which solves the problems of insufficient accuracy in few-shot video generation in the prior art and the inability to effectively extract motion information from training videos.

[0006] The technical solution of the present invention is as follows: The video generation method based on diffusion model optical flow trajectory control of the present invention includes the following steps: S1. Process the training video and model the optical flow trajectory; S2. Calculate the optical flow trajectory and optimize the dynamic attention; S3. Diffusion modeling; S4. Model inference and latent vector reconstruction.

[0007] Optionally, in the above video generation method based on diffusion model optical flow trajectory control, in step S1, sample the training video and extract a continuous sequence of video frames; use the optical flow estimation network to process these video frames, thereby generating a set of optical flow feature maps with a length of ; use the optical flow model to extract the optical flow feature maps of the video frames and calculate the pixel-level motion trajectories to form a complete set of optical flow trajectories; subsequently, input the optical flow trajectory features and the video frame images into the diffusion model to obtain motion and visual features, providing high-quality data for subsequent video generation. ; use the optical flow density clustering algorithm to cluster the trajectories, extract the main features of the motion information,

[0008] Optionally, in the above video generation method based on diffusion model optical flow trajectory control, in step S2, process two adjacent frames of the optical flow feature map to obtain the displacement vector of each pixel and construct the optical flow trajectory , ; finally use , the image latent vector features , the text features to perform dynamic optical flow attention calculation to obtain ; repeat the above steps u times to obtain the output result .

[0009] Optionally, in the above video generation method based on diffusion model optical flow trajectory control, in step S3, model based on the diffusion principle. The diffusion model adds m times of noise to the in step S1 to obtain ; through the diffusion model Unet neural network The calculation of the secondary loop yields , Subtract to obtain the finally generated video frame, and together with constitute the modeling of the loss function to optimize the diffusion model.

[0010] Optionally, in the above video generation method based on the diffusion model optical flow trajectory control, in step S4, given an image , encode it using the variational autoencoder VAE and convert it into a latent vector ; Subsequently, randomly initialize number of , is the video frame length; Replace the first , and add it to the subsequent number of ; Put it into the diffusion model for denoising inference to obtain the final generated video.

[0011] According to the technical solution of the present invention, the beneficial effects are as follows: 1. Personalized and refined video generation. The method of the present invention effectively overcomes the limitations of traditional large-scale video generation models in personalized development and motion information control. Existing diffusion models based on text-to-image (T2I) often have difficulty accurately capturing complex motion patterns when used for video generation, resulting in deviations between the actions of the generated videos and the original training data. The present invention introduces optical flow trajectory control and dynamic optical flow attention mechanism to accurately extract and model the motion information in video training data, enabling the generated videos to more accurately fit the target action trajectory and avoiding problems such as blurred or distorted actions caused by insufficient fitting in traditional methods. In addition, the network structure of the present invention supports few-shot learning and can quickly adapt to different personalized needs under limited training data, improving the quality and stability of the generated videos.

[0012] 2. Efficient motion information modeling and enhancement. Existing video generation methods mainly rely on text descriptions for motion modeling, but text often fails to provide accurate motion information, resulting in videos lacking physical consistency and poor action coherence. The present invention uses optical flow trajectory calculation combined with dynamic attention mechanism to extract pixel-level motion information from video training data, enabling the diffusion model to more accurately understand and learn the motion patterns in video frames.

[0013] 3. Customized video generation based on reference images. In existing few-shot video generation methods, most models mainly rely on text descriptions for action inference and adopt a random sampling strategy for generation when the training data is limited, resulting in unstable video styles and difficulty in controlling the generated content. Through the latent vector optimization strategy, the present invention encodes a specified reference image as a latent vector during the diffusion model inference process and introduces this style information during the denoising process, enabling the generated video to not only conform to the action content described in the text but also maintain a high style consistency with the target image, significantly enhancing the personalization and customization capabilities of video generation.

[0014] To better understand and illustrate the concept, working principle, and invention effect of the present invention, the following will, with reference to the accompanying drawings and through specific embodiments, provide a detailed description of the present invention as follows: BRIEF DESCRIPTION OF THE DRAWINGS

[0015] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art.

[0016] Figure 1 is a flowchart of the video generation method based on diffusion model optical flow trajectory control of the present invention; Figure 2 is a schematic diagram of the network model training diagram of the present invention; Figure 3 is a schematic diagram of the optical flow trajectory calculation of the present invention; Figure 4 is a schematic diagram of the dynamic optical flow attention module of the present invention; Figure 5 is a schematic diagram of reconstructing the diffusion model latent vector of the present invention; Figure 6 is a comparison effect diagram of the existing few-shot video generation method (LAMP) and other video generation methods for the first prompt; Figure 7 is a comparison effect diagram of the existing few-shot video generation method (LAMP) and other video generation methods for the second prompt. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] To make the objectives, technical methods, and advantages of the present invention clearer, the following will, with reference to the accompanying drawings and specific examples, further elaborate on the present invention. These examples are merely illustrative and not restrictive of the present invention.

[0018] By combining a text-to-image generation model (text-to-video model) with optical flow trajectory calculation and a dynamic optical flow attention mechanism, the present invention enables the diffusion model network to accurately extract and capture the features of training videos based on specific motion patterns, thereby significantly improving the accuracy and personalization of video generation and solving the problem in existing few-shot video generation techniques of being unable to effectively extract motion information from training videos. Specifically, the present invention first uses a high-precision optical flow model to extract the optical flow feature map of the training video, and then calculates the motion trajectory of each pixel based on the adjacent frame optical flow feature maps to form a complete description of the motion information. This optical flow trajectory calculation process effectively captures the motion details in the training video and avoids the insufficient capture of motion information by traditional methods in complex scenarios. Then, by introducing a dynamic optical flow attention mechanism and combining convolution and attention calculations, the present invention constructs a deep association between the latent features and the optical flow motion trajectory. This mechanism enables the model to efficiently learn specific action types and optimize the accuracy of action fitting.

[0019] The principle of the present invention is: Its core lies in the innovative integration of the optical flow information and image features of the training video, and through the dynamic optical flow attention mechanism in the diffusion model neural network, it realizes the accurate modeling and efficient integration of motion information. The method of the present invention not only breaks through the limitations of existing few-shot video generation techniques in motion information extraction, but also optimizes the motion consistency and detail performance of the diffusion model in the video generation process. Through a latent vector optimization strategy based on a reference image, a customized video generation network framework with high-precision motion modeling capabilities is constructed, breaking through the limitations of traditional methods in inaccurate action capture, lack of consistency and personalization in generated videos in a few-shot data environment.

[0020] The video generation method based on diffusion model optical flow trajectory control of the present invention includes the following steps: S1. Process the training video and model the optical flow trajectory. First, sample the training video and extract a continuous sequence of video frames; subsequently, use an optical flow estimation network (such as the RAFT optical flow model) to process these video frames to generate a set of optical flow feature maps of length . Use the optical flow model to extract the optical flow feature maps of the video frames and calculate the pixel-level motion trajectories to form a complete set of optical flow trajectories. This process accurately captures the motion information and is particularly suitable for complex backgrounds and multi-object scenarios; subsequently, input the optical flow trajectory features and the video frame images into the diffusion model to obtain motion and visual features, providing high-quality data for subsequent video generation.

[0021] S2. Calculate the optical flow trajectory and optimize the dynamic attention. Process two adjacent frames of the optical flow feature map to obtain the displacement vector of each pixel and construct the optical flow trajectory , Use the optical flow density clustering algorithm to cluster the trajectories and extract the main features of the motion information. . Finally, use , the image latent vector features , the text features , to calculate the dynamic optical flow attention and obtain ; Repeat the above steps u times to get the output result . The motion information modeling adopts the dynamic optical flow attention mechanism, combines convolution and attention calculation, and realizes the efficient fusion of optical flow trajectory features and image features. This mechanism enhances the model's understanding of motion patterns, improves the action consistency and detail performance of video generation, and can more accurately restore specific action trajectories compared with traditional methods.

[0022] S3. Diffusion modeling. Model based on the diffusion principle. The diffusion model adds m times of noise to in step S1 to obtain . After u times of loop calculation through the Unet neural network of the diffusion model, obtain , , Subtract to get the finally generated video frame, and form a loss function with to optimize the modeling of the diffusion model.

[0023] S4. Model inference and latent vector reconstruction. After training the diffusion model, it is the model inference, i.e., video generation, to optimize the inference process of the model and reconstruct the initialization of the latent vector. Given an image , encode it using the variational autoencoder VAE to convert it into a latent vector . Then randomly initialize , is the length of the video frame. Replace the first , and add it to the subsequent ; Put it into the diffusion model for denoising inference to get the final generated video. In the inference optimization of the diffusion model, introduce the latent vector optimization strategy based on the reference image. During the inference process of the diffusion model, use the style image features as additional conditions, so that the diffusion model can generate videos of specific actions based on text descriptions and image styles. This strategy improves the personalized customization ability of the generated videos, ensures unified style and smooth motion, and enhances the accurate reproduction of the target actions.

[0024] Through three core technologies: optical flow trajectory modeling, dynamic attention optimization, and latent vector reconstruction, the present invention improves the accuracy of motion information capture, the detail quality of video generation, and realizes personalized style control, having significant advantages over traditional methods in terms of motion accuracy, video coherence, and personalized customization ability. By using optical flow trajectory information and dynamic optical flow attention mechanism, the motion information features in the training video are effectively extracted. At the same time, the latent vector is inferred through the given image reconstruction to achieve personalized video generation based on specific action patterns. This method solves the problems that existing fine-tuning models cannot effectively extract the motion information of the training video and can only randomly generate videos according to text descriptions. By combining optical flow trajectory calculation with dynamic optical flow attention mechanism, the present invention optimizes the motion information extraction and action fitting in the video generation process, solves the problem of insufficient accuracy in few-shot video generation in the prior art, and provides a new technical path for personalized customized video generation. The specific implementation steps are as follows: S1: Process the training video and optical flow trajectory modeling.

[0025] First, sample consecutive video frames from the training video, and then adjust the images to a unified size, such as , to obtain the adjusted images . Then put the images into the optical flow estimation network (RAFT model) to obtain the optical flow feature map , with a length of n - 1. Pass the adjusted images through the variational encoder for compression, and then put them into the diffusion model Unet neural network to obtain the feature images , where the variational encoder is a neural network model that has been trained on a large public dataset. These VAEs can be trained by oneself or be existing public variational encoders, such as the trained variational autoencoder AutoencderKL provided by StabilityAI in the Stable Diffusion project. In particular, the current read from the video file is dimensionally concatenated. Design the corresponding prompt word features , and jointly input them into the diffusion model for information aggregation.

[0026] S2: Calculate the optical flow trajectory and dynamic attention optimization.

[0027] First, according to the optical flow feature map , with a length of , as shown in Figure 3 , obtain the displacement vector of each pixel in each optical flow feature map , and then calculate the corresponding position subscript of this pixel in the remaining optical flow feature maps through the formula: Calculate the corresponding subscript of each pixel, and form an independent optical flow trajectory with the subscripts of each pixel. The subscripts of each pixel can form a set of optical flow trajectories. Among them and are in the format of image features. According to the optical flow density function, different optical flow trajectories are aggregated so that similar motion information is aggregated (i.e., Figure 3 the trajectory aggregation in ), better highlighting the motion information in the video frame. The processed optical flow trajectory features can be expressed as

[0028] S3: Diffusion Modeling Video generation requires data fitting through a diffusion model, and the neural network algorithm model needs to have strong data generation capabilities. The algorithm model based on the stable diffusion principle has been proven effective in mainstream algorithms such as Stable Diffusion. The present invention expands the input dimension of its data based on the Stable Diffusion model framework. Figure 2 is a schematic diagram of the main training process of the network diffusion model network. First, the training video will be preliminarily processed, such as Figure 2 the "training video" and "reading video" steps in Figure 2 The "Variational Autoencoder VAE" in will process it and convert it into a latent vector, image features Figure 2 . At the same time, text information such as the "prompt word" part in Figure 2 The "optical flow estimation model" in will be combined with the optical flow trajectory features obtained in step S2 Figure 4 and perform fitting diffusion modeling using the optical flow dynamic attention module (as shown in Figure 4 ). Specifically, as shown in Figure 4 the "dynamic convolution module" in extracts features from the "set of optical flow trajectories" to obtain the aggregated optical flow trajectory features Figure 4 . The aggregated optical flow trajectory features are first dot-product matrix multiplied with the "prompt word features" in to obtain the optical flow trajectory information weighted with prompt words Figure 4 . Subsequently, it performs cross-attention calculation with the "image features" in to better fuse these three different feature information. The final output image information and the input image information

[0029] S4: Model Inference and Latent Vector Reconstruction.

[0030] Video generation involves denoising inference of a diffusion model. Given a noise that follows a Gaussian normal distribution, the diffusion model denoises the noise by combining text description information and restores it to a high-definition image. Similarly for video generation, the difference lies in that a noise needs to be denoised to obtain a set of continuous video frames. Step S4 lies in the process of optimizing and reconstructing the latent vector noise . Given an initial image , use a variational autoencoder (VAE) to compress the image into the latent vector space and add a certain amount of noise to it (add noise), that is . Then, randomly sample to construct a latent vector: randomly sample vectors that follow a Gaussian normal distribution , that is . Add to in a decreasing order of subscripts, that is . The schematic diagram of reconstructing the latent vector of the diffusion model is specifically as shown in Figure 5 . Form the optimized reconstructed latent vector of the diffusion model, and then put it into the diffusion model for denoising operation to generate the final personalized action video.

[0031] The present invention designs a video generation algorithm for sampling the optical flow trajectory of a diffusion model. This algorithm uses an optical flow model to extract the motion information in the video and uses the correlation characteristics of the dynamic optical flow attention mechanism and the reconstruction model inference method to realize a multi-style video generation algorithm for specific motion trajectories. The present invention extracts the motion features of a given action video through an optical flow model, calculates the motion pixel trajectory to obtain accurate motion information, and combines the dynamic optical flow attention mechanism to optimize the fitting of motion video data of specific categories. In the subsequent video generation process, through the latent vector inferred by the reconstruction model, a corresponding motion video can be generated according to the specified similar pictures, thus effectively solving the deficiencies of existing few-shot video generation methods in terms of action fitting accuracy and image customization generation. Different from the traditional way of generating videos based on text descriptions, the present invention successfully realizes the function of generating personalized motion videos based on image styles and contents through the process of image encoding and latent vector reconstruction, significantly improving the personalization and customization capabilities of video generation.

[0032] The present invention first uses a high-precision optical flow model (such as RAFT) to extract the optical flow feature map of the video, and forms a complete description of the motion information by calculating the motion trajectories of each pixel. These optical flow trajectories and image features are fused through a dynamic optical flow attention mechanism, so as to accurately capture the motion information in the video and optimize the generation effect. This fusion method makes the generated video more accurate in action details, and can effectively reduce the problems of motion blur and detail loss compared with traditional methods. In addition, the present invention optimizes the reconstruction process of the latent vector by referring to the reference image, and realizes more efficient video generation in the few-shot scenario. Through the innovative optical flow trajectory calculation and information fusion method, combined with the dynamic optical flow attention mechanism and the diffusion model, the present invention solves the limitations in the prior art that cannot accurately extract the video motion information and generate personalized customized motion videos, and significantly improves the quality and flexibility of video generation.

[0033] Figure 6 Compared with Figure 7 is a comparison effect diagram of the existing few-shot video generation method (LAMP) and other video generation methods. The present invention can not only efficiently capture the motion data in the training video, but also realizes the function of generating specific action videos according to text descriptions and reference pictures through the latent vector optimization strategy based on the reference image. This innovation breaks through the limitations of traditional methods and successfully realizes personalized customized output based on the given image style and content, making the generated video not only accurate in action, but also rich in style, meeting higher-level customization requirements.

[0034] The above description is the best embodiment according to the concept and working principle of the invention. The above embodiments should not be construed as limiting the protection scope of the present claims. Combinations of other implementation manners and implementation modes according to the concept of the present invention all belong to the protection scope of the present invention.

Claims

1. A video generation method based on diffusion model optical flow trajectory control, characterized in that: The following steps are involved: S1. Process training videos and optical flow trajectory modeling; S2. Calculate optical flow trajectory and dynamic attention optimization; S3. Diffusion modeling; S4. Model reasoning and latent vector reconstruction.

2. The video generation method based on diffusion model optical flow trajectory control according to claim 1 is characterized in that: In step S1, the training video is sampled to extract a continuous segment Frame video sequence; these video frames are processed using the optical flow estimation network to generate a set of length Optical flow feature map ; The optical flow model is used to extract the optical flow feature map of the video frame, and the pixel-level motion trajectory is calculated to form a complete set of optical flow trajectories; then, the optical flow trajectory features and the video frame image are input into the diffusion model to obtain motion and visual features.

3. The video generation method based on diffusion model optical flow trajectory control according to claim 1 is characterized in that: In step S2, two adjacent frames of the optical flow feature map are processed to obtain the displacement vector of each pixel and construct the optical flow trajectory. , cluster the optical flow trajectory using the optical flow density clustering algorithm to extract the main features of the motion information ; Finally use , image latent vector features , text features Perform dynamic optical flow attention calculation and get ; Repeat the above steps u times to get the output result .

4. The video generation method based on diffusion model optical flow trajectory control according to claim 1 is characterized in that: In step S3, a diffusion model is constructed based on the diffusion principle. The diffusion model is a model of the diffusion principle in step S1. Add m noises to get ; After diffusion model Unet neural network The calculation cycle is repeated, and we get , minus Get the final generated video frame and compare it with Modeling of the diffusion model that constitutes the optimization of the loss function.

5. The video generation method based on diffusion model optical flow trajectory control according to claim 1 is characterized in that: In step S4, given an image , which is encoded using a variational encoder and converted into a latent vector ; Then randomly initialize indivual , is the video frame length; replace The first one , and added to the subsequent indivual middle; The video is put into the diffusion model for denoising reasoning to obtain the final generated video.

Citation Information

Patent Citations

  • Generating video using potential diffusion model

    CN118053090A

  • Video generation method, device and equipment based on diffusion model

    CN118984414B

  • Motion video generation method based on enhanced video diffusion model

    CN119094788A

Cited By

  • Personalized video generation model training and reasoning method based on spatio-temporal representation alignment

    CN121074561A

  • Optical flow estimation method and system

    CN121120683A

  • Video generation method and related equipment

    CN121284363A

  • Video prediction method and system based on dynamic diffusion model

    CN121661573A