Video generation method based on high and low frequency feature fusion
By fusing high and low frequency features into the video generation algorithm and injecting the diffusion transformer model, the problem of object form drift and discontinuity in video generation is solved, and video generation with high consistency and realism is achieved, reducing computing resource consumption.
Patent Information
- Application Number
- CN202510169279.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-27
AI Technical Summary
In the generation of existing video generation algorithms, the object shapes between the generated video frames are prone to drift or discontinuity, resulting in poor consistency, especially in the tracking of characters and objects.
A video generation method based on high and low frequency feature fusion is adopted. The overall structure and object details of the image are extracted by the low frequency feature extractor and the high frequency feature extractor, and injected into the diffusion converter model. By fusion of features from the cross attention layer, videos with high object consistency are generated.
It significantly improves the realism and coherence of the generated video, ensures that the facial features and overall details of the characters are consistent, reduces computing resources and training costs, and improves the practicality and deployment efficiency of the model.
Smart Images

Figure CN120220012A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video processing, and more specifically, to a video generation method based on the fusion of high and low frequency features. Background Art
[0002] Currently, most mainstream video generation algorithms mostly input text features and video noise into the latent space for learning, and perform multiple samplings in the latent space during inference to finally generate a video. However, this method has the following problems: in the video generation task of long time series, the object shapes between the generated video frames are prone to drift or discontinuity, resulting in poor consistency, especially in the aspects of human and object tracking, the problems are particularly obvious. Summary of the Invention
[0003] In response to this problem, this solution proposes a construction method for the wooden structure of ancient architecture.
[0004] A video generation method based on the fusion of high and low frequency features, comprising the following steps:
[0005] Low-frequency feature extraction: Extract low-frequency information from the reference image through a global feature extractor. The low-frequency information includes the overall structure, contour, and core key point map of the image; use traditional visual processing techniques or deep learning algorithms to process the reference image to obtain low-frequency features, where the traditional visual processing techniques include edge detection based on pixel operations, image smoothing processing, and morphological operations, and the deep learning methods include object detection, image recognition, and image segmentation; use a video variational autoencoder (VAE) to extract the latent space representation of the low-frequency features, and splice it with the noise features of the video and then input it into the diffusion transformer (DiT) model.
[0006] High-frequency feature extraction: Use a pre-trained model to extract features from the reference image to obtain high-frequency features containing object detail information; perform feature learning on the high-frequency features through a multi-layer perceptron (MLP) to obtain optimized high-frequency features; in the application scenario of maintaining human consistency, use a face recognition model to extract face features, and fuse the face features and object features through a fusion module to form the final high-frequency features; add a cross-attention layer between the attention layer and the fully connected layer of the diffusion transformer (DiT) model, and inject the final high-frequency features into this layer. The cross-attention layer only trains the newly added key-value (KV) matrix, and keeps the original model parameters of the DiT unchanged.
[0007] Model Training and Inference: During the model training process, a random frame is selected from the input video as a reference image, and the high-frequency features and low-frequency features of this frame are extracted and injected into the specified positions of the model respectively; a diffusion model is used for video generation, and the model training process is optimized by adding noise and denoising methods; since both the high-frequency features and low-frequency features are obtained from the previous preprocessing and pre-trained model, only the key-value (KV) matrix parameters in the cross-attention layer are updated during the training process, resulting in a significant reduction in the number of model parameters compared to the original DiT model, thus reducing the consumption of training resources; after training is completed, it can be directly used for video generation without additional fine-tuning of model parameters, nor the need to use low-rank adaptation (LoRA) or image adapter (IP-Adapter) for consistency adjustment, and a video with high object consistency can be generated.
[0008] Preferably, during the low-frequency feature extraction process, edge detection based on pixel operations, image smoothing, morphological operations, or object detection, image recognition, and image segmentation methods in deep learning are used for low-frequency feature extraction.
[0009] Preferably, during the high-frequency feature extraction process, a pre-trained visual encoding model is used to extract image features, and feature learning is performed through a multi-layer perceptron (MLP) to obtain optimized high-frequency features.
[0010] Preferably, in the scenario of maintaining human consistency, a face recognition model is additionally used to extract face features, and the face features and object high-frequency features are fused through a fusion module to obtain the final high-frequency features.
[0011] Preferably, a cross-attention layer is added between the attention layer and the fully connected layer of the diffusion transformer (DiT) model, and only the key-value (KV) matrix parameters in the cross-attention layer are trained to keep the overall number of model parameters stable.
[0012] Preferably, during the model training process, a random frame is selected from the input video as a reference image, and the high-frequency features and low-frequency features of this frame are extracted to enhance the object consistency of video generation.
[0013] Preferably, during the training process of the video generation model, a diffusion model is used for adding noise and denoising, or a method based on solving ordinary differential equations, a method based on a flow model (Flow-based Model), or a method based on an autoregressive model (VAR model) is used.
[0014] Preferably, the method is applied to the video generation task of maintaining human consistency and is implemented based on the CogVideoX framework.
[0015] Compared with the prior art, the advantages of the present invention are:
[0016] (1) By separately injecting low-frequency features and face high-frequency features extracted and fused through face recognition, the facial features and overall details of the characters in the video are kept consistent, thus significantly enhancing the realism and coherence of the generated video.
[0017] (2) The present invention uses a pre-trained model to extract high-frequency and low-frequency features, and only updates the parameters of the additional cross-attention and fusion modules during the training process, avoiding the need for full update of the original large model, thereby greatly reducing the computing resources and training costs.
[0018] (3) Using the rich features obtained from the early preprocessing and pre-trained model, the present invention can be directly applied to new scenarios after training without additional fine-tuning for each new case, thus improving the practicality and deployment efficiency of the model. Description of the Drawings
[0019] Figure 1 It is the overall method flow chart of the present invention. Detailed Embodiments
[0020] The following further describes the present application in detail with reference to examples, comparative examples and performance detection tests, and these examples should not be construed as limiting the scope claimed by the present application.
[0021] Examples
[0022] A video generation method based on the fusion of high-frequency and low-frequency features, comprising the following steps:
[0023] Low-frequency feature extraction: Extract low-frequency information from the reference image through a global feature extractor. The low-frequency information includes the overall structure, contour and core key point map of the image; Use traditional visual processing techniques and deep learning algorithms to process the reference image to obtain low-frequency features, where the traditional visual processing techniques include edge detection based on pixel operations, image smoothing processing, and morphological operations, and the deep learning methods include object detection, image recognition, and image segmentation; Use a video variational autoencoder to extract the latent space representation of the low-frequency features, and splice it with the noise features of the video and input it into the diffusion transformer model;
[0024] High-frequency feature extraction: Use a pre-trained model to extract features from the reference image to obtain high-frequency features containing object detail information; Perform feature learning on the high-frequency features through a multi-layer perceptron to obtain optimized high-frequency features; In the application scenario of maintaining human consistency, use a face recognition model to extract face features, and fuse the face features and object features through a fusion module to form the final high-frequency features; Add a cross-attention layer between the attention layer and the fully connected layer of the diffusion transformer model, and inject the final high-frequency features into this layer. The cross-attention layer only trains the newly added key-value matrix, keeping the original parameters of the DiT model unchanged;
[0025] Model training and inference: During the model training process, a random frame is selected from the input video as the reference image, and the high-frequency features and low-frequency features of this frame are extracted and injected into the specified positions of the model respectively; a diffusion model is used for video generation, and the model training process is optimized by adding noise and denoising methods.
[0026] The model training and inference steps are as follows:
[0027] Step 1: In each training batch, a random frame is extracted from the input video as the reference image. In practice, a random function can be used to ensure that different frames of each video have the possibility of being selected, so as to cover diverse scene changes. This randomness ensures the richness of conditional information during the training process and helps the model to model the consistency of different scenes.
[0028] Step 2: For the selected reference frame, extract respectively:
[0029] Low-frequency features: Containing information such as overall structure, contour, and core key points, usually obtaining a low-dimensional latent representation through a video VAE encoder.
[0030] High-frequency features: Containing detailed textures, edges, local features, etc., extracted using a pre-trained vision model (such as CLIP) and further learned and optimized by a multi-layer perceptron (MLP).
[0031] Feature injection:
[0032] Inject the low-frequency features into the latent space of the video generation model as conditional information for the global structure. At the same time, by adding a cross-attention module between the attention layer and the fully connected layer of the diffusion transformer (DiT), the high-frequency features are injected to supplement local details. Only the parameters of the newly added cross-attention layer (such as the KV matrix) are updated here, and the remaining main model parameters remain frozen, thereby reducing the overall training cost.
[0033] Step 3: Using the latent representation after feature injection as the basis, add random noise to it according to a predefined noise scheduling strategy (such as linear or cosine scheduling), simulating the process of data gradually changing from a clear state to a noisy state. This step can be regarded as the forward diffusion process, creating conditions for reverse denoising and ensuring that the model learns the mapping relationship between noise and signal.
[0034] Step 4: Input the noisy latent representation and the previously injected high- and low-frequency features into the diffusion model together. The model predicts the current noise residual at each step and gradually removes the noise, making the latent representation gradually transition from a high-noise state to a clear state close to the original data. The cross-attention module is used at each step to help the model make full use of the conditional features and improve the restoration effect of the denoising process on details and the global structure.
[0035] Step 5. Loss function design:
[0036] In each reverse diffusion step, calculate the error between the predicted noise and the actually added noise. The mean squared error (MSE) measures the deviation between the prediction and the true noise.
[0037] Local update:
[0038] To reduce the computational overhead, only perform gradient updates on the newly added cross-attention layers and related MLP modules, keeping the parameters of the backbone DiT model unchanged.
[0039] Using a common optimizer (such as Adam), backpropagate according to the loss to update the parameters of the newly added modules, ensuring that the model can better fuse high-frequency and low-frequency conditional information.
[0040] Step 6. Latent representation decoding:
[0041] After multiple denoising steps, obtain the final clear latent representation.
[0042] Use the pre-trained decoder to restore the low-dimensional latent representation to a high-resolution video frame.
[0043] Additional super-resolution magnification techniques can be used to enhance the details of the output, improving the visual quality and object consistency.
[0044] Step 7. Loop iteration:
[0045] Repeat the above steps (reference frame selection, feature injection, noise addition, reverse denoising, and parameter update) until the training loss converges stably.
[0046] Using different reference frames and noise levels in each batch helps the model have strong generalization ability and consistency maintenance ability.
[0047] Regular evaluation:
[0048] Regularly evaluate the object consistency, detail restoration, and overall visual quality of the generated videos on the validation set. Adjust the noise scheduling strategy, injection ratio, and other hyperparameters according to the evaluation results to further optimize the model performance.
[0049] During the low-frequency feature extraction process, use edge detection based on pixel operations, image smoothing processing, morphological operations, or object detection, image recognition, and image segmentation methods in deep learning for low-frequency feature extraction.
[0050] During the high-frequency feature extraction process, use a pre-trained visual coding model to extract image features and perform feature learning through a multi-layer perceptron to obtain optimized high-frequency features.
[0051] In the scenario of maintaining person consistency, a face recognition model is additionally used to extract face features, and a fusion module is used to fuse the face features and the high-frequency features of the object to obtain the final high-frequency features.
[0052] The steps are as follows:
[0053] Step 1: Preprocess the input image (or video frame). Use a pre-trained face detection model (such as MTCNN, RetinaFace) to detect the face in the image and output the face bounding box. Crop the face region according to the detection result, and perform standardization processing (such as alignment, size normalization, etc.) on the cropped image to ensure the consistency of subsequent feature extraction.
[0054] Step 2: Input the cropped face image into a pre-trained face recognition model (such as FaceNet, ArcFace or InsightFace) to extract the face embedding vector, capture the key details and personality features of the face. The obtained embedding vector is used as a fixed representation, and the model parameters are frozen, and only a small amount of adjustment (such as the mapping layer) is performed in the subsequent fusion stage.
[0055] Step 3: Use a pre-trained vision model (such as CLIP) on the original image (or video frame) to extract global high-frequency detail features, covering edges, textures and local details. Further feature learning and dimensionality reduction processing are performed on the extracted high-frequency features through a multi-layer perceptron (MLP) to make them more suitable for alignment with face features.
[0056] Step 4: Map the face embedding vector and the object high-frequency features to the same feature space (such as through a linear mapping or a small fully connected layer) to ensure that the two parts of the features have the same dimension and scale.
[0057] Step 5: Adopt a cross-attention mechanism or a Q-Former module to perform interactive fusion on the face features and the object high-frequency features:
[0058] a. Feature splicing or weighted fusion: Splice the two parts of the aligned features or combine them through weighted summation;
[0059] b. Attention calculation: Use the cross-attention layer to calculate the importance of the face features in the global high-frequency features and dynamically adjust the contributions of the two parts of the features;
[0060] c. Further processing by MLP: Input the fused features into the MLP layer for non-linear transformation to generate the final high-frequency feature representation.
[0061] Step 6: The high-frequency features output by the fusion module contain complete face details and object details. As the final high-frequency features, inject these high-frequency features into the specified layer of the video generation model to ensure the consistency of the facial and overall details in the generated video.
[0062] During the training process, freeze the pre-trained face recognition model and visual feature extraction model, and only fine-tune the parameters of the fusion module (including the mapping layer, cross-attention layer, and subsequent MLP).
[0063] Use a standard loss function (such as mean squared error, etc.) to calculate the difference between the generated video and the target video, and backpropagation only updates the parameters of the fusion module to achieve efficient training at low cost.
[0064] Add a cross-attention layer between the attention layer and the fully connected layer of the diffusion transformer model, and only train the key-value matrix parameters in the cross-attention layer to keep the overall number of model parameters stable.
[0065] During the model training process, randomly select a certain frame from the input video as the reference image, and extract the high-frequency features and low-frequency features of this frame to enhance the object consistency of video generation.
[0066] During the training process of the video generation model, use a diffusion model for adding noise and denoising, or use methods based on solving ordinary differential equations, flow models, and autoregressive models.
[0067] The method is applied to the video generation task of maintaining human consistency and is implemented based on the CogVideoX framework.
[0068] The above shows and describes the basic principles, main features, and advantages of the present invention; those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed; the scope of protection claimed by the present invention is defined by the appended claims and their effects.
Claims
1. A video generation method based on high- and low-frequency feature fusion, characterized in that: The following steps are involved: Low-frequency feature extraction: Low-frequency information is extracted from the reference image through a global feature extractor. The low-frequency information includes the overall structure, contour, and core key point map of the image. The reference image is processed using traditional visual processing technology and deep learning algorithms to obtain low-frequency features. Traditional visual processing technology includes edge detection based on pixel operations, image smoothing, and morphological operations. Deep learning methods include target detection, image recognition, and image segmentation. A video variational autoencoder is used to extract the latent space representation of low-frequency features, which is then concatenated with the noise features of the video and input into the diffusion transformer model. High-frequency feature extraction: Use the pre-trained model to extract features from the reference image to obtain high-frequency features containing object detail information; use the multi-layer perceptron to learn the high-frequency features to obtain optimized high-frequency features; in the application scenario of maintaining the consistency of the person, use the face recognition model to extract the face features, and use the fusion module to fuse the face features and object features to form the final high-frequency features; A cross attention layer is added between the attention layer and the fully connected layer of the diffusion transformer model, and the final high-frequency features are injected into this layer. The cross attention layer only trains the newly added key-value matrix, keeping the original DiT model parameters unchanged; Model training and reasoning: During the model training process, a frame is randomly selected from the input video as a reference image, and the high-frequency and low-frequency features of the frame are extracted and injected into the specified positions of the model respectively; the diffusion model is used for video generation, and the model training process is optimized by adding noise and denoising methods.
2. The video generation method based on high- and low-frequency feature fusion according to claim 1, characterized in that: In the process of extracting the low-frequency features, edge detection based on pixel operations, image smoothing, morphological operations or target detection, image recognition, and image segmentation methods in deep learning methods are used to extract the low-frequency features.
3. The video generation method based on high- and low-frequency feature fusion according to claim 1, characterized in that: In the high-frequency feature extraction process, a pre-trained visual encoding model is used to extract image features, and feature learning is performed through a multi-layer perceptron to obtain optimized high-frequency features.
4. The video generation method based on high- and low-frequency feature fusion according to claim 1, characterized in that: In the scene of maintaining character consistency, a face recognition model is additionally used to extract facial features, and the facial features and object high-frequency features are fused through a fusion module to obtain the final high-frequency features.
5. The video generation method based on high- and low-frequency feature fusion according to claim 1, characterized in that: A cross-attention layer is added between the attention layer and the fully connected layer of the diffusion transformer model, and only the key-value matrix parameters in the cross-attention layer are trained to keep the overall parameter amount of the model stable.
6. The video generation method based on high- and low-frequency feature fusion according to claim 1 is characterized in that: During the model training process, a frame is randomly selected from the input video as a reference image, and the high-frequency and low-frequency features of the frame are extracted to enhance the consistency of objects generated by the video.
7. The video generation method based on high- and low-frequency feature fusion according to claim 1 is characterized in that: During the training process of the video generation model, a diffusion model is used for noise addition and denoising, or a method based on solving ordinary differential equations, a method based on a flow model, or a method based on an autoregressive model is used.
8. The video generation method based on high- and low-frequency feature fusion according to claim 1, characterized in that: The method is applied to the task of video generation with character consistency preservation and is implemented based on the CogVideoX framework.