Video image frame processing method and system

By adopting low-frequency feature extraction and high-frequency feature injection methods in video generation technology, the problem of difficult to maintain object consistency is solved, and the generation and consistency of high-quality videos are improved, and training resources and costs are reduced.

CN120111278APending Publication Date: 2025-06-06DIGITAL (SHANGHAI) ENTERPRISE DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510234748.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing video generation technology is difficult to effectively maintain object consistency, resulting in the generated video collapse in a short period of time and affecting the overall quality.

Method used

A video image frame processing method is adopted, through technical means of low-frequency feature extraction and high-frequency feature injection, combined with global feature extractor, video VAE, pre-trained large model and DiT model, to ensure that the consistency of objects is improved during video generation.

Benefits of technology

By injecting high and low frequency features, the consistency of objects in video generation is improved, the generated video quality is improved, and no fine-tuning or additional training is required for each new case, reducing training resources and costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120111278A_ABST
    Figure CN120111278A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image frame processing, and discloses a video image frame processing method and system, and the method comprises the following steps: S1, video input: inputting a certain frame of a video to form a picture; s2, low-frequency feature extraction: acquiring features rich in low-frequency information of the picture by adopting a global feature extractor to form low-frequency features of the picture; s3, feature splicing: adopting video VAE to extract potential spatial features, splicing the potential spatial features with noise features and text features of the video, and sending the spliced features into a DiT model; s4, high-frequency feature extraction: extracting picture object features by adopting a pre-trained large model, and then forming picture high-frequency features through MLp layer feature learning; s5, high-frequency feature injection; and S6, generating a video. The method is reasonable in design, and the high-frequency information and the low-frequency information in the picture are respectively injected into the corresponding positions of the model, so that the consistency of objects in the generated video is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image frame processing, and in particular to a method and system for processing video image frames. Background Art

[0002] Maintaining object consistency in video generation is one of the key factors in evaluating video quality, especially in terms of visual coherence, realism, and user experience. Specifically, object consistency ensures that the appearance and movement of objects in the video remain consistent between different time frames, avoiding sudden changes or unnatural transitions in objects, which is crucial to improving the realism and naturalness of the video. For users, consistent object representation significantly improves the viewing experience and is one of the core criteria for measuring the quality of generated videos.

[0003] In application areas such as virtual reality, augmented reality, and filmmaking, maintaining object consistency is not only an important factor in ensuring technical effects, but also the key to improving user satisfaction. In current video generation technology, maintaining object consistency is still a major challenge in the industry, especially because leading generation video models often fail to effectively maintain object consistency, causing the generated video to collapse in a short period of time, thus affecting the overall quality.

[0004] At present, the industry generally uses technologies such as LoRA (Low-Rank Adaptation) and IP-Adapter to maintain object consistency. However, these methods usually require retraining the model, and every time a new object is introduced, a separate LoRA model needs to be trained for the object, resulting in a huge workload. In application scenarios such as object introduction videos or character portraits, maintaining object consistency is particularly important, which has a crucial impact on the quality and performance of the generated video; therefore, we propose a video image frame processing method and system to solve this problem. Summary of the invention

[0005] The purpose of the present invention is to solve the shortcomings mentioned in the above background technology and to propose a video image frame processing method and system.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] A method for processing video image frames, comprising the following steps:

[0008] S1. Video input: input a certain frame of video to form an image;

[0009] S2, low-frequency feature extraction: Use the global feature extractor to obtain the features of the image that are rich in low-frequency information to form the low-frequency features of the image;

[0010] S3, feature splicing: Video VAE is used to extract latent space features, which are then spliced ​​together with the noisy features and text features of the video and fed into the DiT model;

[0011] S4, high-frequency feature extraction: use the pre-trained large model to extract the features of the image objects, and then learn the features through the MLp layer to form the high-frequency features of the image;

[0012] S5, high-frequency feature injection: By adding a cross-attention layer to the original transformer module, high-frequency features are injected into the original Dit model;

[0013] S6. Video generation: Train the Dit model to generate high-quality videos.

[0014] Preferably, in S2, the low-frequency information includes but is not limited to the overall structure, contour and core key point map corresponding to the image, and the global feature extractor uses traditional image processing technology and deep learning processing technology to identify, detect and segment the image, and the traditional image processing technology is not limited to Opencv.

[0015] Preferably, in S3, the DiT model includes an attention layer and a fully connected layer.

[0016] Preferably, in S4, the CLIP pre-trained large model is used to extract image object features. When the scene is a face, the face recognition pre-trained large model is used to extract face features in the image. The face features and object features are fused through the Q-Former module.

[0017] Preferably, in S5, the cross attention layer is set between the attention layer and the fully connected layer of the original DiT. During training, the original DiT large model gradient is locked, the parameters remain unchanged, and only the additional attention layer is trained.

[0018] Preferably, when using the Dit model, the loss function is not limited, and the Dit model includes but is not limited to the classic diffusion model denoising method, the improved version that converts the noise distribution solution into the ordinary differential equation solution, the Flow model or the autoregressive model.

[0019] The present invention also provides a video image frame processing system, comprising:

[0020] The video input module is used to receive each frame of the video and convert it into image format for subsequent processing;

[0021] The low-frequency feature extraction module uses a global feature extractor to extract low-frequency features from the input image. These low-frequency features mainly include the overall structure, outline, and core key point map of the image. The low-frequency feature extraction module ensures that the basic information of the image is obtained to provide support for subsequent feature processing;

[0022] Feature concatenation module, which extracts latent space features through video VAE and concatenates them with noisy features and text features to form a comprehensive set of input features. The concatenated features are fed into the DiT model for further processing and generation;

[0023] The high-frequency feature extraction module is used to extract object features from input images using a pre-trained large model, perform feature learning through the MLP layer, extract high-frequency features of the image, and further extract facial features through the face recognition model in specific scenarios;

[0024] The high-frequency feature injection module is used to add a cross-attention layer to the original DiT model, injecting the high-frequency features obtained from the high-frequency feature extraction module into the original DiT model;

[0025] The DiT model training module is used to train the DiT model after injecting high-frequency features to generate high-quality videos;

[0026] The loss function module is used to set the loss function when training the DiT model.

[0027] Compared with the prior art, the present invention provides a method and system for processing video image frames, which have the following beneficial effects:

[0028] (1) By injecting high-frequency information and low-frequency information in the image into the corresponding positions of the model, the consistency of objects in the generated video can be improved;

[0029] (2) When training the video generation model, a frame is randomly selected from the input video as a reference image, and the high-frequency and low-frequency features of the reference image are extracted and injected into specific positions of the model for training, thereby ensuring full utilization of information and generating high-quality videos.

[0030] (3) Due to the extensive use of pre-processing and pre-training models to obtain high- and low-frequency features, the entire model training requires very few parameters to be updated. Compared with the hundreds of millions or billions of parameters of the original Dit model, the additional parameters required for training are very few, the training resources are small, and the training cost is low.

[0031] The invention is reasonably designed and can be used directly after the model training is completed to generate videos with high object consistency. Since the high-frequency information of the image is input during training, and the low-frequency information is used as an additional input to the model, the model will fully learn the features, so there is no need to fine-tune for each new case, and there is no need to additionally train object lora, or IP-adapter and other methods to maintain consistency. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 A flowchart of a video image frame processing method proposed in this embodiment;

[0033] Figure 2 The original image used in this embodiment;

[0034] Figure 3 This embodiment generates a video in which the 26th frame is captured;

[0035] Figure 4 This embodiment generates a video in which the 51st frame of the image is captured;

[0036] Figure 5 This embodiment generates a video in which the 76th frame of the image is captured;

[0037] Figure 6 This embodiment generates a video in which the 101st frame of the image is captured;

[0038] Figure 7 This embodiment generates a video in which the 126th frame of the image is captured;

[0039] Figure 8 This embodiment generates a video in which the 151st frame of the image is captured;

[0040] Fig. 9 This is a step diagram of a video image frame processing method proposed in this embodiment. DETAILED DESCRIPTION

[0041] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0042] In the description of the present invention, it is necessary to understand that the terms "upper", "lower", "front", "back", "left", "right", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship are based on the orientation or position relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.

[0043] Reference Figure 1-9 , a video image frame processing method, comprising the following steps:

[0044] S1. Video input: input a certain frame of video to form an image;

[0045] S2, low-frequency feature extraction: Use the global feature extractor to obtain the features of the image that are rich in low-frequency information to form the low-frequency features of the image;

[0046] S3, feature splicing: Video VAE is used to extract latent space features, which are then spliced ​​together with the noisy features and text features of the video and fed into the DiT model;

[0047] S4, high-frequency feature extraction: use the pre-trained large model to extract the features of the image objects, and then learn the features through the MLp layer to form the high-frequency features of the image;

[0048] S5, high-frequency feature injection: By adding a cross-attention layer to the original transformer module, high-frequency features are injected into the original Dit model;

[0049] S6. Video generation: Train the Dit model to generate high-quality videos.

[0050] In this embodiment, in S2, the low-frequency information includes but is not limited to the overall structure, contour and core key point map corresponding to the image. The global feature extractor uses traditional image processing technology and deep learning processing technology to identify, detect and segment the image. The traditional image processing technology includes but is not limited to Opencv.

[0051] In this embodiment, in S3, the DiT model includes an attention layer and a fully connected layer.

[0052] In this embodiment, in S4, the CLIP pre-trained large model is used to extract the features of the image objects. When the scene is a face, the face recognition pre-trained large model is used to extract the face features in the image. The face features and object features are fused through the Q-Former module.

[0053] In this embodiment, in S5, the cross attention layer is set between the attention layer and the fully connected layer of the original DiT. During training, the gradient of the original Dit large model is locked, and the parameters remain unchanged. Only the additional attention layer is trained.

[0054] In this embodiment, when using the Dit model, the loss function is not restricted. The Dit model includes but is not limited to the classic diffusion model denoising method, the improved version that converts the noise distribution solution into the ordinary differential equation solution, the Flow model or the autoregressive model.

[0055] The present invention also provides a video image frame processing system, comprising:

[0056] The video input module is used to receive each frame of the video and convert it into image format for subsequent processing;

[0057] The low-frequency feature extraction module uses a global feature extractor to extract low-frequency features from the input image. These low-frequency features mainly include the overall structure, outline, and core key point map of the image. The low-frequency feature extraction module ensures that the basic information of the image is obtained to provide support for subsequent feature processing;

[0058] Feature concatenation module, which extracts latent space features through video VAE and concatenates them with noisy features and text features to form a comprehensive input feature set. The concatenated features are fed into the DiT model for further processing and generation;

[0059] The high-frequency feature extraction module is used to extract object features from input images using a pre-trained large model, perform feature learning through the MLP layer, extract high-frequency features of the image, and further extract facial features through the face recognition model in specific scenarios;

[0060] The high-frequency feature injection module is used to add a cross-attention layer to the original DiT model, injecting the high-frequency features obtained from the high-frequency feature extraction module into the original DiT model. The high-frequency feature injection module strengthens the DiT model's attention to high-frequency details when processing video generation, thereby improving the quality of detail in video generation;

[0061] The DiT model training module is used to train the DiT model after injecting high-frequency features to generate high-quality videos. The DiT model training module optimizes the generation process through training to ensure that the model can accurately generate output that meets the target video quality standards;

[0062] The loss function module is used to set the loss function when training the DiT model. There are no specific restrictions on the design of the loss function module. You can choose the classic diffusion model denoising method, or an improved version that converts the noise distribution solution into an ordinary differential equation solution, or use the Flow flow model, or even the autoregressive model to flexibly respond to different generation requirements.

[0063] In this embodiment, the video file is received by the video input module and the video is split into single-frame images as the basic input for subsequent processing. Then, the low-frequency feature extraction module uses a global feature extractor to extract low-frequency features in the image to capture the overall structure, contour and key point information. Then, the feature splicing module uses the video VAE (variational autoencoder) to extract the latent space features, and splices them with low-frequency features, noisy features and text features to form a richer input feature set, which is then sent to the DiT (Diffusion Transformer) model for preliminary processing. On this basis, the high-frequency feature extraction module extracts object features in the image through a pre-trained large model (such as CLIP), and uses a face recognition pre-training model to extract face features in specific scenarios. The extracted high-frequency features are further learned through the MLP (multi-layer perceptron) layer to enhance the ability to express details. Next, the high-frequency feature injection module injects the extracted high-frequency features into the DiT model, and by adding a cross-attention layer to the DiT model, the model can better focus on details and object consistency during the generation process. Finally, the DiT model training and video generation module optimizes the training generation capability to ensure the generation of high-quality videos. During the training process, the loss function module selects an appropriate loss function based on the generation target and flexibly adjusts the generation strategy, such as using diffusion models to add noise and remove denoising, ordinary differential equations to solve noise distribution or autoregressive models, etc., to continuously optimize model parameters to ensure the fidelity and consistency of the generated video details. Through the collaborative work of these modules, the system can effectively process video image frames and generate high-quality, detailed and consistent videos.

[0064] In one embodiment, the scene is to maintain the consistency of the characters, and the large model base model adopts the CogVideox framework. Figure 2 As shown, the resolution is 1080*716, and the prompt is: "The scene depicts a woman wearing exquisite hybrid armor, the surface of the armor is inlaid with shimmering colored gems, gently gleaming in the soft light. She stands quietly in a gentle cherry blossom rain, and the petals slowly fall around her. Her eyes are sharp and peaceful, exuding a quiet determination, and the breeze gently stirs a strand of hair at the end of her hair. The tranquil courtyard, surrounded by moss-covered stone walls and wooden arches, forms a peaceful background, and the cherry blossom petals cast delicate shadows on the ground. The petals swirl gracefully around her, adding a dreamy atmosphere to the scene, while the blurred edges of the picture focus on her calm figure. The overall atmosphere combines elegance, strength, and a quiet sense of preparation, capturing a brief moment of calm in the face of an impending challenge."

[0065] The results of capturing frames 26, 51, 76, 101, 126, and 151 in the generated video are as follows: Figure 3 ,4 As shown in Figures 5, 6, 7 and 8, it can be seen that the present invention is indeed very effective in maintaining object consistency.

[0066] The standard parts used in the present invention can all be purchased from the market, and special-shaped parts can be customized according to the instructions and the drawings. The specific connection methods of each part adopt conventional means such as mature bolts, rivets, welding, etc. in the prior art. Machinery, parts and equipment all adopt conventional models in the prior art, and the circuit connection adopts the conventional connection method in the prior art, which will not be described in detail here.

Claims

1. A method for processing video image frames, characterized in that: The steps include: S1. Video input: input a certain frame of video to form an image; S2, low-frequency feature extraction: Use the global feature extractor to obtain the features of the image that are rich in low-frequency information to form the low-frequency features of the image; S3, feature splicing: Video VAE is used to extract latent space features, which are then spliced ​​together with the noisy features and text features of the video and fed into the DiT model; S4, high-frequency feature extraction: use the pre-trained large model to extract the features of the image objects, and then learn the features through the MLp layer to form the high-frequency features of the image; S5, high-frequency feature injection: By adding a cross-attention layer to the original transformer module, high-frequency features are injected into the original Dit model; S6. Video generation: Train the Dit model to generate high-quality videos.

2. The video image frame processing method according to claim 1, characterized in that: In S2, the low-frequency information includes but is not limited to the overall structure, contour and core key point map corresponding to the image. The global feature extractor uses traditional image processing technology and deep learning processing technology to identify, detect and segment the image. The traditional image processing technology includes but is not limited to Opencv.

3. The video image frame processing method according to claim 1, characterized in that: In S3, the DiT model includes an attention layer and a fully connected layer.

4. The video image frame processing method according to claim 1, characterized in that: In S4, the CLIP pre-trained large model is used to extract the features of the image objects. When the scene is a face, the face recognition pre-trained large model is used to extract the face features in the image. The face features and object features are fused through the Q-Former module.

5. The video image frame processing method according to claim 1, characterized in that: In S5, the cross attention layer is set between the attention layer and the fully connected layer of the original DiT. During training, the gradient of the original DiT large model is locked, and the parameters remain unchanged. Only the additional attention layer is trained.

6. The video image frame processing method according to claim 1, characterized in that: When using the Dit model, the loss function is not restricted. The Dit model includes but is not limited to the classic diffusion model denoising method, the improved version that converts the noise distribution solution into the ordinary differential equation solution, the Flow model or the autoregressive model.

7. A video image frame processing system, characterized in that: include: The video input module is used to receive each frame of the video and convert it into image format for subsequent processing; The low-frequency feature extraction module uses a global feature extractor to extract low-frequency features from the input image. These low-frequency features mainly include the overall structure, outline, and core key point map of the image. The low-frequency feature extraction module ensures that the basic information of the image is obtained to provide support for subsequent feature processing; Feature concatenation module, which extracts latent space features through video VAE and concatenates them with noisy features and text features to form a comprehensive set of input features. The concatenated features are fed into the DiT model for further processing and generation; The high-frequency feature extraction module is used to extract object features from input images using a pre-trained large model, perform feature learning through the MLP layer, extract high-frequency features of the image, and further extract facial features through the face recognition model in specific scenarios; The high-frequency feature injection module is used to add a cross-attention layer to the original DiT model, injecting the high-frequency features obtained from the high-frequency feature extraction module into the original DiT model; The DiT model training module is used to train the DiT model after injecting high-frequency features to generate high-quality videos; The loss function module is used to set the loss function when training the DiT model.