Video understanding generation method based on picture key frame and motion vector fusion

Through the fusion method based on image keyframes and motion vectors, combined with the Moving tokenizer and LLAMA model, the problem of excessive storage and resource consumption in diversified data processing is solved, and the effect of reducing resource consumption and retaining video timing is achieved.

CN119992406APending Publication Date: 2025-05-13LINKER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411935057.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When existing video encoding technologies process diversified data, they lead to excessive storage and resource consumption, and may lose video timing.

Method used

The video understanding generation method based on image keyframes and motion vector fusion is adopted, and the keyframes are extracted through the scene detection algorithm, motion vector modeling is performed, and the Moving tokenizer is used to encode and decode it, and text or image token generation is combined with the LLAMA model.

Benefits of technology

Reduces the number of tokens of video tokenizers, reduces the storage and resource consumption of computation attention, and retains the timing and action prediction information of the video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005211300810000031
    Figure BDA0005211300810000031
  • Figure FDA0005211300800000021
    Figure FDA0005211300800000021
Patent Text Reader

Abstract

The invention discloses a video understanding generation method based on picture key frame and motion vector fusion. The video understanding generation method comprises the following steps: step 1, extracting a key frame from a video by using a scene detection algorithm; 2, calculating a motion vector; step 3, after the motion vector modeling is completed in the step 2, performing Moving token encoding and decoding on the video data; step 4, obtaining the token of the picture to realize the discretized token of the picture; and step 5, realizing video understanding or generation. According to the video understanding and generating method based on the fusion of the picture key frame and the motion vector, the understanding or generation of the video can be effectively completed by extracting the key frame through the setting of the step 1 to the step 5.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multimodal understanding generation, and more specifically to a video understanding generation method based on the fusion of picture key frames and motion vectors. Background Art

[0002] With the growth of machine learning applications, intelligent platforms are gradually being adopted in the fields of Internet of Vehicles, video surveillance, smart cities, etc. These intelligent platforms generate massive amounts of communication data with a large number of sensors. As the amount of data increases, the inefficiency of traditional video encoding methods becomes increasingly prominent, making it difficult to meet business needs in terms of data latency and data scale. Therefore, it is urgent to launch video encoding technology for intelligent platforms.

[0003] In order to solve the above problems, there is an invention patent with publication number 113810695A in the prior art, entitled Video Coding Method, Device and Computer-readable Storage Medium, which discloses a video coding method. The method mainly solves the above problems by extracting the feature vector and time index of the key frame from the video, and then encoding the feature vector and time index of the key frame. However, with the development of the times, data types are becoming more and more diverse. Therefore, the prior art adopts the method of decomposing the video frame into spatiotemporal coding blocks to achieve video coding. A representative example is Sora. However, Sora adopts a 3D solution. Compared with the single-frame extraction solution, the number of tokens will be very large, which leads to a problem: if there are 1 million tokens, training and calculating their attention will consume a lot of storage and resources. However, the use of the above-mentioned single-frame extraction solution will result in loss of timing. Summary of the invention

[0004] In view of the shortcomings of the prior art, the object of the present invention is to provide a video understanding generation method based on the fusion of picture key frames and motion vectors, which consumes less storage and resources and does not cause loss of timing.

[0005] To achieve the above object, the present invention provides the following technical solution: a video understanding generation method based on image key frame and motion vector fusion, characterized in that it includes the following steps:

[0006] Step 1: Use scene detection algorithm to extract key frames from the video, and use these key frames as representative frames of the video to summarize the main content and scene changes of the video;

[0007] Step 2: Based on the key frames extracted in step 1, the motion vector of the subsequent action is modeled and then the motion vector is calculated;

[0008] Step 3: After completing the motion vector modeling in step 2, the video data is encoded and decoded using Moving tokenizer;

[0009] Step 4: Use anyres and sglip to split the extracted key frame images into different patches according to the resolution, and obtain the image token to realize the image discretization token;

[0010] Step 5: Finally, the text token, image discretization token, and motion vector token are concat together and input into LLAMA, and the corresponding text token or image token is obtained through autoregression, thereby realizing video understanding or generation.

[0011] As a further improvement of the present invention, in step 2, the specific steps of modeling the motion vector of the subsequent action and then calculating the motion vector are:

[0012] Step 21, preprocessing the key frame image extracted in step 1, converting the color image into a grayscale image;

[0013] Step 22, performing Gaussian pyramid decomposition on the grayscale image to generate multiple image pyramids of different scales to construct a Gaussian pyramid;

[0014] Step 23, repeating each pyramid level of the Gaussian pyramid constructed in step 22 to calculate the optical flow;

[0015] Step 24, interpolate the optical flow vectors calculated at each pyramid level to obtain the optical flow field of the entire image;

[0016] Step 25: further process the optical flow field obtained in step 24, remove outliers, perform smoothing, and finally obtain the motion vector.

[0017] As a further improvement of the present invention, the specific steps of calculating the optical flow in steps 2 and 3 are as follows:

[0018] Step a, densely sample the image at the current pyramid level to obtain the optical flow vector of each pixel;

[0019] Step b: For each pixel, the optical flow vector is calculated by comparing the change in grayscale value between adjacent frames. The specific formula is as follows:

[0020]

[0021] Among them, Ix and Iy represent the gradient of the image in the x and y directions respectively, and It represents the gradient of the image in time;

[0022] Step c: use the gradient information calculated in step b to estimate the optical flow vector

[0023] u=mean(I x I x +I y I y ) -1 ·(I x I t )

[0024] v=mean(I x I x +I y I y ) -1 ·(I y I t )

[0025] Mean represents the mean operation on surrounding pixels.

[0026] As a further improvement of the present invention, the specific method of encoding and decoding the Moving tokenizer in step 3 is: designing a coding Tokenizer to encode the actions in the video into tokens, and inputting these tokens into the LLAMA / LLAVA model to convert the vectors into tokens through semantic encoding. As a further improvement of the present invention, the Moving tokenizer in step 3 is trained using MSE loss.

[0027] The beneficial effects of the present invention are as follows: through the settings of step 1 to step 5, the video understanding generation method based on the fusion of image key frames and motion vectors of the present invention applies a moving tokenizer, 1. greatly reduces the number of tokens of the video tokenizer, and reduces the consumption of attention calculation resources. 2. Increases the information of the video's timing and action prediction. DETAILED DESCRIPTION

[0028] The present invention will be further described in detail with reference to the given embodiments below.

[0029] Terminology explanation:

[0030] Moving tokenizer is a mobile tokenizer;

[0031] Token is a tag;

[0032] The LLAMA / LLAVA model is a large language model.

[0033] A video understanding generation method based on image key frame and motion vector fusion in this embodiment includes the following steps:

[0034] 1. Key frame extraction

[0035] First, a scene detection algorithm (such as scenedetect) is used to extract key frames from the video. These key frames, as representative frames of the video, can summarize the main content and scene changes of the video.

[0036] 2. Motion Vector Modeling

[0037] Based on the extracted key frames, the motion vectors of subsequent actions are modeled. In this way, it is not necessary to keep all the frames of the entire video in video processing, only the motion vectors need to be kept. At the same time, this method can preserve the timing information of the video.

[0038] The motion vector can be calculated by an optical flow algorithm (such as Farneback optical flow).

[0039] 2.1 Preprocess the key frame images and convert the color images into grayscale images in order to better handle the changes in pixel grayscale values.

[0040] 2.2 Constructing Gaussian Pyramid: Perform Gaussian pyramid decomposition on the grayscale image to generate multiple image pyramids of different scales. This is done to detect motion in the image at different scales.

[0041] 2.3 Calculate optical flow: For each pyramid level, the following steps are repeated:

[0042] a. Dense sampling: Densely sample the image at the current pyramid level to obtain the optical flow vector of each pixel.

[0043] b. Calculate optical flow: For each pixel, the optical flow vector is calculated by comparing the change in grayscale value between adjacent frames. The formula is as follows: Here, Ix and Iy represent the gradients of the image in the x and y directions respectively, and It represents the gradient of the image in time.

[0044] c. Optical flow estimation: Use the calculated gradient information and estimate the optical flow vector

[0045] u=mean(I x I x +I y I y ) -1 ·(I x I t )

[0046] v=mean(I x I x +I y I y )-1 ·(I y I t )

[0047] Mean represents the mean operation on surrounding pixels.

[0048] 2.4 Interpolation: Interpolate the optical flow vectors calculated at each pyramid level to obtain the optical flow field of the entire image.

[0049] 2.5 Post-processing: Further process the optical flow field, remove outliers, smooth the image, etc.

[0050] 3. Moving tokenizer encoding and decoding

[0051] a. Moving Tokenizer

[0052] Design an encoding tokenizer (called Moving tokenizer) to encode the actions in the video into tokens and input these tokens into the LLAMA / LLAVA model. The training process of Moving tokenizer is very similar to that of LLAVA, both of which convert vectors into tokens through semantic encoding. Specifically, the loss function is trained using MSE loss, which is an offline process. The difference is that the training of Moving tokenizer does not require text alignment and can be completed only by relying on the video itself.

[0053] 4. Image discretization tokenizer preprocessing uses anyres and sglip, which splits the image into different patches according to the resolution and obtains the image token

[0054] Finally, the text token, image discretization token, and motion vector token are concat together and input into LLAMA, and the corresponding text token or image token is obtained through autoregression, thereby achieving video understanding or generation.

[0055] This embodiment provides the following specific examples:

[0056] 1. Key frame extraction

[0057] Use a scene detection algorithm such as `scenedetect` to extract keyframes from the video.

[0058] Split each video into multiple segments, and keep only key frames in each segment.

[0059] Video collection: {V1, V2, ..., V n}

[0060] Keyframe collection:

[0061] 2. Motion Vector Modeling

[0062] Based on the extracted key frames, motion vectors between adjacent key frames are calculated.

[0063] The motion vector can be calculated by an optical flow algorithm (such as Farneback optical flow).

[0064] Motion Vector: MV i,j =OpticalFlow(K i,j , K i,j+1 )

[0065] Motion vector collection:

[0066] 3. Moving Tokenizer encoding and decoding

[0067] Design a Moving Tokenizer to encode motion vectors into tokens.

[0068] Train Moving Tokenizer and use MSE Loss for training.

[0069] Encoding function: T MV (MV i,j )=Tokenizer(MV i,j )

[0070] Loss function:

[0071] Motion vector token set:

[0072] 4. Image discretization Tokenizer preprocessing

[0073] Use anyres and sglip to split the image into different patches according to the resolution. Encode each patch into a token.

[0074] Image segmentation: P i,j =Patch(K i,j , resolution)

[0075] Encoding function: T P (P i,j,k )=Tokenizer(P i,j,k )

[0076] Image patch token collection:

[0077] 5. Model training and inference

[0078] The text token, image discretization token, and motion vector token are combined and input into the LLAMA model.

[0079] Use autoregressive method to generate corresponding text token or image token.

[0080] Merge token: T concat =[T text , T P ,T MV ]

[0081] Input to the LLAMA model:

[0082] Loss function:

[0083] In summary, the video understanding generation method based on the fusion of image key frames and motion vectors in this embodiment reduces the number of tokens of the video tokenizer by extracting key frames, reduces the resource consumption of attention calculation, and adopts image discretization and anyres to adaptively split patches according to the video resolution.

[0084] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. A video understanding generation method based on image key frame and motion vector fusion, characterized by: The steps include: Step 1: Use scene detection algorithm to extract key frames from the video, and use these key frames as representative frames of the video to summarize the main content and scene changes of the video; Step 2: Based on the key frames extracted in step 1, the motion vector of the subsequent action is modeled and then the motion vector is calculated; Step 3: After completing the motion vector modeling in step 2, the video data is encoded and decoded using Moving tokenizer; Step 4: Use anyres and sglip to split the extracted key frame images into different patches according to the resolution, and obtain the image token to realize the image discretization token; Step 5: Finally, the text token, image discretization token, and motion vector token are concat together and input into LLAMA, and the corresponding text token or image token is obtained through autoregression, thereby realizing video understanding or generation.

2. The video understanding generation method based on image key frame and motion vector fusion according to claim 1 is characterized in that: In the step 2, the motion vector modeling of the subsequent action is performed, and then the specific steps of calculating the motion vector are as follows: Step 21, preprocessing the key frame image extracted in step 1, converting the color image into a grayscale image; Step 22, performing Gaussian pyramid decomposition on the grayscale image to generate multiple image pyramids of different scales to construct a Gaussian pyramid; Step 23, repeating each pyramid level of the Gaussian pyramid constructed in step 22 to calculate the optical flow; Step 24, interpolate the optical flow vectors calculated at each pyramid level to obtain the optical flow field of the entire image; Step 25: further process the optical flow field obtained in step 24, remove outliers, perform smoothing, and finally obtain the motion vector.

3. The video understanding generation method based on image key frame and motion vector fusion according to claim 2 is characterized in that: The specific steps for calculating the optical flow in steps 2 and 3 are as follows: Step a: densely sample the image at the current pyramid level to obtain the optical flow vector of each pixel. Step b: for each pixel, calculate the optical flow vector by comparing the change of grayscale value between adjacent frames. The specific formula is as follows: Among them, Ix and Iy represent the gradient of the image in the x and y directions respectively, and It represents the gradient of the image in time; Step c: use the gradient information calculated in step b to estimate the optical flow vector u=mean(I x ·I x +I y ·I y ) -1 ·(I x ·I t ) υ=mean(I x ·I x +I y ·I y ) -1 ·(I y ·I t ) Mean represents the mean operation on surrounding pixels.

4. The video understanding generation method based on image key frame and motion vector fusion according to claim 3 is characterized in that: The specific method of encoding and decoding the Moving tokenizer in step 3 is: designing a coding Tokenizer to encode the actions in the video into tokens, and inputting these tokens into the LLAMA / LLAVA model to convert the vectors into tokens through semantic encoding.

5. The video understanding generation method based on image key frame and motion vector fusion according to claim 4 is characterized in that: The Moving tokenizer in step 3 is trained using MSE loss.