A video understanding method and related device

By introducing weight allocation function and time perception layer into the video understanding big model, the problems of key features loss and confusion in timing information in video understanding are solved, and the effect of video understanding is improved.

CN120047777BActive Publication Date: 2025-07-11HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510511550.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-11
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The existing video understanding model uses videos to process videos to lose key features, retain redundant features due to indiscriminate compression of visual features, and has problems such as confusion in timing information, especially when long video understanding.

Method used

The weight allocation function and time-aware layer are introduced in the video understanding big model. The visual characteristics of the video frame are obtained through the visual encoder, the weight allocation function is run to assign different weights to key features and redundant features, and the timing information is added through the time-aware layer to form the target visual feature.

Benefits of technology

Effectively retain key features in the video, suppress redundant features, solve the problem of confusion in timing information, and improve the accuracy and completeness of video comprehension.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047777B_ABST
    Figure CN120047777B_ABST
Patent Text Reader

Abstract

The present application provides a video understanding method and related devices, which relate to the field of computer vision and are applied to a first device. A weight allocation function and a time perception layer are deployed between the visual encoder and the image connector in the video understanding large model installed on this device. This device receives a video understanding task including a first video and a first prompt word input by a second device; obtains first visual features of a first video frame in the first video through the visual encoder; runs the weight allocation function to assign different weights to key features and redundant features in the first visual features to obtain first key visual features; adds the timing information of the first video frame to the first key visual features through the time perception layer to obtain first target visual features; and returns a first video understanding result output by the video understanding large model based on the first video features and the first prompt word to the second device. The present application can solve the problems of loss of effective information and chaos of timing information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular, to a video understanding method and related devices. Background Art

[0002] Current video understanding large models (such as Qwen-VL, VideoLLaMA, and VideoGPT, etc.) are suitable for short video understanding. When processing videos, since all video frames are processed without discrimination, a large amount of compression of visual features will lead to the loss of key features and the retention of a large number of redundant features, thus seriously affecting the video understanding effect. At the same time, due to the limited memory capacity of the model, it is easy to generate temporal information chaos during video understanding, especially for long videos. Summary of the Invention

[0003] In view of the above problems, this application provides a video understanding method and related devices to solve the problems of poor video understanding effect and temporal information chaos generated by the video understanding large model. The specific solutions are as follows:

[0004] In the first aspect of this application, a video understanding method is provided. The method is applied to a first device, in which a video understanding large model is installed. A weight distribution function and a time perception layer are deployed between the visual encoder and the image connector in the video understanding large model. The method includes:

[0005] Receiving a video understanding task input by a second device, where the video understanding task includes a first video and a first prompt;

[0006] Obtaining a first visual feature of a first video frame in the first video through the visual encoder;

[0007] Running the weight distribution function to assign different weights to key features and redundant features in the first visual feature to obtain a first key visual feature;

[0008] Adding the temporal information of the first video frame to the first key visual feature through the time perception layer to obtain a first target visual feature, where the first target visual feature is the basis for the image connector to extract the first video feature;

[0009] Returning a first video understanding result to the second device, where the first video understanding result is output by the video understanding large model based on the first video feature and the first prompt.

[0010] In a possible implementation, the running the weight distribution function to assign different weights to key features and redundant features in the first visual feature to obtain a first key visual feature includes:

[0011] Through the transformation operation of the two-dimensional feature matrix, the first visual feature is divided into multiple feature blocks;

[0012] For two visually consecutive visual features in the first visual feature, calculate the similarity between two feature blocks with the same spatial position, and assign a target weight to the feature block with a later time point according to the similarity, where the sum of the target weight and the similarity is 1;

[0013] Perform weighted processing on the first visual feature according to the target weights corresponding to each feature block in the first visual feature to obtain the first key visual feature.

[0014] In a possible implementation, the calculating the similarity between two feature blocks with the same spatial position includes:

[0015] Calculate the Euclidean distance between two feature blocks with the same spatial position;

[0016] Calculate the similarity using the Euclidean distance, where the similarity is negatively correlated with the Euclidean distance.

[0017] In a possible implementation, the time perception layer includes a fully connected network and a text encoder in a multi-modal model. The fully connected network is pre-trained. Adding the temporal information of the first video frame to the first key visual feature through the time perception layer to obtain a first target visual feature includes:

[0018] Obtain the first timestamp index text of the first video frame;

[0019] Encode the first timestamp index text into a first timestamp projection vector through the text encoder;

[0020] Map the first timestamp projection vector to the dimensional space of the first visual feature through the fully connected network, and superimpose the mapping result of the first timestamp projection vector and the first visual feature to obtain the first target visual feature.

[0021] In a possible implementation, the process of pre-training the fully connected network includes:

[0022] Obtain the data samples for this training, where the data samples include a second video, a second prompt, and an annotated target video understanding result;

[0023] Obtain the second visual feature of the second video frame in the second video through the visual encoder;

[0024] Run the weight assignment function to assign different weights to the key features and redundant features in the second visual feature to obtain a second key visual feature;

[0025] Obtain the second timestamp index text of the second video frame; encode the second timestamp index text into a second timestamp projection vector through the text encoder; map the second timestamp projection vector to the dimensional space of the second visual feature through the fully connected network, and superimpose the mapping result of the second timestamp projection vector with the second visual feature to obtain a second target visual feature, where the second target visual feature is the basis for the image connector to extract the second video feature;

[0026] Obtain the second video understanding result output by the video understanding large model based on the second video feature and the second prompt word;

[0027] Taking the target video understanding result as the target, calculate the loss function value between the second video understanding result and the target video understanding result;

[0028] When the current training end condition is not satisfied, adjust the network parameters of the fully connected network according to the loss function value, enter the next training, and return to execute the acquisition of the data sample for this training;

[0029] When the current training end condition is satisfied, end the training of the fully connected network.

[0030] The second aspect of this application provides a video understanding method, which is applied to a second device, and the method includes:

[0031] Input a video understanding task to the first device, where the video understanding task includes a first video and a first prompt word;

[0032] Receive the first video understanding result returned by the first device. A video understanding large model is installed in the first device. A weight distribution function and a time perception layer are deployed between the visual encoder and the image connector in the video understanding large model. The visual encoder is used to obtain the first visual feature of the first video frame in the first video. When the weight distribution function runs, it can assign different weights to the key features and redundant features in the first visual feature to obtain a first key visual feature. The time perception layer is used to add the timing information of the first video frame to the first key visual feature to obtain a first target visual feature, where the first target visual feature is the basis for the image connector to extract the first video feature, and the first video understanding result is output based on the first video feature and the first prompt word.

[0033] A third aspect of the present application provides a video understanding device, which is applied to a first device. A video understanding large model is installed in the first device. A weight distribution function and a time perception layer are deployed between the visual encoder and the image connector in the video understanding large model. The device includes:

[0034] A task receiving module, configured to receive a video understanding task input by a second device, where the video understanding task includes a first video and a first prompt

[0035] A video understanding module, configured to obtain a first visual feature of a first video frame in the first video through the visual encoder; run the weight distribution function to assign different weights to key features and redundant features in the first visual feature to obtain a first key visual feature; add the timing information of the first video frame to the first key visual feature through the time perception layer to obtain a first target visual feature, where the first target visual feature is the basis for the image connector to extract the first video feature; and return a first video understanding result to the second device, where the first video understanding result is output by the video understanding large model based on the first video feature and the first prompt.

[0036] A fourth aspect of the present application provides a video understanding device, which includes:

[0037] A task input module, configured to input a video understanding task to a first device, where the video understanding task includes a first video and a first prompt

[0038] A result receiving module, configured to receive a first video understanding result returned by the first device. A video understanding large model is installed in the first device. A weight distribution function and a time perception layer are deployed between the visual encoder and the image connector in the video understanding large model. The visual encoder is configured to obtain a first visual feature of a first video frame in the first video. When the weight distribution function runs, it can assign different weights to key features and redundant features in the first visual feature to obtain a first key visual feature. The time perception layer is configured to add the timing information of the first video frame to the first key visual feature to obtain a first target visual feature, where the first target visual feature is the basis for the image connector to extract the first video feature, and the first video understanding result is output by the video understanding large model based on the first video feature and the first prompt.

[0039] A fifth aspect of the present application provides a computer program product, including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement the video understanding method according to the first aspect or any implementation manner of the first aspect.

[0040] A sixth aspect of the present application provides an electronic device, including at least one processor and a memory connected to the processor, where:

[0041] The memory is used to store computer programs;

[0042] The processor is used to execute the computer program so that the electronic device can implement the video understanding method of the above first aspect or any implementation manner of the first aspect.

[0043] A seventh aspect of the present application provides a computer storage medium carrying one or more computer programs, which can enable the electronic device to implement the video understanding method of the above first aspect or any implementation manner of the first aspect when the one or more computer programs are executed by the electronic device.

[0044] With the above technical solutions, a video understanding method and related device provided by the present application are applied to a first device, and a video understanding large model is installed in the first device. A weight distribution function and a time perception layer are deployed between the visual encoder and the image connector in the video understanding large model. The method includes receiving a video understanding task input by a second device, where the video understanding task includes a first video and a first prompt word; obtaining first visual features of a first video frame in the first video through the visual encoder; running the weight distribution function to assign different weights to key features and redundant features in the first visual features to obtain first key visual features; adding the timing information of the first video frame to the first key visual features through the time perception layer to obtain first target visual features, and the first target visual features are the basis for the image connector to extract first video features; returning a first video understanding result to the second device, and the first video understanding result is output by the video understanding large model based on the first video features and the first prompt word. The present application can assign different weights to key features and redundant features in visual features and inject timing information, thereby solving the problems of loss of effective information and chaotic timing information when the video understanding large model analyzes videos and improving the video understanding ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In combination with the drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale.

[0046] Figure 1 It is a schematic structural diagram of an existing video understanding large model;

[0047] Figure 2 It is a schematic structural diagram of a video understanding large model provided by an embodiment of the present application;

[0048] Figure 3 Schematic flowchart of a video understanding method provided by an embodiment of the present application;

[0049] Figure 4 Partial flowchart of a video understanding method provided by an embodiment of the present application;

[0050] Figure 5 Another partial flowchart of a video understanding method provided by an embodiment of the present application;

[0051] Figure 6 Another partial flowchart of a video understanding method provided by an embodiment of the present application;

[0052] Figure 7 Schematic structural diagram of a video understanding device provided by an embodiment of the present application;

[0053] Figure 8 Another flowchart of a video understanding method provided by an embodiment of the present application;

[0054] Figure 9 Another schematic structural diagram of a video understanding device provided by an embodiment of the present application;

[0055] Figure 10 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0056] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. The terms used in the implementation part of the present application are only used to explain the specific embodiments of the present application, rather than intended to limit the present application.

[0057] The embodiments of the present application will be described below with reference to the accompanying drawings. Those of ordinary skill in the art can know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0058] Terms such as "first" and "second" in the specification of the present application and the above accompanying drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing when describing objects with the same attributes in the embodiments of the present application. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.

[0059] The large video understanding model can be used to develop intelligent editing templates, build intelligent review systems, and generate short drama scripts, etc., which is beneficial to reducing economic costs, improving production efficiency, and enhancing business effects. Refer to Figure 1 , Figure 1 which is a schematic structural diagram of an existing large video understanding model. As Figure 1 shown, the existing large video understanding model consists of a visual encoder, an image connector, a video connector, a tokenizer, and a large language model. Among them, the tokenizer obtains tokenization features from the prompt words, the visual encoder encodes video frames into visual features, the image connector compresses and extracts the visual features of all video frames to obtain video features, the video connector compresses and extracts the video features to map the video features to the feature space of the large language model, and the mapped video features and tokenization features are input into the large language model, and the large language model outputs the video understanding result.

[0060] Taking the large video understanding model Qwen-VL-7B as an example, it uses the visual encoder CLIP-ViT-L / 14 to encode video frames into visual features, then uses the method of merging adjacent tokens to compress and extract the visual features to obtain video features, and then uses 3D convolution to map the video features to the feature space of the large language model, and finally outputs the video understanding result through the Qwen2-7B large language model.

[0061] When the existing large video understanding model extracts video features from visual features, it will perform a large amount of compression on the visual features, and the weights of the visual features of different video frames are the same during compression. However, the key features in the video are few, and the undifferentiated compression will cause the loss of key features and retain a large number of redundant features. This problem will be more serious in the case of long videos. At the same time, the model's memory ability is limited, and it is easy to generate chaotic temporal information during video understanding.

[0062] To solve the above problems, an embodiment of the present application provides a video understanding method. The following will introduce in detail a video understanding method according to an embodiment of the present application with reference to the accompanying drawings.

[0063] An embodiment of the present application provides a video understanding method, which is applied to a first device. A large video understanding model is installed in the first device, and a weight distribution function and a time perception layer are deployed between the visual encoder and the image connector in the large video understanding model. Refer to Figure 2 , Figure 2 which is a schematic structural diagram of a large video understanding model provided by an embodiment of the present application. As Figure 2As shown in the figure, the video understanding large model provided by the embodiments of the present application is constructed based on the existing video understanding large model, and includes a visual encoder, an image connector, a video connector, a tokenizer, and a language large model. Moreover, a weight allocation function and a time perception layer are deployed between the visual encoder and the image connector. For the visual features output by the visual encoder, the weight allocation function can assign low weights to redundant features and high weights to key features among them, so as to solve the problem that key features are lost while a large number of redundant features are retained during the undifferentiated compression of visual features. Moreover, the time perception layer can inject temporal information into the visual features, so as to solve the problem that temporal information confusion is likely to occur during video understanding.

[0064] See Figure 3 , Figure 3 is a schematic flow chart of a video understanding method provided by the embodiments of the present application. As Figure 3 shown, a video understanding method provided by the embodiments of the present application may include steps S301 to S305, and the following will describe these steps in detail respectively.

[0065] S301, Receive a video understanding task input by a second device. The video understanding task includes a first video and a first prompt word. In the embodiments of the present application, the second device is a device held by a user with video understanding requirements, and the second device inputs a video understanding task to the first device in response to the user's input operation. For this, the first device obtains the video understanding task of the first device, and parses and obtains the video to be understood (i.e., the first video) and the prompt word (i.e., the first prompt word) therefrom.

[0066] S302, Obtain the first visual feature of the first video frame in the first video through a visual encoder.

[0067] In the embodiments of the present application, continue to refer to Figure 2 , Input the first video into the visual encoder, and at the same time input the first prompt word into the tokenizer. The tokenizer can obtain the corresponding tokenization features from the first prompt word, and the visual encoder encodes each video frame (i.e., the first video frame) in the first video into visual features (i.e., the first visual feature).

[0068] S303, Run a weight allocation function to assign different weights to the key features and redundant features in the first visual feature to obtain a first key visual feature.

[0069] In the embodiments of the present application, continue to refer to Figure 2 , For each first video frame in the first video, run a weight function to assign different weights to the key features and redundant features in its first visual feature, where the key features are assigned high weights and the redundant features are assigned low weights, so as to perform weighted processing on its first visual feature to obtain a first key visual feature.

[0070] S304. Add the timing information of the first video frame to the first key visual feature through the time perception layer to obtain the first target visual feature, which is the basis for the image connector to extract the first video feature.

[0071] In the embodiment of the present application, continue to refer to Figure 2 , for each first video frame in the first video, obtain the timing information of the first video frame in the first video, and add the timing information of the first video frame to the first key visual feature corresponding to the first video frame through the time perception layer, so as to obtain the first target visual feature.

[0072] The image connector compresses and extracts the first target visual features corresponding to each first video frame to obtain the video feature of the first video (i.e., the first video feature).

[0073] The video connector compresses and extracts the first video feature, and maps the first video feature to the feature space of the language large model. The mapped first video feature and the tokenized feature corresponding to the first prompt are input into the language large model, and the language large model outputs the video understanding result of the first video feature and the first prompt (i.e., the first video understanding result).

[0074] S305. Return the first video understanding result to the second device. The first video understanding result is output by the video understanding large model based on the first video feature and the first prompt.

[0075] In the embodiment of the present application, after the first device obtains the first video understanding result output by the video understanding large model, it can return the first video understanding result to the second device, and the second device displays the first video understanding result to the user.

[0076] In a possible implementation, weights can be assigned to the key features and redundant features in the visual features through feature analysis. Refer to Figure 4 , Figure 4 This is a partial flowchart of a video understanding method provided by the embodiment of the present application. As Figure 4 shown, in a video understanding method provided by the embodiment of the present application, step S303, "Run the weight assignment function to assign different weights to the key features and redundant features in the first visual feature to obtain the first key visual feature", may include steps S401 to S403, and these steps will be described in detail below.

[0077] S401. Divide the first visual feature into multiple feature blocks through the conversion operation of the two-dimensional feature matrix.

[0078] In the embodiments of the present application, for each first video frame in the first video, its first visual feature can be converted into a two-dimensional feature matrix, and the size of the two-dimensional feature matrix is fixed as , where is the number of rows of the two-dimensional feature matrix, is the number of columns of the two-dimensional feature matrix. For a first video including first video frames, its three-dimensional visual feature matrix can be obtained, with a size of , and all visual features are included in the three-dimensional visual feature matrix.

[0079] For each first video frame in the first video, the corresponding two-dimensional feature matrix can be divided into multiple feature blocks, and the size of each feature block is , where are the three dimensions of the three-dimensional space. Among them, and are the sizes of the spatial dimensions, is the size of the time dimension.

[0080] S402. For two visual features with consecutive time points in the first visual feature, calculate the similarity between two feature blocks with the same spatial position, and assign a target weight to the feature block with a later time point according to the similarity. The sum of the target weight and the similarity is 1.

[0081] In the embodiments of the present application, for two consecutive first video frames in the first video, the two corresponding first visual features are two visual features with consecutive time points. Assume they are visual feature 1 and visual feature 2. Two feature blocks with the same spatial position in visual feature 1 and visual feature 2 can be used as two feature blocks with consistent spatial dimensions and adjacent time dimensions. For example, the feature block in visual feature 1 and the feature block in visual feature 2, and another example is the feature block in visual feature 1 and the feature block in visual feature 2.

[0082] For the two determined feature blocks with consistent spatial dimensions and adjacent time dimensions, the similarity between them can be calculated, and a weight (i.e., the target weight) is assigned to the feature block with a later time point. The target weight is equal to the difference between 1 and the similarity. Taking the feature block and the feature block as an example, assume the similarity between them is , then the target weight can be assigned to the feature block .

[0083] It should be noted that if the similarity Greater than a preset threshold , which indicates that the feature block and the feature block are duplicates. Additionally, when calculating the similarity between two feature blocks, the original elements in the feature blocks are calculated, rather than the weighted elements.

[0084] S403. According to the target weights corresponding to the respective feature blocks in the first visual feature, perform weighted processing on the first visual feature to obtain the first key visual feature.

[0085] In the embodiment of the present application, continuing to take visual feature 2 as an example, assume that the target weight corresponding to its feature block is weight 1, and the target weight corresponding to the feature block is weight 2. Then, perform weighted calculation on the elements of visual feature 2 within the feature block according to weight 1, that is, each element of visual feature 2 within the feature block is multiplied by weight 1. Additionally, perform weighted calculation on the elements of visual feature 2 within the feature block according to weight 2, that is, each element of visual feature 2 within the feature block is multiplied by weight 2. After the weighted calculation of visual feature 2 in all feature blocks is completed, the key visual feature corresponding to visual feature 2 (i.e., the first key visual feature) can be obtained.

[0086] Based on this, redundant features with high similarity in the visual feature can be suppressed, and in the three-dimensional visual feature matrix of the video, the effective key features will be retained.

[0087] In a possible implementation, the similarity between two feature blocks can be calculated through the Euclidean distance. In this regard, a video understanding method provided by the embodiment of the present application, wherein, in step S402, "calculate the similarity between two feature blocks with the same spatial position" can include the following steps:

[0088] Calculate the Euclidean distance between two feature blocks with the same spatial position; calculate the similarity using the Euclidean distance, and the similarity is negatively correlated with the Euclidean distance.

[0089] In the embodiment of the present application, for two consecutive visual features in the first visual feature in terms of time point, the following formula (1) can be used to calculate the Euclidean distance between two feature blocks ( and ) with the same spatial position :

[0090] (1)

[0091] Furthermore, use the following formula (2) to calculate the similarity :

[0092] (2)

[0093] In a possible implementation, the time perception layer includes a fully connected network and a text encoder in a multi-modal model, and the fully connected network is pre-trained. It should be noted that the text encoder in the multi-modal model (such as CLIP, ALIGN, FLAVA, etc.) is used because the temporal information needs to be fused with visual features, and an encoder trained with pure text cannot be used. The text encoder trained with multi-modal training can have a similar feature representation space as the visual encoder.

[0094] See Figure 5 , Figure 5 which is another schematic diagram of the process of a video understanding method provided by an embodiment of the present application. As Figure 5 shown, for a video understanding method provided by an embodiment of the present application, in step S304, "adding the temporal information of the first video frame to the first key visual feature through the time perception layer to obtain the first target visual feature" may include steps S501 to S503, and these steps will be described in detail below.

[0095] S501, obtain the first timestamp index text of the first video frame.

[0096] In an embodiment of the present application, for each first video frame in the first video, according to the temporal sequence of the first video frame in the first video, the timestamp index text of the first video frame (i.e., the first timestamp index text) can be obtained, and the first timestamp index text is "the first frame", "the second frame",..., "the Nth frame".

[0097] S502, encode the first timestamp index text into a first timestamp projection vector through the text encoder.

[0098] In an embodiment of the present application, for each first video frame in the first video, the text encoder can encode the first timestamp index text corresponding to the first video frame into a timestamp projection vector (i.e., the first timestamp projection vector).

[0099] S503, map the first timestamp projection vector to the dimensional space of the first visual feature through the fully connected network, and superimpose the mapping result of the first timestamp projection vector and the first visual feature to obtain the first target visual feature.

[0100] In the embodiments of the present application, for each first video frame in the first video, the full connection network can map the first timestamp projection vector corresponding to the first video frame to the dimensional space of the first visual feature, keeping the same dimension as the first visual feature. Then, the mapping result of the first timestamp projection vector and the first visual feature are subjected to element-wise addition to obtain a visual feature containing temporal information (i.e., the first target visual feature).

[0101] The process of timestamp projection can be expressed by the following formula (3):

[0102] (3)

[0103] Wherein, represents the first target visual feature, represents the first visual feature, represents the forward calculation of the full connection network, represents the encoding operation of the text encoder, represents the first timestamp index text.

[0104] It should be noted that during the process of timestamp projection, the first visual feature is a one-dimensional vector. In addition, the full connection network can be single-layer or multi-layer, mapping the dimension of the first timestamp projection vector to be consistent with the dimension of the first visual feature, and then performing element-wise addition to realize the fusion of temporal information and time features.

[0105] See Figure 6 , Figure 6 which is another partial process schematic diagram of a video understanding method provided by the embodiments of the present application. As Figure 6 shown, for a video understanding method provided by the embodiments of the present application, the process of pre-training to obtain the full connection network may include steps S601 to S608, and the following will describe these steps in detail.

[0106] S601, obtain the data samples for this training, and the data samples include the second video, the second prompt word, and the labeled target video understanding result.

[0107] In the embodiments of the present application, obtain the data samples required for local training, and the data samples include the training video (i.e., the second video), the prompt word (i.e., the second prompt word), and the pre-labeled video understanding result (i.e., the target video understanding result).

[0108] S602, obtain the second visual feature of the second video frame in the second video through the visual encoder.

[0109] In the embodiments of the present application, continue to refer to Figure 2, input the second video into the visual encoder, and at the same time input the second prompt into the tokenizer. The tokenizer can obtain the corresponding token features from the second prompt, and the visual encoder encodes each video frame (i.e., the second video frame) in the second video into visual features (i.e., the second visual features).

[0110] S603, run the weight assignment function to assign different weights to the key features and redundant features in the second visual features to obtain the second key visual features.

[0111] In the embodiments of the present application, continue to refer to Figure 2 , for each second video frame in the second video, run the weight function to assign different weights to the key features and redundant features in its second visual features, where the key features are assigned high weights and the redundant features are assigned low weights, so as to perform weighted processing on its second visual features to obtain the second key visual features.

[0112] It should be noted that the implementation process of step S603 in the embodiments of the present application can refer to the implementation process of the above step S303, and the embodiments of the present application will not elaborate on this.

[0113] S604, obtain the second timestamp index text of the second video frame; encode the second timestamp index text into a second timestamp projection vector through a text encoder; map the second timestamp projection vector to the dimensional space of the second visual features through a fully connected network, and superimpose the mapping result of the second timestamp projection vector and the second visual features to obtain the second target visual features, which are the basis for the image connector to extract the second video features.

[0114] In the embodiments of the present application, continue to refer to Figure 2 , for each second video frame in the second video, obtain the timing information of the second video frame in the second video, and add the timing information of the second video frame to the second key visual features corresponding to the second video frame through a time perception layer, so as to obtain the second target visual features.

[0115] It should be noted that the implementation process of step S604 in the embodiments of the present application can refer to the implementation processes of the above steps S501 to S503, and the embodiments of the present application will not elaborate on this.

[0116] The image connector compresses and extracts the second target visual features corresponding to each second video frame to obtain the video features (i.e., the second video features) of the second video.

[0117] The video connector compresses and extracts the video features of the second video to map the video features of the second video to the feature space of the language model. The video features after mapping of the second video and the token features corresponding to the second prompt are input into the language model, and the language model outputs the video understanding result of the second video and the second prompt (i.e., the second video understanding result).

[0118] S605, Obtain the second video understanding result output by the video understanding large model based on the second video features and the second prompt.

[0119] S606, Taking the target video understanding result as the target, calculate the loss function value between the second video understanding result and the target video understanding result.

[0120] S607, When the current training end condition is not met, adjust the network parameters of the fully connected network according to the loss function value, enter the next training, and return to execute step S601.

[0121] In the embodiments of the present application, if the number of training times does not reach the upper limit and the loss function value does not meet the convergence condition, the network parameters of the fully connected network can be adjusted according to the loss function value of this training, enter the next training, and return to execute step S601.

[0122] S608, When the current training end condition is met, end the training of the fully connected network.

[0123] In the embodiments of the present application, if the number of training times reaches the upper limit or the loss function value meets the convergence condition, end the training of the fully connected network.

[0124] Continuing to take the video understanding large model Qwen-VL-7B as an example, after being improved by the present application, the visual encoder CLIP-ViT-L / 14 is used to encode video frames into visual features; the weight assignment function is run to assign different weights to the key features and redundant features in the visual features and suppress the redundant features with high similarity; timestamp projection is used to encode each frame and fuse it with the visual features to obtain visual features containing temporal information;

[0125] The method of Merging adjacent tokens is used to compress and extract the visual features containing temporal information to obtain video features. Since a large number of redundant features in the visual features are suppressed, these features will be compressed, so that the visual features contain a large number of effective features, and since the visual features contain temporal information, the video features contain long-term memory and complete content logic information.

[0126] Through the above description, a video understanding method provided by an embodiment of the present application can solve the problems of loss of effective features and confusion of temporal information when a video understanding large model analyzes a video, and can better understand the picture content, plot logic, character relationships, action details, etc.

[0127] The above introduced a video understanding method provided by an embodiment of the present application. Next, the device for executing the above video understanding method will be introduced.

[0128] See Figure 7 , Figure 7 which is a schematic structural diagram of a video understanding device provided by an embodiment of the present application. As Figure 7 shown, a video understanding device provided by an embodiment of the present application is applied to a first device. A video understanding large model is installed in the first device. A weight distribution function and a time perception layer are deployed between the visual encoder and the image connector in the video understanding large model. The device includes:

[0129] A task receiving module 701, configured to receive a video understanding task input by a second device, where the video understanding task includes a first video and a first prompt word;

[0130] A video understanding module 702, configured to obtain first visual features of a first video frame in the first video through a visual encoder; run a weight distribution function to assign different weights to key features and redundant features in the first visual features to obtain first key visual features; add the temporal information of the first video frame to the first key visual features through a time perception layer to obtain first target visual features, where the first target visual features are the basis for an image connector to extract first video features; and return a first video understanding result to the second device, where the first video understanding result is output by the video understanding large model based on the first video features and the first prompt word.

[0131] In a possible implementation, the video understanding module 702 for running a weight distribution function to assign different weights to key features and redundant features in the first visual features to obtain first key visual features is specifically configured to:

[0132] Through a conversion operation of a two-dimensional feature matrix, divide the first visual features into multiple feature blocks; for two consecutive visual features in terms of time points in the first visual features, calculate the similarity between two feature blocks with the same spatial position, and assign a target weight to the feature block with a later time point according to the similarity, where the sum of the target weight and the similarity is 1; perform weighted processing on the first visual features according to the target weights corresponding to each feature block in the first visual features to obtain first key visual features.

[0133] In a possible implementation, the video understanding module 702 for calculating the similarity between two feature blocks with the same spatial position is specifically configured to:

[0134] Calculate the Euclidean distance between two feature blocks with the same spatial position; calculate the similarity using the Euclidean distance, and the similarity is negatively correlated with the Euclidean distance.

[0135] In a possible implementation, the time perception layer includes a fully connected network and a text encoder in a multi-modal model. The fully connected network is pre-trained and is used to add the temporal information of the first video frame to the first key visual feature through the time perception layer to obtain the video understanding module 702 of the first target visual feature, and is specifically used for:

[0136] Obtain the first timestamp index text of the first video frame; encode the first timestamp index text into a first timestamp projection vector through the text encoder; map the first timestamp projection vector to the dimension space of the first visual feature through the fully connected network, and superimpose the mapping result of the first timestamp projection vector and the first visual feature to obtain the first target visual feature.

[0137] In a possible implementation, the process of pre-training the fully connected network by the video understanding module 702 includes:

[0138] Obtain the second visual feature of the second video frame in the second video through the visual encoder; run the weight assignment function to assign different weights to the key features and redundant features in the second visual feature to obtain the second key visual feature; obtain the second timestamp index text of the second video frame; encode the second timestamp index text into a second timestamp projection vector through the text encoder; map the second timestamp projection vector to the dimension space of the second visual feature through the fully connected network, and superimpose the mapping result of the second timestamp projection vector and the second visual feature to obtain the second target visual feature, which is the basis for the image connector to extract the second video feature; obtain the second video understanding result output by the video understanding large model based on the second video feature and the second prompt word; take the target video understanding result as the target, calculate the loss function value between the second video understanding result and the target video understanding result; in the case that the current training end condition is not satisfied, adjust the network parameters of the fully connected network according to the loss function value, enter the next training, and return to execute to obtain the data sample of this training; in the case that the current training end condition is satisfied, end the training of the fully connected network.

[0139] It should be noted that for the refined functions of each module in the embodiments of the present application, reference can be made to the corresponding disclosed parts in the embodiments of the above video understanding method, which will not be elaborated here.

[0140] See Figure 8 , Figure 8 is another process schematic diagram of a video understanding method provided by the embodiments of the present application. As Figure 8As shown in the figure, a video understanding method provided by an embodiment of the present application is applied to a second device, which may include steps S801 to S802:

[0141] S801. Input a video understanding task to the first device, where the video understanding task includes a first video and a first prompt word.

[0142] S802. Receive the first video understanding result returned by the first device. A video understanding large model is installed in the first device. A weight distribution function and a time perception layer are deployed between the visual encoder and the image connector in the video understanding large model. The visual encoder is used to obtain the first visual features of the first video frames in the first video. When the weight distribution function runs, it can assign different weights to the key features and redundant features in the first visual features to obtain the first key visual features. The time perception layer is used to add the timing information of the first video frames to the first key visual features to obtain the first target visual features. The first target visual features are the basis for the image connector to extract the first video features. The first video understanding result is output based on the first video features and the first prompt word.

[0143] It should be noted that for the specific implementation of the video understanding method provided by the embodiment of the present application, reference can be made to Figure 3 the publicly disclosed part of the corresponding video understanding method shown in the figure, which will not be elaborated here.

[0144] Refer to Figure 9 , Figure 9 which is another structural schematic diagram of a video understanding device provided by an embodiment of the present application. As Figure 9 shown in the figure, a video understanding device provided by an embodiment of the present application is applied to a second device. The device includes:

[0145] A task input module 901, which is used to input a video understanding task to the first device, where the video understanding task includes a first video and a first prompt word.

[0146] A result receiving module 902, which is used to receive the first video understanding result returned by the first device. A video understanding large model is installed in the first device. A weight distribution function and a time perception layer are deployed between the visual encoder and the image connector in the video understanding large model. The visual encoder is used to obtain the first visual features of the first video frames in the first video. When the weight distribution function runs, it can assign different weights to the key features and redundant features in the first visual features to obtain the first key visual features. The time perception layer is used to add the timing information of the first video frames to the first key visual features to obtain the first target visual features. The first target visual features are the basis for the image connector to extract the first video features. The first video understanding result is output by the video understanding large model based on the first video features and the first prompt word.

[0147] An embodiment of the present application also provides an electronic device. Refer to Figure 10 , Figure 10 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device in the embodiment of the present application may include, but is not limited to, fixed terminals such as mobile phones, laptop computers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), desktop computers, and the like. Figure 10 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiment of the present application.

[0148] As Figure 10 shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage device 1008 into the random access memory (RAM) 1003. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 1003. The processing device 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. The input / output (I / O) interface 1005 is also connected to the bus 1004.

[0149] Generally, the following devices may be connected to the I / O interface 1005: an input device 1006 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1007 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1008 including, for example, a memory card, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 10 the electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0150] An embodiment of the present application also provides a computer program product including computer-readable instructions. When the computer-readable instructions run on the electronic device, the electronic device implements any one of the video understanding methods provided by the embodiment of the present application.

[0151] An embodiment of the present application also provides a computer-readable storage medium. The storage medium carries one or more computer programs. When the one or more computer programs are executed by the electronic device, the electronic device can implement any one of the video understanding methods provided by the embodiment of the present application.

[0152] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationships between the modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.

[0153] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits or dedicated circuits. However, for this application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.

[0154] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0155] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center by wired means (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless means (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a training device or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

Claims

1. A video understanding method, characterized in that, The method is applied to a first device, in which a large video understanding model is installed. A weight distribution function and a time perception layer are deployed between the visual encoder and the image connector in the large video understanding model. The method includes: Receiving a video understanding task input by a second device, where the video understanding task includes a first video and a first prompt; Obtaining a first visual feature of a first video frame in the first video through the visual encoder; Running the weight distribution function to assign different weights to the key features and redundant features in the first visual feature to obtain a first key visual feature; Adding the timing information of the first video frame to the first key visual feature through the time perception layer to obtain a first target visual feature, where the first target visual feature is the basis for the image connector to extract the first video feature; Returning a first video understanding result to the second device, where the first video understanding result is output by the large video understanding model based on the first video feature and the first prompt; Among them, the running of the weight distribution function to assign different weights to the key features and redundant features in the first visual feature to obtain a first key visual feature includes: Dividing the first visual feature into multiple feature blocks through a conversion operation of a two-dimensional feature matrix; For two consecutive visual features in terms of time points in the first visual feature, calculating the similarity between two feature blocks with the same spatial position, and assigning a target weight to the feature block with a later time point according to the similarity, where the sum of the target weight and the similarity is 1; Performing weighted processing on the first visual feature according to the target weights corresponding to the respective feature blocks in the first visual feature to obtain the first key visual feature.

2. The video understanding method according to claim 1, wherein The calculating the similarity between two feature blocks with the same spatial position includes: Calculating the Euclidean distance between two feature blocks with the same spatial position; Calculating the similarity using the Euclidean distance, where the similarity is negatively correlated with the Euclidean distance.

3. The video understanding method according to claim 1, characterized in that The time perception layer includes a fully connected network and a text encoder in a multimodal model. The fully connected network is pre-trained. The adding the timing information of the first video frame to the first key visual feature through the time perception layer to obtain a first target visual feature includes: Obtaining a first timestamp index text of the first video frame; Encoding the first timestamp index text into a first timestamp projection vector through the text encoder; Mapping the first timestamp projection vector to the dimensional space of the first visual feature through the fully connected network, and superimposing the mapping result of the first timestamp projection vector and the first visual feature to obtain the first target visual feature.

4. The video understanding method according to claim 3, wherein The process of pre-training the fully connected network includes: Obtaining a data sample for this training, where the data sample includes a second video, a second prompt, and an annotated target video understanding result; Obtaining a second visual feature of a second video frame in the second video through the visual encoder; Run the weight assignment function to assign different weights to the key features and redundant features in the second visual feature, so as to obtain the second key visual feature; Obtain the second timestamp index text of the second video frame; encode the second timestamp index text into a second timestamp projection vector through the text encoder; map the second timestamp projection vector to the dimensional space of the second visual feature through the fully connected network, and superimpose the mapping result of the second timestamp projection vector with the second visual feature to obtain a second target visual feature, and the second target visual feature is the basis for the image connector to extract the second video feature; Obtain the second video understanding result output by the video understanding large model based on the second video feature and the second prompt word; Taking the target video understanding result as the target, calculate the loss function value between the second video understanding result and the target video understanding result; When the current training end condition is not satisfied, adjust the network parameters of the fully connected network according to the loss function value, enter the next training, and return to execute the acquisition of the data sample for this training; When the current training end condition is satisfied, end the training of the fully connected network.

5. A video understanding method, characterized in that, The method is applied to a second device, and the method includes: Input a video understanding task to a first device, where the video understanding task includes a first video and a first prompt word; Receive the first video understanding result returned by the first device. The first device is installed with a video understanding large model. A weight assignment function and a time perception layer are deployed between the visual encoder and the image connector in the video understanding large model. The visual encoder is used to obtain the first visual feature of the first video frame in the first video. When the weight assignment function runs, it can assign different weights to the key features and redundant features in the first visual feature to obtain the first key visual feature. The time perception layer is used to add the timing information of the first video frame to the first key visual feature to obtain a first target visual feature. The first target visual feature is the basis for the image connector to extract the first video feature. The first video understanding result is output based on the first video feature and the first prompt word; Among them, when the weight assignment function runs, it can assign different weights to the key features and redundant features in the first visual feature to obtain the first key visual feature, including: When the weight assignment function runs, through the conversion operation of the two-dimensional feature matrix, the first visual feature is divided into multiple feature blocks. For two visual features with continuous time points in the first visual feature, calculate the similarity between two feature blocks with the same spatial position, and assign a target weight to the feature block with a later time point according to the similarity. The sum of the target weight and the similarity is 1. According to the target weights corresponding to each feature block in the first visual feature, perform weighted processing on the first visual feature to obtain the first key visual feature.

6. A video understanding device, characterized in that, The device is applied to a first device, in which a large video understanding model is installed. A weight distribution function and a time perception layer are deployed between the visual encoder and the image connector in the large video understanding model. The device includes: A task receiving module, configured to receive a video understanding task input by a second device, where the video understanding task includes a first video and a first prompt; A video understanding module, configured to obtain a first visual feature of a first video frame in the first video through the visual encoder; run the weight distribution function to assign different weights to key features and redundant features in the first visual feature to obtain a first key visual feature; add the timing information of the first video frame to the first key visual feature through the time perception layer to obtain a first target visual feature, where the first target visual feature is the basis for the image connector to extract the first video feature; and return a first video understanding result to the second device, where the first video understanding result is output by the large video understanding model based on the first video feature and the first prompt; Wherein, the step of running the weight distribution function to assign different weights to key features and redundant features in the first visual feature to obtain a first key visual feature includes: Through a conversion operation of a two-dimensional feature matrix, the first visual feature is divided into multiple feature blocks. For two consecutive visual features in terms of time points in the first visual feature, the similarity between two feature blocks with the same spatial position is calculated, and a target weight is assigned to the feature block with a later time point according to the similarity. The sum of the target weight and the similarity is 1. According to the target weights corresponding to the feature blocks in the first visual feature, the first visual feature is weighted to obtain the first key visual feature.

7. A video understanding device, characterized in that, The device further includes: A task input module, configured to input a video understanding task to the first device, where the video understanding task includes a first video and a first prompt; A result receiving module, configured to receive the first video understanding result returned by the first device. A large video understanding model is installed in the first device. A weight distribution function and a time perception layer are deployed between the visual encoder and the image connector in the large video understanding model. The visual encoder is configured to obtain a first visual feature of a first video frame in the first video. When the weight distribution function runs, it can assign different weights to key features and redundant features in the first visual feature to obtain a first key visual feature. The time perception layer is configured to add the timing information of the first video frame to the first key visual feature to obtain a first target visual feature, where the first target visual feature is the basis for the image connector to extract the first video feature, and the first video understanding result is output by the large video understanding model based on the first video feature and the first prompt; Wherein, when the weight distribution function runs, it can assign different weights to key features and redundant features in the first visual feature to obtain a first key visual feature, including: When the weight assignment function runs, through the conversion operation of the two-dimensional feature matrix, the first visual feature is divided into multiple feature blocks. For two consecutive visual features in terms of time points in the first visual feature, the similarity between two feature blocks with the same spatial position is calculated, and a target weight is assigned to the feature block with a later time point according to the similarity. The sum of the target weight and the similarity is 1. According to the target weights corresponding to the respective feature blocks in the first visual feature, the first visual feature is weighted to obtain the first key visual feature.

8. A computer program product, characterized in that, It includes computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement the video understanding method according to any one of claims 1 to 5.

9. An electronic device, characterized in that, It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store a computer program; The processor is used to execute the computer program so that the electronic device can implement the video understanding method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Video description method and device and storage medium

    CN115205746A

  • Video language understanding method, device and equipment and readable storage medium

    CN117765450A

  • Video understanding multi-task processing model training method based on large language model

    CN119360262A