Video understanding method and related device
By introducing weight allocation function and time-aware layer into the video understanding big model, the problems of key features loss and confusion in timing information in video understanding are solved, and a more accurate and consistent video understanding effect is achieved.
Patent Information
- Application Number
- CN202510511550.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The existing video understanding model will process all video frames indiscriminately when processing videos, resulting in the loss of key features and the retention of a large number of redundant features, affecting the video understanding effect, and easily confusing timing information when comprehension of long videos.
The weight allocation function and time-aware layer are introduced in the video understanding big model. The weight allocation function assigns different weights to key features and redundant features in visual features. The time-aware layer adds timing information to the visual features to improve the video understanding effect.
By assigning different weights, the influence of redundant features is suppressed and key features is retained, thereby improving the video understanding effect; at the same time, by injecting timing information, the problem of timing information chaos during video understanding is corrected, and the accuracy and consistency of video understanding is improved.
Smart Images

Figure CN120047777A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to a video understanding method and related devices. Background Art
[0002] The current large video understanding models (such as Qwen-VL, VideoLLaMA and VideoGPT) are suitable for short video understanding. When processing videos, since all video frames are processed indiscriminately, a large number of compressed visual features will lead to the loss of key features and retain a large number of redundant features, which will seriously affect the video understanding effect. At the same time, due to the limited memory capacity of the model, it is easy to cause confusion in timing information when understanding videos, especially long videos. Summary of the invention
[0003] In view of the above problems, the present application provides a video understanding method and related devices to solve the problem that the large video understanding model has poor video understanding effect and generates confusion of timing information. The specific scheme is as follows:
[0004] In a first aspect, the present application provides a video understanding method, which is applied to a first device, in which a large video understanding model is installed, and a weight distribution function and a time perception layer are deployed between a visual encoder and an image connector in the large video understanding model. The method includes:
[0005] Receiving a video understanding task input by a second device, wherein the video understanding task includes a first video and a first prompt word;
[0006] Acquire, by the visual encoder, a first visual feature of a first video frame in the first video;
[0007] Running the weight assignment function to assign different weights to the key features and redundant features in the first visual features to obtain a first key visual feature;
[0008] adding the timing information of the first video frame to the first key visual feature through the time perception layer to obtain a first target visual feature, wherein the first target visual feature is a basis for the image connector to extract the first video feature;
[0009] A first video understanding result is returned to the second device, where the first video understanding result is output by the video understanding model based on the first video feature and the first prompt word.
[0010] In a possible implementation, the running the weight assignment function to assign different weights to the key features and redundant features in the first visual features to obtain the first key visual features includes:
[0011] Dividing the first visual feature into a plurality of feature blocks through a conversion operation of a two-dimensional feature matrix;
[0012] For two visual features in the first visual feature that are continuous in time, calculate the similarity between two feature blocks in the same spatial position, and assign a target weight to the feature block with a later time point according to the similarity, where the sum of the target weight and the similarity is 1;
[0013] The first key visual feature is obtained by performing weighted processing on the first visual feature according to the target weight corresponding to each feature block in the first visual feature.
[0014] In a possible implementation, the calculating the similarity between two feature blocks at the same spatial position includes:
[0015] Calculate the Euclidean distance between two feature blocks with the same spatial position;
[0016] The similarity is calculated using the Euclidean distance, and the similarity is negatively correlated with the Euclidean distance.
[0017] In a possible implementation, the time perception layer includes a fully connected network and a text encoder in a multimodal model, the fully connected network is pre-trained, and the adding of the timing information of the first video frame to the first key visual feature through the time perception layer to obtain the first target visual feature includes:
[0018] Obtain a first timestamp index text of the first video frame;
[0019] Encoding the first timestamp index text into a first timestamp projection vector by the text encoder;
[0020] The first timestamp projection vector is mapped to the dimensional space of the first visual feature through the fully connected network, and the mapping result of the first timestamp projection vector is superimposed with the first visual feature to obtain the first target visual feature.
[0021] In a possible implementation, the process of pre-training to obtain the fully connected network includes:
[0022] Obtaining a data sample for this training, wherein the data sample includes the second video, the second prompt word, and annotated target video understanding result;
[0023] Acquire, by the visual encoder, a second visual feature of a second video frame in the second video;
[0024] Running the weight assignment function to assign different weights to the key features and redundant features in the second visual features to obtain a second key visual feature;
[0025] Acquire a second timestamp index text of the second video frame; encode the second timestamp index text into a second timestamp projection vector through the text encoder; map the second timestamp projection vector to the dimensional space of the second visual feature through the fully connected network, and superimpose the mapping result of the second timestamp projection vector with the second visual feature to obtain a second target visual feature, wherein the second target visual feature is the basis for the image connector to extract the second video feature;
[0026] Obtain a second video understanding result output by the video understanding large model based on the second video feature and the second prompt word;
[0027] Taking the target video understanding result as a target, calculating a loss function value between the second video understanding result and the target video understanding result;
[0028] If the training end condition is not met at present, the network parameters of the fully connected network are adjusted according to the loss function value, the next training is started, and the step of obtaining the data samples for this training is returned to be executed;
[0029] When the training end condition is currently met, the training of the fully connected network is ended.
[0030] A second aspect of the present application provides a video understanding method, which is applied to a second device, and includes:
[0031] Inputting a video understanding task into a first device, wherein the video understanding task includes a first video and a first prompt word;
[0032] Receive a first video understanding result returned by the first device, wherein a large video understanding model is installed in the first device, and a weight allocation function and a time perception layer are deployed between a visual encoder and an image connector in the large video understanding model, wherein the visual encoder is used to obtain a first visual feature of a first video frame in the first video, and when the weight allocation function is running, different weights can be allocated to key features and redundant features in the first visual features to obtain a first key visual feature, and the time perception layer is used to add timing information of the first video frame to the first key visual feature to obtain a first target visual feature, and the first target visual feature is the basis for the image connector to extract the first video feature, and the first video understanding result is output based on the first video feature and the first prompt word.
[0033] According to a third aspect of the present application, a video understanding device is provided. The device is applied to a first device. A large video understanding model is installed in the first device. A weight distribution function and a time perception layer are deployed between a visual encoder and an image connector in the large video understanding model. The device includes:
[0034] A task receiving module, used for receiving a video understanding task input by a second device, wherein the video understanding task includes a first video and a first prompt word;
[0035] A video understanding module is used to obtain a first visual feature of a first video frame in the first video through the visual encoder; run the weight allocation function to assign different weights to key features and redundant features in the first visual feature to obtain a first key visual feature; add the timing information of the first video frame to the first key visual feature through the time perception layer to obtain a first target visual feature, the first target visual feature is the basis for the image connector to extract the first video feature; return a first video understanding result to the second device, the first video understanding result is output by the video understanding large model based on the first video feature and the first prompt word.
[0036] A fourth aspect of the present application provides a video understanding device, the device comprising:
[0037] A task input module, used to input a video understanding task to the first device, wherein the video understanding task includes a first video and a first prompt word;
[0038] A result receiving module is used to receive a first video understanding result returned by the first device. A large video understanding model is installed in the first device. A weight allocation function and a time perception layer are deployed between the visual encoder and the image connector in the large video understanding model. The visual encoder is used to obtain a first visual feature of a first video frame in the first video. When the weight allocation function is running, different weights can be allocated to key features and redundant features in the first visual features to obtain a first key visual feature. The time perception layer is used to add the timing information of the first video frame to the first key visual feature to obtain a first target visual feature. The first target visual feature is the basis for the image connector to extract the first video feature. The first video understanding result is output by the large video understanding model based on the first video feature and the first prompt word.
[0039] A fifth aspect of the present application provides a computer program product, including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements the video understanding method of the first aspect or any implementation of the first aspect.
[0040] A sixth aspect of the present application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:
[0041] The memory is used to store computer programs;
[0042] The processor is used to execute the computer program so that the electronic device can implement the video understanding method of the first aspect or any implementation manner of the first aspect.
[0043] The seventh aspect of the present application provides a computer storage medium, which carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement the video understanding method of the first aspect or any implementation method of the first aspect.
[0044] By means of the above technical scheme, a video understanding method and related device provided by the present application are applied to a first device, a video understanding large model is installed in the first device, and a weight distribution function and a time perception layer are deployed between the visual encoder and the image connector in the video understanding large model. The method includes receiving a video understanding task input by a second device, the video understanding task includes a first video and a first prompt word; obtaining a first visual feature of a first video frame in the first video through a visual encoder; running a weight distribution function, assigning different weights to key features and redundant features in the first visual feature to obtain a first key visual feature; adding the timing information of the first video frame to the first key visual feature through a time perception layer to obtain a first target visual feature, the first target visual feature is the basis for the image connector to extract the first video feature; returning the first video understanding result to the second device, the first video understanding result is output by the video understanding large model based on the first video feature and the first prompt word. The present application can assign different weights to key features and redundant features in visual features, and inject timing information, so as to solve the problem of effective information loss and timing information confusion when the video understanding large model analyzes the video, and improve the video understanding ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.
[0046] Figure 1 A schematic diagram of the structure of the existing large model for video understanding;
[0047] Figure 2 A schematic diagram of the structure of a large video understanding model provided in an embodiment of the present application;
[0048] Figure 3 A flowchart of a video understanding method provided in an embodiment of the present application;
[0049] Figure 4 A partial flow chart of a video understanding method provided in an embodiment of the present application;
[0050] Figure 5 A schematic diagram of another part of the flow chart of a video understanding method provided in an embodiment of the present application;
[0051] Figure 6 A schematic diagram of another part of the flow chart of a video understanding method provided in an embodiment of the present application;
[0052] Figure 7 A schematic diagram of the structure of a video understanding device provided in an embodiment of the present application;
[0053] Figure 8 Another flowchart of a video understanding method provided in an embodiment of the present application;
[0054] Fig. 9 Another structural schematic diagram of a video understanding device provided in an embodiment of the present application;
[0055] Fig.10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0056] The following describes the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. The terms used in the implementation method section of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.
[0057] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0058] The terms "first", "second" etc. in the specification of the application and the above-mentioned drawings are used to distinguish similar objects, and need not be used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable in appropriate circumstances, and this is only to describe the distinction mode adopted by the objects of the same attributes when describing in the embodiments of the application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0059] The video understanding model can be used to develop intelligent editing templates, build intelligent review systems, and generate short play scripts, which is conducive to reducing economic costs, improving production efficiency, and enhancing business results. Figure 1 , Figure 1 This is a schematic diagram of the structure of the existing video understanding model. Figure 1 As shown, the existing video understanding model consists of a visual encoder, an image connector, a video connector, a word segmenter and a language model, wherein the word segmenter obtains word segmentation features from the prompt word, the visual encoder encodes the video frame into visual features, the image connector compresses and extracts the visual features of all video frames to obtain video features, the video connector compresses and extracts the video features to map the video features to the feature space of the language model, the mapped video features and word segmentation features are input into the language model, and the language model outputs the video understanding results.
[0060] Taking the video understanding large model Qwen-VL-7B as an example, it uses the visual encoder CLIP-ViT-L / 14 to encode video frames into visual features, and then uses the method of merging adjacent tokens to compress and extract the visual features to obtain video features, and then uses 3D convolution to map the video features to the feature space of the large language model, and finally outputs the video understanding results through the Qwen2-7B language model.
[0061] When existing large video understanding models extract video features from visual features, they will compress the visual features in large quantities, and the weights of the visual features of different video frames are consistent during compression. However, there are only a few key features in the video, and indiscriminate compression will lead to the loss of key features and retain a large number of redundant features. This problem will be more serious in long videos. At the same time, the model's memory capacity is limited, and it is easy to cause confusion in timing information when understanding the video.
[0062] In order to solve the above problems, an embodiment of the present application provides a video understanding method. A video understanding method of an embodiment of the present application is described in detail below in conjunction with the accompanying drawings.
[0063] The embodiment of the present application provides a video understanding method, which is applied to a first device, in which a large video understanding model is installed, and a weight distribution function and a time perception layer are deployed between the visual encoder and the image connector in the large video understanding model. Figure 2 , Figure 2 A schematic diagram of the structure of a large video understanding model provided in an embodiment of the present application. Figure 2As shown, the video understanding model provided by the embodiment of the present application is constructed based on the existing video understanding model, including a visual encoder, an image connector, a video connector, a word segmenter and a language model, and a weight distribution function and a time perception layer are deployed between the visual encoder and the image connector. For the visual features output by the visual encoder, the weight distribution function can assign low weights to the redundant features and high weights to the key features to solve the problem that the key features are lost and a large number of redundant features are retained when the visual features are compressed indiscriminately. In addition, the time perception layer can inject timing information into the visual features to solve the problem of confusion in timing information that is easily generated during video understanding.
[0064] See also Figure 3 , Figure 3 A flowchart of a video understanding method provided in an embodiment of the present application is shown below. Figure 3 As shown, a video understanding method provided by an embodiment of the present application may include steps S301 to S305, and these steps are described in detail below.
[0065] S301, receiving a video comprehension task input by a second device, where the video comprehension task includes a first video and a first prompt word. In the embodiment of the present application, the second device is a device held by a user who has a video comprehension requirement, and the second device inputs the video comprehension task to the first device in response to the user's input operation. In response to this, the first device obtains the video comprehension task of the first device, and parses it to obtain the video to be understood (i.e., the first video) and the prompt word (i.e., the first prompt word).
[0066] S302: Acquire a first visual feature of a first video frame in a first video through a visual encoder.
[0067] In the present application, see Figure 2 , input the first video into the visual encoder, and input the first prompt word into the word segmenter. The word segmenter can obtain the corresponding word segmentation features from the first prompt word, and the visual encoder encodes each video frame in the first video (i.e., the first video frame) into a visual feature (i.e., the first visual feature).
[0068] S303 , running a weight assignment function to assign different weights to the key features and redundant features in the first visual feature to obtain the first key visual feature.
[0069] In the present application, see Figure 2 For each first video frame in the first video, a weight function is run to assign different weights to the key features and redundant features in its first visual features, wherein the key features are assigned high weights and the redundant features are assigned low weights, so as to perform weighted processing on the first visual features to obtain the first key visual features.
[0070] S304, adding the timing information of the first video frame to the first key visual feature through the time perception layer to obtain a first target visual feature, where the first target visual feature is the basis for the image connector to extract the first video feature.
[0071] In the present application, see Figure 2 For each first video frame in the first video, the timing information of the first video frame in the first video is obtained, and the timing information of the first video frame is added to the first key visual feature corresponding to the first video frame through the time perception layer, so as to obtain the first target visual feature.
[0072] The image connector compresses and extracts the first target visual features corresponding to each first video frame to obtain the video features of the first video (ie, the first video features).
[0073] The video connector compresses and extracts the first video features, and maps the first video features to the feature space of the language model. The mapped first video features and the word segmentation features corresponding to the first prompt word are input into the language model, and the language model outputs the video understanding result of the first video features and the first prompt word (i.e., the first video understanding result).
[0074] S305, returning the first video understanding result to the second device, where the first video understanding result is output by the video understanding big model based on the first video feature and the first prompt word.
[0075] In the embodiment of the present application, after obtaining the first video understanding result output by the video understanding large model, the first device can return the first video understanding result to the second device, and the second device will display the first video understanding result to the user.
[0076] In a possible implementation, weights can be assigned to key features and redundant features in visual features through feature analysis. Figure 4 , Figure 4 A partial flow chart of a video understanding method provided in an embodiment of the present application. Figure 4 As shown, a video understanding method provided by an embodiment of the present application, wherein step S303 "runs a weight assignment function to assign different weights to key features and redundant features in the first visual feature to obtain a first key visual feature" may include steps S401 to S403, and these steps are described in detail below.
[0077] S401, dividing the first visual feature into a plurality of feature blocks through a conversion operation of a two-dimensional feature matrix.
[0078] In the embodiment of the present application, for each first video frame in the first video, its first visual feature can be converted into a two-dimensional feature matrix, and the size of the two-dimensional feature matrix is fixed to ,in, is the number of rows of the two-dimensional feature matrix, is the number of columns of the two-dimensional feature matrix. For the first video of the first video frame, we can obtain its three-dimensional visual feature matrix, the size of which is , the three-dimensional visual feature matrix contains all visual features.
[0079] For each first video frame in the first video, the corresponding two-dimensional feature matrix can be divided into multiple feature blocks, and the size of each feature block is ,in, are the three dimensions of three-dimensional space, where and is the size of the spatial dimension, is the size of the time dimension.
[0080] S402, for two visual features in the first visual feature that are continuous in time, calculate the similarity between two feature blocks in the same spatial position, and assign a target weight to the feature block with a later time point according to the similarity, where the sum of the target weight and the similarity is 1.
[0081] In the embodiment of the present application, for two first video frames with consecutive time points in the first video, the two first visual features corresponding to the two first visual features are two visual features with consecutive time points, assuming that they are visual feature 1 and visual feature 2. The two feature blocks with the same spatial position in visual feature 1 and visual feature 2 can be regarded as two feature blocks with consistent spatial dimensions and adjacent time dimensions, such as the feature block in visual feature 1. The feature blocks in visual feature 2 , for example, the feature block in visual feature 1 The feature blocks in visual feature 2 .
[0082] For two feature blocks with the same spatial dimension and adjacent time dimension, the similarity between them can be calculated, and a weight (i.e., target weight) is assigned to the feature block with the later time point. The target weight is equal to the difference between 1 and the similarity. With feature blocks For example, assuming that the similarity between the two is , then it can be a feature block Assigning target weights .
[0083] It should be noted that if the similarity Greater than the preset threshold , which shows that the feature block With feature blocks In addition, when calculating the similarity between two feature blocks, the original elements in the feature blocks are calculated instead of the weighted elements.
[0084] S403: Perform weighted processing on the first visual feature according to the target weight corresponding to each feature block in the first visual feature to obtain a first key visual feature.
[0085] In the embodiment of the present application, continue to take visual feature 2 as an example, assuming that its feature block The corresponding target weight is weight 1, feature block The corresponding target weight is weight 2. Then according to weight 1, visual feature 2 is The elements in the feature block are weighted. Each element in is multiplied by weight 1, and in addition, visual feature 2 is multiplied by weight 2 in feature block The elements in the feature block are weighted. Each element in is multiplied by weight 2 to obtain visual feature 2. After the weighted calculation of all feature blocks is completed, the key visual feature corresponding to visual feature 2 (ie, the first key visual feature) can be obtained.
[0086] Based on this, redundant features with high similarity in visual features can be suppressed, and effective key features will be retained in the three-dimensional visual feature matrix of the video.
[0087] In a possible implementation, the similarity between two feature blocks can be calculated by Euclidean distance. In this regard, an embodiment of the present application provides a video understanding method, wherein in step S402, "calculating the similarity between two feature blocks with the same spatial position" may include the following steps:
[0088] Calculate the Euclidean distance between two feature blocks with the same spatial position; use the Euclidean distance to calculate the similarity, and the similarity is negatively correlated with the Euclidean distance.
[0089] In the embodiment of the present application, for two visual features with consecutive time points in the first visual feature, the following formula (1) can be used to calculate two feature blocks with the same spatial position ( and ) :
[0090] (1)
[0091] Then, the similarity is calculated using the following formula (2): :
[0092] (2)
[0093] In one possible implementation, the temporal perception layer includes a fully connected network and a text encoder in a multimodal model. The fully connected network is pre-trained. It should be noted that the text encoder in the multimodal model (such as CLIP, ALIGN, FLAVA, etc.) is used because the temporal information needs to be fused with the visual features, and the encoder trained with pure text cannot be used. The text encoder trained in multimodal can have a similar feature representation space as the visual encoder.
[0094] See also Figure 5 , Figure 5 Another partial flow chart of a video understanding method provided in an embodiment of the present application. Figure 5 As shown, a video understanding method provided by an embodiment of the present application, wherein in step S304 "adding the timing information of the first video frame to the first key visual feature through the time perception layer to obtain the first target visual feature" may include steps S501 to S503, and these steps are described in detail below.
[0095] S501, obtaining a first timestamp index text of a first video frame.
[0096] In an embodiment of the present application, for each first video frame in the first video, a timestamp index text (i.e., first timestamp index text) of the first video frame can be obtained according to the timing of the first video frame in the first video, and the first timestamp index text is "first frame", "second frame", ..., "Nth frame".
[0097] S502: Encode the first timestamp index text into a first timestamp projection vector through a text encoder.
[0098] In the embodiment of the present application, for each first video frame in the first video, the first timestamp index text corresponding to the first video frame can be encoded into a timestamp projection vector (ie, a first timestamp projection vector) through a text encoder.
[0099] S503: Map the first timestamp projection vector to the dimensional space of the first visual feature through a fully connected network, and superimpose the mapping result of the first timestamp projection vector and the first visual feature to obtain a first target visual feature.
[0100] In an embodiment of the present application, for each first video frame in the first video, the first timestamp projection vector corresponding to the first video frame can be mapped to the dimensional space of the first visual feature through a fully connected network, maintaining the same dimension as the first visual feature, and then the mapping result of the first timestamp projection vector is added to the first visual feature by main element to obtain a visual feature containing timing information (i.e., the first target visual feature).
[0101] The process of timestamp projection can be expressed by the following formula (3):
[0102] (3)
[0103] in, represents the first target visual feature, Indicates the first visual feature, represents the forward computation of the fully connected network, Represents the encoding operation of a text encoder, Represents the first timestamp indexed text.
[0104] It should be noted that, in the process of timestamp projection, the first visual feature is a one-dimensional vector. In addition, the fully connected network can be single-layer or multi-layer, and the dimension of the first timestamp projection vector is mapped to be consistent with the dimension of the first visual feature, and then element-by-element addition is performed to achieve the fusion of time series information and time features.
[0105] See also Figure 6 , Figure 6 This is another partial flow chart of a video understanding method provided in an embodiment of the present application. Figure 6 As shown, an embodiment of the present application provides a video understanding method, wherein the process of pre-training to obtain a fully connected network may include steps S601 to S608, and these steps are described in detail below.
[0106] S601, obtaining a data sample for this training, where the data sample includes a second video, a second prompt word, and an annotated target video understanding result.
[0107] In an embodiment of the present application, a data sample required for local training is obtained, and the data sample includes a video for training (i.e., the second video), a prompt word (i.e., the second prompt word), and a pre-labeled video understanding result (i.e., the target video understanding result).
[0108] S602: Acquire a second visual feature of a second video frame in a second video through a visual encoder.
[0109] In the present application, see Figure 2, input the second video into the visual encoder, and input the second prompt word into the word segmenter. The word segmenter can obtain the corresponding word segmentation features from the second prompt word, and the visual encoder encodes each video frame in the second video (i.e., the second video frame) into a visual feature (i.e., the second visual feature).
[0110] S603, running a weight allocation function to allocate different weights to the key features and redundant features in the second visual features to obtain the second key visual features.
[0111] In the present application, see Figure 2 For each second video frame in the second video, a weight function is run to assign different weights to the key features and redundant features in its second visual features, wherein the key features are assigned high weights and the redundant features are assigned low weights, so as to perform weighted processing on the second visual features to obtain the second key visual features.
[0112] It should be noted that the implementation process of step S603 in the embodiment of the present application can refer to the implementation process of the above-mentioned step S303, and the embodiment of the present application will not be repeated here.
[0113] S604, obtaining a second timestamp index text of a second video frame; encoding the second timestamp index text into a second timestamp projection vector through a text encoder; mapping the second timestamp projection vector to a dimensional space of a second visual feature through a fully connected network, and superimposing the mapping result of the second timestamp projection vector with the second visual feature to obtain a second target visual feature, which is the basis for the image connector to extract the second video feature.
[0114] In the present application, see Figure 2 For each second video frame in the second video, the timing information of the second video frame in the second video is obtained, and the timing information of the second video frame is added to the second key visual feature corresponding to the second video frame through the time perception layer, so as to obtain the second target visual feature.
[0115] It should be noted that the implementation process of step S604 in the embodiment of the present application can refer to the implementation process of the above-mentioned steps S501 to S503, and the embodiment of the present application will not be repeated here.
[0116] The image connector compresses and extracts the second target visual features corresponding to each second video frame to obtain the video features of the second video (ie, the second video features).
[0117] The video connector compresses and extracts the video features of the second video to map the video features of the second video to the feature space of the language model. The mapped video features of the second video and the word segmentation features corresponding to the second prompt word are input into the language model, and the language model outputs the video understanding result of the second video and the second prompt word (i.e., the second video understanding result).
[0118] S605, obtaining a second video understanding result output by the video understanding large model based on the second video feature and the second prompt word.
[0119] S606, taking the target video understanding result as the target, calculating the loss function value between the second video understanding result and the target video understanding result.
[0120] S607, when the training end condition is not met at present, the network parameters of the fully connected network are adjusted according to the loss function value, the next training is started, and the execution of step S601 is returned.
[0121] In an embodiment of the present application, if the number of training times does not reach the upper limit and the loss function value does not meet the convergence condition, the network parameters of the fully connected network can be adjusted according to the loss function value of this training, and the next training can be entered, returning to execute step S601.
[0122] S608, when the training end condition is currently met, end the training of the fully connected network.
[0123] In the embodiment of the present application, if the number of training times reaches the upper limit or the loss function value meets the convergence condition, the training of the fully connected network is terminated.
[0124] Continuing with the video understanding large model Qwen-VL-7B as an example, after being improved by this application, the visual encoder CLIP-ViT-L / 14 is used to encode the video frame into visual features; the weight allocation function is run to assign different weights to the key features and redundant features in the visual features, and the redundant features with high similarity are suppressed; the timestamp projection is used to encode frame by frame and merge with the visual features to obtain the visual features containing time sequence information;
[0125] The method of merging adjacent tokens is used to compress and extract visual features containing temporal information to obtain video features. Since a large number of redundant features in visual features are suppressed, these features will be compressed, so that the visual features contain a large number of effective features. In addition, since the visual features contain temporal information, the video features contain long memory and complete content logic information.
[0126] Through the above description, a video understanding method provided in an embodiment of the present application can solve the problems of loss of effective features and confusion of timing information when analyzing videos by large video understanding models, and can better understand the picture content, plot logic, character relationships, action details, etc.
[0127] A video understanding method provided by an embodiment of the present application is introduced above, and a device for executing the above video understanding method will be introduced below.
[0128] See also Figure 7 , Figure 7 A schematic diagram of the structure of a video understanding device provided in an embodiment of the present application. Figure 7 As shown, an embodiment of the present application provides a video understanding device, which is applied to a first device, in which a large video understanding model is installed, and a weight distribution function and a time perception layer are deployed between a visual encoder and an image connector in the large video understanding model. The device includes:
[0129] A task receiving module 701 is used to receive a video understanding task input by a second device, where the video understanding task includes a first video and a first prompt word;
[0130] The video understanding module 702 is used to obtain the first visual feature of the first video frame in the first video through the visual encoder; run the weight allocation function to assign different weights to the key features and redundant features in the first visual feature to obtain the first key visual feature; add the timing information of the first video frame to the first key visual feature through the time perception layer to obtain the first target visual feature, and the first target visual feature is the basis for the image connector to extract the first video feature; return the first video understanding result to the second device, and the first video understanding result is output by the video understanding large model based on the first video feature and the first prompt word.
[0131] In a possible implementation, the video understanding module 702 for running a weight assignment function to assign different weights to key features and redundant features in the first visual feature to obtain the first key visual feature is specifically configured to:
[0132] Through the conversion operation of the two-dimensional feature matrix, the first visual feature is divided into multiple feature blocks; for two visual features with continuous time points in the first visual feature, the similarity between the two feature blocks with the same spatial position is calculated, and a target weight is assigned to the feature block with a later time point according to the similarity, and the sum of the target weight and the similarity is 1; according to the target weight corresponding to each feature block in the first visual feature, the first visual feature is weighted to obtain the first key visual feature.
[0133] In a possible implementation, the video understanding module 702 for calculating the similarity between two feature blocks with the same spatial position is specifically configured to:
[0134] Calculate the Euclidean distance between two feature blocks with the same spatial position; use the Euclidean distance to calculate the similarity, and the similarity is negatively correlated with the Euclidean distance.
[0135] In a possible implementation, the time perception layer includes a fully connected network and a text encoder in a multimodal model, the fully connected network is pre-trained, and is used to add the timing information of the first video frame to the first key visual feature through the time perception layer to obtain the video understanding module 702 of the first target visual feature, specifically for:
[0136] Acquire a first timestamp index text of a first video frame; encode the first timestamp index text into a first timestamp projection vector through a text encoder; map the first timestamp projection vector to the dimensional space of a first visual feature through a fully connected network, and superimpose the mapping result of the first timestamp projection vector and the first visual feature to obtain a first target visual feature.
[0137] In a possible implementation, the process of pre-training the video understanding module 702 to obtain a fully connected network includes:
[0138] The second visual feature of the second video frame in the second video is obtained through the visual encoder; the weight allocation function is run to assign different weights to the key features and redundant features in the second visual feature to obtain the second key visual feature; the second timestamp index text of the second video frame is obtained; the second timestamp index text is encoded into a second timestamp projection vector through the text encoder; the second timestamp projection vector is mapped to the dimensional space of the second visual feature through the fully connected network, and the mapping result of the second timestamp projection vector is superimposed with the second visual feature to obtain the second target visual feature, which is the basis for the image connector to extract the second video feature; the second video understanding result output by the video understanding large model based on the second video feature and the second prompt word is obtained; with the target video understanding result as the target, the loss function value between the second video understanding result and the target video understanding result is calculated; when the training end condition is not met at present, the network parameters of the fully connected network are adjusted according to the loss function value, the next training is entered, and the data sample of this training is obtained by returning to execute; when the training end condition is met at present, the training of the fully connected network is ended.
[0139] It should be noted that the detailed functions of each module in the embodiment of the present application can be found in the corresponding public part of the above-mentioned video understanding method embodiment, and will not be repeated here.
[0140] See also Figure 8 , Figure 8 Another flowchart of a video understanding method provided in an embodiment of the present application is shown below. Figure 8As shown, a video understanding method provided in an embodiment of the present application is applied to a second device, which may include steps S801 to S802:
[0141] S801, input a video understanding task into a first device, where the video understanding task includes a first video and a first prompt word.
[0142] S802, receiving a first video understanding result returned by a first device, wherein a large video understanding model is installed in the first device, and a weight allocation function and a time perception layer are deployed between the visual encoder and the image connector in the large video understanding model, wherein the visual encoder is used to obtain a first visual feature of a first video frame in a first video, and when the weight allocation function is running, different weights can be allocated to key features and redundant features in the first visual feature to obtain a first key visual feature, and the time perception layer is used to add timing information of the first video frame to the first key visual feature to obtain a first target visual feature, and the first target visual feature is the basis for the image connector to extract the first video feature, and the first video understanding result is output based on the first video feature and the first prompt word.
[0143] It should be noted that the specific implementation of a video understanding method provided in the embodiment of the present application can be found in Figure 3 The video understanding method shown corresponds to the public part and will not be repeated here.
[0144] See also Fig. 9 , Fig. 9 Another structural diagram of a video understanding device provided in an embodiment of the present application. Fig. 9 As shown, an embodiment of the present application provides a video understanding device, which is applied to a second device and includes:
[0145] The task input module 901 is used to input a video understanding task to the first device, where the video understanding task includes a first video and a first prompt word.
[0146] The result receiving module 902 is used to receive the first video understanding result returned by the first device. The first device is installed with a large video understanding model. A weight allocation function and a time perception layer are deployed between the visual encoder and the image connector in the large video understanding model. The visual encoder is used to obtain the first visual feature of the first video frame in the first video. When the weight allocation function is running, different weights can be assigned to the key features and redundant features in the first visual features to obtain the first key visual feature. The time perception layer is used to add the timing information of the first video frame to the first key visual feature to obtain the first target visual feature. The first target visual feature is the basis for the image connector to extract the first video feature. The first video understanding result is output by the large video understanding model based on the first video feature and the first prompt word.
[0147] The present application also provides an electronic device in an embodiment. Fig.10 , Fig.10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device in the embodiment of the present application may include but is not limited to fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Fig.10 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0148] like Fig.10 As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1008 to a random access memory (RAM) 1003. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 1003. The processing device 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0149] Typically, the following devices may be connected to the I / O interface 1005: an input device 1006 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1007 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1008 including, for example, a memory card, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Fig.10 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0150] An embodiment of the present application also provides a computer program product including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any one of the video understanding methods provided in the embodiments of the present application.
[0151] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any video understanding method provided in the embodiment of the present application.
[0152] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.
[0153] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0154] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0155] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a training device, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device, or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.
Claims
1. A video understanding method, characterized in that: The method is applied to a first device, in which a large video understanding model is installed, and a weight distribution function and a time perception layer are deployed between a visual encoder and an image connector in the large video understanding model. The method includes: Receiving a video understanding task input by a second device, wherein the video understanding task includes a first video and a first prompt word; Acquire, by the visual encoder, a first visual feature of a first video frame in the first video; Running the weight assignment function to assign different weights to the key features and redundant features in the first visual features to obtain a first key visual feature; adding the timing information of the first video frame to the first key visual feature through the time perception layer to obtain a first target visual feature, wherein the first target visual feature is a basis for the image connector to extract the first video feature; A first video understanding result is returned to the second device, where the first video understanding result is output by the video understanding model based on the first video feature and the first prompt word.
2. The video understanding method according to claim 1, characterized in that: The step of running the weight assignment function to assign different weights to the key features and redundant features in the first visual features to obtain the first key visual features includes: Dividing the first visual feature into a plurality of feature blocks through a conversion operation of a two-dimensional feature matrix; For two visual features in the first visual feature that are continuous in time, calculate the similarity between two feature blocks in the same spatial position, and assign a target weight to the feature block with a later time point according to the similarity, where the sum of the target weight and the similarity is 1; The first key visual feature is obtained by performing weighted processing on the first visual feature according to the target weight corresponding to each feature block in the first visual feature.
3. The video understanding method according to claim 2, characterized in that: The calculating the similarity between two feature blocks with the same spatial position includes: Calculate the Euclidean distance between two feature blocks with the same spatial position; The similarity is calculated using the Euclidean distance, and the similarity is negatively correlated with the Euclidean distance.
4. The video understanding method according to claim 1, characterized in that The time perception layer includes a fully connected network and a text encoder in a multimodal model, the fully connected network is pre-trained, and the time sequence information of the first video frame is added to the first key visual feature through the time perception layer to obtain a first target visual feature, including: Obtain a first timestamp index text of the first video frame; Encoding the first timestamp index text into a first timestamp projection vector by the text encoder; The first timestamp projection vector is mapped to the dimensional space of the first visual feature through the fully connected network, and the mapping result of the first timestamp projection vector is superimposed with the first visual feature to obtain the first target visual feature.
5. The video understanding method according to claim 4, characterized in that: The process of pre-training to obtain the fully connected network includes: Obtaining a data sample for this training, wherein the data sample includes the second video, the second prompt word, and annotated target video understanding result; Acquire, by the visual encoder, a second visual feature of a second video frame in the second video; Running the weight assignment function to assign different weights to the key features and redundant features in the second visual features to obtain a second key visual feature; Acquire a second timestamp index text of the second video frame; encode the second timestamp index text into a second timestamp projection vector through the text encoder; map the second timestamp projection vector to the dimensional space of the second visual feature through the fully connected network, and superimpose the mapping result of the second timestamp projection vector with the second visual feature to obtain a second target visual feature, wherein the second target visual feature is the basis for the image connector to extract the second video feature; Obtain a second video understanding result output by the video understanding large model based on the second video feature and the second prompt word; Taking the target video understanding result as a target, calculating a loss function value between the second video understanding result and the target video understanding result; If the training end condition is not met at present, the network parameters of the fully connected network are adjusted according to the loss function value, the next training is started, and the step of obtaining the data samples for this training is returned to be executed; When the training end condition is currently met, the training of the fully connected network is ended.
6. A video understanding method, characterized in that: The method is applied to a second device, and the method includes: Inputting a video understanding task into a first device, wherein the video understanding task includes a first video and a first prompt word; Receive a first video understanding result returned by the first device, wherein a large video understanding model is installed in the first device, and a weight allocation function and a time perception layer are deployed between a visual encoder and an image connector in the large video understanding model, wherein the visual encoder is used to obtain a first visual feature of a first video frame in the first video, and when the weight allocation function is running, different weights can be allocated to key features and redundant features in the first visual features to obtain a first key visual feature, and the time perception layer is used to add timing information of the first video frame to the first key visual feature to obtain a first target visual feature, and the first target visual feature is the basis for the image connector to extract the first video feature, and the first video understanding result is output based on the first video feature and the first prompt word.
7. A video understanding device, characterized in that: The apparatus is applied to a first device, a large video understanding model is installed in the first device, a weight distribution function and a time perception layer are deployed between a visual encoder and an image connector in the large video understanding model, and the apparatus includes: A task receiving module, used to receive a video understanding task input by a second device, wherein the video understanding task includes a first video and a first prompt word; A video understanding module is used to obtain a first visual feature of a first video frame in the first video through the visual encoder; run the weight allocation function to assign different weights to key features and redundant features in the first visual feature to obtain a first key visual feature; add the timing information of the first video frame to the first key visual feature through the time perception layer to obtain a first target visual feature, the first target visual feature is the basis for the image connector to extract the first video feature; return a first video understanding result to the second device, the first video understanding result is output by the video understanding large model based on the first video feature and the first prompt word.
8. A video understanding device, characterized in that: The device comprises: A task input module, used to input a video understanding task to the first device, wherein the video understanding task includes a first video and a first prompt word; A result receiving module is used to receive a first video understanding result returned by the first device. A large video understanding model is installed in the first device. A weight allocation function and a time perception layer are deployed between the visual encoder and the image connector in the large video understanding model. The visual encoder is used to obtain a first visual feature of a first video frame in the first video. When the weight allocation function is running, different weights can be allocated to key features and redundant features in the first visual features to obtain a first key visual feature. The time perception layer is used to add the timing information of the first video frame to the first key visual feature to obtain a first target visual feature. The first target visual feature is the basis for the image connector to extract the first video feature. The first video understanding result is output by the large video understanding model based on the first video feature and the first prompt word.
9. A computer program product, characterized in that The method comprises computer-readable instructions, which, when executed on an electronic device, enable the electronic device to implement the video understanding method as claimed in any one of claims 1 to 6.
10. An electronic device, characterized in that: The method comprises at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the video understanding method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Video description method and device and storage medium
CN115205746A
Video language understanding method, device and equipment and readable storage medium
CN117765450A
Video understanding multi-task processing model training method based on large language model
CN119360262A
Video processing method and apparatus
US20220327835A1
Video retrieval method based on attention segment prompt
WO2024001057A1
Cited By
Data processing method, electronic device, storage medium and computer program product
CN121279458A
Frame sequence processing method and device and electronic equipment
CN122160510A