Multi-modal multi-feature end-to-end dense video description method and system

Through the multimodal multi-feature end-to-end dense video description method, combined with visual, audio and text features, the problems of insufficient modal fusion and single features in the prior art are solved, and deeper understanding of video content and more accurate description generation are achieved.

CN120088704APending Publication Date: 2025-06-03HUAZHONG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510196090.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Existing intensive video description methods are difficult to effectively integrate visual, audio and text modalities, resulting in the impact of description accuracy and coherence. At the same time, existing models use only a single feature in visual feature extraction, limiting the depth of video content understanding and the quality of description generation.

Method used

The multimodal multi-feature end-to-end dense video description method is adopted to extract keyframes through frame processing, and multimodal feature fusion is combined with visual, audio and text features. The ViFi-CLIP model and other pre-trained models are used to extract and fusion features to achieve a more comprehensive and accurate understanding and description of video content.

Benefits of technology

It significantly improves the depth of video content understanding and the quality of description generation. Through the effective fusion of multi-feature fusion and multi-modal data, the quality and accuracy of video subtitle generation are improved, and the application of pre-trained models in video description tasks is promoted.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088704A_ABST
    Figure CN120088704A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal multi-feature end-to-end dense video description method and system. The method comprises the following steps: extracting a key frame of a video; a visual area feature and a visual grid feature are extracted from the key frame, the key frame is input into a ViFi-CLIP model to extract a CLIP feature, the visual area feature and the visual grid feature are mapped to the CLIP feature, and two new features obtained through mapping are fused to obtain a visual feature; extracting audio features and text features; respectively inputting the visual features, the audio features and the text features into a coding and decoding module for coding and decoding, and then carrying out feature fusion to obtain fused features; and respectively inputting the fusion features into an event boundary prediction module, a text description prediction module and an event counter to obtain event boundaries in the video, natural language description corresponding to each event and the number of events. According to the method, the video content understanding depth and the description generation quality can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of dense video description in computer vision, and more specifically, relates to a multi-modal multi-feature end-to-end dense video description method and system. Background Art

[0002] In recent years, with the rapid growth of video data and the rapid development of deep learning technology, video understanding and analysis have become one of the hotspots in the field of computer vision. In video understanding, dense video description, as an important task, aims to detect all events in an uncropped video and generate corresponding text descriptions for each event, which can present complex visual information in an easily understandable text form. This not only helps to improve the understanding of video content but also provides important support for applications such as video content search, automatic annotation, and video summary generation.

[0003] Visual, audio, and text information in videos are often interrelated, and single-modal information is difficult to comprehensively reflect the complexity of video content. Audio information provides strong support for the environmental background and the timing of events, while text information helps to strengthen the understanding of video semantics. However, existing models usually only focus on visual features and cannot effectively fuse audio and text modalities, resulting in affected accuracy and coherence of descriptions. At the same time, in the extraction of visual features, most current models only use single features. Therefore, the depth of video content understanding and the quality of description generation of most current models need to be improved.

[0004] In addition, existing two-stage video description methods of first localizing and then describing have the problem of being unable to perform end-to-end training, relying on manually set anchor mechanisms and post-processing steps, which limits the overall performance of the model.

[0005] Therefore, there is an urgent need for an end-to-end dense video description method to improve the depth of video content understanding and the quality of description generation. Summary of the Invention

[0006] In view of the above-mentioned defects or improvement requirements of the prior art, the present invention provides a multi-modal multi-feature end-to-end dense video description method and system, which can improve the depth of video content understanding and the quality of description generation.

[0007] To achieve the above object, according to one aspect of the present invention, there is provided a multi-modal multi-feature end-to-end dense video description method, including the steps of:

[0008] Performing frame division on the video and extracting key frames of the video;

[0009] Input the key frames into the visual feature extraction module, which is used to extract visual region features and visual grid features from the key frames. Input the key frames into the ViFi-CLIP model to extract CLIP features, map the visual region features and visual grid features to the CLIP features respectively, and fuse the two new features obtained by mapping to obtain visual features;

[0010] Input the video into the audio feature extraction module and the text feature extraction module respectively to extract audio features and text features;

[0011] Input the visual features, audio features and text features into the encoding and decoding module for encoding and decoding respectively, and input the decoded visual features, audio features and text features into the multi-modal feature fusion module for feature fusion to obtain fused features;

[0012] Input the fused features into the event boundary prediction module, the text description prediction module and the event counter respectively to obtain the event boundaries in the video, the natural language description corresponding to each event and the number of events.

[0013] Preferably, the step of frame-dividing the video and extracting the key frames of the video includes: evenly dividing the video into multiple segments, then randomly sampling in each segment to extract multiple key frames.

[0014] Preferably, use the TSN model or the Swin Transformer model to extract visual grid features from the key frames. If the TSN model is used, splice the optical flow features and RGB features output by the TSN model as the visual grid features, and use the TransVOD model to extract visual region features from the key frames.

[0015] Preferably, the step of extracting audio features includes: extracting audio data from the video and inputting the audio data into the VGGish model to extract audio features.

[0016] Preferably, the step of extracting text features includes:

[0017] Denote the key frames of the video as F input ={f 1 ,f 2 ,...,f N}, where N is the number of key frames. Use automatic speech recognition technology to obtain the language text descriptions corresponding to the key frames, denoted as {t 1 ,t 2 ,...,t N}. Merge the speech text descriptions of the video into a text paragraph, denoted as T, and use the BERT pre-trained model to extract features from the text paragraph T as the text features.

[0018] Preferably, the multimodal feature fusion module includes a first self-attention module, a second self-attention module, a third self-attention module, a first cross-attention module, a second cross-attention module, a third cross-attention module, a first feed-forward neural network layer, a second feed-forward neural network layer, a third feed-forward neural network layer, and a fourth feed-forward neural network layer. Inputting the decoded visual features, audio features, and text features into the multimodal feature fusion module for feature fusion includes the steps:

[0019] Denote the decoded visual features as X j 、Denote the decoded audio features as Y j , Denote the decoded text features as Z j , Input X j into the first self-attention module to obtain feature X j ′, Input Y j into the second self-attention module to obtain feature Y j ′, Input Z j into the first self-attention module to obtain feature Z j ′, Input X j ′, Y j ′ and Z j ′ into the first cross-attention module, the second cross-attention module, and the third cross-attention module for pairwise cross-attention calculations, and process them through the first feed-forward neural network layer, the second feed-forward neural network layer, and the third feed-forward neural network layer. The three features represent each other pairwise, considering the order of representation, to obtain six interaction features. Concatenate the six interaction features and input them into the fourth feed-forward neural network layer for dimensionality reduction to generate fused features.

[0020] Preferably, take the visual feature extraction module, the audio feature extraction module, the text feature extraction module, the encoding module, the decoding module, the multimodal feature fusion module, the event boundary prediction module, the text description prediction module, and the event counter as an overall dense video description model for the first-stage training. The first-stage training includes the steps:

[0021] Construct a training sample set, and each training sample is labeled with an event boundary label, a text description label, and an event number label;

[0022] Input the training samples into the dense video description model, and output the event boundary prediction result, the text description prediction result, and the event number prediction result;

[0023] Calculate the intersection over union loss between the event boundary label and the event boundary prediction result, calculate the description loss representing the difference between the text description label and the text description prediction result, calculate the counting loss representing the difference between the event number label and the event number prediction result, calculate the classification loss, and calculate the loss function. The calculation formula is:

[0024] L = β giou L giou + β cls L cls + β ec L ec + β cap L cap

[0025] Among them, L represents the loss function, L giou represents the intersection over union loss, L cls represents the classification loss, L ec represents the counting loss, L cap represents the description loss, β giou β cls β ec β cap all represent coefficients.

[0026] Preferably, after the training of the first stage is completed, the training of the second stage is carried out. In the training of the second stage, greedy decoding and soft sampling are respectively used to generate description sentences. The calculation of the description loss L cap includes the steps:

[0027] Calculate the METEOR and CIDEr scores of the text description prediction results generated by greedy decoding and soft sampling respectively, and take the difference between the two scores as the reward. The calculation formula is:

[0028] gen_scores = MS(gen, gt) × MW + CS(gen, gt) × CW

[0029] greedy_scores = MS(greedy, gt) × MW + CS(greedy, gt) × CW

[0030] reward = gen_scores - greedy_scores

[0031] Among them, gen_scores represents the score of the text description prediction result generated by soft sampling, greedy_scores represents the score of the text description prediction result generated by using greedy decoding, gen represents the text description prediction result generated by using soft sampling, gt represents the text description label, greedy represents the text description prediction result generated by using greedy decoding, MW represents the weight assigned to the METEOR score, CW represents the weight assigned to the CIDEr score, MS represents the function for calculating the METEOR score, CS represents the function for calculating the CIDEr score, and reward represents the reward;

[0032] Calculate the description loss L according to the reward lr, the calculation formula is:

[0033] L lr = -sample_logprobs × reward × mask

[0034] Among them, sample_logprobs represents the log probability when the text description prediction module samples from the probability distribution of all possible output words, and mask is a binary mask;

[0035] Based on L lr Calculate the description loss L cap L cap = L lr .mean(L lr ), mean(L lr ) represents calculating the average value for each column of L lr .

[0036] According to another aspect of the present invention, a multi-modal multi-feature end-to-end dense video description system is provided, including:

[0037] A key frame extraction module for frame processing the video and extracting the key frames of the video;

[0038] A visual feature extraction module for extracting visual region features and visual grid features from the key frames, inputting the key frames into the ViFi-CLIP model to extract CLIP features, mapping the visual region features and visual grid features to CLIP features respectively, and fusing the two new features obtained by mapping to obtain visual features;

[0039] An audio feature extraction module for extracting audio features from the video;

[0040] A text feature extraction module for extracting text features from the video;

[0041] An encoding and decoding module for encoding and decoding visual features, audio features, and text features;

[0042] A multi-modal feature fusion module for inputting the decoded visual features, audio features, and text features for feature fusion to obtain fusion features;

[0043] A prediction module for predicting event boundaries, the natural language description corresponding to each event, and the number of events in the video according to the fusion features.

[0044] Generally speaking, compared with the prior art, the above technical solution conceived by the present invention can improve the depth of video content understanding and the quality of description generation. Specifically, it is reflected in the following aspects:

[0045] (1) Improve video understanding ability by complementing multiple features

[0046] This multi-feature fusion method not only makes up for the deficiencies of single features, but also utilizes the complementarity between features, significantly enhancing the model's ability to capture video content, helping to extract more comprehensive and rich visual features, and thus improving the quality and accuracy of dense video caption generation.

[0047] (2) Enrich the fusion method of multi-modal data

[0048] By designing a model that uses the fusion of grid features and region features as visual features and takes three features, namely visual features, audio features, and text features, as inputs, a more comprehensive and accurate description and understanding of dense videos are achieved, enhancing the diversity and accuracy of video content expression, and providing new ideas and technical means for multi-modal data fusion.

[0049] (3) Promote the application of pre-trained models in video description

[0050] By introducing pre-trained models such as ViFi-CLIP into the video description task, their powerful generalization ability and semantic understanding ability are fully utilized. The introduction of pre-trained models not only improves the performance of the model, but also provides reference and inspiration for the application of pre-trained models in other multi-modal tasks.

[0051] (4) Achieve more efficient end-to-end training

[0052] Training is carried out in an end-to-end manner, which simplifies the model development process and improves performance. At the same time, it better captures the complex relationships and context information between data, enabling the model to have better generalization ability and adaptability.

[0053] (5) Practical application value

[0054] Dense video description technology has broad application prospects in fields such as video surveillance, video content retrieval, video question answering, and intelligent security. By generating accurate and coherent natural language descriptions, the intelligence level of these applications can be significantly improved to meet people's needs for video information processing. Brief description of the drawings

[0055] Figure 1 is the flowchart of the multi-modal multi-feature end-to-end dense video description method of the embodiment of the present invention;

[0056] Figure 2 is the network schematic diagram of the dense video description model of the embodiment of the present invention;

[0057] Figure 3 is the schematic diagram of the mapping module of the embodiment of the present invention;

[0058] Figure 4 It is a schematic diagram of the multi-modal feature fusion module according to an embodiment of the present invention;

[0059] Figure 5 It is an application example of the multi-modal multi-feature end-to-end dense video description method according to an embodiment of the present invention. Detailed implementation manners

[0060] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0061] In the description of the embodiments of the present application, the terms "first", "second", "third", and "fourth" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", "third", and "fourth" may explicitly or implicitly include at least one of such features. The meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0062] In the embodiments of the present invention, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product, or equipment that includes a series of steps or modules does not necessarily have to be limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products, or equipment.

[0063] In the embodiments of the present invention, the naming or numbering of steps does not mean that the steps in the method flow must be executed in the time / logical order indicated by the naming or numbering. The named or numbered process steps can be changed in the execution order according to the technical objectives to be achieved, as long as the same or similar technical effects can be achieved.

[0064] Referring to "embodiments" in the embodiments of the present invention means that the specific features, structures, or characteristics described in connection with the embodiments may be included in at least one embodiment of the present invention. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described in the embodiments of the present invention can be combined with other embodiments.

[0065] The present invention provides a multi-modal multi-feature end-to-end dense video description method and system, which will be described separately below.

[0066] As Figure 1 shown, a multi-modal multi-feature end-to-end dense video description method according to an embodiment of the present invention includes the steps of:

[0067] S1, perform frame splitting on the video and extract key frames of the video.

[0068] In this step, first, frame splitting is performed on each video in the dataset. By extracting key frames in the video, it is ensured that the temporal information of the video can be captured. Each frame represents a time point in the video, and subsequently, these frames will be used as the input for subsequent processing. Generally, in order to reduce the computational load and improve the processing efficiency, video frames can be sampled at a certain interval instead of processing every frame of the entire video. For a given video V, in the embodiment of the present invention, the video is first evenly segmented. Then, random sampling is performed in each segment, so that N key frames F input ={f 1 , f 2 ,..., f N} are obtained from the video.

[0069] S2, input the key frames into a visual feature extraction module. The visual feature extraction module is used to extract visual region features and visual grid features from the key frames, input the key frames into the ViFi-CLIP model to extract CLIP features, map the visual region features and visual grid features to CLIP features respectively, and fuse the two new features obtained by mapping to obtain visual features.

[0070] Although grid features can retain background information and fine-grained global content, they are insufficient in capturing specific objects; in contrast, region features can accurately locate target objects but lack an understanding of the environmental context. Both are crucial for generating accurate video descriptions. Therefore, the use of a single feature leads to limitations in description generation, and the multi-feature fusion of combining grid features and region features can improve video understanding performance. At the same time, with the rise of pre-trained vision-language models (such as CLIP), through training with large-scale image-text pairs, these models demonstrate excellent generalization ability and zero-shot performance, especially performing well in the case of unlabeled data. However, directly applying these pre-trained models to the video description task still faces challenges, especially when dealing with complex dynamics and object relationships. ViFi-CLIP effectively improves the accuracy of video understanding by combining frame-level image feature processing and text matching.

[0071] S3, input the video into an audio feature extraction module and a text feature extraction module respectively to extract audio features and text features.

[0072] A visual feature extraction module, an audio feature extraction module, and a text feature extraction module can be integrated into a pre-trained dense video description model to extract features from multiple modalities.

[0073] In the dense video description task, using a grid feature pre-trained model for feature extraction usually causes two problems. First, these pre-trained models may not be able to encode the relationship between objects in the video and image / scene-level information well, resulting in the loss of key information. Second, since the pre-trained model is frozen, the conditional relationship between the calculated features and the input video is fixed and cannot be jointly optimized with the video description task. This may lead to a decline in feature quality, especially considering that these models are trained on different datasets. To solve these problems, an embodiment of the present invention proposes a method to fuse grid / region features with ViFi-CLIP features through a mapping module, and the model structure is as Figure 3 shown. Introducing the ViFi-CLIP model can supplement information for encoding grid / region features without retraining the feature extraction model.

[0074] The specific implementation is as follows: Map the grid feature X g and the region feature X r to the ViFi-CLIP feature X c respectively to obtain two new features X gc and X rc :

[0075] X gc =projector(X g , X c )

[0076] X rc =projector(X r , X c )

[0077] Then, these two features are fused through a multi-feature fusion method F to obtain the final visual feature X:

[0078] X = F(X gc , X rc )

[0079] This design enables the model to capture information at different spatial scales in the video while obtaining higher-level semantic information, thus obtaining a more comprehensive and rich multi-modal information representation and enhancing the understanding and generalization ability of complex video content. For grid features, a TSN pre-trained model or a Swin Transformer pre-trained model can be used. When using the TSN pre-trained model, two feature matrices can be obtained, namely the optical flow feature F flow =TSN flow (F input ) and the RGB feature F RGB =TSN RGB (F input ). To utilize these two features simultaneously, they can be concatenated into the representation X g =[F flow ,F RGB . When using the Swin Transformer pre-trained model, the grid feature is represented as X g =SwinTransformer(F input ). For region features, the TransVOD pre-trained model can be used for feature extraction X r =TransVOD(F input ). Based on the Transformer architecture, TransVOD combines the self-attention mechanism and the visual attention mechanism, and can effectively handle object detection and tracking tasks in the video. Finally, for CLIP features, the ViFi-CLIP pre-trained model (CLIP with video fine-tuning) can be used for feature extraction X c =ViFi-CLIP(F input ). By fine-tuning CLIP, ViFi-CLIP enables the image-based CLIP model to adapt to video-specific tasks, thereby further improving the model's understanding and application performance of video content.

[0080] For audio information, the VGGish pre-trained model can be used for audio feature extraction. First, the FFmpeg tool is used to extract audio data from the video to generate WAV files, and then these audio data are input into the VGGish pre-trained model. The extracted audio feature is represented as Y = VGGish(F input ).

[0081] For text information, an automatic speech recognition (ASR) system is used to obtain the language text descriptions {t input ={f 1 ,f 2 ,...,f N} corresponding to the manually annotated video event frames 1 ,t 2,...,t N}. To verify the effectiveness of text information in improving the performance of the dense video description model, in the embodiments of the present invention, multiple speech text descriptions of each video are first merged into a text paragraph T = Concat(t 1 ,t 2 ,...,t N ). Then, the BERT pre-trained model is used to extract features Z = BERT(T) from the text paragraph, and this feature is used as the input for the subsequent encoding module. Similar to visual and audio features, text paragraph features are also dynamically extracted according to the time-aligned event frames generated by the dense video description model.

[0082] S4. The visual features, audio features, and text features are respectively input into the encoding and decoding module for encoding and decoding, and the decoded visual features, audio features, and text features are input into the multi-modal feature fusion module for feature fusion to obtain the fused features.

[0083] The network diagram of multi-modal feature fusion is as Figure 4 shown, and the specific implementation is as follows.

[0084] To solve the problem that different modal features generate feature vectors with different dimensions after passing through the DeformableTransformer decoding layer, in the embodiments of the present invention, a multi-modal feature fusion module is introduced, as Figure 4 shown. The design of this module aims to merge the decoded visual, audio, and text features into a unified feature vector, so that the subsequent parallel decoding layer can better adapt to feature vectors with different dimensions. This design solves the problem of inconsistent dimensions of different modal features and realizes the effective parallel training of the model.

[0085] After passing through the DeformableTransformer decoder, the decoded visual features X j , audio features Y j , and text features Z j are obtained. First, through the first, second, and third self-attention modules respectively, X j ′, Y j ′, and Z j ′ are obtained. Then, X j ′, Y j ′, and Z jThey respectively perform pairwise cross-attention calculations through the first, second, and third cross-attention modules, and are processed through the first, second, and third feed-forward neural network (FFN) layers. The pairwise interaction representation between the three features is considered, and six interaction features XY, XZ, YX, YZ, ZX, and ZY are obtained in consideration of the order of representation. Subsequently, these six interaction features are concatenated and then dimensionally reduced through the fourth feed-forward neural network layer to finally generate a fused feature Q that contains all three-modal information. j , which is used as the input for the subsequent parallel decoding layer. The entire calculation process of the multi-modal fusion module is as follows:

[0086] X j ' = SelfAttention(Xj)

[0087] Y j ' = SelfAttention(Yj)

[0088] Z' = SelfAention(Z j )

[0089] XY = FFN(CrossAttention(X' j , Y' j , Y' j ))

[0090] XZ = FFN(CrossAttention(X' j , Z' j , Z' j )

[0091] YX = FFN(CrossAttention(Y' j , X' j , X' j )

[0092] YZ = FFN(CrossAttention(Y' j , Z' j , Z' j ))

[0093] ZX = FFN(CrossAttention(Z' j , X' j , X' j ))

[0094] ZY = FFN(CrossAttention(Z' j , Y' j , Y' j ))

[0095] Q j= FFN(Concat(XY,XZ,YX,YZ,ZX,ZY))

[0096] Among them, SelfAttention and CrossAttention use the multi-head attention mechanism, Concat represents matrix concatenation, and FFN represents the feed-forward neural network.

[0097] The specific implementation of the encoding and decoding module is as follows.

[0098] In the embodiments of the present invention, three Deformable Transformers with the same structure are designed as the encoding and decoding modules of the model, and each module processes visual, audio, and text features. In addition, in order to enable the Deformable Transformer encoder to learn the relative or absolute positions of the three-modal features in the sequence, position encoding must be added as input for each modality. The Deformable Transformer adopts a deformable attention module in the Transformer encoder, replacing the self-attention module, and replacing the cross-attention module in the Transformer decoder. This improvement enables the model to better adapt to complex tasks and data features, thereby improving the overall performance and efficiency.

[0099] The Deformable Transformer is an improved version based on the Transformer architecture. By introducing the multi-scale deformable attention module (MSDeformAttn), it enhances the model's ability to model non-uniform spatial features, enabling it to process data sequences with complex structures.

[0100] The formula of MSDeformAttn is as follows:

[0101]

[0102] Among them, represents the input multi-scale feature map, and p q ∈[0,1] 2 represents the normalized coordinates of the reference point of the query element q. m represents the number of attention heads, l represents the level of the input features, k represents the number of sampling points, Vp mlqk and A mlqk represent the sampling offsets and attention weights of the k sampling points and the l-th layer features of the m-th attention head respectively. Vp mlqk and A mlqk are both obtained by linear projection onto the query element. The function rescales the normalized coordinates onto the l-th layer input feature map.

[0103] S5. Input the fused features into the event boundary prediction module (localization head), text description prediction module (captioning head), and event counter respectively to obtain the event boundaries in the video, the natural language description corresponding to each event, and the number of events.

[0104] The decoding module of the model consists of the Deformable Transformer decoding layer of the above-mentioned encoding and decoding module and three parallel heads. These three parallel heads are responsible for predicting event boundaries (localization head), generating text descriptions (captioning head), and event counting (event counter) respectively. The input of the entire decoding module includes the output of the encoding module and learnable query parameters. Among these parameters, N learnable query parameters are used as the initial guesses of event features and positions, called feature initialization. They are iteratively learned and refined in each decoding layer.

[0105] The following specifically describes the specific implementation of the three parallel heads.

[0106] (1) Localization Head:

[0107] The role of the event localization head is to accurately segment the video into multiple event segments, helping to improve the coherence and readability of the predicted text description. The functions of the event localization head include querying q for each event j for boundary prediction and generating location confidence (classification head). The goal of boundary prediction is to predict the two-dimensional relative offset of the true event position, that is, the center position and length of the event, based on the reference point r j (predicted by linearly projecting q j and using the sigmoid activation function). The goal of the classification head is to generate location confidence for each event query, which is similar to a binary classification task. Both modules are implemented by a multi-layer perceptron (MLP). After passing through the boundary prediction and classification head modules, the embodiments of the present invention obtain the detected event set where represents the start time of the video segment, represents the end time of the video segment, is the location confidence of the event query q j . This information helps to better understand the video content and improve the description quality.

[0108] (2) Captioning Head:

[0109] The description head is used to generate text descriptions. In the embodiments of the present invention, Deformable SoftAttention (DSA) is adopted to construct the description head. SoftAttention is a commonly used method in video description generation, mainly used in a two-stage video description generation framework. In the end-to-end model, the embodiments of the present invention propose an improved DSA method based on Soft Attention. DSA is a widely used model in video description generation, which dynamically determines the importance of each frame when generating words. This design helps to better capture the information of key frames in the video, thereby improving the quality and coherence of the generated text description.

[0110] For the DSA model, the embodiments of the present invention use the Transformer model. The input of the model includes the decoded and fused feature Q from the Deformable Transformer decoder j , and the token sequence (seq j ) of the real description. Based on the output of the Transformer model, word prediction is performed through a linear layer and a softmax layer. Initially, based on Q j , K sampling points r l are generated from each f j , where f l represents the feature output from the l-th layer of the encoder. The embodiments of the present invention use L to represent the scale of the feature map, and use K×L sampling points r as key-value pairs. The embodiments of the present invention use BertSelfAttention to implement the soft attention mechanism, and input seq j and attention_mask j to obtain h j . Then, h j is concatenated with Q j and passed as input to the deformable attention layer. Subsequently, this feature passes through a feed-forward neural network layer (FeedForward Network, FFN), and then through the softmax function to obtain a coherent and meaningful text description. The calculation formula of the entire DSA model is as follows:

[0111] h j = BSAttn(seq j , attention_mask j )

[0112]

[0113] h j = Linear(Concat(h j , Q j ))

[0114]

[0115] Among them, Q j represents the fused feature, seq j represents the token of the true description, attention_mask j represents the mask created based on seq j . represents the feature output by the encoder, BSAttn represents the BertSelfAttention layer, MSDeformAttn represents the deformable attention layer, LayerNormLayer represents the layer normalization, FC represents the feed-forward neural network layer, and Log_SoftMax represents the logarithm of the softmax result.

[0116] (3) Event Counter

[0117] Event Counter: To help control the number of generated events and improve the quality and accuracy of the generated text description, the embodiment of the present invention introduces an event counter. The event counter consists of a max pooling layer and a fully connected layer with a softmax activation function. The fully connected layer compresses the significant information in the event query {q j} into a global feature vector, which is then used to predict a fixed-size vector, where each value represents the possibility of a specific quantity. This design helps the model better regulate the number of generated events, thereby improving the quality and accuracy of the generated text description.

[0118] Figure 2 is the network schematic diagram of the multi-modal multi-feature end-to-end dense video description model based on the ViFi-CLIP pre-trained model implemented by the present invention.

[0119] The model proposed in the embodiment of the present invention adopts an encoder-decoder structure. Both the encoder and the decoder are based on the Deformable Transformer architecture to construct an end-to-end model, as Figure 2 shown. The encoder module takes the visual feature X, the audio feature Y, and the text feature Z as input sources. Among them, the extraction of visual features adopts a multi-feature fusion technology and combines the feature map from the ViFi-CLIP model to assist in feature extraction, aiming to obtain more comprehensive and rich feature information. The feature vectors generated by the three-modal features after passing through the encoding module are directly input into the decoding module. The decoded features are processed by the multi-modal fusion module. To achieve end-to-end training, the model introduces parallel modules, including a localization head, a description head, and an event counter, enabling the event localization and text description tasks to be trained in parallel, and finally generating a set of timestamps and their corresponding descriptions.

[0120] To enhance the stability of reinforcement learning, the embodiments of the present invention perform random cropping through data augmentation techniques to increase the number of data samples. Given a video, the embodiments of the present invention randomly crop a time period with a length of r_crop*T, where T represents the duration of the video, and r_crop~U(0.5,1) is the ratio controlling the length of the time period, and U(a,b) here represents a uniform distribution over the range [a,b]. For each video, the embodiments of the present invention repeatedly sample crop_num = 32 time periods to form a batch of samples for RL training.

[0121] The following specifically describes the training process of the above network.

[0122] During the model training process, the model generates a set of results containing M predicted event positions and their corresponding text descriptions. To calculate the position deviation between the real events and the predicted events in the video, the embodiments of the present invention use the Hungarian algorithm to calculate the matching loss between the two sides.

[0123] The visual feature extraction module, audio feature extraction module, text feature extraction module, encoding module, decoding module, multi-modal feature fusion module, event boundary prediction module, text description prediction module, and event counter are used as an overall dense video description model for the first stage of training. The first stage of training includes the steps:

[0124] Construct a training sample set, and each training sample is labeled with an event boundary label, a text description label, and an event number label;

[0125] Input the training samples into the dense video description model to output event boundary prediction results, text description prediction results, and event number prediction results;

[0126] Calculate the intersection over union loss between the event boundary label and the event boundary prediction result, calculate the description loss representing the difference between the text description label and the text description prediction result, calculate the counting loss representing the difference between the event number label and the event number prediction result, and calculate the loss function.

[0127] The overall loss of the model is the weighted sum of multiple parallel task losses, including the GIoU loss (L giou ), classification loss (L cls ), counting loss (L ec ), and description loss (L cap ), where L cap = L lr .mean(). The specific calculation formula is as follows:

[0128] L = β giou L giou+β cls L cls +β ec L ec +β cap L cap

[0129] Among them, L giou represents the IoU (Intersection over Union) loss between the predicted event and the true event of the event boundary label and the event boundary prediction result, L cls represents the focal loss between the predicted classification score and the true event label, L ec represents the count loss of the difference between the event number label and the event number prediction result, L cap represents the description loss of the difference between the text description label and the text description prediction result. During the training of the first stage, the calculation of L cap can be performed using any existing function for describing differences, such as calculating the similarity to calculate L cap , β giou , β cls , β ec , β cap all represent coefficients.

[0130] Preferably, in order to improve the quality of the generated description, after the training of the first stage, a second stage of training, i.e., reinforcement learning, can also be performed to further optimize the performance of the model.

[0131] In order to better adapt to the dynamic characteristics of videos in the dense video description task and generate more contextually and temporally consistent natural language descriptions, the embodiments of the present invention adopt reinforcement learning (RL) as an additional training stage. The embodiments of the present invention use two methods to generate description sentences. The embodiments of the present invention use two methods to generate description sentences: greedy decoding and soft sampling. Greedy decoding is a simple decoding strategy that selects the word with the highest current conditional probability as the output at each time step. Soft sampling is a more flexible sampling method that controls the diversity of the generated text by adjusting the probability distribution.

[0132] Subsequently, the embodiments of the present invention calculate the METEOR and CIDEr scores of these two methods and use the difference between the two scores as the reward. Among them, METEOR is an index for evaluating the quality of machine translation. It not only considers the exact match of vocabulary but also introduces factors such as synonyms, stems, and word order, thus providing a more comprehensive evaluation. CIDEr is an index for automatically evaluating the performance of the image description task. It mainly evaluates the quality of the image description by calculating the similarity between the generated description and a set of reference descriptions. The calculation formula of the reward is as follows:

[0133] gen_scores = MS(gen, gt) × MW + CS(gen, gt) × CW

[0134] greedy_scores = MS(greedy, gt) × MW + CS(greedy, gt) × CW

[0135] reward = gen_scores - greedy_scores

[0136] Among them, gen_scores represents the score of the predicted result of the text description generated by soft sampling, greedy_scores represents the score of the predicted result of the text description generated by using greedy decoding, gen represents the predicted result of the text description generated by using soft sampling, gt represents the text description label, greedy represents the predicted result of the text description generated by using greedy decoding, MW represents the weight assigned to the METEOR score, CW represents the weight assigned to the CIDEr score, MS represents the function for calculating the METEOR score, CS represents the function for calculating the CIDEr score, and reward represents the reward.

[0137] The embodiment of the present invention uses the Policy Gradient reinforcement learning algorithm to construct a reinforcement learning loss function, which is expressed as follows:

[0138] L lr = -sample_logprobs × reward × mask

[0139] Among them, sample_logprobs represents the log probability when the generation model samples from the probability distribution of all possible output words, and mask is a binary mask, which is usually used to mark the valid positions (i.e., positions greater than 0) in the encoded real description.

[0140] Then, based on L lr calculate the description loss L cap . L cap = L lr .mean(L lr ), mean(L lr ) represents calculating the average value for each column of L lr .

[0141] Then, adopt the same total loss calculation formula as in the above first stage to calculate the total loss of the model.

[0142] Figure 5 is an application example of a multi-modal multi-feature end-to-end dense video description method in the embodiment of the present invention.

[0143] The embodiments of the present invention were experimented on the standard dataset ActivityNet Captioning. The evaluation metrics include BLEU, METEOR, ROUGE, and CIDEr, which are calculated based on the matching pairs between the generated text descriptions and the ground-truth text descriptions, where the tIOU thresholds are set to 0.3 / 0.5 / 0.7 / 0.9. Specifically: BLEU: Evaluates the accuracy and fluency of the generated text. METEOR: Considers word order information. ROUGE: Evaluates content relevance. CIDEr: Measures text diversity and accuracy. tIOU: Focuses on temporal alignment. These metrics jointly consider different aspects of text descriptions and help comprehensively evaluate the performance and effectiveness of the model.

[0144] To verify the effectiveness of introducing ViFi-CLIP features, the embodiments of the present invention used TSN and SwinTransformer as visual feature extraction models and compared the experimental results before and after introducing the ViFi-CLIP model. The experimental results show that there is a certain degree of improvement in various metrics, indicating that the visual features assisted by the ViFi-CLIP model can enhance the model performance. For specific experimental results, please refer to Table 1, where the BLEU metric increased by 5%, and the METEOR, ROUGE-L, and CIDEr metrics all increased by 2%.

[0145] In addition, to verify the impact of the fusion of region features and grid features on the results and find the best feature fusion method, the embodiments of the present invention conducted the following experiments. The specific experimental results are shown in Table 2. The embodiments of the present invention used TSN and SwinTransformer to extract grid features and TransVOD to extract region features. Pairwise fusion experiments were conducted on these three types of features. The results show that fusing SwinTransformer features with TransVOD features and fusing TSN features with SwinTransformer features can achieve better performance.

[0146] Specifically, compared with the model using only SwinTransformer features, the model combining TSN and Swin Transformer as visual features has a 17% increase in the BLEU metric, a 3.5% increase in the CIDEr metric, and a 2% increase in both the METEOR and ROUGE-L metrics. Similarly, compared with the model using only Swin Transformer features, the model combining TransVOD and Swin Transformer as visual features has a 12% increase in the BLEU metric, a 7% increase in the CIDEr metric, a 1.3% increase in the SODA metric, and a 3% increase in the ROUGE-L metric.

[0147] The embodiments of the present invention also observe that not only can the fusion of regional features and grid features achieve good results, but the fusion of two grid features also shows relatively ideal experimental effects. This can be attributed to the following two main reasons:

[0148] (1) Feature diversity: Swin Transformer and TSN extract features based on different architectures and methods. Swin Transformer is based on the Transformer architecture, while TSN is based on Temporal Segment Networks. They may capture feature information at different levels and aspects. Therefore, fusing these features can enrich the information expression of the model and enhance the model's understanding and generalization ability of data.

[0149] (2) Feature complementarity: Swin Transformer and TSN may have their own advantages in capturing information in different aspects of the data. For example, Swin Transformer is better at capturing spatial information and local features in images, while TSN is better at capturing dynamic information in time series. Therefore, the features they extract are complementary, and fusing these features can make comprehensive use of their respective advantages.

[0150] Finally, compared with the current mainstream dense video description models, the model of the embodiments of the present invention still shows excellent performance, and the specific results are shown in Table 3. The embodiments of the present invention use two different feature combinations for experiments: one group uses the grid features of TSN and Swin Transformer and the fusion with ViFi-CLIP as visual features; the other group uses the grid features of Swin Transformer and the regional features of TransVOD and the fusion with ViFi-CLIP as visual features. The audio features are extracted by VGGish, and the text features are extracted by BERT.

[0151] Compared with Vid2Seq, the model of the embodiments of the present invention has improved by 6% in the CIDEr metric, 8% in the BLEU-4 metric, and 3% in the SODA metric. Compared with PDVC (prediction-based candidate boxes), the model of the embodiments of the present invention has improved by 10% in the CIDEr metric, 8% in the BLEU-4 metric, 10% in both the METEOR and SODA metrics, and 6.5% in the ROUGE-L metric. In the experiments based on real candidate boxes, the model of the embodiments of the present invention has improved by 5% in the CIDEr metric, 4% in the BLEU-4 metric, and 5%, 5%, and 3% in the METEOR, SODA, and ROUGE-L metrics respectively. Compared with the GIT model, the model of the embodiments of the present invention has improved by 6% in the METEOR metric and 5% in the SODA metric.

[0152] After introducing reinforcement learning, the model further improved by 3% on the METEOR metric, 1.3% on the CIDEr metric, 10% on the ROUGE-L metric, and 3% on the SODA metric. Overall, compared with mainstream dense video description algorithms, the present invention has achieved significant performance improvements on various key metrics.

[0153] Table 1 Comparison of models with and without ViFi-CLIP features

[0154]

[0155] Table 2 Comparison of dual-feature fusion and single feature

[0156]

[0157]

[0158] Table 3 Comparison with current mainstream dense description generation models

[0159]

[0160] An end-to-end dense video description system with multi-modal and multi-feature according to an embodiment of the present invention includes:

[0161] A key frame extraction module for frame processing the video and extracting key frames of the video;

[0162] A visual feature extraction module for extracting visual region features and visual grid features from the key frames, inputting the key frames into the ViFi-CLIP model to extract CLIP features, respectively mapping the visual region features and visual grid features to the CLIP features, and fusing the two new features obtained by the mapping to obtain visual features;

[0163] An audio feature extraction module for extracting audio features from the video;

[0164] A text feature extraction module for extracting text features from the video;

[0165] An encoding and decoding module for encoding and decoding visual features, audio features, and text features;

[0166] A multi-modal feature fusion module for inputting the decoded visual features, audio features, and text features for feature fusion to obtain fusion features;

[0167] A prediction module for predicting event boundaries in the video, the natural language description corresponding to each event, and the number of events according to the fusion features.

[0168] The working principle and technical effects of the above system are the same as those of the above method, and will not be elaborated here.

[0169] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A multi-modal multi-feature end-to-end dense video description method, characterized in that: Includes steps: Divide the video into frames and extract the key frames of the video; The key frame is input into the visual feature extraction module, which is used to extract the visual area feature and the visual grid feature from the key frame. The key frame is input into the ViFi-CLIP model to extract the CLIP feature, and the visual area feature and the visual grid feature are respectively mapped to the CLIP feature. The two new features obtained by mapping are fused to obtain the visual feature; Input the video into the audio feature extraction module and the text feature extraction module respectively to extract the audio features and the text features; The visual features, audio features and text features are respectively input into the encoding and decoding module for encoding and decoding, and the decoded visual features, audio features and text features are input into the multimodal feature fusion module for feature fusion to obtain fusion features; The fused features are input into the event boundary prediction module, text description prediction module and event counter respectively to obtain the event boundaries in the video, the natural language description corresponding to each event and the number of events.

2. A multi-modal multi-feature end-to-end dense video description method as claimed in claim 1, characterized in that: The method of performing frame processing on a video and extracting key frames of the video comprises the steps of evenly dividing the video into a plurality of segments, then performing random sampling in each segment, and extracting a plurality of key frames.

3. The multi-modal multi-feature end-to-end dense video description method according to claim 1, characterized in that: The TSN model or Swin Transformer model is used to extract visual grid features from key frames. If the TSN model is used, the optical flow features and RGB features output by the TSN model are concatenated as visual grid features, and the TransVOD model is used to extract visual area features from key frames.

4. The multi-modal multi-feature end-to-end dense video description method according to claim 1, characterized in that: Extracting audio features includes the steps of extracting audio data from the video and inputting the audio data into the VGGish model to extract audio features.

5. The multi-modal multi-feature end-to-end dense video description method according to claim 1, characterized in that: Extracting text features includes the following steps: The key frame of the video is recorded as F input ={f1,f2,…,f N }, N is the number of key frames, and the automatic speech recognition technology is used to obtain the language text description corresponding to the key frame, which is recorded as {t1, t2, …, t N }, merge the speech text description of the video into a text paragraph, denoted as T, and use the BERT pre-trained model to extract features from the text paragraph T as text features.

6. A multi-modal multi-feature end-to-end dense video description method as claimed in claim 1, characterized in that: The multimodal feature fusion module includes a first self-attention module, a second self-attention module, a third self-attention module, a first cross-attention module, a second cross-attention module, a third cross-attention module, a first feedforward neural network layer, a second feedforward neural network layer, a third feedforward neural network layer and a fourth feedforward neural network layer, and the step of inputting the decoded visual features, audio features and text features into the multimodal feature fusion module for feature fusion includes the following steps: The decoded visual feature is recorded as X j , record the decoded audio features as Y j , the decoded text feature is recorded as Z j , X j Input to the first self-attention module to obtain feature X j ′, Y j Input to the second self-attention module to obtain feature Y j ′, Z j Input to the first self-attention module to obtain feature Z j ′, X j ′、Y j ′ and Z j ’ is input into the first cross-attention module, the second cross-attention module and the third cross-attention module for paired cross-attention calculation, and processed through the first feedforward neural network layer, the second feedforward neural network layer and the third feedforward neural network layer to obtain six interactive features, which are concatenated and input into the fourth feedforward neural network layer for dimensionality reduction to generate fusion features.

7. The multi-modal multi-feature end-to-end dense video description method according to claim 1, characterized in that: The visual feature extraction module, the audio feature extraction module, the text feature extraction module, the encoding module, the decoding module, the multimodal feature fusion module, the event boundary prediction module, the text description prediction module and the event counter are taken as a whole dense video description model for the first stage of training. The first stage of training includes the following steps: Construct a training sample set, where each training sample is annotated with event boundary labels, text description labels, and event quantity labels; Input the training samples into the dense video description model, and output the event boundary prediction results, text description prediction results, and event quantity prediction results; Calculate the intersection-over-union loss between the event boundary label and the event boundary prediction result, calculate the description loss representing the difference between the text description label and the text description prediction result, calculate the count loss representing the difference between the event number label and the event number prediction result, calculate the classification loss, and calculate the loss function. The calculation formula is: L=β giou L giou +b cls L cls +b ec L ec +b cap L cap Among them, L represents the loss function, L giou represents the intersection-over-union loss, L cls represents the classification loss, L ec Represents the counting loss, L cap represents the description loss, β giou , β cls , β ec , β cap All represent coefficients.

8. A multi-modal multi-feature end-to-end dense video description method as claimed in claim 7, characterized in that: After the first stage of training is completed, the second stage of training is carried out. In the second stage of training, greedy decoding and soft sampling are used to generate description sentences, and the description loss L cap The calculation includes the following steps: The METEOR and CIDEr scores of the text description prediction results generated by greedy decoding and soft sampling are calculated respectively, and the difference between the two scores is used as the reward. The calculation formula is: gen_scores=MS(gen,gt)×MW+CS(gen,gt)×CW greedy_scores=MS(greedy,gt)×MW+CS(greedy,gt)×CW reward=gen_scores-greedy_scores Among them, gen_scores represents the score of the text description prediction result generated by soft sampling, greedy_scores represents the score of the text description prediction result generated by greedy decoding, gen represents the text description prediction result generated by soft sampling, gt represents the text description label, greedy represents the text description prediction result generated by greedy decoding, MW represents the weight assigned to the METEOR score, CW represents the weight assigned to the CIDEr score, MS represents the function for calculating the METEOR score, CS represents the function for calculating the CIDEr score, and reward represents the reward; Describe the loss L in terms of reward calculation lr , the calculation formula is: L lr =-sample_logprobs×reward×mask Where sample_logprobs represents the logarithmic probability of the text description prediction module sampling from the probability distribution of all possible output words, and mask is a binary mask; Based on L lr Calculate the description loss L cap , L cap =L lr .mean(L lr ), mean(L lr ) indicates L lr Calculate the mean for each column.

9. A multi-modal multi-feature end-to-end dense video description system, characterized in that: include: A key frame extraction module is used to process the video into frames and extract the key frames of the video; The visual feature extraction module is used to extract visual area features and visual grid features from key frames, input the key frames into the ViFi-CLIP model to extract CLIP features, respectively map the visual area features and visual grid features to CLIP features, and fuse the two new features obtained by mapping to obtain visual features; An audio feature extraction module, used to extract audio features from videos; A text feature extraction module, used to extract text features from videos; The encoding and decoding module is used to encode and decode visual features, audio features and text features; The multimodal feature fusion module is used to fuse the decoded visual features, audio features and text features to obtain fused features; The prediction module is used to predict the event boundaries in the video, the natural language description corresponding to each event, and the number of events based on the fused features.