A video description generation method based on high dynamic multi-layer semantic coding

CN118247704BActive Publication Date: 2026-10-09UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410327726.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-21
Publication Date
2026-10-09
Estimated Expiration
2044-03-21

AI Technical Summary

Technical Problem

一方面视频内容通常具有高动态性,这可能导致受到非关键帧的影响,丢失关键帧所包含的重要视觉信息,另一方面固定的预训练特征提取器存在与描述生成任务不匹配的问题,使得了生成的描述存在细节性差、部分关键对象或内容缺失的问题

Benefits of technology

[0064]The proposed high-dynamic-range multi-layer semantic coding video description generation algorithm, through the design of a parallel-serial combined video feature coding structure, effectively extracts semantic information of intra-frame grid object relationships and inter-frame dynamic changes. Simultaneously, by stacking a multi-layer attention mechanism, it filters and fuses semantic features from multiple video frames, reducing interference from invalid frames on feature coding and fully mining the semantic information of key frames or segments in high-dynamic-range videos, providing highly representative video semantic features for generating natural language descriptions. Compared to existing video description algorithms, this algorithm outperforms existing algorithms in various video scenarios, demonstrating superior video description generation capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118247704B_ABST
    Figure CN118247704B_ABST
Patent Text Reader

Abstract

The application discloses a kind of video description generation method based on high dynamic multilayer semantic coding, video description generation field.The application is extracted and coded by using the powerful semantic feature extraction and coding ability of transformer structure, more abundant visual semantic features are obtained at video frame level, and a feature coding structure combining parallel and series is designed, the semantic information of intra-frame grid object relationship and inter-frame dynamic change semantic information are mined.Meanwhile, the coding structure of multilayer feature attention is designed, the selection and fusion of key frame visual features are carried out, the interference of invalid frames on feature coding is reduced, the coding capability of video semantic information in high dynamic scene is further enhanced, and the accuracy of video description generation is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of deep learning, computer vision, and video captioning, and designs a high-precision video captioning method based on highly dynamic multi-layer semantic coding. Background Technology

[0002] In recent years, against the backdrop of the artificial intelligence era, the field of multimodal scene analysis has achieved rapid development, demonstrating significant application value in areas such as human-computer interaction, multi-source intelligence analysis, medical analysis, and industrial intelligence, attracting widespread attention. Video description generation, as an important research topic in multimodal scene analysis, aims to generate natural language descriptions corresponding to the content of input videos. Unlike static image scenes, videos often span long periods of time and involve numerous scene changes, exhibiting dynamic characteristics. This necessitates not only a thorough understanding of the visual information of the entire input video but also the extraction of the most representative keyframes or video segments to achieve an accurate description of the video content.

[0003] Currently, achieving semantic feature encoding for highly dynamic videos is a key factor in generating accurate video descriptions. Some mainstream methods typically rely on models pre-trained for other visual tasks (such as object detection and video classification) to extract fixed video features. They then employ Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs) to encode and decode these visual features, ultimately generating a language description corresponding to the video content. However, these methods rely on complex pre-trained feature extraction networks, leading to challenges such as poor overall network flexibility, time-consuming feature extraction, and the need for substantial storage space for intermediate features. More importantly, video content is usually highly dynamic, often containing numerous changing scenes. Therefore, it is necessary to selectively extract keyframes or segments from the video. Traditional feature encoding easily overlooks the core visual information contained in keyframes and carries the potential risk of mismatch between pre-trained feature extractors and downstream video description generation tasks. This results in insufficient semantic representation power of the encoded video features, leading to poor detail in the generated descriptions and missing parts of the video content. Summary of the Invention

[0004] This invention proposes a video description generation algorithm based on highly dynamic multi-layer semantic coding. Existing video description generation techniques utilize feature extractors pre-trained on other visual tasks to obtain fixed visual representations, and then perform description generation through the fusion of overall visual representations. On the one hand, video content is usually highly dynamic, which may lead to the loss of important visual information contained in key frames due to the influence of non-keyframes. On the other hand, fixed pre-trained feature extractors have the problem of mismatch with the description generation task, resulting in poor detail in the generated descriptions, and the absence of some key objects or content.

[0005] This invention leverages the powerful semantic feature extraction and encoding capabilities of the transformer structure to obtain richer visual semantic features at the video frame level. It designs a parallel-serial combined feature encoding structure to mine semantic information about intra-frame grid object relationships and inter-frame dynamic changes. Simultaneously, a multi-layer feature attention encoding structure is designed to filter and fuse keyframe visual features, reducing interference from invalid frames and further enhancing the encoding capability of video semantic information in high-dynamic scenes, effectively improving the accuracy of video description generation. Therefore, the technical solution of this invention is: a video description generation method based on high-dynamic multi-layer semantic encoding, which includes:

[0006] Step 1: Video frame feature extraction;

[0007] For the input video, video frames are sampled at frames per second to obtain K frames I. The grid features F of each frame are extracted using a Swing Transformer pre-trained on ImageNet as the backbone network. b :

[0008] F b =SwinT(I)

[0009] in, Z is the number of video frame grids, C is the feature channel dimension, and SwinT(.) is the SwinTransformer feature extractor;

[0010] Step 2: Global visual information extraction and sharing;

[0011] First, average pooling is used to extract visual information F from the video frame level. b Capture global semantic information from the video and use it as shared visual information V. g Then, these features are injected into the visual features of each frame to obtain multi-frame grid features V that share global semantic information. b ;

[0012] V g =FC(AvgPool(Fb ))

[0013]

[0014] V b =Concat(V g V ′ b )

[0015] Where FC represents a fully connected layer that can map visual features to a feature embedding space with the same dimension, AvgPool(.) represents the average pooling operation for dimension K, and Concat(.) represents the feature concatenation operation.

[0016] To provide the location information of video frames within the video for the gridded visual frame features, thereby facilitating the mining of temporal dynamic changes between frames, further supplementing the location information V constructed based on the truncated normal distribution. pos Obtain the visual semantic features V at the initial video frame level. f ;

[0017] V ′ f =V b +V pos

[0018] V f =Dropout(LN(ReLU(V ′ f )))

[0019] Where ReLU(*) represents a non-linear activation function, LN(*) represents layer normalization operation, and Dropout(*) represents random zeroing operation of data;

[0020] Step 3: Construct a parallel multi-layer intra-frame mesh object relation semantic feature encoder

[0021] In the feature encoding stage, a parallel intra-frame intra-grid object relationship semantic feature encoder is designed to mine semantic information of intra-frame grid object relationships.

[0022] Specifically, feature encoding is achieved using a multi-head attention module MHA(Q,K,V):

[0023] MHA(Q,K,V)=Concat(h1,…,h n W1

[0024] h i =Attention(QW2,KW3,VW4)

[0025]

[0026] Among them, W * Let represent the learnable weight matrix, n represent the number of heads in the multi-head attention mechanism, Q, K, and V represent the query vector, key vector, and numerical vector, respectively, Attention(.) represents the attention weighting process, softmax(.) represents the probability distribution calculation process, and d represents the dimension of the key vector.

[0027] To fully model the correlation between different grid objects within the same video frame and maximize the extraction of effective information from the video frame, a parallel N-layer intra-frame grid object relationship semantic feature encoding structure is adopted. For the i-th layer and the l-th video frame, the grid object relationship semantic feature encoding is as follows:

[0028]

[0029]

[0030]

[0031] in, This represents the semantic features of the grid object in the i-th layer and the l-th video frame. Represents the semantic features of the grid object in the (i+1)th layer and the lth video frame. ReLU(.) represents the non-linear activation function, which initializes the input features of the 0th layer. V f This indicates the initialization of the semantic features of the video frame object;

[0032] Based on the semantic feature encoding output of the object relationship in the last frame, the visual semantic feature V of each frame is obtained using an average pooling layer. F ;

[0033]

[0034]

[0035] Where AvgPool(.) represents the average pooling layer, Represents the grid semantic features of the Nth layer and the lth video frame;

[0036] By using grid object relation semantic encoding, the ability of semantic features to represent intra-frame salient objects can be enhanced;

[0037] Step 4: Construct a serial multi-layer inter-frame semantic information feature encoder;

[0038] Furthermore, considering that videos typically have a long time span and varied scenes, exhibiting dynamic characteristics, a serial multi-layer inter-frame semantic information feature encoding was constructed to learn the direct correlations between different video frames. By enhancing the semantic feature responses corresponding to the most representative keyframes or video segments, the dynamics and comprehensiveness of video content are effectively represented during video feature encoding, avoiding the loss of keyframe information and providing more reliable feature representations for subsequent description generation.

[0039] For V F The inter-frame semantic information feature coding structure with serial M layers is adopted. For the i-th layer, the inter-frame semantic information feature coding is as follows:

[0040]

[0041]

[0042]

[0043] in, Representing the semantic features of the video, This represents the semantic features of the video after encoding at the (i+1)th layer.

[0044] Step 5: Construct a natural language description generation decoder;

[0045] During the decoding phase, a Transformer structure with a multi-head attention mechanism is used for word prediction. For time t, in order to fully consider the influence of the semantic information of the words predicted in the previous t-1 times on the word prediction in the current time, language features from the previous t-1 times are fused.

[0046]

[0047] Among them, H t-1 H represents the hidden state of the Transformer at time t-1. 0:t-1 This represents the hidden layer state at time t-1. Represents the semantic features of the text at time t;

[0048] Meanwhile, by calculating the correlation score between the video frame and the hidden state at the previous moment, video frame features are filtered and fused, and the fused video frame features are used as the video semantic feature encoding result for description generation.

[0049]

[0050] in, Represents the semantic features of the video at time t. This represents the semantic features of the video after the Mth layer of encoding;

[0051] Then, the fused visual and linguistic features are fed into the Transformer structure for cross-modal temporal semantic mapping:

[0052]

[0053]

[0054] H t =LN(Dropout(H ′ t-1 W2)+H t-1 )

[0055] Finally, a fully connected layer is used to store the hidden state H of the Transformer. t Mapping to the word vector space, realizing the word w t Probability prediction:

[0056] P t (w t =Softmax(FC(Dropout(H)) t )))

[0057] Based on word w t Predicted probability P t (w t The word generation during the training phase is constrained by cross-entropy loss.

[0058]

[0059] Where T represents the maximum number of words in a sentence, P t (w t () represents the probability value of predicting the word at time t.

[0060] Step 6: Network optimization based on reinforcement learning strategy;

[0061] Based on the description generation model trained in step 5, a reinforcement learning training strategy is further adopted, and a reward mechanism is constructed using CIDEr scores to constrain the caption generation process:

[0062] loss reward =-E 1:T (score(w 1:T ))

[0063] Where score(*) represents the CIDEr score, E 1:T (.) represents the expected value of the predicted score for each word, loss. reward This indicates a strengthening of the learning reward mechanism.

[0064] The proposed high-dynamic-range multi-layer semantic coding video description generation algorithm, through the design of a parallel-serial combined video feature coding structure, effectively extracts semantic information of intra-frame grid object relationships and inter-frame dynamic changes. Simultaneously, by stacking a multi-layer attention mechanism, it filters and fuses semantic features from multiple video frames, reducing interference from invalid frames on feature coding and fully mining the semantic information of key frames or segments in high-dynamic-range videos, providing highly representative video semantic features for generating natural language descriptions. Compared to existing video description algorithms, this algorithm outperforms existing algorithms in various video scenarios, demonstrating superior video description generation capabilities. Attached Figure Description

[0065] Figure 1 This is a flowchart illustrating the overall process framework for generating video descriptions using the high dynamic range multi-layer semantic coding of this invention.

[0066] Figure 2 This is a schematic diagram of the high dynamic multilayer semantic coding algorithm of the present invention. Detailed Implementation

[0067] This invention is implemented on Xmodaler, a multimodal scene parsing framework based on the PyTorch deep learning platform, and specifically includes the following steps:

[0068] Step 1: Database Selection. Choose a suitable video description database, such as the MSVD database, the MSR-VTT database, or a larger-scale VATEX database;

[0069] Step 2: Data Preprocessing. For video data, the videos in the database are acquired at a rate of 1 frame per second, and then scaled proportionally to a size of 384×384 images. For text data, all words are converted to lowercase, and all punctuation marks and low-frequency words appearing less than 5 times are removed.

[0070] Step 3: Model parameter initialization. Except for the pre-trained network used, all parameters in the model are randomly initialized.

[0071] Step 4: Model Hyperparameter Settings. The visual feature dimension, hidden state dimension, and word embedding vector dimension in the model are all set to 512. The AdamW optimizer is used for optimization, with an initial learning rate of 0.0004, adjusted using a step decay mechanism that decreases by a factor of 0.8 every two training epochs.

[0072] Step 5: First-stage model training. After completing the model parameter initialization in Step 3 and the model hyperparameter setting in Step 4, the training data after the data preprocessing in Step 2 is fed into the model in batches, and cross-entropy loss is used for optimization.

[0073] Step Six: Second-stage model training. Based on the model obtained in Step Five, a reinforcement learning strategy is used, with the CIDEr score as the reward function, to further optimize the network.

[0074] Step 7: Final Model Inference Validation. After the two-stage model training is completed, the test data undergoes the data preprocessing operation in Step 2, and then is fed into the final model obtained in Step 6 for inference to obtain the test results.

[0075] The main protection of this invention lies in the design of a parallel and serial combined video feature coding structure, which mines semantic information of intra-frame grid object relationships and inter-frame dynamic change semantic information. Simultaneously, by utilizing a multi-layer semantic feature coding structure, video frame features are filtered and fused to extract semantic information corresponding to key frames or segments in the video, reducing interference from invalid frames on feature coding, enhancing the representational ability of video semantic features, and further improving the performance of video description generation.

[0076] Based on the reinforcement learning strategy, this invention achieved a score of 53.1% on the key description evaluation metric CIDEr.

Claims

1. A video description generation method based on high dynamic range multilayer semantic coding, the method comprising: Step 1: Video frame feature extraction; For the input video, video frames are sampled at frames per second to obtain K frames. The grid features of each frame of the image are extracted using a Swing Transformer pre-trained on ImageNet as the backbone network. : ; in, , The number of video frame grids. For feature channel dimension, For the SwinTransformer feature extractor; Step 2: Global visual information extraction and sharing; First, average pooling is used to extract visual information at the video frame level. Capture global semantic information from the video and use it as shared visual information. Then, these features are injected into the visual features of each frame to obtain multi-frame grid features with shared global semantic information. ; ; ; ; in, This represents a fully connected layer that maps visual features to a feature embedding space with the same dimensions. Indicates the dimension Average pooling operation, Indicates a feature cascade operation; Supplementing location information based on truncated normal distribution Obtain visual semantic features at the initial video frame level. ; ; ; Among them, ReLU LN represents a non-linear activation function. Presentation layer normalization operation, This indicates a random zeroing operation on the data; Step 3: Construct a parallel multi-layer intra-frame mesh object relation semantic feature encoder Utilizing a multi-head attention module Implement feature encoding: ; in, This represents the learnable weight matrix. This indicates the number of heads in a multi-head attention mechanism. These represent the query vector, key vector, and value vector, respectively. This represents the attention-weighted process. This describes the process of calculating the probability distribution. Indicates the key vector dimension; For the Layer, First The semantic feature encoding of grid object relationships in each video frame is as follows: ; ; ; in, Indicates the first Layer, First Semantic features of grid objects in a video frame. Indicates the first Layer, First Semantic features of grid objects in a video frame. Represents a non-linear activation function, used to initialize the input features of layer 0. ∈ , , This indicates the initialization of the semantic features of the video frame object; Based on the semantic feature encoding output of the object relationship in the last frame, the visual semantic features of each frame are obtained using an average pooling layer. ; ; ; in, Indicates the average pooling layer. Indicates the Nth layer, the The grid semantic features of each video frame; Step 4: Construct a serial multi-layer inter-frame semantic information feature encoder; for , using serial The inter-frame semantic information feature encoding structure of the layer, for the first layer The layer, inter-frame semantic information feature encoding is as follows: ; ; ; in, Representing the semantic features of the video, Indicates the first Video semantic features after layer encoding Indicates the first Video semantic features after layer encoding; Step 5: Construct a natural language description generation decoder; Before proceeding Moment-based language feature fusion: ; in, express The hidden state of the Transformer at any given moment. Indicates the preceding The hidden layer state at any given time. express Momentary text semantic features; Meanwhile, by calculating the correlation score between the video frame and the hidden state at the previous moment, video frame features are filtered and fused, and the fused video frame features are used as the video semantic feature encoding result for description generation. ; in, express Moment-time video semantic features This represents the semantic features of the video after the Mth layer of encoding; Then, the fused visual and linguistic features are fed into the Transformer structure for cross-modal temporal semantic mapping: ; ; ; Finally, a fully connected layer is used to hide the Transformer's state. Mapping to the word vector space to realize words Probability prediction: ; Based on words Predicted probability Word generation during the training phase is constrained by cross-entropy loss: ; in, This indicates the maximum number of words in a sentence. Indicates the current Predict the probability value of words at any time; Step 6: Network optimization based on reinforcement learning strategy; Based on the description generation model trained in step 5, a reinforcement learning training strategy is further adopted, and a reward mechanism is constructed using CIDEr scores to constrain the caption generation process: ; in, This represents the score of the CIDEr indicator. This represents the expected value of the predicted score for each word. This indicates a strengthening of the learning reward mechanism.

Citation Information

Patent Citations

  • Video text description method fusing multi-granularity video semantic information

    CN114943921A

  • Video description method based on advanced semantic information feature coding

    CN116091978A