Video abstraction method based on graph model and multi-scale attention mechanism
By introducing graphical models and multi-scale attention mechanisms into the video digest method, the lack of semantic integrity and redundancy of video digests in the prior art is solved, and a more accurate and representative video digest generation is achieved.
Patent Information
- Application Number
- CN202510153403.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-27
AI Technical Summary
Existing video digest methods are difficult to effectively capture the complex timing structure and multi-scale information in video, resulting in the generated digest lacking semantic integrity and accuracy, and there are redundancy problems.
The video digest method based on graph model and multi-scale attention mechanism is adopted to capture inter-frame relationships through global and local attention mechanisms, and redundant frames are removed using non-maximum suppression to generate a more representative video digest.
Improve the quality and accuracy of video digests, reduce redundant content, and generate digests more representative and informative.
Smart Images

Figure CN120050491A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision, and specifically provides a video summarization method based on a graph model and a multi-scale attention mechanism. Background Art
[0002] With the rapid development of the Internet and multimedia technologies, the amount of video data has increased explosively. How to effectively extract key information from a large number of videos has become an urgent problem to be solved. Traditional video summarization methods often rely on simple visual features or manually designed rules, and it is difficult to capture the complex temporal structure and multi-scale information in videos, resulting in the generated summary segments lacking semantic integrity and accuracy. In addition, video content usually has redundancy, and highly similar frames will lead to redundant generated summaries, affecting the user experience.
[0003] It can be seen that most current video summarization methods have obvious deficiencies in processing long video sequences. Specifically, it is difficult to effectively capture global and local key features, resulting in a decrease in the accuracy of the summary. At the same time, the generated summaries also have deficiencies in terms of shot diversity, and content repetition and redundancy are prone to occur. Based on this, a video summarization method based on a graph model and a multi-scale attention mechanism is specifically proposed, which can combine global and local features, effectively model the inter-frame relationship, and remove redundant content, generating a more representative, concise and information-rich video summary. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a video summarization method based on a graph model and a multi-scale attention mechanism, which solves the problems raised in the above background art.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A video summarization method based on a graph model and a multi-scale attention mechanism specifically includes the following steps:
[0006] Step 1: Input the original video into the feature extraction module to obtain a video sequence and extract a frame-level feature sequence;
[0007] Step 2: Add GLS tokens to the frame-level feature sequence;
[0008] Step 3: Input the frame-level feature sequence into the global attention module to construct an inter-frame association graph, and use the attention weights to calculate the aggregated frame-level feature representation to capture the global inter-frame dependence relationship;
[0009] Step 4: Cut all video sequences into several uniformly long subsequences, classify and label them, then connect them to the frame-level feature sequence and input them into the multi-head attention module. Output the attention features of each subsequence, and stack the attention features of each subsequence through the feature fusion module to obtain local attention features;
[0010] Step 5: After the feature fusion module connects the local attention features and the global attention results, through the multi-branch frame comprehensive prediction module, it outputs the importance score, centrality score, and temporal position information of the frame respectively, and obtains the comprehensive score of each frame after weighted fusion;
[0011] Step 6: Remove redundant frames through the non-maximum suppression method, suppress low-contribution frames according to the similarity between frames, screen out the most representative set of video key frames, and then obtain a new video summary sequence through key shot selection.
[0012] The present invention is further configured that: the feature extraction module in the above Step 1 is a GoogleNet network. After the GoogleNet network segments the original video to obtain video frames, it downsamples the video frames to obtain a video sequence.
[0013] The present invention is further configured that: the dimension of the GLS token in the above Step 2 is the same as that of the frame feature vector, denoted as C∈R d , as a learnable parameter, is used to aggregate the global information of the video sequence and globally model all frames, and is used to capture the complex spatio-temporal relationships in the video sequence.
[0014] The present invention is further configured that: the global attention module in the above Step 3 is GATv2, and its attention expression is:
[0015]
[0016] In the formula, α i,j is the normalized attention coefficient of node i to node j, W is a learnable weight matrix, h i h j are the feature vectors of node i and node j respectively, LeakyReLU is a non-linear activation function, a T is the transpose of the learnable weight vector, is the neighbor set of node i.
[0017] The present invention is further configured that: when the feature fusion module in the above Step 5 connects the local attention features and the global attention results, two repeated MetaFormer modules are used. The MetaFormer module includes two residual sub-blocks, which can alleviate the problem of gradient disappearance in the deep network. By introducing residual connections, each sub-block can deepen the depth of the network while retaining the original input features, thereby improving the representation ability of the network.
[0018] The present invention is further configured that: the multi-branch frame comprehensive prediction module includes a shared frame prediction module and three independent loss function branches;
[0019] The shared frame prediction module consists of a fully connected layer, a ReLU activation function, a Dropout layer, and a Layernorm layer;
[0020] The shared frame prediction module is connected through three parallel and independent loss function branches, and each loss function branch is equipped with a fully connected layer. The three loss function branches are respectively used to output: frame importance score, frame centrality score, and frame sequence position deviation. After weighted summation, the comprehensive score of each frame is obtained.
[0021] The present invention is further configured that: the method for screening out the most representative set of video key frames in step six includes:
[0022] Introduce a confidence score:
[0023] c j =s j ×μ j
[0024] In the formula, c j is the confidence score, s j is the frame importance score, and μ j is the frame centrality score;
[0025] Use the non-maximum suppression algorithm to remove redundant frames and retain the video segments corresponding to the frames with high confidence.
[0026] The present invention is further configured that: the method for selecting key shots in step six includes:
[0027] The video sequence composed of all video frames is segmented into several shots with uniform lengths by the KTS algorithm, and the key shots are selected according to the comprehensive scores of the shots by the 0 / 1 knapsack algorithm.
[0028] The present invention provides a video summarization method based on a graph model and a multi-scale attention mechanism. It has the following beneficial effects:
[0029] (1) By introducing a multi-scale attention mechanism, the present invention separately models the local inter-frame relationship and the global long-term dependence relationship, effectively reducing the deviation of attention weights in the calculation process. At the same time, the GATv2 graph attention mechanism is used to extract the local inter-frame correlation, and the CLS identifier is used to aggregate the global feature information to ensure the efficient fusion of global and local information. In addition, the non-maximum suppression is used to remove redundant frames, further improving the quality of the video summary, avoiding the redundant problem caused by the injection of position information, and realizing more accurate inter-frame relationship modeling and video content summary.
[0030] (2) By introducing the GATv2 graph attention mechanism, the present invention effectively models the temporal and semantic complex relationships between video frames, improves the representation ability of frame features, significantly enhances the dynamics and generalization performance of the attention mechanism through non-linear activation reordering, improves the quality of video summaries, combines the features output by the frame-level attention mechanism and GATv2, and uses a multi-branch fusion and comprehensive prediction module to generate more accurate frame importance scores, improving the accuracy of key frame extraction and the overall quality of video summaries. Description of the Drawings
[0031] Figure 1 It is a schematic diagram of the system model distribution of the present invention;
[0032] Figure 2 It is a schematic diagram of the system architecture of the multi-branch frame comprehensive prediction module of the present invention. Detailed Embodiments
[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0034] Please refer to Figure 1-2 , the embodiments of the present invention provide the following technical solutions: A video summary method based on a graph model and a multi-scale attention mechanism, specifically including the following steps:
[0035] Step 1: Input the original video into the feature extraction module. Among them, the original video includes, but is not limited to, the SumMe dataset and the TVSum dataset. After using the GoogleNet network to segment the original video to obtain video frames, downsample the video frames to obtain a video sequence, and extract the frame-level feature sequence.
[0036] Step 2: Add a GLS token to the frame-level feature sequence as a classification marker to obtain the feature set X = CLS, x 1 , x 2 , x 3 ,..., x n .
[0037] Step 3: Input the frame-level feature sequence into the global attention module, construct an inter-frame association graph, and use the attention weights to calculate the aggregated frame-level feature representation to capture the global inter-frame dependence relationship. Specifically, the global attention module is GATv2, and its attention expression is:
[0038]
[0039] In the formula, α i,j is the normalized attention coefficient of node i to node j, W is a learnable weight matrix, h i is the feature vector of node i, h jis the feature vector of node j, LeakyReLU is a non-linear activation function, a T is the transpose of a learnable weight vector, is the set of neighbors of node i.
[0040] Among them, in the video frame structure, a weighted graph is used to represent the frame sequence. A weighted graph is a type of directed graph, where the weight of an edge represents the strength of the relationship between two nodes. After feature extraction, the dot product between feature vectors is calculated to obtain an edge weight matrix. Each element of the matrix reflects the similarity between node features. To avoid a node connecting to itself, the matrix is subtracted by the identity matrix, making the diagonal elements zero. Then, the topk function in Pytorch is used to select several edges with the largest weights. The returned results include the values of the largest weights and their index positions. Then, the two-dimensional structure of the matrix is restored, and the indices of the node pairs of the selected edges are calculated. Finally, the frame sequence is transformed into a weighted directed graph, and the edges connecting the nodes and the corresponding weights are obtained. For GATv2, for each pair of nodes h i and h j , first concatenate their feature vectors: [h i ||h j Then, map the concatenated features through a linear transformation W:
[0041] h ij =W[h i ||h j
[0042] where W is a learnable weight matrix used to map the concatenated node features to a new space;
[0043] Apply the activation function LeakyReLU to the mapped features:
[0044] h’ ij =LeakyReLU(h ij )
[0045] where LeakyReLU is a commonly used non-linear activation function that can effectively introduce non-linear features, enabling the model to learn more complex relationships;
[0046] Use a learnable weight vector a to weight the activated features and calculate the final attention coefficient:
[0047] e i,j =a T h' ij
[0048] where the attention coefficient e i,j Describes the similarity between node i and node j, reflecting the degree of attention of node i to node j.
[0049] Use the softmax function to normalize it to obtain the attention score of node i to node j:
[0050]
[0051] Among them, represents the neighbor set of node i, and α i,j represents the normalized attention coefficient of node i to node j;
[0052] Utilize the normalized attention coefficient α i,j , perform weighted summation on the features of the neighbor nodes of node i to obtain the final feature representation of node i:
[0053]
[0054] Through this process, GATv2 aggregates the information of neighbor nodes and weights according to their importance.
[0055] Step 4. Cut all video sequences into several subsequences with uniform length After classification marking and connecting to the frame-level feature sequence, input it into the multi-head attention module, output the attention features of each subsequence, and stack the attention features of each subsequence through the feature fusion module to obtain the local attention features.
[0056] Specifically, the multi-head attention module maps the input feature sequence into multiple queries, keys, and values. By calculating the similarity between them, use the obtained attention scores for weighted averaging to finally generate the output of the model. Among them, first generate multiple groups of different queries Q, keys K, and values V according to the linear projection of the input data. These projections are realized through learnable weight matrices. Each head will have a set of independent weight matrices. Then, use the scaled dot-product method to calculate the attention weights output by each head. Specifically, for each head, calculate the dot product of the query and the key and divide it by a scaling factor. The scaling factor is usually the square root of the dimension of the key. Then obtain the normalized attention weights through the Softmax function. These attention weights are used for weighted summation of the values to obtain the output of each head:
[0057]
[0058] In the formula, d k is the scaling factor, the softmax function represents the normalization operation on the attention scores, Q is the query vector, and K T is the transpose of the key vector.
[0059] Concatenate the outputs of all heads and perform a linear transformation to obtain the final output:
[0060] MultiHeadAttention(Q, K, V) = Concat(head 1 , …, head n )W o
[0061] Meanwhile, because video frames have sequentiality, in order to preserve the sequential information of the frames and avoid disrupting the order and affecting the understanding of the video content, positional encoding is added when using dot-product attention. The positional encoding is generated by different frequencies of sine and cosine functions to enhance the positional information of each frame in the sequence:
[0062]
[0063] In the formula, pos represents the positional information, i represents the i-th frame in the video frame sequence, and d represents the dimension information corresponding to the feature vector.
[0064] Furthermore, the local attention features are stacked in the 0th dimension to make the shape of the stacked vector the same as the global attention features output by the global attention module in step three. Add the global attention features and the local attention features to obtain the video frame features. Specifically, through the MetaFormer module and the fully connected layer, global aggregation is performed. This process mainly relies on two repeated MetaFormer modules. Each MetaFormer module consists of two residual sub-blocks. The first residual sub-block realizes feature mixing and transmission through TokenMixer and performs standardization processing using layer normalization:
[0065] Y = TokenMixer(Norm(X)) + X
[0066] Among them, the Norm function represents the layer normalization operation, and the TokenMixer module is responsible for mixing the input information and performing feature recombination;
[0067] The second residual sub-block consists of two layers of MLP, and combines the non-linear activation function ReLU for feature mapping and enhancement. The specific calculation includes mapping through the learnable parameter weights W 1 and W 2 to ensure that the features can be more fully expressed in the non-linear space:
[0068] Z = σ(Norm(Y)W 1 )W 2 + Y
[0069] With this structure, PoolFormer effectively reduces the computational complexity while retaining good feature expression ability, enabling it to exhibit efficient and stable performance in video processing tasks.
[0070] Step Five: Through the multi-branch frame synthesis prediction module, output the importance score, centrality score, and temporal position information of the frame respectively, and obtain the comprehensive score of each frame after weighted fusion. Specifically, the multi-branch frame synthesis prediction module includes a shared frame prediction module and three independent loss function branches. The shared frame prediction module consists of a fully connected layer, a ReLU activation function, a Dropout layer, and a Layernorm layer. The shared frame prediction module is connected through three parallel and independent loss function branches, and each loss function branch is equipped with a fully connected layer. The three loss function branches are respectively used to output: the frame importance score s j , the frame centrality score μ j and the frame sequence position deviation δt j . After weighted summation, the comprehensive score of each frame is obtained. The outputs of the three loss function branches use different loss functions to calculate the error between the prediction result and the true value, and the network parameters are adjusted backward in the way of a multi-task loss function.
[0071] During the training process, the frame sequence position deviation is represented as a two-dimensional vector. According to the loss function:
[0072]
[0073] Add the frame importance score prediction term and the frame sequence position deviation term, and use the focal loss L cls to alleviate the class imbalance problem, so as to improve the accuracy of video summarization. Use the intersection over union loss L reg to optimize the segment selection in the video summarization generation process. The λ coefficient is used to balance these two loss terms;
[0074] To ensure that there will not be too many low-quality segments close to the true summary boundary in the video summarization generation process, the centrality score is also introduced. For each frame in the true summary, calculate its centrality score:
[0075]
[0076] After obtaining the centrality score, use the binary cross-entropy BCE loss function to balance the centrality score of the predicted frame. Finally, the three loss functions are added together to obtain the complete loss function expression:
[0077]
[0078] Step Six: Introduce the confidence score:
[0079] cj = s j × μ j
[0080] Wherein, c j is the confidence score, s j is the frame importance score, and μ j is the frame centrality score;
[0081] Redundant frames are removed by the non-maximum suppression method, low contribution frames are suppressed according to the inter-frame similarity, and the most representative set of video key frames is selected;
[0082] A new video summary sequence is obtained through key shot selection. Specifically, segment division is performed by the KTS algorithm, and the average importance of each segment is statistically calculated:
[0083]
[0084] Wherein, N p is the number of frames in the p-th shot, and s p,i is the importance score of the i-th frame in the p-th shot
[0085] The 0 / 1 knapsack algorithm is applied to select the corresponding segments as key shots, and the generated summary length is less than or equal to 15% of the original video length.
[0086] In summary, this embodiment can effectively obtain the summary content in the long video, is highly practical, and has a good promotion prospect.
[0087] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.
[0088] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A video summarization method based on a graph model and a multi-scale attention mechanism, characterized by: The specific steps include: Step 1: Input the original video into the feature extraction module, obtain the video sequence, and extract the frame-level feature sequence; Step 2: Add GLS tokens to the frame-level feature sequence; Step 3: Input the frame-level feature sequence into the global attention module, construct the inter-frame association graph, and use the attention weight calculation to obtain the aggregated frame-level feature representation to capture the global inter-frame dependency; Step 4: Divide all video sequences into several subsequences of uniform length, classify and mark them, connect them to the frame-level feature sequence, input them into the multi-head attention module, output the attention features of each subsequence, and stack the attention features of each subsequence through the feature fusion module to obtain the local attention features; Step 5: After the feature fusion module connects the local attention features with the global attention results, it outputs the importance score, centrality score and time position information of the frame through the multi-branch frame comprehensive prediction module, and obtains the comprehensive score of each frame after weighted fusion; Step 6: Remove redundant frames through non-maximum suppression method, suppress low-contribution frames based on inter-frame similarity, screen out the most representative set of video key frames, and then obtain a new video summary sequence through key shot selection.
2. A video summarization method based on a graph model and a multi-scale attention mechanism according to claim 1, characterized in that: The feature extraction module in step 1 is a GoogleNet network. After the GoogleNet network segments the original video to obtain video frames, it downsamples the video frames to obtain a video sequence.
3. The video summarization method based on a graph model and a multi-scale attention mechanism according to claim 1, characterized in that: The dimension of the GLS token in step 2 is the same as the frame feature vector, denoted by C∈R d , as a learnable parameter, is used to aggregate the global information of the video sequence and globally model all frames to capture the complex spatiotemporal relationships in the video sequence.
4. The video summarization method based on a graph model and a multi-scale attention mechanism according to claim 1, characterized in that: The global attention module in step 3 is GATv2, and its attention expression is: In the formula, α i,j is the normalized attention coefficient of node i to node j, W is the learnable weight matrix, [h i ||h j ] are the feature vectors of node i and node j respectively, LeakyReLU is the nonlinear activation function, a T is the transpose of the learnable weight vector, is the neighbor set of node i.
5. The video summarization method based on a graph model and a multi-scale attention mechanism according to claim 1, characterized in that: In the step 5, the feature fusion module uses two repeated MetaFormer modules when connecting the local attention features with the global attention results, and the MetaFormer module includes two residual sub-blocks.
6. The video summarization method based on a graph model and a multi-scale attention mechanism according to claim 1, characterized in that: The multi-branch frame comprehensive prediction module includes a shared frame prediction module and three independent loss function branches; The shared frame prediction module includes a fully connected layer, a ReLU activation function, a Dropout layer and a Layernorm layer; The shared frame prediction module is connected through three parallel independent loss function branches, and each loss function branch is equipped with a fully connected layer. The three loss function branches are used to output: frame importance score, frame centrality score and frame sequence position deviation respectively. After weighted summation, the comprehensive score of each frame is obtained.
7. The video summarization method based on a graph model and a multi-scale attention mechanism according to claim 6, characterized in that: The method of selecting the most representative video key frame set in step 6 includes: Introducing confidence scores: c j =s j ×μ j In the formula, c j is the confidence score, s j is the frame importance score, μ j is the frame centrality score; Use the non-maximum suppression algorithm to remove redundant frames and retain the video segments corresponding to the frames with high confidence.
8. The video summarization method based on a graph model and a multi-scale attention mechanism according to claim 1, characterized in that: The key lens selection method in step 6 includes: The video sequence composed of all video frames is divided into several shots of uniform length by the KTS algorithm, and the key shots are screened out according to the comprehensive scores of the shots by the 0 / 1 backpack algorithm.
Citation Information
Patent Citations
Novel crop identification and classification system based on Meta-CNN network
CN115311577A
Multi-modal remote sensing data classification method fusing global and local information
CN116863247A
Video abstraction method and device based on graph model and attention mechanism, storage medium and equipment
CN116887012A
Rumor detection method based on graph attention network
CN117112786A
Video abstraction method based on symmetric multi-scale attention
CN117493607A
Cited By
Video abstract generation method based on global multi-scale coding and local sparse attention
CN120296202A
Video summary generation method based on global multi-scale coding and local sparse attention
CN120296202B
Multi-scale time sequence modeling video abstract generation method fusing semantic enhancement and boundary perception
CN121665090A
A multi-scale temporal modeling video summary generation method fusing semantic enhancement and boundary perception
CN121665090B
Video abstraction method based on dynamic multi-dimensional heterogeneous graph and attention aggregation
CN121901454A