Video abstract generation method based on global multi-scale coding and local sparse attention

Through the video digest generation method of global multi-scale encoding and local sparse attention, the problem of high computational complexity in long video processing is solved, efficient and accurate video digest generation is achieved, and the representativeness and diversity of video digests are improved.

CN120296202AActive Publication Date: 2025-07-11SHIJIAZHUANG TIEDAO UNIV

Patent Information

Application Number
CN202510772225.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-11
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The existing video digest methods have high computational complexity in long video processing, making it difficult to effectively model global and local features, resulting in poor digest quality, especially in terms of processing long-distance dependencies and retaining detailed information.

Method used

A video digest generation method using global multi-scale encoding and local sparse attention is adopted. The global context vector is introduced through the video feature optimizer, combined with the multi-head attention mechanism and multi-scale depth separation convolution operation, combined with the local block diagonal sparse attention module, adaptively weighted fusion of global and local feature representations to generate a video digest.

Benefits of technology

It significantly improves the representativeness and accuracy of video summary, can fully capture global long-distance dependence and local details, reduces computing complexity, and generates efficient, diverse and accurate video summary.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296202A_ABST
    Figure CN120296202A_ABST
Patent Text Reader

Abstract

The invention discloses a video abstract generation method based on global multi-scale coding and local sparse attention, and belongs to the technical field of computer vision. The method comprises the following steps: reading an input video, and extracting a frame-level feature vector of the input video; constructing a video abstract generation model, and inputting the frame-level feature vector into the video abstract model; carrying out adaptive fusion on the global multi-scale coding and the output of the local sparse attention module to generate fusion features; and fusing the features, outputting frame importance scores through a regression network, selecting representative frames, and generating a video abstract. According to the video abstraction method provided by the invention, by combining global multi-scale coding and a local sparse attention mechanism, key segments in the video can be accurately identified, and the video browsing efficiency and the user experience are remarkably improved. The experiment is carried out on reference data sets SumMe and TVSum, and the experiment result fully proves the effectiveness of the method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a video summary generation method based on global multi-scale encoding and local sparse attention, belonging to the technical field of computer vision. Background Art

[0002] With the rapid development of multimedia technology, video data has shown explosive growth in many fields such as social media, surveillance and security, education and training, medical imaging, news media, etc. Users are faced with the problem of how to quickly obtain key information from a large amount of video data. Manually watching and screening long videos is not only time-consuming and laborious, but also easy to miss important segments. Therefore, the automated video summary technology has become an important research direction in the field of computer vision in recent years, aiming to automatically extract important frames or segments in the video through algorithm means, generate a concise and representative video summary, and help users quickly understand the video content.

[0003] Early video summary methods mainly relied on low-level visual features, such as color, texture, edge, optical flow, etc., and selected key frames by analyzing the distribution and changes of these features. However, such methods based on handcrafted features lack the understanding of high-level semantics and context relationships, and often can only capture surface information, resulting in the generated summary lacking content integrity and expressiveness. With the rise of deep learning, researchers introduced models such as convolutional neural networks, recurrent neural networks, and long short-term memory networks to automatically learn frame importance scores, and achieved more accurate summary generation through end-to-end training. These methods have significantly improved the model's ability to model high-level semantics, but still face many challenges when dealing with long videos.

[0004] First, the existing deep learning-based methods have deficiencies in long-range dependence modeling. Video is a kind of temporal data, and there are often complex long-term dependence relationships between frames. Only focusing on local information cannot comprehensively understand the overall narrative structure of the video. Although the self-attention mechanism provides a solution, its computational complexity increases quadratically with the growth of the sequence length, making it face computational bottlenecks when dealing with long videos. Second, in order to improve efficiency, some models adopt means such as sampling or dimensionality reduction, but this is likely to lose key local details, especially in scenarios sensitive to short-term events such as sports events and security surveillance, resulting in a decline in summary quality.

[0005] Therefore, how to comprehensively model global and local features, effectively capture long-term and short-term dependencies, and at the same time retain rich context and detail information on the premise of ensuring computational efficiency has become a key issue that urgently needs to be broken through in the current video summary field. Summary of the Invention

[0006] The object of the present invention is to provide a video summary generation method based on global multi-scale coding and local sparse attention, aiming to solve the problem in the prior art that there is a trade-off between global dependence modeling and local detail capture, resulting in high computational complexity in long video processing and unsatisfactory summary quality.

[0007] To achieve the above object, the present invention provides a video summary generation method based on global multi-scale coding and local sparse attention, including the following steps: S1: Read the input video and extract the frame-level feature vectors of the video frames; S2: Input the frame-level feature vectors into the video summary generation model to predict the frame-level importance scores. The video summary generation model includes: Video feature optimizer: The video feature optimizer receives the frame-level feature vectors, generates a global context vector through performing global average pooling, performs attention weighting on the global context vector and the frame-level feature vectors, and outputs an optimized feature sequence; Global multi-scale coding module: The global multi-scale coding module receives the optimized feature sequence, extracts the global feature representation through the multi-head attention mechanism, and extracts the local feature representation through the multi-scale depthwise separable convolution operation, and extracts the global semantic information in the video by fusing the global feature representation and the local feature representation; Local block diagonal sparse attention module: The output of the global multi-scale coding module serves as the input of the local block diagonal sparse attention module. The local block diagonal sparse attention module sparsifies the self-attention matrix into a block diagonal structure, and models the local dependence relationship by splicing the attention-weighted features, the frame-level uniqueness features and the inter-block diversity features; S3: Perform adaptive weighted fusion on the output of the global multi-scale coding module and the output of the local block diagonal sparse attention module to generate a fused feature representation; S4: Input the fused feature representation into the regression network, output the frame-level importance scores, select the most representative frames, and generate the final video summary.

[0008] Preferably, the video feature optimizer includes: Perform global average pooling on the frame-level feature vectors to generate a global context vector C, and the specific calculation formula is as follows: , where represents the feature vector of the i-th frame, and n is the number of video frames; Perform attention weighting on the global context vector and the frame-level feature vectors to generate a weighted feature sequence Y; Perform layer normalization on the weighted feature sequence and output the optimized feature sequence Z. The specific calculation formula is: Z = LayerNorm(Y).

[0009] Preferably, the global multi-scale coding module includes: Calculate the global feature representation G of the optimized feature sequence using the multi-head attention mechanism. The specific calculation formula is: G = Concat(head1,…,head h )W G where head i = A i ’ V i ’ , A i ’ is the attention weight obtained by the global multi-scale coding module modeling the global inter-frame dependence relationship through the multi-head self-attention mechanism, V ’ is the value matrix obtained through linear transformation, h is the number of attention heads, and W G is the output mapping matrix; Extract the local feature representation F using the multi-scale depthwise separable convolution operation. The specific calculation formula is as follows: F = F3+F5 where F3 = Pointwise(D3(V ’ ))、F5 = Pointwise(D5(V ’ )),D k represents the depth convolution operation with a kernel size of k, and Pointwise represents the pointwise convolution; Finally, perform adaptive feature fusion on the global feature representation G and the local feature representation F to generate the final global multi-scale coding representation F global .

[0010] Preferably, the local block diagonal sparse attention module includes: Divide the self-attention matrix into non-overlapping local blocks, and each local block contains a preset number of consecutive frames; Concatenate the attention-weighted features, frame-level uniqueness features, and inter-block diversity features; Fuse the concatenated features through a linear layer and output the frame local importance representation. The frame local importance representation F local is calculated according to the following formula: F local = Linear(Concat(A f ,U f ,D f )) where Af Represents the attention-weighted feature, U f Represents the frame-level distinctiveness feature, D f Represents the inter-block diversity feature. Concat represents the feature concatenation operation, and Linear represents the linear layer fusion operation.

[0011] Preferably, the preset number of frames included in each local block of the local block diagonal sparse attention module is 60 frames.

[0012] Preferably, the frame-level distinctiveness feature U in the local block diagonal sparse attention module f Is obtained by calculating the attention entropy, and the calculation formula is as follows: , Where, p i Represents the attention weight of the i th frame, N represents the number of frames in the local block; the inter-block diversity feature D f Is obtained by calculating the cosine similarity, and the calculation formula is as follows: , Where, F i And F j Respectively represent the feature vectors of the i-th and j-th frames in the local block. (•) represents the vector dot product, and ∣∣ represents the Euclidean norm of the vector.

[0013] Preferably, the calculation formula of the adaptive weighted fusion is as follows: F fused = β × F global + (1-β) × F local Where, F fused Represents the fused feature representation, F global Represents the output of the global multi-scale encoding module, F local Represents the output of the local block diagonal sparse attention module, and β is the adaptive weight coefficient, satisfying 0 <β <1.

[0014] Preferably, the regression network includes the following layers: the first Dropout layer, the first normalization layer, the first fully connected layer, the ReLU activation layer, the second Dropout layer, the second normalization layer, the second fully connected layer, and the Sigmoid activation layer.

[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. The present invention introduces a global context vector through a video feature optimizer, and combines an attention weighting mechanism to dynamically adjust the weights of frame-level feature vectors, effectively strengthening the expression of important frames and key features in the video. This design can significantly improve the model's ability to focus on the core content, making the video summary generation more focused on key information, thereby enhancing the representativeness and accuracy of the summary.

[0016] 2. The global multi-scale encoding module of the present invention combines the multi-head attention mechanism with multi-scale depthwise separable convolution operations, capable of simultaneously capturing the global long-range dependencies and local short-term details of the video. This comprehensive modeling method enables the generated summary to comprehensively and meticulously reflect the semantic content of the video, avoiding the problem of information loss caused by relying solely on a single feature extraction method.

[0017] 3. The local block-diagonal sparse attention module of the present invention sparsifies the self-attention matrix into a block-diagonal structure, only focusing on non-overlapping local frame blocks, significantly reducing the computational complexity. At the same time, the frame-level uniqueness feature and the inter-block diversity feature are introduced to enhance the discriminability of frame importance prediction, thereby significantly improving the diversity and accuracy of the summary, while also improving the computational efficiency.

[0018] 4. The present invention adopts an adaptive weighted fusion mechanism to fuse the outputs of the global multi-scale encoding module and the local block-diagonal sparse attention module, and outputs the frame-level importance scores through an efficient regression network. This design ensures that the generated summary has a high compression rate and expressiveness while retaining important information, meeting the dual requirements of video summary quality and efficiency in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Other features, objectives, and advantages of the present invention will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 is a flowchart of the implementation of the video summary generation method based on global multi-scale encoding and local sparse attention provided by the present invention; Figure 2 is an overall framework diagram of the video summary generation method based on global multi-scale encoding and local sparse attention provided by an embodiment of the present invention; Figure 3 is a schematic structural diagram of the video feature optimizer provided by an embodiment of the present invention; Figure 4 is a schematic structural diagram of the global multi-scale encoding module provided by an embodiment of the present invention; Figure 5 is a schematic structural diagram of the local block-diagonal sparse attention module provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made. These all belong to the protection scope of the present invention.

[0021] As Figure 1 shown, it is a flowchart of the implementation of the video abstract generation method based on global multi-scale coding and local sparse attention provided by the present invention, including the following steps: S1. Read the input video and extract the frame-level feature vectors of the input video; S2. Construct a video abstract generation model and input the frame-level feature vectors into the video abstract model; S3. Perform adaptive fusion on the outputs of the global multi-scale coding and local sparse attention modules to generate fused features; S4. The fused features pass through a regression network to output frame importance scores, select representative frames, and generate a video abstract.

[0022] Embodiment 1:

[0023] The present invention provides a preferred embodiment to execute S1, which is applicable to video data of any duration and type. The specific steps are as follows: First, pre-sample the original video at a sampling rate of two frames per second to obtain a video frame sequence V = {v1, v2,..., v n}, where v i represents the i-th frame image in the video, and n is the total number of frames of the video. Then, use the ImageNet pre-trained GoogLeNet network as a feature extractor to perform visual feature encoding on each frame. Specifically, extract the 1024-dimensional feature vector output by the pool5 layer of the GoogLeNet network as the feature representation of the corresponding frame. Thus, a frame feature sequence X = [x1, x2,..., x n is obtained, where x i represents the feature vector of the i-th frame, and n is the number of video frames, which is used to characterize the visual content contained in this frame.

[0024] Embodiment 2: The present invention provides a preferred embodiment to execute S2. The purpose of this embodiment is to use the designed video abstract generation model to solve the problems of high computational complexity and unsatisfactory abstract quality in the prior art when processing long videos by effectively balancing the trade-off between global dependency modeling and local detail capture, and at the same time improve the model's comprehensive understanding and representation ability of video content, so as to more accurately predict the importance score of each frame. As Figure 2As shown, it is the overall framework diagram of the network model of this embodiment. This network consists of three parts: a video feature optimizer, a global multi-scale encoding module, and a local block diagonal sparse attention module. The specific construction steps of the three parts are as follows: S21. Construct a video feature optimizer. As Figure 3 shown, specifically, this optimizer first receives the input feature sequence X = [x1, x2,…, x n , where x i represents the feature vector of the i-th frame, and n is the number of video frames. The module first performs global average pooling on all frame feature vectors to generate a global context vector C, and its calculation method is: .

[0025] This global context vector captures the overall semantic information of the video and serves as the query vector in the subsequent attention mechanism to guide the modeling of inter-frame relationships.

[0026] Next, the module inputs the input feature sequence and the global context vector into the linear transformation layer through the attention mechanism respectively to generate a key vector K, a query vector Q, and a value vector V. Specifically: K = W K X, Q = W Q C, V = W V X, where W K , W Q , W V are learnable weight matrices. To calculate the attention weights, the module adopts the scaled dot-product attention method, multiplies the query vector Q by the key vector K, scales it, and normalizes it through the softmax function, that is: , where d represents the dimension of the query vector Q and the key vector K.

[0027] In implementation, the module fixes the scaling coefficient to 0.03125.

[0028] After completing the attention weighting, the module calculates the weighted feature output Y, and its calculation method is: Y = AV. To further improve the stability of the feature distribution, the module applies layer normalization to the weighted features, and its normalization formula is: Z = LayerNorm(Y), and the final optimized feature sequence Z is output, with a dimension of n×d. This optimized feature sequence contains the dynamically weighted information guided by the global context, can effectively highlight the important frame features in the video, and provides more discriminative inputs for the subsequent global multi-scale encoding module and local block diagonal sparse attention module.

[0029] S22. Construct a global multi-scale encoding module. As Figure 4As shown, it is a schematic diagram of the global multi-scale encoding module. This module takes the optimized feature sequence Z as input. First, through a linear transformation operation, query matrices , key matrices and value matrices are generated respectively, and the specific definitions are as follows: .

[0030] Subsequently, the module uses the multi-head self-attention mechanism to model the global inter-frame dependencies. The calculation formula for its attention weights is: , where d k is the key vector dimension of each attention head, and the softmax operation is normalized along the last dimension to obtain the relative importance distribution between different frames. The output result of the multi-head attention is integrated as: G = Concat(head1,…,head h )W G where head i = A i ’ V i ’ , h represents the number of attention heads, and W G is the output mapping matrix.

[0031] The module further uses the multi-scale depthwise separable convolution operation to extract the local feature representation F. The specific calculation formula is as follows: F = F3+F5 where F3 = Pointwise(D3(V ’ )) and F5 = Pointwise(D5(V ’ ))), D k represents the depthwise convolution operation with a kernel size of k, and Pointwise represents the pointwise convolution; Finally, the global feature representation G and the local feature representation F are adaptively feature fused to generate the final global multi-scale encoding representation F global , and the specific calculation formula is F global = AFF(G,F), where AFF() represents the adaptive feature fusion function.

[0032] S23, construct a local block diagonal sparse attention module. As Figure 5 shown, it is a schematic diagram of the local block diagonal sparse attention module. Specifically, this module divides the self-attention matrix into non-overlapping local blocks, and each local block contains a preset number of consecutive frames; In this embodiment, the preset number of frames included in each local block is 60 frames, which is the optimal number of frames obtained based on experimental verification and can fully capture the feature correlation between local frames while ensuring computational efficiency. Subsequently, the attention-weighted features, frame-level uniqueness features, and inter-block diversity features are spliced. Among them, the frame-level uniqueness feature U f is calculated through attention entropy, and the calculation formula is as follows: , where, p i represents the attention weight of the i -th frame, and N represents the number of frames in the local block; among them, the inter-block diversity feature D f is calculated through cosine similarity, and the calculation formula is as follows: , where, F i and F j respectively represent the feature vectors of the i-th and j-th frames in the local block, (•) represents the vector dot product, and ∣ ∣ represents the Euclidean norm of the vector.

[0033] Finally, the spliced features are fused through a linear layer to output the frame local importance representation, and the frame local importance representation F local is calculated according to the following formula: F local = Linear(Concat(A f ,U f ,D f )) where, A f represents the attention-weighted feature, U f represents the frame-level uniqueness feature, D f represents the inter-block diversity feature, Concat represents the feature splicing operation, and Linear represents the linear layer fusion operation.

[0034] Embodiment 3: The present invention provides a preferred embodiment to execute S3. The output of the global multi-scale encoding module and the output of the local block diagonal sparse attention module are adaptively weighted and fused to generate a fused feature representation, and the specific calculation formula is as follows: F fused = β × F global + (1-β) × F local where, F fused represents the fused feature representation, F global represents the output of the global multi-scale encoding module, F localIt represents the output of the local block diagonal sparse attention module, and β is the adaptive weight coefficient, satisfying 0 < β < 1.

[0035] Example 4: The present invention provides a preferred embodiment to execute S4. In this embodiment, the regression network is sequentially connected to the following layers: the first Dropout layer, the first normalization layer, the first fully connected layer, the ReLU activation layer, the second Dropout layer, the second normalization layer, the second fully connected layer, and the Sigmoid activation layer. The specific calculation process is as follows: , Among them, y represents the input feature, Dropout_1 and Dropout_2 respectively represent the first and second Dropout layers, Norm_1 and Norm_2 respectively represent the first and second normalization layers, Linear_1 and Linear_2 respectively represent the first and second fully connected layers, ReLU represents the ReLU activation function, σ represents the Sigmoid activation function, represents the output frame-level importance score.

[0036] To verify the effectiveness of the above embodiments, the present invention is applied in practice, and the F-score (%) is calculated to compare with other advanced methods. Specifically, the benchmark datasets SumMe and TVSum are used to evaluate the network.

[0037] The SumMe dataset contains 25 videos with video durations ranging from 1 minute to 6 minutes, covering a variety of scenarios and topics. Each video has been manually annotated by 15 to 18 users, and these annotations provide a high-quality reference standard for the generation of video summaries. The TVSum dataset is even larger, containing 50 videos with durations ranging from 2 minutes to 10 minutes. Each video in this dataset has been annotated by 20 users with frame-level importance scores, and these annotation data provide rich details and diversity for model training and evaluation.

[0038] During the experiment, the present invention adopts the standard 5-fold cross-validation (5FCV) method to ensure the robustness and reliability of the results. Specifically, all videos are evenly divided into 5 subsets. In each experiment, 4 of these subsets are selected as the training set for model training and optimization; the remaining 1 subset is used as the test set to evaluate the performance of the model. In this way, we conduct 5 independent experiments, and each experiment uses a different subset as the test set to ensure the comprehensiveness and objectivity of the evaluation results. Finally, we take the average value of the results of these 5 experiments as the final result to accurately reflect the performance of the method of the present invention.

[0039] Table 1 Comparison with advanced methods Comparison results

[0040] Through experiments on the SumMe and TVSum datasets, the experimental results are shown in Table 1. The comparison results between the method of the present invention and other advanced methods show that the present invention has achieved a significant improvement in the F-score, which fully demonstrates the effectiveness and superiority of the present invention in the video summarization task. By introducing the global multi-scale encoding and local sparse attention mechanism, the present invention can more accurately capture the global semantic information and local detail features in the video, thereby generating a more representative and accurate video summary. In addition, the model of the present invention shows good stability and generalization ability during the training and testing processes, and can adapt to different types of video data, providing a reliable solution for video summarization in practical applications.

Claims

1. A video summary generation method based on global multi-scale encoding and local sparse attention, characterized in that Including the following steps: S1: Read the input video and extract the frame-level feature vectors of the video frames; S2: Input the frame-level feature vectors into a video summary generation model to predict the frame-level importance scores. The video summary generation model includes: Video feature optimizer: The video feature optimizer receives the frame-level feature vectors, generates a global context vector by performing global average pooling, performs attention weighting on the global context vector and the frame-level feature vectors, and outputs an optimized feature sequence; Global multi-scale encoding module: The global multi-scale encoding module receives the optimized feature sequence, extracts the global feature representation through the multi-head attention mechanism, and extracts the local feature representation through the multi-scale depthwise separable convolution operation, and extracts the global semantic information in the video by fusing the global feature representation and the local feature representation; Local block diagonal sparse attention module: The output of the global multi-scale encoding module serves as the input to the local block diagonal sparse attention module. The local block diagonal sparse attention module sparsifies the self-attention matrix into a block diagonal structure, and models the local dependencies by concatenating the attention-weighted features, frame-level uniqueness features, and inter-block diversity features; S3: Perform adaptive weighted fusion on the output of the global multi-scale encoding module and the output of the local block diagonal sparse attention module to generate a fused feature representation; S4: Input the fused feature representation into a regression network, output the frame-level importance scores, select the most representative frames, and generate the final video summary.

2. The video abstract generation method based on global multi-scale coding and local sparse attention according to claim 1, wherein The video feature optimizer includes: Perform global average pooling on the frame-level feature vectors to generate a global context vector C. The specific calculation formula is as follows: , Among them, x i represents the feature vector of the i frame, and n is the number of video frames; Perform attention weighting on the global context vector and the frame-level feature vectors to generate a weighted feature sequence Y; Perform layer normalization processing on the weighted feature sequence and output the optimized feature sequence Z. The specific calculation formula is: Z = LayerNorm(Y).

3. The video abstract generation method based on global multi-scale coding and local sparse attention according to claim 1, wherein The global multi-scale encoding module includes: Calculate the global feature representation G of the optimized feature sequence using the multi-head attention mechanism. The specific calculation formula is: G = Concat(head1,…,head h )W G Among them, head i = A i ’ V i ’ , A i ’ is the attention weight obtained by the global multi-scale encoding module modeling the global inter-frame dependency relationship through the multi-head self-attention mechanism, V ’ is the value matrix obtained through linear transformation, h is the number of attention heads, W G is the output mapping matrix; Extract the local feature representation F using the multi-scale depthwise separable convolution operation. The specific calculation formula is as follows: F = F3+F5 Among them, F3 = Pointwise(D3(V ’ )) and F5 = Pointwise(D5(V ’ )) where D k represents a depth convolution operation with a kernel size of k, and Pointwise represents pointwise convolution; Finally, the global feature representation G and the local feature representation F are adaptively feature fused to generate the final global multi-scale encoded representation F global .

4. The video summary generation method based on global multi-scale coding and local sparse attention according to claim 1, wherein The local block diagonal sparse attention module divides the self-attention matrix into non-overlapping local blocks, and each local block contains a preset number of consecutive frames; concatenate the attention-weighted features, frame-level uniqueness features, and inter-block diversity features; Fuse the features after splicing through a linear layer to output the frame local importance representation, where the frame local importance representation is F local Calculate according to the following formula: F local = Linear(Concat(A f ,U f ,D f )) Among them, A f represents the attention-weighted feature, U f represents the frame-level uniqueness feature, D f represents the inter-block diversity feature, Concat represents the feature concatenation operation, and Linear represents the linear layer fusion operation.

5. The video abstract generation method based on global multi-scale coding and local sparse attention according to claim 4, wherein The preset number of frames contained in each local block in the local block diagonal sparse attention module is 60 frames.

6. The video abstract generation method based on global multi-scale encoding and local sparse attention according to claim 4, wherein The frame-level uniqueness feature U in the local block diagonal sparse attention module f is obtained by calculating the attention entropy, and the calculation formula is as follows: , Among them, p i represents the attention weight of the i th frame, and N represents the number of frames in the local block; the inter-block diversity feature D f is calculated by cosine similarity, and the calculation formula is as follows: , Among them, F i and F j respectively represent the feature vectors of the $i$-th and $j$-th frames in the local block, ($\cdot$) represents the dot product of vectors, and $\vert\vert$ represents the Euclidean norm of the vector.

7. The video abstract generation method based on global multi-scale coding and local sparse attention according to claim 1, wherein The calculation formula for the adaptive weighted fusion is as follows: F fused = β × F global + (1-β) × F local, Among them, F fused represents the fused feature representation, F global represents the output of the global multi-scale encoding module, F local represents the output of the local block diagonal sparse attention module, and β is an adaptive weight coefficient satisfying 0 < β < 1.

8. The video summary generation method based on global multi-scale coding and local sparse attention according to claim 1, wherein The regression network includes the following layers: the first Dropout layer, the first normalization layer, the first fully connected layer, the ReLU activation layer, the second Dropout layer, the second normalization layer, the second fully connected layer, and the Sigmoid activation layer.

Citation Information

Patent Citations

  • Video abstract generation method fusing local target features and global features

    CN113139468A

  • Video abstraction method based on graph model and multi-scale attention mechanism

    CN120050491A

  • Video abstraction method based on global and local convolution attention mechanism

    CN120091197A

Cited By

  • Multi-source data fusion-based intelligent detection method for transportation state of combined transportation of iron and water

    CN120873552A

  • Contact network video abnormal frame detection system based on local multi-head attention mechanism

    CN121747012A

  • Video abstract generation method based on multi-scale kernel pooling and frequency domain interactive attention

    CN122227046A