Video summarization method and system based on multi-temporal granularity feature fusion
Through the G-MMT network, combined with global and multi-granularity temporal feature extractors, the problems of insufficient capture of long-distance dependencies and low computational efficiency in existing methods are solved, and the efficient and accurate extraction and generation of video summaries are achieved.
Patent Information
- Application Number
- CN202410556037.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-07
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2044-05-07
AI Technical Summary
Existing video summarization methods suffer from the vanishing gradient problem of RNN networks when processing long videos, making it difficult to capture long-distance dependencies. Transformer-based methods also ignore differences in shot duration, resulting in the summary lacking important details and high computational cost.
The Transformer-based G-MMT network is used, combined with a global temporal feature extractor and a multi-granularity temporal feature extractor. Through local and global attention modules, the temporal dependencies at different time granularities are modeled to extract the key frames of the video.
The accuracy and computational efficiency of video summarization are improved, and the temporal correlation between video frames can be better captured to generate more representative summaries.
Smart Images

Figure CN118400590B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video summarization, and in particular to a video summarization method based on multi-time granularity feature fusion. Background Art
[0002] In recent years, with the advancement of video capture technology and equipment, the amount of video on the internet has grown rapidly. According to relevant statistics, in 2022, online video accounted for 82% of total internet traffic, with a growing number of video genres, including movies, documentaries, and sports events. Short videos, in particular, have become popular due to their short duration and concise content. The ability to efficiently compress a long video into short clips, enabling users to quickly browse through them, is in high demand. Video summarization is a technology that automatically generates short videos by extracting key information from a video. Applying video summarization to long videos, such as movies and sports events, can significantly compress the video while preserving the main points of the video. Therefore, video summarization is a major research hotspot for enabling faster video browsing.
[0003] The key to video summarization lies in accurately extracting key information from a video. To achieve this, it's first necessary to properly model the temporal relationships within the video content. Previous work has primarily used deep networks centered around recurrent neural networks (RNNs) to model video temporal relationships. These approaches treat a video as a sequence of frames, employing RNNs and their variants (such as LSTM and GRU) to capture the contextual relationships between frames and assess the importance of each frame to effectively extract key information. However, RNNs suffer from the problem of short-term memory. When processing long sequences, the gradients can easily become very small during backpropagation, leading to a vanishing gradient problem. Furthermore, video summarization often involves long videos, making it difficult for RNNs to capture the temporal dependencies between frames that are far apart in time. To circumvent this problem, RNN-based work divides the video into multiple segments and processes each segment separately using an RNN. While these approaches mitigate the short-term memory issue of RNNs to some extent, since the computation of each frame depends on the previous frame, this recurrent structure makes RNN-based approaches computationally inefficient during training and inference. Unlike RNNs, Transformers, based on a self-attention mechanism, can simultaneously interact with information at all positions in a sequence, effectively capturing long-range dependencies within the sequence. In video summarization, Transformer-based methods can better understand the correlations between different frames in a video and consider the context of the entire video, generating more representative summaries than RNN-based methods.
[0004] However, these Transformer-based methods lack consideration for the varying durations of different shots within a video. A video typically contains multiple shots of varying lengths. Modeling temporal relationships from a global perspective tends to ignore local connections within the sequence, especially within short shots, resulting in the final generated summary lacking some important details. Some work, based on the KTS method, segments the video according to the boundaries of different shots and then extracts local connections within each shot sequence through self-attention or combined with LSTM methods. However, KTS segments the video by calculating the similarity between frames, which is computationally expensive and, in turn, affects the computational cost of the entire summary model.
[0005] To address these issues, the present invention considers extracting global and local temporal dependencies from videos at different temporal granularities, enabling the network to better adapt to temporal relationships between different distances. Furthermore, to enhance important local details in the video, the present invention designs a local attention module to model short-range dependencies between video segments. Therefore, the present invention proposes a video summarization method based on multi-temporal granularity feature fusion, which consists of a global temporal feature extractor and three multi-granularity temporal feature extractor branches. Summary of the Invention
[0006] The purpose of the present invention is to overcome the above-mentioned deficiencies in the existing methods and propose a video summarization method based on multi-time granularity feature fusion to achieve accurate extraction of video key frames and generation of video summaries.
[0007] To achieve the above object, the present invention provides the following solutions:
[0008] The present invention provides a video summarization method based on multi-time granularity feature fusion, the method comprising:
[0009] S1. Obtain the original video data and labels, and perform preprocessing to obtain the training set and test set;
[0010] S2. Build the G-MMT network;
[0011] S3. Using the preprocessed training set to train the G-MMT network, wherein the network structure is based on the Transformer structure and, according to the characteristics of the network structure, the G-MMT network is trained using a strategy based on multi-time granularity grouping;
[0012] S4. Use the trained G-MMT network to process the video data and obtain the video summary result.
[0013] Furthermore, in another preferred embodiment, the above step S1 is specifically as follows:
[0014] S11, averaging the frame-level importance labels given by different users for the video to obtain an averaged frame-level importance label;
[0015] S12, downsampling the video and the averaged frame-level importance labels, taking one frame every 15 frames, and extracting the corresponding importance labels;
[0016] S13, normalizing the importance label to the range of [0, 1];
[0017] S14. 80% of all videos are used as training data to form a training set, and the remaining 20% are used as test data to form a test set.
[0018] Furthermore, in another preferred embodiment, the G-MMT network in the above step S2 includes a global temporal feature extractor and three independent multi-granularity temporal feature extractor branches.
[0019] Furthermore, in another preferred embodiment, the global temporal feature extractor first extracts image features from each frame of the training video using a pre-trained GoogLeNet network to obtain extracted temporal features; then uses a Transformer structure to model the temporal features within the video; finally, before modeling the temporal features of the frame sequence, position encoding is added to all frame features, and a multi-granularity temporal grouping scheme for the frame sequence is adopted;
[0020] In the above three independent multi-granularity temporal feature extractor branches, the original feature sequence is grouped non-overlappingly according to different grouping intervals t, and each consecutive t elements in the feature sequence are taken as a row and converted into an array form. Each row in the array is a local sequence group with a time granularity of t, which further models the local features of the original video, and each column is a global sequence group with a time granularity of t, which further models the global features of the original video.
[0021] Furthermore, there is a preferred embodiment in which the multi-granularity temporal feature extractor includes a local attention module and an auxiliary global attention module.
[0022] Furthermore, in another preferred embodiment, the above step S3 is specifically as follows:
[0023] The feature sequences that have been updated at multiple time granularities are fused separately to obtain a new feature sequence, which is then supervised by labels for training. During the training process, the error between the output result of each branch and the labeled result is calculated using the MSE loss function, and the losses of each branch are superimposed with different weights for subsequent training. At the same time, the parameters in the network are updated using the Adam optimization algorithm through the gradient backpropagation method to obtain a trained G-MMT network.
[0024] Furthermore, in a preferred embodiment, the above-mentioned step S4 is specifically as follows: using the trained G-MMT network model to predict the importance score of each frame of the test sample, and calculating the importance score of each segment based on the frame-level importance score combined with the time boundary of each segment of the video; the segment with a higher score is the key segment of the video; and the segments that account for no more than 15% of the total number of frames of the video are selected and combined in chronological order to form a new summary video.
[0025] The video summarization method based on multi-temporal granularity feature fusion described in the present invention can be fully implemented using computer software. Therefore, the present invention also provides a video summarization system based on multi-temporal granularity feature fusion, the system comprising:
[0026] A storage device for obtaining raw video data and labels, and performing preprocessing to obtain training sets and test sets;
[0027] Storage device for building a G-MMT network;
[0028] A storage device for training a G-MMT network using a preprocessed training set, wherein the network structure is based on a Transformer structure and, according to the characteristics of the network structure, a strategy based on multi-time granularity grouping is used to train the G-MMT network;
[0029] A storage device for finally using the trained G-MMT network to process video data and obtain video summary results.
[0030] The present invention also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program executes any one of the above-mentioned video summarization methods based on multi-time granularity feature fusion.
[0031] The present invention also provides a computer device, which includes a memory and a processor, wherein a computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes any one of the above-mentioned video summarization methods based on multi-time granularity feature fusion.
[0032] The beneficial effects of the present invention are:
[0033] 1. The present invention provides a video summarization method based on multi-temporal granularity feature fusion. It adopts a network structure with Transformer as the basic framework, models the long-distance dependency relationship between video frames, uses a global temporal feature extractor to preliminarily extract the global dependency relationship, and further extracts global information at multiple temporal granularities to enhance the global nature of the temporal dependency relationship.
[0034] 2. This paper provides a video summarization method based on multi-temporal feature fusion, and proposes a G-MMT network for keyframe extraction. This network uses a multi-granular temporal feature extractor to model short-range dependencies between video clips, fully extracting local events in the video and enhancing the representativeness of keyframes in the clips.
[0035] The present invention is applicable to the field of video summarization. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 This is a schematic diagram of a specific process of the video summarization method based on multi-time granularity feature fusion according to the first exemplary embodiment of the present invention;
[0037] Figure 2 Schematic diagram of the G-MMT network structure according to the third exemplary embodiment of the present invention;
[0038] Figure 3 Schematic diagram of the structure of the multi-granularity temporal feature extractor described in the third exemplary embodiment of the present invention.
[0039] Among them, Input video is the input video, Positional Encoding is the position encoding, Global Transformer is the global transformer, Goog LeNet is the Google network, Frame Embedding is the frame embedding, Multi-Matrix Transformer is the Multi-Matrix transformer, Fusion is the fusion, Regression Network is the regression network, Local Attcntion is the local attention module, and Secondary Global Attention is the auxiliary global attention module. DETAILED DESCRIPTION
[0040] The specific embodiments of the present invention are further described in detail below with reference to the accompanying drawings and examples. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several changes and improvements can be made without departing from the scope of the present invention, and these all fall within the scope of protection of the present invention.
[0041] Implementation method 1, see Figure 1 This embodiment describes a video summarization method based on multi-time granularity feature fusion, which is as follows:
[0042] S1. Obtain the original video data and labels, and perform preprocessing to obtain the training set and test set;
[0043] S2. Build the G-MMT network;
[0044] S3. Using the preprocessed training set to train the G-MMT network, wherein the network structure is based on the Transformer structure and, according to the characteristics of the network structure, the G-MMT network is trained using a strategy based on multi-time granularity grouping;
[0045] S4. Use the trained G-MMT network to process the video data and obtain the video summary result.
[0046] When this embodiment is actually applied, first of all, this embodiment needs to predict the importance scores of all video frames, so this example requires frame-level importance labels. To this end, this embodiment downsamples the video frames based on the fact that the picture differences between adjacent frames in the actual video are small, reduces the information redundancy between video frames, and reduces the amount of calculation; then, a G-MMT network is constructed with Transformer as the basic structure to achieve multi-level temporal feature extraction, better capture the temporal correlation between video frames, and thus improve the accuracy of key frame extraction. The G-MMT used in this embodiment includes three time granularity branches in the multi-granularity temporal feature extractor to extract temporal features of three different time granularities respectively. The network improves the integrity of video temporal features by modeling the temporal dependency between local sequences and global sequences at different time granularities.
[0047] like Figure 1 As shown, the video summarization method includes the following steps:
[0048] S1. Obtain the original video data and labels, and perform preprocessing to obtain the training set and test set;
[0049] S2. Build the G-MMT network;
[0050] S3. Using the preprocessed training set to train the G-MMT network, the network structure is based on the Transformer structure. Based on the characteristics of the network structure, the G-MMT network is trained using a strategy based on multi-time granularity grouping. During the training process, the model with the best performance on the validation set is saved for testing.
[0051] S4. Use the trained G-MMT network to process the video data and obtain the video summary result.
[0052] This embodiment proposes a G-MMT network for key frame extraction. The network uses a multi-granularity temporal feature extractor to model the short-range dependencies of video clips, fully extracting local events in the video and enhancing the representativeness of key frames in the clips.
[0053] Implementation 2: This implementation is an example of step S1 in the video summarization method based on multi-time granularity feature fusion described in Implementation 1.
[0054] The step S1 is specifically as follows:
[0055] Step S11: averaging the frame-level importance labels given by different users to obtain an averaged frame-level importance label;
[0056] Step S12: downsample the video and the frame-level importance annotations, take one frame every 15 frames, and extract the corresponding importance annotations;
[0057] Step S13: normalize the importance label to the range of [0, 1];
[0058] Step S14: 80% of all videos are used as training data to form a training set, and the remaining 20% are used as test data to form a test set.
[0059] In practical applications of this embodiment, step S1 requires preprocessing the labels of the training set. In addition, different videos vary greatly in length and content. This step preprocesses the video data by downsampling the video frames and normalizing the values to obtain video data suitable for network training. Specifically, it includes the following steps:
[0060] S11. Based on the frame-level importance labels given by different users for the video, the labels are averaged to obtain the averaged frame-level importance labels, so as to predict the importance scores of all video frames.
[0061] S12. Downsample the video and the averaged frame-level importance annotations, take one frame every 15 frames, and extract the corresponding importance annotations, so as to reduce information redundancy between video frames and reduce the amount of calculation.
[0062] S13, normalizing the importance label to the range of [0, 1];
[0063] S14. 80% of all videos are used as training data to form a training set, and the remaining 20% are used as test data to form a test set.
[0064] Implementation method three, see Figure 2 and Figure 3 This embodiment is described as follows: This embodiment is an example of the G-MMT network in step S2 of the video summarization method based on multi-time granularity feature fusion described in the second embodiment;
[0065] The G-MMT network consists of a global temporal feature extractor and three independent multi-granularity temporal feature extractor branches.
[0066] In practical applications, the G-MMT network consists of a global temporal feature extractor and three independent multi-granularity temporal feature extractor branches. When a video is fed into the network, the global temporal feature extractor initially models the temporal relationships between all video frames. After transforming the feature sequence according to different temporal granularities, the feature sequences are fed into three independent multi-granularity temporal feature extractors to extract temporal features at three different temporal granularities. These multi-granularity temporal features are then fused and trained using frame-level importance labels to train the entire keyframe extraction network.
[0067] Among them, the G-MMT network structure is as follows: Figure 2 As shown,
[0068] It uses the Transformer as its basic structure and includes two modules: a global temporal feature extractor and a multi-granularity temporal feature extractor. The global temporal feature extractor includes a pre-trained GoogLeNet and a Transformer. The multi-granularity temporal feature extractor includes three temporal feature extraction branches at different temporal granularities, each of which contains a local attention module and an auxiliary global attention module.
[0069] In the global temporal feature extractor, all video frames of a video are sequentially encoded by GoogLeNet to extract high-level semantic features of each frame. Adding position encoding PE to the frame features helps to strengthen the temporal relationship and relative position relationship of the frames, thereby obtaining the feature sequence E = [e1, e2, ..., e n ].
[0070] E=GoogleNet(V)+PE
[0071] The core of Transformer is the multi-head self-attention mechanism, which uses multiple sets of coefficient matrices W Q ,W K and W V Perform a linear transformation on E to obtain the corresponding Query (Q), Key (K) and Value (V), and calculate the attention score for each head. In this implementation, the feature dimension d is set to 1024, and the number of heads m is set to 12 in the multi-head self-attention mechanism. For each head result H i After concatenation and MLP, we get X = [x1, x2, ..., x n ].
[0072] In the multi-granularity temporal feature extractor, each temporal granularity branch contains a local attention module and an auxiliary global attention module, such as Figure 3As shown. First, the feature sequence X is grouped with a grouping interval of t. When the grouping interval is t, each consecutive t elements in X are taken as a row, and X is transformed. (t) For array X (t) , and sequentially put the i-th row As the input of the local attention module, to extract The local features of each frame in the image are obtained by using the coefficient matrix and Calculate the corresponding and and The product of reflects the similarity between the features of each frame. The similarity matrix is normalized by the softmax function to obtain the attention weight, which is then combined with Multiply and calculate the self-attention features of each head. After concatenating the results of each head and passing through MLP, we get The result of extracting local features through the local attention module
[0073]
[0074] Calculate all After that, arrange them in order, and you can get the same as X (t) Result arrays of the same size
[0075]
[0076] X (t) The i-th column As the input of SGA, we extract x i The global features at the time granularity of t. The process of assisting global attention is similar to local attention, through multiple sets of coefficient matrices and Calculate the corresponding and Then perform self-attention operation and pass MLP to obtain the time granularity of t Global features of each frame
[0077] Calculate all After that, arrange them in order, and you can get the same as X (t) Result arrays of the same size
[0078]
[0079] Then and Fusion is performed to obtain the feature sequence R of the time frame sequence X with a time granularity of t (t) .
[0080] After the multi-granularity temporal feature extractor, local frame features at various time granularities can be obtained. and global frame features The time granularity t obtained and to integrate;
[0081]
[0082] R (t) Indicates the feature fusion result when the time granularity is t, The R obtained at multiple time granularities (t) Fusion is performed to obtain the final frame feature Y=[y1,y2,…,y n ]. Y is used as the input of the regression network, and the regression network predicts the frame-level importance score P = [p1, p2, ..., p n ].
[0083] This embodiment proposes a G-MMT network for key frame extraction. The network uses a multi-granularity temporal feature extractor to model the short-range dependencies of video clips, fully extracting local events in the video and enhancing the representativeness of key frames in the clips.
[0084] Implementation 4: This implementation is an example of a global temporal feature extractor in a video summarization method based on multi-time granularity feature fusion described in Implementation 3.
[0085] The global temporal feature extractor first extracts image features from each frame of the training video using a pre-trained GoogLeNet network to obtain extracted temporal features. It then uses a Transformer structure to model the temporal features within the video. Finally, before modeling the temporal features of the frame sequence, it adds position encoding to all frame features and adopts a multi-granularity temporal grouping scheme for the frame sequence.
[0086] In the three independent multi-granularity temporal feature extractor branches, the original feature sequence is grouped non-overlappingly according to different grouping intervals t, and each consecutive t elements in the feature sequence are taken as a row and converted into an array form. Each row in the array is a local sequence group with a time granularity of t, which further models the local features of the original video, and each column is a global sequence group with a time granularity of t, which further models the global features of the original video.
[0087] In practical applications, the global temporal feature extractor first extracts image features from each frame of the training video using a pre-trained GoogLeNet network. Since the temporal correlation between video frames cannot be captured by the image feature encoder alone, the present embodiment uses the Transformer structure to model the temporal features within the video. To strengthen the temporal relationship between video frames, position encoding is added to all frame features before temporal feature modeling of the frame sequence. Since the features of adjacent frames in a video tend to be similar, while there is often a certain time interval between key frames in a shot, a multi-granularity temporal grouping scheme for the frame sequence can be used to expand the receptive field of frame feature extraction and improve the ability to capture key frames within a shot. Furthermore, the duration of each shot in a video is different, and the global feature focuses on longer-duration shot segments while ignoring shorter shot segments, and lacks attention to important details within local segments. Therefore, the present embodiment uses three temporal feature extractor branches with different time granularities to further model the temporal features between video frames.
[0088] Since the duration of each shot in the video is different, the global features focus on the shot segments with longer duration and ignore the shot segments with shorter duration, and lack attention to the changes in important details in the local segments. Therefore, this embodiment uses three temporal feature extractor branches with different time granularity to further model the temporal features between video frames. In the three independent branch paths, the original feature sequence is grouped non-overlappingly according to different grouping intervals t, and each consecutive t elements in the feature sequence is taken as a row and converted into an array form. Each row in the array is a local sequence group with a time granularity of t, which further models the local features of the original video. Each column is a global sequence group with a time granularity of t, which further models the global features of the original video.
[0089] Implementation 5: This implementation is an example of a multi-granularity temporal feature extractor in a video summarization method based on multi-time granularity feature fusion described in Implementation 3.
[0090] The multi-granularity temporal feature extractor includes a local attention module and an auxiliary global attention module.
[0091] In actual application, this embodiment proposes a local attention module and an auxiliary global attention module for a multi-granularity temporal feature extractor to extract temporal features of different time granularities. Among them, the local attention module takes each row of the feature array as input, and introduces an attention mechanism to construct the temporal dependency between each element in the local feature sequence represented by each row. The auxiliary global attention module takes each column of the feature array as input, and introduces an attention mechanism to construct the temporal dependency between each element in the global feature sequence represented by each column. That is, after converting the frame features from a sequence into different arrays, each array will first pass through the local attention module and the auxiliary global attention module to further extract the temporal information between the frames to identify the key frames. Finally, the features enhanced by the two modules are fused to combine local information and global information.
[0092] Implementation 6: This implementation is an example of step S3 in the video summarization method based on multi-time granularity feature fusion described in Implementation 1.
[0093] The step S3 is specifically as follows:
[0094] The feature sequences that have been updated at multiple time granularities are fused separately to obtain a new feature sequence, which is then supervised by labels for training. During the training process, the error between the output result of each branch and the labeled result is calculated using the MSE loss function, and the losses of each branch are superimposed with different weights for subsequent training. At the same time, the parameters in the network are updated using the Adam optimization algorithm through the gradient backpropagation method to obtain a trained G-MMT network.
[0095] In practical applications, in order to more accurately find the key frames in the video in step S3, this embodiment adopts a supervision strategy based on multi-time granularity feature fusion to train G-MMT. This strategy fuses the feature sequences that have been updated at multiple time granularities, obtains a new feature sequence, and then uses the label for supervised training. During the training process, the error between the output result of each branch and the labeled result is calculated by the MSE loss function, and the losses of each branch are superimposed with different weights for subsequent training. At the same time, the parameters in the network are updated using the Adam optimization algorithm through the gradient backpropagation method to obtain a trained G-MMT network;
[0096] The loss function expression used is as follows:
[0097]
[0098] Where G is the frame-level importance label sequence, g is the importance label of each frame, P is the frame-level importance score sequence predicted by the G-MMT network, and p is each element in P.
[0099] Implementation 7: This implementation is an example of step S4 in the video summarization method based on multi-time granularity feature fusion described in Implementation 1.
[0100] The step S4 is specifically as follows:
[0101] The trained G-MMT network model is used to predict the importance score of each frame of the test sample. The importance score of each segment is calculated based on the frame-level importance score combined with the time boundary of each segment of the video. The segments with higher scores are regarded as the key segments of the video. The segments that account for no more than 15% of the total number of frames in the video are selected and combined in chronological order to form a new summary video.
[0102] In practical application, in step S4, the trained G-MMT network model is used to predict the importance score of each frame in the test video, and the KTS method is used to obtain the segmentation of each video, and then the predicted frame-level importance score p is used. i Calculate the shot-level importance score s i .
[0103]
[0104] Among them, K i Indicates the number of frames contained in the i-th shot. The length of the final generated video summary does not exceed 15% of the original video length. The optimal shot is selected to form the video summary. u i ∈{0,1}, L represents the length of the original video F, l i Represents the length of the i-th shot.
[0105]
[0106] Implementation 9: This implementation is a video summarization method based on multi-time granularity feature fusion described in Implementation 1 to Implementation 8, which extracts key frames from a video and generates a summary video.
[0107] This example uses video data from the public datasets TVSum and SumMe to evaluate the G-MMT model. According to statistics, the TVSum dataset contains a total of 50 videos, with video durations ranging from 2 to 11 minutes, and each video has frame-level importance annotations from 20 users. The SumMe dataset contains a total of 25 videos, with video durations ranging from 1 to 6 minutes, and each video has frame-level importance scores annotated by 15 to 18 users. After preprocessing the experimental data, this example uses 80% of all data as a training set and 20% as a test set to ensure the number of training and test sets in the experiment and avoid overfitting or underfitting.
[0108] During preprocessing, each video frame is first downsampled, extracting one frame every 15 frames from all the frames in a video to reduce the significant redundancy of visual information between consecutive frames. During training, the G-MMT network is trained using the Mean Sequential Error (MSE) loss function to update the parameters of the G-MMT network. This implementation uses the F1 score to evaluate the accuracy of keyframe extraction.
[0109] As shown in the table below, this method demonstrates superior performance compared to existing methods such as Bi-LSTM, SUM-GAN, DR-DSN, SASUM, SUM-FCN, and M-AVS. This demonstrates that the proposed method can more accurately extract key frames from videos, thereby generating more representative summary videos.
[0110] method TVSum SumMe Bi-LSTM 54.2 37.6 SUM-GAN 56.3 41.7 DR-DSN 58.1 42.1 SASUM 58.2 45.3 SUM-FCN 56.8 47.5 M-AVS 61.0 44.4 G-MMT 0.7234 0.7893
[0111] In summary, compared with existing similar methods, the segmentation method described in this embodiment has a more accurate key frame extraction effect and the final summary is more representative.
[0112] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0113] The above description is only a detailed description of the specific embodiments of the present invention, and does not limit the present invention. Various substitutions, modifications and improvements made by those skilled in the relevant art without departing from the principles and scope of the present invention should be included in the scope of protection of the present invention.
Claims
1. A video summarization method based on multi-time granularity feature fusion, characterized in that: The method is: S1. Obtain the original video data and labels, and perform preprocessing to obtain the training set and test set; S2. Build the G-MMT network; The G-MMT network includes a global temporal feature extractor and three independent multi-granularity temporal feature extractor branches; The global temporal feature extractor first extracts image features from each frame of the training video using a pre-trained GoogLeNet network to obtain extracted temporal features; then uses a Transformer structure to model the temporal features within the video. Finally, before modeling the temporal features of the frame sequence, position encoding is added to all frame features, and a multi-granularity temporal grouping scheme for the frame sequence is adopted; In the three independent multi-granularity temporal feature extractor branches, the original feature sequence is divided into different grouping intervals. Perform non-overlapping grouping and group each consecutive elements as a row, converted into an array form, each row in the array is a time granularity of The local sequence group of the original video is further modeled, and each column is the time granularity of The global sequence group is further modeled with the global features of the original video; S3. Using the preprocessed training set to train the G-MMT network, wherein the network structure is based on the Transformer structure and, according to the characteristics of the network structure, the G-MMT network is trained using a strategy based on multi-time granularity grouping; S4. Use the trained G-MMT network to process the video data and obtain the video summary result.
2. The video summarization method based on multi-time granularity feature fusion according to claim 1 is characterized in that: The step S1 is specifically as follows: S11, averaging the frame-level importance labels given by different users for the video to obtain an averaged frame-level importance label; S12, downsampling the video and the averaged frame-level importance labels, taking one frame every 15 frames, and extracting the corresponding importance labels; S13, normalizing the importance label to the range of [0, 1]; S14. 80% of all videos are used as training data to form a training set, and the remaining 20% are used as test data to form a test set.
3. The video summarization method based on multi-time granularity feature fusion according to claim 1 is characterized in that: The multi-granularity temporal feature extractor consists of a local attention module and an auxiliary global attention module.
4. The video summarization method based on multi-time granularity feature fusion according to claim 1 is characterized in that: The step S3 is specifically as follows: The feature sequences that have been updated at multiple time granularities are fused separately to obtain a new feature sequence, which is then supervised by labels for training. During the training process, the error between the output result of each branch and the labeled result is calculated using the MSE loss function, and the losses of each branch are superimposed with different weights for subsequent training. At the same time, the parameters in the network are updated using the Adam optimization algorithm through the gradient backpropagation method to obtain a trained G-MMT network.
5. The video summarization method based on multi-time granularity feature fusion according to claim 1 is characterized in that: The step S4 is specifically as follows: The trained G-MMT network model is used to predict the importance score of each frame of the test sample. The importance score of each segment is calculated based on the frame-level importance score and the time boundary of each segment of the video. The segments with higher scores are considered key segments of the video. Select segments that account for no more than 15% of the total video frames and combine them into a new summary video in chronological order.
6. A video summarization system based on multi-time granularity feature fusion, characterized by: The system comprises: A storage device for obtaining raw video data and labels, and performing preprocessing to obtain training sets and test sets; Storage device for building a G-MMT network; The G-MMT network includes a global temporal feature extractor and three independent multi-granularity temporal feature extractor branches; The global temporal feature extractor first extracts image features from each frame of the training video using a pre-trained GoogLeNet network to obtain extracted temporal features. It then uses a Transformer structure to model the temporal features within the video. Finally, before modeling the temporal features of the frame sequence, it adds position encoding to all frame features and adopts a multi-granularity temporal grouping scheme for the frame sequence. In the three independent multi-granularity temporal feature extractor branches, the original feature sequence is divided into different grouping intervals. Perform non-overlapping grouping and group each consecutive elements as a row, converted into an array form, each row in the array is a time granularity of The local sequence group of the original video is further modeled, and each column is the time granularity of The global sequence group is further modeled with the global features of the original video; A storage device for training a G-MMT network using a preprocessed training set, wherein the network structure is based on a Transformer structure and, according to the characteristics of the network structure, a strategy based on multi-time granularity grouping is used to train the G-MMT network; A storage device for finally using the trained G-MMT network to process video data and obtain video summary results.
7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the video summarization method based on multi-time granularity feature fusion according to any one of claims 1 to 5.
8. A computer device, characterized in that The device includes a memory and a processor, wherein a computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes the video summarization method based on multi-time granularity feature fusion according to any one of claims 1 to 5.
Citation Information
Patent Citations
Video abstraction method based on symmetric multi-scale attention
CN117493607A
Method, apparatus, device and storage medium for training video recognition model
US20230069197A1