Multi-scale spatio-temporal modeling video abstract generation method and device
Through the multi-scale spatiotemporal modeling module, combined with convolutional neural network and importance classifier, the problem that video digests are difficult to capture long-distance time dependence and intra-frame visual significance cues in the prior art is solved, and efficient and low-complexity video digest generation is achieved.
Patent Information
- Application Number
- CN202510840838.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-23
AI Technical Summary
When processing video content, existing video digest technology is difficult to effectively capture long-distance time dependence and intra-frame visual significance cues, and the Transformer-based method has high computational complexity, resulting in inefficiency.
A multi-scale spatiotemporal modeling module, including a multi-scale aggregator, a cascaded time modeling module and a parallel spatial modeling module, is used to extract features through a convolutional neural network, and optimize the importance scores with an importance classifier and a loss function to construct a video summary.
It improves the capture ability of local details and global structure of video, enhances the extraction of long-distance time dependence and intra-frame visual significance cues, reduces the computational complexity, and realizes efficient video summary.
Smart Images

Figure CN120375261A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and particularly to a multi-scale spatio-temporal modeling video summary generation method and apparatus. Background Art
[0002] With the rapid development of Internet technology and the popularization of digital devices, the production and dissemination of video content have shown explosive growth. From short video sharing on social media to a vast amount of film and television resources on online video platforms, videos have become one of the main forms of people's daily information acquisition and entertainment consumption. This growth trend has put forward higher requirements for the management and utilization of video content, and video summary technology has emerged as the times require. The exponential growth of online video content has led to the problem of information overload faced by users. Video summary technology can help users quickly understand the core content of videos, thereby improving the efficiency of video browsing.
[0003] There have been many related studies on video summary technology. Convolutional neural networks (CNNs) play a key role in video understanding. Due to their excellent image feature extraction capabilities, they lay the foundation for video content analysis. Recurrent neural networks (RNNs) and their variants (such as long short-term memory networks LSTMs) have gradually become the mainstream methods in the field of video summary by virtue of their ability to process sequential data and capture temporal dependencies. These networks can effectively utilize the inter-frame correlation to perform frame-level saliency estimation. Transformers can integrate local and global temporal patterns, effectively model the complete video context, while maintaining spatial saliency features. This modeling method endows the video summary system with significantly improved content analysis and summary synthesis capabilities.
[0004] Among these methods, the methods based on RNN and LSTMs mainly focus on local temporal correlations, while ignoring long-distance dependencies and intra-frame visual saliency cues that are consistent with human attention patterns; the methods based on Transformers are limited by the quadratic growth of the inherent computational complexity of Transformers, which poses challenges to sequence-to-sequence applications including video summary. Summary of the Invention
[0005] In view of the defects existing in the prior art, the present invention provides a multi-scale spatio-temporal modeling video summary generation method and apparatus.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows: On the one hand, the present invention provides a multi-scale spatio-temporal modeling video summary generation apparatus, including a feature extraction module, a multi-scale spatio-temporal modeling module, an importance classifier, and a summary generation module; The feature extraction module is used to divide the input video into non-overlapping segments and extract the features of each video segment using a convolutional neural network; The multi-scale spatio-temporal modeling module includes a multi-scale aggregator, a cascaded temporal modeling module, and a parallel spatial modeling module arranged freely in sequence; the multi-scale aggregator uses deconvolution and pooling operations to obtain multi-scale features of the input features, and at the same time adds segment category labels before the input features, and applies a feature fusion strategy to aggregate the multi-scale features to obtain the fused features; the cascaded temporal modeling module consists of multiple cascaded temporal Mamba blocks, and the input features are refined through residual learning in the multiple cascaded temporal Mamba blocks to fuse short-term cues and long-term event dependencies, and the refined features are output; the parallel spatial modeling module consists of multiple parallel spatial Mamba blocks with shared parameters, and each spatial Mamba block is used to process the features of a single frame. After the input features are transformed into frame-level features through a linear layer, they are input into the spatial Mamba blocks, and are processed in parallel in the spatial Mamba blocks through residual learning to extract spatial features, and the spatial features extracted by each spatial Mamba block are aggregated to form a feature representation of the video content; The importance classifier outputs an importance score based on the output features of the multi-scale spatio-temporal modeling module, and optimizes the importance score in combination with a loss function; The summary generation module is used to construct a summary by solving a constrained optimization problem according to the optimized importance score.
[0007] Further, the fused features are obtained according to the following formula: ; where, is the low-scale feature; is the basic-scale feature; is the high-scale feature; is the fused feature; is the weight of the low-scale feature; is the weight of the basic-scale feature; is the weight of the high-scale feature; is a learnable linear frame projection, L represents the low scale, B represents the basic scale, H represents the high scale.
[0008] Further, the input features are refined according to the following formula: ; where, is the input feature of the l th temporal Mamba block, ; is the l th temporal Mamba block; is the l- Input features of 1 temporal Mamba block.
[0009] Furthermore, the feature representation of the video content is formed according to the following formula: ; where, is the feature representation of the video content; is the parallel spatial Mamba block of shared parameters; is the t th frame-level feature, .
[0010] Furthermore, the importance score is output according to the following steps: Perform layer normalization on the output features of the multi-scale spatio-temporal modeling module; Extract higher-level semantic patterns of the layer-normalized features through a fully-connected layer with an activation function; Use function to normalize the higher-level semantic patterns of the features into importance scores and output.
[0011] Furthermore, the loss function is calculated according to the following formula: ; where, is the loss function; is the cross-entropy loss function; is the mean squared error loss function; is an adjustable hyperparameter.
[0012] Furthermore, the cross-entropy loss function is calculated according to the following formula: ; where, t is the index value of the frame; c is the result label of the classification; is the true label of the t th frame; is the prediction result of the frame importance classifier for the t th frame.
[0013] Furthermore, the mean squared error loss function is calculated according to the following formula: ; where, is the predicted importance score; is the true normalized importance score.
[0014] Furthermore, the abstract is constructed according to the following steps: The mean importance score is calculated by weighted averaging the optimized importance scores of each video segment. ; Solve the constrained optimization problem to construct the summary: ; where i is the index value of the video segment; M is the number of video segments; is the i number of frames in the th video segment; When i = 1, it indicates that the th video segment is selected into the video summary, i When = 0, it indicates that the
[0015] th video segment is not important and not selected; Build the above multi-scale spatio-temporal modeling module; Divide the input video into non-overlapping segments, and use a convolutional neural network to extract the features of each video segment; The multi-scale aggregator uses transposed convolution and pooling operations to obtain the multi-scale features of the input features. At the same time, segment category labels are added in front of the input features, and a feature fusion strategy is applied to aggregate the multi-scale features to obtain the fused features; The cascaded time modeling module consists of multiple cascaded temporal Mamba blocks. The input features are refined through residual learning in multiple cascaded temporal Mamba blocks to fuse short-term cues and long-term event dependencies, and the refined features are output; The parallel spatial modeling module consists of multiple parallel spatial Mamba blocks with shared parameters. Each spatial Mamba block is used to process the features of a single frame. The input features are transformed into frame-level features through a linear layer and then input into the spatial Mamba block. They are processed in parallel in the spatial Mamba block through residual learning to extract spatial features, and the spatial features extracted by each spatial Mamba block are aggregated to form a feature representation of the video content; According to the output features of the multi-scale spatio-temporal modeling module, output the importance scores, and optimize the importance scores in combination with the loss function; According to the optimized importance scores, solve the constrained optimization problem to construct the summary.
[0016] Compared with the prior art, the beneficial technical effects of the present invention are as follows: The multi-scale spatio-temporal modeling video abstract generation method and device provided by the present invention comprehensively capture the local details and global structure of the video by constructing a multi-scale spatio-temporal modeling module including a multi-scale aggregator, a cascaded time modeling module, and a parallel space modeling module. Among them, the multi-scale aggregator effectively aggregates multi-scale features through a feature fusion strategy to generate fused features, thereby capturing information at different granularity levels in the video content, and adding segment category tags before the input features, so that the model can aggregate and summarize the temporal information across frames from video segments. In the cascaded time modeling module, the temporal Mamba block is used to capture the temporal dependencies between frames, enhancing the model's ability to capture long-distance temporal dependencies. The parallel space modeling module captures the structural features within the frame through the parallel space Mamba block, improving the model's ability to extract visual saliency cues within the frame, and its parallel processing ability can efficiently model the spatial relationships and patterns within the frame, significantly improving the computational efficiency. On the other hand, the order of the multi-scale aggregator, the cascaded time modeling module, and the parallel space modeling module in the multi-scale spatio-temporal modeling module can be freely arranged, improving the model's adaptability to different video features, and enabling the model to have a low computational complexity while maintaining high performance.
[0017] The multi-scale spatio-temporal modeling video abstract generation method and device described in the present invention have the advantages of low cost, simple structure, good performance, and strong adaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0019] Figure 1 Schematic diagram of the multi-scale spatio-temporal modeling video abstract generation device provided for an embodiment; Figure 2 Schematic diagram of the combination of the multi-scale spatio-temporal modeling modules provided for an embodiment, where Figure 2 (a) Schematic diagram of the combination of the multi-scale aggregator - cascaded time modeling module - parallel space modeling module, Figure 2 (b) Schematic diagram of the combination of the parallel space modeling module - multi-scale aggregator - cascaded time modeling module - concat, Figure 2 (c) Schematic diagram of the combination of the parallel space modeling module - multi-scale aggregator - cascaded time modeling module - pool; Figure 3Schematic diagram of comparison of F1-score, floating-point operation count, and number of parameters of different algorithms provided for one embodiment on the TVSum dataset; Figure 4 Schematic diagram of the qualitative results of different video summarization algorithms provided for one embodiment. Detailed implementation manners
[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0021] Referring to Figure 1 , one embodiment provides a multi-scale spatio-temporal modeling video summarization generation device, including a feature extraction module, a multi-scale spatio-temporal modeling module, an importance classifier, and a summary generation module; The feature extraction module is used to divide the input video into non-overlapping segments and extract the features of each video segment using a convolutional neural network; The multi-scale spatio-temporal modeling module includes a multi-scale aggregator, a cascaded time modeling module, and a parallel space modeling module arranged in a free order; the multi-scale aggregator uses deconvolution and pooling operations to obtain multi-scale features of the input features, and at the same time adds segment category labels before the input features, and applies a feature fusion strategy to aggregate the multi-scale features to obtain the fused features; the cascaded time modeling module is composed of multiple cascaded temporal Mamba blocks, and the input features are refined through residual learning in multiple cascaded temporal Mamba blocks to fuse short-term cues and long-term event dependencies, and output the refined features; the parallel space modeling module is composed of multiple parallel spatial Mamba blocks with shared parameters, each spatial Mamba block is used to process the features of a single frame, the input features are transformed into frame-level features through a linear layer and then input into the spatial Mamba block, and are processed in parallel in the spatial Mamba block through residual learning to extract spatial features, and the spatial features extracted by each spatial Mamba block are aggregated to form a feature representation of the video content; The importance classifier outputs an importance score according to the output features of the multi-scale spatio-temporal modeling module, and optimizes the importance score in combination with a loss function; The summary generation module is used to construct a summary by solving a constrained optimization problem according to the optimized importance score.
[0022] The multi-scale aggregator is used to integrate video features at different granularity scales, so as to ensure a more comprehensive understanding of the video sequence. The multi-scale aggregator uses deconvolution and pooling operations to obtain the input features Multi-scale features to generate low-scale features , basic-scale features , high-scale features . The input features can be expressed as , where T is the number of frames in the video clip, the spatial size of the input features is , i is the number of input features, H is the height, W is the width. Moreover, a segment category marker is added before the input features, so that the model can aggregate and summarize the temporal information across frames from the video clip.
[0023] In a preferred embodiment, the fused features are obtained according to the following formula: ; where is the fused feature; is the weight of the low-scale feature; is the weight of the basic-scale feature; is the weight of the high-scale feature; is a learnable linear frame projection, L represents the low scale, B represents the basic scale, H represents the high scale.
[0024] The fused feature of the multi-scale aggregator , where D is the number of channels.
[0025] The cascaded temporal modeling module consists of multiple cascaded temporal Mamba blocks. The bidirectional nature of the temporal Mamba blocks enables the model to integrate the information of past and future frames, thus ensuring that the generated video summary is contextually relevant and temporally coherent; the cascaded temporal modeling module effectively improves the quality of the video summary by capturing the key information in the input video features. Each temporal Mamba block selectively compresses the information through its parallel scan algorithm, enabling the model to process long sequence data without significantly increasing the computational overhead; by stacking multiple temporal Mamba blocks, the ability of the model to understand the semantic content of the video from both forward and backward perspectives is enhanced.
[0026] In a preferred embodiment, the input features are refined according to the following formula: ; where is the input feature of the l th temporal Mamba block, , ; is the l th temporal Mamba block; is the input feature of the l -1 th temporal Mamba block.
[0027] The output feature of the cascaded temporal modeling module is denoted as , which fuses short-term cues and long-term event dependencies.
[0028] The parallel spatial modeling module consists of multiple parallel spatial Mamba blocks with shared parameters, and each spatial Mamba block is used to process the features of a single frame. By focusing on a single frame, the parallel spatial modeling module can capture the complex spatial details and context information crucial for accurately summarizing visual content, ensuring that the generated summary is spatially coherent and contextually relevant; at the same time, the parallel nature of multiple spatial Mamba blocks allows for the simultaneous processing of multiple frames, significantly improving computational efficiency.
[0029] In a preferred embodiment, the feature representation of the video content is formed according to the following formula: ; where is the feature representation of the video content; is the parallel spatial Mamba block with shared parameters; is the t th frame-level feature, .
[0030] Referring to Figure 2 , in one embodiment, three different combinations of the multi-scale spatio-temporal modeling module are shown.
[0031] Figure 2 (a) shows that the multi-scale spatio-temporal modeling module is a combination of a multi-scale aggregator - cascaded temporal modeling module - parallel spatial modeling module; the multi-scale aggregator uses deconvolution and pooling operations to obtain multi-scale features of the input features, and then T the feature embeddings of the frames are merged into dimension to form the input of the cascaded temporal modeling module; adding a segment category token to learn the global importance of the video segment, forming dimension T , the cascaded temporal modeling module uses temporal Mamba blocks to capture the temporal correlations between frames and outputs features with dimension
[0032] Figure 2(b) shows that the multi-scale spatio-temporal modeling module is a parallel spatial modeling module - multi-scale aggregator - cascaded temporal modeling module - concat combination; after the parallel spatial modeling module extracts the spatial features within each frame, the multi-scale aggregator uses deconvolution and pooling operations to obtain the multi-scale features of the input features; then the multi-scale features are concatenated in the spatial dimension to form features with a dimension of ; and then the feature dimension is reduced to through a linear layer, and finally the cascaded temporal modeling module is used to learn the temporal information.
[0033] Figure 2 (c) shows that the multi-scale spatio-temporal modeling module is a parallel spatial modeling module - multi-scale aggregator - cascaded temporal modeling module - pool combination; after the parallel spatial modeling module extracts the spatial features within each frame, the multi-scale aggregator uses deconvolution and pooling operations to obtain the multi-scale features of the input features, forming features with a dimension of ; then the cascaded temporal modeling module is used to capture the T temporal correlation between frames; then average pooling is performed on the output of each scale, and concatenation is performed in the channel dimension to form a feature representation with a dimension of .
[0034] The outputs of the above three multi-scale spatio-temporal modeling modules are input to the importance classifier to obtain the importance scores at the frame level.
[0035] In a preferred embodiment, the importance scores are output according to the following steps: Perform layer normalization on the output features of the multi-scale spatio-temporal modeling module; Extract the higher-level semantic patterns of the layer-normalized features through a fully connected layer with a activation function; Use the function to normalize the higher-level semantic patterns of the features into importance scores and output them.
[0036] In a preferred embodiment, the loss function is calculated according to the following formula: ; where is the loss function; is the cross-entropy loss function, which is used to measure the difference between the predicted class probabilities and the true labels; is the mean squared error loss function, which is used to penalize the bias in continuous score predictions; is an adjustable hyperparameter used to control the loss balance.
[0037] The cross-entropy loss function is calculated according to the following formula: ; Among them, t is the index value of the frame; c is the result label of the classification; is the t true label of the frame; is the prediction result of the frame importance classifier for the t frame.
[0038] The mean squared error loss function is calculated according to the following formula: ; Among them, is the predicted importance score; is the true normalized importance score.
[0039] The abstract is constructed according to the following steps: The mean importance score is calculated by weighted averaging the optimized importance scores of each video segment ; Solve the constrained optimization problem to construct the abstract: ; Among them, i is the index value of the video segment; M is the number of video segments; is the i number of frames in the th video segment; When i = 1, it indicates that the th video segment is selected into the video abstract, i = 0 indicates that the th video segment is unimportant and not selected;
[0040] In one embodiment, a multi-scale spatio-temporal modeling video abstract generation method is provided, including the following steps: Build the multi-scale spatio-temporal modeling module described in any of the above embodiments; Divide the input video into non-overlapping segments, and use a convolutional neural network to extract the features of each video segment; The multi-scale aggregator obtains multi-scale features of the input features through deconvolution and pooling operations. Meanwhile, segment category labels are added before the input features, and a feature fusion strategy is applied to aggregate the multi-scale features to obtain the fused features. The cascaded temporal modeling module consists of multiple cascaded temporal Mamba blocks. The input features are refined through residual learning in multiple cascaded temporal Mamba blocks to fuse short-term cues and long-term event dependencies, and the refined features are output. The parallel spatial modeling module consists of multiple parallel spatial Mamba blocks with shared parameters. Each spatial Mamba block is used to process the features of a single frame. After the input features are transformed into frame-level features through a linear layer, they are input into the spatial Mamba blocks and processed in parallel in the spatial Mamba blocks through residual learning to extract spatial features, and the spatial features extracted by each spatial Mamba block are aggregated to form a feature representation of the video content. Based on the output features of the multi-scale spatio-temporal modeling module, importance scores are output, and the importance scores are optimized in combination with a loss function. Based on the optimized importance scores, a constrained optimization problem is solved to construct a summary.
[0041] In one embodiment, to verify the effectiveness of the present invention, comparisons with existing methods were made on the SumMe and TVSum datasets. The solution adopted in this embodiment is Solution A, as Figure 2 shown in (a), and the evaluation metrics are F1-score, Kendall, and Spearman.
[0042] The comparison results of F1-score are shown in Table 1.
[0043] Table 1 Comparison results of different algorithms in terms of F1-score
[0044] It can be seen that compared with the existing methods, the evaluation results of the present invention on the SumMe and TVSum datasets both exceed the existing methods. The present invention can more effectively learn the temporal correlation between frames and the spatial attention within frames to classify the importance of frames. The present invention can effectively provide top-level video summary results on the two datasets.
[0045] Kendall ( ), and Spearman ( ), and the comparison results of the correlation coefficients are shown in Table 2.
[0046] Table 2 Comparison results of different algorithms in terms of Kendall ( ), and Spearman ( ).
[0047] As can be seen from Table 2, the present invention achieves the best and the second-best performance, indicating that there is a stronger consistency between the rankings provided by the frame importance scores of human annotation and model prediction. The excellent performance of the present invention in terms of Kendall ( ) and Spearman ( ) metrics demonstrates the high effectiveness of the present invention in learning complex spatio-temporal control and spatial information in video frames.
[0048] Referring to Figure 3 and Table 3, the performance-complexity trade-offs of various methods on the TVSum dataset are compared. In the table, Scheme A, Scheme B, and Scheme C correspond to the combinations shown in Figure 2 (a), Figure 2 (b), Figure 2 (c) respectively. It can be seen that the setting shown in Figure 2 (a) in the present invention achieves the best performance in terms of F1-score, Kendall, and Spearman coefficients. At the same time, compared with other methods, it maintains a relatively small model size and low computational complexity. The setting shown in Figure 2 (c) in the present invention achieves an F1-score of 66.6% while having only 25.20M parameters, demonstrating excellent computational efficiency. Compared with the Transformer method, the number of parameters is reduced by 72.2%. Moreover, in the case of comparable FLOPs, it is superior to the CNN architecture by 3.8% in terms of F1-score; compared with lightweight methods (such as RR-STG), its F1-score is 22.7% higher. In summary, the present invention achieves the best balance between computational efficiency and summary accuracy.
[0049] Table 3 Comparison results of different combinations and different algorithms in terms of Kendall ( ) and Spearman ( )
[0050] Referring to Figure 4 , in order to visually compare the video summary results, the qualitative results of the present invention are compared with those of three methods: spatio-temporal vision transformer, anchor-based flexible detection summary network, and anchor-free flexible detection summary network. As can be seen from Figure 4 , in the qualitative results of the 32nd video of the TVSum dataset, the method of the present invention shows a higher similarity in terms of frame index and summary length compared with the ground truth label, further demonstrating the effectiveness of the present invention in maintaining consistency with the reference data.
[0051] Referring to Table 4, the performance of the three combinations shown in Figure 2 is presented.Figure 2 The combination shown in (a) exhibits superior performance on both datasets, achieving a F1-score of 56.0%, a Kendall ( ), and a Spearman ( ), on SumMe, and reaching a F1-score of 67.5% on TVSum; indicating that its combination effectively captures key spatial features while maintaining temporal coherence. Figure 2 The combination shown in (b) compromises in temporal modeling ability on SumMe, with a F1-score of 51.8% and lower correlation coefficients (a Kendall ( ), and a Spearman ( )); it maintains a F1-score of 66.7% on TVSum, showing reasonable spatial understanding ability. Figure 2 The combination shown in (c) has a F1-score of 48.2% on SumMe.
[0052] Table 4 Comparison results of different combinations on SumMe and TVSum
[0053] In one embodiment, the impacts of the cascaded temporal modeling module (CTMM) and the parallel spatial modeling module (PSMM) are evaluated through ablation experiments, and the results are shown in Table 5.
[0054] Table 5 Ablation experiments of CTMM and PSMM on TVSum dataset
[0055] As can be seen from Table 5, when CTMM is disabled, the F1-score performance on the TVSum dataset drops by 0.8%, and the Kendall ( ), and Spearman ( ) drop by 0.049 and 0.062 respectively, indicating that the cascaded temporal modeling module plays a key role in establishing the temporal dependencies of non-adjacent frames. When the parallel spatial modeling module is removed, the F1-score decreases by 0.5%, the Kendall ( ) drops by 0.067, and the Spearman ( ) drops by 0.091, highlighting the ability of the parallel spatial modeling module to effectively extract discriminative spatial features through a multi-branch attention scheme.
[0056] Matters not covered by this invention are well-known techniques.
[0057] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0058] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application.
[0059] The above is only the preferred embodiment of the present invention and is not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. Multi-scale spatio-temporal modeling video abstract generation device, characterized in that It includes a feature extraction module, a multi-scale spatio-temporal modeling module, an importance classifier, and a summary generation module; The feature extraction module is used to divide the input video into non-overlapping segments and extract the features of each video segment using a convolutional neural network; The multi-scale spatio-temporal modeling module includes a multi-scale aggregator, a cascaded temporal modeling module, and a parallel spatial modeling module arranged in a sequence-free manner; the multi-scale aggregator uses deconvolution and pooling operations to obtain multi-scale features of the input features, adds segment category labels in front of the input features at the same time, and applies a feature fusion strategy to aggregate the multi-scale features to obtain the fused features; the cascaded temporal modeling module consists of multiple cascaded temporal Mamba blocks, and the input features are refined through residual learning in multiple cascaded temporal Mamba blocks to fuse short-term cues and long-term event dependencies, and the refined features are output; the parallel spatial modeling module consists of multiple parallel spatial Mamba blocks with shared parameters, each spatial Mamba block is used to process the features of a single frame, the input features are transformed into frame-level features through a linear layer and then input into the spatial Mamba block, and are processed in parallel in the spatial Mamba block through residual learning to extract spatial features, and the spatial features extracted by each spatial Mamba block are aggregated to form a feature representation of the video content; The importance classifier outputs an importance score based on the output features of the multi-scale spatio-temporal modeling module and optimizes the importance score in combination with a loss function; The summary generation module is used to construct a summary by solving a constrained optimization problem according to the optimized importance score.
2. The multi-scale spatio-temporal modeling video abstract generation device according to claim 1, wherein The fused features are obtained according to the following formula: Among them, is a low-scale feature; is a basic-scale feature; is a high-scale feature; is the fused feature; is the weight of the low-scale feature; is the weight of the basic-scale feature; is the weight of the high-scale feature; is a learnable linear frame projection, L represents the low scale, B represents the basic scale, H represents the high scale.
3. The multi-scale spatio-temporal modeling video abstract generation device according to claim 1, wherein The input features are refined according to the following formula: Among them, is the input feature of the l th temporal Mamba block, ; is the l th temporal Mamba block; is the input feature of the l -1th temporal Mamba block.
4. The multi-scale spatio-temporal modeling video abstract generation device according to claim 1, wherein, The feature representation of the video content is formed according to the following formula: Among them, is the feature representation of the video content; is the parallel spatial Mamba block of the shared parameters; is the t th frame-level feature, .
5. The multi-scale spatio-temporal modeling video abstract generation device according to claim 1, characterized in that, The importance score is output according to the following steps: Perform layer normalization on the output features of the multi-scale spatio-temporal modeling module; Extract more advanced semantic patterns of the features after layer normalization through a fully connected layer with an activation function; Adopt The function normalizes the higher-level semantic patterns of the features into importance scores and outputs them.
6. The multi-scale spatio-temporal modeling video abstract generation device according to claim 1, wherein, The loss function is calculated according to the following formula: Among them, is the loss function; is the cross-entropy loss function; is the mean squared error loss function; is an adjustable hyperparameter.
7. The multi-scale spatio-temporal modeling video abstract generation device according to claim 6, characterized in that, The cross-entropy loss function is calculated according to the following formula: Among them, t is the index value of the frame; c is the result label of the classification; is the t ground truth label of the frame; is the prediction result of the frame importance classifier for the t frame.
8. The multi-scale spatio-temporal modeling video abstract generation device according to claim 6, wherein The mean squared error loss function is calculated according to the following formula: Among them, is the predicted importance score; is the true normalized importance score.
9. The multi-scale spatio-temporal modeling video abstract generation device according to claim 1, wherein The summary is constructed according to the following steps: The mean importance score is calculated by weighted averaging the optimized importance scores of each video segment ; Solve the constrained optimization problem to construct a summary: wherein, i is the index value of the video segment; M is the number of video segments; is the i number of frames in the th video segment; is the video segment summary metric, when i = 1, it indicates that the th video segment is selected into the video summary, when i = 0, it indicates that the i th video segment is unimportant and not selected; is the total length of the original video.
10. A multi-scale spatio-temporal modeling video abstract generation method, characterized in that, It includes the following steps: Build the multi-scale spatio-temporal modeling module as described in claim 1; Divide the input video into non-overlapping segments and extract the features of each video segment using a convolutional neural network; The multi-scale aggregator obtains multi-scale features of the input features by using deconvolution and pooling operations. At the same time, segment class tokens are added before the input features, and a feature fusion strategy is applied to aggregate the multi-scale features to obtain the fused features. The cascaded temporal modeling module consists of multiple cascaded temporal Mamba blocks. The input features are refined through residual learning in multiple cascaded temporal Mamba blocks to fuse short-term cues and long-term event dependencies, and the refined features are output. The parallel spatial modeling module consists of multiple parallel spatial Mamba blocks with shared parameters. Each spatial Mamba block is used to process the features of a single frame. After the input features are transformed into frame-level features through a linear layer, they are input into the spatial Mamba blocks and processed in parallel in the spatial Mamba blocks through residual learning to extract spatial features, and the spatial features extracted by each spatial Mamba block are aggregated to form a feature representation of the video content. Based on the output features of the multi-scale spatio-temporal modeling module, importance scores are output, and the importance scores are optimized in combination with a loss function. According to the optimized importance scores, a constrained optimization problem is solved to construct a summary.
Citation Information
Patent Citations
Video abstraction method based on symmetric multi-scale attention
CN117493607A
High-resolution remote sensing image target detection method based on multi-scale network
CN118485927A
Crowd counting method based on adaptive global perception and multi-scale feature fusion
CN119048993A