A video summarization method integrating cross-modal semantic information

Through cross-modal feature extraction and fusion, combined with the spatiotemporal convolutional correlation attention mechanism and semantic consistency corrector, the problem of insufficient utilization of motion information in the existing video digest method is solved, and higher quality video digest generation is achieved.

CN120126056BActive Publication Date: 2025-08-12SHIJIAZHUANG TIEDAO UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510248933.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-08-12
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

The existing video digest methods fail to fully combine the motion information characteristics in the space-time stream network, resulting in room for improvement in the quality of the digest, especially when processing dynamic scenes and motion information.

Method used

The cross-modal feature extraction network is used to extract static and dynamic features respectively, and the space-time importance attention map is generated through the spatiotemporal convolutional correlation attention mechanism, and a cross-modal dynamic fusion module and a semantic consistency corrector are introduced to construct an objective function for training to generate a video summary.

Benefits of technology

Improves the comprehensiveness and accuracy of video digests, enhances the ability to capture critical events and dynamic content, reduces computational costs and reduces noise interference, and improves the semantic accuracy and content coherence of the digests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126056B_ABST
    Figure CN120126056B_ABST
Patent Text Reader

Abstract

The present invention discloses a video summarization method that integrates cross-modal semantic information, which belongs to the field of computer vision technology. The method first extracts an image frame sequence and a motion frame sequence from an input video, and then uses a cross-modal feature extraction network to extract static features and dynamic features respectively. Then, the frame features are processed by a spatiotemporal convolutional association attention mechanism to generate an attention map that reflects the spatiotemporal importance of the frame features, while capturing the intra-frame spatial information and the inter-frame temporal information. In addition, a cross-modal dynamic fusion module and a semantic consistency corrector are introduced to optimize the video summary generation process, reduce noise interference, and improve the summary quality. Finally, an objective function is constructed, and a video summary generation model is trained through unsupervised or supervised learning to generate a dynamic video summary based on the predicted importance score. The method comprehensively utilizes the static and dynamic features in the video to improve the semantic accuracy and content coherence of the summary.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a video summarization method for fusing cross-modal semantic information, and belongs to the technical field of image communication. Background Art

[0002] With the rapid development of the internet and social media, the amount of video data is growing exponentially. Efficiently extracting key information from this massive amount of video data and generating concise and representative video summaries has become a key research topic. Video summarization not only has important applications in video content retrieval, management, and presentation, but also plays an indispensable role in automated editing, intelligent surveillance, and other fields. However, existing video summarization methods still have many limitations, especially when dealing with dynamic scenes and capturing motion information in videos.

[0003] Traditional video summarization techniques can be roughly divided into two categories: keyframe selection based on static images and dynamic event recognition based on video content. Keyframe extraction methods based on static images typically rely on image features such as color, texture, and shape to select representative images as video summaries. However, this method ignores the motion information in the video and cannot fully reflect the dynamic changes in the video content. Although dynamic event recognition methods based on video content can take motion information into account, they are often limited by temporal and spatial characteristics when dealing with complex scenes, resulting in the generated summary being unable to effectively reflect the important events or plot points in the video.

[0004] Capturing motion information is crucial in generating video summaries. Motion information in a video, such as the trajectory of objects and the dynamic changes in scenes, is often crucial for conveying the core content of the video. However, traditional methods often fail to fully consider the dynamic evolution of motion information, resulting in a summary that lacks sufficient detail and accuracy in conveying the video's plot.

[0005] In recent years, with the rapid development of deep learning technology, spatiotemporal flow networks, as an emerging model structure, have demonstrated significant advantages in processing video data. They can simultaneously capture both the spatial characteristics and temporal dynamics of a video, achieving outstanding results in fields such as action recognition and video analysis. Using spatiotemporal flow networks, motion information in videos can be more accurately modeled, enabling video summaries to more comprehensively reflect key events and dynamic content within the video.

[0006] However, existing spatiotemporal flow network methods typically focus on specific tasks (such as action recognition and video classification) and are not specifically optimized for video summarization. Traditional video summarization methods fail to fully incorporate the motion information characteristics of spatiotemporal flow networks, resulting in significant room for improvement in summary quality. Therefore, how to effectively integrate spatiotemporal flow networks with motion information to improve the quality of video summarization has become a pressing research challenge. Summary of the Invention

[0007] The purpose of the present invention is to provide a video summarization method that integrates cross-modal semantic information, aiming to solve the problem that the existing technology does not fully utilize the static features and motion features in the video, resulting in the model being unable to fully and accurately understand the video content.

[0008] To achieve the above object, the present invention provides a video summarization method integrating cross-modal semantic information, comprising the following steps:

[0009] S1: Read the input video and extract the image frame sequence used to represent the static visual content and the motion frame sequence that reflects the dynamic motion state changes;

[0010] S2: extracting static features and dynamic features of the video frame respectively through a cross-modal feature extraction network, wherein the cross-modal feature extraction network includes a temporal stream network and a spatial stream network, wherein the spatial stream network is used to extract static features and the temporal stream network is used to extract dynamic features;

[0011] S3: Generate spatiotemporal importance attention map through spatiotemporal convolutional association attention mechanism;

[0012] S4: Introducing a cross-modal dynamic fusion module, dynamically adjusting the weight ratio of static and dynamic modalities based on the semantic features of the current frame, and generating a hybrid feature representation that integrates cross-modal semantics;

[0013] S5: Introduce a semantic consistency corrector to optimize the semantic consistency between cross-modal features and static features;

[0014] S6: Construct an objective function, train a video summary generation model, and generate a video summary based on the importance score predicted by the model.

[0015] Furthermore, the spatial stream network in the video summarization method for fusing cross-modal semantic information is used to extract static features that reflect object categories, scene semantics and visual content in video frames.

[0016] Furthermore, the temporal stream network in the video summarization method that integrates cross-modal semantic information is used to extract two dynamic features: motion RGB features and optical flow features. The motion RGB features are used to capture scene switching and color changes of dynamic targets, and the optical flow features are used to describe the direction and speed of motion between frames. By simultaneously extracting the two dynamic features, the model can more comprehensively capture the motion information in the video and effectively distinguish different types of dynamic changes.

[0017] Furthermore, in the video summarization method that integrates cross-modal semantic information, the spatiotemporal convolutional association attention mechanism first stacks the video frame features to form a frame feature representation with a two-dimensional structure; then a convolutional neural network is used to generate an attention map that reflects the spatiotemporal importance of frame features, while capturing the spatial information within the frame and the temporal information between frames.

[0018] Furthermore, the cross-modal dynamic fusion module in the video summarization method for integrating cross-modal semantic information is a key improvement introduced into the spatiotemporal convolutional attention mechanism. It uses a small multi-layer perceptron to dynamically calculate the ratio of static and dynamic modal weights (α and β) based on the features of the current frame. This module receives embedded frame features as input and, after processing by the multi-layer perceptron, outputs a normalized weight ratio, where α + β = 1. In this way, the model can adaptively adjust the distribution of static and dynamic modal weights based on the video content, thereby more flexibly processing different types of video content and enhancing the ability to recognize key frames.

[0019] Furthermore, the semantic consistency corrector in the video summarization method that integrates cross-modal semantic information can dynamically adjust the semantic matching between cross-modal features and video static features, effectively reducing the noise caused by motion features, thereby enhancing the semantic accuracy and content coherence of the generated video summary.

[0020] Furthermore, the objective function in the video summarization method for fusing cross-modal semantic information includes the following items:

[0021] The objective function is defined as:

[0022] ,

[0023] in: is a reward function term used to evaluate the importance and coverage of frames in the generated summary; is a regularization term used to prevent the model from overfitting; It is a semantic consistency loss term used to optimize the semantic consistency between cross-modal features and appearance features; and is a hyperparameter used to balance the weights of different loss terms.

[0024] The reward function is used to evaluate the similarity between the generated summary and the annotated summary, thereby measuring the quality of the summary generated by the model. The specific calculation formula is as follows:

[0025] ,

[0026] in: is the total number of video frames; It is The annotation importance score of the frame; The model predicts The importance score of the frame, ranging from [0, 1].

[0027] The regularization term is used to prevent the model from overfitting, using the weight decay method. The specific calculation formula of the weight decay regularization term is as follows:

[0028] ,

[0029] in represents all parameters of the model, Represents the L2 norm of the parameter. By adding a weight decay term, the scale of the model parameters can be limited, thereby improving the generalization ability of the model.

[0030] The semantic consistency loss term is used to optimize the semantic consistency between cross-modal features and appearance features and reduce noise interference in motion features. The specific calculation formula is as follows:

[0031] ,

[0032] in, It is the mixed feature vector after fusing time and space features. is the static feature vector extracted from the time stream, N is the total number of video frames; Represents the Euclidean distance of vectors.

[0033] Furthermore, the training steps of the model proposed in the motion information enhanced video summarization method based on the spatiotemporal flow network include: constructing a video summary generation network; constructing a training set containing original video frames and their corresponding annotated importance scores; inputting the training set into the network for training; the network outputs the predicted importance scores of the frames; calculating the loss between the predicted scores and the annotated scores; when the loss is minimized, the model converges and the training is stopped to obtain a trained network.

[0034] Compared with the prior art, the present invention has the following beneficial effects:

[0035] 1. The present invention not only considers the static visual content in the video, but also comprehensively analyzes the dynamic motion state, so that the video summary can more comprehensively reflect the video content.

[0036] 2. By extracting and utilizing motion RGB features and optical flow features, the present invention can more accurately capture important dynamic events in the video and improve the accuracy of the summary.

[0037] 3. The spatiotemporal convolutional attention mechanism proposed in this paper improves the ability to capture key information in video summaries by enhancing the spatiotemporal importance of frame features, while reducing computational costs.

[0038] 4. The cross-modal dynamic fusion module introduced in this invention enables the model to adaptively balance static and dynamic features, optimizes the feature fusion process, and improves the accuracy of the summary.

[0039] 5. The semantic consistency corrector proposed in this invention enhances the semantic accuracy and content coherence of video summaries by reducing noise interference. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:

[0041] Figure 1 A flowchart of the video summarization method for fusing cross-modal semantic information provided by the present invention;

[0042] Figure 2 A network framework for a video summarization method integrating cross-modal semantic information provided by an embodiment of the present invention;

[0043] Figure 3 Schematic diagram of the structure of the spatiotemporal convolutional attention mechanism provided by an embodiment of the present invention;

[0044] Figure 4 is a schematic structural diagram of a cross-modal dynamic fusion module provided by an embodiment of the present invention;

[0045] Figure 5 4 is a schematic diagram of the structure of the semantic consistency corrector provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0046] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several variations and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.

[0047] like Figure 1 FIG. 1 is a flowchart of an implementation of a video summarization method for fusing cross-modal semantic information provided by the present invention, comprising the following steps:

[0048] S1, reads the input video and extracts the image frame sequence and motion frame sequence;

[0049] S2, extracts static and dynamic features of video frames through a cross-modal feature extraction network;

[0050] S3, generates a spatiotemporal importance attention map through the spatiotemporal convolutional association attention mechanism;

[0051] S4, introduces a cross-modal dynamic fusion module into the spatiotemporal convolutional attention mechanism;

[0052] S5, introduces a semantic consistency corrector to optimize the video summary generation process;

[0053] S6, constructs the objective function, trains the video summary model, predicts the importance of new video frames, and generates summaries.

[0054] Example 1: The present invention provides a preferred embodiment to execute S1, first read the input video, and obtain two types of frame sequences: one is an image frame sequence reflecting static visual content, and the other is a motion frame sequence reflecting the dynamic motion state change of the video. This method is suitable for processing video content of various types and different lengths, such as movies, surveillance videos, sports event videos, etc., and can efficiently process both long and short videos. The specific operating steps are as follows: The extraction of video frame sequences is achieved by uniform sampling. First, two frames per second are selected as the sampling rate, which means that for a video, one frame is extracted every 0.5 seconds. Then, according to the sampling rate, frames are evenly extracted from the beginning to the end of the video to form a frame sequence. ,in T is the total number of frames extracted, Indicates the t Finally, the extracted frame sequence is fed into the pre-trained neural network model to extract static and dynamic features.

[0055] Example 2: The present invention provides a preferred embodiment to perform S2, using a neural network model to extract features from image frame sequences and motion frame sequences to obtain frame-level static features and frame-level motion features. Figure 2 The figure shows the network framework of the video summarization method for integrating cross-modal semantic information provided by an embodiment of the present invention. The specific steps are as follows:

[0056] S21, extract static features. The spatial stream network focuses on extracting static features from video frames. Static features are mainly used to capture semantic information about object categories, scene semantics, and visual content in video frames. First, the sampled frames are input into a pre-trained deep convolutional neural network model. Here, GoogleNet is used as the base model of the spatial stream network. Then, feature vectors of the pool5 layer with a dimension of 1024 are extracted from the model. These feature vectors reflect the semantic information about the object categories, scene semantics, and visual content in the frame, which can help the model understand the static content within the frame.

[0057] S22, extract dynamic features. The temporal stream network focuses on extracting dynamic features from video frames, including motion RGB features and optical flow features. First, the sampled frames are input into a pre-trained I3D model. The I3D model is pre-trained on large-scale video datasets such as Kinetics and can effectively capture dynamic information in the video. Then, two types of motion features are extracted from the model: motion RGB features. X r and optical flow features X f .

[0058] Motion RGB features capture color changes and scene transitions within video frames. These features reflect visual changes in the video, such as scene changes or object color changes. First, the acquired frame sequence is fed into the RGB stream of the I3D model for processing. Next, the Mixed_5c layer is selected to extract feature points. The extracted feature vectors are then normalized to ensure a consistent value range.

[0059] The optical flow feature focuses on the pixel movement between frames, which can directly reflect the motion direction and speed information in the video, such as the walking of a person, the movement of an object, or the dynamic changes of the background. First, the optical flow field is calculated in the input video. , where each o t is the optical flow map between frame t and frame t+1. The calculated optical flow field sequence O is input into the optical flow stream of the I3D model. The optical flow stream specializes in processing optical flow information and can capture motion patterns in the video.

[0060] Example 3: The present invention provides a preferred embodiment to perform S3, using a spatiotemporal convolutional attention mechanism to fuse the spatial information within a frame and the temporal information between frames to generate an attention map reflecting the spatiotemporal importance of frame features. Figure 3 The figure shows the structure of the spatiotemporal convolutional attention mechanism provided by an embodiment of the present invention. The specific steps are as follows:

[0061] First, the features extracted from the spatial and temporal streams are stacked to form a two-dimensional frame feature representation. The static features of the spatial stream and the dynamic features of the temporal stream are concatenated along the feature dimensions to obtain a high-dimensional frame feature representation. A convolutional neural network is then used to process these stacked frame features. Convolutional neural networks can simultaneously capture both spatial information within a frame and temporal information between frames, generating an attention map that reflects the spatiotemporal importance of frame features. Each element of the attention map represents the importance of the corresponding frame feature. The attention map can be used for subsequent feature weighting and importance score prediction.

[0062] Example 4: The present invention provides a preferred embodiment to perform S4, by introducing a cross-modal dynamic fusion module into the spatiotemporal convolutional attention mechanism, the model can dynamically calculate the ratio of static and dynamic modal weights (α and β) based on the features of the current frame. Figure 4 The figure shows a schematic diagram of the structure of the cross-modal dynamic fusion module provided by an embodiment of the present invention. The specific steps are as follows:

[0063] First, a multi-layer perceptron dynamically calculates the ratio of static and dynamic modal weights based on the features of the current frame. The input of this network is the feature representation E of the current frame, and the output is the weight ratio α and β. The specific formula is as follows:

[0064] ,

[0065] Among them, α and β represent the ratio of time weight and space weight respectively, and satisfy α+β=1.

[0066] Then, the time and space weights are adjusted according to the calculated weight ratio. Assume that the original time and space weights are and , the adjusted weights are:

[0067] ,

[0068] Fuse the adjusted weights with the frame features to generate a hybrid feature representation M The specific formula is as follows:

[0069] ,

[0070] in, is the frame feature representation, Represents element-wise multiplication.

[0071] Finally, using hybrid feature representation M Predict the importance score of each frame. Use a fully connected layer and Sigmoid activation function to map the mixed features into importance scores, ranging from 0 to 1. The specific formula is as follows:

[0072] ,

[0073] in, represents the Sigmoid activation function, is the importance score of the prediction.

[0074] Example 5: The present invention provides a preferred embodiment to perform S5, designing a semantic consistency corrector for optimizing the video summary generation process, by dynamically correcting the semantic consistency between cross-modal features and static features, reducing noise interference in motion features, thereby improving the semantic accuracy and coherence of the video summary. Figure 5 The following is a schematic diagram of the structure of the semantic consistency corrector provided by an embodiment of the present invention. The specific steps are as follows:

[0075] S51, semantic alignment of cross-modal features and static features: First, the semantic consistency between the cross-modal features and the static features is calculated. Semantic consistency is then measured by calculating the cosine similarity between the two. Cosine similarity is a method for measuring the directional similarity between two vectors and is applicable to high-dimensional feature vectors. The calculation formula is as follows:

[0076] ,

[0077] in, It is the mixed feature vector after fusing time and space features. is the static feature vector extracted from the time stream, ∙ represents the dot product of vectors, Represents the modulus of the vector. By calculating cosine similarity, the semantic consistency between cross-modal features and appearance features can be quantified, providing a basis for subsequent optimization processes.

[0078] S52, semantic consistency loss calculation: Define the semantic consistency loss function to optimize the semantic consistency of the model. The calculation formula is as follows:

[0079] ,

[0080] in, N is the total number of video frames; The Euclidean distance between representation vectors. This loss function aims to minimize the difference between cross-modal features and appearance features, thereby improving the model's ability to understand the semantics of video content. By optimizing this loss function, the semantic match between cross-modal and appearance features can be dynamically adjusted, reducing noise interference in motion features.

[0081] S53, Noise Suppression and Optimization: By minimizing semantic consistency loss and dynamically adjusting the semantic match between cross-modal features and static features, we can reduce noise interference caused by motion features. The optimized model can generate more semantically accurate and coherent video summaries.

[0082] Example 6: The present invention provides a preferred embodiment for executing S6 to construct an objective function and train the video summary generation model in an unsupervised or supervised learning manner. The specific steps are as follows:

[0083] The objective function consists of a reward function term, a regularization term, and a semantic consistency loss term. The reward function term is used to evaluate the importance and coverage of frames in the summary, the regularization term is used to prevent model overfitting, and the semantic consistency loss term is used to optimize the semantic consistency between cross-modal features and static features.

[0084] The objective function is defined as:

[0085] ,

[0086] in: is a reward function term used to evaluate the importance and coverage of frames in the generated summary; is a regularization term used to prevent the model from overfitting; is a semantic consistency loss term used to optimize the semantic consistency between cross-modal features and appearance features (defined in Example 5); and is a hyperparameter used to balance the weights of different loss terms.

[0087] The reward function is used to evaluate the similarity between the generated summary and the annotated summary, thereby measuring the quality of the summary generated by the model. The specific calculation formula is as follows:

[0088] ,

[0089] in: N is the total number of video frames; It is i The annotation importance score of the frame; The model predicts i The importance score of the frame, ranging from [0, 1]. This loss function encourages the model to generate summaries similar to the annotation summaries by minimizing the cross entropy between the annotation scores and the prediction scores.

[0090] Regularization is used to prevent overfitting of the model, usually using methods such as weight decay or dropout. The regularization term of weight decay can be expressed as:

[0091] ,

[0092] in represents all parameters of the model, Represents the L2 norm of the parameter. By adding a weight decay term, the scale of the model parameters can be limited, thereby improving the generalization ability of the model.

[0093] The semantic consistency loss term has been defined in Example 5, and its purpose is to reduce noise interference in motion features by optimizing the semantic consistency between cross-modal features and appearance features.

Claims

1. A video summarization method integrating cross-modal semantic information, characterized in that: The following steps are involved: S1: Read the input video and extract the image frame sequence used to represent the static visual content and the motion frame sequence that reflects the dynamic motion state changes; S2: extracting static features and dynamic features of the video frame respectively through a cross-modal feature extraction network, wherein the cross-modal feature extraction network includes a temporal stream network and a spatial stream network, wherein the spatial stream network is used to extract static features and the temporal stream network is used to extract dynamic features; The spatial stream network is used to extract static semantic features that reflect the object categories, scene semantics, and visual content in the video frames; the temporal stream network is used to extract dynamic features, including motion RGB features and optical flow features. The motion RGB features are used to capture scene switching and color changes of dynamic targets, and the optical flow features are used to describe the direction and speed of motion between frames. S3: Generate spatiotemporal importance attention map through spatiotemporal convolutional association attention mechanism; The spatiotemporal convolutional association attention mechanism is used to fuse the spatial information within a frame and the temporal information between frames to generate an attention map that reflects the spatiotemporal importance of frame features. First, the features extracted from the spatial stream and the temporal stream are stacked to form a two-dimensional frame feature representation. The static features of the spatial stream and the dynamic features of the temporal stream are spliced along the feature dimension to obtain a high-dimensional frame feature representation. The stacked frame features are then processed using a convolutional neural network. The convolutional neural network can simultaneously capture the spatial information within the frame and the temporal information between frames to generate an attention map that reflects the spatiotemporal importance of frame features. Each element of the attention map represents the importance of the corresponding frame feature. The attention map is used for subsequent feature weighting and importance score prediction. S4: Introducing a cross-modal dynamic fusion module, dynamically adjusting the weight ratio of static and dynamic modalities based on the semantic features of the current frame, and generating a hybrid feature representation that integrates cross-modal semantics; S5: Introduce a semantic consistency corrector to optimize the semantic consistency between cross-modal features and static features; S6: Construct an objective function, train a video summary generation model, and generate a video summary based on the importance score predicted by the model.

2. The video summarization method for fusing cross-modal semantic information according to claim 1, wherein: The cross-modal dynamic fusion module dynamically calculates the ratio of static and dynamic modal weights according to the semantic features of the current frame, adjusts the static and dynamic modal weights according to the ratio, and fuses the adjusted weights with the frame features. Generate a hybrid feature representation that integrates cross-modal semantics and then predict the importance score of each frame.

3. The video summarization method for fusing cross-modal semantic information according to claim 1, wherein: The semantic consistency corrector dynamically adjusts the semantic matching degree between cross-modal features and video static features.

4. The video summarization method for fusing cross-modal semantic information according to claim 1, wherein: The objective function is defined as: , in L reward is a reward function term used to evaluate the importance and coverage of frames in the generated summary; L reg is a regularization term used to prevent the model from overfitting; L sem It is a semantic consistency loss term used to optimize the semantic consistency between cross-modal features and static features; and is a hyperparameter used to balance the weights of different loss terms; The reward function is used to evaluate the similarity between the generated summary and the annotated summary, thereby measuring the quality of the summary generated by the model. The specific calculation formula is as follows: , in N is the total number of video frames; It is i The annotation importance score of the frame; is the importance score of the i-th frame predicted by the model, ranging from [0, 1]; The regularization term is used to prevent the model from overfitting. The weight decay method is used. The specific calculation formula of the weight decay regularization term is as follows: , in Represents all parameters of the model; The semantic consistency loss term is used to optimize the semantic consistency between cross-modal features and static features and reduce noise interference in motion features. The specific calculation formula is as follows: , in It is a mixed feature vector after fusing static and dynamic features. is the static feature vector extracted from the spatial stream, N is the total number of video frames, Represents the Euclidean normal form of a vector.

Citation Information

Patent Citations

  • Multi-modal information fusion identification method and system based on attention mechanism

    CN114332573A

  • Video abstract generation method based on motion information assistance

    CN116233569A