A video summarization method in a self-supervised manner
Through the inverse optimal transmission problem of self-supervised method, the alignment of video frames and text descriptions is achieved, and the frame-level pseudo-importance scores is generated, which solves the problems of insufficient training data and poor generalization capabilities in the existing video digest methods, and improves model performance and summary reliability.
Patent Information
- Application Number
- CN202311104554.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-30
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-08-30
AI Technical Summary
The existing video digest methods rely on supervised learning to require a large amount of artificial annotation, which has problems such as insufficient training data and poor generalization ability. The unsupervised method lacks semantic guidance, poor model effect, and low universality.
The self-supervision method is adopted to achieve alignment of video frames and text descriptions through the inverse optimal transmission problem, generate frame-level pseudo-importance scores, and use pseudo-fractions to train the keyframe selector to construct a video summary.
Semantic information is provided without fine-grained annotation, improving model performance and generating video digests is more reliable.
Smart Images

Figure CN117150076B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and particularly to a method for video summarization in a self-supervised manner. Background Art
[0002] Video summarization technology is a technology that analyzes and processes videos to extract key information and streamline content, and is mainly applied to fields such as video retrieval, browsing, and video processing. When given a video composed of a number of video frames, the video summarization task is to select a set of video frames from the input video to construct a video summary, so that the generated summary can reflect the main content of the input video while being sufficiently concise.
[0003] In implementation, video summarization technology mainly relies on technical means such as image processing, machine learning, and natural language processing. By collecting, analyzing, processing, and mining video data, important visual information and semantic information are extracted to help users more quickly understand and obtain the key information of the video. Currently, common methods include: (1) Shot change detection method, which divides the video into different shot segments by detecting shot changes between adjacent frames, so as to quickly understand the video content. Typical shot change detection methods include color histogram matching, scene similarity, etc.; (2) Video key frame extraction method, which extracts the most expressive frames from the video to represent the entire video content. This method usually analyzes and compares the features of the frames (such as brightness, color histogram, saliency, etc.) to extract key frames; (3) Method for extracting objects with distinct backgrounds, which extracts the main content of the video by detecting objects with distinct backgrounds in the video. This method usually uses methods such as motion tracking, color segmentation, and edge detection; (4) PCA-based video dimensionality reduction method, which uses the PCA (Principal Component Analysis) algorithm to reduce the video dimension, so as to achieve fast browsing and summarization. This method first extracts the feature vectors of each frame, then calculates the projection weights through PCA technology, and finally filters out unnecessary video frames through a certain threshold; (5) Machine learning-based video summarization method, which uses machine learning algorithms to train a video analysis model to achieve tasks such as video classification and key frame extraction. For example, a convolutional neural network (CNN) is used to classify and extract key frames from the video.
[0004] In existing machine learning-based video summarization methods, the methods can generally be divided into unsupervised methods and supervised methods. In unsupervised methods, the system selects representative frames or segments according to some heuristic principles (e.g., visual differences between frames); while in supervised methods, the system learns these principles from labeled data. Most of the existing models are trained by supervised learning. When given frame-level annotations of a video, supervised learning methods usually abstract the video summarization task into a label prediction task, and by training different models, such as Determinantal Point Process, Long Short-Term Memory Network (LSTM), attention models, etc., to directly predict the key frames of the video. However, supervised learning methods require a large number of frame-level annotated videos as training data. Obtaining fine-grained annotations of videos through manual annotation is time-consuming and laborious, and at the same time has a certain degree of subjectivity. Therefore, video summarization methods based on supervised learning often suffer from the problems of insufficient training data and poor generalization ability. Methods based on unsupervised learning usually use adversarial learning techniques to predict key frames by training sequence models, and use discriminators to measure the quality of the selected key frames. Models based on unsupervised learning often have worse performance than those based on supervised learning due to the lack of semantic guidance information. In addition, Narasimhan M et al. proposed a weakly supervised IV-SUM model, which is a video summarization method dedicated to teaching video summarization. It calculates pseudo-video summaries (frame-level pseudo-importance scores) for a large number of teaching videos through weakly supervised means, and then uses this pseudo-video summary to train a Segment Scoring Transformer, which can output frame-level importance scores. According to the importance scores, the model can select key frames to construct a video summary. The method for generating pseudo-video summaries in the IV-SUM model is mainly based on two points: (1) Task relevance: When given multiple videos belonging to the same task, the important steps of this task should appear in multiple videos; (2) Cross-modal importance: When an important step appears, the narrator is likely to mention this step before, during, or after this step. Therefore, these key steps are likely to appear in the caption text obtained through Automatic Speech Recognition.The IV-SUM model generates pseudo video summaries based on these two assumptions, and then uses the pseudo summaries to train the segment scoring transformer to achieve model training (Narasimhan M, Nagrani A, Sun C, et al. TL; DW? Summarizing Instructional Videos with Task Relevance and Cross-Modal Saliency[C] / / Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXXIV. Cham: Springer Nature Switzerland, 2022: 540-557.). For models based on the weakly supervised learning method, auxiliary information is generally used as supervision information for model training, such as video-level labels (categories), user comments (bullet screens), etc. Although such auxiliary information is easy to obtain, this method can only be applied to some specific scenarios and has low universality. SUMMARY OF THE INVENTION
[0005] In view of this, the present invention provides a video summarization method in a self-supervised manner, which can extract key frames in a self-supervised manner by using the characteristics of video data itself.
[0006] The video summarization method in a self-supervised manner of the present invention includes:
[0007] Step 1, construct a training set, where the training set contains N videos; for each video in the training set, divide it into multiple video segments with the same duration, and generate the text of each video segment to obtain a set of "video-text" pairs where, represents the nth video contains I n video frames; represents the text corresponding to the nth video which contains a total of J texts, corresponding to the J n video segments of the video respectively; n
[0008] Step 2, for each "video-text" pair in the set D, extract the visual representation V of the video and the text representation W of the text;
[0009] Step 3, for Align the visual representation and the text representation of each "video - text" pair, specifically as follows:
[0010] Transform the alignment problem into an inverse optimal transport problem, that is, learn a text projection module f to achieve the projection from the text representation to the visual representation, and obtain a set of optimal transport solutions
[0011]
[0012]
[0013] Among them, T n represents the optimal transport matrix between the set of video frames and the set of texts where the n - th video is located, f represents the text projection module from the text domain to the visual domain, d is the Euclidean distance, f(w n,j ) is the projected text representation obtained by passing the text representation w n,j through the text projection module; "<A, B>" represents the dot product between matrix A and matrix B; τ is a hyperparameter, τ > 0; The KL - divergence regularization term is used to measure the distance between two probability distribution functions P(X) and Q(X), defined as are uniform distributions with lengths J n and I n and all values equal to 1; is the visual representation of the video segment corresponding to the j - th text in the n - th video;
[0014] Among them,
[0015] is the distance matrix between the projected text representations, and the Euclidean distance matrix is adopted; f(W n ) is the set of text representations obtained after being projected by the text projection module; is an edge - constraint matrix with a shape of J×J, where the diagonal elements of the matrix are 0 and the rest of the elements are 1;
[0016] Step 4, generate the frame - level pseudo - importance score of the video according to the optimal transport solution obtained in step 3; the frame - level pseudo - importance score is the weighted sum of the alignment score and the representation score;
[0017] Among them, the alignment score is: for a video with a total of I frames and J texts generated for this video, its optimal transport solution is denoted as Then the alignment score of the i - th video frame is: s a,i = maxj∈{1,...,J} s ij ,s ij is the relative importance between the j-th text and the i-th video frame, where u j is the text representation after projection of the j-th text, and v i is the visual representation of the i-th video frame.
[0018] The representation score is: for the i-th video frame, centered on this frame, there are a total of K' adjacent frames around it, and the adjacent frame set is defined as Then the representation score of the i-th video frame is:
[0019] Step 5, construct a key frame selector g. The key frame selector g takes the visual representation of the video as input and outputs the frame-level importance scores of each video frame; using the frame-level pseudo-importance scores of each video in the training set generated in Step 4 as pseudo-labels, train the key frame selector g to obtain the trained key frame selector g;
[0020] Step 6, for the video from which the summary is to be extracted, first divide the video into multiple video segments and extract the visual representation; then input the visual representation into the trained key frame selector g to obtain the frame-level importance scores of each video frame; finally, use the 0 / 1 knapsack algorithm to select the key frames to form the video summary.
[0021] Preferably, in Step 1, a text generator is used to generate a text that can describe the main content of the video segment.
[0022] Preferably, the text generator uses a hierarchical temporal-aware video-language pre-training framework or a bimodal transformer for text generation.
[0023] Preferably, in Step 2, a contrastive language-image pre-training model or a separable 3D convolutional neural network is used to extract the visual representation and the text representation.
[0024] Preferably, in Step 3, an alternating optimization strategy is adopted to solve the inverse optimal transport problem, specifically:
[0025] First, when the current text projection module f is given, solve N unbalanced Wasserstein distances D n (f), T n , n = 1, 2,..., N respectively according to the Bregman alternating direction method of multipliers to obtain N optimal transport schemes;
[0026] Then, update the text projection module f in a gradient-based manner through the loss function of the text projection module f. Among them, for the n-th video, d uw (V n ,f(W n )) = <D n (f),T n > + τR T (T n ),where γ and τ are hyperparameters for controlling the degree.
[0027] Preferably, the text projection module f adopts a multi - layer perceptron structure, specifically two fully - connected layers and a ReLU activation layer between the two fully - connected layers.
[0028] Preferably, in step 4, the frame - level pseudo - importance score of the i - th video frame is:
[0029]
[0030] where and are the normalized alignment score and the representation score.
[0031] Preferably, in step 5, the key - frame selector g is a Transformer encoder considering positional encoding.
[0032] Preferably, in step 5,
[0033] When training the key - frame selector, the loss function is:
[0034]
[0035] where loss is the mean - squared error or binary cross - entropy, is the frame - level importance score output by the key - frame selector, s n is the frame - level pseudo - importance score of the visual representation .
[0036] Preferably, in step 6, a series of scene change points of the video are detected by the temporal kernel segmentation algorithm, and the video is divided into M video segments by the scene change points, where M is the number of scene change points plus one.
[0037] Beneficial effects:
[0038] The present invention realizes the extraction of key frames in a self-supervised manner by mining the features inherent in video data as pseudo-labels. It neither requires fine-grained video labels nor needs to provide necessary semantic information through the features of the video itself. Taking video clips and their texts as representations, the alignment between video representations and text representations is achieved by solving the inverse optimal transport problem. Then, based on the characteristic that the generated text can reflect the main content of a certain clip, an alignment score is proposed based on the matching results between video frames and text; according to the characteristic that key frames in the video summary contain more information than ordinary frames, a representation score is proposed to characterize the amount of information contained in a frame based on the reconstruction error between the visual representation of the current frame and the visual representation of adjacent frames in the latent space. The alignment score and the representation score together constitute the frame-level pseudo-importance score proposed in the present invention. Finally, by using these pseudo-scores as offline supervision information to train the key frame selector, the performance of the model is greatly improved compared with previous models, and the obtained video summary is more reliable.
[0039] The present invention projects the text representation near the corresponding visual representation by solving the inverse optimal transport problem, realizing the alignment of different modality features. At the same time, an alternating optimization algorithm is proposed to effectively solve the inverse optimal transport problem, and a boundary regularization term is designed to avoid generating trivial solutions. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a flowchart of the method of the present invention;
[0041] Figure 2 is a schematic structural diagram of the video summary framework in the self-supervised manner of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The following examples are given in conjunction with the accompanying drawings to describe the present invention in detail.
[0043] The present invention provides a video summary method in a self-supervised manner, which is used to assist model training by mining the inherent feature information of video data. Specifically, the present invention first generates a series of texts according to the video data, and each text corresponds to a clip of the video. Then, the matching between video frames and text descriptions is realized by using the relevant theories and technologies of optimal transport. In particular, the present invention learns a text projection module by solving an inverse optimal transport problem, and through this module, the text representation can be projected to its corresponding visual representation in the latent space to achieve the alignment of modality representations. Then, based on the obtained optimal transport matrix, a method for generating a frame-level pseudo-importance score is proposed, and this pseudo-score consists of two types of sub-scores: alignment score and representation score. By using the obtained pseudo-scores as supervision information, a key frame selector can be trained. During testing, based on the frame-level importance scores output by the key frame selector, key frames can be selected to construct a video summary.
[0044] The model framework constructed by the present invention is as shown in Figure 2 and the method flow is as shown in Figure 1 which specifically includes:
[0045] S1. Prepare model input: video data.
[0046] Specifically, when given a video composed of I-frames, the video is denoted as where v i represents the i-th video frame and I represents the total number of frames in the video. The video summarization task is to select a set of key frames from all the I-frames, that is, {v i} i∈S and this set of key frames can reflect the main content of the original video.
[0047] S2. Generate a set of text descriptions through a pre-trained video caption model.
[0048] Specifically, when given an input video first, the video is divided into multiple video segments with the same duration. For each video segment, it is used as input to a text generator, which is a pre-trained video caption model. The output of the text generator is a text sentence that can describe the main content of the video segment. The set of all generated texts is used as the text set of the input video, denoted as In the present invention, different video caption models can be adopted, such as the Hierarchical Temporal-Aware Video-Language Pre-training framework (HiTeA), the Bi-Modal Transformer (BMT), etc.
[0049] S3. Learn the text projection module to achieve the alignment of model representations.
[0050] Specifically, assume that the set of training videos contains N training videos, denoted as According to the steps in S2, the generated text sets corresponding to each training video can be obtained respectively, so as to obtain a paired "video-text" set denoted as where represents that the n-th video has I n video frames, represents that there are J n texts, corresponding to the J n video segments of the n-th video respectively.
[0051] For each pair of video and text in the "video - text" collection Visual and text representations are respectively extracted from each pair of video and text in it through a pre - trained vision - language model. Different pre - trained vision - language models can be used to meet the requirements, such as the Contrastive Language - Image Pre - training (CLIP) model, the Separable 3D Convolutional Neural Network (S3D), etc. Specifically, taking the pre - trained CLIP model as an example, when given a video with a total of I frames and a set of J texts generated from this video Pass them through the pre - trained CLIP model to obtain their D - dimensional visual representations and text representations where, f v is the visual encoder of the pre - trained vision - language model, and f w is the text encoder of the pre - trained vision - language model. In order to build the semantic association between the visual representation and the text representation, it is necessary to align the two in the latent space. A direct way is to calculate the Wasserstein distance between the visual representation and the text representation, and this distance can be defined as:
[0052]
[0053] where is the distance matrix, and each element d(v i , w j ) of the matrix represents the distance between the element v i in the set V and the element w j in the set W. For the Wasserstein distance, the Euclidean distance matrix is often used to calculate the distance matrix. Π(u, μ) = {T≥0|T1 J = u, T T 1 I = μ} is the set of doubly - stochastic matrices, and its marginal distributions must be on the simplex, such as u ∈ Δ I-1 and μ ∈ Δ J-1 . Generally, the Wasserstein metric sets the marginal distributions as uniform distributions, such as and and the optimal transport matrix corresponding to the distance d w (V, W), denoted as is the optimal joint distribution between the two sets under the condition of minimizing the expected distance. The element * of the optimal transport matrix T can be interpreted as the element v in the set V i and the element w in the set W j The probability of consistency between them. The sets V and W can be matched with each other according to the probability of consistency. For example, for the element v in the set V i , the corresponding element in the set W can be determined by .
[0054] In practical applications, the Wasserstein distance may have two problems: (1) The semantic association between the two modalities highly depends on the quality of the visual and text representations extracted by the pre-trained vision-language model. However, when visualizing the two-modal representations extracted by multiple vision-language models, the present invention finds that the distributions of the two representations are generally very different. The text representations usually cluster together, while the visual representations show some clustering structures according to different video segments. The visualization results indicate that the semantic consistency of the two-modal representations extracted by the pre-trained vision-language model under the current data distribution is not strong. (2) When given segment-level text representations and frame-level visual representations, a mechanism is needed to achieve partial matching between them, while the Wasserstein distance cannot achieve this mechanism. There are many redundant and meaningless frames in the video. Therefore, when matching video frames with text, these redundant frames should not be matched with the text. So a mechanism is needed to achieve partial matching.
[0055] To solve the above two problems, the present invention proposes a method of Inverse Optimal Transport (IOT). The IOT problem corresponds to a two-layer optimization problem. The upper-layer problem corresponds to the optimization of the unbalanced Wasserstein distance between the projected text representation and the frame-level visual representation. The lower-layer problem corresponds to the optimization of the projection module that projects the text representation near the visual representation of the corresponding video segment. The upper-layer problem corresponds to solving N optimal transport schemes by minimizing the sum of N unbalanced Wasserstein distances. Compared with the Wasserstein distance, the unbalanced Wasserstein distance used in the upper-layer problem introduces a regularization term acting on the marginal distribution of the transport scheme to replace the strict equality constraint originally imposed between the marginal distribution and the uniform distribution, enabling the method to achieve partial matching. The lower-layer problem corresponds to learning the text projection module by minimizing the sum of the distances between the representations of each text and the representations of its corresponding video segments. The connection between the upper and lower-layer problems lies in that the optimization of the projection module in the lower-layer problem affects the projected text representation, and thus affects the distance matrix in the unbalanced Wasserstein distance between the projected text representation and the frame-level visual representation in the upper-layer problem, ultimately affecting the solved unbalanced Wasserstein distance and the optimal transport scheme. Therefore, the two problems need to be considered comprehensively to achieve joint optimization. The IOT method aims at optimal transport, optimizes the underlying distance matrix, and avoids trivial solutions through a boundary regularization term based on KL divergence.
[0056] The inverse optimal transport method can be formally described by the formula. When given N video-text pairs and their corresponding representations We can learn a text projection module f to achieve the projection from the text domain to the visual domain and obtain a set of optimal transport schemes by solving an inverse optimal transport problem.
[0057]
[0058]
[0059] The IOT problem corresponds to a two-layer optimization problem. The upper-layer problem corresponds to the sum of N unbalanced Wasserstein distances, and each unbalanced Wasserstein distance corresponds to a video-text pair. For each video-text pair, there are two differences between the present invention and the Wasserstein distance: (1) Before calculating the distance matrix, we will first apply the text projection module f to achieve the projection of the text representation, that is, the distance matrix calculation formula is By using the text projection model, the semantic consistency between the visual representation and the text representation is enhanced, the distance between the representations is more realistic, and thus the optimal transport scheme is more reliable. (2) The present invention introduces an unbalanced regularization term acting on the marginal distributions of each transport scheme At the same time, a hyperparameter τ > 0 is used to control the weight of this regularization term. This regularization term uses the KL divergence between the marginal distribution of the transport scheme and the uniform distribution ( and ) as a constraint condition, rather than imposing a strict equality constraint. Therefore, it can support partial matching, that is, only some video frames will be matched with the text, so that meaningful video frames can be effectively selected and redundant video frames can be filtered out.
[0060] The lower-level problem corresponds to learning the text projection module. Specifically, since the text description is generated based on the video clip, the correspondence between the text description and the video clip is clear. Therefore, in the lower-level problem, the learning of the text projection module is achieved by minimizing the sum of the distances between the representations of each text and the representations of its corresponding video clip. Among them, the representation of the video clip corresponds to the mean of the representations of all frames within the video clip, denoted as In addition, the process of splitting continuous videos into fixed-length video clips may bring the problem of semantic ambiguity, that is, the two text representations corresponding to adjacent video clips will get closer in the latent space. To prevent the text projection module from projecting different text representations to the same point, for each video, a boundary regularization term based on KL divergence is proposed:
[0061]
[0062] Among them, is the distance matrix between the projected text representations, and the Euclidean distance matrix is used. Matrix is an edge constraint matrix with a shape of J×J. The values of the diagonal elements in the matrix are 0, and the values of the remaining elements are 1. The purpose of this regularization term is to enhance the diversity of the projected text representations.
[0063] The optimization problem can be solved by using an alternating optimization strategy. Specifically, when the current text projection module f is given, the N unbalanced Wasserstein distances can be solved respectively according to the Bregman Alternating Direction Method of Multipliers (B-ADMM) algorithm, and N optimal transport schemes can be obtained. The text projection module f can be updated based on the gradient of the loss function of the projection module Among them, for the nth video, d uw (V n ,f(W n )) = <D n (f), T n > + τR T (T n ). In the present invention, the architecture of the text projection module f is a multi-layer perceptron structure (Multi-Layer Perceptron, MLP), specifically two fully connected layers, and a ReLU activation layer between the two fully connected layers.
[0064] S4. Generate frame-level pseudo-importance scores for the video according to the optimal transport matrix.
[0065] After step S3, a trained text projection module and a set of optimal transport schemes can be obtained In this step, a method for calculating the frame-level pseudo-importance scores of the video is proposed. The pseudo-scores include two parts: alignment scores and representation scores.
[0066] Alignment scores: Considering that the generated text can capture the main content of a video clip, the matching result between the generated text and the video frames in a video clip can reflect the relative importance of each frame in the video clip. After solving the IOT problem in step S3, the learned optimal transport scheme has achieved cross-modal semantic alignment. Therefore, this optimal transport scheme can be used to characterize the importance of each frame. Specifically, when given a video with a total of I frames and J texts generated from this video, the corresponding optimal transport scheme is denoted as For each text w j , a set of K video frames that best match this text can be obtained Denote the alignment score as The calculation process of the alignment score is as follows: For each text w j , a set of video frames that best match it can be calculated Define the relative importance between the j-th text and the i-th video frame as where u j =f(w j ) is the text representation projected by the text projection module. After calculating all J texts, the importance scores of all video frames can be obtained. For example, the importance score s a,i of the i-th video frame is maxj∈{1,...,J} s ij . This calculation method can ensure that the video frames matching the generated text all have relatively high alignment scores.
[0067] Representation scores: The representation scores measure the representativeness of each video frame. This score is denoted as Considering that the video summarization task requires the video summary to reflect the story line of the original video, that is, the selected key frames can reconstruct adjacent frames within a tolerable error, that is, the key frames contain more information than ordinary frames. Therefore, the present invention uses the form of reconstruction error in the latent space to define the representation score. For the i-th video frame, considering the K' adjacent frames centered around this frame, the adjacent frame set is defined as Then the representation score of the i-th video frame is defined as:
[0068]
[0069] Among them, the representation score calculates the mean of the Euclidean distances between the representation of this video frame and the representations of its adjacent video frames. The lower the representation score, the more information this video frame contains and the stronger its representation ability.
[0070] Taking the above two scores into comprehensive consideration, we obtain the final frame-level pseudo-importance score through weighted averaging:
[0071]
[0072] Among them, and are the normalized alignment score and representation score, that is, for a score vector s a (s r ), we subtract the minimum value of this vector from it and divide it by the dynamic range of this vector (the maximum value minus the minimum value is the dynamic range), so that the value range of the normalized score is [0, 1]. The value range of the final pseudo-score is also [0, 1].
[0073] S5. Train the key frame selector according to the pseudo-score labels.
[0074] The pseudo-importance scores generated in step S4 can be used as pseudo-labels to train the key frame selector g proposed in the present invention. The architecture of the key frame selector is an encoder-only transformer that takes into account positional encoding. When given N training videos, their corresponding visual representations are The generated pseudo-importance scores are We can train the key frame selector through the following loss function:
[0075]
[0076] Among them, the loss function can be mean squared error or binary cross-entropy loss, is the frame-level importance score output by the key frame selector.
[0077] S6. Construct a video summary according to the 0 / 1 knapsack algorithm.
[0078] In step S5, the training of the key frame selector is completed. When the visual representation of the video frame is input into the key frame selector, the key frame selector will output the frame-level importance score. In the test phase, when a test video is given, we first use the Kernel Temporal Segmentation (KTS) algorithm to detect a series of scene change points of the video. Each scene change point is the subscript of the video frame where the scene changes. Through the scene change points, we can divide the test video into M video segments (M is not a fixed value, and its value is the number of change points plus one). After obtaining the frame-level importance scores [s i ∈ [0, 1] I of the test video through the key frame selector, the importance score of each video segment can be obtained by weighted averaging the importance scores of all video frames within each video segment:
[0079]
[0080] where is the set of all video frames within the m-th video segment, is the number of video frames included in the m-th video segment.
[0081] When obtaining the segment-level importance scores we use the 0 / 1 knapsack algorithm to select the key video segments, that is, under the condition that the total number of frames selected does not exceed a certain set threshold, the sum of the importance scores of the selected segments is maximized. It can be formally defined in formula as:
[0082]
[0083] where is a binary classification vector (i.e., the component values are 0 or 1), used to indicate whether the video segment is selected as a key segment. is a vector, used to represent the number of frames included in each video segment. L is the set length of the video summary. By solving this problem, we can select the key video segments. By concatenating these video segments, the final video summary can be obtained.
[0084] In summary, the above is only the preferred embodiment of the present invention and is not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A video summarization method in a self-supervised manner, characterized in that, Including: Step 1, construct a training set that contains N videos; for each video in the training set, divide it into multiple video segments of the same duration, and generate the text of each video segment to obtain a set of "video-text" pairs Among them, represents the nth video with I n video frames; represents the nth video corresponding text with a total of J n texts, respectively corresponding to the J video segments of the video n ; Step 2: For each "video - text" pair in set D, extract the visual representation V of the video and the text representation W of the text; Step 3, align the visual representation and the text representation of each "video-text" pair in specifically as follows: Convert the alignment problem into an inverse optimal transport problem, that is, learn a text projection module f to project from text representations to visual representations and obtain a set of optimal transport solutions Among them, T n represents the optimal transport matrix between the video frame set and the text set of the nth video, f represents the text projection module from the text domain to the visual domain, d is the Euclidean distance, f(w n,j ) is the projected text representation obtained by passing the text representation w n,j through the text projection module; "<A,B>" represents the dot product between matrices A and B; τ is a hyperparameter, τ > 0; The KL divergence regularization term is used to measure the distance between two probability distribution functions P(X) and Q(X), defined as are uniform distributions with lengths J n and I n and all values equal to 1; is the visual representation of the video segment corresponding to the jth text in the nth video; Among them, is the distance matrix between the projected text representations, using the Euclidean distance matrix; f(W n ) is the set of text representations obtained after being projected by the text projection module; is an edge constraint matrix with a shape of J×J, where the values of the diagonal elements in the matrix are 0 and the values of the remaining elements are 1; Step 4, according to the optimal transmission scheme obtained in Step 3 Generate the frame-level pseudo-importance score of the video; the frame-level pseudo-importance score is the weighted sum of the alignment score and the characterization score; Among them, the alignment score is: for a video with a total of I frames and J texts generated by the video, its optimal transmission scheme is denoted as Then the alignment score of the i-th video frame is: s a,i = max j∈{1,…,J} s ij , s ij is the relative importance between the j-th text and the i-th video frame, where u j is the text representation after projection of the j-th text, and v i is the visual representation of the i-th video frame. The representation score is: For the $i$-th video frame, with this frame as the center, there are a total of $K'$ adjacent frames around it. The set of adjacent frames is defined as Then the representation score of the $i$-th video frame is: Step 5: Construct a key - frame selector g. The key - frame selector g takes the visual representation of the video as input and outputs the frame - level importance scores of each video frame. Using the frame - level pseudo - importance scores of each video in the training set generated in Step 4 as pseudo - labels, train the key - frame selector g to obtain the trained key - frame selector g; Step 6: For the video from which the summary is to be extracted, first divide the video into multiple video segments and extract the visual representation. Then input the visual representation into the trained key - frame selector g to obtain the frame - level importance scores of each video frame. Finally, use the 0 / 1 knapsack algorithm to select the key frames to form the video summary.
2. The method according to claim 1, wherein In Step 1, use a text generator to generate a sentence that can describe the main content of the video segment.
3. The method according to claim 2, wherein The text generator uses a hierarchical temporal - aware video - language pre - training framework or a dual - modal transformer for text generation.
4. The method according to claim 1, wherein In Step 2, use a contrastive language - image pre - training model or a separable 3D convolutional neural network to extract the visual representation and the text representation.
5. The method according to claim 1, wherein In Step 3, use an alternating optimization strategy to solve the inverse optimal transport problem. Specifically: First, when the current text projection module f is given, the N unbalanced Wasserstein distances D n (f), T n are solved respectively according to the Bregman alternating direction multiplier method for n = 1, 2, …, N, and N optimal transport plans are obtained respectively; Then, update the text projection module f in a gradient-based manner through the loss function of the text projection module f, where, for the nth video, γ and τ are hyperparameters for controlling the degree.
6. The method according to claim 1 or 5, characterized in that The text projection module f adopts a multi - layer perceptron structure, specifically two fully - connected layers and a ReLU activation layer between the two fully - connected layers.
7. The method according to claim 1, characterized in that In Step 4, the frame - level pseudo - importance score of the i - th video frame is: Among them, and are the aligned score and the characterization score after standardization.
8. The method according to claim 1, characterized in that In Step 5, the key - frame selector g is a transformer encoder that considers positional encoding.
9. The method according to claim 1 or 8, characterized in that, In Step 5, When training the key - frame selector, the loss function is: where loss is the mean squared error or binary cross - entropy, is the frame - level importance score output by the key - frame selector, s n is the visual representation for the frame - level pseudo - importance score.
10. The method according to claim 1, characterized in that, In Step 6, use a temporal kernel segmentation algorithm to detect a series of scene change points of the video. Divide the video into M video segments through the scene change points, where M is the number of scene change points plus one.
Citation Information
Patent Citations
Video abstract generation method and device, computer device and medium
CN113052149A
Video understanding method
CN115578680A