A noise-robust learning method for multi-modal video highlight detection

By combining visual and audio features with low-dimensional compression and global-local enhancement, using multimodal interactive fusion module and clean sample training, the problems of semantic gaps and noise labels between modes in multimodal video highlight detection are solved, and the detection accuracy and robustness are improved.

CN117058571BActive Publication Date: 2025-07-04NINGBO UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310850066.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-12
Publication Date
2025-07-04
Estimated Expiration
2043-07-12

AI Technical Summary

Technical Problem

The existing multimodal video highlight detection methods have the problem that semantic gaps between modes are difficult to integrate, and the lack of full utilization of global-local spatial feature information and noise labels lead to poor training results.

Method used

By linearly mapping visual and audio features to low dimensions, the global-local enhancement dual-transformer module and multimodal interactive fusion module are used, combining global-local spatial features and timing feature information, and training by screening clean samples, the target loss function optimization network is constructed.

Benefits of technology

It improves the accuracy and robustness of video highlight detection, reduces parameter amount and video memory usage, enhances the integration of modal representation information, and ensures the stability and efficiency of the training process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117058571B_ABST
    Figure CN117058571B_ABST
Patent Text Reader

Abstract

The present invention relates to a noise-robust multi-modal video highlight detection learning method. This method takes the visual and audio multi-modal representation information of the video as input, and then compresses the high-dimensional features into a low-dimensional feature through linear mapping to reduce the number of parameters and video memory occupancy, while improving the running speed. Then, it combines the internal global-local spatial feature information of video segments and the temporal feature information between video segments to enhance the modal representation information. Next, through a multi-modal interaction and fusion module, the visual and audio modal representation information is interacted and fused. Finally, a target loss function is constructed to optimize the network. This method can effectively improve the accuracy and robustness of video highlight detection, has a wide range of applications, and is particularly suitable for video highlight detection of noisy label data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video highlight detection, and in particular to a noise-robust multi-modal video highlight detection learning method. Background Art

[0002] Video highlight detection aims to automatically and quickly detect the highlight moment segments of a video through various machine learning and computer vision methods, and can be applied to fields such as video editing, video search, and video recommendation. Since a video is composed of multi-modal components, this provides important clues for video highlight detection. However, most existing methods focus on only using visual single-modal representation information for video highlight detection; compared with previous single-modal video highlight detection methods, there are currently some multi-modal video highlight detection methods, and the current existing multi-modal video highlight detection methods have the following three challenges: 1) For different modalities, there is a semantic gap between modalities, and it is difficult to fully fuse the modal representation information; 2) Previous methods only consider the temporal feature relationship between segments within a modality and do not fully utilize the global-local spatial feature information; 3) Since the annotation of highlight segments has a certain subjectivity, the annotation information has a certain uncertainty and there are noise labels, which leads to misguidance of the labels and thus poor supervised training effects. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a noise-robust multi-modal video highlight detection learning method that can fully fuse modal representation information, fully utilize global-local spatial feature information, and improve the accuracy and robustness of video highlight detection.

[0004] The technical solution adopted by the present invention is a noise-robust multi-modal video highlight detection learning method, and the method includes the following steps:

[0005] S1. Divide a video into several video segments, and for each video segment, extract its visual features and audio features;

[0006] S2. Uniformly compress the visual features and audio features extracted in step S1 to low-dimensional visual features and low-dimensional audio features of the same dimension through a linear mapping fully connected layer;

[0007] S3. Use a global-local enhancement dual transformer module to enhance the modal representation information of the low-dimensional visual features and the modal representation information of the low-dimensional audio features respectively, and obtain the enhanced visual modal representation and the enhanced audio modal representation where i represents the i-th video segment and n represents the total number of video segments;

[0008] S4. Use the multi-modal interaction fusion module to perform interaction fusion on the enhanced visual modality representation and the enhanced audio modality representation to output the post-interaction visual modality representation and the post-interaction audio modality representation and by and fuse and output the multi-modal representation

[0009] S5. Predict the saliency scores of each visual feature based on the visual modality representation to obtain the saliency score A1 corresponding to each visual feature, normalize and sort the loss values of each visual feature according to the obtained saliency score A1, set the clean sample ratio, and select the visual features with small loss values according to the ratio to form the visual clean sample dataset N v ; Predict the saliency scores of each audio feature based on the audio modality representation to obtain the saliency score A2 corresponding to each audio feature, normalize and sort the loss values of each audio feature according to the obtained saliency score A2, set the clean sample ratio, and select the audio features with small loss values according to the ratio to form the audio clean sample dataset N a ;

[0010] S6. Predict the saliency score of the multi-modal feature based on the multi-modal representation to obtain the saliency score A3; Use the union N v of the selected visual clean sample dataset N a and the audio clean sample dataset N m as the training sample of the multi-modal representation to train the multi-modal representation to obtain the trained multi-modal representation network; Use the selected visual clean sample dataset N v as the training sample of the visual modality representation to train the visual modality representation to obtain the trained visual modality representation network; Use the selected audio clean sample dataset N a as the training sample of the audio modality representation to train the audio modality representation to obtain the trained audio modality representation network;

[0011] S7. Construct the target loss function according to the three-way saliency scores A1, A2, and A3 obtained in step S5, and optimize the trained visual modality representation network, trained audio modality representation network, and trained multi-modal representation network with the target loss function.

[0012] The beneficial effects of the present invention are as follows: By taking the visual and audio modal representation information of the video as inputs, the present invention then compresses the high-dimensional features into low-dimensional features through linear mapping to reduce the number of parameters and video memory occupancy, while improving the running speed. Then, by combining the global-local spatial feature information within the video segment and the temporal feature information between video segments, the modal representation information is enhanced; through the multi-modal interaction and fusion module, the visual and audio modal representation information is fully interacted and fused; according to the principle that the loss trained with noise samples often has large fluctuations, while the loss trained with clean samples is often small and stable, the loss of each sample is normalized and sorted, and then the samples with smaller loss are selected as clean samples. The union of the two-way selected clean samples is used as the samples for the fusion prediction branch. To prevent the interference of noise samples on training, the network is trained with the screened and cleaned clean samples, so that the entire training process is more robust and easier to converge to the optimum, and the accuracy of video highlight detection is improved.

[0013] Preferably, in step S3, the global-local enhanced dual-transformer module is composed of a Transformer structure with a global-local attention layer and a classical Transformer structure connected in series; the specific process of enhancing the modal representation information of the low-dimensional visual features and the low-dimensional audio feature modal representation information by the global-local attention layer is as follows: The modal representation information of the low-dimensional visual features and the modal representation information of the low-dimensional audio features are input into the global-local attention layer. Each input modal representation information is divided into two branches, global and local, for operation, and then the outputs after the two-branch operations are concatenated. After concatenation, the global and local spatial features are fused through convolution as the final output. By combining the global-local spatial feature information within the video segment and the temporal feature information between video segments, the modal representation information is enhanced, combining the advantages of spatial feature information and temporal feature information, and making more full use of the modal representation information.

[0014] Preferably, in step S4, the specific process of using the multi-modal interaction and fusion module to interact and fuse the enhanced visual modal representation and the enhanced audio modal representation includes the following steps:

[0015] S4.1, introducing a trainable token sequence The trainable token sequence is respectively subjected to multi-head cross-attention with the visual modal representation information and the audio modal representation information to perform multi-head cross-attention on the visual modal representation information and the audio modal representation information Compress them into trainable token sequences respectively to obtain the compressed visual modality representation and the compressed audio modality representation;

[0016] S4.2. Add the representation information of the three parts: the compressed visual modality representation, the compressed audio modality representation, and the trainable token sequence to obtain the intermediate compressed modality feature

[0017] S4.3. Expand the intermediate compressed modality feature and use multi-head attention to propagate it to the visual modality representation information and the audio modality representation information respectively.

[0018] Preferably, in step S7, the objective loss function is expressed as:

[0019] where g i represents the true label of the i-th video segment, y i represents the predicted score of the i-th video segment, and N represents the number of video segments;

[0020] represents the saliency loss function of the visual modality representation, represents the saliency loss function of the audio modality representation, represents the branch saliency loss function of the multi-modal representation, L cons represents the consistency loss function, and β represents the hyperparameter; y vi represents the predicted score of the i-th visual feature segment, y ai represents the predicted score of the i-th audio feature segment, y mi represents the predicted score of the i-th multi-modal feature segment, and N' represents the number of filtered clean samples. Description of the Drawings

[0021] Figure 1 is a process diagram of a noise-robust multi-modal video highlight detection learning method of the present invention;

[0022] Figure 2 is a structural diagram of the global-local enhanced dual-transformer module in the present invention;

[0023] Figure 3 is a structural diagram of the global-local attention layer of the global-local enhanced dual-transformer module in the present invention;

[0024] Figure 4It is the structural diagram of the multi-modal interaction fusion module in the present invention;

[0025] Figure 5 It is the schematic diagram of the visualization result obtained through experiments in the present invention. Specific implementation manners

[0026] The following further describes the invention with reference to the accompanying drawings and in combination with specific implementation manners, so that those skilled in the art can implement it according to the text of the specification. The protection scope of the present invention is not limited to this specific implementation manner.

[0027] The present invention relates to a noise-robust multi-modal video highlight detection learning method, as Figure 1 shown, the method includes the following steps:

[0028] S1. Divide a video into several video segments. For each video segment, extract its visual features and audio features, and perform feature extraction by a pre-trained feature extractor; in the implementation example, the visual feature extractor uses the I3D network for feature extraction, and the audio feature uses the PANN model pre-trained on AudioSets for feature extraction;

[0029] S2. Compress the visual features and audio features extracted in step S1 into low-dimensional visual features and low-dimensional audio features through a linear mapping fully connected layer; in the implementation example, the compression unified dimension is 256 dimensions;

[0030] S3. As Figure 2 shown, adopt the global-local enhancement dual transformer module to enhance the modal representation information of the low-dimensional visual features and the modal representation information of the low-dimensional audio features respectively, and obtain the enhanced visual modal representation and the enhanced audio modal representation where i represents the i-th video segment and n represents the total number of video segments; the global-local enhancement dual transformer module combines the global-local spatial feature information within the video segment and the temporal feature information between video segments to enhance the modal representation information;

[0031] As Figure 2 shown, the global-local enhancement dual transformer module is composed of a Transformer structure with a global-local attention layer and a classical Transformer structure connected in series; Figure 2 In, the Transformer structure with a global-local attention layer includes a global-local attention layer, a first layer normalization, a pre-feedforward layer, and a second layer normalization; the classical Transformer structure includes a multi-head attention layer, a first layer normalization, a pre-feedforward layer, and a second layer normalization; as Figure 3As shown, the specific process of the global-local attention layer for enhancing the modal representation information of low-dimensional visual features and the modal representation information of low-dimensional audio features is as follows: The modal representation information of low-dimensional visual features and the modal representation information of low-dimensional audio features are input into the global-local attention layer. Each input modal representation information is divided into two branches, global and local, for operation. Then, the outputs after the operations of the two branches are concatenated, and the global and local spatial features are fused through convolution as the final output. In the implementation example, the local branch convolution layer uses dilated convolution, batch normalization, and the activation function is the Relu function.

[0032] S4. Use the multi-modal interaction fusion module to perform interaction fusion on the enhanced visual modal representation and the enhanced audio modal representation to output the visual modal representation after interaction and the audio modal representation after interaction and and fuse the output to obtain the multi-modal representation As Figure 4 shown, the specific process of using the multi-modal interaction fusion module to perform interaction fusion on the enhanced visual modal representation and the enhanced audio modal representation is as follows:

[0033] S4.1. Introduce a trainable token sequence Perform multi-head cross-attention on the trainable token sequence with the visual modal representation information and the audio modal representation information respectively. Compress the visual modal representation information and the audio modal representation information into the trainable token sequence respectively to obtain the compressed visual modal representation and the compressed audio modal representation. In the implementation example, N b is set to 4.

[0034] S4.2. Add the three parts of the representation information of the compressed visual modal representation, the compressed audio modal representation, and the trainable token sequence to obtain the intermediate compressed modal feature

[0035] S4.3. Expand the intermediate compressed modal feature and use multi-head attention to spread it to the visual modal representation information and the audio modal representation information respectively.

[0036] S5. According to the visual modality representation Predict the saliency scores of each visual feature to obtain the saliency score A1 corresponding to each visual feature. According to the obtained saliency score A1, perform binary cross-entropy loss on the saliency score A1 and the label annotation information, record the loss value of each visual feature and perform normalized sorting. Due to the principle that noise samples often have large fluctuations in loss while clean samples often have small and stable loss values, set the clean sample ratio, and select visual features with small loss values according to the ratio to form the visual clean sample dataset N v ; According to the audio modality representation Predict the saliency scores of each audio feature to obtain the saliency score A2 corresponding to each audio feature. According to the obtained saliency score A2, perform binary cross-entropy loss on the saliency score A2 and the label annotation information, record the loss value of each audio feature and perform normalized sorting, set the clean sample ratio, and select audio features with small loss values according to the ratio to form the audio clean sample dataset N a ;

[0037] S6. According to the multi-modal representation Predict the saliency scores of the multi-modal features to obtain the saliency score A3; adopt the union N v of the selected visual clean sample dataset N a and the audio clean sample dataset N m as the training samples of the multi-modal representation to train the multi-modal representation to obtain the trained multi-modal representation network; adopt the selected visual clean sample dataset N v as the training samples of the visual modality representation to train the visual modality representation to obtain the trained visual modality representation network; adopt the selected audio clean sample dataset N a as the training samples of the audio modality representation to train the audio modality representation to obtain the trained audio modality representation network;

[0038] S7. Construct an objective loss function according to the three-way saliency scores A1, A2, and A3 obtained in step S5, and optimize the trained visual modality representation network, the trained audio modality representation network, and the trained multi-modal representation network by the objective loss function; the objective loss function is expressed as:

[0039] where g i represents the true label of the i-th video segment, yi denotes the predicted score of the \(i\)-th video segment, and \(N\) denotes the number of video segments;

[0040] denotes the saliency loss function of the visual modality representation denotes the saliency loss function of the audio modality representation

[0041] denotes the branch saliency loss function of the multimodal representation, \(L\) cons denotes the consistency loss function, and \(\beta\) denotes the hyperparameter;

[0042] \(y\) vi denotes the predicted score of the \(i\)-th visual feature segment, \(y\) ai denotes the predicted score of the \(i\)-th audio feature segment, \(y\) mi denotes the predicted score of the \(i\)-th multimodal feature segment, and \(N'\) denotes the number of filtered clean samples.

[0043] The following uses specific experiments to illustrate the effectiveness of a noise-robust multimodal video highlight detection learning method of the present invention:

[0044] Dataset:

[0045] This experiment selects the publicly available datasets YouTube Highlights and TvSum commonly used in the field of video highlight detection for experiments. The dataset YouTube Highlights is a popular video highlight detection dataset that collects videos from six different fields; each field contains 50 to 90 videos of different durations. Each video is divided into several video segments, containing approximately 100 frames, and each video segment has three different labels: 1 - selected as a highlight by the user; 0 - borderline case, -1 - non-highlight; borderline cases are regarded as non-prominent samples. This dataset contains six different categories: dogs, gymnastics, parkour, skating, skiing, and surfing. Each category has approximately 100 videos; the total length of all videos is 1430 minutes; segment-level annotations are provided to indicate whether the segment is a highlight moment; the dataset TvSum is a video summary dataset that contains 50 user videos collected from YouTube, contains 10 query labels for specific categories, contains 10 different categories, and each category has 5 videos. Following the convention, a 0.8 / 0.2 training / test split is randomly performed, and the titles annotating these videos are used as text queries.

[0046] Parameter configuration:

[0047] In all experiments, we used the Adam optimizer with a learning rate of 1e-3 and a weight decay of 1e-4. All global-local enhanced dual-transformer modules used learnable positional encoding, layer normalization, 8 heads, and a dropout of 0.1; the number of trainable token sequences N b was set to 4; the model was trained for 100 epochs with a batch size of 4 on the YouTube Highlights dataset and for 500 epochs with a batch size of 1 on the TvSum dataset.

[0048] Comparison with similar video highlight detection methods:

[0049] Video highlight detection methods can be classified into two categories: weakly / unsupervised and supervised methods, depending on whether there is supervision information. Supervised methods can be further divided into unimodal and multimodal methods. The following table presents a comparison of the video highlight detection results of existing different mainstream video highlight detection methods on the publicly available YouTube and TvSum datasets. In Table 1, LSVM, LIM-S, GNN, SL, SA, and PLD are unimodal methods, while MINI-Net, TGG, Join-VA, UMT, and Ours are multimodal methods.

[0050] Table 1. Results of video highlight detection using different methods on the YouTube dataset

[0051]

[0052] Table 2. Results of video highlight detection using different methods on the TvSum dataset

[0053]

[0054]

[0055] As can be seen from Table 1, the method of the present invention was compared with existing methods for highlight detection of six YouTube video categories. For the six video categories, the method of the present invention achieved the best performance in four video categories and the best average performance; moreover, it was also superior to four multimodal-based algorithms, such as MININet, TGG, JVA, and UMT.

[0056] As can be seen from Table 2, the method of the present invention was compared with the existing multimodal method UMT for highlight detection in 10 video categories. The method of the present invention was superior to the UMT method in 9 video categories, and these results verify the effectiveness of the method proposed in the present invention.

[0057] Visualization results:

[0058] To further evaluate the performance of the method of the present invention, as Figure 5 shown, the visualization results of the segment image frames corresponding to the videos of 6 categories in the YouTube dataset and the predicted saliency scores of the segments are presented; the higher the predicted score, the higher the possibility of representing the highlight segment of the video, thus proving the effectiveness of the method of the present invention.

Claims

1. A noise-robust multi-modal video highlight detection learning method, characterized in that: The method includes the following steps: S1. Divide a video into several video segments, and for each video segment, extract its visual features and audio features; S2. Uniformly compress the visual features and audio features extracted in step S1 into low-dimensional visual features and low-dimensional audio features of the same dimension via a linear mapping fully connected layer; S3. Use the global-local enhancement dual-transformer module to enhance the modal representation information of the low-dimensional visual features and the modal representation information of the low-dimensional audio features respectively, to obtain the enhanced visual modal representation and the enhanced audio modal representation , where represents the th video segment, and represents the total number of video segments; S4. Use the multi-modal interaction fusion module to process the enhanced visual modality representation and the enhanced audio modality representation for interactive fusion, and output the interactive visual modality representation and the interactive audio modality representation , and by and fuse and output the multi-modal representation ; S5. According to the visual modality representation Predict the saliency scores of each visual feature to obtain the saliency score A1 corresponding to each visual feature. Normalize and rank the loss values of each visual feature according to the obtained saliency score A1. Set the clean sample ratio, and select the visual features with small loss values according to the ratio to form a visual clean sample data set ; According to the audio modality representation Predict the saliency scores of each audio feature to obtain the saliency score A2 corresponding to each audio feature. Normalize and rank the loss values of each audio feature according to the obtained saliency score A2. Set the clean sample ratio, and select the audio features with small loss values according to the ratio to form an audio clean sample data set ; S6. According to the multi-modal representation Predict the saliency score of the multi-modal features to obtain the saliency score A3; use the visually clean sample dataset obtained by screening And the audio clean sample dataset Of the union As the training samples of the multi-modal representation Train the multi-modal representation To obtain the trained multi-modal representation network; use the visually clean sample dataset obtained by screening As the training samples of the visual modal representation Train the visual modal representation To obtain the trained visual modal representation network; use the audio clean sample dataset obtained by screening As the training samples of the audio modal representation Train the audio modal representation To obtain the trained audio modal representation network; S7. Construct a target loss function according to the three-way saliency scores A1, A2, and A3 obtained in step S5, and optimize the trained visual modality representation network, the trained audio modality representation network, and the trained multi-modal representation network by the target loss function.

2. A noise-robust multi-modal video highlight detection learning method according to claim 1, characterized in that: In step S3, the global-local enhanced dual transformer module is composed of a Transformer structure with a global-local attention layer and a classical Transformer structure connected in series; the specific process of enhancing the modality representation information of the low-dimensional visual features and the modality representation information of the low-dimensional audio features by the global-local attention layer is as follows: the modality representation information of the low-dimensional visual features and the modality representation information of the low-dimensional audio features are input into the global-local attention layer, and each input modality representation information is divided into two branches, global and local, for operation, and then the outputs after the operations of the two branches are spliced, and the global and local spatial features are fused through convolution after splicing as the final output.

3. A noise-robust multi-modal video highlight detection learning method according to claim 2, characterized in that: In step S4, the multi-modal interaction fusion module is used to perform interaction fusion on the enhanced visual modality representation and the enhanced audio modality representation The specific process of the interaction fusion includes the following steps: S4.

1. Introduce a trainable token sequence , and perform multi-head cross-attention on the trainable token sequence with the visual modality representation and the audio modality representation respectively, and compress the visual modality representation and the audio modality representation into the trainable token sequence respectively to obtain the compressed visual modality representation and the compressed audio modality representation; S4.

2. Add the representation information of the three parts, namely the compressed visual modality representation, the compressed audio modality representation, and the trainable token sequence to obtain the intermediate compressed modality feature ; S4.

3. Expand the intermediate compression modality features and use multi-head attention to propagate them to the visual modality representation and the audio modality representation respectively.

4. A noise-robust multi-modal video highlight detection learning method according to claim 3, characterized in that: In step S7, the target loss function is expressed as: ; where , represents the true label of the th video segment, is expressed as the predicted score of the th video segment, represents the number of video segments; represents the saliency loss function of the visual modality representation, represents the saliency loss function of the audio modality representation, represents the branch saliency loss function of the multimodal representation, represents the consistency loss function, represents a hyperparameter; , is expressed as the predicted score of the th visual feature segment, represents the th predicted score of the audio feature segment, is expressed as the predicted score of the th multimodal feature segment, represents the number of filtered clean samples.

Citation Information

Patent Citations

  • Multi-mode video highlight detection method and system based on commodity perception

    CN112801762A

  • Video detection method and system, storage medium and server

    CN114581821A