Video time retrieval and highlight detection method and system based on double semantic alignment
By adopting dual semantic alignment methods in video moment retrieval and highlight detection, using significance comparison learning and center distance regression, the shortcomings of time-level and fragment-level semantic alignment in the prior art are solved, and more accurate and complete semantic alignment is achieved, and model performance is significantly improved.
Patent Information
- Application Number
- CN202510072981.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
Existing video moment retrieval and highlight detection methods have shortcomings in time-level and fragment-level semantic alignment, especially when dealing with partial intersections, it is difficult to achieve accurate semantic alignment.
Using a dual semantic alignment method, semantic alignment at the fragment level and time level is achieved through significant comparison learning and central distance regression. The specific steps include: extracting the features of the video and text, performing cross-attention operations, extracting joint features using the encoder and decoder, achieving semantic alignment through significant comparison learning and center distance regression, and optimizing time matching through Hungarian algorithms.
It improves the accuracy of time-level semantic alignment and the integrity of fragment-level semantic alignment, and significantly improves the performance of video moment retrieval and highlight detection.
Smart Images

Figure CN120011593A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a video moment retrieval and highlight detection method based on dual semantic alignment. Background Art
[0002] The task of video moment retrieval and highlight detection aims to find one or more video moments that are closest to the semantics of the text query given by the user from the original video, and predict a saliency score for each unit video clip. However, it is very time-consuming to manually find key information from the video and understand the highlights in the video. Therefore, there is an urgent need for an automated video moment retrieval and highlight detection tool to help users understand the key information in the video.
[0003] In recent years, this task has received extensive attention and in-depth exploration from many researchers. The cross-modal video moment retrieval method based on cross-modal dynamic convolutional network proposed by Xu et al. focuses on the core task of video moment retrieval. Although its scaled intersection-over-union loss has taken into account the diverse values of the intersection-over-union ratio, in actual application scenarios, in the face of partial intersections, the change in the loss value is still slightly slow, which is not conducive to achieving moment-level semantic alignment. The joint moment retrieval and highlight detection method and system based on multi-task reciprocity proposed in Publication No. CN117648463A and the joint moment retrieval and highlight detection method and system based on multi-scale difference proposed in Publication No. CN 117668293A are both committed to overcoming the joint task difficulties of video moment retrieval and highlight detection. However, in the regression process, only L1 loss is used. When the loss value is small, oscillation is prone to occur, resulting in difficulty in convergence, and it is difficult to achieve accurate alignment of moment-level semantics; in the highlight detection task, only a few boundary fragments are used to penalize the model, which makes it difficult to achieve segment-level semantic alignment of the entire video.
[0004] Although existing methods have made some progress, there are still many unsatisfactory aspects in the semantic alignment of moment-level and segment-level video text. Specifically, in terms of moment-level semantic alignment, existing methods pay too much attention to the situation where the predicted moment and the real annotated moment are completely non-intersecting, and are not sensitive enough to the situation where they are partially intersecting, which greatly reduces the performance of the model in video moment retrieval. However, during the training phase, more than 90% of the predicted moments have partial intersections with the real moments, so these methods cannot achieve accurate moment-level semantic alignment. In terms of segment-level semantic alignment, existing methods only select a few unit segments, such as: a high-scoring segment and a low-scoring segment within the real annotated moment, or a segment within the range of the real annotated moment and a segment outside the range of the real annotated moment. Obviously, this selective strategy is usually highly random, and these methods cannot accurately constrain all segments in the entire video, resulting in incomplete segment-level semantic alignment in highlight detection, making the model easily disturbed by other segments with similar semantics.
[0005] In general, existing methods for video moment retrieval and highlight detection have the following shortcomings: 1) In regression, only the case where the predicted moment and the real moment are completely disjoint is considered, while the case where they partially intersect is ignored, so that moment-level semantic alignment cannot be achieved. 2) In the saliency loss, only 4 special clips are randomly selected for penalty judgment, without considering the saliency of all clips in the entire video, so accurate clip-level semantic alignment cannot be achieved. Summary of the invention
[0006] Purpose of the invention: The purpose of the present invention is to address the deficiencies in the prior art and to provide a method and system for video moment retrieval and highlight detection based on dual semantic alignment, which performs segment-level alignment and moment-level feature representation through a saliency contrast learning method and a center distance regression method, so that the model can learn more accurate moment-level semantic information, perform highlight detection by accurately predicting saliency scores, and output more accurate moment retrieval.
[0007] Technical solution: The video moment retrieval and highlight detection method based on dual semantic alignment of the present invention comprises the following steps:
[0008] S1. Using pre-trained visual encoder E v and text encoder E t Extract visual features X from the video respectively v and the text feature X in the text query t ;
[0009] S2, the visual feature X v and the text feature X tPerform cross attention operation to obtain the joint feature X;
[0010] S3, extracting the joint feature X by using an encoder, outputting the encoded feature F and a segment-level saliency representation for highlight detection; processing the feature F output by the encoder by using a decoder, and outputting a predicted moment representation for video moment retrieval;
[0011] S4, divide the feature F into positive sample features and negative sample features through the significance score threshold, and construct a significance contrast learning loss function to achieve segment-level semantic alignment;
[0012] S5. Calculate the center distance d between the predicted time output by the decoder and the real time t To perceive the position relationship at different times, use the distance-time intersection loss Achieve moment-level semantic alignment;
[0013] S6, based on the predicted time output by the decoder, the Hungarian algorithm is used to obtain the best correspondence between the predicted time and the real time by minimizing the total matching cost;
[0014] S7: Jointly optimize the highlight detection loss and the moment retrieval loss to update the encoder and decoder parameters, and use the updated encoder and decoder to output the highlight detection results and the moment prediction results.
[0015] To further improve the above technical solution, the S1 includes: the visual encoder is SlowFast and CLIP, the text encoder is CLIP; using the visual encoder E v By N v A video sequence consisting of Perform visual feature extraction and obtain the visual features as follows: Using the text encoder E t By N t The query text consists of Perform text feature extraction, and the obtained text features are
[0016] Further, the S2 includes:
[0017] Set video feature X v As the query Q v , the text feature X t As key K t Sum value V t , perform cross attention operation according to the following formula:
[0018]
[0019] Among them, d is the dimension of the query, key and value after projection, and the joint feature X is obtained through attention operation.
[0020] Furthermore, the encoder is stacked by 6 identical layers, each layer includes a multi-head self-attention mechanism sublayer and a fully connected feedforward network sublayer, and residual connections and layer normalization are used around the multi-head self-attention mechanism sublayer and the fully connected feedforward network sublayer;
[0021] The decoder is composed of 6 identical stacked layers, each of which includes a multi-head self-attention mechanism sublayer, a masked multi-head self-attention sublayer and a fully connected feedforward network sublayer, and residual connections and layer normalization are used around the multi-head self-attention mechanism sublayer, the masked multi-head self-attention sublayer and the fully connected feedforward network sublayer.
[0022] Furthermore, the multi-head self-attention mechanism sublayer is calculated using the following formula:
[0023] MultiHead(Q,K,V)=Concat(head1,head2,…,head h )W O
[0024] Among them, W O is a learnable weight, Concat(·) represents a concatenation operation, head i It is expressed as follows:
[0025] head i =Attention(QW i Q ,KW i K ,VW i V )
[0026] Among them, W i Q , W i K and W i V Represents the learnable weight matrix inside each self-attention unit;
[0027] The masked multi-head self-attention sublayer is calculated using the following formula:
[0028]
[0029] Where M is the mask matrix;
[0030] The fully connected feedforward network sublayer is calculated using the following formula:
[0031] FFN(x)=max(0,xW1+b1)W2+b2
[0032] Among them, x is the input feature, W1 is the weight matrix, mapping the input feature to the hidden layer, W2 is the weight matrix, mapping the hidden layer output to the final output, b1 is the bias vector, used to adjust the linear transformation result of xW1, b2 is the bias vector, used to adjust the linear transformation result of the hidden layer output max(0,xW1+b1);
[0033] The specific operation process of the residual connection and layer normalization is as follows:
[0034] each_sublayer_output=LayerNorm(x+Sublayer(x))
[0035] Among them, Sublayer(x) refers to the function implemented by the multi-head self-attention mechanism sublayer, the masked multi-head self-attention sublayer and the fully connected feedforward network sublayer.
[0036] Further, the S4 includes:
[0037] S401, divide feature F into positive sample feature F according to significance score threshold S0 + And negative sample features F - , the calculation process is as follows:
[0038]
[0039] Among them, S l is the true annotation significance score; S0 is the partition threshold;
[0040] S402, calculating the cosine similarity of the positive-positive feature pair and the positive-negative feature pair, the calculation process is as follows:
[0041]
[0042] Among them, cos(·) is the cosine similarity, and are the cosine similarities of positive-positive pairs and positive-negative pairs, respectively, N + and N - are the number of positive samples and the number of negative samples respectively;
[0043] Construct a similarity score set, and the calculation process is as follows:
[0044]
[0045] Among them, S +,+ and S +,- are sets of similarity scores of positive-positive and positive-negative sample pairs respectively;
[0046] S403, calculating the weight of each element in saliency contrast learning based on the similarity score set, the calculation process is as follows:
[0047]
[0048] in, and They are and In S +,+ and S +,- rank in the set, α is a hyperparameter that controls the smoothness of the exponential function;
[0049] S404, using similarity and weights for comparative learning, so that the features between positive samples are closer, and the features between positive samples and negative samples are farther apart, the calculation process is as follows:
[0050]
[0051] in, As an indicator function, it takes 1 when i≠j and 0 otherwise; and are the loss values between positive-positive and positive-negative, respectively.
[0052] Further, the S5 includes:
[0053] S501, calculate the center distance between the predicted time and the real time, the process is as follows:
[0054] Define the prediction time as t P =(p b ,p e ), defining the real time as T G =(g b ,g e ), center distance d t The calculation process is as follows:
[0055]
[0056] S502, calculate the length s of the minimum closure moment t :
[0057] s t =max(p e ,g e )-min(p b ,g b );
[0058] S503: Calculate the center distance loss for center distance regression based on the center distance and length
[0059]
[0060] Furthermore, the Hungarian algorithm obtains the best correspondence between the predicted time and the real time by minimizing the total matching cost, including:
[0061] Will Represented as a set of N prediction moments, Represented as a set of N real moments, the matching cost between the predicted moment and the real moment It is expressed as:
[0062]
[0063] in, It is expressed as (c i ,a i ), c i is the category label representing the foreground and background, a i ∈[0,1] 2 is a normalized vector used to define the center coordinates and width of the moment; a and They are the predicted moment and the real moment; Indicates background; As an indicator function, When established, take 1; Is in arrangement The best bipartite match between GT and prediction is It is the moment to retrieve the loss.
[0064] Furthermore, the moment retrieval loss Including cross entropy loss L1 loss for video moment regression and distance time intersection loss It is expressed as:
[0065]
[0066] Among them, λ CE and λ DTIoU is the balance weight, a and They are the predicted moment and the real moment;
[0067] The highlight detection loss Including significant loss and saliency contrast loss It is expressed as:
[0068]
[0069] Among them, λ SCis the weight of the significant contrast loss, It is expressed as:
[0070]
[0071] Among them, t low and t high A low-score segment and a high-score segment are randomly selected in the real moment; t out and t in is a segment outside and a segment inside the true annotation moment; Δ is a hyperparameter set to 0.2; S(·) represents the significance score value;
[0072] The total loss Including time retrieval loss and highlight detection loss It is expressed as:
[0073]
[0074] Among them, λ highlight is the weight of the highlight detection loss.
[0075] The system for implementing the above-mentioned video moment retrieval and highlight detection method based on dual semantic alignment includes:
[0076] A feature extraction module, configured to receive a video and a text query, and extract visual features from the video and text features from the text query using a pre-trained visual encoder and a pre-trained text encoder, respectively;
[0077] A cross-attention fusion module is used to perform cross-attention calculation based on visual features and text features, realize the fusion of visual features and text features, and generate joint features;
[0078] The feature encoding module, including an encoder and a decoder, is used to process joint features, where:
[0079] The encoder generates a segment-level saliency feature representation for highlight detection;
[0080] The decoder generates a predicted time representation for video time retrieval;
[0081] The fragment saliency comparison module is used to divide positive and negative samples based on the saliency score threshold, calculate the similarity of positive and negative sample features, and achieve fragment-level semantic alignment through saliency comparison learning;
[0082] The time prediction module is used to calculate the center distance and time intersection-to-union ratio based on the time prediction output of the decoder to achieve time-level semantic alignment;
[0083] The matching optimization module is used to perform one-to-one matching between the predicted time and the real time using the Hungarian algorithm, and optimize the matching results based on minimizing the matching cost;
[0084] The output module is used to output highlight detection results and moment retrieval results, wherein the highlight detection results represent the importance significance score of each video segment; and the moment retrieval results represent the start time and end time related to the text query.
[0085] Beneficial effects: Compared with the prior art, the advantages of the present invention are:
[0086] (1) In view of the problem that the existing methods are too focused on the situation where the predicted moment and the real annotated moment are completely disjoint and cannot achieve moment-level semantic alignment, the present invention introduces the center distance between the predicted moment and the real moment to reflect their positional relationship, so that the model can learn more accurate moment-level semantic information, thereby achieving more accurate moment-level semantic alignment. Compared with the existing methods, this model has better performance.
[0087] (2) Compared with the existing method that only randomly selects 4 special segments for penalty judgment in the saliency loss, the saliency contrast learning of the present invention takes into account all the segments in the video and divides all the segments into positive sample segments and negative sample segments, wherein the encoder output corresponding to the positive sample segment is used as the positive sample feature, and the encoder output corresponding to the negative sample segment is used as the negative sample feature. The idea of contrast learning is used to increase the distance between the positive sample feature and the negative sample feature, while reducing the distance within the positive sample feature. In this way, the model can learn the difference in features between segments with high and low saliency scores, so as to achieve segment-level semantic alignment in the feature space, and then accurately predict the saliency score. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] Figure 1 It is a schematic diagram of the overall process of the present invention.
[0089] Figure 2 : is a schematic diagram of the model network structure of the present invention. Where X is the joint feature obtained by cross-attention calculation, F is the feature further extracted by the encoder, and F + and F - They are respectively the positive sample features and negative sample features obtained by saliency contrast learning. FC is the fully connected layer, and FFN is the feedforward neural network. is a significant loss, is the saliency contrast loss, It is the moment to retrieve the loss.
[0090] Figure 3 It is a video moment retrieval and highlight detection effect diagram of the present invention. DETAILED DESCRIPTION
[0091] The technical solution of the present invention is described in detail below with reference to the accompanying drawings, but the protection scope of the present invention is not limited to the embodiments.
[0092] Example 1: Figure 1 The video moment retrieval and highlight detection method based on dual semantic alignment shown includes the following steps:
[0093] Step S1: Use the frozen pre-trained visual encoder E v and text encoder E t The visual features in the video and the text features in the text query are extracted respectively. The extracted visual features are: The extracted text features are:
[0094] Step S2: Visual features and text features Perform cross attention operation to obtain the joint features after interaction
[0095] Step S3, using the encoder and decoder to further extract joint features X to meet different tasks, where the encoder output is used for highlight detection, the decoder output is used for moment retrieval, and the encoder outputs the encoded feature F;
[0096] Step S4: Divide feature F into positive sample feature F by using saliency contrast learning partitioning strategy + And negative sample features F - Then, contrastive learning is used to bring feature segments with high saliency scores closer together and feature segments with large saliency scores farther apart, thereby achieving segment-level semantic alignment.
[0097] Step S5, by introducing the center distance d between the time t To perceive the position relationship at different times, use the distance-time intersection loss Close the distance between the predicted moment and the actual moment, thus achieving moment-level semantic alignment;
[0098] Step S6: using the Hungarian algorithm to perform one-to-one matching between the predicted time and the real time, optimizing the matching result in a manner of minimizing the matching cost, and obtaining the best corresponding relationship between the predicted time and the real time;
[0099] Step S7: Using cross entropy loss L1 loss Distance-time intersection loss Significant loss and saliency contrast loss Update the model parameters to train the model to the best performance. The trained model outputs highlight detection segments and video retrieval moments.
[0100] Specifically, step S1 uses the pre-trained visual encoder E v , for N v A video sequence consisting of The extracted visual features are Using the text encoder E t , for N t The query text consists of The extracted text features are
[0101] In step S2, the visual feature X v and text feature X t Perform cross-attention operation.
[0102] Set video feature X v As a query (Q v ), text feature X t As a key (K t ) and value (V t ). The calculation is as follows:
[0103]
[0104] Among them, d is the dimension of the query, key and value after projection, and the joint feature X is obtained through 6 layers of attention operation.
[0105] In step S3, the encoder is composed of 6 identical layers stacked together, each of which is composed of two sub-layers. The first sub-layer is a multi-head self-attention mechanism, which is used to calculate the degree of correlation between elements in the input sequence, assign attention weights, and output features containing long-distance dependencies; the second sub-layer is a fully connected feedforward network, which is used to perform nonlinear transformations on the input, restore the dimension, learn complex patterns, and output features of appropriate dimensions; residual connections and layer normalization are used around these two sub-layers to ensure information transmission and gradient flow, so that the model can be trained stably and converge faster. The encoder output features are The decoder is also stacked with 6 identical layers. Unlike the encoder, in addition to the multi-head self-attention mechanism and the fully connected feedforward network sub-layers of each layer, the decoder also inserts a third sub-layer called masked multi-head self-attention before the two sub-layers. This layer uses masks to prevent the decoder from seeing future information in the process of generating sequences. Similar to the encoder, the decoder uses residual connections and layer normalization around each sub-layer.
[0106] The multi-head self-attention mechanism is calculated as follows:
[0107] MultiHead(Q,K,V)=Concat(head1,head2,…,head h )W O (2)
[0108] Among them, W O is a learnable weight, Concat(·) represents a concatenation operation, head i It is expressed as follows:
[0109] head i =Attention(QWi i Q ,KW i K ,VW i V ) (3)
[0110] Among them, Wi i Q , W i K and W i V Represents the learnable weight matrix inside each self-attention unit.
[0111] By introducing the mask matrix M based on the self-attention mechanism, the masked self-attention matrix is obtained, which is calculated as follows:
[0112]
[0113] The fully connected feedforward network consists of two linear transformations, and a ReLU activation function is used between the two for nonlinear transformation. The calculation is as follows:
[0114]
[0115] The specific operation process of residual connection and layer normalization is as follows:
[0116] each_sublayer_output=LayerNorm(x+Sublayer(x)) (6)
[0117] Among them, Sublayer(x) refers to the function implemented by the sublayer (multi-head self-attention mechanism, masked self-attention mechanism, fully connected feedforward network).
[0118] In step S4, segment-level semantic alignment is achieved through saliency contrast learning.
[0119] According to the significance score threshold S0, feature F is divided into positive sample features F + And negative sample features F - , calculated as follows:
[0120]
[0121] Among them, S l is the true annotation significance score; S0 is the partition threshold, which is set to 2.
[0122] According to the positive and negative sample features, the cosine similarity of the positive-positive feature pair and the positive-negative feature pair is calculated as follows:
[0123]
[0124] Among them, cos(·) is the cosine similarity, and are the cosine similarities of positive-positive pairs and positive-negative pairs, respectively, N + and N - are the number of positive samples and the number of negative samples respectively.
[0125] Based on the similarity, two sets of similarity scores can be constructed for positive-positive and positive-negative sample pairs, expressed as follows:
[0126]
[0127] Among them, S +,+ and S +,- are sets of similarity scores for positive-positive and positive-negative sample pairs, respectively.
[0128] Based on the above two sets, the weight of each element in the set in saliency contrast learning is calculated as follows:
[0129]
[0130] in, and They are and In S +,+ and S +,- rank in the set, α is a hyperparameter that controls the smoothness of the exponential function and is set to 0.25.
[0131] Finally, the features of positive and negative samples are contrastively learned using similarity and weights to make the features of positive samples closer and the features of positive and negative samples farther apart. The loss function used in the contrastive learning process is calculated as follows:
[0132]
[0133] in, As an indicator function, it takes 1 when i≠j and 0 otherwise; and are the loss values between positive-positive and positive-negative, respectively.
[0134] In step S5, the center distance d between the introduction time t Achieve moment-level semantic alignment.
[0135] The prediction time is defined as T P =(p b ,p e ) and the real time is defined as T G =(g b ,g e ), center distance d t The calculation is as follows:
[0136]
[0137] The length of the minimum closure moment s t The calculation is as follows:
[0138]
[0139] Center distance loss for center distance regression The calculation is as follows:
[0140]
[0141] In step S6, the Hungarian algorithm is used to determine the best bipartite match between the actual result and the prediction, including the following steps: constructing an N×N cost matrix, where N is the number of moments; calculating the matching cost between each pair of predicted moments and the actual moments; using the Hungarian algorithm to find the perfect match with the minimum total cost; establishing a one-to-one correspondence between the predicted moments and the actual moments to ensure that each actual moment matches only one predicted moment.
[0142] Will Represented as a set of N predictions derived from the query at time, Represented as a set of real annotated moments, each predicted moment and real moment contains two key information: category label c i (indicates foreground or background), time vector a i ∈[0,1] 2 (Defines the normalized vector of the center coordinates and width at the moment).
[0143] Matching cost between predicted moment and real moment It is expressed as:
[0144]
[0145] in, It can be expressed as (c i ,a i), c i is the category label representing the foreground and background, a i ∈[0,1] 2 is a normalized vector used to define the center coordinates and width of the moment; a and They are the predicted moment and the real moment; Indicates background; As an indicator function, When established, take 1; Is in arrangement ( is the set of all permutations of N elements) and the best bipartite match between GT and prediction is It is the moment to retrieve the loss.
[0146] In step S7, the total loss is calculated for joint optimization. The total loss includes the time retrieval loss. and highlight detection loss Two parts, moment retrieval loss Through back propagation, the multi-head self-attention mechanism weight matrix and the fully connected feedforward network weight matrix in the encoder are mainly updated, and the high-light detection loss Through back-propagation, the multi-head self-attention mechanism weight matrix, masked multi-head self-attention weight matrix and fully connected feed-forward network weight matrix in the decoder are mainly updated.
[0147] Total loss It is expressed as follows:
[0148]
[0149] Among them, λ highlight is the weight of the highlight detection loss.
[0150] For moment retrieval loss, it mainly includes cross entropy loss for foreground and background classification And the L1 loss for video moment regression and DTIoU loss Therefore, the loss is retrieved at any moment It is expressed as follows:
[0151]
[0152] Among them, λ CE and λ DTIoU is the balance weight, a and They are the predicted moment and the actual moment.
[0153] For highlight detection loss, it mainly includes the significance loss used for coarse screening and segment-level saliency contrast loss for fine segment-by-segment filtering Therefore, the highlight detection loss It is expressed as follows:
[0154]
[0155] Among them, λ SC is the weight of the saliency contrast loss.
[0156] Significance loss for coarse screening It is expressed as follows:
[0157]
[0158] Among them, t low and t hig A low-score segment and a high-score segment are randomly selected in the real moment; t out and t in is a segment outside and a segment inside the true annotation moment; Δ is a hyperparameter set to 0.2; S(·) represents the significance score value.
[0159] Embodiment 2: The video moment retrieval and highlight detection system in this embodiment is as follows Figure 2 As shown, including:
[0160] A feature extraction module, configured to receive video and text queries, and extract visual features of the video and text features of the text through a pre-trained visual encoder and a pre-trained text encoder, respectively;
[0161] A cross-attention fusion module is used to perform cross-attention calculation based on visual features and text features, realize the fusion of visual features and text features, and generate joint features;
[0162] The feature encoding module, including an encoder and a decoder, is used to further process the joint features, where:
[0163] The encoder generates segment-level saliency feature representation for highlight detection;
[0164] The decoder generates a predicted moment representation for video moment retrieval;
[0165] The saliency comparison module is used to divide positive and negative samples based on the saliency score threshold, calculate the similarity of positive and negative sample features, and achieve segment-level semantic alignment through saliency comparison learning;
[0166] The center distance regression module is used to calculate the center distance and time intersection-to-union ratio based on the decoder's moment prediction output to achieve moment-level semantic alignment;
[0167] The matching optimization module is used to perform one-to-one matching between the predicted time and the real time using the Hungarian algorithm, and optimize the matching results based on minimizing the total loss;
[0168] The output module is used to output highlight detection results and moment retrieval results, where:
[0169] The highlight detection result indicates the importance and significance score of each video clip;
[0170] The time search results represent the start time and end time associated with the text query.
[0171] The feature extraction module and the feature encoding module are implemented with the help of existing pre-trained models, wherein the feature extraction module uses pre-trained SlowFast and CLIP as visual feature extractors, and CLIP as a text feature extractor.
[0172] The encoder is implemented by stacking 6 layers of the same structure, each layer contains: the first sublayer: a multi-head self-attention mechanism for calculating the correlation between sequence elements; the second sublayer: a fully connected feedforward network for nonlinear transformation; each sublayer uses residual connections and layer normalization to ensure information transfer and training stability; the encoder input includes the joint feature X obtained from the cross-attention module, and the encoder outputs the feature matrix F, which is used for two tasks: predicting the saliency score through the FC layer for highlight detection of people; dividing the features into positive sample features F+ and negative sample features F- for saliency contrast learning; the encoder output participates in the calculation and loss.
[0173] Similarly, the decoder uses a 6-layer stacked structure, each layer contains: the first sublayer: multi-head self-attention mechanism; the second sublayer: masked multi-head self-attention (to prevent seeing future information); the third sublayer: fully connected feedforward network; each sublayer uses residual connection and layer normalization. The input of the decoder includes the feature F output by the encoder and the learnable auxiliary vector a μ and a σ The output of the decoder is used to predict the time position of the video clip and calculate the center distance d from the true time (GT). t , generate the final moment prediction result, the decoder output participates in the calculation loss.
[0174] The segment-level saliency comparison module and the moment-level center distance regression module are proposed in the present invention to solve the problem of difficulty in aligning segment-level and moment-level dual semantics in video moment retrieval and highlight detection tasks.
[0175] Specifically, firstly, the visual encoder and text encoder in the feature extraction module are used to extract visual features and text features respectively; secondly, the cross-attention fusion module is used to perform cross-attention operations on visual features and text features to obtain joint features of visual and text interaction; then, the encoder and decoder are used to further extract joint features to meet the needs of different tasks; then, the segment-level saliency contrast module and the moment-level center distance regression module are used to perform segment-level alignment and moment-level feature representation respectively. Among them, the segment-level saliency contrast module divides the features into positive sample features and negative sample features based on the saliency contrast learning partitioning strategy, and performs contrastive learning in the feature space. By using contrastive learning, feature segments with high saliency scores are brought closer to each other, and feature segments with different saliency scores are appropriately distanced, so that the encoder can establish a better relationship between high-score segments and high-low-score segments, fully realizing segment-level feature alignment and effectively improving the accuracy of highlight detection; the moment-level center distance regression module introduces the center distance between moments to perceive the positional relationship between different moments, and uses the distance-time intersection loss to pull the distance between the prediction and the true label to achieve moment-level semantic alignment. Subsequently, the loss function is used to jointly optimize the model. Finally, the encoder output predicts the saliency score to represent the highlight detection result in the video, and the decoder output predicts the video moment retrieval result through classification and regression operations combined with a binary matching algorithm.
[0176] The present invention has the following effects on video moment retrieval and highlight detection: Figure 3 shown.
[0177] This implementation is experimentally verified on the QVHighlights, Charades-STA and TVSum datasets. Table 1 shows the experimental results on the QVHighlights dataset, which verifies that the present invention is significantly superior to other methods in both video moment retrieval tasks and video highlight detection tasks.
[0178] Table 1 Experimental results on the QVHighlights dataset
[0179]
[0180] Table 2 shows the experimental results on the Charades-STA dataset, which verifies that the present invention is superior to other methods in the video moment retrieval task.
[0181] Table 2 Experimental results on the Charades-STA dataset
[0182]
[0183] Table 3 shows the experimental results on the TVSum dataset, which verifies that the present invention is superior to other methods in the task of video highlight detection.
[0184] Table 3 Experimental results on TVSum dataset
[0185]
[0186] As described above, although the present invention has been shown and described with reference to specific preferred embodiments, it should not be construed as limiting the present invention itself. Various changes in form and details may be made without departing from the spirit and scope of the present invention as defined in the appended claims.
Claims
1. A video moment retrieval and highlight detection method based on dual semantic alignment, characterized in that: The following steps are involved: S1. Using pre-trained visual encoder E v and text encoder E t Extract visual features X from the video respectively v and the text feature X in the text query t ; S2, the visual feature X v and the text feature X t Perform cross attention operation to obtain the joint feature X; S3, extracting the joint feature X by using an encoder, outputting the encoded feature F and a segment-level saliency representation for highlight detection; processing the feature F output by the encoder by using a decoder, and outputting a predicted moment representation for video moment retrieval; S4, divide the feature F into positive sample features and negative sample features through the significance score threshold, and construct a significance contrast learning loss function to achieve segment-level semantic alignment; S5. Calculate the center distance between the predicted time output by the decoder and the real time to perceive the positional relationship between different times, and use the distance time intersection loss to achieve time-level semantic alignment; S6, based on the predicted time output by the decoder, the Hungarian algorithm is used to obtain the best correspondence between the predicted time and the real time by minimizing the total matching cost; S7: Jointly optimize the highlight detection loss and the moment retrieval loss to update the parameters of the encoder and decoder, and use the updated encoder and decoder to output the highlight detection results and the moment prediction results.
2. The video moment retrieval and highlight detection method based on dual semantic alignment according to claim 1 is characterized in that: The visual encoder is SlowFast and CLIP, and the text encoder is CLIP; using the visual encoder E v By N v A video sequence consisting of Perform visual feature extraction and obtain the visual features as follows: Among them, L v is the dimension of visual features; using the text encoder E t By N t The query text consists of Perform text feature extraction, and the obtained text features are Among them, L t is the dimension of text features.
3. The video moment retrieval and highlight detection method based on dual semantic alignment according to claim 2 is characterized in that: The S2 includes: Set video feature X v As the query Q v , the text feature X t As key K t Sum value V t , perform cross attention operation according to the following formula: Among them, d is the dimension of the query, key and value after projection, and the joint feature X is obtained through attention operation.
4. The video moment retrieval and highlight detection method based on dual semantic alignment according to claim 3 is characterized in that: The encoder is composed of 6 identical layers stacked together, each of which includes a multi-head self-attention mechanism sublayer and a fully connected feedforward network sublayer, and residual connections and layer normalization are used around the multi-head self-attention mechanism sublayer and the fully connected feedforward network sublayer; The decoder is composed of 6 identical stacked layers, each of which includes a multi-head self-attention mechanism sublayer, a masked multi-head self-attention sublayer and a fully connected feedforward network sublayer, and residual connections and layer normalization are used around the multi-head self-attention mechanism sublayer, the masked multi-head self-attention sublayer and the fully connected feedforward network sublayer.
5. The video moment retrieval and highlight detection method based on dual semantic alignment according to claim 4 is characterized in that: The multi-head self-attention mechanism sublayer is calculated using the following formula: MultiHead(Q,K,V)=Concat(head1,head2,…,head h )W O Among them, W O is a learnable weight, Concat(·) represents a concatenation operation, head i It is expressed as follows: in, and Represents the learnable weight matrix inside each self-attention unit; The masked multi-head self-attention sublayer is calculated using the following formula: Where M is the mask matrix; The fully connected feedforward network sublayer is calculated using the following formula: FFN(x)=max(0,xW1+b1)W2+b2 Among them, x is the input feature, W1 is the weight matrix, mapping the input feature to the hidden layer, W2 is the weight matrix, mapping the hidden layer output to the final output, b1 is the bias vector, used to adjust the linear transformation result of xW1, b2 is the bias vector, used to adjust the linear transformation result of the hidden layer output max(0,xW1+b1); The specific operation process of the residual connection and layer normalization is as follows: each_sublayer_outout=LayerNorm(x+Sublayer(x)) Among them, Sublayer(x) refers to the function implemented by the multi-head self-attention mechanism sublayer, the masked multi-head self-attention sublayer and the fully connected feedforward network sublayer.
6. The video moment retrieval and highlight detection method based on dual semantic alignment according to claim 5 is characterized in that: The S4 includes: S401, divide feature F into positive sample feature F according to significance score threshold S0 + And negative sample features F - , the calculation process is as follows: Among them, S l is the true annotation significance score; S0 is the partition threshold; S402, calculating the cosine similarity of the positive-positive feature pair and the positive-negative feature pair, the calculation process is as follows: Among them, cos(·) is the cosine similarity, and are the cosine similarities of positive-positive pairs and positive-negative pairs, respectively, N + and N - are the number of positive samples and the number of negative samples respectively; Construct a similarity score set, and the calculation process is as follows: Among them, S +,+ and S +,- are sets of similarity scores of positive-positive and positive-negative sample pairs respectively; S403, calculating the weight of each element in saliency contrast learning based on the similarity score set, the calculation process is as follows: in, and They are and In S +,+ and S +,- rank in the set, α is a hyperparameter that controls the smoothness of the exponential function; S404, using similarity and weights for comparative learning, so that the features between positive samples are closer, and the features between positive samples and negative samples are farther apart, the calculation process is as follows: in, As an indicator function, it takes 1 when i≠j and 0 otherwise; and are the loss values between positive-positive and positive-negative, respectively.
7. The video moment retrieval and highlight detection method based on dual semantic alignment according to claim 5 is characterized in that: The S5 includes: S501, calculate the center distance between the predicted time and the real time, the process is as follows: The prediction time is defined as T P =(p b ,p e ), defining the real time as T G =(g b ,g e ), center distance d t The calculation process is as follows: S502, calculate the length s of the minimum closure moment t : s t =max(p e ,g e )-min(p b ,g b ); S503: Calculate the center distance loss for center distance regression based on the center distance and length 8. The video moment retrieval and highlight detection method based on dual semantic alignment according to claim 6, characterized in that: The Hungarian algorithm obtains the best correspondence between the predicted time and the real time by minimizing the total matching cost, including: Will Represented as a set of N prediction moments, Represented as a set of N real moments, the matching cost between the predicted moment and the real moment It is expressed as: in, It is expressed as (c i ,a i ), c i is the category label representing the foreground and background, a i ∈[0,1] 2 is a normalized vector used to define the center coordinates and width of the moment; a and They are the predicted moment and the real moment; Indicates background; As an indicator function, When established, take 1; Is in arrangement The best bipartite match between GT and prediction is It is the moment to retrieve the loss.
9. The video moment retrieval and highlight detection method based on dual semantic alignment according to claim 5, characterized in that: The moment to retrieve the loss Including cross entropy loss L1 loss for video moment regression and distance time intersection loss It is expressed as: Among them, λ CE and λ DTIoU is the balance weight, a and They are the predicted moment and the real moment; The highlight detection loss Including significant loss and saliency contrast loss It is expressed as: Among them, λ SC is the weight of the significant contrast loss, It is expressed as: Among them, t low and t hig A low-score segment and a high-score segment are randomly selected in the real moment; t out and t in is a segment outside the true annotation moment and a segment inside the true annotation moment; Δ is a hyperparameter set to 0.2; S(·) represents the significance score value; The total loss Including time retrieval loss and highlight detection loss It is expressed as: Among them, λ highlight is the weight of the highlight detection loss.
10. A system for implementing the video moment retrieval and highlight detection method based on dual semantic alignment as claimed in claim 1, characterized in that: include: A feature extraction module, configured to receive a video and a text query, and extract visual features from the video and text features from the text query using a pre-trained visual encoder and a pre-trained text encoder, respectively; A cross-attention fusion module is used to perform cross-attention calculation based on visual features and text features, realize the fusion of visual features and text features, and generate joint features; The feature encoding module, including an encoder and a decoder, is used to process joint features, where: The encoder generates a segment-level saliency feature representation for highlight detection; The decoder generates a predicted time representation for video time retrieval; The fragment saliency comparison module is used to divide positive and negative samples based on the saliency score threshold, calculate the similarity of positive and negative sample features, and achieve fragment-level semantic alignment through saliency comparison learning; The time prediction module is used to calculate the center distance and time intersection-to-union ratio based on the time prediction output of the decoder to achieve time-level semantic alignment; The matching optimization module is used to perform one-to-one matching between the predicted time and the real time using the Hungarian algorithm, and optimize the matching results based on minimizing the matching cost; The output module is used to output highlight detection results and moment retrieval results, wherein the highlight detection results represent the importance significance score of each video segment; and the moment retrieval results represent the start time and end time related to the text query.
Citation Information
Patent Citations
Joint time retrieval and highlight detection method and system based on multitask reciprocity
CN117648463A
Joint time retrieval and highlight detection method and system based on multi-scale difference
CN117668293A
Cited By
Video clip retrieval method, video clip retrieval device and storage medium
CN120316307A
Automatic construction method and equipment for fragment-level alignment data and readable storage medium
CN120743282A
Method for time retrieval and highlight detection and related device
CN120783171A