Weakly supervised cross-modal video localization based on modal feature alignment
By optimizing candidate segment training through a global feature extraction module and positive-negative sample comparison learning, the problems of insufficient modal feature alignment and interaction in cross-modal video localization are solved, thereby improving computational efficiency and localization accuracy.
Patent Information
- Application Number
- CN202310888432.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-19
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2043-07-19
AI Technical Summary
In existing cross-modal video localization methods, the initial spatial distance between video features and text features is too large, making alignment difficult. This results in insufficient interaction between modal features and fails to fully utilize the differences between videos during training, leading to problems such as low computational efficiency and the impact of annotation errors.
A global feature extraction module was designed for modal feature alignment. During training, positive and negative samples were compared and learned. Gaussian masks and intersection-union ratios were used to optimize the training weights of candidate segments and construct negative video samples to enhance modal feature alignment and interaction.
It effectively reduces the spatial distance of modal features, improves computational efficiency and positioning accuracy, reduces the impact on annotation errors, and enhances the effect of cross-modal video positioning.
Smart Images

Figure CN116935274B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of video positioning, and particularly relates to a modal feature alignment model for weakly supervised cross-modal video positioning. The application aligns the two modal features before fusing the video features and the text features, solves the problem of too far initial spatial distance of the modal features, can strengthen the interaction of the two modalities, and improves the positioning result. The application is in the form of end-to-end, and has good effect on cross-modal video positioning. BACKGROUND
[0002] With the development of Internet technology and the increase in the possession rate of electronic devices such as mobile phones and computers, a large amount of information is uploaded to the Internet every day. People communicate with each other on the Internet through audio, picture, video and other information carriers. How to process these data and extract the content needed by people has become a difficult problem that needs to be solved by many researchers. Compared with picture, audio and other information carriers, video has rich information content because it can be understood as a collection of pictures, audio and text, and the related research results can be used in many fields such as monitoring security and unmanned driving, so in recent years, research on video has attracted a large number of researchers.
[0003] Research on video includes video behavior detection and video question answering. Video behavior detection only uses simple action labels for training, but actions in the real world are very complex, many actions are continuous and cannot be simply explained by a word, so the annotation of actions may be endless, which leads to great limitations of video behavior detection. If natural language is used for annotation, this problem can be effectively solved, which leads to the task of cross-modal video positioning.
[0004] The definition of the task of cross-modal video positioning is to find the start time and end time of the related segment in the video given a text description. Cross-modal video positioning has many practical scenarios, such as video editing, inputting the approximate description of the desired video segment to get the approximate editing segment time. It can also be used in the field of monitoring search to search for important segments in many long monitoring videos and help the police collect evidence. This task not only requires learning text features but also requires learning video features, and how to make the two modalities interact. If this task can be solved well, it will bring good benefits in the fields of monitoring security and video editing.
[0005] Since 2012, deep learning has developed rapidly, and current cross-modal video positioning is based on deep learning. There are four types of methods for current cross-modal video positioning: two-stage based, end-to-end based, reinforcement learning based and weakly supervised learning based.
[0006] The two-stage based learning method is further divided into sliding window based and candidate frame generation based. The sliding window based method is to generate a series of candidate segment feature files of different lengths in advance using a sliding window. However, this method has the disadvantages of occupying a large amount of storage space and causing a large amount of redundant calculation, and the efficiency is very low. In addition, each candidate segment is too independent, and it is difficult to exchange information between them, so the performance is also poor. The end-to-end based method also includes two types, one based on anchor frame and the other based on anchor-free frame. The anchor frame based method still needs to generate a large number of candidate segments, which requires a large amount of calculation and is low in efficiency. The anchor-free frame based method can directly predict the start time and end time of the relevant segment, which can reduce a large amount of calculation. The weakly supervised learning method is mainly divided into two categories: multi-instance learning based and reconstruction based. The multi-instance learning based method sets all sentences of the video pair as positive packages and all sentences of other videos as negative packages, which may occupy a large amount of video memory. In addition, for the data set with unbalanced number of annotations, the batch size needs to be set according to the maximum number of sentences, which reduces the GPU utilization. Because the positive and negative packages need to be trained, a large amount of training time is required. The reconstruction based method masks the text and then uses the candidate segment to predict the masked words to reconstruct the sentence.
[0007] Now more and more researches are based on weak supervision learning. Because the strong supervision based cross-modal video positioning needs a large amount of time annotation, the start time and end time of each text description need to be annotated, which consumes a lot of manpower and financial resources. However, the weak supervision based cross-modal video positioning only needs the sentence annotation of the video, without the time annotation, which greatly reduces the manpower and financial resources. Because the annotation often has subjective factors, different annotators will have some errors in annotation due to subjective factors, which will affect the training. Therefore, the weak supervision method can also avoid the influence caused by annotation errors.
[0008] Whether the weak supervision based cross-modal video positioning is based on multi-instance learning or reconstruction, most of them still use the candidate segments obtained by the sliding window to calculate the matching score of each candidate segment, which still requires a large amount of calculation. The latest research proposes to use Gaussian mask to directly obtain a small number of candidate regions, which effectively improves the calculation speed and prediction result.
[0009] The weak supervision based cross-modal video positioning based on Gaussian mask still has some disadvantages:
[0010] (1)Current video field task data set is extracted in advance feature, so the initial space distance of video feature and text feature is far, which is not conducive to the alignment between two kinds of modal features.
[0011] (2)Only the video internal negative sample mining strategy is used, and the difference between different videos is not utilized
[0012] (3)In the training, the latest research only considers the candidate segment with the minimum reconstruction sentence loss, but the insufficient early training may deviate the network model.
[0013] The present application is from the shortcomings of the latest research, provides a new idea for the existing method, and finally forms a complete cross-modal video positioning system. SUMMARY
[0014] The present application proposes a weakly supervised cross-modal video positioning method based on modal feature alignment, which has solved the above three shortcomings. 1. A global feature extraction module is designed, which can extract the global features of the input video and text into the network through the global feature extraction module each time a video text pair is input, and the spatial distance between the two modalities is reduced through training, thereby solving the problem of difficult alignment of two modal features.
[0015] 2. In the same training batch, the paired video text pair is a positive sample pair, and the unpaired video text pair is a negative sample pair, and the global feature matrix obtained can be used to simply construct a video negative sample, thereby training the difference between the text and other videos.
[0016] 3. After obtaining the candidate segment, the loss of the reconstructed sentence obtained by the candidate segment is used, which is different from the latest research which only uses the candidate segment with the minimum reconstruction loss for training. The present application calculates the Intersection over Union (IoU) of other candidate segments and the segment, and according to the intersection over union, the training weight of other candidate segments is allocated. Because if a candidate segment is close to the segment, the candidate segment should also be able to well reconstruct the sentence, so the candidate segment with a larger intersection over union will have a larger training weight.
[0017] The description of cross-modal video positioning is that, given a video V and a text T, the most matched segment in the video according to the semantic information of the text is found.
[0018] The technical solution adopted by the present application to solve its technical problems comprises the following steps:
[0019] Step (1), data preprocessing, extracting the initial modal features of video and text;
[0020] Step (2), constructing the overall architecture of the network and designing the loss function;
[0021] Step (3), model training, optimizing network parameters;
[0022] Step (4), generating the positioning detection result according to the trained network model.
[0023] Further, the step (1) is specifically implemented as follows:
[0024] Regarding video feature extraction, the video is first converted into a set of picture sequences with time sequence, and then a pre-trained C3D network is used to extract video features. A 4096-dimensional vector is extracted every 16 frames, and after the processing is completed, a vector sequence, i.e. video features, is finally obtained.
[0025] Regarding text feature extraction, a Glove word vector data file is used to obtain the word vector corresponding to each word, and the finally obtained word vector sequence is the text feature.
[0026] Further, the step (2) is specifically implemented as follows:
[0027] The network model mainly includes three parts: a feature alignment module, a candidate segment generation module and a text reconstruction module; the feature alignment module is used to extract deep features of two modalities and corresponding global modal features, and the global modal features are used for feature alignment and positive and negative video text pair learning; the candidate segment generation module is used to fuse the deep features of two modalities to generate a fixed number of Gaussian distributions, so as to obtain corresponding candidate segments and Gaussian mask attention; and the text reconstruction module uses the candidate segments to reconstruct the masked text, and obtains the optimal candidate segment through the reconstruction loss.
[0028] Further, the feature alignment module is specifically implemented as follows:
[0029] In the data preprocessing stage, the initial features of two modalities have been obtained, the video initial feature is N is the number of extracted video frames, D V is the video initial feature dimension, and the text initial feature is M is the number of words, D T is the text initial feature dimension.
[0030] Before being input into the feature alignment module, three learnable tokens need to be added to the initial features. In order to obtain the global feature of the fused feature in the subsequent candidate segment generation module, a learnable Gaussian token v gauss is added at the end of the video feature. Because the feature alignment module needs to obtain the global features of two modalities, a learnable classification token vcls , t cls . Finally, the video feature V = {v1, v2,..., v N , v gauss , v cls} and the text feature T = {t1, t2,..., t M , t cls} are obtained.
[0031] The two modal features after adding the wordpiece are input into the corresponding modal self-attention structure to extract deep modal features, i.e., modal deep features. Since the self-attention structure is not sensitive to the time sequence of the features, the position information corresponding to each word is obtained by using the sine and cosine functions, and the position encoding is represented by the following formula:
[0032] PE (pos,2a) = sin(pos / 10000 2a / d model )
[0033] PE (pos,2a+1) = cos(pos / 10000 2a / d model )
[0034] where a represents the a-th dimension of the position encoding vector, pos represents the position of the current word in the text, d model represents the dimension of the word vector.
[0035] After inputting the video feature into the video modal self-attention structure, the deep video feature is obtained, and after adding the text feature and the position encoding, the deep text feature is obtained by inputting into the text modal self-attention structure. The self-attention structure is represented by the following formula:
[0036]
[0037] where X is the input modal feature, is the dimension of X.
[0038] At this time, the classification wordpiece has obtained the global modal feature. After mapping the classification wordpiece to a low dimension and normalizing, the similarity of the matching video text pair and the similarity of the non-matching video text pair s = g v (v cls ) T g t (t cls ), g v , g t represent the mapping and normalization operations on the video and the text. In each training batch, the similarity between each video and each text can be calculated, and the similarity matrix s(V, T) = gv (v cls ) T g t (t cls ) and s(T, V) = g t (t cls ) T g v (v cls ), respectively, represent the similarity matrix of video-text pairs and the similarity matrix of text-video pairs. The similarity matrix is further optimized by softmax normalization:
[0039]
[0040]
[0041] where B represents the number of samples of video-text pairs in a training batch.
[0042] The contrastive learning loss of video-text is:
[0043]
[0044] where CE is the cross entropy loss, y v2t , y t2v is the true label of the image-text pair, and if the image and text do not match, the true label is 0, and if they match, the true label is 1.
[0045] Before inputting the extracted features into the candidate segment generation module, the global features are deleted from the feature sequence, and the video features obtained are and the text features are
[0046] Further, the candidate segment generation module is implemented as follows:
[0047] The essence of the candidate segment generation module is a decoder structure of a transformer. After inputting the video features and the text features into the module, the fusion features that fuse visual information and text information are obtained.
[0048]
[0049] where D represents the decoder operation. At this time all feature information is integrated, and the center and the width of the candidate segment are predicted by a fully connected layer and a Sigmoid activation function. K is the number of candidate segments. After obtaining the center and width of the candidate segment, the Gaussian mask m is calculated p ∈R K×N :
[0050]
[0051] where c is the video feature position, k is the index of the candidate segment, represents the Gaussian mask value of the frame number c in the positive candidate segment with index k. The Gaussian mask is the Gaussian distribution value representing each position of the video, and σ is a hyperparameter for controlling the width of the Gaussian curve.
[0052] In order to make the candidate segments as little overlap as possible to improve the recall rate of the positioning result, a diversity loss is used to optimize this module, which is expressed as follows:
[0053]
[0054] where, represents the Frobenius norm of the matrix, λ is a hyperparameter for controlling the overlap of the candidate segments, and I represents the unit matrix.
[0055] Further, the text reconstruction module is implemented as follows:
[0056] The text reconstruction module is also a transformer decoder structure. After obtaining the candidate segments, the video features corresponding to the candidate segments are extracted through the video modality self-attention structure of the feature alignment module, and then the deep features of the candidate segments are input into the text reconstruction module.
[0057] Then the original text is randomly masked, and the number of masks is 1 / 3 of the number of text words. After text feature extraction of the masked text, deep feature extraction is performed through the text modality self-attention structure of the feature alignment module, and the deep features of the candidate segments are input into the text reconstruction module together. But different from the candidate box generation module, when the two modalities interact, the attention distribution of the two modalities also needs to be multiplied by the Gaussian mask, so that the text pays more attention to the video frames related to the positive candidate segment and suppresses the irrelevant video frame information. Finally, the fusion feature that integrates text information and visual information is obtained The formula is expressed as:
[0058]
[0059] where, represents the deep feature of the masked text, is the deep feature of the candidate segment.
[0060] After the fusion features are input into the full connection layer, the probability of the masked word mapping to all words can be obtained by softmax calculation Q is the total number of words in the text, and the process is expressed by the formula:
[0061]
[0062] where W is the weight of the full connection layer, b' is the bias of the full connection layer, d ∈ M, The dth fusion feature.
[0063] Each mask position takes the word with the maximum corresponding probability as the prediction result. The reconstruction loss L is calculated according to the prediction result ce , the negative logarithmic probability of each masked word is calculated, and they are added:
[0064]
[0065] where L ce represents the reconstruction loss calculated based on the video candidate segment and p(t d+1 |t 1:d ) (the probability of predicting the d+1th word according to the 1st word to the dth word).
[0066] Further, the complete loss function is as follows:
[0067] First, the optimal candidate segment with the minimum reconstruction loss is obtained:
[0068]
[0069] where k * is the index of the optimal candidate segment, represents the loss of the positive candidate segment indexed by k. The optimal candidate segment is used as a pseudo label, and then the intersection over union of each positive candidate segment and the optimal candidate segment is calculated:
[0070]
[0071] where, is the optimal candidate segment indexed by k * , m k is the positive candidate segment indexed by k. and represent the end time and start time of the optimal candidate segment indexed by k * , e k and s k represent the end time and start time of the positive candidate segment indexed by k.
[0072] The intersection-over-union is normalized to obtain the pseudo-label value of each candidate segment:
[0073]
[0074] wherein o min and o max are hyperparameters, respectively representing the minimum value and the maximum value of the intersection-over-union.
[0075] According to the pseudo-label value and the reconstruction loss of other positive candidate segments, the positive candidate segment learning loss function is calculated as:
[0076]
[0077] wherein, is the final total learning loss of the positive candidate segment.
[0078] In order to learn more fine-grained information, negative candidate segments also need to be selected for contrastive learning. Since the entire video contains the semantic information of the text, the entire video is taken as a positive candidate segment, but it contains a large amount of redundant information, so the entire video is taken as a global candidate segment. Since the two segments at the ends of the positive candidate segment are quite different from the positive sample, the two segments are taken as negative candidate segments. Now there are three types of video segments, the positive candidate segment m p generated by the candidate segment generation module, the global candidate segment m r of the entire video, and the negative candidate segments m n at the two ends of the positive candidate segment. The loss of reconstructing the text from the global candidate segment is between the positive and negative:
[0079]
[0080] The negative candidate segments are selected from the two ends of the optimal candidate segment, the starting time of the left end negative candidate segment is fixed as the video starting time, and the ending time of the right end negative candidate segment is fixed as the video ending time. With the increase of the number of training times, the width of the negative candidate segments at the two ends will become larger, and the negative candidate segments will gradually approach the optimal candidate segment. Under the initial condition, the width of the negative candidate segments at the left and right ends is:
[0081]
[0082] wherein w 1 is the width of the left end negative candidate segment, and w 2 is the width of the right end negative candidate segment.
[0083] The growth rate of the width is:
[0084]
[0085] wherein, represents the width of the current left end negative candidate segment, represents the width of the current right end negative candidate segment.
[0086] where η will become larger and larger with the number of training times:
[0087]
[0088] where e max is the total number of training times, e is the current training time, and the center of the negative sample is:
[0089]
[0090] After obtaining the width and center of the negative candidate frame, the Gaussian mask of the negative candidate frame is calculated. The Gaussian mask of the entire video is 1.
[0091] After calculating the reconstruction loss of the positive candidate segment the reconstruction loss of the global candidate segment and the reconstruction loss of the left and right negative candidate segments a more complete loss function is designed. Because the entire video also contains positive candidate segment information, the entire video can also be regarded as a rough positive candidate segment. The positive sample learning loss function is designed as follows:
[0092]
[0093] In order to make the network learn the difference between the positive candidate segment and the negative candidate segment, the positive and negative sample contrast learning loss function is designed as follows:
[0094]
[0095] where β1, β2 are hyperparameters for controlling the difference of sample reconstruction loss.
[0096] The final complete total learning loss is:
[0097] L = L rec + α1L IVC + α2L div + α3L con
[0098] where α1, α2, α3 are hyperparameters for balancing the total learning loss.
[0099] Further, the specific implementation of step (3) model training is as follows:
[0100] According to the designed loss function, in the training process, the model parameters are updated through the back propagation algorithm (Back-Propagation, BP), until the model converges and the model file is saved.
[0101] Further, step (4) is specifically implemented as follows:
[0102] After the network loads the saved model file, the test set is tested, one text and to-be-positioned video are input each time, finally K candidate clips and corresponding reconstruction losses are obtained, and the optimal candidate clip with the minimum reconstruction loss is selected as the positioning result.
[0103]
[0104]
[0105] Wherein, Duration is the duration of the video.
[0106] The present application has the following advantages:
[0107] A new weakly supervised cross-video positioning method is proposed, and a network for aligning features before cross-modal interaction is designed.
[0108] We introduce a feature alignment module before cross-modal interaction, which can effectively reduce the spatial distance of the two modalities, and also learn the difference between non-matching samples. And when learning the positive candidate clip, the optimal positive candidate clip is used as a pseudo label, so that the positive candidate clip with a large intersection over union with the optimal positive candidate clip can also participate in training, increasing the interactivity between positive candidate clips.
[0109] A large number of experimental results on two benchmark datasets prove the effectiveness of the method. Our cross-modal video positioning method achieves very effective results. BRIEF DESCRIPTION OF DRAWINGS
[0110] Figure 1 The present application is based on a weakly supervised cross-modal video positioning network model based on modal feature alignment;
[0111] Figure 2 The present application is a whole flow chart for realizing the cross-modal video positioning task. DETAILED DESCRIPTION
[0112] The present application method and its detailed parameters are further specifically described below in combination with the drawings and examples.
[0113] For the description of cross-modal video positioning, given a video V and a text T, find the most matching clip in the video according to the semantic information of the text.
[0114] As shown in Figure 1 and 2 , a cross-modal video positioning method based on feature alignment, the specific steps are as follows:
[0115] Step (1), data preprocessing, extracting initial modal features of video and text;
[0116] The video dataset includes ActivityNet Caption and Charades-STA, wherein the training set of the ActivityNet Caption dataset contains 37421 video-text pairs, and the test set contains 17031 video-text pairs. The training set of the Charades-STA contains 13898 video-text pairs, and the test set contains 4233 video-text pairs.
[0117] Regarding video feature extraction, the video is first converted into a sequence of pictures with time sequence, and then a pre-trained C3D network is used to extract video features. A 4096-dimensional vector is extracted every 16 frames, and after processing, a final vector sequence, i.e., video features, is obtained.
[0118] Regarding text feature extraction, a Glove word vector data file open sourced by Stanford University is used. The file is essentially a one-to-one mapping of words and word vectors, and the corresponding word vector of each word can be directly obtained through the file. The final word vector sequence is the text feature.
[0119] Step (2), constructing the overall architecture of the network and designing the loss function;
[0120] The network model mainly includes three parts: a feature alignment module, a candidate segment generation module, and a text reconstruction module. The feature alignment module is used to extract deep features of two modalities and corresponding global modal features, which are used for feature alignment and positive and negative video-text pair learning. The candidate segment generation module is used to fuse the deep features of the two modalities to generate a fixed number of Gaussian distributions, so as to obtain the corresponding candidate segments and Gaussian mask attention. The text reconstruction module uses the candidate segments to reconstruct the masked text, and the optimal candidate segment is obtained through the reconstruction loss.
[0121] The feature alignment module is specifically implemented as follows:
[0122] In the data preprocessing stage, the initial features of the two modalities are obtained, the video initial feature is N is the number of extracted video frames, D V is the dimension of the video initial feature, and the text initial feature is M is the number of words, D T is the dimension of the text initial feature.
[0123] Before inputting into the feature alignment module, three learnable tokens are added to the initial feature. In order to get the global feature of the fused feature in the candidate segment generation module, a learnable Gaussian token v gauss is added at the end of the video feature. Because the feature alignment module needs to get the global feature of the two modalities, a learnable classification token v cls is added after the video feature and the text feature respectively. cls . Finally, the video feature V = {v1, v2,..., v N , v gauss , v cls} and the text feature T = {t1, t2,..., t M , t cls} are obtained.
[0124] The two modalities of features after adding tokens are input into the corresponding modal self-attention structure to extract deep modal features, i.e. modality deep features. Since the self-attention structure is not sensitive to the time sequence of the features, the sine and cosine functions are used to obtain the position information corresponding to each word. The position encoding is represented by the following formula:
[0125] PE (pos,2a) = sin(pos / 10000 2a / d model )
[0126] PE (pos,2a+1) = Cos(pos / 10000 2a / d model )
[0127] where a represents the a-th dimension of the position encoding vector, pos represents the position of the current word in the text, d model represents the dimension of the word vector.
[0128] After inputting the video feature into the video modality self-attention structure, the deep video feature is obtained. After adding the text feature and the position encoding, the deep text feature is obtained after inputting into the text modality self-attention structure. The self-attention structure is represented by the following formula:
[0129]
[0130] where X is the input modality feature, is the dimension of X.
[0131] At this time, the classification token has obtained the global modality feature. After mapping the classification token to a low dimension and normalizing, the similarity of the matching video text pair and the similarity of the non-matching video text pair s = g v (vcls T g t (t cls ), g v , g t denote the mapping and normalization operation on video and text. In each training batch, the similarity between each video and each text can be calculated, and the similarity matrix s(V, T) = g v (v cls ) T g t (t cls ) and s(T, V) = g t (t cls ) T g v (v cls ), respectively, denote the video pair to text similarity matrix and the text to video similarity matrix. The pair similarity matrix is further optimized by softmax normalization:
[0132]
[0133]
[0134] where B denotes the number of video-text pairs in a training batch.
[0135] The contrastive learning loss of video-text is:
[0136]
[0137] where CE is the cross entropy loss, y v2t , y t2v is the true label of the image-text pair, and if the image and text do not match, the true label is 0, and if they match, the true label is 1.
[0138] The role of the feature alignment module is to solve the shortcomings (1) and (2) described above. Using the similarity matrix can directly learn the difference between positive and negative samples and align the modal features.
[0139] Before inputting the extracted features into the candidate segment generation module, the global features are deleted from the feature sequence, and the obtained video features are and the text features are
[0140] The specific implementation of the candidate segment generation module is as follows:
[0141] The essence of the candidate segment generation module is a transformer decoder structure, which takes the video features and the text features After inputting into the module, the fusion feature that integrates visual information and text information will be obtained D H D is the dimension of the fused feature, and the formula is expressed as:
[0142]
[0143] where D represents the decoder operation. At this time After all the feature information is converged, the center of the candidate segment is predicted through a fully connected layer and a Sigmoid activation function and width K is the number of candidate segments. After obtaining the center and width of the candidate segment, the Gaussian mask m is calculated p ∈R K×N :
[0144]
[0145] where c is the video feature position, k is the index of the candidate segment, represents the Gaussian mask value of the frame number c in the positive candidate segment with index k. The Gaussian mask is the Gaussian distribution value representing each position of the video, and sigma is a hyperparameter used to control the width of the Gaussian curve.
[0146] In order to make the candidate segments as little overlapping as possible to improve the recall rate of the positioning result, a diversity loss is used to optimize the module, and the diversity loss is expressed as follows:
[0147]
[0148] where represents the Frobenius norm of the matrix, and lambda is a hyperparameter used to control the overlap degree of the candidate segments, and I represents the unit matrix.
[0149] The text reconstruction module is implemented as follows:
[0150] The text reconstruction module is also a transformer decoder structure. After obtaining the candidate segment, the video feature corresponding to the candidate segment is extracted through the video modality self-attention structure of the feature alignment module, and then the deep feature of the candidate segment is input into the text reconstruction module.
[0151] Then the original text is randomly masked, and the number of masks is 1 / 3 of the number of text words. After text feature extraction of the masked text, deep feature extraction is performed through the text modal self-attention structure of the feature alignment module, and the deep features of the candidate segment are input into the text reconstruction module together. But unlike the candidate box generation module, when the two modalities interact, the attention distribution of the two modalities also needs to be multiplied by a Gaussian mask. The purpose is to make the text pay more attention to the video frames related to the positive candidate segment and suppress irrelevant video frame information. Finally, the fusion feature that integrates text information and visual information is obtained The formula is expressed as:
[0152]
[0153] Wherein, represents the deep feature of the masked text, is the deep feature of the candidate segment.
[0154] After the fusion feature is input into the full connection layer, the probability of the masked word mapping to all words can be obtained by softmax calculation Q is the total number of words in the text. This process is expressed by the formula as:
[0155]
[0156] Wherein, W is the weight of the full connection layer, b' is the bias of the full connection layer, and d is in M, The dth fusion feature.
[0157] The word with the maximum corresponding probability is taken as the prediction result at each mask position. The reconstruction loss L is calculated according to the prediction result ce , the negative logarithmic probability of each masked word is calculated, and they are added together:
[0158]
[0159] Wherein, L ce represents the reconstruction loss calculated based on the video candidate segment and p(t d+1 |t 1:d (predict the probability of the d+1th word according to the 1st word to the dth word). The reconstruction loss provides an unsupervised information for network training. We believe that the candidate segment related to the text should be able to ideally reconstruct the text, and if the reconstruction loss is smaller, the candidate segment is closer to the real positioning.
[0160] Disadvantage (3) proposes that if the optimal candidate segment with the minimum reconstruction loss is taken for training each time, then each candidate segment is independent and there is no exchange of information. We believe that if a candidate segment has a large intersection-over-union with the optimal candidate segment, then the candidate segment can also ideally reconstruct the text. The specific operation is to first obtain the optimal candidate segment with the minimum reconstruction loss:
[0161]
[0162] where k is the index of the optimal candidate segment, * is the loss of the positive candidate segment indexed by k. The optimal candidate segment is taken as a pseudo label, and then the intersection-over-union of each positive candidate segment and the optimal candidate segment is calculated:
[0163]
[0164] where, is the optimal candidate segment indexed by k * , m k is the positive candidate segment indexed by k. and represent the end time and the start time of the optimal candidate segment indexed by k * , e k and s k represent the end time and the start time of the positive candidate segment indexed by k.
[0165] The pseudo label value of each candidate segment is obtained by normalizing the intersection-over-union:
[0166]
[0167] where o min and o max are hyperparameters, respectively representing the minimum value and the maximum value of the intersection-over-union.
[0168] According to the pseudo label value and the reconstruction loss of other positive candidate segments, the positive candidate segment learning loss function is calculated as:
[0169]
[0170] where, is the final total learning loss of the positive candidate segment.
[0171] To learn more fine-grained information, negative candidate clips also need to be selected for contrastive learning. Since the entire video contains the semantic information of the text, the entire video is taken as the positive candidate clip, but it contains a lot of redundant information, so the entire video is taken as the global candidate clip. Since the clips at both ends of the positive candidate clip are quite different from the positive sample, the clips at both ends are taken as negative candidate clips. Now there are three types of video clips, the positive candidate clip m p , the global candidate clip of the entire video m r , and the negative candidate clips at both ends of the positive candidate clip m n . We want the loss of the positive candidate clip to reconstruct the text to be as small as possible, and the loss of the negative candidate clip to reconstruct the text to be as large as possible. Since the global candidate clip contains both positive and negative candidate clips, the loss of the global candidate clip to reconstruct the text can be between the positive and negative:
[0172]
[0173] The negative candidate clip is selected from both ends of the optimal candidate clip. The start time of the left end negative candidate clip is fixed as the start time of the video, and the end time of the right end negative candidate clip is fixed as the end time of the video. As the number of training times increases, the width of the negative candidate clip at both ends will continuously increase, and the negative candidate clip will continuously approach the optimal candidate clip. Under the initial condition, the width of the negative candidate clip at both ends is:
[0174]
[0175] where w 1 is the width of the left end negative candidate clip, and w 2 is the width of the right end negative candidate clip.
[0176] The growth rate of the width is:
[0177]
[0178] where w represents the width of the current left end negative candidate clip, and w represents the width of the current right end candidate clip.
[0179] where η will continuously increase with the number of training times:
[0180]
[0181] where e max is the total number of training times, e is the current training time, and the center of the negative sample is:
[0182]
[0183] After obtaining the width and center of the negative candidate frame, the Gaussian mask of the negative candidate frame is calculated, and the Gaussian mask of the entire video is 1.
[0184] After calculating the reconstruction loss of the positive candidate segment The reconstruction loss of the global candidate segment And the reconstruction loss of the left and right negative candidate segments After that, a more complete loss function is designed. Because the entire video also contains positive candidate segment information, the entire video can also be regarded as a rough positive candidate segment. The positive sample learning loss function is designed as follows:
[0185]
[0186] In order to make the network learn the difference between the positive candidate segment and the negative candidate segment, the positive and negative sample comparison learning loss function is designed as follows:
[0187]
[0188] Where β1, β2 are hyperparameters for controlling the difference of sample reconstruction loss.
[0189] The final complete total learning loss is:
[0190] L = L rec + α1L IVC + α2L div + α3L con
[0191] Where α1, α2, α3 are hyperparameters for balancing the total learning loss.
[0192] Step (3), model training, optimizing network parameters;
[0193] According to the loss function designed before, in the training process, the model parameters are updated through the back propagation algorithm (Back-Propagation, BP), until the model converges and the model file is saved.
[0194] Step (4) generates the positioning detection result according to the trained network model, which is implemented as follows:
[0195] After the network loads the saved model file, the test set is tested, and each time a text and a video to be positioned are input, K candidate segments and the corresponding reconstruction loss are finally obtained, and the optimal candidate segment with the minimum reconstruction loss is selected as the positioning result. The start time st and the end time en of the candidate segment are:
[0196]
[0197]
[0198] where Duration is the length of the video.
[0199] Embodiments
[0200] A cross-modal video localization method based on feature alignment, the specific steps are as follows:
[0201] Step (1), data preprocessing stage
[0202] Video preprocessing. We use the public pre-trained C3D model to extract video features from each video. In order to reduce the computational complexity, the features of the ActivityNet Caption dataset are reduced to 500, and the features of the Charades-STA dataset are reduced to 1024.
[0203] Text preprocessing. For each text, we use the NLTK natural language toolkit of Pytorch to split it into words, and use the pre-trained Glove to get the word vector features of each word, and the dimension of the word vector features is 300. We keep the most common words in the training set, so the vocabulary of the ActivityNet Caption word table is 8000, and the vocabulary of the Charades-STA word table is 1111, and unknown words are filled with 0.
[0204] Step (2), network model stage
[0205] The maximum number of frames of the video input to the network is 200, and the maximum number of text words is 20. The video features and text features are both reduced to 256, and the hidden layer state size of the Transformer is also 256. The number of layers is set to 3. In the feature alignment mode, the global feature will be reduced from 256 to 128. The number of generated candidate clips is 8. The maximum number of training is 30 times.
[0206] In the loss part, the α1 of the ActivityNet Caption dataset is 2, the α2 is 0.1, the α3 is 0.01, the β1 is 0.1, the β2 is 0.15, the o min is 0.9, and the λ in L div is 0.135 or α2=1, o min is 0.95, λ=0.136, the α1 of the Charades-STA dataset is 2, the α2 is 0.1, the α3 is 0.01, the β1 is 0.1, the β2 is 0.15, the o min is 1.0, and the λ in L div is 0.150.
[0207] The implementation results on the Charades-STA dataset are as follows:
[0208]
[0209] The implementation results on the ActivityNet Caption dataset are as follows:
[0210]
[0211] Those skilled in the art will understand that the embodiments in the present application can be provided as a method, a system, or a computer program product. Therefore, the embodiments in the present application can adopt a completely hardware embodiment, a completely software embodiment, or an embodiment in the form of combination of software and hardware aspects.
Claims
1. A weakly supervised cross-modal video localization method based on modal feature alignment, characterized in that, The steps include the following: Step (1), data preprocessing, extracting initial modal features of video and text; Step (2), constructing the overall architecture of the network and designing the loss function; Step (3), model training, optimizing network parameters; Step (4), generating positioning detection results according to the trained network model; The step (1) is specifically implemented as follows: Regarding video feature extraction, first convert the video into a set of picture sequences with time sequence, and then use a pre-trained C3D network to extract video features; every 16 frames extract a 4096-dimensional vector, and after processing, finally obtain a vector sequence, that is, the video feature; Regarding text feature extraction, use a Glove word vector data file to obtain the word vector corresponding to each word, and finally obtain the word vector sequence as the text feature; The step (2) is specifically implemented as follows: The network model mainly includes three parts: a feature alignment module, a candidate segment generation module and a text reconstruction module; the feature alignment module is used to extract deep features of two modalities and corresponding global modal features, and the global modal features are used for feature alignment and positive and negative video and text pair learning; the candidate segment generation module is used to fuse the deep features of the two modalities to generate a fixed number of Gaussian distributions to obtain the corresponding candidate segments and Gaussian mask attention; the text reconstruction module uses the candidate segments to reconstruct the masked text, and the optimal candidate segment is obtained through the reconstruction loss; The feature alignment module is specifically implemented as follows: In the data preprocessing stage, the initial features of two modalities have been obtained, the video initial feature is N is the number of extracted video frames, D V is the dimension of the video initial feature, and the text initial feature is M is the number of words, D T is the dimension of the text initial feature; Before inputting into the feature alignment module, three learnable tokens are added to the initial features; in order to get the global feature of the fused feature in the candidate segment generation module, a learnable Gaussian token v gauss is added at the end of the video feature; because the feature alignment module needs to get the global feature of the two modalities, a learnable classification token v cls , cls is added after the video feature and the text feature respectively; finally, the video feature V = {v1, v2, …, v N , gauss , cls} and the text feature T = {t1, t2, …, t M , cls} are obtained. The two modal features after adding word elements are respectively input into the corresponding modal self-attention structure to extract deep modal features; since the self-attention structure is not sensitive to the time sequence of the features, the sine and cosine functions are used to obtain the position information corresponding to each word, and the position encoding is represented by the following formula: PE (pos,2a) = sin (pos / 10000 2a / d model ) PE (pos,2a+1) = cos (pos / 10000 2a / d model ) where a denotes the athdimension of the position encoding vector, pos denotes the position of the current word in the text, d model denotes the dimension of the word vector; After the video features are input into the video modal self-attention structure, the deep video features are obtained, and the text features and position encoding are added and input into the text modal self-attention structure to obtain the deep text features; the self-attention structure is represented by the following formula: wherein X is the input modal feature, is the dimension of X; At this time, the classification word units have obtained global modal characteristics. After mapping the classification word units to low dimensions and normalization, the similarity of matched video text pairs and the similarity of non-matched video text pairs can be calculated by calculating the similarity of classification s=g v (v cls ) T g t (t cls ), g v , g t denote the mapping and normalization operations on videos and texts; in each training batch, the similarity of each video and each text can be calculated, and the similarity matrix s(V,T)=g v (v cls ) T g t (t cls ) and s(T,V)=g t (t cls ) T g v (v cls ), respectively, denote the similarity matrix of the video pair and the text and the similarity matrix of the text pair and the video; the pair similarity matrix is further optimized by softmax normalization: Wherein, B represents the number of video and text sample pairs in a training batch; The contrastive learning loss of video and text is: where CE is the cross entropy loss, y v2t is the true label of the image-text pair, which is 0 if the image and text do not match, and 1 if they do match. t2v is the true label of the image-text pair, which is 0 if the image and text do not match, and 1 if they do match. Before the extracted features are input into the candidate segment generation module, the global features are deleted from the feature sequence, and the obtained video features are The text features are 2.The weakly supervised cross-modal video localization method based on modal feature alignment according to claim 1, characterized in that, The candidate segment generation module is specifically implemented as follows: The essence of the candidate segment generation module is a decoder structure of a transformer, which inputs video features and text features After inputting into the module, the fusion features D H of the fused features, which is expressed by the formula: Where D denotes the decoder operation; at this time All feature information is gathered to predict the center of the candidate segment through a fully connected layer and a sigmoid activation function And the width K is the number of candidate segments; after obtaining the center and width of the candidate segment, the Gaussian mask m is calculated p ∈R K×N : where c is the video feature position, k is the index of the candidate segment, represents the Gaussian mask value of the frame number c in the positive candidate segment with index k, the Gaussian mask is the value representing the Gaussian distribution of each position of the video, and σ is a hyperparameter for controlling the width of the Gaussian curve. In order to make the candidate segments as little overlapping as possible to improve the recall rate of the positioning result, a diversity loss is used to optimize the module, and the diversity loss is expressed as follows: wherein, denotes the Frobenius norm of a matrix, λ is a hyperparameter used to control the degree of overlap of the candidate segments, and I represents the identity matrix. 3.The weakly supervised cross-modal video localization method based on modal feature alignment of claim 1, characterized in that, The text reconstruction module is specifically implemented as follows: The text reconstruction module is also a transformer decoder structure, after obtaining the candidate segment, the video features corresponding to the candidate segment are input into the video modal self-attention structure of the feature alignment module for deep feature extraction, and then the deep features of the candidate segment are input into the text reconstruction module; Then the original text is randomly masked, and the number of masks is 1 / 3 of the number of text words; after the masked text is extracted, the text features are extracted through the text modal self-attention structure of the feature alignment module, and the deep features are input into the text reconstruction module together with the deep features of the candidate segment. But unlike the candidate box generation module, when the two modalities interact, the attention distribution of the two modalities also needs to be multiplied by a Gaussian mask, so that the text pays more attention to the video frames related to the positive candidate segment and suppresses the irrelevant video frame information; finally, the fusion features that fuse the text information and the visual information are obtained The formula is expressed as: wherein, denotes the masked text deep features, is the candidate segment deep features; After the fusion features are input into the full connection layer, the probability of the masked word mapping to all words can be obtained by softmax calculation Q is the total number of words in the text, which is expressed by the formula: wherein W is the weight of the fully connected layer, b' is the bias of the fully connected layer, d e M, the dth fused feature; Each mask position takes the word with the largest corresponding probability as the prediction result; the reconstruction loss L is calculated according to the prediction result ce The negative logarithm probabilities of each masked word are calculated and added together: Among them, L ce Indicates based on video candidate segments and p(t) d+1 |t 1:d The reconstruction loss is calculated as p(t). d+1 |t 1:d ) represents the probability of predicting the (d+1)th word based on the first to the dth words.
4. The weakly supervised cross-modal video localization method based on modal feature alignment according to any one of claims 1-3, characterized in that, The complete loss function is specifically as follows: First, the optimal candidate segment with the minimum reconstruction loss is obtained: where k * is the index of the optimal candidate segment, represents the loss of the positive candidate segment indexed by k; the optimal candidate segment is taken as the pseudo label, and then the intersection over union of each positive candidate segment and the optimal candidate segment is calculated: wherein, is the end time of the optimal candidate segment indexed by k * is the start time of the optimal candidate segment indexed by k k is the positive candidate segment indexed by k and is the end time of the positive candidate segment indexed by k * is the start time of the positive candidate segment indexed by k k and s k is the end time of the positive candidate segment indexed by k The intersection over union is normalized to obtain the pseudo label value of each candidate segment: where o min and o max are hyperparameters, representing the minimum and maximum values of the intersection over union, respectively; According to the pseudo label value and the reconstruction loss of other positive candidate segments, the positive candidate segment learning loss function is calculated as: wherein, is the total learning loss for the final positive candidate segment; To be able to learn more fine-grained information, negative candidate segments also need to be selected for contrastive learning; since the entire video contains the semantic information of the text, the entire video is taken as the positive candidate segment, but it contains a large amount of redundant information, so the entire video is taken as the global candidate segment; since the segments at both ends of the positive candidate segment are different from the positive sample, the two segments are taken as negative candidate segments; now there are three types of video segments, the positive candidate segment m p generated by the candidate segment generation module, the global candidate segment m r of the entire video, and the negative candidate segment m n at both ends of the positive candidate segment; the loss of the global candidate segment reconstructing the text is between the positive and negative: The negative candidate segment is selected from both ends of the optimal candidate segment, the start time of the left end negative candidate segment is fixed as the video start time, the end time of the right end negative candidate segment is fixed as the video end time, with the increase of the training times, the width of the negative candidate segment at both ends will become larger, and the negative candidate segment will be closer to the optimal candidate segment; under the initial condition, the width of the negative candidate segment at both ends is: where w 1 is the width of the left end negative candidate segment, w 2 is the width of the right end negative candidate segment; The growth rate of the width is: wherein, represents the width of the current left end negative candidate segment, represents the width of the current right end candidate segment; Wherein η will become larger with the increase of the training times: where e max is the total number of training, e is the current training number, the center of the negative sample is: After obtaining the width and center of the negative candidate frame, the Gaussian mask of the negative candidate frame is calculated, and the Gaussian mask of the entire video is 1; The reconstruction loss of the positive candidate segment is calculated The reconstruction loss of the global candidate segment The reconstruction loss of the left and right negative candidate segments After that, a more complete loss function is designed; because the entire video also contains positive candidate segment information, the entire video can also be regarded as a rough positive candidate segment, and the positive sample learning loss function is designed as follows: In order to make the network learn the difference between the positive candidate segment and the negative candidate segment, the positive and negative sample contrast learning loss function is designed as follows: Wherein β1, β2 are hyperparameters for controlling the difference of sample reconstruction loss; The final complete total learning loss is: L = L rec + α1L IVC + α2L div + α3L con Wherein α1, α2, α3 are hyperparameters for balancing the total learning loss.
5. The weakly supervised cross-modal video localization method based on modal feature alignment according to claim 4, characterized in that, The specific implementation of step (3) model training is as follows: According to the designed loss function, in the training process, the model parameters are updated through the back propagation algorithm until the model converges and the model file is saved.
6. The weakly supervised cross-modal video localization method based on modal feature alignment according to claim 5, characterized in that, The specific implementation of step (4) is as follows: After the network loads the saved model file, the test set is tested, one text and video to be positioned is input each time, and finally K candidate segments and corresponding reconstruction losses are obtained, the optimal candidate segment with the minimum reconstruction loss is selected as the positioning result; the start time st and the end time en of the candidate segment are respectively: Wherein Duration is the duration of the video.