A Cross-Scene Video Semantic Localization Method and Device Based on Sample Weight Adjustment

A dual-model framework addresses single-modal biases in video semantic localization by learning scene-specific biases and adjusting weights, enhancing precision and generalization across diverse scenes.

CN115761560BActive Publication Date: 2025-07-15PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111026168.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-02
Publication Date
2025-07-15
Estimated Expiration
2041-09-02

AI Technical Summary

Technical Problem

The existing video semantic positioning model has a single-modal preference, which leads to a decrease in positioning accuracy and generalization ability under cross-scene test data, and it is impossible to effectively understand the common semantic information of the two modalities of video and language.

Method used

A dual twin model architecture is adopted, in which one model only reads video input to learn preference information of visual content, and the other model processes video and language input at the same time, eliminates single-modal preferences through the sample weight adjustment mechanism, and improves cross-scene positioning accuracy and generalization capabilities.

Benefits of technology

The generalization ability and positioning accuracy of the model are significantly improved under cross-scenario conditions, and the problem of single-modal preference is alleviated, without additional data annotation and computing resources are required, and a flexible training method is realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115761560B_ABST
    Figure CN115761560B_ABST
Patent Text Reader

Abstract

The present invention relates to a cross-scenario video semantic localization method and device based on sample weight adjustment. The present invention simultaneously uses two twin models with the same backbone network structure. The first model only reads the video input without reading the sentence, and the second model reads the complete video input and the sentence at the same time. The first model is used to learn preference information, predict the localization result based only on the single video modality, and adjust the weights of the training samples according to the learned preference information, so that the training samples received by the second model do not have data preference information, forcing the second model to simultaneously understand the common semantic information in both the video and language modalities. The present invention provides a training framework to prevent the model from overfitting to the preference information in the video segment, enabling it to truly understand both the video and sentence modalities and perform semantic localization in the video according to the semantic information of both. The present invention has obvious advantages in the generalization ability under cross-scenario conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for cross-scenario video semantic localization, and particularly to a method and device for improving the localization accuracy and generalization ability of a cross-scenario video semantic localization model by using a sample weight adjustment method, belonging to the field of computer vision. Technical Background

[0002] Video semantic localization is one of the most important problems in the field of computer vision and has received increasing attention in recent years. Visual semantic localization models have great application potential in many fields such as video surveillance, robotics, and multimedia retrieval. For a given long unclipped video and a sentence composed of natural language, the goal of video semantic localization is to locate the start and end moments of the event described in the sentence in the video. In visual semantic localization, the use of natural language not only makes the action content to be located unrestricted by a pre-defined action label list, but also allows for flexible description of object attributes and relationships. For example, people can use natural language to locate video segments with complex semantic information such as "a man in red clothes takes a cup out of the refrigerator and drinks the water in the cup".

[0003] Visual semantic localization is a typical multi-modal understanding task. The model needs to simultaneously understand the semantic information of both the video and language modalities in order to give the correct localization result. However, existing video semantic localization models have a significant single-modal preference phenomenon: the model can directly make predictions based on only the video single modality without understanding the specific content of the sentence. This leads to a decrease in the localization accuracy of the model under cross-scene test data. The reason for this single-modal preference phenomenon is the obvious data preference in the annotation of the training data. The visual content and time intervals of the localized segments in the training data annotation show a long-tail distribution, that is, some visual content and time intervals frequently appear in the segments to be localized in the training data. For example, in the training data, the visual content "stand" is localized much more frequently than "cut"; similarly, there is a similar preference phenomenon in the time interval distribution: for example, video segments with longer time intervals appear in the training data annotation much more frequently than segments with shorter time intervals. When this data preference in the training data is strong enough, the model will also have a single-modal preference, that is, it tends to guess the video segment to be localized by only using the visual content and time intervals in the video single modality without understanding the sentence. This results in a biased model that cannot simultaneously understand both the video and sentence modalities, and thus cannot perform mutual reasoning on the common semantic information involved in the two modalities. Although the biased model performs well on the test set with the same scene as the training data and similar data preferences to the training data, once the model is used for real application data across scenes, the distribution of the visual content and time intervals of the segment to be localized will change, and its data preference will no longer be the same as that of the training data. Since the data preference during the training stage no longer exists when applying across scenes, and the biased semantic localization model cannot truly perform semantic localization across both the video and language modalities, its generalization ability and localization accuracy will be impaired. This single-modal preference problem of the video semantic localization model severely limits its potential and prospects for industrial applications. Summary of the Invention

[0004] Aiming at the single-modal preference problem, low generalization ability, and low localization accuracy in cross-scene video semantic localization, the purpose of the present invention is to provide a training framework to prevent the model from overfitting to the preference information in the video segment, enabling it to truly understand both the video and sentence modalities simultaneously and perform semantic localization in the video based on the semantic information of both. The core content of the present invention is to use two twin models simultaneously. One model is designed to learn the preference information of the video segment from the training data and is used to further adaptively eliminate the video single-modal preference problem of the other model, thereby improving its cross-scene generalization ability and localization ability.

[0005] Different from previous semantic localization models that are trained using only a single model, the present invention simultaneously uses two twin models with the same backbone network structure, and eliminates the single-modal preference information such as visual content and time interval learned during training in the model through these two twin models. The difference between these two twin models is that the first model only reads the video input without reading the sentence, while the other model normally reads the complete video input and the query sentence at the same time. The role of the first model is to learn the preference information, predict the localization result based on only the video single modality, and further adjust the weights of the training samples according to the learned preference information, so that the training samples received by the second model do not have data preference information and cannot guess based on only the video single modality, forcing the second model to simultaneously understand the common semantic information in both the video and language modalities.

[0006] Specifically, the technical solution adopted by the present invention is as follows:

[0007] A cross-scene video semantic localization method based on sample weight adjustment, comprising the following steps:

[0008] Use a video encoder to extract the visual feature representation of the video candidate window from the input video;

[0009] Use a language encoder to encode the sentence to obtain the feature representation of the sentence;

[0010] Fuse the visual feature representation of the video candidate window and the feature representation of the sentence to obtain the visual-semantic feature representation of the video candidate window;

[0011] Use a visual locator to predict the localization result based only on the visual feature representation of the video candidate window, learn the preference information of the video segment from the training samples, and adjust the weights of the training samples according to the learned preference information;

[0012] Use a visual-semantic locator to predict the localization result based on the visual-semantic feature representation of the video candidate window, and use the training samples with adjusted weights to train the visual-semantic locator to obtain a de-preferred visual-semantic locator;

[0013] For the video and sentence to be localized, use the trained visual-semantic locator for video semantic localization.

[0014] Further, the using a video encoder to extract the visual feature representation of the video candidate window from the input video includes:

[0015] The video encoder divides the input video into multiple video clips, samples the video clips at fixed intervals to obtain N video basic clips, and uses a pre-trained I3D model to extract a series of I3D basic features for each video basic clip;

[0016] For all the I3D basic features included in the video candidate window, apply the boundary matching operator to obtain the visual feature representation of the video candidate window.

[0017] Further, the applying the boundary matching operator to obtain the visual feature representation of the video candidate window includes:

[0018] Perform bilinear interpolation and sampling on all the I3D basic features covered by the video candidate window (a, b) with the starting time a and the ending time b, and obtain K basic feature vectors through sampling, where K is a preset hyperparameter;

[0019] Pass the K basic feature vectors through a convolutional layer with a convolutional kernel size of K and a ReLU layer of a non-linear function to obtain 1 feature vector As the visual feature representation of this video candidate window;

[0020] Repeat the above process for all 1 ≤ a ≤ b ≤ N to obtain the feature vectors of all video candidate windows

[0021] Further, the using the language encoder to encode the sentence to obtain the feature representation of the sentence includes:

[0022] Take the sentence sequence composed of multiple word features as the input and send it to the long short-term memory network LSTM to extract the feature representation of the sentence.

[0023] Further, the training process of the visual locator and the visual-semantic locator includes:

[0024] In the visual locator, the visual feature representation of the video candidate window is directly passed to a fully connected layer and a sigmoid layer to generate the predicted value of the visual score map; in the visual-semantic locator, the visual-semantic feature representation of the video candidate window is input into a fully connected layer and a sigmoid layer to generate the predicted value of the visual-semantic score map of the candidate window; obtain the localization results of the visual locator and the visual-semantic locator from the visual score map and the visual-semantic score map respectively;

[0025] Calculate the loss functions of the visual locator and the visual-semantic locator respectively, and adjust the weights of the training samples according to the localization result of the visual locator, and perform end-to-end training on the visual locator and the visual-semantic locator to obtain a de-preferred visual-semantic locator.

[0026] Further, the calculating the loss functions of the visual locator and the visual-semantic locator respectively, and adjusting the weights of the training samples according to the localization result of the visual locator, and performing end-to-end training on the visual locator and the visual-semantic locator includes:

[0027] For each candidate window in the visual score map and the visual-semantic score map, calculate the IoU score IoU between its temporal boundaries (a, b) and the annotation T ab , and then, according to the preset hyperparameters μ min and μ max , assign a soft label gt ab to IoU ab ;

[0028] Train the visual locator and the visual-semantic locator respectively according to the obtained soft label gt ab , and at the same time adjust the weights of the training samples in the visual-semantic locator according to the localization results of the visual locator, including:

[0029] Use the cross-entropy function as the loss function to train the visual locator , and the definition of which is:

[0030]

[0031] Use the cross-entropy function to calculate the unweighted loss function of the visual-semantic locator:

[0032]

[0033] Calculate the cosine similarity s between the predicted value p′ = {p′ ab} and the true value gt = {gt ab} of the visual locator, estimate the importance 1 - s of the training samples according to the cosine similarity s α , and adjust the loss function of the visual-semantic locator to where α is a hyperparameter that controls the weight change;

[0034] consists of and to form the final loss function According to Train the visual locator and the visual-semantic locator in an end-to-end manner simultaneously to mitigate the video unimodal preference in the model.

[0035] A cross-scene video semantic localization device based on sample weight adjustment adopting the above method, which includes:

[0036] A video encoder, configured to extract the visual feature representation of the video candidate window from the input video;

[0037] A language encoder, configured to encode the sentence to obtain the feature representation of the sentence;

[0038] A visual locator, which is used to predict a localization result only based on the visual feature representation of a video candidate window and learn the preference information of video clips from training samples;

[0039] A visual-semantic locator, which is used to fuse the visual feature representation of a video candidate window and the feature representation of a sentence to obtain the visual-semantic feature representation of the video candidate window, and then predict the localization result based on the visual-semantic feature representation of the video candidate window;

[0040] A sample weight adjustment module, which is used to adjust the weights of training samples according to the preference information learned by the visual locator, and use the training samples with adjusted weights to train the visual-semantic locator to obtain a de-preference visual-semantic locator.

[0041] Compared with the prior art, the advantages of the present invention are as follows:

[0042] (1) Enhanced generalization ability under cross-scene conditions: In the past, video semantic localization models were affected by the single-modal preference problem. Since the data preference in the training dataset no longer exists when applied to cross-scene datasets, their generalization ability under cross-scene settings will be severely impaired. After training with sample weight adjustment, the single-modal preference problem of the model is corrected, and thus the generalization ability of the model under cross-scene conditions is significantly improved.

[0043] (2) No need for additional balanced data annotation: Adding additional balanced training data annotation is beneficial to reducing the single-modal preference of the video semantic localization model and improving its cross-scene generalization ability. However, adding annotation will cost a large amount of additional human resources, and it is difficult to collect balanced video semantic localization data. This method can adaptively adjust the weights and distribution of training samples according to the distribution characteristics of training data, and correct the single-modal preference problem of the video semantic localization model at the algorithm level rather than the data level.

[0044] (3) Flexible and simple training method: This method is based on a dual-model training framework to eliminate the video single-modal preference problem, which can be carried out flexibly end-to-end, that is, by learning the importance of training samples during the training process to adaptively balance the training data and automatically eliminate the influence of video clip preferences on the video semantic localization model.

[0045] (4) Compared with the existing semantic localization technology, no additional computing and storage resources are required in the testing stage: Although two semantic localization models need to be trained simultaneously in the training stage of the present invention, in the testing stage, the visual locator is deleted, and only a single visual-semantic locator is required. Although the present invention eliminates the single-modal preference of the model, the computational complexity in the testing stage is the same as that of the existing semantic localization technology, and no additional computing and storage resources are required. Description of the Drawings

[0046] Figure 1 Schematic diagram of the method flow of the present invention;

[0047] Figure 2 Video processing flow chart of the video encoder;

[0048] Figure 3 Flow chart for sentence feature extraction of the language encoder;

[0049] Figure 4 Soft label assignment strategy diagram;

[0050] Figure 5 Sample weight adjustment diagram. Specific implementation manner

[0051] The present invention will be further described in detail below through specific embodiments and drawings.

[0052] The specific process of the cross-scene video semantic localization method based on sample weight adjustment of the present invention is as Figure 1 shown and includes the following steps:

[0053] Step1: The video encoder divides the input video into multiple video clips, samples these video clips at fixed intervals to obtain N video basic clips. For each sampled video basic clip, a series of I3D basic features are extracted using a pre-trained I3D model.

[0054] Step2: The language encoder encodes each word in the sentence, and takes the sentence sequence composed of multiple word features as input and sends it to a long short-term memory network (LSTM) to extract sentence features.

[0055] Step3: Starting from the I3D basic features, construct the visual feature representation of each video candidate window, that is, apply a boundary matching operator to all the I3D basic features contained in the candidate window to obtain the visual feature representation of the video candidate window.

[0056] Step4: Fuse the visual feature representation of the video candidate window and the feature representation of the sentence, and construct the visual-semantic feature of the video candidate window after their interaction.

[0057] Step5: In the visual locator, the visual features of each video candidate window are directly passed to a fully connected layer and a sigmoid layer to directly generate the predicted value of the visual score map. At the same time, the visual-semantic of the video candidate window is input to another fully connected layer and a sigmoid layer to generate the predicted value of the visual-semantic score map of the candidate window. The localization results of the two locators are obtained from the visual score map and the visual-semantic score map respectively.

[0058] Step 6: Calculate the loss functions of the visual locator and the visual-semantic locator respectively, and adjust the weights of the training samples according to the localization results of the visual locator. Use the Adam algorithm to train the two locators end-to-end, and use the training samples with adjusted weights to obtain a debiased visual-semantic locator.

[0059] Step 7: Discard the visual locator in the test phase, and only use the debiased visual-semantic locator during the training process for semantic localization.

[0060] As Figure 1 shown, this method is based on 5 constructed basic modules to learn the importance of training samples, so as to rebalance the training data and eliminate the influence of video segment preferences for the video semantic localization model. The names and functions of these 5 basic modules are as follows:

[0061] 1. Language encoder: Encode the sentence to be located, and obtain the feature representation of the sentence for effectively retrieving multiple interesting video segments described in the paragraph in the video.

[0062] 2. Video encoder: Extract the visual features in time and space of the part covered by the video candidate window from the original video frames, so as to obtain the feature representation that can depict the visual structure of different video candidate windows

[0063] 3. Visual locator: The visual locator predicts the localization result only based on the visual features of the video candidate window extracted by the video encoder, without the need to input the sentence to be queried. The visual locator can learn the preference information of video segments from the training data and use it to further adjust the loss function of the visual-semantic locator.

[0064] 4. Visual-semantic locator: The visual-semantic locator inputs both the video and the sentence to be queried in a complete manner, and predicts the localization result according to the visual-semantic feature representation of the video candidate window.

[0065] 5. Sample weight adjustment module: The sample weight adjustment module uses the predicted output of the visual locator to predict the importance of the training samples, and adjusts their weights in the loss function of the visual-semantic locator accordingly, so as to reduce the data preference problem of the training samples received by the visual-semantic locator, that is, to achieve debiasing processing.

[0066] The implementation process of each step of the present invention will be specifically described below.

[0067] 1. Video preprocessing

[0068] This step preprocesses the unclipped long video to be located, cuts it into multiple video clips, samples and extracts its visual features from them for subsequent understanding of its semantic content and location. When extracting the visual features of the video, the I3D model pre-trained on the Kinetics dataset is used. The preprocessing process is as Figure 2 shown and includes the following steps:

[0069] (Step1) Split the input video into multiple video clips, where each video clip contains T frames.

[0070] (Step2) Sample the video clips at fixed intervals to obtain N video base segments.

[0071] (Step3) For each sampled video base segment, use the pre-trained I3D model to extract the I3D basic feature vectors respectively. A total of N I3D basics can be obtained. The feature V = {v i} (i = 1..N), and these basic features represent the visual content of the video from the 1st to the Nth video segments.

[0072] 2. Sentence Feature Extraction

[0073] To locate the semantic content described by the sentence in the video, it is necessary to extract the semantic features of the sentence to be queried and vectorize them. This method uses a word vector model pre-trained on large-scale text data and a long short-term memory (LSTM) network to extract the semantic features of the sentence. Let the number of words in the sentence be V s , and the specific process of its feature extraction is as Figure 3 shown and includes the following steps:

[0074] (Step1) Use a word embedding model pre-trained on large-scale text data to encode each word in the sentence, and each word obtains a word feature vector w i . A total of V s word features are extracted from this sentence, which can be represented as a sentence sequence {w i} (i = 1..V s ).

[0075] (Step2) Take the sentence sequence {w i} (i = 1..V s ) composed of multiple word features as the input and send it to the LSTM network in turn, and output the last hidden state of the LSTM network.

[0076] (Step 3) Input the last hidden state of the LSTM network into a fully connected layer to extract the final sentence feature f S

[0077] 3. Generate the visual features of the candidate windows

[0078] Starting from the I3D features of the video base segments obtained by preprocessing, the semantic features of the candidate windows can be generated. The role of the visual features of the candidate windows is to characterize and represent the video visual content covered by the candidate windows. For all the I3D basic features included in this candidate window, in this method, the boundary matching operator (Boundary Matching Operation) BM is applied to obtain the visual feature representation of the video candidate window.

[0079] The boundary matching operator BM can effectively generate the features of the candidate windows from the video base segments through a series of bilinear sampling and convolution operations. For the video candidate window (a, b) with the starting time a and the ending time b, the specific steps of the BM operator are as follows:

[0080] (Step 1) Perform bilinear interpolation and sampling on all the I3D basic features covered by (a, b), and obtain K basic feature vectors through sampling. Where K is a preset hyperparameter.

[0081] (Step 2) Pass these K basic features through a convolutional layer with a convolutional kernel size of K and a ReLU layer of a nonlinear function to obtain 1 feature vector As the visual feature representation of this video candidate window.

[0082] (Step 3) Repeat the above process for all 1 ≤ a ≤ b ≤ N to obtain the feature vectors of all video candidate windows Where N represents the number of video base segments.

[0083] 4. Generate the visual-semantic features of the candidate windows

[0084] Specifically for generating the visual-semantic features of the candidate windows, then we fuse the features of the two modalities of vision and language to generate the visual-semantic features of each candidate window.

[0085] (Step 1) For the candidate window (a, b), take the visual features of this candidate window And interact with the language feature f S That is, multiply f S And Point by point to obtain the visual-semantic feature

[0086] (Step 2) Take the visual-semantic feature Normalize it with its L2 norm: Obtain the visual-semantic feature M of the candidate window ab ;

[0087] (Step3) Repeat the above process for all candidate windows (a, b) to obtain their visual-semantic features

[0088] 5. Localize using the visual locator and the visual-semantic locator

[0089] In this method, two twin locator models, the visual locator and the visual-semantic locator, are used to localize video clips simultaneously. The role of the visual locator is as follows: directly guess the most likely predicted clip from the visual features of all candidate windows according to the unimodal preference information of the video, without the need to input the sentence to be queried, so as to learn the labeled unimodal preference from the training data. The visual-semantic locator inputs the visual-semantic features of the candidate windows and locates the clip described by the sentence in the video by simultaneously understanding the visual and semantic contents of both the video and the sentence modalities.

[0090] The processes of localizing using the visual locator and the visual-semantic locator respectively are described as follows:

[0091] For the visual locator, the visual features of each video candidate window are input into a fully connected layer and a sigmoid layer to directly generate the predicted value p′ of the visual score map for each candidate window (a, b) ab . The candidate window corresponding to its maximum value argmax (a,b) p′ ab is used as the final visual localization result. For the visual-semantic locator, the localization process of its video clip is similar to that of the visual locator, and the difference between them lies in the input features. The input feature of the visual-semantic locator is the visual-semantic feature {M ab} of each video candidate window, and this feature is passed to another fully connected layer and a sigmoid layer to generate the predicted value p ab of the visual-semantic score map for each candidate window (a, b). The candidate window corresponding to its maximum value argmax (a,b) p ab is used as the final visual-semantic localization result.

[0092] 6. Adjust the sample weights during the training process

[0093] 6.1 Determine the labels of the training samples

[0094] During the training phase, each training sample contains an input video V, a sentence S, and a temporal segment annotation T corresponding to the sentence. During training, it is necessary to determine which temporal segment in the visual-semantic score map corresponds to the ground truth of the annotation and train the model accordingly. During training, first calculate the IoU score between each candidate window and the annotated temporal segment, and use soft labels to specify which candidate window belongs to the true result according to the IoU score. For each candidate window in the visual score map and the visual-semantic score map, calculate the IoU score IoU between its temporal boundaries (a, b) and the annotation T ab . Then, according to the pre-set hyperparameter μ min and μ max , assign a soft label gt ab to IoU ab . The assignment process is as shown in Figure 4 and includes the following steps:

[0095] (i) When IoU ab ≤μ min , gt ab =0

[0096] (ii) When μ min ≤IoU ab ≤μ max ,

[0097] (iii) When μ max ≤IoU ab , gt ab =1

[0098] 6.2 Calculate the loss function and adjust the sample weights

[0099] In the sample weight adjustment module, train the visual locator and the visual-semantic locator respectively according to the obtained soft label gt ab . At the same time, adjust the weights of the training samples in the visual-semantic locator according to the localization results of the visual locator. The process is described as shown in Figure 5 and includes the following steps:

[0100] (Step1) Use the traditional cross-entropy function as the loss function to train the visual locator for each training sample, which is defined as:

[0101] (Step2) Similar to Step1, use the traditional cross-entropy function to calculate the unweighted loss function of the visual-semantic locator:

[0102] (Step3) Calculate the predicted value p′ of the visual locator = {p′ab} and the ground truth gt = {gt ab} of the cosine similarity:

[0103]

[0104] (Step4) Estimate the importance 1 - s of the training samples according to the cosine similarity α , and adjust the loss function of the visual - semantic locator to where α is a hyperparameter that controls the weight change.

[0105] (Step5) Balance the two loss functions. The final loss function is composed of the loss function of the visual locator and the adjusted loss function of the visual - semantic locator

[0106] (Step6) Use the Adam algorithm as the optimization algorithm of the model. According to the above loss function Train the visual locator and the visual - semantic locator in an end - to - end manner to reduce the video unimodal preference in the model.

[0107] 7. Testing phase

[0108] The above training phase involves two locators, namely the visual locator and the visual - semantic locator. In the testing phase, the visual locator is removed, and only the visual - semantic locator is involved. After being trained by the above method, the video unimodal preference of the visual - semantic locator has been reduced after the adjustment of the training sample weights. Therefore, the biased visual locator can be removed in the testing phase, and only the de - preferred visual - semantic locator is used for video semantic localization. The visual - semantic locator finally outputs the prediction scores p of all candidate windows ab . Take the maximum value max (a,b) p ab , and the corresponding time segment argmax (a,b) p ab is used as the final localization result.

[0109] ​The present invention evaluates the positioning accuracy of the above-mentioned model in a cross-scenario. The most commonly used evaluation metric for evaluation is R@K,θ. The meaning of R@K,θ is the proportion that the IoU of the overlapping degree between at least one time segment in the top K positioning results and the true value exceeds θ. The commonly used existing positioning video semantic positioning datasets are ActivityNet Captions, Charades-STA, and DiDeMo. The present invention trains and evaluates the model on datasets in two different scenarios respectively. Considering that the dataset ActivityNet Captions is large in scale and involves various scenarios and activity types, this article uses ActivityNet Captions as the training set to train the model, and then tests the model on the datasets Charades-STA and DiDeMo (referred to as AcNet2Charades and AcNet2DiDeMo respectively). Tables 1 and 2 show the comparison of the positioning accuracy with other existing methods under the cross-scenario settings of AcNet2Charades and AcNet2DiDeMo.

[0110] Table 1: Comparison of model performance under the cross-scenario setting of AcNet2Charade

[0111] Method R@1, IoU = 0.5 R@1, IoU = 0.7 R@5, IoU = 0.5 R@5, IoU = 0.7 PFGA 5.75 1.53 - - SCDM 15.91 6.19 54.04 30.39 2D-TAN 15.81 6.30 59.06 31.53 The present invention 21.45 10.38 62.34 32.90

[0112] Table 2: Comparison of model performance under the cross-scenario setting of AcNet2DiDeMo

[0113] Method R@1, IoU = 0.5 R@1, IoU = 0.7 R@5, IoU = 0.5 R@5, IoU = 0.7 PFGA 6.24 2.01 - - SCDM 10.88 4.34 43.30 18.40 2D-TAN 12.50 5.50 44.88 20.73 The present invention 13.11 7.70 44.98 21.32

[0114] In the present invention, the weights of different training samples are adjusted in the loss function according to the cosine similarity s In fact, the adjustment of the weights in the loss function is not limited to the above method. As long as the training samples with low similarity are given higher weights in the loss function, and the samples with high similarity are given lower weights, the purpose of balancing the training data distribution can be achieved.

[0115] Based on the same inventive concept, another embodiment of the present invention provides a cross-scenario video semantic positioning device based on sample weight adjustment using the above method, which includes:

[0116] A video encoder for extracting the visual feature representation of video candidate windows from the input video;

[0117] A language encoder for encoding sentences to obtain the feature representation of the sentences;

[0118] A visual locator for predicting the positioning result only based on the visual feature representation of the video candidate window and learning the preference information of video segments from the training samples;

[0119] A vision-semantic locator that fuses the visual feature representation of a video candidate window and the feature representation of a sentence to obtain the visual-semantic feature representation of the video candidate window, and then predicts a localization result based on the visual-semantic feature representation of the video candidate window;

[0120] A sample weight adjustment module that adjusts the weights of training samples according to the preference information learned by the visual locator, and uses the training samples with adjusted weights to train the vision-semantic locator to obtain a preference-free vision-semantic locator.

[0121] Based on the same inventive concept, another embodiment of the present invention provides an electronic device (such as a computer, a server, a smart phone, etc.), which includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for performing the steps in the method of the present invention.

[0122] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, a magnetic disk, an optical disk). When the computer program stored in the computer-readable storage medium is executed by a computer, the various steps of the method of the present invention are implemented.

[0123] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and implement it accordingly. Those of ordinary skill in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the content disclosed in the embodiments of this specification, and the protection scope of the present invention is defined by the scope of the claims.

Claims

1. A cross-scenario video semantic localization method based on sample weight adjustment, characterized in that Including the following steps: Using a video encoder to extract the visual feature representation of a video candidate window from an input video; Using a language encoder to encode a sentence to obtain the feature representation of the sentence; Fusing the visual feature representation of the video candidate window and the feature representation of the sentence to obtain the visual-semantic feature representation of the video candidate window; Using a visual locator to predict a localization result only based on the visual feature representation of the video candidate window, learning the preference information of video segments from training samples, and adjusting the weights of the training samples according to the learned preference information; Using a visual-semantic locator to predict a localization result based on the visual-semantic feature representation of the video candidate window, and training the visual-semantic locator using the training samples with adjusted weights to obtain a de-preferred visual-semantic locator; For a video and a sentence to be localized, using the trained visual-semantic locator to perform video semantic localization.

2. The method according to claim 1, wherein The using a video encoder to extract the visual feature representation of a video candidate window from an input video includes: The video encoder divides the input video into multiple video clips, samples the video clips at fixed intervals to obtain N video base clips, and extracts a series of I3D basic features for each video base clip using a pre-trained I3D model; Applying a boundary matching operator to all the I3D basic features included in the video candidate window to obtain the visual feature representation of the video candidate window.

3. The method according to claim 2, wherein The applying a boundary matching operator to obtain the visual feature representation of the video candidate window includes: Performing bilinear interpolation and sampling on all the I3D basic features covered by a video candidate window (a, b) with a start time of a and an end time of b, and obtaining K basic feature vectors through sampling, where K is a preset hyperparameter; Pass the K basic feature vectors through a convolutional layer with a convolutional kernel size of K and a ReLU non-linear function layer to obtain 1 feature vector. As the visual feature representation of the candidate window of this video; Repeat the above process for all \(1\leq a\leq b\leq N\) to obtain the feature vectors of all video candidate windows 4. The method according to claim 1, wherein The using a language encoder to encode a sentence to obtain the feature representation of the sentence includes: Taking a sentence sequence composed of multiple word features as input and sending it to a long short-term memory network (LSTM) to extract the feature representation of the sentence.

5. The method according to claim 1, characterized in that The training processes of the visual locator and the visual-semantic locator include: In the visual locator, the visual feature representation of the video candidate window is directly passed to a fully connected layer and a sigmoid layer to generate the predicted value p′ of the visual score map. ab In the visual-semantic locator, the visual-semantic feature representation of the video candidate window is input into a fully connected layer and a sigmoid layer to generate the predicted value p of the visual-semantic score map of the candidate window. ab The localization results of the visual locator and the visual-semantic locator are obtained from the visual score map and the visual-semantic score map respectively. Calculating the loss functions of the visual locator and the visual-semantic locator respectively, adjusting the weights of the training samples according to the localization result of the visual locator, and training the visual locator and the visual-semantic locator end-to-end to obtain a de-preferred visual-semantic locator.

6. The method according to claim 5, wherein The calculating the loss functions of the visual locator and the visual-semantic locator respectively, adjusting the weights of the training samples according to the localization result of the visual locator, and training the visual locator and the visual-semantic locator end-to-end includes: For each candidate window in the visual score map and the visual-semantic score map, calculate the IoU score IoU between its temporal boundaries (a, b) and the annotation T ab , and then according to the preset hyperparameters μ min and μ max , assign a soft label gt ab to IoU ab ; According to the obtained soft label gt ab Train the visual locator and the visual-semantic locator respectively, and at the same time adjust the weights of the training samples in the visual-semantic locator according to the positioning results of the visual locator, including: Use the cross-entropy function as the loss function Train the visual locator, The definition of which is: Using a cross-entropy function to calculate the loss function of the visual-semantic locator without weight adjustment: Calculate the cosine similarity s between the predicted value p′ = {p′ ab} and the ground truth gt = {gt ab}, estimate the importance of the training sample as 1 - s based on the cosine similarity s α , and adjust the loss function of the vision-semantic locator to where α is a hyperparameter that controls the change of the weight; Consisting of and The two parts form the final loss function According to The visual locator and the vision-semantic locator are trained in an end-to-end manner simultaneously to mitigate the video unimodal preference in the model.

7. The method according to claim 6, wherein The soft label gt ab The allocation process includes: (i) When IoU ab ≤ μ min , gt ab = 0; (ii) When μ min ≤ IoU ab ≤ μ max At this time, where μ min and μ max are pre-set hyperparameters; (iii) When μ max ≤ IoU ab , gt ab = 1.

8. A cross-scenario video semantic localization device based on sample weight adjustment using the method according to any one of claims 1 to 7, characterized in that, Including: A video encoder for extracting the visual feature representation of a video candidate window from an input video; A language encoder for encoding a sentence to obtain the feature representation of the sentence; A visual locator for predicting a localization result only based on the visual feature representation of the video candidate window and learning the preference information of video segments from training samples; A vision-semantic locator that fuses the visual feature representation of a video candidate window and the feature representation of a sentence to obtain the vision-semantic feature representation of the video candidate window, and then predicts a localization result based on the vision-semantic feature representation of the video candidate window; A sample weight adjustment module that adjusts the weights of training samples according to the preference information learned by the vision locator, and uses the training samples with adjusted weights to train the vision-semantic locator to obtain a vision-semantic locator with de-preference processing.

9. An electronic device, characterized in that, It includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for executing the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • RGBD image semantic segmentation method

    CN107403430A

  • Double-stream video classification method and device based on cross-mode attention mechanism

    CN110188239A