False dependency elimination method for video clip retrieval

By constructing a structured attribution model and backdoor adjustment method, splitting and optimizing video features, the time domain position distribution bias problem of VCMR model on large-scale video data sets is solved, the correct dependence of the model on the target time domain position is achieved, and the generalization ability and retrieval accuracy of the model are improved.

CN119149778BActive Publication Date: 2025-05-09BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411277266.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-12
Publication Date
2025-05-09
Estimated Expiration
2044-09-12

AI Technical Summary

Technical Problem

The existing video set clip retrieval (VCMR) model has a bias problem in the time domain position distribution when dealing with large-scale video data sets, resulting in the model's erroneous dependence on the location of the target area, affecting the inference authenticity of the model.

Method used

The model inference path is analyzed by constructing a structured attribution model (SCM), and the model inference path is adjusted through attribution intervention, split the video features into content features and position features, and reconstruct the model inference path using the backdoor adjustment method to reduce the erroneous dependence on the target time domain location.

Benefits of technology

Effectively eliminate the model's erroneous dependence on the time domain location distribution of the data set, improve the model's generalization ability on different distributed data sets, and significantly improve the model's retrieval accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119149778B_ABST
    Figure CN119149778B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for eliminating the false dependency of video set segment retrieval, and belongs to the field of multimodal data retrieval. The method of the present invention comprises: re-segmenting the video data set currently applied by the VCMR model, evaluating the bias dependency of the VCMR model on the target time domain position in the video data set through out-of-distribution testing on the re-segmented data set, if the performance of the model is significantly reduced compared with the original data set, it means that the model has obvious false dependency on the bias in the data set; constructing a structured attribution model to analyze the reasoning path of the model, adjusting the reasoning path of the model through attribution intervention, and alleviating and eliminating the false bias dependency of the VCMR model on the target time domain position. The method of the present invention realizes a fair test of the generalization ability of the model on data sets with different distributions, can correct the original reasoning path that makes false dependency on the bias of the data set, and significantly improves the retrieval and generalization ability of the model on data sets with different distributions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of multimodal data retrieval, relates to data mining technology, and specifically relates to a method for eliminating error dependencies in video set segment retrieval. Background Art

[0002] Nowadays, with the rapid increase of online video content, people have an increasing demand to retrieve target moments from a large number of video collections. People can quickly locate the video where the described content is located and the temporal location in the video through language description. Therefore, the task of Video Corpus Moment Retrieval (VCMR) has become a popular research trend. Based on natural language queries, VCMR includes two subtasks: video retrieval and moment location, which retrieve specific video moments from a large number of unedited and unsegmented video collections. In the most widely studied methods of VCMR, features of two modalities—video features and query features—are projected into a common embedding space and cross-modal feature matching is performed. According to the early or late fusion of features of different modalities, this type of work can be divided into early fusion and late fusion. At present, the late fusion strategy has received more extensive research attention due to its comparable accuracy and significantly superior efficiency compared to the early fusion strategy. In the late fusion strategy mode, the features of the two modalities are optimized separately after mapping and then fused for inference.

[0003] Although existing studies have given seemingly good experimental results, we believe that these results do not truly reflect the model's ability to perform multimodal semantic understanding, but rather rely on dataset bias. Single video segment retrieval (VMR) is a subtask of VCMR, which only requires locating the temporal position of the description content from a single video. Recent studies on the VMR task have found that many state-of-the-art models have implicit distribution biases when training common datasets. Similarly, we believe that the VCMR task may also be affected by various dataset biases. In large-scale video datasets, this bias may originate from the process of selecting the benchmark truth moment: segments near the beginning or end of the video are more likely to be selected as the benchmark truth. Therefore, when the model is reasoning, it will have a false reliance on the location of the target area, that is, it is more inclined to generate results from these areas, which in turn affects the model's reasoning authenticity.

[0004] In recent years, many research works have been carried out around the false dependency in VMR tasks. They have a set of mature methodological processes: (1) verifying the impact of false dependency on model reasoning through experiments; (2) proposing new model designs to solve this false dependency. However, the specific scheme used in the VMR task cannot be directly used in the VCMR task. The main reason is that the retrieval scope of the VMR task is limited, resulting in very loose restrictions on the scale of the model. In comparison, the VCMR task needs to retrieve massive video content and uses a completely different reasoning path from the VMR task. Therefore, the debiasing method of VMR cannot be used in the VCMR task.

[0005] Dataset and model debiasing is a common problem in the field of data mining. For example, reference document 1 (Chinese invention patent application with publication number CN118312653A published on July 9, 2024, "Project recommendation method and system based on causal popularity debiasing") quantifies user consistency by integrating project popularity and user interest, and alleviates the impact of popularity bias on recommendation results; reference document 2 (Chinese invention patent application with publication number CN118247003A published on June 25, 2024, "A debiased sequence recommendation method and system for separating interest and conformity representation") solves popularity bias by introducing self-supervised separation representation learning into sequence recommendation; these technical solutions all focus on model debiasing in the field of recommendation systems, but there is currently a lack of relevant solutions for the temporal segment position bias in VCMR tasks. For example, reference document 3 (G. Nanet al., "Interventional Video Grounding with Dual Contrastive Learning," 2021 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR)) proposed an intervention video anchoring paradigm to eliminate selection bias and introduced a dual contrastive learning method to align text and video; reference document 4 (Yang X, Feng F, Ji W, et al. Deconfounded Video Moment Retrieval with Causal Intervention [J]. ACM, 2021. DOI: 10.1145 / 3404835.3462823.) uses a deconfusion cross-modal matching method to eliminate the confusion effect of moment position to capture the true impact of query and video content; these technical solutions solve the bias problem in VMR, but due to different reasoning paths, they cannot be transferred to VCMR tasks. For example, reference document 5 (Yoon S, Hong JW, Yoon E, et al. Selective Query-Guided Debiasing for Video Corpus Moment Retrieval [C] / / European Conference on Computer Vision. Springer, Cham, 2022. DOI: 10.1007 / 978-3-031-20059-5_11.) attempts to solve the predicate-object combination bias in the VCMR task, but also does not solve a more common and significant bias in the VCMR task, namely the time domain segment position bias. Summary of the invention

[0006] In view of the bias problem of the temporal position distribution of VCMR datasets that has not been solved in the prior art, the present invention provides a method for eliminating the incorrect dependence of video set fragment retrieval, and for VCMR task scenarios, provides an evaluation system for model bias dependence, verifies the incorrect dependence of existing models on the temporal position distribution of datasets, and provides a method for alleviating and eliminating the bias dependence of models, thereby eliminating the incorrect dependence of existing VCMR models on the temporal position distribution of datasets.

[0007] The present invention provides a method for eliminating error dependencies in video set segment retrieval, comprising:

[0008] Step 1, re-segment the video data set currently used by the VCMR model, and evaluate the bias dependence of the VCMR model on the target time domain position in the video data set through out-of-distribution testing on the re-segmented data set;

[0009] Step 2: When it is necessary to alleviate and eliminate the erroneous paranoid dependence of the VCMR model, the reasoning path of the VCMR model is analyzed by constructing a structured attribution model, and the reasoning path of the VCMR model is adjusted through attribution intervention, including:

[0010] (1) The initial confused video feature v 0 Through two linear layers g c and g l Split into features c representing content v and the feature l representing the location v ; Among them, two linear layers g are trained c and g l Including: position features l v Set a non-learnable positional encoding p and train a linear layer g c Make l v Similar to the positional encoding p, training the linear layer g l Make the content feature c v Close to the initial confused video features, train the linear layer g c and g l Make the content feature c v and position feature l v Try to stay away;

[0011] (2) Use the do operation to intervene in the query text and video content features, reconstruct the reasoning path of the VCMR model, and alleviate and eliminate the VCMR model's incorrect biased dependence on the target temporal position;

[0012] The calculation method for setting the do operation is as follows:

[0013]

[0014] The calculation method of the reconstructed reasoning process is as follows:

[0015]

[0016] Among them, the conditional probability P(X|do(Q,V)) observes the output X of the VCMR model after executing the do operation do(Q,V), X is the matching score ML of the query text Q and the video V in different sub-segments, or the matching score VR of the query text Q and the video V; the three variable parameters q, v and l represent the text features, video content features, and position features respectively; f*(q,v,l) means using the linear method f* to calculate the score ML or VR for the combination (q,v,l); E l (f*(q,v,l)) is the expectation of ML or VR at the specified expected position l; φ * It is the max operation in VR tasks or the one-dimensional convolution operation for calculating the start and end point scores in ML tasks; q * Query text features used for different tasks; is the reorganized video feature, W 1 is the weight parameter; h c (l) is a location feature improved by integrating video content information; is to calculate h at the specified position l c (l) expectations;

[0017] The final query statement-video time domain segment matching score is calculated based on the ML and VR scores observed by the do operation, and the task output result is obtained after final sorting.

[0018] The step 1 is to re-split the video data set of the current application by estimating the frequency of the video samples, construct an out-of-distribution test OOD-test set, and re-split the test set and the training set; wherein, when calculating the frequency estimation value of the video, it is necessary to first calculate the probability density of the target time sequence position, determine the relative position of the start position and the end position of the target time sequence position corresponding to the query statement in the entire video, use a Gaussian kernel to perform kernel density estimation, and obtain the probability density of the target time sequence position; then calculate the frequency estimation value S of each video q in the current data set v as follows:

[0019]

[0020] Among them, N q represents the number of query-target temporal positions contained in the video, where the relative starting position of the i-th target temporal position is The relative end position is is the probability density of the i-th target time series position.

[0021] Compared with the prior art, the present invention has the following advantages and positive effects: (1) The method of the present invention uses an OOD segmentation method to re-segment the data set, which can fairly test the generalization ability of the model on data sets with different distributions. (2) The method of the present invention constructs a model reasoning path diagram based on SCM, analyzes and corrects the reasoning path of the model, and designs a reasoning process to correct the original reasoning path that incorrectly relies on the data set bias. (3) The method of the present invention proposes a specific model correction scheme for the model reasoning path in theory through a backdoor adjustment method, so that the model no longer relies too much on the wrong data set bias, thereby significantly improving the generalization ability of the model on data sets with different distributions. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 The schematic diagram of the SCM that can be constructed for the VCML task reasoning process and the SCM after error correction;

[0023] Figure 2 The present invention is a structural schematic diagram of alleviating and eliminating the model's dependence on the target time domain position bias. DETAILED DESCRIPTION

[0024] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0025] The method of the present invention includes two main invention points: (1) A system for evaluating the model's bias dependence is proposed. For an existing VCMR dataset and inference model, a system for evaluating the degree of erroneous reliance of the model on the bias of the time domain position where the benchmark truth value is located in the dataset is constructed through a dataset re-segmentation method. (2) A method for alleviating and eliminating the model's bias dependence is proposed. The model reasoning path is analyzed by constructing a structured attribution model (SCM), and the model reasoning path is adjusted through attribution intervention, thereby reducing the erroneous reliance on the target time domain position, and finally optimizing the generalization ability of the VCMR model on datasets with different distributions.

[0026] To achieve the above-mentioned invention point (1), the method of the present invention designs a debiased evaluation scheme for VCMR, namely, the out-of-distribution test (OOD test), which changes the distribution of the data of the training set and the test set in the target time domain by re-segmenting the existing data set, and evaluates the debiased reasoning ability of the VCMR model on the re-segmented data set.

[0027] The out-of-distribution test designed by the method of the present invention first merges the training set, test set and validation set of the original video subtitle segment set to form a merged set. A VCMR model can perform retrieval tasks well on a certain video subtitle segment set, but when applied to a user-defined video subtitle segment set, the retrieval effect may not be good, so it is necessary to evaluate the bias dependence of the VCMR model on the time domain position of the reference true value in the current data set. The embodiment of the present invention uses the TVR data set to evaluate the bias dependence of the VCMR model, and the specific use is based on the data set provided by the user for evaluation.

[0028] Then, a Gaussian kernel is used to perform kernel density estimation (KDE) on the relative positions of the start and end points of the target time domain segment in the merged dataset. The relative position is the proportional position of the start / end point in the total length of the video. The kernel density is calculated as follows:

[0029]

[0030] in, is the KDE value at the data point with the relative starting position x∈[0,y] and the relative ending position y∈(x,1], n is the number of samples, and the i-th sample (x i ,y i ) records the relative starting position x of the target segment i and the relative end position y i , h is the bandwidth estimated using Scott's method, K(·) is the Gaussian kernel function, u is the parameter, Represents the sample standard deviation.

[0031] Then, to complete the OOD test, the method of the present invention re-splits the existing data set to construct new subsets with different data distributions. Usually, for the unevenly distributed related attributes, the most atypical part is extracted to form a test data set, and the rest is divided into other randomly distributed data sets. With respect to the time domain position bias that the present invention is concerned with, the TVR data set is re-split according to the position distribution at the target moment. In order to cope with the situation where the original TVR data set is organized by video rather than by query statement-target time domain segment, the present invention re-splits the original TVR data set by video-level sample frequency estimation, thereby constructing a TVR-OOD data set.

[0032] Finally, for any existing VCMR inference model, train and test the model on the re-segmented dataset. If the performance of the model does not drop significantly compared to the original dataset, it means that the model has no or little false reliance on the target temporal position bias of the dataset. On the contrary, if the performance of the model drops significantly on the reconstructed dataset with the target temporal position distribution changed, such as a drop higher than a preset threshold, it can be proved that the model has a significant false reliance on the bias in the dataset.

[0033] To achieve the above invention point (2), the present invention designs an SCM framework for analyzing and modifying the reasoning path of the VCMR model, and analyzes the incorrect dependence on the data set bias caused by the reasoning path defects of the existing VCMR method. Then, by modifying the reasoning path in the SCM and processing the input data in a different way, the entanglement is resolved, and finally the model's dependence on the data set bias is alleviated.

[0034] like Figure 1 As shown in the figure, in the existing VCMR reasoning model, there are implicit and obfuscated reasoning paths V←L→VR and V←L→ML, where V represents the obfuscated / purified video (content) features, L represents the area where the target temporal position in the reference truth appears in the entire video, L is the implicit / explicit location information set, and VR is the query statement Q V The matching score of the entire video V, ML is the query statement Q M The matching scores of the video V in different sub-clip segments, VR and ML are the output values ​​of the model inference, that is, the final calculation result of the task Y. Existing work directly uses the video feature V for cross-modal matching with text information, which causes the model to implicitly introduce dataset bias as the basis for reasoning, resulting in a significant decrease in the performance of the model in datasets with different distributions.

[0035] exist Figure 1 In the above example, the reasoning path of L→V is the erroneous knowledge learned by the model through the distribution bias of the data set, that is, the reasoning path that needs attribution intervention. Therefore, the method of the present invention proposes an attribution intervention scheme to eliminate or weaken the reasoning path of L→V through backdoor adjustment, and eliminate the confusion of the implicit and confused V←L→VR and V←L→ML reasoning paths and construct an explicit reasoning path.

[0036] In order to eliminate the entanglement and confusion of the content information in the video features, that is, the real video features that are expected to be obtained, and the implicit self-position information, the initial mixed feature v 0 Split into content-representing features c through two different linear layers v and the feature l representing the location vThen, the parameters of the two linear layers are adjusted through training, and for the position feature l v , creating a non-learnable positional encoding p and using it in training The loss function is l v Similar to position encoding; for content feature c v , use L in training indep Function makes it and position feature l v Try to keep away from it to reduce the position information in the content vector. At the same time, use Function makes content feature c v and the initial confused video feature v 0 Close to ensure that it does not lose the ability to represent the content. The loss function is defined as follows:

[0037]

[0038] The “^” on the character refers to the softmax operation, and KL is the KL divergence.

[0039] Finally, the do operation is used to implement backdoor adjustment and reconstruct the reasoning path of the model. Specifically, the expected approximation of the position vector is performed through the Bayesian formula, and the final reasoning of the model is completed through the text feature Q, the video content feature V and the expected position feature l:

[0040]

[0041] In the above formula, X is the ML result output or VR result output, and the video text matching calculation of both uses cosine similarity calculation. P(l) represents the probability of the expected position feature l, do(Q,V) refers to intervening in the content features of the query text Q and the video V, performing the do operation, and using the conditional probability P(X|do(Q,V)) to observe the output of the VCMR model after performing the do operation. The final output of the task is calculated by the scores of ML and VR.

[0042] The implementation steps of the method for eliminating false dependencies in video set segment retrieval according to the embodiment of the present invention include the following seven steps.

[0043] Step 1) Normalize the target temporal position of the benchmark truth value in the dataset using the total video duration as the unit.

[0044] Step 2) Use the Scott bandwidth method to estimate the bandwidth and use the Gaussian kernel to perform kernel probability density analysis.

[0045] Step 3) Taking the video as a unit, the frequency estimation value of each video in the data set is calculated according to the probability density of the target temporal position contained therein. The specific calculation method is as follows:

[0046]

[0047] Among them, Nq represents the number of query statements-target time domains contained in video q, represents the probability density of the i-th target time domain estimated using the Gaussian kernel, and the relative starting position of the i-th target time domain is The relative end position is S v is the frequency estimate of video q.

[0048] Step 4) Calculate the frequency estimation value of all videos and convert the frequency estimation value S v The lowest part of the videos is combined into the OOD-test set, and the remaining videos are randomly divided into new training sets and test sets in a set ratio.

[0049] In the embodiment of the present invention, all videos are sorted according to the evaluation value, and the frequency evaluation value S v The lowest 5% of the videos in the video set are divided into the OOD-test set, and the remaining videos in the video set are randomly divided into new training sets and test sets in proportion. The OOD-test set obtained by the division is used as the validation set.

[0050] Step 5) Retrain the VCMR model on the re-segmented training set, and perform testing and verification to detect the performance difference between the VCMR model on the original test set and on the OOD-test set, evaluate the degree of VCMR model's incorrect dependence on the dataset bias, and also indicate the model's generalization ability on datasets with different distributions. When it is necessary to alleviate and eliminate the incorrect bias dependence of the VCMR model.

[0051] Step 6) Design and implement a debiasing model based on the constructed SCM model and the analyzed reasoning path correction scheme. The specific steps are shown in 6a to 6d below.

[0052] Step 6a) Use a linear layer to transform the initial obfuscated video features v 0 Decomposed into video content features and video location information:

[0053] c v =g c (v 0 ),e v =g l (v 0 )

[0054] Among them, gc and gl represent two different linear layers.

[0055] Step 6b) Define the calculation method of the do operation:

[0056]

[0057] Where q is a text feature variable, v is a video content feature variable, and l is a position feature variable; the left side of the equation is the ML or VR matching score calculated by the do operation; f*(q,v,l) means using the linear method f* to calculate the ML score or VR score for the combination (q,v,l); E l is the expected value; E l (f*(q,v,l)) is the expectation of the ML score or VR score computed using f*(q,v,l) given the desired location l.

[0058] Step 6c) Reconstruct the calculation method of the reasoning process:

[0059]

[0060] Among them, φ * It is the max operation in VR tasks or the one-dimensional convolution operation for calculating the start and end point scores in ML tasks; q * Query text features used for different tasks; is the reorganized video feature; the position information is the interference factor in the original video feature, then represents the expectation of the interference factor l, specifically It is the expectation of the improved position feature under the expected position l. 1 is the weight parameter, h c (l) represents the position feature improved by integrating video content information.

[0061] Step 6d) Definition The approximate calculation of is as follows:

[0062]

[0063] in, represents the expectation of the improved position feature calculated approximately; Q l ,K c ,V l is the intermediate parameter, Q l =W 2 l v , K c =W 3 c v , V l =W 4 l v , Q l ,K c ,V l As the query matrix, key matrix, and value matrix input of point attention, W 2 , W 3 , W 4is the weight parameter; d is the feature dimension of the query matrix; P(l) represents the probability of the expected position feature l, and P(l) is set to a fixed value As each time domain segment is selected with equal probability, N l is the total number of sub-segments the video contains.

[0064] like Figure 2 As shown, the segmented c v and l v Optimize the processing and v After being processed by the FFN (fully connected feedforward network) layer, it is then combined with c v Added to get optimized video content features c v and l v Input the hybrid expectation approximate calculation module to implement the approximate calculation in step 6d, purify the untangled position information, and obtain the expected value of the optimized position feature l as the subsequent input; further and Add up to get the entanglement-free video content features Perform the calculation in step 6c to obtain the ML score or VR score. The CLDM network eliminates the entanglement of content information and location information in video features through two steps. First, a video feature is extracted from its content and location information parts, and then the two parts are refined and optimized as input features of the subsequent reasoning model.

[0065] Step 7) Calculate the final query statement-video temporal segment matching score based on the ML and VR scores calculated in steps 6b to 6d, and obtain the task output s after final sorting. vcmr as follows:

[0066]

[0067] Among them, s vr is the confidence score of the match between the overall video feature v and the text feature q, For the i-th segment As a confidence score for the starting point, For the jth segment As the confidence score of the end point, α is a hyperparameter for adjusting the effects of VR and ML confidence, and is set to 20 in the embodiment of the present invention.

[0068] The method of the present invention is experimented with using the TVR dataset commonly used in the industry, and the retrieval models used are XML and ReLoNet. An experimental result is shown in Table 1.

[0069] Table 1 Comparison of experimental results

[0070]

[0071] In Table 1, VCMR refers to multi-video segment matching, and VMR refers to single video segment retrieval, and comparison is made for the two detection tasks. w / o means that the method of the present invention is not used, and w / , i.e. w / CLDM, means that the method of the present invention is used. The existing models XML and ReLoNet perform VCMR tasks and VMR tasks on the TVR dataset, and the performance is compared with and without the method of the present invention. IoU indicates whether the percentage of the length of the overlap between the target segment and the benchmark truth segment reaches the set value. Recall@k and R1, etc. indicate that the target video segment is correctly matched within the top few positions in the retrieval list. For example, R10, IoU = 0.5 means that there is a correctly matched video in the target segment ranked in the top ten, and the IoU between the matched segment and the benchmark truth value is greater than or equal to 0.5. It can be seen from the above table that the error detection accuracy and retrieval accuracy of the existing model after using the method of the present invention are improved compared with those without the method of the present invention, which shows that the method of the present invention has the effect of correcting the model error dependence on the data set with distribution changes.

[0072] In summary, the false dependency elimination scheme for video set segment retrieval proposed in the present invention can evaluate the false dependency degree of the VCMR model on the data set bias. At the same time, a method is proposed to optimize the reasoning path of the traditional VCMR model, and a model is implemented to disentangle the confusion of the interference factors of the input features, thereby completing the elimination of the dependency on the time domain positioning bias and improving the generalization ability of the model on data sets with different distributions.

Claims

1. A method for eliminating false dependencies in video set segment retrieval, characterized in that: The steps include: Step 1, re-segment the video data set currently used by the VCMR model, and evaluate the bias dependence of the VCMR model on the target time domain position in the video data set through out-of-distribution testing on the re-segmented data set; VCMR stands for Video Collection Clip Retrieval; Step 2: When it is necessary to alleviate and eliminate the erroneous paranoid dependence of the VCMR model, the reasoning path of the VCMR model is analyzed by constructing a structured attribution model, and the reasoning path of the VCMR model is adjusted through attribution intervention, including: (1) The initial confused video feature v0 is passed through two linear layers g c and g l Split into features c representing content v and the feature l representing the location v ; Among them, two linear layers g are trained c and g l Including: position features l v Set a non-learnable positional encoding p and train a linear layer g c Make l v Similar to the positional encoding p, training the linear layer g l Make the content feature c v Close to the initial confused video features, train the linear layer g c and g l Make the content feature c v and position feature l v Try to stay away; (2) Use the do operation to intervene in the query text and video content features, reconstruct the reasoning path of the VCMR model, and alleviate and eliminate the VCMR model's incorrect biased dependence on the target temporal position; The calculation method for setting the do operation is as follows: The calculation method of the reconstructed reasoning process is as follows: Among them, the conditional probability P(X|do(Q,V)) observes the output X of the VCMR model after executing the do operation do(Q,V), X is the matching score ML of the query text Q and the video V in different sub-segments, or the matching score VR of the query text Q and the video V; the three variables q, v and l represent the text features, video content features, and position features respectively; f*(q,v,l) represents the use of the linear method f* to calculate the score ML or VR for the combination (q,v,l); E l (f*(q,v,l)) is the expectation of ML or VR at the specified expected position l; φ * It is the max operation in VR computing tasks or the one-dimensional convolution operation for calculating the start and end point scores in ML computing tasks; q * Query text features used for different tasks; is the reorganized video feature, W1 is the weight parameter; h c (l) is a location feature improved by integrating video content information; is to calculate h at the specified position l c (l) expectations; The final query statement-video time domain segment matching score is calculated based on the ML and VR scores observed by the do operation, and the task output result is obtained after final sorting.

2. The method according to claim 1, characterized in that The step 1 is to re-split the video data set of the current application by estimating the frequency of the video samples, construct an out-of-distribution test OOD-test set, and re-split the test set and the training set; When calculating the frequency estimate of a video, the probability density of the target time sequence position needs to be calculated first, as follows: Determine the relative positions of the start and end positions of the target temporal position corresponding to the query statement in the entire video, use the Gaussian kernel for kernel density estimation, and obtain the probability density of the target temporal position as follows: in, is the probability density of the target time series position relative to the starting position x∈[0,y] and the ending position y∈(x,1], n is the number of samples, h is the bandwidth; K(·) is the Gaussian kernel function, and u is the parameter; represents the sample standard deviation; x i and i is the relative start position and relative end position of the target segment of the i-th sample; Then calculate the frequency estimate S of each video q in the current dataset v as follows: Among them, N q represents the number of query-target temporal positions contained in the video, where the relative starting position of the i-th target temporal position is The relative end position is is the probability density of the i-th target time series position estimated using the Gaussian kernel.

3. The method according to claim 1 or 2, characterized in that: The step 1 comprises the following steps: Step 1.1) The training set, test set, and validation set of the VCMR model currently applied are merged into one video set, and the target temporal position of the reference truth value in the merged video set is normalized by the total video duration; Step 1.2) Use the Scott bandwidth method to estimate the bandwidth h, and use the Gaussian kernel to estimate the probability density of the target time series position; Step 1.3) Taking the video as a unit, the frequency estimation value of each video in the merged video set is calculated according to the probability density of the target temporal position contained therein; Step 1.4) Split the x% of videos with the lowest frequency estimation values ​​in the merged video set into the OOD-test set, and randomly split the remaining videos in the merged video set into new training sets and test sets in proportion; Step 1.5) Use the re-segmented training set, OOD-test set and test set to train and test the VCMR model, and compare the performance difference of the VCMR model on the original test set and on the OOD-test set. When the performance of the VCMR model on the reconstructed dataset decreases and the degree of decrease is higher than the preset threshold, it means that the VCMR model needs to alleviate and eliminate the error bias dependence of the video dataset in the current application.

4. The method according to claim 3, characterized in that In step 1.4, the 5% videos with the lowest frequency estimation values ​​in the merged video set are selected as the OOD-test set.

5. The method according to claim 1, characterized in that In step 2, when training two linear layers, the loss function is used Make l v Similar to p, using loss function Make the content feature c v Close to the initial confused video features, use the loss function L indep Make the content feature c v and position feature l v Try to stay away; the loss function is as follows: Among them, KL is to find the KL divergence, and Respectively represent c v Perform softmax operation with v0.

6. The method according to claim 1 or 5, characterized in that: The step 2 comprises the following steps: Step 2.1) Use a linear layer to decompose the initial confused video feature v0 into video content features c v and video position feature l v , which is expressed as follows: Among them, g c and g l Represents different linear layers; Step 2.2) Execute the do operation to reconstruct the reasoning process, where The approximate calculation is as follows in, Represents an approximate calculation Intermediate parameter Q l =W2l v , K c =W3c v , V l =W4l v , W2, W3, W4 are weight parameters; P(l) represents the probability of the expected position feature l, and P(l) is set to a fixed value N l is the total number of sub-segments of the video; Step 2.3) Calculate the final query statement-video time domain segment matching score based on the ML and VR scores obtained by performing the do operation in step 2.2, and obtain the task output s after the final sorting. vcmr as follows: Among them, s vr is the confidence score of the match between the overall video feature v and the text feature q, For the i-th segment As a confidence score for the starting point, For the jth segment As the confidence score of the end point, α is a hyperparameter that adjusts the effects of VR and ML confidence.

Citation Information

Patent Citations

  • Depolarization sequence recommendation method and system for interest separation and crowd representation

    CN118247003A

  • Project recommendation method and system based on causal popularity depolarization

    CN118312653A