A video text retrieval method, device, system, and storage medium

By building a video encoder and visual semantic supervised encoder, combining the data processing of the training set and test set, the problem of inaccurate semantic alignment in cross-modal retrieval of video and text is solved, and more efficient video text retrieval effect and model stability are achieved.

CN115757873BActive Publication Date: 2025-07-18GUILIN UNIV OF ELECTRONIC TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211477636.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-23
Publication Date
2025-07-18
Estimated Expiration
2042-11-23

AI Technical Summary

Technical Problem

In the prior art, the cross-modal retrieval method between video and text ignores local semantics, resulting in inaccurate semantic alignment and ineffective in improving the ability to distinguish fine-grained semantic differences.

Method used

By randomly dividing video and natural language text descriptions into training sets and test sets, a video encoder and visual semantic supervised encoder are constructed, and the video encoder is trained using visual semantic supervised encoder. Combined with loss function analysis and parameter updates, the updated encoder is obtained to realize video text retrieval.

Benefits of technology

It achieves more precise semantic alignment, improves video text retrieval effect, and improves the reliability and stability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115757873B_ABST
    Figure CN115757873B_ABST
Patent Text Reader

Abstract

The present invention provides a video text retrieval method, device, system and storage medium, belonging to the field of video processing. The method includes: randomly dividing a video into a training set and a test set; preprocessing the video and a natural language text description to obtain a target video frame block sequence; constructing a video encoder and a visual semantic supervision encoder, and training the video encoder by using the visual semantic supervision encoder and the target video frame block sequence to obtain a trained video encoder and a video text distance. The present invention ensures the high efficiency of the encoder, can effectively mine the spatio-temporal information of video data and the context information of text data, realizes more accurate semantic alignment, can effectively improve the effect of video text retrieval, and has a certain generalization ability, improving the reliability and stability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to the technical field of video processing, and particularly relates to a video text retrieval method, device, system and storage medium. Background Art

[0002] Due to the rapid emergence of videos on the Internet, cross-modal retrieval between videos and texts has attracted more and more attention. In the mainstream methods, although the two separate encoder architectures maintain the efficiency of retrieval, they ignore many important local semantics. At the same time, a simple joint embedding space is not sufficient to achieve accurate semantic alignment, nor can it effectively improve the ability to distinguish fine-grained semantic differences. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a video text retrieval method, device, system and storage medium in view of the deficiencies of the prior art.

[0004] The technical solution of the present invention for solving the above technical problem is as follows: A video text retrieval method includes the following steps:

[0005] Import a plurality of videos and a plurality of natural language text descriptions corresponding to each of the videos one by one, and randomly divide all the videos into a training set and a test set;

[0006] Preprocess each of the videos and the corresponding natural language text descriptions in the training set to obtain a plurality of target video frame block sequences for each of the videos in the training set;

[0007] Construct a video encoder and a visual semantic supervision encoder, and use the visual semantic supervision encoder and the plurality of target video frame block sequences of each of the videos in the training set to train the video encoder to obtain a trained video encoder and the video text distances of each of the target video frame block sequences;

[0008] Use the trained video encoder to encode the video text distances of each of the target video frame block sequences to obtain the video features of each of the target video frame block sequences;

[0009] Use a text encoder to encode each of the target video frame block sequences to obtain the text features of each of the target video frame block sequences;

[0010] Analyze the loss function respectively according to the video features and text features of each of the target video frame block sequences to obtain a plurality of loss functions for each of the target video frame block sequences;

[0011] Update the parameters of the visual semantic supervision encoder and the trained video encoder respectively according to the multiple loss functions of each of the target video frame block sequences to obtain an updated visual semantic supervision encoder and an updated video encoder;

[0012] Use the updated visual semantic supervision encoder and the updated video encoder to perform video text retrieval processing on the test set to obtain video text retrieval results.

[0013] Another technical solution of the present invention to solve the above technical problems is as follows: A video text retrieval device, comprising:

[0014] A partitioning module, configured to import a plurality of videos and a plurality of natural language text descriptions corresponding to each of the videos one by one, and randomly partition all the videos into a training set and a test set;

[0015] A preprocessing module, configured to preprocess each of the videos and the corresponding natural language text descriptions in the training set to obtain a plurality of target video frame block sequences of each of the videos in the training set;

[0016] A training module, configured to construct a video encoder and a visual semantic supervision encoder, and use the visual semantic supervision encoder and a plurality of target video frame block sequences of each of the videos in the training set to train the video encoder to obtain a trained video encoder and the video text distances of each of the target video frame block sequences;

[0017] A video encoding module, configured to encode the video text distances of each of the target video frame block sequences respectively by using the trained video encoder to obtain video features of each of the target video frame block sequences;

[0018] A text encoding module, configured to encode each of the target video frame block sequences respectively by using a text encoder to obtain text features of each of the target video frame block sequences;

[0019] A loss function analysis module, configured to perform loss function analysis respectively according to the video features and text features of each of the target video frame block sequences to obtain a plurality of loss functions of each of the target video frame block sequences;

[0020] A parameter update module, configured to update the parameters of the visual semantic supervision encoder and the trained video encoder respectively according to the multiple loss functions of each of the target video frame block sequences to obtain an updated visual semantic supervision encoder and an updated video encoder;

[0021] A retrieval result acquisition module, configured to perform video text retrieval processing on the test set by using the updated visual semantic supervision encoder and the updated video encoder, so as to obtain a video text retrieval result.

[0022] Based on the above video text retrieval method, the present invention further provides a video text retrieval system.

[0023] Another technical solution for the present invention to solve the above technical problems is as follows: A video text retrieval system includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above video text retrieval method is implemented.

[0024] Based on the above video text retrieval method, the present invention further provides a computer-readable storage medium.

[0025] Another technical solution for the present invention to solve the above technical problems is as follows: A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the above video text retrieval method is implemented.

[0026] The beneficial effects of the present invention are as follows: By randomly dividing the video into a training set and a test set, preprocessing the video and the natural language text description to obtain a target video frame block sequence, using the visual semantic supervision encoder and the target video frame block sequence to train the video encoder to obtain a trained video encoder and a video text distance, using the trained video encoder to encode the video text distance to obtain video features, using the text encoder to encode the target video frame block sequence to obtain text features, analyzing the loss function according to the video features and the text features to obtain a loss function, updating the parameters of the visual semantic supervision encoder and the trained video encoder according to the loss function to obtain an updated visual semantic supervision encoder and an updated video encoder, and performing video text retrieval processing on the test set by using the updated visual semantic supervision encoder and the updated video encoder to obtain a video text retrieval result, which ensures the high efficiency of the encoder, can effectively extract the spatio-temporal information of the video data and the context information of the text data, realizes more accurate semantic alignment, can effectively improve the effect of video text retrieval, and has a certain generalization ability, improving the reliability and stability of the model. Description of the Drawings

[0027] Figure 1 It is a schematic flowchart of a video text retrieval method provided by an embodiment of the present invention;

[0028] Figure 2 It is a block diagram of modules of a video text retrieval device provided by an embodiment of the present invention. Specific Embodiments

[0029] The principles and features of the present invention will be described below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0030] Figure 1 It is a schematic flowchart of a video text retrieval method provided by an embodiment of the present invention.

[0031] As Figure 1 shown, a video text retrieval method includes the following steps:

[0032] Import multiple videos and multiple natural language text descriptions corresponding to each of the videos one by one, and randomly divide all the videos into a training set and a test set;

[0033] Preprocess each of the videos and the corresponding natural language text descriptions in the training set to obtain multiple target video frame block sequences for each of the videos in the training set;

[0034] Construct a video encoder and a visual semantic supervision encoder, and use the visual semantic supervision encoder and the multiple target video frame block sequences of each of the videos in the training set to train the video encoder to obtain a trained video encoder and the video text distances of each of the target video frame block sequences;

[0035] Use the trained video encoder to encode the video text distances of each of the target video frame block sequences to obtain the video features of each of the target video frame block sequences;

[0036] Use a text encoder to encode each of the target video frame block sequences to obtain the text features of each of the target video frame block sequences;

[0037] Conduct loss function analysis based on the video features and text features of each of the target video frame block sequences respectively to obtain multiple loss functions for each of the target video frame block sequences;

[0038] Update the parameters of the visual semantic supervision encoder and the trained video encoder respectively according to the multiple loss functions of each of the target video frame block sequences to obtain an updated visual semantic supervision encoder and an updated video encoder;

[0039] Use the updated visual semantic supervision encoder and the updated video encoder to perform video text retrieval processing on the test set to obtain a video text retrieval result.

[0040] It should be understood that the original data (i.e., the multiple videos and the multiple natural language text descriptions corresponding to each of the videos one by one) is collected and these data are divided into a test set and a training set.

[0041] Specifically, the original data requires at least 10,000 of the videos and the corresponding natural language text descriptions, and there are at least 20 sentences in the natural language text description corresponding to each video; the original data (i.e., the multiple videos and the multiple natural language text descriptions corresponding to each of the videos one by one) is split, 9,000 videos are used for the training process and 1,000 videos are used for the testing process.

[0042] Specifically, the video encoder and the visual-semantic supervision encoder have exactly the same architecture and are composed of a stack of spatio-temporal self-attention modules (i.e., the Bert encoder).

[0043] It should be understood that the test set is input into the trained model (i.e., the updated visual-semantic supervision encoder and the updated video encoder) to implement video-text retrieval.

[0044] It should be understood that during the testing process, the visual-semantic supervision encoder (i.e., the updated visual-semantic supervision encoder) will be frozen.

[0045] It should be understood that the video coding module (i.e., the visual-semantic supervision encoder and the video encoder) is used to encode the spatio-temporal information of the video frame block sequence (i.e., the target video frame block sequence) to obtain the global event features, local entity features, and action features of the video (i.e., the video features).

[0046] Specifically, the video encoder (i.e., the trained video encoder) obtains the global features V g 、W e and W a , of the video through g 、local entity features V e and action features V a (i.e., the video features). Among them,

[0047] V g =V*W g

[0048] V e =V*W e

[0049] V a =V*W a

[0050] V g ∈{v g1 ,vg2 ,…,v gk}

[0051] V e ∈{v e1 ,v e2 ,…,v ek}

[0052] V a ∈{v a1 ,v a2 ,…,v ak}

[0053] Here, k refers to the number of frames of a video.

[0054] It should be understood that by updating the visual semantic supervision encoder and calculating the loss function, the entire model (i.e., the updated visual semantic supervision encoder and the updated video encoder) is trained on the test set data (i.e., the test set).

[0055] In the above embodiments, by randomly dividing the video into a training set and a test set, preprocessing the video and the natural language text description to obtain a target video frame block sequence, training the video encoder using the visual semantic supervision encoder and the target video frame block sequence to obtain a trained video encoder and a video-text distance, encoding the video-text distance using the trained video encoder to obtain video features, encoding the target video frame block sequence using the text encoder to obtain text features, analyzing the loss function based on the video features and text features to obtain the loss function, updating the parameters of the visual semantic supervision encoder and the trained video encoder according to the loss function to obtain an updated visual semantic supervision encoder and an updated video encoder, and performing video-text retrieval processing on the test set using the updated visual semantic supervision encoder and the updated video encoder to obtain a video-text retrieval result, it ensures the high efficiency of the encoder while effectively mining the spatio-temporal information of video data and the context information of text data, achieving more accurate semantic alignment, effectively improving the effect of video-text retrieval, and having a certain generalization ability, improving the reliability and stability of the model.

[0056] Optionally, as an embodiment of the present invention, the process of respectively preprocessing each of the videos and the corresponding natural language text descriptions in the training set to obtain multiple target video frame block sequences for each of the videos includes:

[0057] Using the natural language text description corresponding to each of the videos in the training set as a preset text description as a dividing line, respectively segmenting each of the videos in the training set to obtain multiple to-be-mapped video frame block sequences for each of the videos in the training set;

[0058] Map the multiple video frame block sequences of each video in the training set respectively to obtain the multiple target video frame block sequences of each video.

[0059] It should be understood that the videos in the training set are segmented and projected into a series of video frame block sequences (i.e., the target video frame block sequences).

[0060] In the above embodiments, the multiple target video frame block sequences of each video in the training set are preprocessed respectively, laying a foundation for subsequent data processing. While ensuring the high efficiency of the encoder, it can effectively mine the spatio-temporal information of video data and the context information of text data, achieving more accurate semantic alignment.

[0061] Optionally, as an embodiment of the present invention, the process of training the video encoder by using the visual-semantic supervision encoder and the multiple target video frame block sequences of each video in the training set to obtain the trained video encoder and the video-text distances of each target video frame block sequence includes:

[0062] Mask each target video frame block sequence of each video in the training set respectively to obtain the masked video frame block sequences of each target video frame block sequence;

[0063] Perform positional encoding on the masked video frame block sequences of each target video frame block sequence respectively to obtain the encoded video frame block sequences of each target video frame block sequence;

[0064] Use the video encoder to encode the encoded video frame block sequences of each target video frame block sequence respectively to obtain the first encoded video features of each target video frame block sequence;

[0065] Use the visual-semantic supervision encoder to encode each target video frame block sequence of each video respectively to obtain the second encoded video features of each target video frame block sequence;

[0066] Based on the first formula, calculate the video-text distances of each target video frame block sequence according to the first encoded video features and the second encoded video features of each target video frame block sequence. The first formula is:

[0067] L = ||V - Q||,

[0068] where L is the video-text distance, V is the first encoded video feature, and Q is the second encoded video feature;

[0069] Update the parameters of the video encoder according to the video text distances of all the target video frame blocks sequences to obtain a trained video encoder.

[0070] It should be understood that a part of the video frame block sequence (i.e., the target video frame block sequence) is masked in the spatial and temporal dimensions.

[0071] It should be understood that positional embedding is performed on the video frame block sequence (i.e., the masked video frame block sequence) in the spatial and temporal dimensions to obtain the input video frame block sequence of the video (i.e., the encoded video frame block sequence).

[0072] Specifically, the input token sequence (i.e., the encoded video frame block sequence) is input into the video encoder for encoding, and the video encoder automatically learns the features V of the masked video frame blocks (i.e., the first encoded video features) through the visible video frame blocks of the spatial and temporal neighbors.

[0073] Specifically, the unmasked original video frame block sequence (i.e., the target video frame block sequence) is input into the visual semantic supervision encoder to obtain the features Q of the masked video blocks (i.e., the second encoded video features).

[0074] It should be understood that by minimizing the distance between the above-mentioned V and Q, visual semantic supervision of the video encoder is realized at the video frame block level, so as to obtain the spatio-temporal information of the video frame.

[0075] Specifically, taking minimizing the distance between the above-mentioned V and Q as the objective, the encoder model is optimized. The specific formula is as follows:

[0076] L = ||V - Q||.

[0077] In the above embodiment, the visual semantic supervision encoder and the target video frame block sequence are used to train the video encoder to obtain the trained video encoder and the video text distance, realizing the visual semantic supervision of the video encoder, so as to obtain the spatio-temporal information of the video frame.

[0078] Optionally, as an embodiment of the present invention, the text encoder includes a multi-layer bidirectional transformer encoder;

[0079] The process of encoding the video text distances of each of the target video frame block sequences by using the trained video encoder to obtain the video features of each of the target video frame block sequences includes:

[0080] Use the graph inference mechanism algorithm and the multi-layer bidirectional Transformer encoder to encode each of the target video frame block sequences, respectively, to obtain the text global feature, text entity feature, and text action feature of each of the target video frame block sequences;

[0081] The video feature of the target video frame block sequence includes the text global feature, text entity feature, and text action feature of the target video frame block sequence.

[0082] It should be understood that the context information encoding of the text data (i.e., the target video frame block sequence) by using the text encoder obtains the global event feature (i.e., the text global feature), local entity feature (i.e., the text entity feature), and action feature (i.e., the text action feature) of the text.

[0083] Specifically, the text encoder adopts a multi-layer bidirectional Transformer structure encoder to obtain a text feature representation with context information, and obtains the global feature C g (i.e., the text global feature), local entity feature C e (i.e., the text entity feature), and action feature C a (i.e., the text action feature) of the text feature representation. Among them,

[0084] C e ∈{c e1 , c e2 , …, c ek},

[0085] C a ∈{c a1 , c a2 , …, c ak}.

[0086] In the above embodiment, by using the trained video encoder to encode the video text distance of each target video frame block sequence respectively to obtain the video feature of each target video frame block sequence, a text feature representation with context information can be obtained.

[0087] Optionally, as an embodiment of the present invention, the video feature includes multiple video sub-features, and the text feature includes multiple text sub-features;

[0088] The process of respectively analyzing the loss function according to the video feature and text feature of each target video frame block sequence to obtain multiple loss functions of each target video frame block sequence includes:

[0089] Analyze the video-text similarity based on each of the video sub-features and each of the text sub-features of each of the target video frame sequences, and obtain multiple video-text similarities for each of the target video frame sequences;

[0090] Based on the second formula, calculate the loss function according to the multiple video-text similarities of each of the target video frame sequences, and obtain multiple loss functions for each of the target video frame sequences. The second formula is:

[0091] Loss(v a ,v a ,c b ,c b ,α)=[β+S(v a ,c b )-S(v a ,c a )] + +[β+S(v b ,c a )-

[0092] S(v a ,c a )] + ,

[0093] where Loss(v a ,v a ,c b ,c b ,α) is the loss function of the a-th video sub-feature and the b-th text sub-feature, S(v a ,c b ) is the video-text similarity of the a-th video sub-feature and the b-th text sub-feature, S(v a ,c a ) is the video-text similarity of the a-th video sub-feature and the a-th text sub-feature, S(v b ,c a ) is the video-text similarity of the b-th video sub-feature and the a-th text sub-feature, β is a preset hyperparameter, a ∈ [1, i], b ∈ [1, j], i is the number of video sub-features, j is the number of text sub-features, [.] + is max(·, 0).

[0094] It should be understood that v a can be the a-th video sub-feature, c a can be the a-th text sub-feature, v b can be the b-th video sub-feature, c b can be the b-th text sub-feature.

[0095] It should be understood that the contrastive ranking loss is used as the training objective to maximize the similarity between video-text combinations in positive samples and minimize the similarity between combinations in negative samples.

[0096] It should be understood that the distance between the randomly sampled negative samples and the positive samples should be greater than a fixed margin β, where β is a preset hyperparameter.

[0097] It should be understood that i is not equal to j, [...] + denotes max(·, 0), that is, the model is updated by maximizing the similarity between positive video-text combinations and minimizing the similarity between negative combinations.

[0098] In the above embodiments, the loss function is obtained by analyzing the video features and text features, which can maximize the similarity between video-text combinations in positive samples and minimize the similarity between combinations in negative samples, effectively improving the effect of video-text retrieval, and having a certain generalization ability, thus improving the reliability and stability of the retrieval model.

[0099] Optionally, as an embodiment of the present invention, the video sub-features include video global sub-features, video entity sub-features, and video action sub-features, and the text sub-features include text global sub-features, text entity sub-features, and text action sub-features;

[0100] The process of analyzing the video-text similarity according to each of the video sub-features and each of the text sub-features of each of the target video frame sequences to obtain multiple video-text similarities of each of the target video frame sequences includes:

[0101] Based on the third formula, calculate the global matching scores according to each of the video global sub-features and each of the text global sub-features of each of the target video frame sequences to obtain multiple global matching scores of each of the target video frame sequences. The third formula is:

[0102] S g(i,j) = cos(v g,i , c g,j ),

[0103] where S g(i,j) is the global matching score between the i-th video global sub-feature and the j-th text global sub-feature, v g,i is the i-th video global sub-feature, and c g,j is the j-th text global sub-feature;

[0104] Based on the fourth formula, entity matching scores are calculated according to the respective video entity sub - features and text entity sub - features of each of the target video frame block sequences, obtaining multiple entity matching scores for each of the target video frame block sequences. The fourth formula is:

[0105] S e(i,j) =cos(v e,i ,c e,j ),

[0106] where S e(i,j) is the entity matching score between the i - th video entity sub - feature and the j - th text entity sub - feature, v e,i is the i - th video entity sub - feature, and c e,j is the j - th text entity sub - feature;

[0107] Based on the fifth formula, action matching scores are calculated according to the respective video action sub - features and text action sub - features of each of the target video frame block sequences, obtaining multiple action matching scores for each of the target video frame block sequences. The fifth formula is:

[0108] S a(i,j) =cos(v a,i ,c a,j ),

[0109] where S a(i,j) is the action matching score between the i - th video action sub - feature and the j - th text action sub - feature, v a,i is the i - th video action sub - feature, and c a,j is the j - th text action sub - feature;

[0110] Normalization processing is respectively performed on the respective global matching scores, entity matching scores, and action matching scores of each of the target video frame block sequences, corresponding to obtaining global attention weight parameters for each of the global matching scores, entity attention weight parameters for each of the entity matching scores, and action attention weight parameters for each of the action matching scores;

[0111] Based on the sixth formula, target global matching scores are calculated according to the respective global matching scores of each of the target video frame block sequences and the global attention weight parameters of each of the global matching scores, obtaining multiple target global matching scores for each of the target video frame block sequences. The sixth formula is:

[0112] S g,i,j =r g(i,j) S g(i,j) ,

[0113] where S g,i,jis the target global matching score between the i-th video global sub-feature and the j-th text global sub-feature, S g(i,j) is the global matching score between the i-th video global sub-feature and the j-th text global sub-feature, r g(i,j) is the global attention weight parameter between the i-th video global sub-feature and the j-th text global sub-feature;

[0114] Based on the seventh formula, calculate the target entity matching scores according to the entity matching scores of each of the target video frame block sequences and the entity attention weight parameters of each of the entity matching scores, to obtain multiple target entity matching scores of each of the target video frame block sequences. The seventh formula is:

[0115] S e,i,j = r e(i,j) S e(i,j) ,

[0116] wherein, S e,i,j is the target entity matching score between the i-th video entity sub-feature and the j-th text entity sub-feature, S e(i,j) is the entity matching score between the i-th video entity sub-feature and the j-th text entity sub-feature, r e(i,j) is the entity attention weight parameter between the i-th video entity sub-feature and the j-th text entity sub-feature;

[0117] Based on the eighth formula, calculate the target action matching scores according to the action matching scores of each of the target video frame block sequences and the action attention weight parameters of each of the action matching scores, to obtain multiple target action matching scores of each of the target video frame block sequences. The eighth formula is:

[0118] S a,i,j = r a(i,j) S a(i,j) ,

[0119] wherein, S a,i,j is the target action matching score between the i-th video action sub-feature and the j-th text action sub-feature, S a(i,j) is the action matching score between the i-th video action sub-feature and the j-th text action sub-feature, r a(i,j) is the action attention weight parameter between the i-th video action sub-feature and the j-th text action sub-feature;

[0120] Based on the ninth formula, calculate the video-text similarity according to the target global matching scores, the target entity matching scores, and the target action matching scores of each of the target video frame block sequences, to obtain multiple video-text similarities of each of the target video frame block sequences. The ninth formula is:

[0121] S(v i ,c j ) = (S g,i,j + S e,i,j + S a,i,j ) / 3,

[0122] where S(v i ,c j ) is the video - text similarity between the i - th video action sub - feature and the j - th text action sub - feature, S g,i,j is the target global matching score between the i - th video global sub - feature and the j - th text global sub - feature, S e,i,j is the target entity matching score between the i - th video entity sub - feature and the j - th text entity sub - feature, S a,i,j is the target action matching score between the i - th video action sub - feature and the j - th text action sub - feature.

[0123] It should be understood that the above - mentioned video and text global features and local features are projected onto an aligned common network to calculate the similarity between the video and text features (i.e., the video - text similarity).

[0124] Specifically, the cosine similarity is used to calculate the similarity between the global event features (i.e., the video global sub - feature and the text global sub - feature), and the obtained global matching score has the following formula:

[0125] S g(i,j) = cos(v g,i ,c g,j ),

[0126] The cosine similarity is used to calculate the similarity between the local entity features (i.e., the video entity sub - feature and the text entity sub - feature), and the obtained matching score of the local entity features (i.e., the entity matching score) has the following formula:

[0127] S e(i,j) = cos(v e,i ,c e,j ),

[0128] The cosine similarity is used to calculate the similarity between the local action features (i.e., the video action sub - feature and the text action sub - feature), and the obtained matching score of the local action features (i.e., the action matching score) has the following formula:

[0129] S a(i,j) = cos(v a,i ,c a,j ).

[0130] Specifically, S g(i,j) , S e(i,j) and S a(i,j)Normalize them respectively to obtain the attention weight parameter r e(i,j) (i.e., the entity attention weight parameter), r g(i,j) (i.e., the global attention weight parameter), and r a(i,j) (i.e., the action attention weight parameter), and dynamically realize the semantic alignment between text and video on local features, where

[0131]

[0132]

[0133]

[0134] It should be understood that the matching scores S e(i,j) , S g(i,j) , and S a(i,j) are weighted and averaged to obtain the final matching scores (i.e., the target global matching score, the target action matching score, and the target action matching score), and the formula is as follows:

[0135] S g,i,j = r g(i,j) S g(i,j) ,

[0136] S e,i,j = r e(i,j) S e(i,j) ,

[0137] S a,i,j = r a(i,j) S a(i,j) .

[0138] It should be understood that during training, the average value S of the feature matching scores at three levels (i.e., the target global matching score, the target action matching score, and the target action matching score) is regarded as the final video-text similarity (i.e., the video-text similarity).

[0139] In the above embodiments, analyzing the video-text similarity according to the video sub-features and the text sub-features can effectively improve the effect of video-text retrieval, and has a certain generalization ability, improving the reliability and stability of the retrieval model.

[0140] Optionally, as an embodiment of the present invention, the process of respectively updating the parameters of the visual semantic supervision encoder and the trained video encoder according to the multiple loss functions of each of the target video frame block sequences to obtain the updated visual semantic supervision encoder and the updated video encoder includes:

[0141] Based on the exponential moving average mechanism, and using the trained video encoder to update the parameters of the visual semantic supervision encoder, an updated visual semantic supervision encoder is obtained;

[0142] According to multiple loss functions of each of the target video frame sequences, the parameters of the trained video encoder are updated to obtain an updated video encoder.

[0143] It should be understood that during training, the video encoder of the previous epoch (the trained video encoder) is used as the visual semantic supervision encoder to update the visual semantic supervision encoder.

[0144] Specifically, based on the exponential moving average (EMA) mechanism, the encoder is frozen in one epoch and its parameters are updated in the k-th epoch. This mechanism is expressed as

[0145] {θ q} k =β{θ q} k +(1 - β){θ v} k-1

[0146] where {θ v} k-1 represents the parameters of the video encoder at the end of the (k - 1)-th epoch, and {θ q} k-1 represents the parameters of the video supervision encoder at the end of the (k - 1)-th epoch.

[0147] In the above embodiments, the parameters of the visual semantic supervision encoder and the trained video encoder are updated respectively according to the loss function to obtain the updated visual semantic supervision encoder and the updated video encoder. While ensuring the high efficiency of the encoder, it can effectively extract the spatio-temporal information of video data and the context information of text data, achieve more accurate semantic alignment, effectively improve the video-text retrieval effect, and have a certain generalization ability, improving the reliability and stability of the retrieval model.

[0148] Optionally, as another embodiment of the present invention, the present invention collects original data and divides these data into a test set and a training set; divides the videos in the test set and projects them into a series of video frame block sequences; uses a video encoding module to encode the spatio-temporal information of the video frame block sequences to obtain the global event features, local entity features, and action features of the video; uses a text encoder to encode the context information of the text data to obtain the global event features, local entity features, and action features of the text; projects the above-mentioned global and local features of the video and text into an aligned common network to calculate the similarity between the video and text features; trains the entire model on the test set data by updating the visual semantic supervision encoder and calculating the loss function; inputs the test set into the trained model to achieve video-text retrieval. The present invention ensures the high efficiency of the encoder while effectively mining the spatio-temporal information of video data and the context information of text data, achieving more accurate semantic alignment, effectively improving the video-text retrieval effect, and having a certain generalization ability, improving the reliability and stability of the retrieval model.

[0149] Figure 2 It is a module block diagram of a video-text retrieval device provided by an embodiment of the present invention.

[0150] Optionally, as another embodiment of the present invention, as Figure 2 shown, a video-text retrieval device includes:

[0151] A division module, configured to import a plurality of videos and a plurality of natural language text descriptions corresponding to each of the videos one by one, and randomly divide all the videos into a training set and a test set;

[0152] A preprocessing module, configured to preprocess each of the videos and the corresponding natural language text descriptions in the training set to obtain a plurality of target video frame block sequences of each of the videos in the training set;

[0153] A training module, configured to construct a video encoder and a visual semantic supervision encoder, and use the visual semantic supervision encoder and the plurality of target video frame block sequences of each of the videos in the training set to train the video encoder to obtain the trained video encoder and the video-text distances of each of the target video frame block sequences;

[0154] A video encoding module, configured to use the trained video encoder to encode the video-text distances of each of the target video frame block sequences to obtain the video features of each of the target video frame block sequences;

[0155] A text encoding module, configured to encode each of the target video frame block sequences respectively by using a text encoder to obtain text features of each of the target video frame block sequences;

[0156] A loss function analysis module, configured to perform loss function analysis respectively according to the video features and text features of each of the target video frame block sequences to obtain multiple loss functions of each of the target video frame block sequences;

[0157] A parameter update module, configured to update parameters of the visual semantic supervision encoder and the trained video encoder respectively according to the multiple loss functions of each of the target video frame block sequences to obtain an updated visual semantic supervision encoder and an updated video encoder;

[0158] A retrieval result obtaining module, configured to perform video text retrieval processing on the test set by using the updated visual semantic supervision encoder and the updated video encoder to obtain a video text retrieval result.

[0159] Optionally, another embodiment of the present invention provides a video text retrieval system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the video text retrieval method as described above is implemented. The system may be a computer or the like.

[0160] Optionally, another embodiment of the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the video text retrieval method as described above is implemented.

[0161] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.

[0162] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0163] In several embodiments provided by this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0164] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present invention.

[0165] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0166] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0167] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A video text retrieval method, characterized in that, It includes the following steps: Import multiple videos and multiple natural language text descriptions corresponding to each of the videos one by one, and randomly divide all the videos into a training set and a test set; Preprocess each of the videos and the corresponding natural language text descriptions in the training set to obtain multiple target video frame block sequences for each of the videos in the training set; Construct a video encoder and a visual-semantic supervision encoder, and use the visual-semantic supervision encoder and the multiple target video frame block sequences for each of the videos in the training set to train the video encoder, obtaining a trained video encoder and the video-text distances for each of the target video frame block sequences; Use the trained video encoder to encode the video-text distances for each of the target video frame block sequences to obtain the video features for each of the target video frame block sequences; Use a text encoder to encode each of the target video frame block sequences to obtain the text features for each of the target video frame block sequences; Conduct loss function analysis based on the video features and text features for each of the target video frame block sequences respectively to obtain multiple loss functions for each of the target video frame block sequences; Update the parameters of the visual-semantic supervision encoder and the trained video encoder respectively according to the multiple loss functions for each of the target video frame block sequences to obtain an updated visual-semantic supervision encoder and an updated video encoder; Use the updated visual-semantic supervision encoder and the updated video encoder to perform video-text retrieval processing on the test set to obtain video-text retrieval results.

2. The video text retrieval method according to claim 1, wherein The process of preprocessing each of the videos and the corresponding natural language text descriptions in the training set to obtain multiple target video frame block sequences for each of the videos includes: Taking the natural language text descriptions corresponding to each of the videos in the training set as preset text descriptions as dividing lines, and respectively segmenting each of the videos in the training set to obtain multiple to-be-mapped video frame block sequences for each of the videos in the training set; Map the multiple to-be-mapped video frame block sequences for each of the videos in the training set respectively to obtain multiple target video frame block sequences for each of the videos.

3. The video text retrieval method according to claim 1, wherein The process of using the visual-semantic supervision encoder and the multiple target video frame block sequences for each of the videos in the training set to train the video encoder, obtaining a trained video encoder and the video-text distances for each of the target video frame block sequences includes: Perform masking processing on each of the target video frame block sequences for each of the videos in the training set to obtain masked video frame block sequences for each of the target video frame block sequences; Perform position encoding on the masked video frame block sequences for each of the target video frame block sequences respectively to obtain encoded video frame block sequences for each of the target video frame block sequences; Use the video encoder to encode the encoded video frame blocks of each of the target video frame block sequences respectively to obtain the first encoded video features of each of the target video frame block sequences; Use the visual semantic supervision encoder to encode each of the target video frame block sequences of each of the videos respectively to obtain the second encoded video features of each of the target video frame block sequences; Based on the first formula, calculate the video text distance according to the first encoded video features and the second encoded video features of each of the target video frame block sequences to obtain the video text distance of each of the target video frame block sequences, and the first formula is: L = ||V - Q||, where L is the video text distance, V is the first encoded video feature, and Q is the second encoded video feature; Update the parameters of the video encoder according to the video text distances of all the target video frame block sequences to obtain the trained video encoder.

4. The video text retrieval method according to claim 1, wherein The text encoder includes a multi-layer bidirectional Transformer encoder; The process of using the trained video encoder to encode the video text distances of each of the target video frame block sequences respectively to obtain the video features of each of the target video frame block sequences includes: Use the graph reasoning mechanism algorithm and the multi-layer bidirectional Transformer encoder to encode each of the target video frame block sequences respectively to obtain the text global feature, text entity feature, and text action feature of each of the target video frame block sequences; The video feature of the target video frame block sequence includes the text global feature, text entity feature, and text action feature of the target video frame block sequence.

5. The video text retrieval method according to claim 1, wherein The video feature includes multiple video sub-features, and the text feature includes multiple text sub-features; The process of analyzing the loss function according to the video features and text features of each of the target video frame block sequences respectively to obtain multiple loss functions of each of the target video frame block sequences includes: Analyze the video text similarity according to each of the video sub-features and each of the text sub-features of each of the target video frame block sequences respectively to obtain multiple video text similarities of each of the target video frame block sequences; Based on the second formula, calculate the loss function according to the multiple video text similarities of each of the target video frame block sequences to obtain multiple loss functions of each of the target video frame block sequences, and the second formula is: Loss(v a ,v a ,c b ,c b ,α) = [β + S(v a ,c b ) - S(v a ,c a )] + + [β + S(v b ,c a ) - S(v a ,c a )] + , Among them, Loss(v a , v a , c b , c b , α) is the loss function of the a-th video sub-feature and the b-th text sub-feature, S(v a , c b ) is the video-text similarity between the a-th video sub-feature and the b-th text sub-feature, S(v a , c a ) is the video-text similarity between the a-th video sub-feature and the a-th text sub-feature, S(v b , c a ) is the video-text similarity between the b-th video sub-feature and the a-th text sub-feature, β is a preset hyperparameter, a ∈ [1, i], b ∈ [1, j], i is the number of video sub-features, j is the number of text sub-features, [.] + is max(·, 0).

6. The video text retrieval method according to claim 5, wherein The video sub-feature includes a video global sub-feature, a video entity sub-feature, and a video action sub-feature, and the text sub-feature includes a text global sub-feature, a text entity sub-feature, and a text action sub-feature; The process of analyzing the video text similarity according to each of the video sub-features and each of the text sub-features of each of the target video frame block sequences respectively to obtain multiple video text similarities of each of the target video frame block sequences includes: Based on the third formula, calculate the global matching scores for each of the target video frame block sequences according to each of the video global sub-features and each of the text global sub-features of each of the target video frame block sequences. The third formula is: S g(i,j) = cos(v g,i , c g,j ), Among them, S g(i,j) is the global matching score between the i-th video global sub-feature and the j-th text global sub-feature, v g,i is the i-th video global sub-feature, c g,j is the j-th text global sub-feature; Based on the fourth formula, calculate the entity matching scores for each of the target video frame block sequences according to each of the video entity sub-features and each of the text entity sub-features of each of the target video frame block sequences. The fourth formula is: S e(i,j) = cos(v e,i , c e,j ), Among them, S e(i,j) is the entity matching score between the i-th video entity sub-feature and the j-th text entity sub-feature, v e,i is the i-th video entity sub-feature, c e,j is the j-th text entity sub-feature; Based on the fifth formula, calculate the action matching scores for each of the target video frame block sequences according to each of the video action sub-features and each of the text action sub-features of each of the target video frame block sequences. The fifth formula is: S a(i,j) = cos(v a,i , c a,j ), Among them, S a(i,j) is the action matching score between the i-th video action sub-feature and the j-th text action sub-feature, v a,i is the i-th video action sub-feature, c a,j is the j-th text action sub-feature; Normalize each of the global matching scores, each of the entity matching scores, and each of the action matching scores for each of the target video frame block sequences, respectively, to obtain the global attention weight parameters for each of the global matching scores, the entity attention weight parameters for each of the entity matching scores, and the action attention weight parameters for each of the action matching scores; Based on the sixth formula, calculate the target global matching scores for each of the target video frame block sequences according to each of the global matching scores and the global attention weight parameters of each of the global matching scores. The sixth formula is: S g,i,j = r g(i,j) S g(i,j) , Among them, S g,i,j is the target global matching score between the i-th video global sub-feature and the j-th text global sub-feature, and S g(i,j) is the global matching score between the i-th video global sub-feature and the j-th text global sub-feature, and r g(i,j) is the global attention weight parameter between the i-th video global sub-feature and the j-th text global sub-feature; Based on the seventh formula, calculate the target entity matching scores for each of the target video frame block sequences according to each of the entity matching scores and the entity attention weight parameters of each of the entity matching scores. The seventh formula is: S e,i,j = r e(i,j) S e(i,j) , Among them, S e,i,j is the target entity matching score between the i-th video entity sub-feature and the j-th text entity sub-feature, S e(i,j) is the entity matching score between the i-th video entity sub-feature and the j-th text entity sub-feature, r e(i,j) is the entity attention weight parameter between the i-th video entity sub-feature and the j-th text entity sub-feature; Based on the eighth formula, calculate the target action matching scores for each of the target video frame block sequences according to each of the action matching scores and the action attention weight parameters of each of the action matching scores. The eighth formula is: S a,i,j = r a(i,j) S a(i,j) , Among them, S a,i,j is the target action matching score between the i-th video action sub-feature and the j-th text action sub-feature, and S a(i,j) is the action matching score between the i-th video action sub-feature and the j-th text action sub-feature, and r a(i,j) is the action attention weight parameter between the i-th video action sub-feature and the j-th text action sub-feature; Based on the ninth formula, calculate the video-text similarity for each of the target video frame block sequences according to each of the target global matching scores, each of the target entity matching scores, and each of the target action matching scores. The ninth formula is: S(v i ,c j ) = (S g,i,j + S e,i,j + S a,i,j ) / 3, Among them, S(v i , c j ) is the video-text similarity between the i-th video action sub-feature and the j-th text action sub-feature, S g,i,j is the target global matching score between the i-th video global sub-feature and the j-th text global sub-feature, S e,i,j is the target entity matching score between the i-th video entity sub-feature and the j-th text entity sub-feature, S a,i,j is the target action matching score between the i-th video action sub-feature and the j-th text action sub-feature.

7. The video text retrieval method according to claim 1, wherein The process of updating the parameters of the visual-semantic supervision encoder and the trained video encoder according to the multiple loss functions for each of the target video frame block sequences includes: Based on the exponential moving average mechanism, update the parameters of the visual-semantic supervision encoder using the trained video encoder to obtain the updated visual-semantic supervision encoder; Update the parameters of the trained video encoder according to multiple loss functions of each of the target video frame block sequences to obtain an updated video encoder.

8. A video text retrieval device, characterized in that, Including: A partitioning module, configured to import multiple videos and multiple natural language text descriptions corresponding to each of the videos one by one, and randomly partition all the videos into a training set and a test set; A preprocessing module, configured to preprocess each of the videos and the corresponding natural language text descriptions in the training set respectively to obtain multiple target video frame block sequences of each of the videos in the training set; A training module, configured to construct a video encoder and a visual-semantic supervision encoder, and use the visual-semantic supervision encoder and the multiple target video frame block sequences of each of the videos in the training set to train the video encoder to obtain a trained video encoder and the video-text distances of each of the target video frame block sequences; A video encoding module, configured to encode the video-text distances of each of the target video frame block sequences respectively by using the trained video encoder to obtain video features of each of the target video frame block sequences; A text encoding module, configured to encode each of the target video frame block sequences respectively by using a text encoder to obtain text features of each of the target video frame block sequences; A loss function analysis module, configured to perform loss function analysis respectively according to the video features and text features of each of the target video frame block sequences to obtain multiple loss functions of each of the target video frame block sequences; A parameter update module, configured to update the parameters of the visual-semantic supervision encoder and the trained video encoder respectively according to the multiple loss functions of each of the target video frame block sequences to obtain an updated visual-semantic supervision encoder and an updated video encoder; A retrieval result obtaining module, configured to perform video-text retrieval processing on the test set by using the updated visual-semantic supervision encoder and the updated video encoder to obtain a video-text retrieval result.

9. A video text retrieval system, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the video-text retrieval method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the video-text retrieval method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Video natural language text retrieval method based on space time sequence characteristics

    CN113704546A

  • Text retrieval method, system and device and storage medium

    CN114003698A