A text-video cross-modal matching method based on asymmetric semantic optimization
Through an asymmetric semantically optimized text-video cross-modal matching method, cross-modal attention and multi-layer linear perceptron are used to interact global and fine-grained features. Combined with a knowledge-driven text editing mechanism and loss function optimization, the problems of insufficient exploration of semantic associations and redundant visual content in text-video retrieval are solved, and efficient and accurate text-video matching is achieved.
Patent Information
- Application Number
- CN202411868548.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-12-18
AI Technical Summary
Existing technologies fail to fully exploit the detailed semantic associations between language and vision in text-video retrieval, and the presence of redundant visual content in videos causes discriminative features to be submerged, affecting matching accuracy.
A text-video cross-modal matching method based on asymmetric semantic optimization is adopted. Global interaction is performed through the cross-modal attention module and the multi-layer linear perceptron. Global and fine-grained feature fusion is combined. A knowledge-driven text editing mechanism is used to generate negative samples. A comprehensive loss function is designed for optimization.
It significantly improves the accuracy and efficiency of text-video retrieval, can accurately capture key clues in video content, reduce redundant visual information interference, and improve matching accuracy and reliability.
Smart Images

Figure CN119719800B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video understanding and multimodal learning, and specifically relates to a text-video cross-modal matching method based on asymmetric semantic optimization. Background Art
[0002] With the rapid growth of multimedia data, text-video retrieval technology has attracted increasing attention. Current mainstream research focuses on constructing a unified embedding space for text and video to measure the similarity between the two. Although significant progress has been made through carefully designed strategies such as multi-granularity feature alignment and multimodal knowledge distillation, the detailed semantic connections between language and vision have not yet been fully explored. Specifically, textual expressions usually convey important semantic information through a few keywords or phrases, which play a dominant role in cross-modal text-video retrieval. In contrast, videos contain a large amount of redundant visual content that is irrelevant to the corresponding text description, resulting in the submergence of discriminative features. Summary of the Invention
[0003] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a text-video cross-modal matching method based on asymmetric semantic optimization.
[0004] The purpose of the present invention can be achieved by the following technical solutions:
[0005] The present invention provides a text-video cross-modal matching method based on asymmetric semantic optimization, comprising the following steps:
[0006] Step S1: Given a video set and a text query set, the video set includes multiple videos, and the text query set includes multiple text descriptions;
[0007] Step S2: extracting global features and local features of each text description through a text encoding model;
[0008] Step S3: extracting video frame-level features and image block-level features of each video through an image visual coding model;
[0009] Step S4: using a cross-modal attention module and a multi-layer linear perceptron to interact the global features of the text description with the video frame-level features of the video, and obtain the global interaction features of each video and the global matching similarity score of each frame in the video;
[0010] Step S5: Based on the global matching similarity scores of each frame in the video, the frames with the top-K global matching similarity scores are selected as candidates, and the local image block features of the candidates are interacted with the local text features of each text description to obtain fine-grained interaction features of each video;
[0011] Step S6: fuse the global interaction features and the fine-grained interaction features of each video to obtain the final video features of each video;
[0012] Step S7: calculate the cosine similarity between the global features of each text description and the final video features of each video, and match each text description with each video according to the cosine similarity.
[0013] Further, the video set is The text query set is Where n and m represent the number of videos and texts.
[0014] Further, the global feature and the local feature of each text description, the global feature of each text description is t ∈ R d , representing the overall semantic information of the text, and the local feature is h ∈ R L×d , representing the fine-grained semantic information in the text, where d is the dimension of the feature vector, and L represents the length of the label in the text sequence.
[0015] Further, the video frame level feature and the image block level feature of each video, the video frame level feature is f ∈ R T×d , representing the overall feature extracted from each frame of image in the video, capturing the global visual information of the video, and the image block level feature is p ∈ R T×M×d , representing the fine-grained feature extracted after dividing each frame of image into multiple image blocks, capturing the local visual information, where T represents the number of frames in the video, M represents the number of image blocks divided in each frame, and d is the dimension of the feature vector.
[0016] Further, the step S4 includes the following steps:
[0017] The global feature t ∈ R d of each text description and the video frame level feature f ∈ R T×d of each video are calculated by cross-modal attention, to obtain the global matching similarity score between the video frame and the text, and the formula is:
[0018] Q = LN(t)W Q
[0019] K = LN(f)W K
[0020] V = LN(f)W V
[0021]
[0022] Where LN() represents layer normalization, W Q ∈ R d×d, W K ∈R d×d and W V ∈R d×d is a learnable mapping matrix, Q is the query vector, which is normalized by the layer and passed through the mapping matrix W Q The converted text feature vector, K is the key vector, which means it is normalized by the layer and passed through the mapping matrix W K The converted video frame feature matrix, V is a value vector, which means it is normalized by the layer and passed through the mapping matrix W V The transformed video frame feature matrix, S∈R 1×T is the attention score matrix, which is the similarity score calculated by the dot product of the query Q and the key K;
[0023] The interactive features are further processed by the multi-layer linear perceptron MLP to obtain the global interactive features of each video. The formula is:
[0024] out=SV
[0025] v g =MLP(out)
[0026] Among them, out is the intermediate feature of the output, which is the weighted summation result obtained by multiplying the attention score matrix S with the value vector V, containing weighted video frame information, where the weight of each frame is determined by the corresponding score in S, v g is the global interaction feature of each video, and MLP() is a multi-layer linear perceptron operation.
[0027] Furthermore, based on the global matching similarity scores of the frames in the video, frames with top-K global matching similarity scores are selected as candidates, including the following steps:
[0028] The global matching similarity scores of each frame in the video are sorted from large to small, and the top K video frames are selected as candidates for the video, where K is a constant.
[0029] Furthermore, the local image block features of the candidate are interacted with the local text features of each text description to obtain the fine-grained interactive features of each video, and the formula is:
[0030]
[0031] Among them, v l is the fine-grained interaction feature of each video, L is the length of the token in the text sequence, K is the number of candidates for selecting frames, and p i ∈R M×d represents the M image block features of the i-th video frame, A is the index set of candidate frames with top-K similarity scores, h jRepresents the local text feature corresponding to the jth token in the text, and T is the transposition.
[0032] Furthermore, the global interaction features of each video are fused with the fine-grained interaction features, and the formula is:
[0033]
[0034] Among them, v is the final video feature of the video, v g is the global interaction feature of each video, v l is the fine-grained interaction feature of each video.
[0035] Furthermore, the text encoding model, image visual encoding model, cross-modal attention module and multi-layer linear perception mechanism cost method model, the overall loss function of the model is:
[0036]
[0037]
[0038] in, is the overall loss function of the model, is the cross entropy loss, is the text-to-video loss, is the loss from video to text, γ is a predefined hyperparameter, q(t i ,v i ) is the global feature t of the i-th text description i and the final video feature v of the i-th video i The cosine similarity between i ,v j ) is the global feature t of the i-th text description i and the final video feature v of the jth video j The cosine similarity between j ,v i ) is the global feature t of the j-th text description j and the final video feature v of the i-th video i The cosine similarity between i is the feature of the i-th negative sample, N represents the number of negative samples, λ is a learnable scaling factor, and z is a predefined spacing constant.
[0039] Furthermore, the negative samples are obtained by the following steps:
[0040] Use language tools to identify adjectives, nouns, and verbs in text descriptions, replace these words using a pre-trained language model, RoBERTa, and use the replaced synthetic text descriptions as negative samples;
[0041] The features of the negative samples are processed by the text encoder to obtain the features g of each negative sample. i .
[0042] Compared with the prior art, the present invention has the following advantages:
[0043] (1) The present invention proposes an asymmetric semantic optimization mechanism designed specifically for text-video retrieval, which can significantly alleviate the asymmetric characteristics of semantic information between text and video modal data, thereby achieving more accurate and efficient retrieval effects.
[0044] (2) This invention pioneered a multi-granular text-video interaction mechanism. This mechanism ingeniously refines and integrates learned knowledge from a global perspective, using this knowledge as a guide to deeply analyze subtle yet critical clues in video content. This mechanism can accurately highlight highly valuable clues in video content, thereby cleverly resolving the common problem of large amounts of redundant visual content in videos that is irrelevant to the text description, and significantly improving the efficiency and accuracy of text-video interaction.
[0045] (3) This paper innovatively proposes a knowledge-driven text editing mechanism, applying it to the field of text-video retrieval for the first time. By synthesizing structurally similar negative samples, it optimizes reliable language features from a metric learning perspective. This mechanism effectively extracts key fine-grained semantic knowledge, addressing the problem of traditional methods that under-exploit detailed semantic connections between language and vision.
[0046] (4) The cross-modal matching method proposed in the present invention overcomes the asymmetry of semantic information between video and text in traditional methods through the hierarchical interaction of global features and fine-grained features. First, a cross-modal attention mechanism is used to globally interact the global features of the text with the features of the video frame, and the global matching similarity score is calculated to ensure that the semantic association between the video and the text can be accurately captured. Then, by selecting candidate video frames with high similarity, the image block features are further interacted with the local features of the text to mine fine-grained semantic clues. This multi-level interaction mechanism effectively eliminates the interference of redundant visual information and improves the accuracy of cross-modal matching.
[0047] (5) This invention innovatively introduces an asymmetric semantic optimization mechanism to address the semantic information asymmetry between video and text. Under this mechanism, the global matching similarity of text and video and fine-grained features are combined to achieve multi-level and multi-granular semantic optimization. In particular, by designing a knowledge-driven text editing mechanism and synthesizing challenging negative samples, the discriminability of language features is optimized, thereby improving the model's sensitivity to semantic details and enhancing the reliability of matching.
[0048] (6) To prevent redundant information in the video from affecting the matching effect, the present invention introduces a knowledge-based text editing mechanism in fine-grained feature modeling. By identifying keywords in the text (such as adjectives, nouns, verbs, etc.) and generating negative samples through a pre-trained language model (such as RoBERTa), the model's ability to perceive fine-grained text semantics is further enhanced. This processing enables the model to more finely capture key details in the text description, improving the accuracy of text-video matching.
[0049] (7) This paper combines metric learning loss functions for both text-to-video and video-to-text to design a comprehensive loss function. This loss function, through bidirectional metric optimization, further improves the model's cross-modal matching performance. Furthermore, the metric learning-based loss function can effectively distinguish relevant and irrelevant text-video pairs during training, optimizing the model's matching accuracy.
[0050] (8) In the final video feature generation stage, the present invention ensures that the multi-level semantic information of the video can be fully expressed by fusing global interaction features with fine-grained interaction features. The final fused video features have stronger distinguishing capabilities and can achieve higher accuracy in text-video matching. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is a flow chart of the method of the present invention;
[0052] Figure 2 This is the overall technical framework diagram of the present invention. DETAILED DESCRIPTION
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0054] Example 1:
[0055] This embodiment provides a text-video cross-modal matching method based on asymmetric semantic optimization, such as Figure 1 As shown, the following steps are included:
[0056] Step S1: Given a video set and a text query set, the video set includes multiple videos, and the text query set includes multiple text descriptions;
[0057] Step S2: extracting global features and local features of each text description through a text encoding model;
[0058] Step S3: extracting video frame-level features and image block-level features of each video through an image visual coding model;
[0059] Step S4: using a cross-modal attention module and a multi-layer linear perceptron to interact the global features of the text description with the video frame-level features of the video, and obtain the global interaction features of each video and the global matching similarity score of each frame in the video;
[0060] Step S5: Based on the global matching similarity scores of each frame in the video, the frames with the top-K global matching similarity scores are selected as candidates, and the local image block features of the candidates are interacted with the local text features of each text description to obtain fine-grained interaction features of each video;
[0061] Step S6: Fusing the global interaction features and fine-grained interaction features of each video to obtain the final video features of each video;
[0062] Step S7: Calculate the cosine similarity between the global features of each text description and the final video features of each video, and match each text description with each video based on the cosine similarity.
[0063] The video collection is The text query set is Where n and m represent the number of videos and texts.
[0064] Among them, the global features and local features of each text description, the global feature of each text description is t∈R d , represents the overall semantic information of the text, and the local feature is h∈R L×d , represents the fine-grained semantic information in the text, where d is the feature vector dimension and L represents the length of the token in the text sequence.
[0065] Among them, the video frame level features and image block level features of each video, the video frame level features are f∈R T×d , represents the overall features extracted from each frame of the video, capturing the global visual information of the video, and the image block level feature is p∈R T×M×d, represents the fine-grained features extracted after dividing each frame into multiple image blocks, capturing local visual information, where T represents the number of frames in the video, M represents the number of image blocks in each frame, and d is the dimension of the feature vector.
[0066] Wherein, it is characterized in that step S4 includes the following steps:
[0067] The global features t∈R for each text description d And the video frame level features f∈R of each video T×d Perform cross-modal attention calculation to obtain the global matching similarity score between the video frame and the text. The formula is:
[0068] Q=LN(t)W Q
[0069] K=LN(f)W K
[0070] V=LN(f)W V
[0071]
[0072] Among them, LN() represents layer regularization, W Q ∈R d×d , W K ∈R d×d and W V ∈R d×d is a learnable mapping matrix, Q is the query vector, which is normalized by the layer and passed through the mapping matrix W Q The converted text feature vector, K is the key vector, which means it is normalized by the layer and passed through the mapping matrix W K The converted video frame feature matrix, V is a value vector, which means it is normalized by the layer and passed through the mapping matrix W V The transformed video frame feature matrix, S∈R 1×T is the attention score matrix, which is the similarity score calculated by the dot product of the query Q and the key K;
[0073] The interactive features are further processed by the multi-layer linear perceptron MLP to obtain the global interactive features of each video. The formula is:
[0074] out=SV
[0075] v g =MLP(out)
[0076] Among them, out is the intermediate feature of the output, which is the weighted summation result obtained by multiplying the attention score matrix S with the value vector V, containing weighted video frame information, where the weight of each frame is determined by the corresponding score in S, vg is the global interaction feature of each video, and MLP() is a multi-layer linear perceptron operation.
[0077] Among them, based on the global matching similarity scores of each frame in the video, the frames with the top-K global matching similarity scores are selected as candidates, including the following steps:
[0078] The global matching similarity scores of each frame in the video are sorted from large to small, and the top K video frames are selected as candidates for the video, where K is a constant.
[0079] Among them, the local image block features of the candidate are interacted with the local text features of each text description to obtain the fine-grained interactive features of each video. The formula is:
[0080]
[0081] Among them, v l is the fine-grained interaction feature of each video, L is the length of the token in the text sequence, K is the number of candidates for selecting frames, and p i ∈R M×d represents the M image block features of the i-th video frame, A is the index set of candidate frames with top-K similarity scores, h j Represents the local text feature corresponding to the jth token in the text, and T is the transposition.
[0082] Among them, the global interaction features of each video are fused with the fine-grained interaction features. The formula is:
[0083]
[0084] Among them, v is the final video feature of the video, v g is the global interaction feature of each video, v l is the fine-grained interaction feature of each video.
[0085] Among them, the total model of the text encoding model, image visual encoding model, cross-modal attention module and multi-layer linear perception mechanism cost method, the overall loss function of the total model training is:
[0086]
[0087] in, is the overall loss function for total model training, is the cross entropy loss, is the text-to-video loss, is the loss from video to text, γ is a predefined hyperparameter, q(t i ,v icosine similarity between the global feature t i of the i-th text description and the final video feature v i of the i-th video, q(t i ,v j ) is the cosine similarity between the global feature t i of the i-th text description and the final video feature v j of the j-th video, q(t j ,v i ) is the cosine similarity between the global feature t j of the i-th text description and the final video feature v i of the i-th video, g i is the feature of the i-th negative sample, N represents the number of negative samples, and λ is a learnable scaling factor, and z is a predefined margin constant.
[0088] where the negative samples are obtained by the following steps:
[0089] Adjectives, nouns and verbs in the text description are identified using a language tool, and a pre-trained language model RoBERTa is used to replace these words. The synthesized text description after replacement is used as a negative sample;
[0090] The features of the negative samples are obtained by processing each negative sample through a text encoder to obtain the features g i of each negative sample.
[0091] Embodiment 2:
[0092] The parts not mentioned in this embodiment are the same as in Embodiment 1.
[0093] The text-video cross-modal matching method based on asymmetric semantic optimization has the overall technical framework as shown in Figure 2 , which can be realized by the following steps:
[0094] Step 1: Given a set of videos and a set of text queries where n and m represent the number of videos and texts. The goal of the text-video retrieval task is to learn a metric function B = (v i ,t i ) that measures the similarity between text and video. First, the global feature t ∈ R d and the local feature h ∈ R L×d of the text are extracted by a text encoding model, where L represents the length of the token in the text sequence. In addition, the frame-level feature f ∈ R T×d and the patch-level feature p ∈ R T×M×d of the video are extracted by an image visual encoding model.where T and M represent the number of frames and the number of image patches, respectively.
[0095] Step 2: To model the global interaction information between video and text description, a cross-modal attention module and a Multi-Layer Perceptron (MLP) are applied on the global text feature t and video frame feature f to obtain the video-level feature v g and the global matching similarity score S ∈ R 1×T , the calculation process is as follows:
[0096] Q = LN(t)W Q
[0097] K = LN(f)W K
[0098] V = LN(f)W V
[0099]
[0100] out = SV
[0101] v g = MLP(out)
[0102] where LN denotes layer normalization, W Q ∈ R d×d , W K ∈ R d×d and W V ∈ R d×d are learnable mapping matrices, Q is the query vector, which represents the text feature vector after layer normalization and conversion through the mapping matrix W Q , K is the key vector, which represents the video frame feature matrix after layer normalization and conversion through the mapping matrix W K , V is the value vector, which represents the video frame feature matrix after layer normalization and conversion through the mapping matrix W V , S ∈ R 1×T is the attention score matrix, which is calculated by the dot product of the query Q and the key K.
[0103] Step 3: To further encode fine-grained semantic cues, the global matching similarity score S generated in step 2 is used to guide the modeling of fine-grained features. Specifically, the frames with top-K similarity scores are selected as candidates, and the interaction between the local image patch features of these candidates and the local text features h is modeled to produce fine-grained interaction features v l , the specific process is as follows:
[0104]
[0105] where p i ∈R M×d represents the M image block features of the i-th video frame, and A is the index set of candidate frames with top-K similarity scores.
[0106] Step 4: Global interaction feature v g With fine-grained interaction features v l are fused to obtain the final video features
[0107] Step 5: Finally, the cosine similarity between the text feature t and the video feature v is used to obtain their matching degree, which is used as the selection criterion for the retrieval results.
[0108] Step 6: To prevent fine-grained discriminative video features from being overwhelmed by the honorific video information, the present invention further mines fine-grained text semantic information and designs a knowledge-based text editing mechanism and fine-grained text semantic perception module. It uses language tools to identify adjectives, nouns, and verbs in text descriptions, and then uses a pre-trained language model RoBERTa to replace these words. These replaced synthetic text descriptions are considered negative samples, and they are processed by the text encoder to obtain feature g. Next, a metric loss is used to distinguish relevant and irrelevant text-video pairs. The specific calculation method is as follows:
[0109]
[0110] Here, z is a predefined spacing constant and N represents the number of negative samples.
[0111] Step 7: In order to associate the paired text and video, this embodiment considers the retrieval metrics of the "text-video" direction and the "video-text" direction at the same time, and obtains the loss function and Finally, the loss function is obtained by combining the two Conduct supervised training:
[0112]
[0113] The overall loss function for model training is as follows:
[0114]
[0115] in, is the overall loss function for total model training, is the cross entropy loss, is the text-to-video loss, is the loss from video to text, γ is a predefined hyperparameter, q(t i ,vi cosine similarity between the global feature t i of the i-th text description and the final video feature v i of the i-th video, q(t i ,v j cosine similarity between the global feature t i of the i-th text description and the final video feature v j of the j-th video, q(t j ,v i cosine similarity between the global feature t j of the j-th text description and the final video feature v i of the i-th video, g i is the feature of the i-th negative sample, N represents the number of negative samples, λ is a learnable scaling factor, and m is a pre-defined margin constant. By introducing multi-granularity feature interaction, cross-modal attention mechanism and asymmetric semantic optimization strategy, the application significantly improves the matching accuracy between text and video, especially in handling complex cross-modal retrieval tasks, it can more accurately capture the deep semantic association between text and video, and reduce the interference of redundant information. In addition, the application of negative sample generation strategy enables the model to better distinguish relevant and irrelevant text-video pairs, further optimizing the retrieval effect.
[0116] Through the comprehensive application of the above technical features, the application realizes efficient and accurate text-video cross-modal matching, providing an innovative solution for text-video retrieval tasks, with significant technical advantages and application prospects.
[0117] The above functions, if realized in the form of software function units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the application or parts of the existing technology that essentially contribute or parts of the technical solutions can be embodied in the form of software products, which are stored in a storage medium and include instructions to make a computer device (which can be a personal computer, server, or network device, etc.) execute all or part of the steps of the method described in various embodiments of the application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, and various program code storage media.
[0118] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A text-video cross-modal matching method based on asymmetric semantic optimization, characterized in that: The following steps are involved: Step S1: Given a video set and a text query set, the video set includes multiple videos, and the text query set includes multiple text descriptions; Step S2: extracting global features and local features of each text description through a text encoding model; Step S3: extracting video frame-level features and image block-level features of each video through an image visual coding model; Step S4: using a cross-modal attention module and a multi-layer linear perceptron to interact the global features of the text description with the video frame-level features of the video, and obtain the global interaction features of each video and the global matching similarity score of each frame in the video; Step S5: Based on the global matching similarity scores of each frame in the video, top-K Frames with global matching similarity scores are selected as candidates, and the local image block features of the candidates are interacted with the local text features of each text description to obtain fine-grained interaction features of each video; Step S6: Fusing the global interaction features and fine-grained interaction features of each video to obtain the final video features of each video; Step S7: Calculate the cosine similarity between the global features of each text description and the final video features of each video, and match each text description with each video based on the cosine similarity; The step S4 comprises the following steps: Global features of each text description and the video frame level features of each video Perform cross-modal attention calculation to obtain the global matching similarity score between the video frame and the text. The formula is: in, Representation layer regularization, , and is a learnable mapping matrix, is the query vector, which is normalized by the layer and passed through the mapping matrix The converted text feature vector, is the key vector, which means it is normalized by the layer and passed through the mapping matrix The converted video frame feature matrix, Is a value vector, which means it is normalized by the layer and passed through the mapping matrix The converted video frame feature matrix, is the attention score matrix, by querying and key The similarity score is calculated by the dot product of The interactive features are further processed by the multi-layer linear perceptron MLP to obtain the global interactive features of each video. The formula is: in, As the intermediate feature of the output, the attention score matrix With value vector The weighted summation result obtained by multiplication contains weighted video frame information, where the weight of each frame is determined by The corresponding score in is the global interaction feature of each video, It is a multi-layer linear perceptron operation; The candidate's local image block features are interacted with the local text features of each text description to obtain the fine-grained interactive features of each video. The formula is: in, is the fine-grained interaction feature of each video, represents the length of the token in the text sequence, is the number of candidates for selecting frames, Indicates the video frames M image patch features, Is a top-K The index set of candidate frames with similarity scores, Indicates the first j The local text features corresponding to the tags, T is transposed.
2. A text-video cross-modal matching method based on asymmetric semantic optimization according to claim 1, characterized in that: The video collection is U= , the text query set is C= ,in and Indicates the number of videos and texts.
3. The text-video cross-modal matching method based on asymmetric semantic optimization according to claim 1, characterized in that: The global features and local features of each text description, the global features of each text description are , represents the overall semantic information of the text, and the local features are , represents the fine-grained semantic information in the text, where is the feature vector dimension, Indicates the length of tokens in a text sequence.
4. The text-video cross-modal matching method based on asymmetric semantic optimization according to claim 1, characterized in that: The video frame level features and image block level features of each video, the video frame level features are , represents the overall features extracted from each frame of the video, capturing the global visual information of the video, and the image block level features are , represents the fine-grained features extracted after dividing each frame image into multiple image blocks, capturing local visual information, where, Indicates the number of frames in the video, Indicates the number of image blocks in each frame, is the feature vector dimension.
5. The text-video cross-modal matching method based on asymmetric semantic optimization according to claim 1, characterized in that: The global matching similarity score based on each frame in the video has top-K Frames with globally matching similarity scores are selected as candidates, which includes the following steps: Sort the global matching similarity scores of each frame in the video from large to small, and select the top ranked K The video frames of the bit are selected as candidates for the video, K is a constant.
6. The text-video cross-modal matching method based on asymmetric semantic optimization according to claim 1, characterized in that: The global interaction features of each video are fused with the fine-grained interaction features. The formula is: in, is the final video feature of the video, is the global interaction feature of each video, is the fine-grained interaction feature of each video.
7. The text-video cross-modal matching method based on asymmetric semantic optimization according to claim 1, characterized in that: The text encoding model, image visual encoding model, cross-modal attention module and multi-layer linear perception mechanism cost method model, the overall loss function of the model is: in, is the overall loss function of the model, is the cross entropy loss, is the text-to-video loss, is the video-to-text loss, are predefined hyperparameters, For the i Global features of text description With the i Final video features of videos The cosine similarity between For the i Global features of text description With the j Final video features of videos The cosine similarity between For the j Global features of text description With the i Final video features of videos The cosine similarity between For the i The characteristics of negative samples, represents the number of negative samples, is a learnable scaling factor, is a predefined spacing constant.
8. The text-video cross-modal matching method based on asymmetric semantic optimization according to claim 7, characterized in that: The negative samples are obtained by the following steps: Use language tools to identify adjectives, nouns, and verbs in text descriptions, replace these words using a pre-trained language model, RoBERTa, and use the replaced synthetic text descriptions as negative samples; The features of the negative samples are processed by the text encoder to obtain the features of each negative sample. .
Citation Information
Patent Citations
Cross-modal retrieval method based on multi-granularity feature interaction
CN114037945A
Cross-modal text-video retrieval method based on space-time relationship enhancement
CN114048351A