Video-text cross-modal prompt hash retrieval method based on global-local video attention

The video-text cross-modal cue hash retrieval method using global-local video attention optimizes video and text feature representations by leveraging a visual-language pre-trained model and a global-local video attention mechanism. This addresses the problem in existing methods that fail to fully utilize the temporal continuity of video and the cross-modal semantic gap, achieving efficient and accurate cross-modal video-text retrieval.

CN121833995APending Publication Date: 2026-04-10NANJING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING UNIV OF SCI & TECH
Filing Date
2025-12-26
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing deep cross-modal video-text hashing methods fail to fully utilize the temporal continuity of videos and the cross-modal semantic gap, resulting in a lack of robustness in complex scenarios and reduced performance of cross-modal video-text hashing retrieval.

Method used

A video-text cross-modal cue hash retrieval method with global-local video attention is adopted. The visual-language pre-trained model CLIP is used for encoding. The method combines visual-language guided cue learning and global-local video attention mechanism to optimize video and text feature representations. The feature quality is improved by an adaptive feature fusion module.

Benefits of technology

It effectively bridges the semantic gap, improves the performance and retrieval efficiency of cross-modal video-text hash retrieval, and enhances the accuracy and efficiency of retrieving videos from big data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121833995A_ABST
    Figure CN121833995A_ABST
Patent Text Reader

Abstract

The invention discloses a video-text cross-modal prompt hash retrieval method based on global-local video attention. The method comprises the steps that a CLIP model serves as a video and text encoder, a vision-language guided video-text prompt technology is used for generating a vision guided text prompt and a language guided local frame prompt, and a global video prompt is designed; a global-local video attention mechanism is integrated in each layer of Transform of a visual encoder, text and video prompt learning is used to obtain text and video depth feature representations, and the text and text depth feature representations are input into an adaptive feature fusion module to obtain optimized video and text depth feature representations; inputting the optimized video and text depth feature representation into a Hash network to obtain continuous real value representation of the video and the text, and training the CLIP and the Hash network by using an overall objective function; and performing cross-modal video-text retrieval by using the trained model. According to the method, the efficiency and the accuracy of retrieving the video from the big data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and multimodal retrieval technology, and in particular to a video-text cross-modal cue hash retrieval method based on global-local video attention. Background Technology

[0002] With the rapid development of the internet and multimedia technologies, video data is growing exponentially. This growth makes efficient and accurate retrieval from massive video datasets a significant technical challenge. In practical retrieval scenarios, users typically prefer to use natural language text, such as keywords or sentence descriptions, for searching, which significantly increases the demand for cross-modal video-text retrieval. However, traditional real-valued vector-based retrieval methods face significant limitations when processing massive amounts of data, making them difficult to implement in real-world applications.

[0003] The emergence of hashing technology provides an effective method for solving the above problems. By mapping high-dimensional video and text to a compact Hamming space, hashing methods can improve retrieval speed and reduce storage costs, thus becoming a core technology for large-scale cross-modal video-text retrieval.

[0004] Despite some progress in feature extraction and retrieval performance from deep cross-modal video-text hashing, several challenges remain. First, some existing methods directly apply image hashing to videos, neglecting the unique temporal continuity of videos and failing to fully utilize discriminative temporal information. Second, the inherent semantic gap between video and text is not adequately addressed during cross-modal alignment. Consequently, the joint embedding space learned by current methods lacks robustness in complex scenarios, weakening the discriminative power of hash codes and reducing the performance of cross-modal video-text hash retrieval. Summary of the Invention

[0005] The purpose of this invention is to provide a video-text cross-modal cue hash retrieval method based on global-local video attention, which has high retrieval efficiency and high retrieval performance.

[0006] The solution to achieve the purpose of this invention is: a video-text cross-modal cue hash retrieval method based on global-local video attention, comprising the following steps:

[0007] Step 1: Use the visual-language pre-trained model CLIP as the encoder for video and text. Use visual-language guided video-text prompting technology to generate visually guided text prompts for text, generate language-guided local frame prompts for video, and design global video prompts.

[0008] Step 2: Integrate a global-local video attention mechanism into each layer of the Transformer in the visual encoder, which is divided into local frame attention and global video attention;

[0009] Step 3: Use text prompt learning and video prompt learning to obtain deep feature representations of text and video respectively. Then, input the video and text deep feature representations together into the adaptive feature fusion module to obtain optimized video and text deep feature representations.

[0010] Step 4: Input the optimized video and text deep feature representations into the hash network to obtain continuous real-valued representations of the video and text, and binarize them into a unified hash code.

[0011] Step 5: Design the overall objective function, including supervised intermodal contrast loss, quantization loss, and bit balance loss, and train the CLIP and hash networks;

[0012] Step 6: Use the trained model to perform cross-modal video-text retrieval.

[0013] An electronic device includes a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the video-text cross-modal cue hash retrieval method based on global-local video attention.

[0014] Compared with the prior art, the present invention has the following significant advantages: (1) It uses a visual language pre-trained model to encode video and text modalities, resulting in a more powerful network structure that can capture the temporal features of video and the latent semantic features of text. It also makes full use of the cross-modal semantic alignment performance of the large model to effectively bridge the semantic gap. (2) It adopts visual-language guided video-text prompt learning, optimizes text representation through visual-guided text prompt learning to make it more consistent with video content, focuses on relevant semantic regions and learns local frame information through language-guided local frame prompt learning, integrates video features of global video semantics through global video prompt learning, adopts a global-local video attention module to achieve a comprehensive understanding of the video and effective learning of cross-frame temporal dynamics, and uses an adaptive feature fusion module to adaptively filter and weight redundant information in video and text features, effectively improving feature quality. This improves the performance of cross-modal video-text hash retrieval and enhances the efficiency and accuracy of retrieving videos from big data. Attached Figure Description

[0015] Figure 1 This is a flowchart illustrating a video-text cross-modal cue hash retrieval method based on global-local video attention according to the present invention. Detailed Implementation

[0016] It is readily understood that, based on the technical solution of this invention, those skilled in the art can conceive of various embodiments of this invention without altering its essential spirit. Therefore, the following specific embodiments and accompanying drawings are merely illustrative examples of the technical solution of this invention and should not be considered as the entirety of this invention or as limitations or restrictions on its technical solution.

[0017] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of the invention.

[0018] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the invention or its application or use.

[0019] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.

[0020] In all the examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.

[0021] Combination Figure 1 This invention is a video-text cross-modal cue hash retrieval method based on global-local video attention, comprising the following steps:

[0022] Step 1: Use the visual-language pre-trained model CLIP as the encoder for video and text. Use visual-language guided video-text prompting technology to generate visually guided text prompts for text, generate language-guided local frame prompts for video, and design global video prompts.

[0023] Step 2: Integrate a global-local video attention mechanism into each layer of the Transformer in the visual encoder, which is divided into local frame attention and global video attention;

[0024] Step 3: Use text prompt learning and video prompt learning to obtain deep feature representations of text and video respectively. Then, input the video and text deep feature representations together into the adaptive feature fusion module to obtain optimized video and text deep feature representations.

[0025] Step 4: Input the optimized video and text deep feature representations into the hash network to obtain continuous real-valued representations of the video and text, and binarize them into a unified hash code.

[0026] Step 5: Design the overall objective function, including supervised intermodal contrast loss, quantization loss, and bit balance loss, and train the CLIP and hash networks;

[0027] Step 6: Use the trained model to perform cross-modal video-text retrieval.

[0028] As a specific example, step 1 describes using the visual-language pre-trained model CLIP as an encoder for both video and text, employing visual-language guided video-text prompting technology to generate visually guided text prompts for the text, and generating language-guided local frame prompts and designing global video prompts for the video, as detailed below:

[0029] Step 1.1: Process the video into frame embeddings and the text into word embeddings;

[0030] Step 1.2: Generate visually guiding text prompts for the text by embedding video frames;

[0031] Step 1.3: Generate local frame cues with language guidance for the video through word embedding, and design global video cues for the video.

[0032] As a specific example, step 1.1, which involves processing video into frame embeddings and text into word embeddings, is detailed as follows:

[0033] Step 1.1.1: Sample the original video as... Each frame is divided into several frames. A fixed-size block is defined, and these blocks are linearly projected into an embedding, i.e. ,in This indicates the dimension of the frame embedding, and k represents the frame index;

[0034] Step 1.1.2: Segment the text to generate a text sequence, and then project the text sequence into the embedding space to generate word embeddings. ,in The length of the word tag, The dimension of word embedding.

[0035] As a specific example, step 1.2, which involves embedding video frames to generate visually guiding text prompts, is as follows:

[0036] Step 1.2.1: Embed all frames in a video and aggregate them into a single frame using the time aggregation module;

[0037] Step 1.2.2: Send a single frame into the text prompt generation module. In the text prompt generation module, the single frame is first processed into text modulation parameters with video information. Then, the text modulation parameters are used to randomly initialize learnable text tags to finally obtain visually guided text prompts.

[0038] As a specific example, step 1.3, which involves generating language-guided local frame cues for the video through word embedding and designing global video cues, is detailed as follows:

[0039] Step 1.3.1: The word embedding is sent to the visual cue generation module. In the visual cue generation module, the word embedding is first processed into visual modulation parameters with semantic information. Then, the visual modulation parameters are used for randomly initialized learnable visual tags to finally obtain language-guided local frame cue.

[0040] Step 1.3.2: Design shallow global video cues based on video characteristics and obtain global information of the video. Shallow global video cues are randomly initialized learnable visual tags.

[0041] As a specific example, step 2, which integrates a global-local video attention mechanism in each layer of the visual encoder's Transformer, is divided into local frame attention and global video attention, as detailed below:

[0042] Step 2.1, Design Global Video Attention: Global cues need to learn global discriminative information. Therefore, global cues need to pay attention to all video frame sequences and frame cue embeddings. For global video attention, the global cue is used as a query, and the global cue is concatenated with the CLS tag, frame cue, and frame sequence as a key and value. The process of global video attention is as follows:

[0043]

[0044]

[0045]

[0046] in, This indicates a global video prompt. The CLS tag representing the frame, Indicates frame hint embedding, Indicates frame embedding, Indicates the index of the frame;

[0047] Step 2.2, Designing Local Frame Attention: For local frame attention, each frame cue needs to perceive information from every local frame. Therefore, the CLS tag, frame cue, and frame sequence are concatenated as a query. To ensure that each frame cue can perceive global information, the global video cue needs to be repeated. Next, it is concatenated with the CLS tag, frame cue, and frame sequence as a key and value; the process of local frame attention is as follows:

[0048]

[0049]

[0050]

[0051] As a specific example, step 3 involves using text prompt learning and video prompt learning to obtain deep feature representations of text and video respectively. Then, the video and text deep feature representations are input together into the adaptive feature fusion module to obtain optimized video and text deep feature representations, as detailed below:

[0052] Step 3.1: Learn the deep feature representation of the text through text prompts;

[0053] Step 3.2: Learn the deep feature representation of the video through video prompts;

[0054] Step 3.3: Input the output video and text deep feature representations together into the adaptive feature fusion module to obtain the optimized video and text deep feature representations.

[0055] As a specific example, the deep feature representation of the text obtained through text prompts in step 3.1 is as follows:

[0056] Step 3.1.1: Input the word embeddings and text prompts as input into the CLIP text encoder. The The layer output is:

[0057]

[0058] in, This indicates the total number of layers in the CLIP text encoder. This indicates a text prompt. Indicates word embedding;

[0059] Step 3.1.2: By outputting the last layer [EOS] token Projecting onto the latent embedding space, the final text features are obtained:

[0060]

[0061] in This represents the final text features. This indicates the text projection layer.

[0062] As a specific example, the deep feature representation of the video obtained through video prompts in step 3.2 is as follows:

[0063] Step 3.2.1: Concatenate the frame embedding, global video cues, and frame cues with the learnable CLS tags for each frame, and input them together into the CLIP visual encoder. The The layer output is:

[0064]

[0065] in, This indicates a global video prompt. The CLS tag representing the frame, Indicates frame hint embedding, Indicates frame embedding, Indicates the index of the frame;

[0066] Step 3.2.2: The global cue perceives information from each frame, so it can be considered a combination of fine-grained frame representation and global video representation. Therefore, the first label of the global cue obtained from the last layer output is projected into the latent embedding space to obtain the final video features:

[0067]

[0068] in This represents the final video features. This represents the video projection layer.

[0069] As a specific example, step 3.3 involves inputting the output video and text deep feature representations together into the adaptive feature fusion module to obtain optimized video and text deep feature representations, as detailed below:

[0070] Step 3.3.1: Obtain the deep feature representations of the video and text sequentially through Global Response Normalization (GRN) and Multilayer Perceptron (MLP). and The results are then concatenated to obtain the fused features. This process can be represented as:

[0071]

[0072]

[0073] Where Concat represents the concatenation operation, and SiLU is the activation function;

[0074] Step 3.3.2: Flip the fused features and feed them into the State-Space Model (SSM). Then perform the inverse flipping operation to obtain the flipped features. A total of three flipping operations are performed to obtain the final three flipped features. The process can be represented as follows:

[0075]

[0076] in , A value of 1 indicates a left-right flip. A value of 2 indicates flipping vertically. (1,2) indicates that the elements are flipped simultaneously in all directions (up, down, left, and right). This indicates a flipping operation performed along different axes. This indicates its inverse operation;

[0077] Step 3.3.3: Perform adaptive weighted filtering on the three flipped features and the fused features to obtain the features. The process can be represented as:

[0078]

[0079] in Represents a set of learnable parameters. This represents a temperature coefficient used to balance the relative importance of flipped and non-flipped characteristics;

[0080] Step 3.3.4: Divide the features obtained after weighted filtering into two parts, and then compare them with... and Concatenate and sequentially pass through GRN and MLP to obtain and The process can be represented as:

[0081]

[0082]

[0083] Where chunk represents the operation of dividing the feature into two parts;

[0084] Step 3.3.5, will and After concatenation, the data is separated again through a Transformer encoder layer to obtain the final video and text features. and The process can be represented as:

[0085]

[0086] Split represents the operation of separating the final video features and text features from the fused features.

[0087] As a specific example, step 4 involves inputting the optimized video and text deep feature representations into a hash network to obtain continuous real-valued representations of the video and text, which are then binarized into a unified hash code, as detailed below:

[0088] Step 4.1: After the CLIP encoder, design a hash network consisting of an MLP and a hash layer. Input the obtained video and text deep feature representations into the hash network to obtain continuous real-valued representations of the video and text. and :

[0089]

[0090]

[0091] Step 4.2: Binarize the continuous real-valued representations of the video and text to obtain a unified binary hash code:

[0092]

[0093] in Represents a symbolic function.

[0094] As a specific example, the overall objective function designed in step 5 includes supervised intermodal contrast loss, quantization loss, and bit balance loss. The CLIP and hash networks are trained as follows:

[0095] Step 5.1: Design the supervised inter-modal contrast loss, as follows:

[0096] Given the first Video and text representation and Taking video modality as an example, set It is a collection of video indexes. Then the positive sample set is defined as ,in and They represent the first The and the first The tags corresponding to each video;

[0097] by The definition of inter-modal contrast loss as an anchor point is as follows:

[0098]

[0099] in Temperature coefficient;

[0100] Similarly, with The definition of inter-modal contrast loss as an anchor point is as follows:

[0101]

[0102] Therefore, the supervised modal contrast loss is:

[0103]

[0104] Step 5.2: Design the quantification loss, as follows:

[0105] set up , as well as The quantization loss is then defined as:

[0106]

[0107] Step 5.3: Design the bit balance loss, which is defined as:

[0108]

[0109] Step 5.4, the overall objective function is:

[0110]

[0111] in and It is a trade-off parameter for the relative importance of controlling losses.

[0112] The present invention provides an electronic device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the video-text cross-modal cue hash retrieval method based on global-local video attention.

[0113] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0114] Example

[0115] This embodiment first provides a given... Video-Text Dataset of Examples , .in and They represent the first One video and one text, This represents the label corresponding to the training data, where This indicates the number of categories in the dataset. If Belongs to the kind, ,otherwise The goal of the proposed method is to learn a unified... Bit hash codes serve as a compact representation of each video and text, and efficient retrieval can be performed using the learned hash codes.

[0116] Combination Figure 1 This embodiment uses a video-text cross-modal cue hash retrieval method based on global-local video attention, which includes the following steps:

[0117] Step 1: Use the visual-language pre-trained model CLIP as the encoder for video and text. Use visual-language guided video-text prompting technology to generate visually guided text prompts for text, generate language-guided local frame prompts for video, and design global video prompts.

[0118] Visual guidance is used to generate text cues that optimize text representation to better fit the video content. Language guidance is used to generate local frame cues that focus on relevant semantic regions and learn local frame information. Global video cues are used to learn and integrate video features with global video semantics. Specifically:

[0119] Step 1.1: Process the video into frame embeddings and the text into word embeddings, as follows:

[0120] Step 1.1.1: Sample the original video as... Each frame is divided into several frames. A fixed-size block is defined, and these blocks are linearly projected into an embedding, i.e. ,in This indicates the dimension of the frame embedding, and k represents the frame index;

[0121] Step 1.1.2: Segment the text to generate a text sequence, and then project the text sequence into the embedding space to generate word embeddings. ,in The length of the word tag, The dimension of word embedding;

[0122] Step 1.2: Generate visually guiding text prompts for the text by embedding video frames, as follows:

[0123] Step 1.2.1: Embed all frames in a video and aggregate them into a single frame using the time aggregation module;

[0124] Step 1.2.2: Send a single frame into the text prompt generation module. In the text prompt generation module, the single frame is first processed into text modulation parameters with video information. Then, the text modulation parameters are used for randomly initialized learnable text tags to finally obtain visually guided text prompts.

[0125] Step 1.3: Generate local frame cues with language guidance for the video through word embedding, and design global video cues for the video, as follows:

[0126] Step 1.3.1: The word embedding is sent to the visual cue generation module. In the visual cue generation module, the word embedding is first processed into visual modulation parameters with semantic information. Then, the visual modulation parameters are used for randomly initialized learnable visual tags to finally obtain language-guided local frame cue.

[0127] Step 1.3.2: Design shallow global video cues based on video characteristics and obtain global information of the video. Shallow global video cues are randomly initialized learnable visual tags.

[0128] Step 2: Integrate a global-local video attention mechanism into each layer of the Transformer in the visual encoder, which is divided into local frame attention and global video attention;

[0129] In text-video matching, the relevance between text and video content varies significantly. Some texts describe frame-level features or short segments, requiring precise local matching, while other texts summarize the overall video semantics, necessitating robust global representations. To simultaneously meet these two requirements, a dual-stream attention module is employed, combining local frame attention with global video attention, as detailed below:

[0130] Step 2.1, Design Global Video Attention: Global cues need to learn global discriminative information. Therefore, global cues need to pay attention to all video frame sequences and frame cue embeddings. For global video attention, the global cue is used as a query, and the global cue is concatenated with the CLS tag, frame cue, and frame sequence as a key and value. The process of global video attention is as follows:

[0131]

[0132]

[0133]

[0134] in, This indicates a global video prompt. The CLS tag representing the frame, Indicates frame hint embedding, Indicates frame embedding, Indicates the index of the frame;

[0135] Step 2.2, Designing Local Frame Attention: For local frame attention, each frame cue needs to perceive information from every local frame. Therefore, the CLS tag, frame cue, and frame sequence are concatenated as a query. To ensure that each frame cue can perceive global information, the global video cue needs to be repeated. Next, it is concatenated with the CLS tag, frame cue, and frame sequence as a key and value; the process of local frame attention is as follows:

[0136]

[0137]

[0138]

[0139] Step 3: Use text prompt learning and video prompt learning to obtain deep feature representations of text and video respectively. Then, input the video and text deep feature representations together into the adaptive feature fusion module to obtain optimized video and text deep feature representations.

[0140] Video is a sequence of video frames, and text is a sequence of words. CLIP's visual encoder and text encoder are both based on the Transformer network structure, which can effectively capture the correlation between sequences and fully utilize the natural cross-modal alignment capability of visual language pre-trained models to effectively bridge the semantic gap. Text cue learning and visual cue learning are applications of cue learning on CLIP. By concatenating cues with modal data and inputting them into the CLIP encoder, the CLIP backbone is frozen, and only cue parameters are trained, resulting in deep feature representations of video and text. These deep feature representations are then input together into an adaptive feature fusion module to obtain optimized video and text deep features. The adaptive feature fusion module can adaptively filter and weight redundant information in video and text features, effectively improving feature quality, as detailed below:

[0141] Step 3.1: Learn the deep feature representation of the text through text prompts, as follows:

[0142] Step 3.1.1: Input the word embeddings and text prompts as input into the CLIP text encoder. The The layer output is:

[0143]

[0144] in, This indicates the total number of layers in the CLIP text encoder. This indicates a text prompt. Indicates word embedding;

[0145] Step 3.1.2: By outputting the last layer [EOS] token Projecting onto the latent embedding space, the final text features are obtained:

[0146]

[0147] in This represents the final text features. Indicates the text projection layer;

[0148] Step 3.2: Learn the deep feature representation of the video through video prompts, as follows:

[0149] Step 3.2.1: Concatenate the frame embedding, global video cues, and frame cues with the learnable CLS tags for each frame, and input them together into the CLIP visual encoder. The The layer output is:

[0150]

[0151] in, This indicates a global video prompt. The CLS tag representing the frame, Indicates frame hint embedding, Indicates frame embedding, Indicates the index of the frame;

[0152] Step 3.2.2: The global cue perceives information from each frame, so it can be considered a combination of fine-grained frame representation and global video representation. Therefore, the first label of the global cue obtained from the last layer output is projected into the latent embedding space to obtain the final video features:

[0153]

[0154] in This represents the final video features. Indicates the video projection layer;

[0155] Step 3.3: Input the output video and text deep feature representations together into the adaptive feature fusion module to obtain the optimized video and text deep feature representations, as follows:

[0156] Step 3.3.1: Obtain the deep feature representations of the video and text sequentially through Global Response Normalization (GRN) and Multilayer Perceptron (MLP). and The results are then concatenated to obtain the fused features. This process can be represented as:

[0157]

[0158]

[0159] Where Concat represents the concatenation operation, and SiLU is the activation function;

[0160] Step 3.3.2: Flip the fused features and feed them into the State-Space Model (SSM). Then perform the inverse flipping operation to obtain the flipped features. A total of three flipping operations are performed to obtain the final three flipped features. The process can be represented as follows:

[0161]

[0162] in , A value of 1 indicates a left-right flip. A value of 2 indicates flipping vertically. (1,2) indicates that the elements are flipped simultaneously in all directions (up, down, left, and right). This indicates a flipping operation performed along different axes. This indicates its inverse operation;

[0163] Step 3.3.3: Perform adaptive weighted filtering on the three flipped features and the fused features to obtain the features. The process can be represented as:

[0164]

[0165] in Represents a set of learnable parameters. This represents a temperature coefficient used to balance the relative importance of flipped and non-flipped characteristics;

[0166] Step 3.3.4: Divide the features obtained after weighted filtering into two parts, and then compare them with... and Concatenate and sequentially pass through GRN and MLP to obtain and The process can be represented as:

[0167]

[0168]

[0169] Where chunk represents the operation of dividing the feature into two parts;

[0170] Step 3.3.5, will and After concatenation, the data is separated again through a Transformer encoder layer to obtain the final video and text features. and The process can be represented as:

[0171]

[0172] Split represents the operation of separating the final video features and text features from the fused features.

[0173] Step 4: Input the optimized video and text deep feature representations into the hash network respectively to obtain continuous real-valued representations of the video and text, and binarize them into a unified hash code, as follows:

[0174] Step 4.1: After the CLIP encoder, design a hash network consisting of an MLP and a hash layer. Input the obtained video and text deep feature representations into the hash network to obtain continuous real-valued representations of the video and text. and :

[0175]

[0176]

[0177] Step 4.2: Binarize the continuous real-valued representations of the video and text to obtain a unified binary hash code:

[0178]

[0179] in Represents a symbolic function.

[0180] Step 5: Design the overall objective function, including supervised intermodal contrast loss, quantization loss, and bit balance loss, and train the CLIP and hash networks as follows:

[0181] Step 5.1: Design a supervised inter-modal contrastive loss. Contrastive learning is based on a triple consisting of an anchor point, positive samples, and negative samples. Its goal is to group the anchor point with positive samples and separate the anchor point from negative samples. Using a supervised inter-modal contrastive loss can fully utilize label information for training while maintaining the similarity structure between modalities, as detailed below:

[0182] Given the first Video and text representation and Taking video modality as an example, set It is a collection of video indexes. Then the positive sample set is defined as ,in and They represent the first The and the first The tags corresponding to each video;

[0183] by The definition of inter-modal contrast loss as an anchor point is as follows:

[0184]

[0185] in Temperature coefficient;

[0186] Similarly, with The definition of inter-modal contrast loss as an anchor point is as follows:

[0187]

[0188] Therefore, the supervised modal contrast loss is:

[0189]

[0190] Step 5.2: Design quantization loss; In order to achieve the learning of unified hash codes, quantization loss is adopted, which plays a key role in bridging the gap between continuous hash representations and discrete hash codes;

[0191] set up , as well as The quantization loss is then defined as:

[0192]

[0193] Step 5.3: Design bit balance loss; In order to balance the hash code, most hashing methods require that the number of +1 and -1 in each bit is almost equal in all training samples. This constraint can maximize the amount of information provided by each bit.

[0194] Position balance loss is defined as:

[0195]

[0196] Step 5.4, the overall objective function is:

[0197]

[0198] in and It is a trade-off parameter for the relative importance of controlling losses.

[0199] Step 6: Use the trained model to perform cross-modal video-text retrieval.

[0200] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

[0201] It should be understood that, in order to simplify the present invention and help those skilled in the art understand its various aspects, in the above description of exemplary embodiments of the present invention, various features of the present invention are sometimes described in a single embodiment or with reference to a single figure. However, the present invention should not be construed as including all features in the exemplary embodiments as essential technical features of the claims of this patent.

Claims

1. A video-text cross-modal prompt hash retrieval method based on global-local video attention, characterized in that, Includes the following steps: Step 1: Use the visual-language pre-trained model CLIP as the encoder for video and text. Use visual-language guided video-text prompting technology to generate visually guided text prompts for text, generate language-guided local frame prompts for video, and design global video prompts. Step 2: Integrate a global-local video attention mechanism into each layer of the Transformer in the visual encoder, which is divided into local frame attention and global video attention; Step 3: Use text prompt learning and video prompt learning to obtain deep feature representations of text and video respectively. Then, input the video and text deep feature representations together into the adaptive feature fusion module to obtain optimized video and text deep feature representations. Step 4: Input the optimized video and text deep feature representations into the hash network to obtain continuous real-valued representations of the video and text, and binarize them into a unified hash code. Step 5: Design the overall objective function, including supervised intermodal contrast loss, quantization loss, and bit balance loss, and train the CLIP and hash networks; Step 6: Use the trained model to perform cross-modal video-text retrieval.

2. The video-text cross-modal prompt hash retrieval method based on global-local video attention according to claim 1, wherein, Step 1 specifically includes: Step 1.1: Process the video into frame embeddings and the text into word embeddings, as follows: Step 1.1.

1. Sample the original video into frames, then divide each frame into fixed-size patches and project them linearly into an embedding, i.e. where denotes the dimension of the frame embedding and k denotes the index of the frame. Step 1.1.

2. Tokenize the text to generate a text sequence, and then project the text sequence to an embedding space to generate word embeddings where is the length of the word token, is the dimension of the word embedding; Step 1.2: Generate visually guiding text prompts for the text by embedding video frames, as follows: Step 1.2.1: Embed all frames in a video and aggregate them into a single frame using the time aggregation module; Step 1.2.2: Send a single frame into the text prompt generation module. In the text prompt generation module, the single frame is first processed into text modulation parameters with video information. Then, the text modulation parameters are used for randomly initialized learnable text tags to finally obtain visually guided text prompts. Step 1.3: Generate local frame cues with language guidance for the video through word embedding, and design global video cues for the video, as follows: Step 1.3.1: The word embedding is sent to the visual cue generation module. In the visual cue generation module, the word embedding is first processed into visual modulation parameters with semantic information. Then, the visual modulation parameters are used for randomly initialized learnable visual tags to finally obtain language-guided local frame cue. Step 1.3.2: Design shallow global video cues based on video characteristics and obtain global information of the video. Shallow global video cues are randomly initialized learnable visual tags.

3. The video-text cross-modal prompt hash retrieval method based on global-local video attention according to claim 1, wherein, Step 2 describes integrating a global-local video attention mechanism into each layer of the visual encoder's Transformer, which is divided into local frame attention and global video attention, as detailed below: Step 2.1, Design Global Video Attention: Global cues need to learn global discriminative information. Therefore, global cues need to pay attention to all video frame sequences and frame cue embeddings. For global video attention, the global cue is used as a query, and the global cue is concatenated with the CLS tag, frame cue, and frame sequence as a key and value. The process of global video attention is as follows: ; ; ; wherein, represents a global video hint, represents a CLS flag of a frame, represents a frame hint embedding, represents a frame embedding, represents an index of a frame; Step 2.2, design local frame attention: for local frame attention, each frame prompt needs to perceive the information of each local frame, so the CLS token, frame prompt and frame sequence are connected as the query, in order to ensure that each frame prompt can perceive global information, the global video prompt needs to be repeated twice and connected with the CLS token, frame prompt and frame sequence as the key and value; the process of local frame attention is: ; ; 。 4. The video-text cross-modal cue hash retrieval method based on global-local video attention according to claim 1, characterized in that, Step 3 specifically includes: Step 3.1: Learn the deep feature representation of the text through text prompts; Step 3.2: Learn the deep feature representation of the video through video prompts; Step 3.3: Input the output video and text deep feature representations together into the adaptive feature fusion module to obtain the optimized video and text deep feature representations.

5. The video-text cross-modal cue hash retrieval method based on global-local video attention according to claim 4, characterized in that, Step 3.1, which involves learning the deep feature representation of the text through text prompts, is as follows: Step 3.1.1: Input the word embeddings and text prompts as input into the CLIP text encoder. The The layer output is: ; in, This indicates the total number of layers in the CLIP text encoder. This indicates a text prompt. Indicates word embedding; Step 3.1.2: By outputting the last layer [EOS] token Projecting onto the latent embedding space, the final text features are obtained: ; in This represents the final text features. This indicates the text projection layer.

6. The video-text cross-modal cue hash retrieval method based on global-local video attention according to claim 5, characterized in that, Step 3.2, which involves learning the deep feature representation of the video through video prompts, is as follows: Step 3.2.1: Concatenate the frame embedding, global video cues, and frame cues with the learnable CLS tags for each frame, and input them together into the CLIP visual encoder. The The layer output is: ; in, This indicates a global video prompt. The CLS tag representing the frame, Indicates frame hint embedding, Indicates frame embedding, Indicates the index of the frame; Step 3.2.2: The global cue perception of each frame is a combination of fine-grained frame and global video representation. Therefore, the first label of the global cue obtained from the last layer output is projected into the latent embedding space to obtain the final video features: ; in This represents the final video features. This represents the video projection layer.

7. The video-text cross-modal cue hash retrieval method based on global-local video attention according to claim 6, characterized in that, Step 3.3 involves inputting the output video and text deep feature representations together into the adaptive feature fusion module to obtain optimized video and text deep feature representations, as detailed below: Step 3.3.1: Obtain the deep feature representations of the video and text sequentially through Global Response Normalization (GRN) and Multilayer Perceptron (MLP). and The results are then concatenated to obtain the fused features. The process is represented as: ; ; Where Concat represents the concatenation operation, and SiLU is the activation function; Step 3.3.2: Flip the fused features and feed them into the State-Space Model (SSM). Then perform the inverse flipping operation to obtain the flipped features. A total of three flipping operations are performed to obtain the final three flipped features. The process is represented as follows: ; in , A value of 1 indicates a left-right flip. A value of 2 indicates flipping vertically. (1,2) indicates that the elements are flipped simultaneously in all directions (up, down, left, and right). This indicates a flipping operation performed along different axes. This indicates its inverse operation; Step 3.3.3: Perform adaptive weighted filtering on the three flipped features and the fused features to obtain the features. The process is represented as follows: ; in Represents a set of learnable parameters. This represents a temperature coefficient used to balance the relative importance of flipped and non-flipped characteristics; Step 3.3.4: Divide the features obtained after weighted filtering into two parts, and then compare them with... and Concatenate and sequentially pass through GRN and MLP to obtain and The process is represented as follows: ; ; Where chunk represents the operation of dividing the feature into two parts; Step 3.3.5, will and After concatenation, the data is separated again through a Transformer encoder layer to obtain the final video and text features. and The process is represented as follows: ; Split represents the operation of separating the final video features and text features from the fused features.

8. The video-text cross-modal cue hash retrieval method based on global-local video attention according to claim 1, characterized in that, Step 4 involves inputting the optimized video and text deep feature representations into a hash network to obtain continuous real-valued representations of the video and text, which are then binarized into a unified hash code, as detailed below: Step 4.1: After the CLIP encoder, design a hash network consisting of an MLP and a hash layer. Input the obtained video and text deep feature representations into the hash network to obtain continuous real-valued representations of the video and text. and : ; ; Step 4.2: Binarize the continuous real-valued representations of the video and text to obtain a unified binary hash code: ; in Represents a symbolic function.

9. The video-text cross-modal cue hash retrieval method based on global-local video attention according to claim 1, characterized in that, Step 5 describes the design of the overall objective function, which includes supervised intermodal contrast loss, quantization loss, and bit balance loss. The CLIP and hash networks are trained as follows: Step 5.1: Design the supervised inter-modal contrast loss, as follows: Given the first Video and text representation and Taking video modality as an example, set It is a collection of video indexes. Then the positive sample set is defined as ,in and They represent the first The and the first The tags corresponding to each video; by The definition of inter-modal contrast loss as an anchor point is as follows: ; in Temperature coefficient; by The definition of inter-modal contrast loss as an anchor point is as follows: ; Therefore, the supervised modal contrast loss is: ; Step 5.2: Design the quantification loss, as follows: set up , as well as The quantization loss is then defined as: ; Step 5.3: Design the bit balance loss, which is defined as: ; Step 5.4, the overall objective function is: ; in and It is a trade-off parameter for the relative importance of controlling losses.

10. An electronic device, characterized in that, include: A memory and a processor are communicatively connected, the memory storing computer instructions, and the processor executing the computer instructions to perform the video-text cross-modal cue hash retrieval method based on global-local video attention as described in any one of claims 1 to 9.