A text-video retrieval method inspired by the human brain's episodic memory pathway

Through a multi-grained information fusion method inspired by the human brain episodic memory pathway, multi-scale representations and tokens are extracted using content encoding components and context encoding components, and multi-modal information fusion is combined with hyperbolic graph neural networks, which solves the problems of multi-modal data retrieval accuracy and inefficiency in the existing technology, and achieves efficient and accurate matching of text and video retrieval.

CN119938985BActive Publication Date: 2025-07-01NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510416357.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-01
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

Existing multimodal data retrieval technology cannot achieve multi-grained alignment during cross-modal alignment, resulting in low retrieval accuracy and efficiency, especially when processing video data, the computing resource requirements are high.

Method used

The text video retrieval method inspired by the human brain episodic memory pathway is adopted, and multi-scale representations and tokens are extracted through content encoding components and context encoding components, and multi-grained information fusion is combined with hyperbolic neural network to capture the complex relationship between text and video.

Benefits of technology

It significantly improves the accuracy and efficiency of text video retrieval, and enhances the robustness and generalization capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938985B_ABST
    Figure CN119938985B_ABST
Patent Text Reader

Abstract

The present invention discloses a text-video retrieval method inspired by the human brain's episodic memory pathway. The method includes using a content encoding component to extract content representations from target text data or target video data to obtain multi-scale target representations; using a context encoding component to extract context representations from target text data or target video data to obtain target tokens; inputting the multi-scale target representations and target tokens into a hyperbolic graph neural network to obtain target scene representations; using the target scene representations as target indexes; calculating the similarity between the representations of the text or video to be retrieved and the target indexes, and screening the text or video to be retrieved according to the similarity to obtain target retrieval results. The present invention comprehensively captures multi-level semantic features through multi-granularity information fusion, and fuses multi-modal and multi-granularity high-order information through hyperbolic graph convolution operations, which can better capture the complex relationships between text and video, and significantly improve the accuracy and efficiency of text-video retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network model analysis, and particularly relates to a text-video retrieval method for multi-granularity information fusion. Background Art

[0002] With the rapid development of the Internet, the quantity of multimodal data (such as text, images, videos, etc.) has grown explosively. How to efficiently and accurately retrieve the information required by users from the vast amount of multimodal data has become an important research direction. Among them, the text-video retrieval task is particularly challenging because it needs to simultaneously process two highly heterogeneous data modalities, namely text and video.

[0003] Multimodal data retrieval is an information retrieval method involving multiple media modalities (such as text, images, audio, videos, etc.). Current multimodal retrieval technologies mainly convert data into vector representations through deep learning models and extract common features through modality fusion, and sort the retrieval results through similarity measurement. However, existing methods can often only process coarse-grained or fine-grained information during the cross-modal alignment process, and cannot achieve multi-granularity alignment, resulting in insufficient cross-modal alignment. At the same time, video data has the characteristics of high dimensionality and high redundancy. Existing methods often require a large amount of computing resources when processing video data and cannot improve the retrieval efficiency while ensuring the retrieval accuracy. Summary of the Invention

[0004] The present invention provides a text-video retrieval method inspired by the human brain's episodic memory pathway. By comprehensively capturing multi-level semantic features in text and video through multi-granularity information fusion, and fusing multi-modal and multi-granularity high-order information through hyperbolic graph convolution operations, it can better capture the complex relationships between text and video, and significantly improve the accuracy and efficiency of text-video retrieval.

[0005] To achieve the above object, the technical solution adopted by the present invention is:

[0006] The first aspect of the present invention provides a text-video retrieval method inspired by the human brain's episodic memory pathway, including:

[0007] Obtain target text data or target video data and input it into a text-video retrieval model, where the text-video retrieval model includes a content encoding component, a context encoding component, and a hyperbolic graph neural network;

[0008] Use the content encoding component to extract content representations of the target text data or target video data to obtain multi-scale target text representations or multi-scale target visual representations;

[0009] Use the context encoding component to extract context representations of the target text data or target video data to obtain target text tokens or target visual tokens;

[0010] Input the multi-scale target text representation and target text tokens into a hyperbolic graph neural network to obtain the target text scene representation; or input the multi-scale target visual representation and target visual tokens into a hyperbolic graph neural network to obtain the target visual scene representation; use the target text scene representation or the target visual scene representation as the target index;

[0011] Calculate the similarity between the representation of the text or video to be retrieved and the target index, and screen the text or video to be retrieved according to the similarity to obtain the target retrieval result.

[0012] Further, the training process of the text-video retrieval model includes:

[0013] Obtain text training data and video training data and input them into the content encoding component to obtain a word matrix mask, text event representation, text semantic unit representation, visual event representation, and visual semantic unit representation;

[0014] Input the video training data and the text training data and the word matrix mask into the context encoding component respectively to obtain text token representation and visual token representation;

[0015] Map the text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation, and visual token representation as node features to the hyperbolic space to construct an adjacency matrix, input the adjacency matrix and node features into the hyperbolic graph neural network, and obtain the text scene representation and visual scene representation through hyperbolic graph convolution operations and pooling operations; calculate the training loss value according to the text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation, visual token representation, text scene representation, and visual scene representation; optimize the weight parameters of the text-video retrieval model according to the training loss value, and repeat the iteration until the training termination condition is reached, and output the trained text-video retrieval model.

[0016] Further, obtaining text training data and video training data and inputting them into the content encoding component to obtain a word matrix mask, text event representation, text semantic unit representation, visual event representation, and visual semantic unit representation; specifically includes:

[0017] The content encoding component includes a first content encoding component, a second content encoding component, and a third content encoding component;

[0018] Input the text training data and the video training data into the first content encoding component to obtain text global representation and visual global representation;

[0019] Obtain phrases and a word matrix mask from the text training data through a syntactic analyzer; input the phrases into the second content encoding component to obtain text semantic unit representation;

[0020] Segment the visual global representation into visual semantic unit representations through the K-means algorithm;

[0021] After inputting the text semantic unit representation and the visual semantic unit representation into the third content encoding component, add them to the text global representation and the visual global representation to obtain the text event representation and the visual event representation.

[0022] Further, after inputting the text semantic unit representation and the visual semantic unit representation into the third content encoding component, add them to the text global representation and the visual global representation to obtain the text event representation and the visual event representation, specifically including:

[0023] The third content encoding component includes an event visual encoder and an event text encoder;

[0024] After performing layer normalization on the visual semantic unit representation, input it into the multi-head attention layer in the event visual encoder to obtain visual event extraction features, and after performing layer normalization on the visual event extraction features, input them into the multi-layer perceptron in the event visual encoder to obtain visual event perception features;

[0025] After performing average pooling on the visual global representation, concatenate it with the visual event perception features and the visual event extraction features to obtain the visual event representation;

[0026] After performing layer normalization on the text semantic unit representation, input it into the multi-head attention layer in the event text encoder to obtain text event extraction features, and after performing layer normalization on the text event extraction features, input them into the multi-layer perceptron in the event text encoder to obtain text event perception features;

[0027] After adding a classification token to the text global representation, concatenate it with the text event perception features and the text event extraction features to obtain the text event representation.

[0028] Further, input the video training data into the context encoding component to obtain visual tokens, specifically including:

[0029] The context encoding component includes a context visual encoder; input the video training data into the context visual encoder, perform layer normalization on the video training data to obtain visual standard data, and add a classification label to the visual standard data to obtain visual initial tokens;

[0030] Move the visual initial token back and forth along the direction of the video frame sequence, and input it into the multi-head attention layer in the context visual encoder to obtain the visual extraction token. Concatenate the visual extraction token with the visual initial token to obtain the visual fusion token; perform layer normalization on the visual fusion token and then input it into the multi-layer perceptron in the context visual encoder to obtain the first visual perception token; concatenate the first visual perception token with the visual fusion token to obtain the visual refinement token;

[0031] Perform importance scoring on the visual refinement token through the token selection layer in the context visual encoder, and then select the top K visual refinement tokens in each video frame as the visual key tokens according to the importance scoring;

[0032] Perform layer normalization on the visual key tokens, and then input them into the multi-head attention layer in the context visual encoder to obtain the visual key refinement tokens. Concatenate the visual key refinement tokens with the visual key tokens to obtain the visual key fusion tokens; perform layer normalization on the visual key fusion tokens and then input them into the multi-layer perceptron in the context visual encoder to obtain the second visual perception token, and then concatenate the second visual perception token with the visual key fusion tokens to obtain the visual tokens.

[0033] Further, perform importance scoring on the visual refinement token through the token selection layer in the context visual encoder, and then select the top K visual refinement tokens in each video frame as the visual key tokens according to the importance scoring; specifically including:

[0034] Input the visual refinement token into the multi-layer perceptron in the token selection layer, and compress the visual refinement token to a set ratio to obtain the first visual compression token;

[0035] Add a classification marker to the first visual compression token and then input it into the multi-layer perceptron in the token selection layer again to obtain the second visual compression token;

[0036] Perform Softmax function calculation on the second visual compression token to obtain the importance scoring, and then select the top K visual refinement tokens in each video frame as the visual key tokens according to the importance scoring.

[0037] Further, input the text training data and the word matrix mask into the context encoding component to obtain the text tokens, specifically including:

[0038] The context encoding component includes a first neural network architecture and a second neural network architecture;

[0039] Input the text training data into the first neural network architecture. After performing layer normalization on the text training data, input it into the multi-head attention layer within the first neural network architecture to obtain the first text extraction token. Concatenate the first text extraction token with the text training data to obtain the first text fusion token; perform layer normalization on the first text fusion token and then input it into the multi-layer perceptron within the first neural network architecture to obtain the first text perception token; concatenate the first text perception token with the first text fusion token to obtain the text refinement token.

[0040] Input the text refinement token into the second neural network architecture. Perform layer normalization on the text refinement token to obtain the text normalization token. Input the text normalization token and the word matrix mask into the multi-head attention layer within the second neural network architecture to obtain the second text extraction token. Concatenate the second text extraction token with the text refinement token to obtain the second text fusion token; perform layer normalization on the second text fusion token and then input it into the multi-layer perceptron within the second neural network architecture to obtain the second text perception token. Concatenate the second text perception token with the second text fusion token to obtain the text token.

[0041] Further, map the text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation, and visual token representation as node features to the hyperbolic space to construct the adjacency matrix, specifically including:

[0042] Map the visual event representation and the text event representation to the node features at the first-level granularity in the hyperbolic space; map the visual semantic unit representation and the text semantic unit representation to the node features at the second-level granularity in the hyperbolic space; map the visual token representation and the text token representation to the node features at the third-level granularity in the hyperbolic space;

[0043] Connect the nodes with the same-level granularity, and establish connections between each node feature at the second-level granularity and all node features at the first-level granularity; establish connections between the node features at the second-level granularity and the node features at the third-level granularity according to the semantic subordination relationship; when there is a connection between the th node feature and the th node feature, the connection edge ; otherwise, the connection edge ; establish the adjacency matrix according to the connection relationship between the node features .

[0044] Further, input the adjacency matrix and the node features into the hyperbolic graph neural network, and obtain the text scene representation and the visual scene representation through hyperbolic graph convolution operations and pooling operations, specifically including:

[0045] Capture the hyperbolic space hidden representation by performing feature transformation on the node features. The calculation formula is:

[0046]

[0047]

[0048]

[0049] Among them, represents the hyperbolic space hidden representation of the -th node feature of the -th layer, is the representation mapping function from Euclidean space to hyperbolic space; represents the Euclidean space hidden representation of the -th node feature of the -th layer, is the hyperbolic tangent function, is the inverse hyperbolic tangent function, represents the learnable parameter of the -th layer, represents the curvature of the -th layer in hyperbolic space, is the -th node feature of the -th layer in hyperbolic space; is the representation mapping function from hyperbolic space to Euclidean space;

[0050] The hyperbolic space aggregation representation is obtained by aggregating node features according to the adjacency matrix, and the expression formula is:

[0051]

[0052]

[0053]

[0054] Among them, represents the hyperbolic space aggregation representation of the -th node feature of the -th layer, is the node information aggregation function; represents the neighbor node set of the -th node feature, represents the aggregation weight between the -th node feature and the -th node feature, [ ; ] represents the tensor concatenation operation, is the learnable matrix; represents the hyperbolic space hidden representation of the -th node feature of the Distance from the hidden representation in hyperbolic space ; Denote the curvature of the -th layer in hyperbolic space; is the hyperbolic function; is the leaky rectified linear unit; is the logistic function;

[0055] Input the aggregated representation in hyperbolic space into the activation function to obtain the hyperbolic space representation, and the expression formula is:

[0056]

[0057] where, Denote the hyperbolic space representation of the -th node in the -th layer; is the activation function of the hyperbolic graph neural network;

[0058] Perform a pooling operation on the hyperbolic space representation to obtain the text scene representation and the video scene representation.

[0059] Furthermore, calculate the training loss value according to the text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation and visual token representation, text scene representation and visual scene representation, specifically including:

[0060] Calculate the event retrieval loss, unit representation retrieval loss, token retrieval loss and scene retrieval loss respectively according to the text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation, visual token representation, text scene representation and visual scene representation;

[0061] Add the parent-child relationship between the node features at each level of granularity, and calculate the hierarchical structure loss of the hyperbolic space, and the expression formula is:

[0062]

[0063] In the formula, is the hyperbolic representation of the text child node, is the hyperbolic representation of the visual child node, is the hyperbolic representation of the text father node, Denote the hyperbolic representation of the visual father node, is the hyperbolic representation to the hyperbolic representation Distance loss between; is the hyperbolic representation and the hyperbolic representation Distance loss between; A set of node features indicating the existence of a parent-child relationship, Indicates the set of nodes that have no parent-child relationship with the th node feature; Is the hyperbolic representation To the hyperbolic representation The positional loss between; Is the hyperbolic representation And the hyperbolic representation The positional loss between; Is a hyperparameter, Is the second norm, Is to take the maximum value; Is the serial number of the text child node or visual child node; Is the serial number of the text father node or visual father node;

[0064] According to the event retrieval loss, unit representation retrieval loss, token retrieval loss, scene retrieval loss, distance loss 、Distance loss 、Positional loss And positional loss Calculate the training loss value.

[0065] Compared with the prior art, the beneficial effects of the present invention are:

[0066] In the present invention, a content encoding component is used to extract content representations from target text data or target video data to obtain multi-scale target text representations or multi-scale target visual representations; a context encoding component is used to extract context representations from target text data or target video data to obtain target text tokens or target visual tokens; through multi-granularity information fusion, multi-level semantic features in text and video are comprehensively captured, significantly improving the accuracy of text-video retrieval.

[0067] In the present invention, the multi-scale target text representation and the target text token are input into a hyperbolic graph neural network to obtain a target text scene representation; or the multi-scale target visual representation and the target visual token are input into a hyperbolic graph neural network to obtain a target visual scene representation; the target text scene representation or the target visual scene representation is used as a target index; through hyperbolic graph convolution operations, multi-modal and multi-granularity high-order information is fused, which can better capture the complex relationship between text and video, and enhance the robustness and generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 Is a flowchart of a text-video retrieval method based on the inspiration of the human brain's episodic memory pathway provided by Embodiment 1 of the present invention;

[0069] Figure 2 Is a structural diagram of the first content encoding component provided by Embodiment 2 of the present invention;

[0070] Figure 3 It is the structural diagram of the third content encoding component provided in Embodiment 2 of the present invention

[0071] Figure 4 It is the structural diagram of the scenario visual encoder provided in Embodiment 2 of the present invention;

[0072] Figure 5 It is the structural diagram of the scenario text encoder provided in Embodiment 2 of the present invention;

[0073] Figure 6 It is the structural diagram of the hyperbolic graph convolutional neural network provided in Embodiment 2 of the present invention;

[0074] Figure 7 It is the schematic diagram of the Poincaré disk provided in Embodiment 2 of the present invention. Detailed implementation manners

[0075] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention, and cannot be used to limit the protection scope of the present invention.

[0076] Brain-inspired computing is an emerging research direction in the field of artificial intelligence. The core lies in drawing on the information processing mode or structure of the biological nervous system, and then constructing corresponding computing theories, chip architectures, and application models and algorithms. In recent years, cognitive science has made objective progress in the research on the human brain's episodic memory pathway. The human brain's episodic memory pathway decomposes external perceptual signals into information of multiple granularities and then fuses them into complex scene representations, which is somewhat different from the traditional encoding, alignment, and retrieval methods for text and video in the field of artificial intelligence, providing a new reference for the model design of text and video retrieval tasks.

[0077] Inspired by the human brain's episodic memory pathway, the present invention comprehensively captures multi-level semantic features in text and video through multi-granularity information fusion, and fuses multi-modal and multi-granularity high-order information through hyperbolic graph convolution operations, which can better capture the complex relationships between text and video, and significantly improve the accuracy and efficiency of text and video retrieval.

[0078] Embodiment 1

[0079] As Figure 1 shown, this embodiment provides a text and video retrieval method inspired by the human brain's episodic memory pathway, including:

[0080] Obtain target text data or target video data and input it into the text and video retrieval model to obtain a target index; calculate the similarity between the representation of the text or video to be retrieved and the target index, and screen the text or video to be retrieved according to the similarity to obtain the target retrieval result; specifically including:

[0081] Obtain target text data or target video data and input it into a text-video retrieval model, where the text-video retrieval model includes a content encoding component, a context encoding component, and a hyperbolic graph neural network;

[0082] Use the content encoding component to extract content representations from the target text data or target video data to obtain multi-scale target text representations or multi-scale target visual representations; the multi-scale target text representations include target text event representations, target text global representations, and target text semantic unit representations; the multi-scale target visual representations include target visual event representations, target visual global representations, and target visual semantic unit representations.

[0083] Use the context encoding component to extract context representations from the target text data or target video data to obtain target text tokens or target visual tokens;

[0084] Input the multi-scale target text representations and target text tokens into the hyperbolic graph neural network to obtain target text scene representations; or input the multi-scale target visual representations and target visual tokens into the hyperbolic graph neural network to obtain target visual scene representations; use the target text scene representations or target visual scene representations as target indexes;

[0085] Calculate the similarity between the representation of the text or video to be retrieved and the target index, and screen the text or video to be retrieved according to the similarity to obtain target retrieval results. The target retrieval results include relevant videos and relevant texts; that is, retrieve relevant videos according to the target text data; retrieve relevant texts according to the target video data.

[0086] The text-video retrieval model includes a content encoding component, a context encoding component, and a hyperbolic graph neural network; the training process of the text-video retrieval model includes:

[0087] Obtain text training data and video training data and input them into the content encoding component to obtain a word matrix mask, text event representations, text semantic unit representations, visual event representations, and visual semantic unit representations; specifically including:

[0088] The content encoding component includes a first content encoding component, a second content encoding component, and a third content encoding component;

[0089] Input the text training data and video training data into the first content encoding component to obtain text global representations and visual global representations;

[0090] Obtain phrases and a word matrix mask from the text training data through a syntactic analyzer; input the phrases into the second content encoding component to obtain text semantic unit representations;

[0091] Segment the visual global representation into visual semantic unit representations through the K-means algorithm;

[0092] After inputting the text semantic unit representation and the visual semantic unit representation into the third content encoding component, add them to the text global representation and the visual global representation to obtain the text event representation and the visual event representation.

[0093] Input the video training data, the text training data, and the word matrix mask into the context encoding component respectively to obtain the text token representation and the visual token representation;

[0094] Map the text event representation, the text semantic unit representation, the visual event representation, the visual semantic unit representation, the text token representation, and the visual token representation as node features to the hyperbolic space to construct an adjacency matrix. Input the adjacency matrix and the node features into the hyperbolic graph neural network. Through hyperbolic graph convolution operations and pooling operations, obtain the text scene representation and the visual scene representation; Calculate the training loss value according to the text event representation, the text semantic unit representation, the visual event representation, the visual semantic unit representation, the text token representation, the visual token representation, the text scene representation, and the visual scene representation; Optimize the weight parameters of the text-video retrieval model according to the training loss value, and repeat the iteration until the training termination condition is reached, and output the trained text-video retrieval model.

[0095] Embodiment 2

[0096] As Figures 2 to 5 shown, this embodiment provides a text-video retrieval method inspired by the human brain's episodic memory pathway, including:

[0097] The text-video retrieval model includes a content encoding component, a context encoding component, and a hyperbolic graph neural network; The training process of the text-video retrieval model includes:

[0098] Obtain the text training data and the video training data and input them into the content encoding component to obtain the word matrix mask, the text event representation, the text semantic unit representation, the visual event representation, and the visual semantic unit representation; The content encoding component includes a first content encoding component, a second content encoding component, and a third content encoding component; In this embodiment, the first content encoding component, the second content encoding component, and the third content encoding component are Content CLIP (Content Contrastive Language-Image Pretrained Model);

[0099] Input the text training data and the video training data into the first content encoding component to obtain the text global representation and the visual global representation, specifically including:

[0100] As Figure 2 shown, the first content encoding component includes a convolutional neural network, a global vision, and a global text encoder;

[0101] Extract a sequence of image patches from video training data through a convolutional neural network; after performing layer normalization on the sequence of image patches, input it into the multi-head attention layer in the global visual encoder to obtain global visual extraction features, and obtain global visual fusion features by concatenating the global visual extraction features with the sequence of image patches. After performing layer normalization on the global visual fusion features, input them into the multi-layer perceptron in the global visual encoder to obtain global visual perception features; concatenate the global visual perception features with the global visual fusion features to obtain a visual global representation;

[0102] After performing layer normalization on the text training data, input it into the multi-head attention layer in the global text encoder to obtain global text extraction features, and obtain global text fusion features by concatenating the global text extraction features with the text training data. After performing layer normalization on the global text fusion features, input them into the multi-layer perceptron in the global text encoder to obtain global text perception features; concatenate the global text perception features with the global text fusion features to obtain a text global representation.

[0103] Convert the visual global representation into a visual semantic unit representation through the K-means algorithm; obtain phrases and word matrix masks from the text training data through a syntactic analyzer; input the phrases into the second content encoding component to obtain text semantic unit representations; the second content encoding component includes a unit text encoder; a multi-layer perceptron and a multi-head attention layer are configured in the unit text encoder.

[0104] After inputting the text semantic unit representation and the visual semantic unit representation into the third content encoding component, and adding them to the text global representation and the visual global representation to obtain a text event representation and a visual event representation, specifically including:

[0105] The third content encoding component includes an event visual encoder and an event text encoder;

[0106] After performing layer normalization on the visual semantic unit representation, input it into the multi-head attention layer in the event visual encoder to obtain visual event extraction features. After performing layer normalization on the visual event extraction features, input them into the multi-layer perceptron in the event visual encoder to obtain visual event perception features;

[0107] After performing average pooling on the visual global representation, concatenate it with the visual event perception features and the visual event extraction features to obtain a visual event representation;

[0108] After performing layer normalization on the text semantic unit representation, input it into the multi-head attention layer in the event text encoder to obtain text event extraction features. After performing layer normalization on the text event extraction features, input them into the multi-layer perceptron in the event text encoder to obtain text event perception features;

[0109] After adding classification tokens to the global representation of the text, it is concatenated with the text event perception features and the text event extraction features to obtain the text event representation.

[0110] The described scenario encoding component includes a scenario visual encoder and a scenario text encoder; the scenario visual encoder and the scenario text encoder are Context CLIP (Context Contrastive Language-Image Model).

[0111] Inputting the video training data into the scenario visual encoder to obtain visual tokens, specifically including:

[0112] Inputting the video training data into the scenario visual encoder, performing layer normalization on the video training data to obtain visual standard data, and adding classification tokens to the visual standard data to obtain visual initial tokens;

[0113] Moving the visual initial tokens forward and backward along the direction of the video frame sequence, and inputting them into the multi-head attention layer in the scenario visual encoder to obtain visual extraction tokens, concatenating the visual extraction tokens with the visual initial tokens to obtain visual fusion tokens; performing layer normalization on the visual fusion tokens and then inputting them into the multi-layer perceptron in the scenario visual encoder to obtain the first visual perception tokens; concatenating the first visual perception tokens with the visual fusion tokens to obtain visual refinement tokens;

[0114] Inputting the visual refinement tokens into the token selection layer in the scenario visual encoder, compressing the visual refinement tokens to a set ratio through the multi-layer perceptron in the token selection layer to obtain the first visual compression tokens; adding classification tokens to the visual compression tokens and then inputting them again into the multi-layer perceptron in the token selection layer to obtain the second visual compression tokens; performing Softmax function calculation on the second visual compression tokens to obtain importance scores, and then selecting the top K visual refinement tokens (Top K) in each video frame as visual key tokens according to the importance scores;

[0115] Performing layer normalization on the visual key tokens, and inputting them into the multi-head attention layer in the scenario visual encoder to obtain visual key refinement tokens, concatenating the visual key refinement tokens with the visual key tokens to obtain visual key fusion tokens; performing layer normalization on the visual key fusion tokens and then inputting them into the multi-layer perceptron in the scenario visual encoder to obtain the second visual perception tokens, and then concatenating the second visual perception tokens with the visual key fusion tokens to obtain visual tokens.

[0116] Inputting the text training data and the word matrix mask into the scenario text encoder to obtain text tokens, specifically including:

[0117] The described scenario text encoder includes a first neural network architecture and a second neural network architecture; the first neural network architecture and the second neural network architecture are Transformer neural network architectures;

[0118] Input the text training data into the first neural network architecture. After performing layer normalization on the text training data, input it into the multi-head attention layer within the first neural network architecture to obtain the first text extraction token. Concatenate the first text extraction token with the text training data to obtain the first text fusion token; perform layer normalization on the first text fusion token and then input it into the multi-layer perceptron within the first neural network architecture to obtain the first text perception token; concatenate the first text perception token with the first text fusion token to obtain the text refinement token.

[0119] Input the text refinement token into the second neural network architecture. Perform layer normalization on the text refinement token to obtain the text normalization token. Input the text normalization token and the word matrix mask into the multi-head attention layer within the second neural network architecture to obtain the second text extraction token. Concatenate the second text extraction token with the text refinement token to obtain the second text fusion token; perform layer normalization on the second text fusion token and then input it into the multi-layer perceptron within the second neural network architecture to obtain the second text perception token. Concatenate the second text perception token with the second text fusion token to obtain the text token.

[0120] Map the text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation, and visual token representation as node features to the hyperbolic space to construct an adjacency matrix, specifically including:

[0121] As Figure 7 shown, in this embodiment, the hyperbolic space is represented by the Poincaré disk, with the space centered at the origin and the outer space capacity growing exponentially.

[0122] Map the visual event representation and the text event representation to the node features of the first-level granularity in the hyperbolic space; map the visual semantic unit representation and the text semantic unit representation to the node features of the second-level granularity in the hyperbolic space; map the visual token representation and the text token representation to the node features of the third-level granularity in the hyperbolic space;

[0123] Connect the nodes with the same-level granularity, and establish connections between each node feature of the second-level granularity and all node features of the first-level granularity; establish connections between the node features of the second-level granularity and the node features of the third-level granularity according to the semantic subordination relationship; when there is a connection between the th node feature and the th node feature, the connection edge ; otherwise, the connection edge ; establish an adjacency matrix according to the connection relationships between the node features.

[0124] As Figure 6As shown, the adjacency matrix and node features are input into the hyperbolic graph neural network, and the text scene representation and visual scene representation are obtained through hyperbolic graph convolution operations and pooling operations, specifically including:

[0125] Capture the hyperbolic space hidden representation by performing feature transformation on the node features. The calculation formula is:

[0126]

[0127]

[0128]

[0129] Among them, represents the hyperbolic space hidden representation of the -th node feature in the -th layer. is the representation mapping function from the Euclidean space to the hyperbolic space; represents the Euclidean space hidden representation of the -th node feature in the -th layer. is the hyperbolic tangent function. is the inverse hyperbolic tangent function. represents the learnable parameter of the -th layer. represents the curvature of the -th layer in the hyperbolic space. is the -th node feature in the -th layer in the hyperbolic space; is the representation mapping function from the hyperbolic space to the Euclidean space;

[0130] Aggregate the node features according to the adjacency matrix to obtain the hyperbolic space aggregation representation. The expression formula is:

[0131]

[0132]

[0133]

[0134] Among them, represents the hyperbolic space aggregation representation of the -th node feature in the -th layer. is the node information aggregation function; represents the neighbor node set of the -th node feature. represents the aggregation weight between the i-th node feature and the j-th node feature, and [ ; ] represents the tensor concatenation operation. is a learnable matrix; represents the in the hyperbolic space hidden representation of the -th node feature of the -th layer; is the distance between the hyperbolic space hidden representation and the hyperbolic space hidden representation ; represents the hyperbolic function; is the leaky rectified linear unit function; is the logistic function;

[0135] The hyperbolic space aggregation representation is input into the activation function to obtain the hyperbolic space representation, and the expression formula is:

[0136]

[0137] where, represents the -th layer of the -th node's hyperbolic space representation; is the activation function of the hyperbolic graph neural network;

[0138] Perform a pooling operation on the hyperbolic space representation to obtain the text scene representation and the visual scene representation.

[0139] Align the visual event representation and the text event representation and calculate to obtain the first similarity; calculate the event retrieval loss according to the first similarity, and the event retrieval loss includes the retrieval loss from the visual event representation to the text event representation and the retrieval loss from the text event representation to the visual event representation; the expression formula is:

[0140]

[0141] In the formula, is the normalization coefficient; is the natural exponential function; B is the batch size; is the retrieval loss from the text event representation to the visual event representation; is the retrieval loss from the visual event representation to the text event representation; is the positive sample similarity between the text event representation and the visual event representation; is the negative sample similarity of retrieving the visual event representation from the text event representation; is the negative sample similarity of retrieving the text event representation from the visual event representation;

[0142] Align the visual semantic units and the text semantic units and calculate to obtain the second similarity; calculate the semantic unit retrieval loss according to the second similarity; the semantic unit retrieval loss includes the retrieval loss from the visual semantic units to the text semantic units and the retrieval loss from the text semantic units to the visual semantic units; the expression formula is:

[0143]

[0144] In the formula, is the retrieval loss from the text semantic units to the visual semantic units; is the retrieval loss from the text semantic units to the visual semantic units; is the positive sample similarity between the text semantic units and the visual semantic units; is the negative sample similarity for retrieving visual semantic units from text semantic units; is the negative sample similarity for retrieving text semantic units from visual semantic units.

[0145] Align the text token representations and the visual token representations and calculate to obtain the third similarity; calculate the token retrieval loss according to the third similarity; the token retrieval loss includes the retrieval loss from the text tokens to the visual tokens and the retrieval loss from the visual tokens to the text tokens; the expression formula is:

[0146]

[0147] In the formula, is the retrieval loss from the text tokens to the visual tokens; is the retrieval loss from the visual tokens to the text tokens; is the positive sample similarity between the text tokens and the visual tokens; is the negative sample similarity for retrieving visual tokens from text tokens; is the negative sample similarity for retrieving text tokens from visual tokens.

[0148] Align the text scene representations and the visual scene representations and calculate to obtain the fourth similarity; calculate the scene retrieval loss according to the fourth similarity, and the scene retrieval loss includes the retrieval loss from the text scene representations to the visual scene representations and the retrieval loss from the visual scene representations to the text scene representations; the expression formula is:

[0149]

[0150] In the formula, is the retrieval loss from the text scene representations to the visual scene representations; is the retrieval loss from the visual scene representations to the text scene representations; is the positive sample similarity between the visual scene representations and the text scene representations; Retrieve the negative sample similarity of the visual scene representation from the text scene representation; Retrieve the negative sample similarity of the text scene representation from the visual scene representation.

[0151] Add parent-child relationships between node features at various levels of granularity, calculate the hierarchical structure loss in the hyperbolic space, and the expression formula is:

[0152]

[0153] In the formula, is the hyperbolic representation of the text child node, is the hyperbolic representation of the visual child node, is the hyperbolic representation of the text parent node, represents the hyperbolic representation of the visual parent node, is the hyperbolic representation to the hyperbolic representation the distance loss between; is the hyperbolic representation and the hyperbolic representation the distance loss between; represents the set of node features with parent-child relationships, represents the set of nodes that do not have a parent-child relationship with the th node feature; is the hyperbolic representation to the hyperbolic representation the position loss between; is the hyperbolic representation and the hyperbolic representation the position loss between; is a hyperparameter, is the second norm, is to take the maximum value; is the serial number of the text child node or the visual child node; is the serial number of the text parent node or the visual parent node;

[0154] According to the event retrieval loss, semantic unit retrieval loss, token retrieval loss, scene retrieval loss, distance loss distance loss position loss and position loss calculate the training loss value.

[0155] Optimize the weight parameters of the text-video retrieval model according to the training loss value, and repeat the iteration until the training termination condition is reached, and output the trained text-video retrieval model.

[0156] In the inference stage, this embodiment adopts a two-stage retrieval strategy. First, only the content encoding component is activated, and the target text data or target video data is input into the content encoding component to extract the semantic unit representation and event representation; the semantic unit representation and event representation are used as the target index to quickly screen out the candidate set in the video database;

[0157] The text or video to be retrieved in the candidate set is sent into the text-video retrieval model again to calculate the similarity between the text or video to be retrieved in the candidate set and the target index, and the text or video to be retrieved in the candidate set is screened according to the similarity to obtain the target retrieval result and re-rank; this embodiment not only ensures the retrieval accuracy but also ensures the model efficiency.

[0158] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0159] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0160] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device realizes the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0161] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, thereby providing instructions for implementing the steps specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 in one block or a plurality of blocks.

[0162] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A text-video retrieval method inspired by human brain episodic memory pathways, characterized in that: include: Obtaining target text data or target video data and inputting the data into a text-video retrieval model, wherein the text-video retrieval model includes a content encoding component, a context encoding component, and a hyperbolic graph neural network; Using the content encoding component to extract content representation of target text data or target video data to obtain multi-scale target text representation or multi-scale target visual representation; Using the context encoding component to extract context representation of target text data or target video data to obtain target text tokens or target visual tokens; Input the multi-scale target text representation and the target text token into the hyperbolic graph neural network to obtain the target text scene representation; Alternatively, the multi-scale target visual representation and the target visual token are input into a hyperbolic graph neural network to obtain a target visual scene representation; Using a target text scene representation or a target visual scene representation as a target index; Calculate the similarity between the representation of the text or video to be retrieved and the target index, and filter the text or video to be retrieved according to the similarity to obtain the target retrieval result; The training process of the text video retrieval model includes: Obtain text training data and video training data and input them into the content encoding component to obtain word matrix masks, text event representations, text semantic unit representations, visual event representations, and visual semantic unit representations; Input the video training data and the text training data and the word matrix mask into the context encoding component to obtain the text token representation and the visual token representation respectively; The text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation and visual token representation are mapped as node features to the hyperbolic space to construct an adjacency matrix, and the adjacency matrix and node features are input into the hyperbolic graph neural network, and the text scene representation and visual scene representation are obtained through hyperbolic graph convolution and pooling operations; the training loss value is calculated according to the text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation, visual token representation, text scene representation and visual scene representation; the weight parameters of the text video retrieval model are optimized according to the training loss value, and the iteration is repeated until the training termination condition is reached, and the trained text video retrieval model is output.

2. The text video retrieval method according to claim 1, characterized in that: The content encoding component includes a first content encoding component, a second content encoding component and a third content encoding component; The acquiring of text training data and video training data and inputting them into the content encoding component to obtain word matrix masks, text event representations, text semantic unit representations, visual event representations, and visual semantic unit representations specifically includes: Inputting text training data and video training data into a first content encoding component to obtain text global representation and visual global representation; The text training data is converted into phrases and word matrix masks through a syntactic analyzer; Inputting the phrase into a second content encoding component to obtain a text semantic unit representation; The K-means algorithm is used to segment the global visual representation into visual semantic unit representations; After the text semantic unit representation and the visual semantic unit representation are input into the third content encoding component, they are added with the text global representation and the visual global representation to obtain the text event representation and the visual event representation.

3. The text video retrieval method according to claim 2, characterized in that: The first content encoding component includes a convolutional neural network, a global visual encoder, and a global text encoder; Inputting the text training data and the video training data into the first content encoding component to obtain a text global representation and a visual global representation, specifically including: Extracting an image block sequence from video training data through a convolutional neural network; performing layer normalization on the image block sequence and inputting it into a multi-head attention layer in the global visual encoder to obtain a global visual extraction feature; performing layer normalization on the global visual fusion feature and inputting it into a multi-layer perceptron in the global visual encoder to obtain a global visual perception feature; splicing the global visual perception feature with the global visual fusion feature to obtain a visual global representation; The text training data is layer-normalized and input into the multi-head attention layer in the global text encoder to obtain global text extraction features, the global text extraction features are spliced ​​with the global text fusion features of the text training data, the global text fusion features are layer-normalized and input into the multi-layer perceptron in the global text encoder to obtain global text perception features; the global text perception features are spliced ​​with the global text fusion features to obtain a global representation of the text.

4. The text video retrieval method according to claim 2, characterized in that: The third content encoding component includes an event visual encoder and an event text encoder; After the text semantic unit representation and the visual semantic unit representation are input into the third content encoding component, they are added with the text global representation and the visual global representation to obtain the text event representation and the visual event representation, which specifically include: The visual semantic unit representation is subjected to layer normalization processing and then input into the multi-head attention layer in the event visual encoder to obtain visual event extraction features, and the visual event extraction features are subjected to layer normalization processing and then input into the multi-layer perceptron in the event visual encoder to obtain visual event perception features; After average pooling of the global visual representation, it is concatenated with the visual event perception features and the visual event extraction features to obtain the visual event representation; The text semantic unit representation is subjected to layer normalization processing and then input into the multi-head attention layer in the event text encoder to obtain text event extraction features, and the text event extraction features are subjected to layer normalization processing and then input into the multi-layer perceptron in the event text encoder to obtain text event perception features; After adding classification tags to the global text representation, it is concatenated with the text event perception features and text event extraction features to obtain the text event representation.

5. The text video retrieval method according to claim 1, characterized in that: The context encoding component includes a context visual encoder; Input the video training data into the context encoding component to obtain visual tokens, including: The video training data is input into the context vision encoder, the video training data is layer-normalized to obtain the visual standard data, and the classification label is added to the visual standard data to obtain the visual initial token; The visual initial token is moved forward and backward along the direction of the video frame sequence to capture fine-grained temporal information, and is input into the multi-head attention layer in the context visual encoder to obtain a visual extraction token, and the visual extraction token is spliced ​​with the visual initial token to obtain a visual fusion token; the visual fusion token is layer-normalized and then input into the multi-layer perceptron in the context visual encoder to obtain a first visual perception token; the first visual perception token is spliced ​​with the visual fusion token to obtain a visual refinement token; Inputting the visual refinement token into the multi-layer perceptron in the token selection layer, compressing the visual refinement token to a set ratio to obtain a first visual compression token; After adding a classification mark to the first visual compression token, the token is input again into the multi-layer perceptron in the token selection layer to obtain a second visual compression token; Perform Softmax function calculation on the second visual compression token to obtain an importance score, and then select the first K visual refinement tokens in each video frame as visual key tokens according to the importance score; The visual key token is layer-normalized and input into the multi-head attention layer in the contextual visual encoder to obtain the visual key refinement token, and the visual key refinement token is concatenated with the visual key token to obtain the visual key fusion token; the visual key fusion token is layer-normalized and input into the multi-layer perceptron in the contextual visual encoder to obtain the second visual perception token, and the second visual perception token is concatenated with the visual key fusion token to obtain the visual token.

6. The text video retrieval method according to claim 1, characterized in that: The context encoding component includes a first neural network architecture and a second neural network architecture; The text training data and word matrix mask are input into the context encoding component to obtain text tokens, including: Input text training data into the first neural network architecture, perform layer normalization on the text training data and then input it into the multi-head attention layer in the first neural network architecture to obtain a first text extraction token, concatenate the first text extraction token with the text training data to obtain a first text fusion token; perform layer normalization on the first text fusion token and then input it into the multi-layer perceptron in the first neural network architecture to obtain a first text perception token; concatenate the first text perception token with the first text fusion token to obtain a text refinement token; The text refinement token is input into the second neural network architecture, the text refinement token is layer-normalized to obtain the text standardization token, the text standardization token and the word matrix mask are input into the multi-head attention layer in the second neural network architecture to obtain the second text extraction token, the second text extraction token is concatenated with the text refinement token to obtain the second text fusion token; the second text fusion token is layer-normalized and input into the multi-layer perceptron in the second neural network architecture to obtain the second text perception token, and the second text perception token is concatenated with the second text fusion token to obtain the text token.

7. The text video retrieval method according to claim 1, characterized in that: Mapping text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation and visual token representation as node features to the hyperbolic space to construct an adjacency matrix, specifically including: Mapping visual event representation and text event representation to node features of the first level of granularity in the hyperbolic space; mapping visual semantic unit representation and text semantic unit representation to node features of the second level of granularity in the hyperbolic space; mapping visual token representation and text token representation to node features of the third level of granularity in the hyperbolic space; Connect nodes of the same level of granularity to each other, connect each node feature of the second level of granularity with all node features of the first level of granularity; connect the node features of the second level of granularity with the node features of the third level of granularity based on semantic affiliation; The node characteristics and When there is a connection between node features, the connecting edge Otherwise, connect the edges ; Establish an adjacency matrix based on the connection relationship between each node feature .

8. The text video retrieval method according to claim 7, characterized in that: The adjacency matrix and node features are input into the hyperbolic graph neural network, and the text scene representation and visual scene representation are obtained through the hyperbolic graph convolution operation and pooling operation, including: The node features are transformed to capture the hidden representation in the hyperbolic space. The calculation formula is: ; ; ; in, Indicates Layer The hyperbolic space hidden representation of node features, It is the characterization mapping function from Euclidean space to hyperbolic space; Indicates Layer The Euclidean space hidden representation of node features, is the hyperbolic tangent function, is the inverse hyperbolic tangent function, Indicates The learnable parameters of the layer, represents the first The curvature of the layer, is the first Layer Node features; It is the characterization mapping function from hyperbolic space to Euclidean space; According to the adjacency matrix, the node features are aggregated to obtain the hyperbolic space aggregation representation, which is expressed as follows: ; ; ; in, Indicates Layer Hyperbolic space aggregation representation of node features, It is the node information aggregation function; Indicates The neighbor node set of node features, represents the aggregation weight between the i-th node feature and the j-th node feature, [;] represents the tensor concatenation operation, is a learnable matrix; Indicates Layer Hyperbolic space hidden representation of node features; Hidden representation for hyperbolic space Hidden representation with hyperbolic space The distance between represents the first The curvature of the layer; is a hyperbolic function; is a leaky linear rectification function; is a logical function; The hyperbolic space aggregation representation is input into the activation function to obtain the hyperbolic space representation, which is expressed as: ; in, Indicates Layer Hyperbolic space representation of nodes; is the activation function of the hyperbolic graph neural network; Representation of hyperbolic space Pooling operations are performed to obtain text scene representation and video scene representation.

9. The text video retrieval method according to claim 1, characterized in that: The training loss value is calculated based on the text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation and visual token representation, text scene representation and visual scene representation, including: The event retrieval loss, unit representation retrieval loss, token retrieval loss, and scene retrieval loss are calculated based on the text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation, visual token representation, text scene representation, and visual scene representation, respectively. Add parent-child relationships between node features at each level of granularity and calculate the hierarchical structure loss in the hyperbolic space. The expression formula is: ; In the formula, is the hyperbolic representation of the text child nodes, is the hyperbolic representation of the visual child node, is the hyperbolic representation of the text father node, represents the hyperbolic representation of the visual father node, Hyperbolic characterization To hyperbolic representation Distance loss between Hyperbolic characterization With hyperbolic representation Distance loss between A node feature set that represents a parent-child relationship. Indicates A node set whose node features do not have a parent-child relationship; Hyperbolic characterization To hyperbolic representation Position loss between Hyperbolic characterization With hyperbolic representation Position loss between is a hyperparameter, is the two-norm, To obtain the maximum value; It is the sequence number of the text child node or visual child node; is the serial number of the textual parent node or visual parent node; According to event retrieval loss, unit representation retrieval loss, token retrieval loss, scene retrieval loss, distance loss , distance loss , Position loss and position loss Calculate the training loss.

Citation Information

Patent Citations

  • Cross-modal video text retrieval method, system and equipment and medium

    CN116910307A

  • Video text retrieval method based on BEiT-3 multi-mode large model

    CN118377930A