Text video retrieval method based on human brain scene memory pathway inspiration
Through the multi-grained information fusion method and hyperbolic graph convolution operation inspired by the human brain episodic memory pathway, the problems of insufficient multi-grained alignment and low video data processing efficiency in the prior art are solved, and efficient and accurate text video retrieval is achieved.
Patent Information
- Application Number
- CN202510416357.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The existing multimodal data retrieval technology is difficult to achieve the alignment of multi-grained information, resulting in insufficient cross-modal alignment, and the computing resources consumed when processing high-dimensional and highly redundant video data, which cannot improve efficiency while ensuring retrieval accuracy.
The text video retrieval method inspired by the human brain episodic memory pathway is adopted to comprehensively capture the multi-level semantic features in text and videos through multi-grained information fusion, and the hyperbolic graph convolution operation is used to fuse multi-modal and multi-grained high-order information, thereby better capturing the complex relationship between text and videos.
It significantly improves the accuracy and efficiency of text video retrieval, enhances the robustness and generalization capabilities of the model, and can better process multi-grained information and high-dimensional video data.
Smart Images

Figure CN119938985A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of network model analysis, and in particular relates to a text video retrieval method of multi-granularity information fusion. Background Art
[0002] With the rapid development of the Internet, the amount of multimodal data (such as text, images, videos, etc.) has exploded. How to efficiently and accurately retrieve the information required by users from massive multimodal data has become an important research direction. Among them, the text and video retrieval task is particularly challenging because it requires processing two highly heterogeneous data modalities, text and video, at the same time.
[0003] Multimodal data retrieval is an information retrieval method involving multiple media modalities (such as text, images, audio, video, etc.). Current multimodal retrieval technology mainly converts data into vector representations through deep learning models and extracts common features through modal fusion, and sorts retrieval results through similarity metrics. However, existing methods can only process coarse-grained or fine-grained information in the cross-modal alignment process, and cannot achieve multi-granular alignment, resulting in insufficient cross-modal alignment. At the same time, video data has the characteristics of high dimensionality and high redundancy. Existing methods often require a lot of computing resources when processing video data, and cannot improve retrieval efficiency while ensuring retrieval accuracy. Summary of the invention
[0004] The present invention provides a text-video retrieval method inspired by the episodic memory pathway of the human brain. It comprehensively captures the multi-level semantic features in text and video through multi-granularity information fusion, and fuses multi-modal and multi-granularity high-order information through hyperbolic graph convolution operations, which can better capture the complex relationship between text and video and significantly improve the accuracy and efficiency of text-video retrieval.
[0005] In order to achieve the above object, the technical solution adopted by the present invention is:
[0006] The first aspect of the present invention provides a text video retrieval method based on human brain episodic memory pathway inspiration, comprising:
[0007] Obtaining target text data or target video data and inputting the data into a text-video retrieval model, wherein the text-video retrieval model includes a content encoding component, a context encoding component, and a hyperbolic graph neural network;
[0008] Using the content encoding component to extract content representation of target text data or target video data to obtain multi-scale target text representation or multi-scale target visual representation;
[0009] Using the context encoding component to extract context representation of target text data or target video data to obtain target text tokens or target visual tokens;
[0010] Inputting the multi-scale target text representation and the target text token into the hyperbolic graph neural network to obtain the target text scene representation; or inputting the multi-scale target visual representation and the target visual token into the hyperbolic graph neural network to obtain the target visual scene representation; using the target text scene representation or the target visual scene representation as the target index;
[0011] Calculate the similarity between the representation of the text or video to be retrieved and the target index, and filter the text or video to be retrieved based on the similarity to obtain the target retrieval result.
[0012] Furthermore, the training process of the text video retrieval model includes:
[0013] Obtain text training data and video training data and input them into the content encoding component to obtain word matrix masks, text event representations, text semantic unit representations, visual event representations, and visual semantic unit representations;
[0014] Input the video training data and the text training data and the word matrix mask into the context encoding component to obtain the text token representation and the visual token representation respectively;
[0015] The text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation and visual token representation are mapped as node features to the hyperbolic space to construct an adjacency matrix, and the adjacency matrix and node features are input into the hyperbolic graph neural network, and the text scene representation and visual scene representation are obtained through hyperbolic graph convolution and pooling operations; the training loss value is calculated according to the text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation, visual token representation, text scene representation and visual scene representation; the weight parameters of the text video retrieval model are optimized according to the training loss value, and the iteration is repeated until the training termination condition is reached, and the trained text video retrieval model is output.
[0016] Further, text training data and video training data are obtained and input into the content encoding component to obtain word matrix mask, text event representation, text semantic unit representation, visual event representation, and visual semantic unit representation; specifically including:
[0017] The content encoding component includes a first content encoding component, a second content encoding component and a third content encoding component;
[0018] Inputting text training data and video training data into a first content encoding component to obtain text global representation and visual global representation;
[0019] The text training data is obtained through a syntactic analyzer to obtain phrases and word matrix masks; the phrases are input into a second content encoding component to obtain text semantic unit representations;
[0020] The K-means algorithm is used to segment the global visual representation into visual semantic unit representations;
[0021] After the text semantic unit representation and the visual semantic unit representation are input into the third content encoding component, they are added with the text global representation and the visual global representation to obtain the text event representation and the visual event representation.
[0022] Furthermore, after the text semantic unit representation and the visual semantic unit representation are input into the third content encoding component, they are added with the text global representation and the visual global representation to obtain the text event representation and the visual event representation, which specifically includes:
[0023] The third content encoding component includes an event visual encoder and an event text encoder;
[0024] The visual semantic unit representation is subjected to layer normalization processing and then input into the multi-head attention layer in the event visual encoder to obtain visual event extraction features, and the visual event extraction features are subjected to layer normalization processing and then input into the multi-layer perceptron in the event visual encoder to obtain visual event perception features;
[0025] After average pooling of the global visual representation, it is concatenated with the visual event perception features and the visual event extraction features to obtain the visual event representation;
[0026] The text semantic unit representation is subjected to layer normalization processing and then input into the multi-head attention layer in the event text encoder to obtain text event extraction features, and the text event extraction features are subjected to layer normalization processing and then input into the multi-layer perceptron in the event text encoder to obtain text event perception features;
[0027] After adding classification tags to the global text representation, it is concatenated with the text event perception features and text event extraction features to obtain the text event representation.
[0028] Furthermore, the video training data is input into the context encoding component to obtain visual tokens, including:
[0029] The context encoding component includes a context visual encoder; inputting video training data into the context visual encoder, performing layer normalization on the video training data to obtain visual standard data, and adding classification labels to the visual standard data to obtain visual initial tokens;
[0030] The visual initial token is moved forward and backward along the direction of the video frame sequence, and is input into the multi-head attention layer in the context visual encoder to obtain a visual extraction token, and the visual extraction token is spliced with the visual initial token to obtain a visual fusion token; the visual fusion token is layer-normalized and then input into the multi-layer perceptron in the context visual encoder to obtain a first visual perception token; the first visual perception token is spliced with the visual fusion token to obtain a visual refinement token;
[0031] The visual refinement tokens are scored for importance by the token selection layer in the contextual visual encoder, and then the top K visual refinement tokens in each video frame are selected as visual key tokens based on the importance scores;
[0032] The visual key token is layer-normalized and input into the multi-head attention layer in the contextual visual encoder to obtain the visual key refinement token, and the visual key refinement token is concatenated with the visual key token to obtain the visual key fusion token; the visual key fusion token is layer-normalized and input into the multi-layer perceptron in the contextual visual encoder to obtain the second visual perception token, and the second visual perception token is concatenated with the visual key fusion token to obtain the visual token.
[0033] Furthermore, the importance of visual refinement tokens is scored by the token selection layer in the contextual visual encoder, and then the first K visual refinement tokens in each video frame are selected as visual key tokens according to the importance scores; specifically, the following steps are performed:
[0034] Inputting the visual refinement token into the multi-layer perceptron in the token selection layer, compressing the visual refinement token to a set ratio to obtain a first visual compression token;
[0035] After adding a classification mark to the first visual compression token, the token is input again into the multi-layer perceptron in the token selection layer to obtain a second visual compression token;
[0036] The Softmax function is calculated for the second visual compression token to obtain the importance score, and then the top K visual refinement tokens in each video frame are selected as visual key tokens according to the importance score.
[0037] Furthermore, the text training data and the word matrix mask are input into the context encoding component to obtain text tokens, specifically including:
[0038] The context encoding component includes a first neural network architecture and a second neural network architecture;
[0039] Input text training data into the first neural network architecture, perform layer normalization on the text training data and then input it into the multi-head attention layer in the first neural network architecture to obtain a first text extraction token, concatenate the first text extraction token with the text training data to obtain a first text fusion token; perform layer normalization on the first text fusion token and then input it into the multi-layer perceptron in the first neural network architecture to obtain a first text perception token; concatenate the first text perception token with the first text fusion token to obtain a text refinement token;
[0040] The text refinement token is input into the second neural network architecture, the text refinement token is layer-normalized to obtain the text standardization token, the text standardization token and the word matrix mask are input into the multi-head attention layer in the second neural network architecture to obtain the second text extraction token, the second text extraction token is concatenated with the text refinement token to obtain the second text fusion token; the second text fusion token is layer-normalized and input into the multi-layer perceptron in the second neural network architecture to obtain the second text perception token, and the second text perception token is concatenated with the second text fusion token to obtain the text token.
[0041] Furthermore, the text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation and visual token representation are mapped as node features to the hyperbolic space to construct an adjacency matrix, specifically including:
[0042] Mapping visual event representation and text event representation to node features of the first level of granularity in the hyperbolic space; mapping visual semantic unit representation and text semantic unit representation to node features of the second level of granularity in the hyperbolic space; mapping visual token representation and text token representation to node features of the third level of granularity in the hyperbolic space;
[0043] Connect nodes of the same level of granularity to each other, connect each node feature of the second level of granularity with all node features of the first level of granularity; connect the node features of the second level of granularity with the node features of the third level of granularity based on semantic affiliation; The node characteristics and When there is a connection between node features, the connecting edge Otherwise, connect the edges ; Establish an adjacency matrix based on the connection relationship between each node feature .
[0044] Furthermore, the adjacency matrix and node features are input into the hyperbolic graph neural network, and the text scene representation and visual scene representation are obtained through the hyperbolic graph convolution operation and pooling operation, including:
[0045] The node features are transformed to capture the hidden representation in the hyperbolic space. The calculation formula is:
[0046]
[0047]
[0048]
[0049] in, Indicates Layer The hyperbolic space hidden representation of node features, It is the characterization mapping function from Euclidean space to hyperbolic space; Indicates Layer The Euclidean space hidden representation of node features, is the hyperbolic tangent function, is the inverse hyperbolic tangent function, Indicates The learnable parameters of the layer, represents the first The curvature of the layer, is the first Layer Node features; It is the characterization mapping function from hyperbolic space to Euclidean space;
[0050] According to the adjacency matrix, the node features are aggregated to obtain the hyperbolic space aggregation representation, which is expressed as follows:
[0051]
[0052]
[0053]
[0054] in, Indicates Layer Hyperbolic space aggregation representation of node features, It is the node information aggregation function; Indicates The neighbor node set of node features, represents the aggregation weight between the i-th node feature and the j-th node feature, [;] represents the tensor concatenation operation, is a learnable matrix; Indicates Layer Hyperbolic space hidden representation of node features; Hidden representation for hyperbolic space Hidden representation with hyperbolic space The distance between represents the first The curvature of the layer; is a hyperbolic function; is a leaky linear rectification function; is a logical function;
[0055] The hyperbolic space aggregation representation is input into the activation function to obtain the hyperbolic space representation, which is expressed as:
[0056]
[0057] in, Indicates Layer Hyperbolic space representation of nodes; is the activation function of the hyperbolic graph neural network;
[0058] Representation of hyperbolic space Pooling operations are performed to obtain text scene representation and video scene representation.
[0059] Further, the training loss value is calculated based on the text event representation, the text semantic unit representation, the visual event representation, the visual semantic unit representation, the text token representation and the visual token representation, the text scene representation and the visual scene representation, specifically including:
[0060] The event retrieval loss, unit representation retrieval loss, token retrieval loss, and scene retrieval loss are calculated based on the text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation, visual token representation, text scene representation, and visual scene representation, respectively.
[0061] Add parent-child relationships between node features at each level of granularity and calculate the hierarchical structure loss in the hyperbolic space. The expression formula is:
[0062]
[0063] In the formula, is the hyperbolic representation of the text child node, is the hyperbolic representation of the visual child node, is the hyperbolic representation of the text father node, represents the hyperbolic representation of the visual father node, Hyperbolic characterization To hyperbolic representation Distance loss between Hyperbolic characterization With hyperbolic representation Distance loss between A node feature set that represents a parent-child relationship. Indicates A node set whose node features do not have a parent-child relationship; Hyperbolic characterization To hyperbolic representation Position loss between Hyperbolic characterization With hyperbolic representation Position loss between is a hyperparameter, is the two-norm, To obtain the maximum value; It is the sequence number of the text child node or visual child node; is the serial number of the textual parent node or visual parent node;
[0064] According to event retrieval loss, unit representation retrieval loss, token retrieval loss, scene retrieval loss, distance loss , distance loss , Position loss and position loss Calculate the training loss.
[0065] Compared with the prior art, the present invention has the following beneficial effects:
[0066] In the present invention, a content encoding component is used to extract content representation of target text data or target video data to obtain a multi-scale target text representation or a multi-scale target visual representation; a context encoding component is used to extract context representation of target text data or target video data to obtain a target text token or a target visual token; and multi-granularity information fusion is used to comprehensively capture multi-level semantic features in text and video, thereby significantly improving the accuracy of text and video retrieval.
[0067] In the present invention, a multi-scale target text representation and a target text token are input into a hyperbolic graph neural network to obtain a target text scene representation; or a multi-scale target visual representation and a target visual token are input into a hyperbolic graph neural network to obtain a target visual scene representation; the target text scene representation or the target visual scene representation is used as a target index; and multi-modal and multi-granular high-order information is fused through a hyperbolic graph convolution operation, which can better capture the complex relationship between text and video and enhance the robustness and generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 is a flowchart of a text video retrieval method based on human brain situational memory pathway inspiration provided by Example 1 of the present invention;
[0069] Figure 2 is a structural diagram of a first content encoding component provided by Embodiment 2 of the present invention;
[0070] Figure 3 This is a structural diagram of the third content encoding component provided in Example 2 of the present invention.
[0071] Figure 4 is a structural diagram of a contextual visual encoder provided by Embodiment 2 of the present invention;
[0072] Figure 5 is a structural diagram of a context text encoder provided by Embodiment 2 of the present invention;
[0073] Figure 6 is a structural diagram of a hyperbolic graph convolutional neural network provided in Example 2 of the present invention;
[0074] Figure 7 Schematic diagram of the Poincare disk provided in Example 2 of the present invention. DETAILED DESCRIPTION
[0075] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.
[0076] Brain-like computing is an emerging research direction in the field of artificial intelligence. Its core lies in learning from the information processing mode or structure of the biological nervous system, and then building corresponding computing theories, chip architectures, and application models and algorithms. In recent years, cognitive science has made objective progress in the study of human brain episodic memory pathways. The human brain episodic memory pathway decomposes external sensory signals into information of multiple granularities and then fuses them into complex scene representations. This is different from the traditional encoding, alignment, and retrieval methods of text and video in the field of artificial intelligence, and provides a new reference for the model design of text and video retrieval tasks.
[0077] Inspired by the episodic memory pathway of the human brain, the present invention comprehensively captures the multi-level semantic features in text and video through multi-granularity information fusion, and fuses multi-modal and multi-granularity high-order information through hyperbolic graph convolution operations, which can better capture the complex relationship between text and video and significantly improve the accuracy and efficiency of text and video retrieval.
[0078] Example 1
[0079] like Figure 1 As shown, this implementation provides a text video retrieval method inspired by the human brain's episodic memory pathway, including:
[0080] Obtain target text data or target video data and input it into the text video retrieval model to obtain a target index; calculate the similarity between the representation of the text or video to be retrieved and the target index, and screen the text or video to be retrieved according to the similarity to obtain the target retrieval result; specifically including:
[0081] Obtaining target text data or target video data and inputting the data into a text-video retrieval model, wherein the text-video retrieval model includes a content encoding component, a context encoding component, and a hyperbolic graph neural network;
[0082] The content encoding component is used to extract content representation of target text data or target video data to obtain a multi-scale target text representation or a multi-scale target visual representation; the multi-scale target text representation includes a target text event representation, a target text global representation and a target text semantic unit representation; the multi-scale target visual representation includes a target visual event representation, a target visual global representation and a target visual semantic unit representation.
[0083] Using the context encoding component to extract context representation of target text data or target video data to obtain target text tokens or target visual tokens;
[0084] Inputting the multi-scale target text representation and the target text token into the hyperbolic graph neural network to obtain the target text scene representation; or inputting the multi-scale target visual representation and the target visual token into the hyperbolic graph neural network to obtain the target visual scene representation; using the target text scene representation or the target visual scene representation as the target index;
[0085] Calculate the similarity between the representation of the text or video to be retrieved and the target index, and screen the text or video to be retrieved according to the similarity to obtain the target retrieval result. The target retrieval result includes related videos and related texts; that is, retrieve the related videos according to the target text data; retrieve the related text according to the target video data.
[0086] The text video retrieval model includes a content encoding component, a context encoding component and a hyperbolic graph neural network; the training process of the text video retrieval model includes:
[0087] Obtain text training data and video training data and input them into the content encoding component to obtain word matrix masks, text event representations, text semantic unit representations, visual event representations, and visual semantic unit representations; specifically including:
[0088] The content encoding component includes a first content encoding component, a second content encoding component and a third content encoding component;
[0089] Inputting text training data and video training data into a first content encoding component to obtain text global representation and visual global representation;
[0090] The text training data is obtained through a syntactic analyzer to obtain phrases and word matrix masks; the phrases are input into a second content encoding component to obtain text semantic unit representations;
[0091] The K-means algorithm is used to segment the global visual representation into visual semantic unit representations;
[0092] After the text semantic unit representation and the visual semantic unit representation are input into the third content encoding component, they are added with the text global representation and the visual global representation to obtain the text event representation and the visual event representation.
[0093] Input the video training data and the text training data and the word matrix mask into the context encoding component to obtain the text token representation and the visual token representation respectively;
[0094] The text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation and visual token representation are mapped as node features to the hyperbolic space to construct an adjacency matrix, and the adjacency matrix and node features are input into the hyperbolic graph neural network, and the text scene representation and visual scene representation are obtained through hyperbolic graph convolution and pooling operations; the training loss value is calculated according to the text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation, visual token representation, text scene representation and visual scene representation; the weight parameters of the text video retrieval model are optimized according to the training loss value, and the iteration is repeated until the training termination condition is reached, and the trained text video retrieval model is output.
[0095] Example 2
[0096] like Figures 2 to 5 As shown, this implementation provides a text video retrieval method inspired by the human brain's episodic memory pathway, including:
[0097] The text video retrieval model includes a content encoding component, a context encoding component and a hyperbolic graph neural network; the training process of the text video retrieval model includes:
[0098] Acquire text training data and video training data and input them into a content encoding component to obtain a word matrix mask, a text event representation, a text semantic unit representation, a visual event representation, and a visual semantic unit representation; the content encoding component includes a first content encoding component, a second content encoding component, and a third content encoding component; in this embodiment, the first content encoding component, the second content encoding component, and the third content encoding component are Content CLIP (Content Comparative Language-Image Pre-training Model);
[0099] Inputting the text training data and the video training data into the first content encoding component to obtain a text global representation and a visual global representation, specifically including:
[0100] like Figure 2 As shown, the first content encoding component includes a convolutional neural network, a global vision and a global text encoder;
[0101] Extracting an image block sequence from video training data through a convolutional neural network; performing layer normalization on the image block sequence and inputting it into a multi-head attention layer in the global visual encoder to obtain a global visual extraction feature; performing layer normalization on the global visual fusion feature and inputting it into a multi-layer perceptron in the global visual encoder to obtain a global visual perception feature; splicing the global visual perception feature with the global visual fusion feature to obtain a visual global representation;
[0102] The text training data is layer-normalized and input into the multi-head attention layer in the global text encoder to obtain global text extraction features, the global text extraction features are spliced with the global text fusion features of the text training data, the global text fusion features are layer-normalized and input into the multi-layer perceptron in the global text encoder to obtain global text perception features; the global text perception features are spliced with the global text fusion features to obtain a global representation of the text.
[0103] The visual global representation is converted into a visual semantic unit representation through the K-means algorithm; the text training data is converted into phrases and word matrix masks through a syntactic analyzer; the phrases are input into a second content encoding component to obtain a text semantic unit representation; the second content encoding component includes a unit text encoder; the unit text encoder is configured with a multi-layer perceptron and a multi-head attention layer.
[0104] After the text semantic unit representation and the visual semantic unit representation are input into the third content encoding component, they are added with the text global representation and the visual global representation to obtain the text event representation and the visual event representation, which specifically include:
[0105] The third content encoding component includes an event visual encoder and an event text encoder;
[0106] The visual semantic unit representation is subjected to layer normalization processing and then input into the multi-head attention layer in the event visual encoder to obtain visual event extraction features, and the visual event extraction features are subjected to layer normalization processing and then input into the multi-layer perceptron in the event visual encoder to obtain visual event perception features;
[0107] After average pooling of the global visual representation, it is concatenated with the visual event perception features and the visual event extraction features to obtain the visual event representation;
[0108] The text semantic unit representation is subjected to layer normalization processing and then input into the multi-head attention layer in the event text encoder to obtain text event extraction features, and the text event extraction features are subjected to layer normalization processing and then input into the multi-layer perceptron in the event text encoder to obtain text event perception features;
[0109] After adding classification tags to the global text representation, it is concatenated with the text event perception features and text event extraction features to obtain the text event representation.
[0110] The context encoding component includes a context visual encoder and a context text encoder; the context visual encoder and the context text encoder are Context CLIP (Contextual Contrastive Language Image Model).
[0111] Input the video training data into the context visual encoder to obtain the visual token, which includes:
[0112] The video training data is input into the context vision encoder, the video training data is layer-normalized to obtain the visual standard data, and the classification label is added to the visual standard data to obtain the visual initial token;
[0113] The visual initial token is moved forward and backward along the direction of the video frame sequence, and is input into the multi-head attention layer in the context visual encoder to obtain a visual extraction token, and the visual extraction token is spliced with the visual initial token to obtain a visual fusion token; the visual fusion token is layer-normalized and then input into the multi-layer perceptron in the context visual encoder to obtain a first visual perception token; the first visual perception token is spliced with the visual fusion token to obtain a visual refinement token;
[0114] The visual refinement token is input into the token selection layer in the context visual encoder, and the visual refinement token is compressed to a set ratio by the multi-layer perceptron in the token selection layer to obtain the first visual compression token; the visual compression token is added with a classification mark and then input into the multi-layer perceptron in the token selection layer again to obtain the second visual compression token; the second visual compression token is calculated by the Softmax function to obtain an importance score, and then the top K visual refinement tokens (Top K) in each video frame are selected as visual key tokens according to the importance score;
[0115] The visual key token is layer-normalized and input into the multi-head attention layer in the contextual visual encoder to obtain the visual key refinement token, and the visual key refinement token is concatenated with the visual key token to obtain the visual key fusion token; the visual key fusion token is layer-normalized and input into the multi-layer perceptron in the contextual visual encoder to obtain the second visual perception token, and the second visual perception token is concatenated with the visual key fusion token to obtain the visual token.
[0116] Input text training data and word matrix mask into the context text encoder to obtain text tokens, including:
[0117] The context text encoder includes a first neural network architecture and a second neural network architecture; the first neural network architecture and the second neural network architecture are Transformer neural network architectures;
[0118] Input text training data into the first neural network architecture, perform layer normalization on the text training data and then input it into the multi-head attention layer in the first neural network architecture to obtain a first text extraction token, concatenate the first text extraction token with the text training data to obtain a first text fusion token; perform layer normalization on the first text fusion token and then input it into the multi-layer perceptron in the first neural network architecture to obtain a first text perception token; concatenate the first text perception token with the first text fusion token to obtain a text refinement token;
[0119] The text refinement token is input into the second neural network architecture, the text refinement token is layer-normalized to obtain the text standardization token, the text standardization token and the word matrix mask are input into the multi-head attention layer in the second neural network architecture to obtain the second text extraction token, the second text extraction token is concatenated with the text refinement token to obtain the second text fusion token; the second text fusion token is layer-normalized and input into the multi-layer perceptron in the second neural network architecture to obtain the second text perception token, and the second text perception token is concatenated with the second text fusion token to obtain the text token.
[0120] Mapping text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation and visual token representation as node features to the hyperbolic space to construct an adjacency matrix, specifically including:
[0121] like Figure 7 As shown, in this embodiment, the hyperbolic space is represented by a Poincare disk, the space is centered at the origin, and the capacity of the space increases exponentially outward.
[0122] Mapping visual event representation and text event representation to node features of the first level of granularity in the hyperbolic space; mapping visual semantic unit representation and text semantic unit representation to node features of the second level of granularity in the hyperbolic space; mapping visual token representation and text token representation to node features of the third level of granularity in the hyperbolic space;
[0123] Connect nodes of the same level of granularity to each other, connect each node feature of the second level of granularity with all node features of the first level of granularity; connect the node features of the second level of granularity with the node features of the third level of granularity based on semantic affiliation; The node characteristics and When there is a connection between node features, the connecting edge Otherwise, connect the edges ; Establish an adjacency matrix based on the connection relationship between each node feature .
[0124] like Figure 6As shown in the figure, the adjacency matrix and node features are input into the hyperbolic graph neural network, and the text scene representation and visual scene representation are obtained through the hyperbolic graph convolution operation and pooling operation, which specifically include:
[0125] The node features are transformed to capture the hidden representation in the hyperbolic space. The calculation formula is:
[0126]
[0127]
[0128]
[0129] in, Indicates Layer The hyperbolic space hidden representation of node features, It is the characterization mapping function from Euclidean space to hyperbolic space; Indicates Layer The Euclidean space hidden representation of node features, is the hyperbolic tangent function, is the inverse hyperbolic tangent function, Indicates The learnable parameters of the layer, represents the first The curvature of the layer, is the first Layer Node features; It is the characterization mapping function from hyperbolic space to Euclidean space;
[0130] According to the adjacency matrix, the node features are aggregated to obtain the hyperbolic space aggregation representation, which is expressed as follows:
[0131]
[0132]
[0133]
[0134] in, Indicates Layer Hyperbolic space aggregation representation of node features, It is the node information aggregation function; Indicates The neighbor node set of node features, represents the aggregation weight between the i-th node feature and the j-th node feature, [;] represents the tensor concatenation operation, is a learnable matrix; Indicates Layer Hyperbolic space hidden representation of node features; Hidden representation for hyperbolic space Hidden representation with hyperbolic space The distance between represents the first The curvature of the layer; is a hyperbolic function; is a leaky linear rectification function; is a logical function;
[0135] The hyperbolic space aggregation representation is input into the activation function to obtain the hyperbolic space representation, which is expressed as:
[0136]
[0137] in, Indicates Layer Hyperbolic space representation of nodes; is the activation function of the hyperbolic graph neural network;
[0138] Representation of hyperbolic space Pooling operations are performed to obtain text scene representation and visual scene representation.
[0139] The visual event representation and the text event representation are aligned and a first similarity is calculated; the event retrieval loss is calculated according to the first similarity, and the event retrieval loss includes the retrieval loss from the visual event representation to the text event representation and the retrieval loss from the text event representation to the visual event representation; the expression formula is:
[0140]
[0141] In the formula, is the normalization coefficient; is the natural exponential function; B is the batch size; The retrieval loss for textual event representation to visual event representation; The retrieval loss for visual event representation to textual event representation; is the positive sample similarity between text event representation and visual event representation; Retrieving negative sample similarity of visual event representation from textual event representation; Retrieving negative sample similarity of textual event representation from visual event representation;
[0142] The visual semantic unit and the text semantic unit are aligned and the second similarity is calculated; the semantic unit retrieval loss is calculated according to the second similarity; the semantic unit retrieval loss includes the retrieval loss from the visual semantic unit to the text semantic unit and the retrieval loss from the text semantic unit to the visual semantic unit; the expression formula is:
[0143]
[0144] In the formula, is the retrieval loss from text semantic units to visual semantic units; is the retrieval loss from text semantic units to visual semantic units; is the positive sample similarity between text semantic unit and visual semantic unit; The negative sample similarity of retrieving visual semantic units from textual semantic units; The negative sample similarity of retrieving text semantic units from visual semantic units.
[0145] The text token representation and the visual token representation are aligned and the third similarity is calculated; the token retrieval loss is calculated according to the third similarity; the token retrieval loss includes the retrieval loss from the text token to the visual token and the retrieval loss from the visual token to the text token; the expression formula is:
[0146]
[0147] In the formula, is the retrieval loss from textual tokens to visual tokens; is the retrieval loss from visual token to textual token; is the positive sample similarity between text tokens and visual tokens; Similarity of negative samples for retrieving visual tokens from textual tokens; Negative similarity for retrieving text tokens from visual tokens.
[0148] The text scene representation and the visual scene representation are aligned and a fourth similarity is calculated; the scene retrieval loss is calculated according to the fourth similarity, and the scene retrieval loss includes the retrieval loss from the text scene representation to the visual scene representation and the retrieval loss from the visual scene representation to the text scene representation; the expression formula is:
[0149]
[0150] In the formula, The retrieval loss from textual scene representation to visual scene representation; The retrieval loss for visual scene representation to textual scene representation; is the positive sample similarity between visual scene representation and text scene representation; Retrieve the similarity of negative samples of visual scene representation from textual scene representation; Similarity of negative samples for retrieving textual scene representations from visual scene representations.
[0151] Add parent-child relationships between node features at each level of granularity and calculate the hierarchical structure loss in the hyperbolic space. The expression formula is:
[0152]
[0153] In the formula, is the hyperbolic representation of the text child node, is the hyperbolic representation of the visual child node, is the hyperbolic representation of the text father node, represents the hyperbolic representation of the visual father node, Hyperbolic characterization To hyperbolic representation Distance loss between Hyperbolic characterization With hyperbolic representation Distance loss between A node feature set that represents a parent-child relationship. Indicates A node set whose node features do not have a parent-child relationship; Hyperbolic characterization To hyperbolic representation Position loss between Hyperbolic characterization With hyperbolic representation Position loss between is a hyperparameter, is the two-norm, To obtain the maximum value; It is the sequence number of the text child node or visual child node; is the serial number of the textual parent node or visual parent node;
[0154] According to event retrieval loss, semantic unit retrieval loss, token retrieval loss, scene retrieval loss, distance loss , distance loss , Position loss and position loss Calculate the training loss.
[0155] The weight parameters of the text-video retrieval model are optimized according to the training loss value, and the iteration is repeated until the training termination condition is reached to output the trained text-video retrieval model.
[0156] In the inference phase, this embodiment adopts a two-stage retrieval strategy. First, only the content encoding component is started, and the target text data or target video data is input into the content encoding component to extract the semantic unit representation and event representation; the semantic unit representation and event representation are used as target indexes to quickly screen out candidate sets in the video database;
[0157] The text or video to be retrieved in the candidate set is sent to the text and video retrieval model again, the similarity between the text or video to be retrieved in the candidate set and the target index is calculated, and the text or video to be retrieved in the candidate set is screened according to the similarity to obtain the target retrieval results and re-sorted; this embodiment not only ensures the retrieval accuracy but also ensures the model efficiency.
[0158] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0159] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0160] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0161] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0162] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A text-video retrieval method inspired by human brain episodic memory pathways, characterized in that: include: Obtaining target text data or target video data and inputting the data into a text-video retrieval model, wherein the text-video retrieval model includes a content encoding component, a context encoding component, and a hyperbolic graph neural network; Using the content encoding component to extract content representation of target text data or target video data to obtain multi-scale target text representation or multi-scale target visual representation; Using the context encoding component to extract context representation of target text data or target video data to obtain target text tokens or target visual tokens; Input the multi-scale target text representation and the target text token into the hyperbolic graph neural network to obtain the target text scene representation; Alternatively, the multi-scale target visual representation and the target visual token are input into a hyperbolic graph neural network to obtain a target visual scene representation; Using a target text scene representation or a target visual scene representation as a target index; Calculate the similarity between the representation of the text or video to be retrieved and the target index, and filter the text or video to be retrieved based on the similarity to obtain the target retrieval result.
2. The text video retrieval method according to claim 1, characterized in that: The training process of the text video retrieval model includes: Obtain text training data and video training data and input them into the content encoding component to obtain word matrix masks, text event representations, text semantic unit representations, visual event representations, and visual semantic unit representations; Input the video training data and the text training data and the word matrix mask into the context encoding component to obtain the text token representation and the visual token representation respectively; The text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation and visual token representation are mapped as node features to the hyperbolic space to construct an adjacency matrix, and the adjacency matrix and node features are input into the hyperbolic graph neural network, and the text scene representation and visual scene representation are obtained through hyperbolic graph convolution and pooling operations; the training loss value is calculated according to the text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation, visual token representation, text scene representation and visual scene representation; the weight parameters of the text video retrieval model are optimized according to the training loss value, and the iteration is repeated until the training termination condition is reached, and the trained text video retrieval model is output.
3. The text video retrieval method according to claim 2, characterized in that: The content encoding component includes a first content encoding component, a second content encoding component and a third content encoding component; The acquiring of text training data and video training data and inputting them into the content encoding component to obtain word matrix masks, text event representations, text semantic unit representations, visual event representations, and visual semantic unit representations specifically includes: Inputting text training data and video training data into a first content encoding component to obtain text global representation and visual global representation; The text training data is converted into phrases and word matrix masks through a syntactic analyzer; Inputting the phrase into a second content encoding component to obtain a text semantic unit representation; The K-means algorithm is used to segment the global visual representation into visual semantic unit representations; After the text semantic unit representation and the visual semantic unit representation are input into the third content encoding component, they are added with the text global representation and the visual global representation to obtain the text event representation and the visual event representation.
4. The text video retrieval method according to claim 3, characterized in that: The first content encoding component includes a convolutional neural network, a global visual encoder, and a global text encoder; Inputting the text training data and the video training data into the first content encoding component to obtain a text global representation and a visual global representation, specifically including: Extracting an image block sequence from video training data through a convolutional neural network; performing layer normalization on the image block sequence and inputting it into a multi-head attention layer in the global visual encoder to obtain a global visual extraction feature; performing layer normalization on the global visual fusion feature and inputting it into a multi-layer perceptron in the global visual encoder to obtain a global visual perception feature; splicing the global visual perception feature with the global visual fusion feature to obtain a visual global representation; The text training data is layer-normalized and input into the multi-head attention layer in the global text encoder to obtain global text extraction features, the global text extraction features are spliced with the global text fusion features of the text training data, the global text fusion features are layer-normalized and input into the multi-layer perceptron in the global text encoder to obtain global text perception features; the global text perception features are spliced with the global text fusion features to obtain a global representation of the text.
5. The text video retrieval method according to claim 3, characterized in that: The third content encoding component includes an event visual encoder and an event text encoder; After the text semantic unit representation and the visual semantic unit representation are input into the third content encoding component, they are added with the text global representation and the visual global representation to obtain the text event representation and the visual event representation, which specifically include: The visual semantic unit representation is subjected to layer normalization processing and then input into the multi-head attention layer in the event visual encoder to obtain visual event extraction features, and the visual event extraction features are subjected to layer normalization processing and then input into the multi-layer perceptron in the event visual encoder to obtain visual event perception features; After average pooling of the global visual representation, it is concatenated with the visual event perception features and the visual event extraction features to obtain the visual event representation; The text semantic unit representation is subjected to layer normalization processing and then input into the multi-head attention layer in the event text encoder to obtain text event extraction features, and the text event extraction features are subjected to layer normalization processing and then input into the multi-layer perceptron in the event text encoder to obtain text event perception features; After adding classification tags to the global text representation, it is concatenated with the text event perception features and text event extraction features to obtain the text event representation.
6. The text video retrieval method according to claim 2, characterized in that: The context encoding component includes a context visual encoder; Input the video training data into the context encoding component to obtain visual tokens, including: The video training data is input into the context vision encoder, the video training data is layer-normalized to obtain the visual standard data, and the classification label is added to the visual standard data to obtain the visual initial token; The visual initial token is moved forward and backward along the direction of the video frame sequence to capture fine-grained temporal information, and is input into the multi-head attention layer in the context visual encoder to obtain a visual extraction token, and the visual extraction token is spliced with the visual initial token to obtain a visual fusion token; the visual fusion token is layer-normalized and then input into the multi-layer perceptron in the context visual encoder to obtain a first visual perception token; the first visual perception token is spliced with the visual fusion token to obtain a visual refinement token; Inputting the visual refinement token into the multi-layer perceptron in the token selection layer, compressing the visual refinement token to a set ratio to obtain a first visual compression token; After adding a classification mark to the first visual compression token, the token is input again into the multi-layer perceptron in the token selection layer to obtain a second visual compression token; Perform Softmax function calculation on the second visual compression token to obtain an importance score, and then select the first K visual refinement tokens in each video frame as visual key tokens according to the importance score; The visual key token is layer-normalized and input into the multi-head attention layer in the contextual visual encoder to obtain the visual key refinement token, and the visual key refinement token is concatenated with the visual key token to obtain the visual key fusion token; the visual key fusion token is layer-normalized and input into the multi-layer perceptron in the contextual visual encoder to obtain the second visual perception token, and the second visual perception token is concatenated with the visual key fusion token to obtain the visual token.
7. The text video retrieval method according to claim 2, characterized in that: The context encoding component includes a first neural network architecture and a second neural network architecture; The text training data and word matrix mask are input into the context encoding component to obtain text tokens, including: Input text training data into the first neural network architecture, perform layer normalization on the text training data and then input it into the multi-head attention layer in the first neural network architecture to obtain a first text extraction token, concatenate the first text extraction token with the text training data to obtain a first text fusion token; perform layer normalization on the first text fusion token and then input it into the multi-layer perceptron in the first neural network architecture to obtain a first text perception token; concatenate the first text perception token with the first text fusion token to obtain a text refinement token; The text refinement token is input into the second neural network architecture, the text refinement token is layer-normalized to obtain the text standardization token, the text standardization token and the word matrix mask are input into the multi-head attention layer in the second neural network architecture to obtain the second text extraction token, the second text extraction token is concatenated with the text refinement token to obtain the second text fusion token; the second text fusion token is layer-normalized and input into the multi-layer perceptron in the second neural network architecture to obtain the second text perception token, and the second text perception token is concatenated with the second text fusion token to obtain the text token.
8. The text video retrieval method according to claim 2, characterized in that: Mapping text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation and visual token representation as node features to the hyperbolic space to construct an adjacency matrix, specifically including: Mapping visual event representation and text event representation to node features of the first level of granularity in the hyperbolic space; mapping visual semantic unit representation and text semantic unit representation to node features of the second level of granularity in the hyperbolic space; mapping visual token representation and text token representation to node features of the third level of granularity in the hyperbolic space; Connect nodes of the same level of granularity to each other, connect each node feature of the second level of granularity with all node features of the first level of granularity; connect the node features of the second level of granularity with the node features of the third level of granularity based on semantic affiliation; The node characteristics and When there is a connection between node features, the connecting edge Otherwise, connect the edges ; Establish an adjacency matrix based on the connection relationship between each node feature .
9. The text video retrieval method according to claim 8, characterized in that: The adjacency matrix and node features are input into the hyperbolic graph neural network, and the text scene representation and visual scene representation are obtained through the hyperbolic graph convolution operation and pooling operation, including: The node features are transformed to capture the hidden representation in the hyperbolic space. The calculation formula is: ; ; ; in, Indicates Layer The hyperbolic space hidden representation of node features, It is the characterization mapping function from Euclidean space to hyperbolic space; Indicates Layer The Euclidean space hidden representation of node features, is the hyperbolic tangent function, is the inverse hyperbolic tangent function, Indicates The learnable parameters of the layer, represents the first The curvature of the layer, is the first Layer Node features; It is the characterization mapping function from hyperbolic space to Euclidean space; According to the adjacency matrix, the node features are aggregated to obtain the hyperbolic space aggregation representation, which is expressed as follows: ; ; ; in, Indicates Layer Hyperbolic space aggregation representation of node features, It is the node information aggregation function; Indicates The neighbor node set of node features, represents the aggregation weight between the i-th node feature and the j-th node feature, [;] represents the tensor concatenation operation, is a learnable matrix; Indicates Layer Hyperbolic space hidden representation of node features; Hidden representation for hyperbolic space Hidden representation with hyperbolic space The distance between represents the first The curvature of the layer; is a hyperbolic function; is a leaky linear rectification function; is a logical function; The hyperbolic space aggregation representation is input into the activation function to obtain the hyperbolic space representation, which is expressed as: ; in, Indicates Layer Hyperbolic space representation of nodes; is the activation function of the hyperbolic graph neural network; Representation of hyperbolic space Pooling operations are performed to obtain text scene representation and video scene representation.
10. The text video retrieval method according to claim 2, characterized in that: The training loss value is calculated based on the text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation and visual token representation, text scene representation and visual scene representation, including: The event retrieval loss, unit representation retrieval loss, token retrieval loss, and scene retrieval loss are calculated based on the text event representation, text semantic unit representation, visual event representation, visual semantic unit representation, text token representation, visual token representation, text scene representation, and visual scene representation, respectively. Add parent-child relationships between node features at each level of granularity and calculate the hierarchical structure loss in the hyperbolic space. The expression formula is: ; In the formula, is the hyperbolic representation of the text child nodes, is the hyperbolic representation of the visual child node, is the hyperbolic representation of the text father node, represents the hyperbolic representation of the visual father node, Hyperbolic characterization To hyperbolic representation Distance loss between Hyperbolic characterization With hyperbolic representation Distance loss between A node feature set that represents a parent-child relationship. Indicates A node set whose node features do not have a parent-child relationship; Hyperbolic characterization To hyperbolic representation Position loss between Hyperbolic characterization With hyperbolic representation Position loss between is a hyperparameter, is the two-norm, To obtain the maximum value; It is the sequence number of the text child node or visual child node; is the serial number of the textual parent node or visual parent node; According to event retrieval loss, unit representation retrieval loss, token retrieval loss, scene retrieval loss, distance loss , distance loss , Position loss and position loss Calculate the training loss.
Citation Information
Patent Citations
Cross-modal video text retrieval method, system and equipment and medium
CN116910307A
Video text retrieval method based on BEiT-3 multi-mode large model
CN118377930A
Video text retrieval method based on time sequence token combination
CN119066222A
Method and device for determining multi-modal retrieval model and electronic equipment
CN119336947A
Method And Apparatus For Retrieving Video, Device And Medium
US20210209155A1
Cited By
Video content detection method
CN120356136A
Video understanding method and device and computer program product
CN120598060A
A method, apparatus, and computer program product for video understanding
CN120598060B
Text video retrieval method and device, equipment and storage medium
CN120892601A
Medical image case retrieval system and method based on multi-modal knowledge graph
CN122220550A