Cross-modal Similarity-based Text Mining Data Query Method and System
Through multi-level feature extraction and adaptive cross-modal contrast learning network, combined with multi-level screening model and graph neural network, the problem of correlation capture and timing information ignorance in the existing cross-modal search methods is solved, efficient and accurate cross-modal query is achieved, and structured query reports are generated.
Patent Information
- Application Number
- CN202411620278.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-14
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-11-14
AI Technical Summary
Existing cross-modal retrieval methods are difficult to effectively capture the complex relationships between different modalities, ignore timing information, and lack structured expressions of the returned results, making it difficult to quickly understand relevant information.
The text mining data query method based on cross-modal similarity is adopted, and text, image and timing features are obtained through multi-level feature extraction networks, and adaptive cross-modal contrast learning networks are built for feature fusion, and structured query reports are generated using multi-level screening models and graph neural networks.
It improves the accuracy and efficiency of the query, can find the target text related to the query text more accurately, and generates a structured query report, enhancing the interpretability of the results.
Smart Images

Figure CN119311854B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to data mining technology, and particularly to a method and system for text mining data query based on cross-modal similarity. Background Art
[0002] Traditional cross-modal retrieval methods usually fuse features of different modalities in a simple way such as concatenation or weighted averaging, and it is difficult to effectively capture the complex associations between different modalities. Different modality data has heterogeneity, and simple feature fusion methods are prone to information loss or noise interference, thus affecting the accuracy of retrieval. For example, directly concatenating the semantic features of text and the visual features of images cannot reflect the fine-grained correspondence between the text description and the image content.
[0003] Many multimedia data, such as videos, contain rich temporal information. Existing cross-modal retrieval methods often ignore the importance of temporal information, resulting in the inability to effectively handle retrieval tasks involving temporal relationships. For example, if a user hopes to retrieve a video clip of "a person runs first and then drinks water", it is difficult to return accurate results if the retrieval method cannot understand the temporal relationship between "running" and "drinking water".
[0004] Existing cross-modal retrieval methods usually only return a simple sorted list, lacking structured expression and in-depth analysis of the retrieval results. Users often need to spend a lot of time screening and sorting the required information from the returned results. For example, if a user hopes to retrieve graphic and text information related to "healthy lifestyle", it is difficult for the user to quickly understand the associations between different healthy lifestyles and their respective characteristics. Ideally, the retrieval results should be presented in a more structured way, such as a knowledge graph or topic clustering, so that users can more effectively understand and utilize the retrieved information. Summary of the Invention
[0005] Embodiments of the present invention provide a method and system for text mining data query based on cross-modal similarity, which can solve the problems in the prior art.
[0006] In the first aspect of the embodiments of the present invention,
[0007] A method for text mining data query based on cross-modal similarity is provided, including:
[0008] Obtain the text to be queried and the target text in the preset text library, where the target text includes image feature data and temporal relationship data; construct a multi-level feature extraction network, which includes a text feature extraction module, an image feature extraction module, and a temporal feature extraction module; input the text to be queried into the text feature extraction module to obtain a text multi-dimensional feature vector, input the image feature data into the image feature extraction module to obtain an image multi-dimensional feature vector, and input the temporal relationship data into the temporal feature extraction module to obtain a temporal multi-dimensional feature vector;
[0009] Construct an adaptive cross-modal contrastive learning network, which includes a feature fusion sub-network and a dynamic weight adjustment module; input the text multi-dimensional feature vector, the image multi-dimensional feature vector, and the temporal multi-dimensional feature vector into the feature fusion sub-network, and perform fusion mapping on different modal features through the dynamic weight adjustment module to obtain an adaptive fusion feature; use a contrastive learning loss function to optimize based on the adaptive fusion feature and construct a cross-modal similarity matrix;
[0010] Construct a multi-level screening model, which includes a coarse-grained filtering layer and a fine-grained evaluation layer; use the multi-level screening model to hierarchically screen the target text based on the cross-modal similarity matrix to obtain a target candidate text set; use a graph neural network to construct a text relationship reasoning module and establish a text relationship graph based on the target candidate text set; adopt a multi-task learning framework to perform text clustering and relationship prediction on the text relationship graph to generate a structured query report.
[0011] Construct a multi-level feature extraction network, including:
[0012] Construct a hierarchical attention network for the text feature extraction module. Convert the text to be queried into an initial word vector through a word vector embedding layer; construct a self-attention matrix based on the initial word vector, and obtain word-level attention scores through the scaled dot product operation of the query vector, key vector, and value vector; perform normalization processing on the word-level attention scores and then perform weighted combination with the value vector to obtain word-level context representations; divide the text data into multiple sentence units, perform pooling operations on the word-level features within each sentence unit to obtain sentence representations; construct an inter-sentence attention network to calculate the semantic dependency relationship between sentences and obtain a text multi-dimensional feature vector;
[0013] Construct a spatial attention network for the image feature extraction module, and extract multi-scale feature maps of the image feature data through a multi-layer convolutional network; construct a channel attention branch and a spatial attention branch for the multi-scale feature maps respectively; in the channel attention branch, obtain channel descriptors through global average pooling and max pooling operations, and input the channel descriptors into a multi-layer perceptron to generate channel weights; in the spatial attention branch, generate a two-dimensional attention map through convolutional operations; perform feature recalibration on the channel weights and the two-dimensional attention map to obtain an image multi-dimensional feature vector;
[0014] Construct a bidirectional long short-term memory network for the temporal feature extraction module, and map the timestamp information of the temporal relationship data into a time embedding vector; construct a forward long short-term memory network layer and a backward long short-term memory network layer, and both the forward long short-term memory network layer and the backward long short-term memory network layer include an input gate unit, a forget gate unit, and an output gate unit; control the input ratio of the current moment information through the input gate unit, adjust the retention degree of historical information through the forget gate unit, and determine the state output information volume through the output gate unit; construct a temporal attention layer on top of the long short-term memory network layer, calculate the association strength of different time steps, and perform weighted combination on the historical states according to the attention weights to obtain an image multi-dimensional feature vector.
[0015] Construct an adaptive cross-modal contrastive learning network, including:
[0016] Construct a feature fusion sub-network, and divide the feature fusion sub-network into a text feature transformation branch, an image feature transformation branch, and a temporal feature transformation branch; set a projection layer in each feature transformation branch, and the projection layer includes a fully connected layer for feature dimension transformation, a batch normalization layer for feature distribution standardization, and an activation function for non-linear mapping; set a cross-modal attention layer after the projection layer, and the cross-modal attention layer includes a query mapping unit, a key-value mapping unit, and an attention calculation unit;
[0017] Input the input text multi-dimensional feature vector, the image multi-dimensional feature vector, and the temporal multi-dimensional feature vector into the projection layers of the text feature transformation branch, the image feature transformation branch, and the temporal feature transformation branch respectively; the projection layer maps feature vectors of different dimensions to the same-dimensional feature space through a fully connected layer, a batch normalization layer, and an activation function to obtain unified feature vectors; the cross-modal attention layer uses the unified feature vector of each modality as a query vector, and the unified feature vectors of other modalities as key-value pairs to calculate the inter-modal attention weights to obtain attention-enhanced modality features;
[0018] Construct a dynamic weight adjustment module, where the dynamic weight adjustment module includes a feature quality evaluation network and a weight generation network; the feature quality evaluation network receives the statistics and distribution information of the attention-enhanced modal features and outputs a feature reliability score; the weight generation network generates dynamic fusion weights based on the feature reliability score and the attention-enhanced modal features through a multi-layer perceptron, and normalizes the dynamic fusion weights; the dynamic fusion weights are weighted and combined with the attention-enhanced modal features to obtain an adaptive fusion feature.
[0019] Optimize based on the adaptive fusion features using a contrastive learning loss function to construct a cross-modal similarity matrix, including:
[0020] Perform data augmentation processing on the adaptive fusion features to generate positive sample pairs, and randomly sample negative sample pairs from the adaptive fusion features of the same batch; construct a unimodal contrastive loss based on the positive sample pairs and the negative sample pairs; construct a modal alignment loss and a modal contrast loss, where the modal alignment loss realizes feature alignment by minimizing the feature distance of the same adaptive fusion feature in different modalities, and the modal contrast loss realizes feature discrimination by increasing the feature distance between different adaptive fusion features, obtaining optimized discriminative features;
[0021] Construct a dynamic neighborhood for each query sample based on the spatial local density of the optimized discriminative features; calculate the correlation weights between the sample and other samples in the neighborhood through an attention mechanism to obtain neighborhood attention weights; calculate the cosine similarity between sample pairs in the space of the optimized discriminative features as the initial similarity;
[0022] Construct a graph attention network, input the neighborhood attention weights and the initial similarity into the graph attention network for feature propagation and aggregation to obtain a similarity representation that fuses local structure information; iteratively update the similarity representation to obtain the final similarity, and construct a cross-modal similarity matrix based on the final similarities of all sample pairs, where the construction of the cross-modal similarity matrix includes direct similarity relationships and indirect similarity information based on local structures.
[0023] Construct a graph attention network, including:
[0024] Construct the input samples into a graph structure, where the nodes in the graph structure represent sample features and the edges represent the association strength between samples; construct an adjacency matrix of the graph based on the pre-obtained neighborhood attention weights, and establish connection relationships between each node and the nodes in its dynamic neighborhood; map the node features to multiple semantic subspaces through a multi-head attention mechanism, and each semantic subspace is used to independently learn the feature association pattern;
[0025] For node pairs in the graph structure, input node features into a weight matrix for linear transformation to obtain a query vector and a key vector. The query vector is used to represent the feature query information of the target node, and the key vector is used to represent the feature response information of the source node; perform matrix multiplication on the query vector and the key vector to obtain an original attention score; multiply the original attention score element-wise with the pre-computed neighborhood attention weight to obtain a local structure constraint score; perform non-linear transformation on the local structure constraint score through an activation function with a negative slope, and use a function for normalization processing to obtain the attention coefficient between node pairs;
[0026] Construct a weighted matrix based on the attention coefficient, and perform matrix multiplication on the weighted matrix and the neighborhood features of the node to obtain an aggregated feature; concatenate the aggregated feature and the original node feature in the feature dimension, and perform non-linear feature transformation through a multi-layer perceptron to obtain a preliminary updated feature; perform a residual connection on the preliminary updated feature and the original node feature, and adjust the feature distribution through a layer normalization network to obtain a normalized feature; selectively retain the effective information of the original feature and the updated feature for the normalized feature through an adaptive gating mechanism to obtain a similarity representation that fuses local structure information.
[0027] Using the multi-level screening model, hierarchically screen the target text based on the cross-modal similarity matrix to obtain a set of target candidate texts, including:
[0028] Perform bidirectional normalization on the cross-modal similarity matrix. The bidirectional normalization includes normalizing the row vectors and applying a normalization function to the column vectors to obtain a preliminarily normalized similarity matrix; calculate the average similarity of each text with its K nearest neighbors as a local similarity feature, and weight-adjust the preliminarily normalized similarity matrix based on the local similarity feature to obtain a similarity matrix that strengthens the local structure;
[0029] Map the similarity matrix with the strengthened local structure to a low-dimensional space and perform product quantization discretization processing to construct a multi-layer index tree structure; set an adaptive search radius based on the multi-layer index tree structure, perform approximate nearest neighbor search within the adaptive search radius, and record the search path information to obtain a preliminary screened candidate text; extract features from the preliminary screened candidate text at multiple semantic levels and calculate the feature similarity score, and fuse the feature similarity score with the search path information to obtain a multi-dimensional evaluation index; construct a weighted ranking model based on the multi-dimensional evaluation index to re-rank the preliminary screened candidate text to obtain a re-ranked candidate text;
[0030] Statistically analyze the similarity distribution characteristics of the re-ranked candidate texts, and adaptively adjust the screening threshold based on the similarity distribution characteristics; calculate the similarity between the re-ranked candidate texts, and perform diversity constraints on the re-ranked candidate texts based on the maximum margin relevance algorithm to obtain an intermediate candidate text set; calculate the confidence scores of the candidate texts in the intermediate candidate text set based on the feature consistency of multiple modalities; preferentially retain the candidate texts with confidence scores higher than the preset threshold to obtain a target candidate text set.
[0031] Adopt a multi-task learning framework to perform text clustering and relationship prediction on the text relationship graph, and generate a structured query report, including:
[0032] Extract word-level features, sentence-level features, and document-level features from the text nodes to obtain word-level features, sentence-level features, and document-level features; perform weighted fusion on the word-level features, the sentence-level features, and the document-level features to obtain text node features; based on the text node features, adopt bidirectional feature propagation to extract features from the relationship edges to obtain initial relationship edge features; calculate information aggregation weights according to the initial relationship edge features, and perform weighted aggregation on the initial relationship edge features to obtain optimized relationship edge features;
[0033] Input the text node features and the optimized relationship edge features into the multi-task learning framework, and the multi-task learning framework includes a text clustering sub-task and a relationship prediction sub-task; for the text clustering sub-task, calculate the node similarity matrix based on the text node features; for the relationship prediction sub-task, calculate the edge feature vector based on the optimized relationship edge features; calculate the probability distribution of the samples to the cluster centers based on the node similarity matrix, and construct a distance metric function in combination with the graph structure information of the text relationship graph; cluster the text nodes according to the distance metric function and the probability distribution to obtain a text clustering result;
[0034] Construct a relationship feature representation from the semantic dimension, the structural dimension, and the temporal dimension based on the edge feature vector, and use a multi-head attention mechanism to fuse the relationship feature representation; generate a relationship prediction result according to the fused relationship feature representation; organize the text clustering result into a topic hierarchy, and use the relationship prediction result as an entity relationship structure; generate a structured query report based on the topic hierarchy and the entity relationship structure.
[0035] In the second aspect of the embodiments of the present invention,
[0036] Provide a cross-modal similarity-based text mining data query system, including:
[0037] The first unit is used to obtain the text to be queried and the target text in the preset text library, where the target text includes image feature data and time series relationship data; construct a multi-level feature extraction network, which includes a text feature extraction module, an image feature extraction module, and a time series feature extraction module; input the text to be queried into the text feature extraction module to obtain a text multi-dimensional feature vector, input the image feature data into the image feature extraction module to obtain an image multi-dimensional feature vector, and input the time series relationship data into the time series feature extraction module to obtain a time series multi-dimensional feature vector;
[0038] The second unit is used to construct an adaptive cross-modal contrast learning network, which includes a feature fusion sub-network and a dynamic weight adjustment module; input the text multi-dimensional feature vector, the image multi-dimensional feature vector, and the time series multi-dimensional feature vector into the feature fusion sub-network, and perform fusion mapping on different modal features through the dynamic weight adjustment module to obtain an adaptive fusion feature; optimize based on the adaptive fusion feature using a contrast learning loss function to construct a cross-modal similarity matrix;
[0039] The third unit is used to construct a multi-level screening model, which includes a coarse-grained filtering layer and a fine-grained evaluation layer; use the multi-level screening model to perform hierarchical screening on the target text based on the cross-modal similarity matrix to obtain a target candidate text set; use a graph neural network to construct a text relationship reasoning module, and establish a text relationship graph based on the target candidate text set; use a multi-task learning framework to perform text clustering and relationship prediction on the text relationship graph to generate a structured query report.
[0040] In the third aspect of the embodiments of the present invention,
[0041] A kind of electronic device is provided, including:
[0042] A processor;
[0043] A memory for storing instructions executable by the processor;
[0044] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0045] In the fourth aspect of the embodiments of the present invention,
[0046] A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0047] The beneficial effects of this application are as follows:
[0048] The present invention constructs a cross-modal similarity matrix by combining text, images, and temporal relationship data, and uses a multi-level screening model for hierarchical screening, effectively filtering out irrelevant information, improving the query accuracy and efficiency, and being able to more accurately find the target text related to the query text. The introduction of the adaptive cross-modal contrast learning network enables the model to adaptively fuse according to the importance of different modal features through dynamic weight adjustment, further enhancing the accuracy of the query.
[0049] The present invention uses a graph neural network to construct a text relationship reasoning module for clustering and relationship prediction of candidate texts, which can not only provide relevant target texts, but also reveal the relationships between target texts, generate a structured query report, enhance the interpretability of the query results, and make it easier for users to understand the logic and associations behind the query results.
[0050] The present invention combines three different types of data (text, images, temporal relationships), enabling it to be applied to a wider range of query scenarios, such as event retrieval and multimedia archive retrieval that simultaneously contain text, pictures, and time information, breaking through the limitations of traditional single-modal text queries. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is a schematic flow chart of the cross-modal similarity text mining data query method according to an embodiment of the present invention;
[0052] Figure 2 is a schematic structural diagram of the cross-modal similarity text mining data query system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0054] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0055] Figure 1 is a schematic flow chart of the cross-modal similarity text mining data query method according to an embodiment of the present invention, as Figure 1 shown, the method includes:
[0056] Obtain the text to be queried and the target text in the preset text library, where the target text includes image feature data and temporal relationship data; construct a multi-level feature extraction network, which includes a text feature extraction module, an image feature extraction module, and a temporal feature extraction module; input the text to be queried into the text feature extraction module to obtain a text multi-dimensional feature vector, input the image feature data into the image feature extraction module to obtain an image multi-dimensional feature vector, and input the temporal relationship data into the temporal feature extraction module to obtain a temporal multi-dimensional feature vector;
[0057] Construct an adaptive cross-modal contrastive learning network, which includes a feature fusion sub-network and a dynamic weight adjustment module; input the text multi-dimensional feature vector, the image multi-dimensional feature vector, and the temporal multi-dimensional feature vector into the feature fusion sub-network, and perform fusion mapping on different modal features through the dynamic weight adjustment module to obtain an adaptive fusion feature; optimize based on the adaptive fusion feature using a contrastive learning loss function to construct a cross-modal similarity matrix;
[0058] Construct a multi-level screening model, which includes a coarse-grained filtering layer and a fine-grained evaluation layer; use the multi-level screening model to hierarchically screen the target text based on the cross-modal similarity matrix to obtain a target candidate text set; use a graph neural network to construct a text relationship reasoning module, and establish a text relationship graph based on the target candidate text set; adopt a multi-task learning framework to perform text clustering and relationship prediction on the text relationship graph to generate a structured query report.
[0059] Obtain the text to be queried and the target text in the preset text library. The target text contains image feature data and temporal relationship data. For example, the text to be queried could be "A red sports car is driving on the highway", and the target text in the preset text library could be a text description containing "A red Ferrari is speeding on the racetrack", corresponding vehicle image feature data (such as color histograms, edge detection features, etc.), and a timestamp sequence indicating the vehicle's driving.
[0060] Construct a multi-level feature extraction network. This network includes a text feature extraction module, an image feature extraction module, and a temporal feature extraction module. The text feature extraction module can use a pre-trained BERT model to convert the query text "A red sports car is driving on the highway" into a sequence of 768-dimensional word vectors, and then obtain a multi-dimensional feature vector of the text through average pooling. The image feature extraction module can use a convolutional neural network (CNN), such as ResNet50, to extract features such as the color, texture, and shape of the image, and convert the image feature data into a 2048-dimensional feature vector. The temporal feature extraction module can adopt a recurrent neural network (RNN), such as LSTM, to process the timestamp sequence, capture the temporal relationship, and output a 128-dimensional temporal feature vector.
[0061] Construct an adaptive cross-modal contrastive learning network. This network includes a feature fusion sub-network and a dynamic weight adjustment module. Input the text multi-dimensional feature vector, the image multi-dimensional feature vector, and the temporal multi-dimensional feature vector into the feature fusion sub-network. The feature fusion sub-network can adopt a multi-layer perceptron (MLP) to map the feature vectors of different modalities to the same dimensional space, such as 512 dimensions. The dynamic weight adjustment module dynamically adjusts the weights according to the importance of different modality features. For example, if the text description is highly relevant to the image content, the weights of the text and image features will be higher. The dynamic weight adjustment module can use an attention mechanism to calculate the weights based on the correlation between different modality features. Concatenate the weighted feature vectors to obtain an adaptive fusion feature. Use the contrastive learning loss function to optimize based on the adaptive fusion feature to construct a cross-modal similarity matrix. For example, construct a similarity matrix by calculating the cosine similarity between the query text and the adaptive fusion features of each target text.
[0062] Construct a multi-level screening model. This model includes a coarse-grained filtering layer and a fine-grained evaluation layer. Use the multi-level screening model to hierarchically screen the target texts based on the cross-modal similarity matrix to obtain a set of target candidate texts. For example, the coarse-grained filtering layer sets a threshold to screen out texts with a similarity lower than the threshold; the fine-grained evaluation layer uses a more complex sorting algorithm, such as a sorting algorithm based on BPR, to sort the remaining texts and select the top K texts as the set of target candidate texts. Assume K = 10, and 10 candidate texts are obtained.
[0063] Use a graph neural network to construct a text relationship reasoning module and establish a text relationship graph based on the set of target candidate texts. For example, take the 10 candidate texts as the nodes of the graph, and the edges between the nodes represent the similarity between the texts. A graph convolutional network (GCN) can be used to model the text relationship graph.
[0064] Use a multi-task learning framework to perform text clustering and relationship prediction on a text relationship graph, and generate a structured query report. For example, simultaneously perform two tasks of text clustering and relationship prediction, and perform weighted summation on the loss functions of the two tasks as the final loss function. Text clustering groups similar texts, and relationship prediction predicts the relationships between texts, such as "contains", "belongs to", etc. Finally, a structured query report containing the text clustering results and relationship prediction results is generated.
[0065] Beneficial effects:
[0066] By integrating text, image, and temporal information, the user's query intention can be understood more accurately, thereby improving the accuracy of query results.
[0067] This method supports queries based on multiple modalities such as text, image, and temporal relationships, expands the query methods, and is convenient for users to use.
[0068] Through the text relationship reasoning module, a structured query report can be generated, providing more comprehensive and easier-to-understand query results.
[0069] In an optional implementation manner, a multi-level feature extraction network is constructed, including:
[0070] Construct a hierarchical attention network for the text feature extraction module. Convert the text to be queried into an initial word vector through a word vector embedding layer; construct a self-attention matrix based on the initial word vector, and obtain word-level attention scores through the scaled dot product operation of the query vector, key vector, and value vector; after normalizing the word-level attention scores, perform weighted combination with the value vector to obtain word-level context representations; divide the text data into multiple sentence units, perform pooling operations on the word-level features within each sentence unit to obtain sentence representations; construct an inter-sentence attention network to calculate the semantic dependency relationship between sentences, and obtain a multi-dimensional text feature vector;
[0071] Construct a spatial attention network for the image feature extraction module. Extract multi-scale feature maps of the image feature data through a multi-layer convolutional network; respectively construct a channel attention branch and a spatial attention branch for the multi-scale feature maps; in the channel attention branch, obtain channel descriptors through global average pooling and max pooling operations, and input the channel descriptors into a multi-layer perceptron to generate channel weights; in the spatial attention branch, generate a two-dimensional attention map through convolutional operations; perform feature recalibration on the channel weights and the two-dimensional attention map to obtain a multi-dimensional image feature vector;
[0072] Construct a bidirectional long short-term memory network for the temporal feature extraction module to map the timestamp information of the temporal relationship data into a temporal embedding vector; construct a forward long short-term memory network layer and a backward long short-term memory network layer, both of which include an input gate unit, a forget gate unit, and an output gate unit; control the input ratio of the current moment information through the input gate unit, adjust the retention degree of historical information through the forget gate unit, and determine the state output information volume through the output gate unit; construct a temporal attention layer on top of the long short-term memory network layer, calculate the association strength of different time steps, and perform weighted combination on the historical states according to the attention weights to obtain an image multi-dimensional feature vector.
[0073] To more effectively extract key information from multi-modal data, this embodiment proposes a multi-level feature extraction network that integrates multi-dimensional features of text, images, and temporal data.
[0074] First, perform text feature extraction. Input the text to be queried into the word vector embedding layer. For example, using Word2Vec or the glove model, convert each word into an initial word vector of a fixed dimension, such as 300 dimensions. Suppose the input text is "The weather is fine today and suitable for outdoor sports". "Today", "weather", "fine", "suitable", "outdoor", "sports" will be converted into corresponding word vectors respectively.
[0075] Then, construct a self-attention matrix based on the initial word vectors. Use the word vector of "today" as the query vector, and perform dot product operations with the word vectors of "today", "weather", "fine", "suitable", "outdoor", "sports" (as key vectors and value vectors) respectively to obtain word-level attention scores. Perform such operations on all words to obtain a self-attention matrix. To avoid the gradient vanishing problem, divide the attention scores by the square root of the word vector dimension for scaling.
[0076] Next, perform normalization processing on the word-level attention scores, such as using the softmax function, to make the sum of the attention scores of each word equal to 1. Then, perform weighted combination of the normalized attention scores with the corresponding value vectors to obtain the word-level context representation of each word. For example, the word-level context representation of "today" is the weighted sum of the word vectors of all other words and their corresponding attention scores.
[0077] Divide the text data into multiple sentence units according to punctuation marks, such as the two sentences "The weather is fine today" and "Suitable for outdoor sports". Perform pooling operations on the word-level features within each sentence unit, such as average pooling or max pooling, to obtain the sentence representation of each sentence.
[0078] Furthermore, construct an inter-sentence attention network to calculate the semantic dependency relationships between sentences. Take the sentence representation of each sentence as input, calculate the attention scores between sentences, and then perform a weighted combination of the attention scores and the sentence representations to obtain the final multi-dimensional feature vector of the text.
[0079] Secondly, perform image feature extraction. Input the image data into a multi-layer convolutional network, such as using a residual network or a VGG network, to extract multi-scale feature maps. Assume that the input image is a color image of 224x224x3. After passing through the convolutional network, multiple feature maps of different scales are obtained, such as 112x112x64, 56x56x128, 28x28x256, etc.
[0080] Construct a channel attention branch and a spatial attention branch for each scale of the feature map. In the channel attention branch, perform global average pooling and max pooling operations on the feature map to obtain two channel descriptors. Concatenate these two channel descriptors and input them into a multi-layer perceptron to generate channel weights.
[0081] In the spatial attention branch, perform a convolutional operation on the feature map to generate a two-dimensional attention map.
[0082] Multiply the channel weights by the two-dimensional attention map to perform feature recalibration on the feature map and obtain the multi-dimensional feature vector of the image.
[0083] Finally, perform temporal feature extraction. Map the timestamp information of the temporal relationship data into a time embedding vector. For example, use sine and cosine functions to encode the timestamp into a vector of a fixed dimension. Assume that the timestamps are 1, 2, 3, 4, 5, and each timestamp is converted into a 128-dimensional time embedding vector.
[0084] Construct a forward long short-term memory network layer and a backward long short-term memory network layer. Both of these network layers contain an input gate unit, a forget gate unit, and an output gate unit. The input gate unit controls the input ratio of the current moment information, the forget gate unit adjusts the retention degree of historical information, and the output gate unit determines the amount of state output information.
[0085] Taking timestamp 1 as an example, the forward LSTM receives the embedding vector of timestamp 1 as input, and the backward LSTM receives the embedding vector of timestamp 5 as input. Control the flow of information through the input gate, forget gate, and output gate to obtain the forward and backward hidden states respectively.
[0086] Concatenate the hidden states of the forward and backward LSTMs to obtain the complete state representation at each time step.
[0087] Construct a temporal attention layer on top of the LSTM layer to calculate the association strength at different time steps. Weightedly combine the historical states according to the attention weights to obtain the final temporal multi-dimensional feature vector.
[0088] The beneficial effects are reflected in three aspects:
[0089] Through the multi-level attention mechanism, the network can more effectively capture key information in text, images, and temporal data, improving the feature expression ability.
[0090] The attention mechanism can highlight the contributions of important features, enhancing the interpretability of the model and helping to understand the model's decision-making process.
[0091] By fusing multi-modal information, the model can better handle complex scenarios and improve the generalization performance.
[0092] In an optional implementation, construct an adaptive cross-modal contrastive learning network, including:
[0093] Construct a feature fusion sub-network, divide the feature fusion sub-network into a text feature transformation branch, an image feature transformation branch, and a temporal feature transformation branch; set a projection layer in each feature transformation branch, and the projection layer includes a fully connected layer for feature dimension transformation, a batch normalization layer for feature distribution standardization, and an activation function for non-linear mapping; set a cross-modal attention layer after the projection layer, and the cross-modal attention layer includes a query mapping unit, a key-value mapping unit, and an attention calculation unit;
[0094] Input the input text multi-dimensional feature vector, the image multi-dimensional feature vector, and the temporal multi-dimensional feature vector into the projection layers of the text feature transformation branch, the image feature transformation branch, and the temporal feature transformation branch respectively; the projection layer maps feature vectors of different dimensions to the same-dimensional feature space through a fully connected layer, a batch normalization layer, and an activation function to obtain unified feature vectors; the cross-modal attention layer uses the unified feature vector of each modality as a query vector, and the unified feature vectors of other modalities as key-value pairs to calculate the inter-modal attention weights, obtaining attention-enhanced modality features;
[0095] Construct a dynamic weight adjustment module, and the dynamic weight adjustment module includes a feature quality evaluation network and a weight generation network; the feature quality evaluation network receives the statistics and distribution information of the attention-enhanced modality features and outputs a feature reliability score; the weight generation network generates dynamic fusion weights through a multi-layer perceptron based on the feature reliability score and the attention-enhanced modality features, and normalizes the dynamic fusion weights; weightedly combine the dynamic fusion weights with the attention-enhanced modality features to obtain adaptive fusion features.
[0096] To achieve adaptive cross-modal contrastive learning, the present invention proposes a new network structure. This network mainly consists of a feature fusion sub-network and a dynamic weight adjustment module.
[0097] First, construct the feature fusion sub-network. This sub-network is divided into a text feature transformation branch, an image feature transformation branch, and a temporal feature transformation branch, and the three branch structures are similar. Taking the text feature transformation branch as an example, it includes a projection layer and a cross-modal attention layer. The projection layer consists of a fully connected layer, a batch normalization layer, and an activation function. The input text multi-dimensional feature vector, for example, a 768-dimensional vector, first undergoes dimension transformation through the fully connected layer, assuming it is transformed to 256 dimensions. Then, the batch normalization layer normalizes the transformed features so that their mean is zero and variance is one. Finally, the activation function, such as the rectified linear unit function, performs a non-linear mapping on the normalized features to enhance the model's expressive power. The other two branches are processed similarly to map the image features and temporal features to the same 256-dimensional feature space.
[0098] After being processed by the projection layer, the three branches respectively obtain 256-dimensional unified feature vectors. Next, the cross-modal attention layer processes these feature vectors. Taking the text feature as an example, the text feature is used as the query vector, and the image feature and temporal feature are used as the key-value pair. The attention calculation unit calculates the attention weights between the text feature and the image feature, and between the text feature and the temporal feature. For example, the scaled dot-product attention mechanism can be used to calculate the dot product between the query vector and each key vector, and then convert the dot product into a probability distribution, that is, the attention weight, through the softmax function. Finally, the value vectors are weighted and summed using the attention weights to obtain the attention-enhanced text feature. The other two branches are processed similarly to obtain the attention-enhanced image feature and temporal feature.
[0099] Then, construct the dynamic weight adjustment module. This module includes a feature quality evaluation network and a weight generation network. The feature quality evaluation network receives the statistics and distribution information of the attention-enhanced modal features, such as mean, variance, kurtosis, etc., and outputs a feature reliability score. For example, if the feature distribution of a modality is relatively concentrated and the variance is small, its reliability score may be high. The weight generation network generates dynamic fusion weights through a multi-layer perceptron based on the feature reliability score and the attention-enhanced modal features. For example, the feature reliability score and the attention-enhanced modal features are concatenated together as the input of the multi-layer perceptron, and the weights corresponding to the three modalities are output. Finally, the dynamic fusion weights are normalized, for example, using the softmax function, to ensure that the sum of the three weights is 1.
[0100] Finally, the dynamically fused weights are weighted and combined with the attention-enhanced modal features to obtain the adaptive fusion features. For example, assuming that the dynamically fused weights of text, image, and temporal features are 0.2, 0.5, and 0.3 respectively, then the attention-enhanced text features are multiplied by 0.2, the attention-enhanced image features are multiplied by 0.5, and the attention-enhanced temporal features are multiplied by 0.3. Then, the three results are added together to obtain the final adaptive fusion features.
[0101] For example, input a text describing a cat playing, a picture of a cat playing, and a video clip of a cat playing. The dimension of the text feature vector is 768, the dimension of the image feature vector is 1024, and the dimension of the temporal feature vector is 512. After being processed by the feature fusion sub-network, the feature vectors of the three modalities are all mapped to a 256-dimensional space and enhanced through the cross-modal attention mechanism. Suppose the feature quality evaluation network determines that the quality of the image modality is the highest, followed by the text modality, and the quality of the temporal modality is the lowest. Then the weight generation network may generate a weight vector such as 0.2, 0.6, 0.2. Finally, the fused feature vector will be more affected by the image modality.
[0102] The beneficial effects provided by the present invention can be summarized into the following three aspects:
[0103] Through the cross-modal attention mechanism, the features of different modalities can complement and enhance each other, thus obtaining a more comprehensive and expressive feature representation.
[0104] The dynamic weight adjustment module can dynamically adjust the fusion weights according to the quality of different modal features, avoiding the limitations of the fixed-weight fusion method, and thus making better use of multi-modal information.
[0105] By using the feature quality evaluation network to evaluate the quality of different modal features, the influence of noise and low-quality features can be effectively reduced, and the robustness of the model can be enhanced.
[0106] In an alternative embodiment, a contrastive learning loss function is used to optimize based on the adaptive fusion features, and a cross-modal similarity matrix is constructed, including:
[0107] Performing data augmentation processing on the adaptive fusion features to generate positive sample pairs, and randomly sampling from the adaptive fusion features of the same batch to obtain negative sample pairs; constructing a unimodal contrastive loss based on the positive sample pairs and the negative sample pairs; constructing a modal alignment loss and a modal contrast loss, where the modal alignment loss realizes feature alignment by minimizing the feature distance of the same adaptive fusion feature under different modalities, and the modal contrast loss realizes feature discrimination by increasing the feature distance between different adaptive fusion features, so as to obtain the optimized discriminative features;
[0108] Construct a dynamic neighborhood for each query sample based on the spatial local density of the optimized discriminative features; calculate the correlation weights between the sample and other samples in the neighborhood through an attention mechanism within the dynamic neighborhood to obtain neighborhood attention weights; calculate the cosine similarity between sample pairs in the optimized discriminative feature space as the initial similarity;
[0109] Construct a graph attention network, input the neighborhood attention weights and the initial similarity into the graph attention network for feature propagation and aggregation, and obtain a similarity representation that fuses local structure information; iteratively update the similarity representation to obtain the final similarity, and construct a cross-modal similarity matrix based on the final similarities of all sample pairs. The construction of the cross-modal similarity matrix includes direct similarity relationships and indirect similarity information based on local structures.
[0110] To construct a cross-modal similarity matrix, this embodiment proposes a scheme based on a contrastive learning loss function and adaptive fusion feature optimization. This scheme uses contrastive learning to optimize the feature representation and combines a graph attention network to fuse local structure information, thereby improving the performance of cross-modal retrieval.
[0111] First, extract data features of different modalities and project them into a common feature space to obtain adaptive fusion features. For example, for the two modalities of images and texts, use a pre-trained convolutional neural network and a transformer model to extract features respectively, and then project the extracted features through a fully connected layer to the same dimension, such as 512 dimensions. Assume that a batch contains 16 samples, and each sample contains data of both image and text modalities.
[0112] Next, perform data augmentation on the adaptive fusion features to generate positive sample pairs. The ways of data augmentation can include random cropping, color jittering, adding noise, etc. For example, perform random cropping on the image modality, and use the cropped image and the original image as a positive sample pair. Randomly sample from other adaptive fusion features in the same batch to obtain negative sample pairs. For example, randomly select one sample from the remaining 15 samples and pair it with the current sample to form a negative sample pair.
[0113] Then, a unimodal contrastive loss is constructed based on positive sample pairs and negative sample pairs. The goal of the contrastive loss is to reduce the distance between positive sample pairs while increasing the distance between negative sample pairs. In addition, a modality alignment loss and a modality contrast loss are constructed. The modality alignment loss realizes feature alignment by minimizing the feature distance of the same adaptive fusion feature in different modalities. For example, the Euclidean distance between the image modality feature and the text modality feature is calculated and used as part of the loss function. The modality contrast loss realizes feature discrimination by increasing the feature distance between different adaptive fusion features. For example, the Euclidean distance between the image modality features of different samples is calculated and used as part of the loss function. By minimizing the weighted sum of these three loss functions, the optimized discriminative features are obtained.
[0114] Next, a dynamic neighborhood is constructed for each query sample based on the spatial local density of the optimized discriminative features. For example, the Euclidean distance between each sample and other samples in the feature space is calculated, and the K samples with the closest distance are selected as its neighborhood according to the distance ranking. Assume K = 5.
[0115] Within the dynamic neighborhood, the correlation weights between the sample and other samples in the neighborhood are calculated through an attention mechanism to obtain the neighborhood attention weights. For example, the multi-head self-attention mechanism can be used to calculate the weights.
[0116] Then, the cosine similarity between the sample pairs in the optimized discriminative feature space is calculated as the initial similarity.
[0117] A graph attention network is constructed, and the neighborhood attention weights and the initial similarity are input into the graph attention network for feature propagation and aggregation to obtain a similarity representation that fuses local structure information.
[0118] The similarity representation is iteratively updated to obtain the final similarity. The way of iterative update can be a multi-layer graph attention network.
[0119] Finally, a cross-modal similarity matrix is constructed based on the final similarities of all sample pairs. This matrix contains direct similarity relationships and indirect similarity information based on local structures. For example, a 16x16 matrix, where each element represents the final similarity of the corresponding sample pair.
[0120] Suppose there are two samples, and their image features are [0.1, 0.2, 0.3] and [0.2, 0.3, 0.4] respectively, and their text features are [0.4, 0.5, 0.6] and [0.5, 0.6, 0.7] respectively. After the above steps, the value corresponding to these two samples in the similarity matrix is 0.95.
[0121] The beneficial effects of this solution can be summarized in the following three aspects:
[0122] More discriminative feature representations are learned through contrastive learning and modality alignment, thereby improving the accuracy of cross-modal retrieval.
[0123] Through the graph attention network, local structural information is fused, thus better capturing the relationships between samples and further enhancing the retrieval performance.
[0124] The adaptive fusion of features and data augmentation strategies improve the robustness of the model to noise and data variations.
[0125] In an optional implementation manner, a graph attention network is constructed, including:
[0126] The input samples are constructed into a graph structure, where the nodes in the graph structure represent sample features and the edges represent the association strength between samples; an adjacency matrix of the graph is constructed based on the pre-obtained neighborhood attention weights, and each node establishes a connection relationship with the nodes in its dynamic neighborhood; the node features are mapped to multiple semantic subspaces through a multi-head attention mechanism, and each semantic subspace is used to independently learn the feature association pattern;
[0127] For a pair of nodes in the graph structure, the node features are input into a weight matrix for linear transformation to obtain a query vector and a key vector. The query vector is used to represent the feature query information of the target node, and the key vector is used to represent the feature response information of the source node; the query vector and the key vector are subjected to matrix multiplication operation to obtain an original attention score; the original attention score and the pre-computed neighborhood attention weights are multiplied element-wise to obtain a local structure constraint score; the local structure constraint score is non-linearly transformed through an activation function with a negative slope and is normalized using a function to obtain the attention coefficient between the node pairs;
[0128] A weighted matrix is constructed based on the attention coefficients, and the weighted matrix and the neighborhood features of the nodes are subjected to matrix multiplication operation to obtain an aggregated feature; the aggregated feature and the original node features are concatenated in the feature dimension, and non-linear feature transformation is performed through a multi-layer perceptron to obtain a preliminary updated feature; the preliminary updated feature and the original node features are subjected to residual connection, and the feature distribution is adjusted through a layer normalization network to obtain a normalized feature; the normalized feature selectively retains the effective information of the original feature and the updated feature through an adaptive gating mechanism to obtain a similarity representation that fuses local structural information.
[0129] A method for constructing a graph attention network for learning the similarity representation between samples is as follows in its specific implementation manner:
[0130] First, construct the input samples into a graph structure. The features of each sample are represented as a node in the graph. The association strength between samples is represented as an edge in the graph. For example, assume there are three samples with feature vectors [0.2, 0.5, 0.8], [0.1, 0.7, 0.2], and [0.3, 0.2, 0.5] respectively. The association strength between samples can be calculated based on the cosine similarity between feature vectors. Suppose the calculated similarities are 0.8, 0.6, and 0.7 respectively, then these three samples can be constructed into a graph, where nodes represent sample features and edges represent the association strength between samples.
[0131] Next, construct the adjacency matrix of the graph based on the pre - obtained neighborhood attention weights. Each node establishes a connection relationship with the nodes in its dynamic neighborhood. The neighborhood attention weights can be obtained by calculating the similarity or distance between nodes. For example, a threshold can be set, and only nodes with a similarity greater than this threshold will be considered neighbor nodes. Assume the threshold is 0.7, then in the above example, sample 1 and sample 3 are neighbors, and sample 2 and sample 3 are also neighbors. Thus, an adjacency matrix can be constructed, where 1 represents connection and 0 represents non - connection.
[0132] Then, map the node features to multiple semantic sub - spaces through the multi - head attention mechanism. Each semantic sub - space is used to independently learn the feature association pattern. For example, 8 attention heads can be used to map the feature vector of each node into 8 different semantic sub - spaces.
[0133] For node pairs in the graph structure, input the node features into the weight matrix for linear transformation to obtain the query vector and the key vector. The query vector is used to represent the feature query information of the target node, and the key vector is used to represent the feature response information of the source node. For example, for sample 1 and sample 3, their feature vectors can be multiplied by different weight matrices respectively to obtain the corresponding query vector and key vector.
[0134] Perform matrix multiplication on the query vector and the key vector to obtain the original attention score. For example, perform matrix multiplication on the query vector of sample 1 and the key vector of sample 3 to obtain a score, which represents the attention degree of sample 1 to sample 3.
[0135] Multiply the original attention score element - by - element with the pre - calculated neighborhood attention weights to obtain the local structure constraint score. For example, multiply the original attention scores of sample 1 and sample 3 with their neighborhood attention weight (0.7) to obtain the local structure constraint score.
[0136] The local structural constraint scores are non-linearly transformed through an activation function with a negative slope and normalized using the softmax function to obtain the attention coefficients between node pairs. For example, the leaky rectified linear unit activation function can be used to non-linearly transform the local structural constraint scores, and then the softmax function is used for normalization to obtain the attention coefficient between sample 1 and sample 3.
[0137] A weighted matrix is constructed based on the attention coefficients, and the weighted matrix is multiplied with the neighborhood features of the nodes to obtain the aggregated features. For example, the feature vector of sample 3 is multiplied by the attention coefficient between sample 1 and sample 3, and then the results are added to obtain the aggregated feature of sample 3.
[0138] The aggregated features and the original features of the nodes are concatenated in the feature dimension, and non-linear feature transformation is performed through a multi-layer perceptron to obtain the preliminary updated features. For example, the aggregated feature of sample 3 and its original feature are concatenated together, and then non-linear transformation is performed through a two-layer multi-layer perceptron to obtain the preliminary updated features.
[0139] The preliminary updated features and the original features of the nodes are connected by residual connection, and the feature distribution is adjusted through a layer normalization network to obtain the normalized features. For example, the preliminary updated features are added to the original features of sample 3, and then adjusted through the layer normalization network to obtain the normalized features.
[0140] The effective information of the original features and the updated features is selectively retained for the normalized features through an adaptive gating mechanism to obtain the similarity representation that fuses local structural information. For example, the sigmoid function can be used to control the retention ratio of the original features and the updated features, and finally the similarity representation that fuses local structural information is obtained.
[0141] The beneficial effects of this method can be summarized in the following three aspects:
[0142] By constructing the samples as a graph structure and using the graph attention network to learn the correlation relationships between nodes, the complex relationships between samples can be better captured, thereby improving the expression ability of the model.
[0143] The attention mechanism can help us understand which nodes and which features the model focuses on, thereby enhancing the interpretability of the model.
[0144] By fusing local structural information, the similarity representation of the samples can be better learned, thereby improving the performance of the model in various downstream tasks, such as classification, clustering, etc.
[0145] In an alternative embodiment, the multi-level screening model is used to hierarchically screen the target text based on the cross-modal similarity matrix to obtain a target candidate text set, including:
[0146] Perform bidirectional normalization on the cross-modal similarity matrix. The bidirectional normalization includes normalizing the row vectors and applying a normalization function to the column vectors to obtain a preliminarily normalized similarity matrix; calculate the K average similarities between each text and its neighbors as local similarity features, and based on the local similarity features, perform weighted adjustment on the preliminarily normalized similarity matrix to obtain a similarity matrix that strengthens the local structure;
[0147] Map the similarity matrix with strengthened local structure to a low-dimensional space and perform product quantization discretization processing to construct a multi-layer index tree structure; set an adaptive search radius based on the multi-layer index tree structure, perform approximate nearest neighbor search within the adaptive search radius, and record the search path information to obtain a preliminary screened candidate text; extract features and calculate feature similarity scores for the preliminary screened candidate text at multiple semantic levels, and fuse the feature similarity scores with the search path information to obtain a multi-dimensional evaluation index; construct a weighted ranking model based on the multi-dimensional evaluation index, and re-rank the preliminary screened candidate text to obtain a re-ranked candidate text;
[0148] Statistically analyze the similarity distribution characteristics of the re-ranked candidate text, and adaptively adjust the screening threshold based on the similarity distribution characteristics; calculate the similarities between the re-ranked candidate text, perform diversity constraints on the re-ranked candidate text based on the maximum margin relevance algorithm to obtain an intermediate candidate text set; calculate the confidence scores of each candidate text in the intermediate candidate text set based on the feature consistency of multiple modalities; preferentially retain the candidate texts with confidence scores higher than the preset threshold to obtain a target candidate text set.
[0149] A cross-modal text retrieval method aims to quickly and accurately find candidate texts semantically similar to the target text from a large amount of text data. This method uses a multi-level screening model, and through a series of processing and screening on the cross-modal similarity matrix, finally obtains a high-quality target candidate text set.
[0150] First, perform bidirectional normalization on the input cross-modal similarity matrix. This step includes normalizing each row of the matrix so that the sum of the values in each row is 1, and applying a normalization function, such as the commonly used normalization exponential function, to each column of the matrix to convert the values in each column into a probability distribution.
[0151] For example, a 3x3 similarity matrix:
[0152] [[0.1,0.2,0.3],[0.4,0.5,0.6],[0.7,0.8,0.9]], after row normalization, it becomes [[0.14,0.29,0.43],[0.27,0.33,0.40],[0.26,0.30,0.34]],
[0153] After column normalization (such as the normalization exponential function), it becomes:
[0154] [[0.24,0.24,0.26],[0.33,0.33,0.34],[0.43,0.43,0.40]]. The preliminary normalized similarity matrix is obtained.
[0155] Next, calculate the average similarity between each text and its K nearest neighbor texts as the local similarity feature. Assume K = 2. For the first text, its nearest neighbors are the second and third texts, then its local similarity feature is (0.24 + 0.26) / 2 = 0.25. And so on, calculate the local similarity features of all texts. Then, based on these local similarity features, the preliminary normalized similarity matrix is weighted and adjusted. For example, multiply the local similarity feature of each text by its corresponding row vector to obtain a similarity matrix that strengthens the local structure. Taking the first text as an example, its corresponding row vector [0.24, 0.24, 0.26] is multiplied by its local similarity feature 0.25 to get [0.06, 0.06, 0.065].
[0156] Then, map the similarity matrix that strengthens the local structure to a low-dimensional space, such as using dimensionality reduction methods like PCA or t-SNE. And perform product quantization and discretization processing on the data in the low-dimensional space, such as using the K-means clustering algorithm. Based on these discretized data points, construct a multi-level index tree structure, such as a KD tree or a Ball tree.
[0157] On the constructed multi-level index tree structure, set an adaptive search radius. For example, according to the position of the target text in the low-dimensional space and the distribution of the data, dynamically adjust the size of the search radius. Perform approximate nearest neighbor search within the adaptive search radius and record the search path information, such as the tree nodes passed through and the distance to each node. The preliminary screened candidate texts are obtained.
[0158] Extract features from the preliminary screened candidate texts at multiple semantic levels, such as word vectors, sentence vectors, and paragraph vectors, etc. Calculate the similarity scores of these features, such as using cosine similarity. Integrate these feature similarity scores with the search path information, such as weighted summation of them, to obtain a multi-dimensional evaluation metric.
[0159] Construct a weighted ranking model based on multi-dimensional evaluation metrics, such as using linear regression or learning-to-rank algorithms. Re-rank the initially screened candidate texts to obtain re-ranked candidate texts.
[0160] Statistically analyze the similarity distribution characteristics of the re-ranked candidate texts, such as the mean and variance of similarity. Adaptively adjust the screening threshold based on these distribution characteristics. For example, dynamically adjust the size of the screening threshold according to the mean and variance of similarity.
[0161] Calculate the similarity between the re-ranked candidate texts, such as using cosine similarity. Apply diversity constraints to the re-ranked candidate texts based on the maximum marginal relevance algorithm, such as selecting texts with lower similarity but still relevant to the target text. Obtain an intermediate candidate text set.
[0162] Calculate the confidence scores of each candidate text in the intermediate candidate text set based on the feature consistency of multiple modalities. For example, perform weighted averaging on the feature similarities of different modalities to obtain the final confidence scores.
[0163] Prioritize and retain the candidate texts with confidence scores higher than the preset threshold to obtain a target candidate text set.
[0164] The beneficial effects of this method can be summarized in the following three aspects:
[0165] Through the multi-level index tree structure and adaptive search radius, candidate texts can be quickly screened out from massive text data, significantly improving the retrieval efficiency.
[0166] Through the multi-level screening model and multi-dimensional evaluation metrics, the similarity between candidate texts and the target text can be evaluated more accurately, thereby improving the retrieval accuracy.
[0167] Through diversity constraints using the maximum marginal relevance algorithm, it is possible to avoid returning overly similar candidate texts, ensuring the diversity of the results, and thus meeting the diverse needs of users.
[0168] In an alternative implementation, a multi-task learning framework is adopted to perform text clustering and relationship prediction on the text relationship graph, generating a structured query report, including:
[0169] Extract word-level features, sentence-level features, and document-level features from the text nodes to obtain word-level features, sentence-level features, and document-level features; perform weighted fusion on the word-level features, the sentence-level features, and the document-level features to obtain text node features; based on the text node features, use bidirectional feature propagation to extract features from the relationship edges to obtain initial relationship edge features; calculate information aggregation weights according to the initial relationship edge features, and perform weighted aggregation on the initial relationship edge features to obtain optimized relationship edge features;
[0170] Input the text node features and the optimized relationship edge features into a multi-task learning framework, which includes a text clustering sub-task and a relationship prediction sub-task; for the text clustering sub-task, calculate the node similarity matrix based on the text node features; for the relationship prediction sub-task, calculate the edge feature vector based on the optimized relationship edge features; calculate the probability distribution of samples to the cluster centers based on the node similarity matrix, and construct a distance metric function in combination with the graph structure information of the text relationship graph; cluster the text nodes according to the distance metric function and the probability distribution to obtain the text clustering result;
[0171] Construct a relationship feature representation from the semantic dimension, structural dimension, and temporal dimension based on the edge feature vector, and use the multi-head attention mechanism to fuse the relationship feature representation; generate a relationship prediction result according to the fused relationship feature representation; organize the text clustering result into a topic hierarchy, and use the relationship prediction result as the entity relationship structure; generate a structured query report based on the topic hierarchy and the entity relationship structure.
[0172] To achieve effective analysis and query of the text relationship graph, a text clustering and relationship prediction method based on a multi-task learning framework is proposed, and a structured query report is generated accordingly.
[0173] First, extract features from the text nodes. Taking a text relationship graph containing multiple documents as an example, where the document content includes "Artificial intelligence is developing rapidly", "Deep learning is a key technology of artificial intelligence", "Machine learning is the foundation of artificial intelligence", etc. Preprocess each document by word segmentation, part-of-speech tagging, named entity recognition, etc. At the word level, a word embedding model (such as Word2Vec, GloVe) can be used to convert each word into a vector representation as the word-level feature. At the sentence level, a sentence embedding model (such as Sentence-BERT) can be used to convert each sentence into a vector representation as the sentence-level feature. At the document level, the word frequency, TF-IDF value, etc. in the document can be counted, or a model such as Doc2Vec can be used to convert the document into a vector representation as the document-level feature. Then, perform weighted fusion on the word-level feature, sentence-level feature, and document-level feature. For example, weights can be set according to experience or experimental results, and the three feature vectors can be linearly combined to obtain the final text node features.
[0174] Next, feature extraction is performed on the relational edges. Using the text node features obtained in the previous step, a bidirectional feature propagation algorithm is used to extract features from the relational edges. Specifically, for the relational edge connecting node A and node B, the feature vectors of node A and node B are concatenated or combined in other ways to obtain the initial relational edge features. Then, the information aggregation weights are calculated based on the initial relational edge features. For example, the weights can be calculated according to the cosine similarity of the two node feature vectors. Finally, the initial relational edge features are weighted and aggregated to obtain the optimized relational edge features.
[0175] Then, the text node features and the optimized relational edge features are input into the multi-task learning framework. This framework includes a text clustering sub-task and a relationship prediction sub-task. For the text clustering sub-task, a similarity matrix between nodes is calculated based on the text node features. For example, cosine similarity or Euclidean distance can be used to calculate the similarity between nodes. For the relationship prediction sub-task, edge feature vectors are calculated based on the optimized relational edge features.
[0176] In the text clustering sub-task, the probability distribution of samples to the cluster centers is calculated based on the similarity matrix between nodes. At the same time, a distance metric function is constructed by combining the graph structure information of the text relationship graph. For example, the shortest path length between nodes can be incorporated into the distance metric function. The text nodes are clustered according to the distance metric function and the probability distribution. For example, the K-Means algorithm is used to obtain the text clustering results. For example, documents such as "Artificial intelligence is developing rapidly", "Deep learning is a key technology of artificial intelligence", and "Machine learning is the foundation of artificial intelligence" are clustered under the "Artificial intelligence" theme.
[0177] In the relationship prediction sub-task, relationship feature representations are constructed from the semantic dimension, structural dimension, and temporal dimension based on the edge feature vectors. For example, in the semantic dimension, the word vector similarity of the two nodes connected by the edge can be used; in the structural dimension, the connection relationship between the two nodes in the graph can be used; in the temporal dimension, the time order of the appearance of the two nodes can be used. A multi-head attention mechanism is used to fuse the relationship feature representations. For example, different attention heads are used to focus on different feature dimensions, and then the outputs of multiple attention heads are concatenated or weighted averaged. Relationship prediction results are generated based on the fused relationship feature representations. For example, it is predicted that there is a "sub-field" relationship between the document "Deep learning is a key technology of artificial intelligence" and the document "Machine learning is the foundation of artificial intelligence".
[0178] Finally, organize the text clustering results into a thematic hierarchy and use the relationship prediction results as the entity relationship structure. For example, take "Artificial Intelligence" as the upper-level theme, and "Deep Learning" and "Machine Learning" as the lower-level themes, and predict the relationship between "Deep Learning" and "Artificial Intelligence" as a "sub-field" relationship. Generate a structured query report based on the thematic hierarchy and the entity relationship structure. For example, generate a JSON-format report containing the thematic hierarchy and the entity relationship structure to facilitate user query and browsing.
[0179] The beneficial effects of this method are reflected in three aspects:
[0180] Improve query efficiency: By organizing text information into a structured thematic hierarchy and entity relationship structure, it is convenient for users to quickly locate the required information and improve query efficiency.
[0181] Enhance the interpretability of query results: By presenting the thematic hierarchy and entity relationship structure, the association relationships between text information can be clearly shown, enhancing the interpretability of query results.
[0182] Support more complex query requirements: By combining the thematic hierarchy and entity relationship structure, more complex query requirements can be supported, such as joint queries based on themes and relationships.
[0183] Figure 2 This is a schematic structural diagram of the cross-modal similarity text mining data query system according to an embodiment of the present invention, as Figure 2 shown, the system includes:
[0184] A first unit for obtaining the text to be queried and the target text in a preset text library, where the target text includes image feature data and time series relationship data; constructing a multi-level feature extraction network, the multi-level feature extraction network including a text feature extraction module, an image feature extraction module, and a time series feature extraction module; inputting the text to be queried into the text feature extraction module to obtain a text multi-dimensional feature vector, inputting the image feature data into the image feature extraction module to obtain an image multi-dimensional feature vector, and inputting the time series relationship data into the time series feature extraction module to obtain a time series multi-dimensional feature vector;
[0185] A second unit for constructing an adaptive cross-modal contrast learning network, the adaptive cross-modal contrast learning network including a feature fusion sub-network and a dynamic weight adjustment module; inputting the text multi-dimensional feature vector, the image multi-dimensional feature vector, and the time series multi-dimensional feature vector into the feature fusion sub-network, performing fusion mapping on different modal features through the dynamic weight adjustment module to obtain an adaptive fusion feature; optimizing based on the adaptive fusion feature using a contrast learning loss function to construct a cross-modal similarity matrix;
[0186] A third unit for constructing a multi-level screening model, the multi-level screening model including a coarse-grained filtering layer and a fine-grained evaluation layer; using the multi-level screening model to perform hierarchical screening on the target text based on the cross-modal similarity matrix to obtain a target candidate text set; using a graph neural network to construct a text relationship reasoning module, and establishing a text relationship graph based on the target candidate text set; adopting a multi-task learning framework to perform text clustering and relationship prediction on the text relationship graph to generate a structured query report.
[0187] In a third aspect of the embodiments of the present invention,
[0188] There is provided an electronic device, including:
[0189] A processor;
[0190] A memory for storing instructions executable by the processor;
[0191] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0192] In a fourth aspect of the embodiments of the present invention,
[0193] There is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0194] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.
[0195] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data query method based on cross-modal similarity text mining, characterized in that: include: Acquire a text to be queried and a target text in a preset text library, wherein the target text includes image feature data and temporal relationship data; construct a multi-level feature extraction network, wherein the multi-level feature extraction network includes a text feature extraction module, an image feature extraction module, and a temporal feature extraction module; Input the text to be queried into the text feature extraction module to obtain a text multidimensional feature vector, input the image feature data into the image feature extraction module to obtain an image multidimensional feature vector, and input the time series relationship data into the time series feature extraction module to obtain a time series multidimensional feature vector; Constructing an adaptive cross-modal contrastive learning network, the adaptive cross-modal contrastive learning network includes a feature fusion subnetwork and a dynamic weight adjustment module; inputting the text multidimensional feature vector, the image multidimensional feature vector and the time series multidimensional feature vector into the feature fusion subnetwork, fusing and mapping different modal features through the dynamic weight adjustment module to obtain adaptive fusion features; optimizing based on the adaptive fusion features using a contrastive learning loss function to construct a cross-modal similarity matrix; A multi-level screening model is constructed, wherein the multi-level screening model includes a coarse-grained filtering layer and a fine-grained evaluation layer; the target text is hierarchically screened based on the cross-modal similarity matrix using the multi-level screening model to obtain a target candidate text set; a text relationship reasoning module is constructed using a graph neural network, and a text relationship graph is established based on the target candidate text set; a multi-task learning framework is used to perform text clustering and relationship prediction on the text relationship graph to generate a structured query report.
2. The method according to claim 1, characterized in that Construct a multi-level feature extraction network, including: Constructing a hierarchical attention network of a text feature extraction module, converting the query text into an initial word vector through a word vector embedding layer; constructing a self-attention matrix based on the initial word vector, wherein the self-attention matrix obtains a word-level attention score through a scaled dot product operation of a query vector, a key vector, and a value vector; normalizing the word-level attention score and performing a weighted combination with the value vector to obtain a word-level context representation; dividing the text data into a plurality of sentence units, performing a pooling operation on the word-level features in each sentence unit to obtain a sentence representation; constructing an inter-sentence attention network to calculate the semantic dependency between sentences, and obtaining a multi-dimensional feature vector of the text; Constructing a spatial attention network of an image feature extraction module, extracting a multi-scale feature map of the image feature data through a multi-layer convolutional network; constructing a channel attention branch and a spatial attention branch for the multi-scale feature map; in the channel attention branch, obtaining a channel descriptor through global average pooling and maximum pooling operations, and inputting the channel descriptor into a multi-layer perceptron to generate a channel weight; in the spatial attention branch, generating a two-dimensional attention map through a convolution operation; performing feature recalibration on the channel weight and the two-dimensional attention map to obtain a multi-dimensional feature vector of the image; A bidirectional long short-term memory network of a temporal feature extraction module is constructed to map the timestamp information of the temporal relationship data into a time embedding vector; a forward long short-term memory network layer and a backward long short-term memory network layer are constructed, and both the forward long short-term memory network layer and the backward long short-term memory network layer include an input gating unit, a forgetting gating unit and an output gating unit; the input gating unit is used to control the current moment information input ratio, the forgetting gating unit is used to adjust the degree of historical information retention, and the output gating unit is used to determine the state output information amount; a temporal attention layer is constructed on top of the long short-term memory network layer, the correlation strength of different time steps is calculated, and the historical states are weightedly combined according to the attention weights to obtain a multi-dimensional feature vector of the image.
3. The method according to claim 1, characterized in that: Construct an adaptive cross-modal contrastive learning network, including: Constructing a feature fusion subnetwork, dividing the feature fusion subnetwork into a text feature transformation branch, an image feature transformation branch, and a temporal feature transformation branch; setting a projection layer in each feature transformation branch, wherein the projection layer includes a fully connected layer for feature dimension transformation, a batch normalization layer for feature distribution standardization, and an activation function for nonlinear mapping; setting a cross-modal attention layer after the projection layer, wherein the cross-modal attention layer includes a query mapping unit, a key-value mapping unit, and an attention calculation unit; The input text multidimensional feature vector, the image multidimensional feature vector and the time series multidimensional feature vector are respectively input into the projection layers of the text feature transformation branch, the image feature transformation branch and the time series feature transformation branch; the projection layer maps the feature vectors of different dimensions to the feature space of the same dimension through the fully connected layer, the batch normalization layer and the activation function to obtain a unified feature vector; the cross-modal attention layer uses the unified feature vector of each modality as a query vector and the unified feature vectors of other modalities as a key-value pair to calculate the inter-modal attention weights, and obtains the attention-enhanced modal features; A dynamic weight adjustment module is constructed, and the dynamic weight adjustment module includes a feature quality assessment network and a weight generation network; the feature quality assessment network receives the statistics and distribution information of the attention-enhanced modal features, and outputs a feature reliability score; the weight generation network generates dynamic fusion weights based on the feature reliability score and the attention-enhanced modal features through a multi-layer perceptron, and normalizes the dynamic fusion weights; the dynamic fusion weights are weightedly combined with the attention-enhanced modal features to obtain an adaptive fusion feature.
4. The method according to claim 1, characterized in that: The contrastive learning loss function is used to optimize based on the adaptive fusion features to construct a cross-modal similarity matrix, including: Performing data enhancement processing on the adaptive fusion features to generate positive sample pairs, and randomly sampling from the adaptive fusion features of the same batch to obtain negative sample pairs; constructing a single modality contrast loss based on the positive sample pairs and the negative sample pairs; constructing a modality alignment loss and a modality contrast loss, wherein the modality alignment loss realizes feature alignment by minimizing the feature distance of the same adaptive fusion feature in different modalities, and the modality contrast loss realizes feature discrimination by increasing the feature distance between different adaptive fusion features, thereby obtaining an optimized discrimination feature; Based on the spatial local density of the optimized discriminant feature, a dynamic neighborhood is constructed for each query sample; in the dynamic neighborhood, the correlation weight between the sample and other samples in the neighborhood is calculated by the attention mechanism to obtain the neighborhood attention weight; and the cosine similarity of the sample pair in the optimized discriminant feature space is calculated as the initial similarity; A graph attention network is constructed, and the neighborhood attention weights and the initial similarities are input into the graph attention network for feature propagation and aggregation to obtain a similarity representation that integrates local structural information; the similarity representation is iteratively updated to obtain a final similarity, and a cross-modal similarity matrix is constructed based on the final similarities of all sample pairs, wherein the constructed cross-modal similarity matrix includes direct similarity relationships and indirect similarity information based on local structures.
5. The method according to claim 4, characterized in that Construct a graph attention network, including: The input samples are constructed into a graph structure, in which nodes represent sample features and edges represent the strength of association between samples; an adjacency matrix of the graph is constructed based on pre-acquired neighborhood attention weights, and each node is connected to nodes in its dynamic neighborhood; the node features are mapped to multiple semantic subspaces through a multi-head attention mechanism, and each semantic subspace is used to independently learn feature association patterns; For the node pairs in the graph structure, the node features are input into the weight matrix for linear transformation to obtain a query vector and a key vector, wherein the query vector is used to represent the feature query information of the target node, and the key vector is used to represent the feature response information of the source node; the query vector and the key vector are matrix multiplied to obtain the original attention score; the original attention score is element-wise multiplied with the pre-calculated neighborhood attention weight to obtain the local structure constraint score; the local structure constraint score is nonlinearly transformed through an activation function with a negative slope, and normalized using a function to obtain the attention coefficient between the node pairs; A weighted matrix is constructed based on the attention coefficient, and matrix multiplication operation is performed on the weighted matrix and the neighborhood features of the node to obtain aggregated features; the aggregated features are spliced with the original features of the node in the feature dimension, and nonlinear feature transformation is performed through a multi-layer perceptron to obtain preliminary updated features; the preliminary updated features are residually connected with the original features of the node, and the feature distribution is adjusted through a layer normalization network to obtain standardized features; the standardized features are selectively retained through an adaptive gating mechanism to obtain the effective information of the original features and the updated features, so as to obtain a similarity representation that integrates local structural information.
6. The method according to claim 1, characterized in that The target text is hierarchically screened by using the multi-level screening model based on the cross-modal similarity matrix to obtain a target candidate text set, including: Performing bidirectional normalization processing on the cross-modal similarity matrix, wherein the bidirectional normalization processing includes normalizing the row vectors and applying a normalization function to the column vectors to obtain a preliminary normalized similarity matrix; calculating K average similarities between each text and its neighbors as local similarity features, and performing weighted adjustment on the preliminary normalized similarity matrix based on the local similarity features to obtain a similarity matrix that strengthens the local structure; The similarity matrix of the enhanced local structure is mapped to a low-dimensional space and subjected to product quantization discretization processing to construct a multi-layer index tree structure; an adaptive search radius is set based on the multi-layer index tree structure, an approximate nearest neighbor search is performed within the adaptive search radius, and search path information is recorded to obtain a preliminary screening candidate text; features are extracted from the preliminary screening candidate text at multiple semantic levels and feature similarity scores are calculated, and the feature similarity scores are merged with the search path information to obtain a multi-dimensional evaluation index; a weighted ranking model is constructed based on the multi-dimensional evaluation index, and the preliminary screening candidate text is re-ranked to obtain a re-ranked candidate text; The similarity distribution characteristics of the reordered candidate texts are counted, and the screening threshold is adaptively adjusted based on the similarity distribution characteristics; the similarity between the reordered candidate texts is calculated, and the diversity constraints are performed on the reordered candidate texts based on the maximum marginal correlation algorithm to obtain an intermediate candidate text set; the confidence score of each candidate text in the intermediate candidate text set is calculated based on the feature consistency of multiple modalities; the candidate texts with confidence scores higher than a preset threshold are preferentially retained to obtain a target candidate text set.
7. The method according to claim 1, characterized in that A multi-task learning framework is used to perform text clustering and relationship prediction on the text relationship graph to generate a structured query report, including: Perform word-level feature extraction, sentence-level feature extraction and document-level feature extraction on the text node to obtain word-level features, sentence-level features and document-level features; perform weighted fusion on the word-level features, sentence-level features and document-level features to obtain text node features; based on the text node features, perform feature extraction on the relationship edge using bidirectional feature propagation to obtain initial relationship edge features; calculate information aggregation weights based on the initial relationship edge features, perform weighted aggregation on the initial relationship edge features, and obtain optimized relationship edge features; The text node features and the optimized relationship edge features are input into a multi-task learning framework, which includes a text clustering subtask and a relationship prediction subtask; for the text clustering subtask, a node similarity matrix is calculated based on the text node features; for the relationship prediction subtask, an edge feature vector is calculated based on the optimized relationship edge features; based on the node similarity matrix, a probability distribution from a sample to a cluster center is calculated, and a distance metric function is constructed in combination with the graph structure information of the text relationship graph; the text nodes are clustered according to the distance metric function and the probability distribution to obtain a text clustering result; Based on the edge feature vector, a relation feature representation is constructed from the semantic dimension, the structural dimension and the temporal dimension, and the relation feature representation is fused using a multi-head attention mechanism; a relation prediction result is generated according to the fused relation feature representation; the text clustering result is organized into a topic hierarchy, and the relation prediction result is used as an entity relationship structure; and a structured query report is generated based on the topic hierarchy and the entity relationship structure.
8. A cross-modal similarity text mining data query system, used to implement the method described in any one of claims 1 to 7, characterized in that: include: The first unit is used to obtain the target text in the query text and the preset text library, wherein the target text includes image feature data and time series relationship data; construct a multi-level feature extraction network, wherein the multi-level feature extraction network includes a text feature extraction module, an image feature extraction module and a time series feature extraction module; Input the text to be queried into the text feature extraction module to obtain a text multidimensional feature vector, input the image feature data into the image feature extraction module to obtain an image multidimensional feature vector, and input the time series relationship data into the time series feature extraction module to obtain a time series multidimensional feature vector; The second unit is used to construct an adaptive cross-modal contrastive learning network, which includes a feature fusion subnetwork and a dynamic weight adjustment module; the text multidimensional feature vector, the image multidimensional feature vector and the time series multidimensional feature vector are input into the feature fusion subnetwork, and different modal features are fused and mapped through the dynamic weight adjustment module to obtain adaptive fusion features; and the contrastive learning loss function is used to optimize based on the adaptive fusion features to construct a cross-modal similarity matrix; The third unit is used to construct a multi-level screening model, which includes a coarse-grained filtering layer and a fine-grained evaluation layer; using the multi-level screening model, the target text is hierarchically screened based on the cross-modal similarity matrix to obtain a target candidate text set; using a graph neural network to build a text relationship reasoning module, and establishing a text relationship graph based on the target candidate text set; using a multi-task learning framework to perform text clustering and relationship prediction on the text relationship graph to generate a structured query report.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Multi-modal named entity recognition method based on dependency syntax and graph neural network
CN118673922A
Cross-modal data fusion user psychological portrait acquisition method and system
CN118885766A