Knowledge Graph-based Retrieval Method, Apparatus, Computer Device, and Storage Medium
Through the cross-modal information retrieval method based on knowledge graph, the problem of insufficient alignment of image and text modal semantics in the community property management system is solved, and higher information retrieval accuracy and model generalization ability are achieved.
Patent Information
- Application Number
- CN202510671809.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-05-23
AI Technical Summary
In the existing community property management system, image recognition methods cannot process intention information described in natural language, and traditional text retrieval technology is difficult to perceive image content, resulting in low accuracy of information retrieval.
A cross-modal information retrieval method based on knowledge graph is adopted. By obtaining the knowledge graph, the node relationship vector is output using the relationship learning module, external knowledge representation results are generated, and feature fusion is used for feature fusion. The two-way semantic alignment of the graph and text is performed based on the attention mechanism, and the joint feature vector is generated, and the matching score of the degree of semantic correlation is finally calculated.
It improves the accuracy of information retrieval in the community property management system, enhances the semantic alignment between the image and text modality, and improves the generalization ability of the model and the accuracy of information retrieval.
Smart Images

Figure CN120196767B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a retrieval method, device, computer device and storage medium based on a knowledge graph. Background Art
[0002] With the development of intelligent property management, community property management systems have gradually introduced means such as video surveillance, image acquisition, and text repair reporting in order to improve service efficiency and management level. However, existing image recognition methods are mostly based on single-modal computer vision models and cannot process the intention information described by users in natural language; traditional text retrieval technologies are also difficult to perceive the complex scene content contained in images, and thus cannot achieve deep semantic reasoning based on multi-modal input.
[0003] In recent years, although multi-modal learning technologies have made some progress, their actual application in community property management systems still faces problems such as insufficient alignment of graphic and text semantics, weak dependence on external knowledge, and insufficient model generalization ability, resulting in low accuracy of information retrieval in community property management systems. Summary of the Invention
[0004] Embodiments of the present invention provide a retrieval method, device, computer device and storage medium based on a knowledge graph, aiming to solve the problem of low accuracy of information retrieval in existing community property management systems.
[0005] In a first aspect, embodiments of the present invention provide a cross-modal information retrieval method based on a knowledge graph, including:
[0006] Obtain a knowledge graph, use a preset relationship learning module to output node relationship vectors, and retrieve nodes related to the user's query content from the knowledge graph, and generate an external knowledge representation result;
[0007] Use a correlation matrix to perform feature fusion on the user's query content and the external knowledge representation result to obtain a fused feature vector;
[0008] Based on an attention mechanism, perform two-way graphic and text semantic alignment on the fused feature vector to generate a joint feature vector for cross-modal retrieval;
[0009] Based on the joint feature vector, perform a matching calculation to obtain a matching score for characterizing the semantic correlation degree between the user's query content and candidate graphic and text data for cross-modal retrieval.
[0010] In a second aspect, embodiments of the present invention provide a retrieval device based on a knowledge graph, including:
[0011] A node retrieval unit, configured to obtain a knowledge graph, output node relationship vectors by using a preset relationship learning module, retrieve nodes related to the user query content from the knowledge graph, and generate an external knowledge representation result;
[0012] A feature fusion unit, configured to perform feature fusion on the user query content and the external knowledge representation result by using a correlation matrix to obtain a fused feature vector;
[0013] A semantic alignment unit, configured to perform two-way text-image semantic alignment on the fused feature vector based on an attention mechanism to generate a joint feature vector for cross-modal retrieval;
[0014] A feature matching unit, configured to perform matching calculations based on the joint feature vector to obtain a matching score for characterizing the semantic correlation degree between the user query content and the candidate text-image data for cross-modal retrieval.
[0015] In a third aspect, an embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the cross-modal information retrieval method based on a knowledge graph in the first aspect is implemented.
[0016] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, where a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the cross-modal information retrieval method based on a knowledge graph in the first aspect is implemented.
[0017] An embodiment of the present invention provides a cross-modal information retrieval method based on a knowledge graph, including obtaining a knowledge graph, outputting node relationship vectors by using a preset relationship learning module, retrieving nodes related to the user query content from the knowledge graph, and generating an external knowledge representation result; performing feature fusion on the user query content and the external knowledge representation result by using a correlation matrix to obtain a fused feature vector; performing two-way text-image semantic alignment on the fused feature vector based on an attention mechanism to generate a joint feature vector for cross-modal retrieval; performing matching calculations based on the joint feature vector to obtain a matching score for characterizing the semantic correlation degree between the user query content and the candidate text-image data. The present invention fuses the user query content and the generated external knowledge representation result through a correlation matrix, then aligns the text-image semantics through an attention mechanism, generates a joint feature vector, and calculates a matching score, thereby realizing cross-modal retrieval. In this way, the information retrieval accuracy of the community property management system is greatly improved.
[0018] An embodiment of the present invention further provides a cross-modal information retrieval device, a computer device, and a storage medium based on a knowledge graph, which also have the above beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0020] Figure 1 A schematic diagram of a process flow of a cross-modal information retrieval method based on a knowledge graph provided by an embodiment of the present invention;
[0021] Figure 2 A schematic block diagram of a cross-modal information retrieval device based on a knowledge graph provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0023] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0024] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0025] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0026] See below Figure 1 , Figure 1 A flow chart of a cross-modal information retrieval method based on a knowledge graph provided in an embodiment of the present invention specifically includes: steps S101 to S104.
[0027] S101. Obtain a knowledge graph, use a preset relationship learning module to output node relationship vectors, retrieve nodes related to the user's query content from the knowledge graph, and generate an external knowledge representation result;
[0028] S102. Use a correlation matrix to perform feature fusion on the user's query content and the external knowledge representation result to obtain a fused feature vector;
[0029] S103. Based on an attention mechanism, perform image-text bidirectional semantic alignment on the fused feature vector to generate a joint feature vector for cross-modal retrieval;
[0030] S104. Perform a matching calculation based on the joint feature vector to obtain a matching score for characterizing the semantic correlation degree between the user's query content and candidate image-text data for cross-modal retrieval.
[0031] In one embodiment, before the step S101, the following steps are further included:
[0032] Construct an image-text pairing data set in the target scenario; wherein, the image-text pairing data set includes a plurality of image training data with scene semantics and text description training data corresponding to the image training data;
[0033] Extract features from the image training data and the text description training data respectively to obtain an image training feature vector and a text training feature vector;
[0034] Extract nodes and semantic representations related to the target scenario from the knowledge graph to obtain a node semantic set;
[0035] Based on a node fusion mechanism, fuse the image training feature vector, the text training feature vector, and the node semantic set to generate a fused node vector for characterizing multi-modal information;
[0036] Input the fused node vector into the original relationship learning module to construct the node relationship vectors between multiple nodes;
[0037] Construct a loss function, and use the loss function to train the prediction ability of the original relationship learning module to obtain the preset relationship learning module.
[0038] In this embodiment, to construct the data resources required for training a cross-modal knowledge retrieval model applicable to the community property scenario, it is necessary to first execute the construction process of the image-text pairing data set. The construction process includes three stages: acquisition of image training data, generation of text description training data, and manual verification and revision. Finally, an image-text pairing data set is generated as the basic support for model training and feature learning.
[0039] In the acquisition stage of image training data, image training data is collected based on cameras installed in different functional areas within the property community. To ensure the comprehensiveness and diversity of the data, multiple target scene locations are covered, including public areas, around the swimming pool, children's playground, garage entrances and exits, etc. Further, environmental changes under different weather conditions (such as sunny days, cloudy and rainy days, nights, etc.) and time periods (such as daytime, evening, late at night, etc.) are considered to enhance the adaptability and generalization ability of the data. During the process of acquiring image training data, a frame extraction strategy is set according to the surveillance video sources within the community, and image frames are extracted at a predetermined time interval. Video segments that trigger specific events (such as target movement, abnormal behavior, etc.) are preferentially collected, so as to ensure the full capture and coverage of community event-related targets.
[0040] For the collected image training data, according to 22 types of target detection tasks in community property management, the image training data is initially labeled and scene-classified. The specific classification tasks include: garbage in public areas, employees leaving their posts in the command center, trash cans overflowing, children playing by the pool, congestion at the entrance and exit, people lingering at the garage entrance, motor vehicles staying put, electric vehicles parked randomly, intrusion into the community perimeter, motor vehicles parked illegally, trash beside the trash can, electric vehicles entering the lobby, entering the pool during non-business hours, electric vehicles entering the elevator, the pool safety officer leaving their post, people falling, pets walking alone, children walking alone, face recognition, detection of open flames, playing with mobile phones on duty, and sundries piled up in the fire passage.
[0041] In the generation stage of text description training data, based on the actual business requirements of community property management and the goal of semantic understanding of scene events, multi-angle semantic question-and-answer tasks are constructed for each image. For example, 5 standard questions can be generated for each image, including dimensions such as environmental state description, personnel behavior discrimination, traffic condition assessment, implicit risk inference, and scene applicability judgment. Example questions are as follows: (1) What is the environmental situation of this image? Please describe it in detail; (2) Are there any people in the image? Is there any obvious social interaction between them? (3) Is there any traffic flow or potential traffic hazards in the picture? (4) Does the phenomenon reflected in this picture have potential risks in the context of community management? (5) Which specific scene in property management does this image apply to? Does it have a warning or reference significance? Based on the above set of questions, using a pre-trained multi-modal large language model (such as Qwen2-VL-7b), each image and the corresponding questions are input to generate preliminary text answers, forming the machine understanding results of the image semantic information. This process can automatically complete the preliminary semantic annotation of a large number of images, improving the efficiency of dataset construction.
[0042] In the manual verification and revision stage, to ensure the accuracy and usability of the final text description, after the large model generation is completed, manual review and revision can be arranged for each of the 5 text answers corresponding to each image item by item. The content of manual verification includes the coherence of semantic logic, the accuracy of event and object descriptions, the standardization of grammatical structures, and the supplementation of background details such as time and space. During the verification process, if it is found that the content generated by the model is ambiguous, incomplete, or does not match the image, the manual will make modifications or rewrites based on the facts to ensure that the text and image semantics strictly correspond and are trainable. Finally, each image corresponds to 5 text descriptions revised manually, forming the complete semantic annotation of the image.
[0043] Based on the above process, the finally formed image-text paired dataset is as follows: The image part includes T static image frames collected from real community monitoring devices, covering 22 types of typical scenario events; the text part correspondingly generates 5T semantic description statements, covering information such as event status, environmental information, object attributes, and risk metaphors, fully reflecting the multi-modal semantic complementarity and task-driven relevance.
[0044] Furthermore, feature extraction is performed on the image training data and text description data respectively. For the image modality, a convolutional neural network based on ResNet or a Faster-RCNN structure can be used to extract region-level semantic embedding features to obtain the image training feature vector; for the text modality, each description statement is encoded at the Token level through a pre-trained language model (such as BERT) to obtain the text training feature vector. Then, semantic node and structural relationship information related to the target scenario are extracted from the knowledge graph to form a node semantic set.
[0045] To solve the problems of semantic heterogeneity and cross-modal alignment difficulties between image and text modalities in the community property scenario, this embodiment fuses image information, text information, and knowledge graph structured semantics through a unified modeling method, thereby providing a more generalizable and reasoning-capable feature basis for subsequent cross-modal retrieval. Specifically, this unified representation method uses the Visualsem knowledge graph as the semantic enhancement carrier, and comprehensively fuses and vectorizes the text annotations, image instances, and semantic nodes contained in the graph. Visualsem is an open knowledge graph for multi-language and multi-modal information processing, covering approximately 1.3 million lexical annotations and 938,000 image resources, forming a multi-modal triple set consisting of 89,896 unique semantic nodes and approximately 1.48 million structural relationships between them , representing node connected to node through relationship .
[0046] Since Visualsem covers a wide range of content and contains a large amount of general knowledge information related to non-community properties, in order to improve the scenario adaptability and accuracy of the model, the graph needs to be pre-screened. The specific processing flow includes: first constructing a vocabulary list in the field of property management, and then filtering out a set of nodes related to the property management scenario from the original graph through word vector semantic matching and image content analysis. These nodes are recorded as M target nodes, as well as the common nodes associated with these M nodes. Comment statements and Then, for each target node, its associated text annotations and description images are collected, and after manual confirmation of their semantic validity, they are retained as the original materials for cross-modal fusion modeling. Comment statements and Then, feature extraction is performed on the graphic data of each node in Visualsem to construct a unified cross-modal representation for subsequent model learning. All nodes in the entire knowledge graph are interconnected in a similar manner. Taking the node "Air quality index" in the Visualsem knowledge graph as an example, for the M target nodes selected in the previous step, the specific data processing method is as follows:
[0047] First, for any of the M nodes nodes, and use LSTM and ResNet networks to analyze the nodes related to the glossary and Extract features from each image and obtain the corresponding high-level semantic representation vectors of the original text and image. and ,in Indicates The node related Notes, Indicates The node related images; then, for all text vectors and image vectors , and the average aggregation method is used to obtain and , to indicate the The final text features and image features of each node. The formula for average aggregation of text and image is as follows:
[0048] .
[0049] Furthermore, a node fusion mechanism is constructed to unify the modeling of the image training feature set, the text training feature set, and the node semantic set. It includes: using LSTM and ResNet to extract the original semantic vectors for the vocabulary annotations and images related to each node respectively, and forming the node text features and image features through average aggregation; then, through the GraphSAGE graph neural network structure, combined with the adjacency structure, performing structure-aware feature aggregation on the node's own information to obtain a node vector containing multi-modal features and structural semantics; subsequently, through the node gating unit, the image training feature set, the text training feature set, and the node semantic vector are fused into a fused node vector.
[0050] Specifically, in the node feature aggregation stage, an improved graph neural network structure - GraphSAGE (Graph Sample and Aggregate) can be used to perform feature aggregation and update on the semantic nodes in the Visualsem knowledge graph. Different from the traditional GCN full-graph convolutional calculation, GraphSAGE takes each node as the center during the calculation, samples its neighbor nodes to construct a subgraph, and generates a context-aware expression of the target node through local aggregation operations, effectively improving the computational efficiency on large-scale graphs and the dynamics of node expressions. For any target node A in the Visualsem knowledge graph, an Aggregator function is used to learn its local neighbor aggregation feature information. Each Aggregator function aggregates information from different hops (here taking 1) or different search depths of the node, and its role is to convert a set of vectors into a vector. The specific formula of the Aggregator function is as follows:
[0051] ;
[0052] Among them, represents the total number of 1-hop nodes of the target node, represents the weight matrix, represents the initial feature vectors of the 1-hop nodes , is the activation function, represents the target node feature vector finally generated after the neighbor aggregation operation.
[0053] After completing the structure aggregation, the target node A not only contains its own feature information but also integrates the context semantic features of adjacent nodes, and has a preliminary graph structure perception ability.
[0054] Further, the fused node vectors are input into the original relationship learning module to construct the node relationship vectors between multiple said nodes. The original relationship learning module constructs a set of node pairs with actual semantic dependencies based on a predefined head and tail node screening strategy, and models the potential semantic relationships between each pair of node pairs. The original relationship learning module uses a unified scale feature mapping and ReLU activation structure to generate relationship vectors. The N generated node pair relationship vectors form a node relationship matrix, which has the ability to recognize 13 predefined semantic relationships.
[0055] The main processes of the original relationship learning module include three stages: feature mapping and scale unification, head and tail node screening and pairing, and relationship vector fusion and non-linear activation. The specific implementation is as follows:
[0056] In the stage of feature mapping and scale unification, let the node features with multimodal fusion information have a dimension of (1, ), and multiply all M with a weight matrix W of dimension ( ). The purpose is to uniformly map node features from different sources into a standardized scale space, ensuring that information from images and texts can be effectively fused and processed in subsequent relationship learning processes, thereby eliminating possible differences between the original feature spaces. After mapping, the dimension of the M node features is (1, ).
[0057] In the stage of head and tail node screening and matching, after the node features are uniformly mapped to the same scale space, it is necessary to screen and pair the head and tail node concepts. The screened node pairs are those that have clear semantic correspondence relationships in the VisualSem knowledge graph, rather than being randomly paired. In this way, the interference of irrelevant nodes is avoided, ensuring that the relationship learning module can focus on the truly existing semantic relationships, rather than irrelevant hypothetical connections. Based on these learned correct relationships, the knowledge understanding breadth of the subsequent model in retrieval and reasoning tasks can be broadened, and by extending and associating relationships with the user's input, deeper possible intentions can be captured. After screening and pairing, there are a total of N pairs of head and tail node pairs (2N M). The 13 inherent relationship types between nodes that already exist in the VisualSem knowledge graph are: is-a; has-part; related-to; used-for; used-to; subject-of; receives-action; made-of; has-property; gloss-related; synonym; part-of and located-at.
[0058] In the relationship vector fusion and nonlinear activation stage, the selected N head and tail nodes are linearly transformed and mapped into The high-dimensional vector of dimension is generated, and the two paired vectors are added and averaged. Through this operation, the information of the head and tail nodes can be effectively integrated to form a comprehensive feature vector to represent the relationship between the two. Subsequently, the ReLU function is used to perform nonlinear activation on the comprehensive feature vector. The ReLU function can further enhance the relationship representation ability of the model by suppressing negative values and keeping positive values unchanged, making the relationship vectors between different nodes more distinguishable and discriminative. Finally, N relationship vectors between the head and tail nodes predicted by the model are obtained, which will be used for subsequent loss function calculations and gradient updates. The purpose is to cultivate the model's ability to accurately associate external inputs with Visualsem's internal knowledge through continuous training.
[0059] The overall process is expressed by the following formula:
[0060] ;
[0061] in, and Represents the head and tail node features of any pair of M nodes, W is the weight matrix, represents a linear transformation, Indicates that the dimension of the final output is (N, ) is a relationship matrix (i.e., N node relationship vectors).
[0062] Furthermore, to optimize the semantic modeling performance of the original relational learning module, a multi-dimensional loss function is constructed for training optimization. The training goal of the original relational learning module is to cultivate the model's external generalization capability under the constraints of known relational groundtruth, so that when external information is input to the model, it can still accurately and efficiently connect relationships with the model's internal knowledge.
[0063] First, we need to map the 13 inherent relationship types in Visualsem into vectors as the relationship groundtruth in the loss function, use one-hot encoding to uniquely represent each relationship (13-dimensional vector), and then map it into a trainable MLP. The specific formula for embedding is as follows:
[0064] ;
[0065] in, Indicates the One-Hot encoding of a certain relationship, Is a trainable weight matrix responsible for learning the embedding of the One-Hot relationship, is the final relation embedding, and is also the relation groundtruth in the loss function below. The following formula is the correlation loss , contrast loss and triplet loss The specific content forms of the three functions are:
[0066] ;
[0067] in, Represents the relationship vector between a pair of nodes predicted by the model, Represents the true relationship vector (Groundtruth) between the node pairs obtained by One-Hot encoding, Indicates any non type of relation vector, represents a constant, represents the sigmoid activation function, ( represents element-wise multiplication) is a scoring function for a triple.
[0068] Among the three loss functions mentioned above, It uses the conventional cross entropy loss function to narrow the distance between the predicted relationship and the true relationship vector; Based on the cosine similarity consideration, the cosine similarity between the predicted relationship vector and the corresponding true relationship vector is made as high as possible (close to 1), while the cosine similarity with other irrelevant true relationship vectors is made as low as possible, and the gap is made as close to the edge value as possible. ; The role of is to force the model to obtain the same tail entity from the predicted relationship of the head node as that obtained from the real relationship.
[0069] By minimizing the above three loss functions, the model of the original relationship learning module will dynamically adjust the weights and feature representations during the training process to achieve the best ability to predict the true relationship between nodes.
[0070] Furthermore, the N head-and-tail node relationship vectors (i.e., the predicted relationship matrix) predicted by the original relationship learning module are multiplied with the 13 categories of true relationship vectors generated by one-hot encoding (the resulting relationship matrix). This results in an N×13 similarity matrix, where each row represents the similarity score between a predicted relationship and one of the 13 true relationship categories. Based on the idea that more similar vectors have larger dot products, the similarity scores in each row are softmax-normalized, where the subscript of the maximum value corresponds to the specific category of the predicted relationship, ranging from 1 to 13.
[0071] Furthermore, to verify the effectiveness and robustness of the original relationship learning module in node relationship modeling tasks, after completing the original relationship learning module construction and multiple loss function optimization, two evaluation metrics commonly used in knowledge graph link prediction can be set: Mean Reciprocal Rank (MRR) and Hits@n. These metrics intuitively reflect the model's accuracy and ranking ability in predicting head-tail entity relationships. Specifically, the MRR metric measures the average degree to which the model ranks true head-tail node relationships at the top of all prediction results. Its specific calculation method is as follows:
[0072] ;
[0073] Among them, N represents the total number of predicted relationships between head and tail entity pairs, It indicates the ranking of the true head-tail relationship in the candidate list output by the model in the i-th sample. The value range of MRR is (0,1]. The larger the value, the more likely the model is to rank the correct relationship prediction in the front position, and has a stronger relationship recognition and ranking ability.
[0074] The Hits@n indicator focuses on whether the model can predict the top n results (n 13) to accurately identify the true head-tail relationship, we can select n=1, n=3, or n=10 to observe the prediction accuracy at different granularities. The higher the Hits@n, the more the model is able to identify the correct relationship within a limited candidate range. The specific formula for Hits@n is as follows:
[0075] ;
[0076] Through the above two indicators, the performance of the model in the task of predicting semantic relationships between nodes can be systematically evaluated from the two dimensions of accuracy and sorting ability. With the help of continuous correction and supervision of existing relationships in the knowledge graph, the model will gradually learn how to rely on the semantic features of the two nodes themselves to infer the most appropriate relationship type between them. This ability will also be transferred to nodes outside the graph to achieve cross-domain semantic connection and information completion, laying the foundation for subsequent diverse intelligent retrieval and semantic extension understanding. At this point, the original relationship learning module that has been trained and optimized is obtained, which is the preset relationship learning module.
[0077] In one embodiment, the node fusion mechanism is used to fuse the image training feature vector, the text training feature vector, and the node semantic set to generate a fused node vector for representing multimodal information, including:
[0078] Inputting the image training feature vector and the text training feature vector into corresponding projection matrices according to the node semantic set to generate image-side feature representation and text-side feature representation;
[0079] Inputting the image-side feature representation and the text-side feature representation into corresponding multi-layer perceptrons respectively to generate an image-side gating factor and a text-side gating factor;
[0080] Adjusting the image-side feature representation using the image-side gating factor, and adjusting the text-side feature representation using the text-side gating factor, to generate image-gated features and text-gated features, respectively;
[0081] The image gated features and the text gated features are spliced and input into a hybrid projection matrix to generate the fusion node vector.
[0082] In this embodiment, the node fusion mechanism includes a node gating unit, which integrates the text training feature vector and image training feature vector related to the node into the features of the node semantic set itself, thereby achieving a unified representation of cross-modal data in the knowledge graph. The specific formula of the node gating unit is as follows:
[0083] ;
[0084] in, represents the node gating function, Indicates concatenation of vectors. and They are respectively about text training feature vectors and image training feature vector The projection matrix, and Multilayer perceptrons representing text-side feature representation and image-side feature representation, respectively. and They are the text side gating factor and the image side gating factor, and They are node features that integrate text training feature vector information and image training feature vector information. Finally, and After stitching, multiply it by the mixed projection matrix , obtain node features with multimodal data blending information (i.e. fusion node vector).
[0085] In step S101, first, obtain a pre-constructed knowledge graph, which includes node entities, attribute information, and various semantic relationships related to the community property scenario. The knowledge graph is preferably a multi-modal knowledge graph that integrates image and text modal information, such as the Visualsem knowledge graph, whose nodes have multilingual annotations, illustrative images, and semantic relationship definitions. Use a preset relationship learning module to output node relationship vectors, perform semantic parsing on the query content input by the user, and retrieve target nodes related thereto in the knowledge graph. The preset relationship learning module introduces a graph neural network (such as GraphSAGE) to fuse the text features, image features, and structural features of the nodes, and then generates an external knowledge representation result expressing the query semantics for supporting subsequent cross-modal semantic modeling.
[0086] In one embodiment, step S101 includes:
[0087] When the user query content is an image, input the image into a preset visual feature extraction model to extract image feature vectors of multiple regions in the image. At the same time, input the image into the knowledge graph to retrieve multiple first nodes related to the image and obtain corresponding first node indexes;
[0088] Extract the fused node vectors corresponding to the first nodes based on the first node indexes;
[0089] Input the image feature vectors and the fused node vectors into the preset relationship learning module for relationship modeling, and output the first node relationship vector between the image and the first nodes;
[0090] When the user query content is text, input the text into a preset text feature extraction model to extract text feature vectors at the Token level in the text. At the same time, input the text into the knowledge graph to retrieve multiple second nodes related to the text and obtain corresponding second node indexes;
[0091] Extract the fused node vectors corresponding to the second nodes based on the second node indexes;
[0092] Input the text feature vectors and the fused node vectors into the preset relationship learning module for relationship modeling, and output the second node relationship vector between the text and the second nodes;
[0093] Integrate the first node relationship vector and the second node relationship vector to obtain the external knowledge representation result.
[0094] In this embodiment, image paths and text paths are processed separately based on the user query content to achieve external knowledge enhancement and semantic relationship modeling under the dual-modality conditions of image and text. In the case that the user query content is an image, the following operations are performed:
[0095] The image is input into a pre-defined visual feature extraction model, preferably a Faster-RCNN architecture, comprising a Region Proposal Network (RPN) and ROI feature extraction modules. Several candidate regions of interest (ROIs) are generated from the image, and a semantic embedding vector is calculated for each region. This yields a set of region-level image feature vectors, each corresponding to a local semantic region node (i.e., the first node) in the image, capturing fine-grained spatial semantic information. The image is then input into a knowledge graph, a multimodal semantic graph called Visualsem, which uses its accompanying image-text retrieval capabilities to semantically associate nodes with the knowledge graph. A CLIP model (such as RN50x16) is used to semantically encode the input image and selected property scene image nodes from the knowledge graph. By calculating the cosine similarity of the feature vectors between the images, the graph nodes most relevant to the input image are retrieved, resulting in the corresponding first node indexes. Based on the first node indexes, the corresponding fused node vectors are extracted from a pre-built set of fused node vectors. The image feature vectors and fused node vectors are then input into a pre-defined relational learning module. The preset relationship learning module learns the explicit semantic relationships that may exist between image features and knowledge nodes (such as "has-part", "used-for", "located-at", etc.) based on the aforementioned trained structure, and outputs the first node relationship vector between the image and each retrieval node as the external knowledge representation of the image modality side.
[0096] If the user query content is text, perform the following operations:
[0097] The text is input into a preset text feature extraction model, preferably using the pre-trained language model BERT, to extract token-level semantic representations of the input text, resulting in a set of text feature vectors, each corresponding to a semantic unit (e.g., a word or phrase) that characterizes the semantic granularity of the text. The text is then input into a knowledge graph, where a knowledge enhancement operation similar to the image path is performed. The specific steps include using the Sentence-BERT model to semantically encode the input text and a collection of text annotations related to the property sector in the knowledge graph. Next, the cosine similarity between the text semantics is calculated to select the annotations (i.e., second nodes) with the highest relevance to the input text. Finally, the corresponding node indexes are determined based on the annotation mapping relationship to obtain a set of second node indexes. Based on the second node indexes, the corresponding fused node vectors are extracted. Subsequently, the text's token-level feature vectors and the fused node vectors are input into a preset relationship learning module. The preset relationship learning module learns the semantic relationship connection between the input text content and the graph nodes, outputs the second node relationship vector between the text modality and the second node, and represents the possible semantic relationship between the text semantic unit and the graph knowledge entity, such as "subject-of", "has-property", "related-to", etc.
[0098] Finally, the first node relationship vector obtained from the image path and the second node relationship vector obtained from the text path are combined and integrated to generate the external knowledge representation corresponding to the user's query content (this can be two: one for the retrieved node vector and one for the relationship vector between the retrieved node vector and the retrieved node vector). Through the above processing steps, the knowledge structure and semantic context associated with the input content can be fully mined and modeled in the pre-processing stage of image-text query, effectively improving the understanding and reasoning ability of fuzzy expressions, implicit relationships, and contextual semantics in the subsequent image-text matching process. This is particularly suitable for diverse and complex community property intelligent service scenarios.
[0099] In step S102, based on the correlation matrix construction mechanism (the core goal is to embed the input stream (region / token embedding), the knowledge graph retrieval node (retrieved KG node), and the relationship between the two (relation embedding) into a unified representation space), the user query content is fused with the generated external knowledge representation results to obtain a fused multimodal feature vector (i.e., a fused feature vector). This process structurally balances the diversity of semantic expression with the ability to perceive external knowledge, improving the model's understanding of complex scenarios.
[0100] In one embodiment, step S102 includes:
[0101] Inputting multiple vectors in the user query content as multiple heterogeneous entities; wherein the vectors include image feature vectors, text feature vectors, fusion node vectors and node relationship vectors;
[0102] Embedding the vectors into a unified representation space and calculating semantic relevance scores between heterogeneous entities to construct a relevance matrix having relationships between multiple entities;
[0103] Taking the image feature vector or the text feature vector as a query vector based on the correlation matrix;
[0104] Mapping the fused node vector and the node relationship vector to a preset weight matrix based on the correlation matrix to obtain corresponding key vectors and value vectors;
[0105] An attention score is calculated based on the query vector, the key vector, and the value vector to generate the fused feature vector.
[0106] In this embodiment, multiple semantic representation vectors obtained from user query content are organized into multiple heterogeneous entities as input. This input set includes: multiple region-level image feature vectors extracted in the image modality (such as ROI embeddings extracted by Faster-RCNN), multiple token-level text feature vectors extracted in the text modality (such as TokenEmbedding output by BERT), fused node vectors obtained through knowledge graph retrieval (containing image and text semantics and structural information), and node relationship vectors obtained through a preset relationship learning module (representing the semantic relationship between the input features and the knowledge graph). All of these vectors participate in the subsequent unified modeling process as independent semantic entities.
[0107] Furthermore, the system embeds these vectors into a unified representation space. In order to eliminate the inconsistency between the feature dimensions and semantic scales between modalities, a unified embedding transformation matrix is preset, and linear mapping is performed on various types of vectors so that they are projected into the same embedding space. Then, based on the above embedding representation, the semantic relevance scores between any entities are calculated, and a correlation matrix between entities is constructed. For all input entities, the values in the matrix are filled by calculating the similarity between each pair (such as dot product or cosine similarity), and finally a correlation matrix is formed. After the correlation matrix is constructed, a multimodal attention interaction fusion operation is performed based on the correlation matrix. Based on the attention mechanism, the fusion of external enhanced knowledge and the original input features is divided into image stream and text stream. Here, the processing method of the image stream is taken as an example. The processing method of the text stream is exactly the same as that of the image stream, as shown below:
[0108] The size of the correlation matrix of the image stream is × ,in , let the extracted data from CNN Backbone be The region-level image features are expressed as , ,…… , retrieved from the knowledge graph The node features are expressed as , ,…… , both of which are predicted by the preset relationship learning module The relationship features are expressed as , ,…… .
[0109] Based on the attention mechanism, the query, key, and value are constructed separately, referring to the following formula:
[0110] ;
[0111] ;
[0112] ;
[0113] Among them, query Only the input image Region-level features and weight matrices Multiply by , that is, the query vector only focuses on the original input features, and for the key Sum In addition to focusing on the original features, Node features and relationship features (total features) are weighted (with and multiply).
[0114] Furthermore, the information fusion process based on the attention mechanism is carried out according to the following formula:
[0115] ;
[0116] ;
[0117] The final output … It is a region-level image feature (i.e., fused feature vector) that integrates external enhanced information (i.e., external knowledge representation result). Compared with the original input, it not only retains its own representation ability, but also integrates richer external semantic-related knowledge, which provides more comprehensive support and guidance for the complex and diverse retrieval questions and answers in subsequent community property scenarios.
[0118] The relevance matrix of a text stream is of size q×q, where , let the extracted from Bert The token-level features are expressed as , ,…… , after similar operations to the image stream, the final output feature is represented as … .
[0119] In step S103, a bidirectional image-text semantic alignment mechanism is further implemented, leveraging the co-attention architecture to perform image-text feature interaction and joint encoding on the fused feature vector. This mechanism uses image modality features as the query and text modality features as the key and value, respectively, and employs a reverse alignment process to achieve semantic guidance and feature perception between image and text features, enhancing the semantic alignment between image and text. Through the multi-layered interaction modules, a joint feature vector is ultimately generated that represents the semantic consistency between the query content and the candidate image and text samples, providing high-level semantic support for subsequent matching calculations.
[0120] In one embodiment, step S103 includes the following steps:
[0121] Step 1: Based on the attention mechanism, the fused image feature vector is used as the first query vector, and the fused text feature vector is used as the first key vector and the first value vector, and a first attention calculation is performed to obtain a first alignment vector;
[0122] Step 2: Based on the attention mechanism, the fused text feature vector is used as the second query vector, and the fused image feature vector is used as the second key vector and the second value vector, and a second attention calculation is performed to obtain a second alignment vector;
[0123] Step 3: Perform residual connection and regularization processing on the first alignment vector to generate image feature representation;
[0124] Step 4: Perform residual connection and regularization processing on the second alignment vector to generate text feature representation;
[0125] Step 5: Input the image feature representation and text feature representation into a feedforward neural network to generate a joint feature representation after semantic fusion of image and text;
[0126] Step 6: Repeat steps 1 to 5 at least twice to enhance the deep semantic alignment effect between the image and the text;
[0127] Step 7: Integrate the joint feature representation to obtain the joint feature vector.
[0128] In this embodiment, step 1: First, based on the multimodal attention mechanism, the fused image feature vector is used as the first query vector (Query), and the fused text feature vector is used as the first key vector (Key) and the first value vector (Value), to build an image-to-text interaction path. Performing the first attention calculation through this path enables the image features to perceive and capture the semantic information expressed in the text, and then generate the first alignment vector on the image side. In the specific implementation, the attention mechanism adopts a multi-head structure and parallel modeling in different semantic subspaces to enhance semantic coverage and contextual interaction capabilities;
[0129] Step 2: Perform reverse attention path processing, using the fused text feature vector as the second query vector (Query), and the fused image feature vector as the second key vector (Key) and second value vector (Value), to construct the interaction path of the context-oriented graph. By performing the second attention calculation, the text features can be focused on the image area information that matches their semantics, thereby outputting the second alignment vector on the text side;
[0130] Step 3: For the first alignment vector on the image side, a residual connection mechanism is used to perform element-by-element summation of this alignment vector and the original image input features to preserve the low-level representation information in the original image features. Layer normalization is then performed on the concatenated result to generate the image-side feature representation. This processing step improves model stability and prevents feature dissipation, forming a preliminary image representation after image-text interaction.
[0131] Step 4: For the second alignment vector on the text side, perform residual connection and regularization operations, connect the alignment vector with the initial text embedding vector, and then normalize it to obtain a text feature representation that includes visual information perception capabilities;
[0132] Step 5: The image and text feature representations are fed into a feedforward neural network (FFN), which consists of a two-layer fully connected structure with integrated activation functions (such as ReLU) and projection mechanisms to further extract nonlinear semantic combinations between cross-modal features. The feedforward output undergoes another residual connection and normalization process to output a joint feature representation of the current layer's image and text semantic fusion.
[0133] Step 6: To enhance the deep semantic alignment between the image and text modalities and improve the hierarchical nature of contextual understanding, construct steps 1 to 5 above into a complete image-text interaction submodule and stack this module in multiple layers. Preferably, the stacking number is 5, and the above submodule is repeated at least five times, progressively achieving a gradual penetration of image and text semantics layer by layer. The final image features contain text semantics, and the text features are integrated with image context, thus forming a high-level cross-modal joint representation.
[0134] Step 7: Finally, the joint feature representations of the image and text output from each interaction layer are integrated (e.g., through concatenation, weighted summation, or adaptive aggregation strategies) to generate a high-dimensional joint feature vector, which serves as the basic input for cross-modal similarity calculation and retrieval matching in subsequent steps. This joint feature vector not only preserves the original modal information but also incorporates the deeper semantics of the other modality. Through external knowledge graph relationship enhancement and attention mechanisms, it achieves detailed context perception, improving the matching accuracy and generalization ability between image and text.
[0135] Specifically, you can refer to the following formula for calculation: Let the input , , taking visual flow as an example (the same applies to text flow):
[0136] ;
[0137] ;
[0138] ;
[0139] in, 、 、 、 、 and Both represent weight coefficients, and represents the bias coefficient, h represents the number of attention heads, and multi-head attention The output of is represented as x after residual connection and regularization, which is input into a feedforward neural network FFN, and then obtained after residual connection and regularization. , similarly the text stream will get .
[0140] For example, stacking 5 layers of the above interaction modules to gradually enhance the deep fusion capability of image and text features, the final output of the image and text streams are and , then the image features Encoded text semantics, text features The visual context is also included in , forming a joint feature vector.
[0141] In step S104, a matching calculation is performed based on the joint feature vector, and a weighted aggregation strategy is used to globally encode the image and text features, generating a matching score. This matching score is used to measure the semantic relevance between the user's query content and each candidate image and text sample, and thus complete the cross-modal retrieval task. Preferably, a ranking loss function can be introduced in the matching calculation process to enhance the semantic distinction between positive and negative samples, thereby improving the model's efficiency in intelligent question answering, event location, and inspection response in community property application scenarios.
[0142] In one embodiment, step S104 includes:
[0143] Performing linear mapping on the image local features and the text local features respectively to obtain corresponding image mapping features and text mapping features;
[0144] Normalizing the image mapping features and the text mapping features respectively to obtain corresponding image attention weights and text attention weights;
[0145] Multiplying the image attention weight by the image mapping feature to generate an image global feature;
[0146] Multiplying the text attention weight by the text mapping feature to generate a text global feature;
[0147] The image global feature and the text global feature are added to generate the matching score.
[0148] In this embodiment, the matching score may be generated according to the following formula:
[0149] ;
[0150] ;
[0151] ;
[0152] ;
[0153] Will and Through the weight matrix and Perform linear mapping, and then obtain the attention weight factors after softmax and , the attention weight factor is then combined with and Multiplying them together gives the overall image features after aggregating all sub-features. and overall text characteristics ; Then, and After addition, the results are fed into an MLP, and the output is normalized to [0, 1] using a sigmoid activation function to represent the probability of image-text matching (i.e., matching score), where 1 indicates that the image-text matches, and 0 indicates that the image-text does not match.
[0154] Finally, based on this matching score, we create a ranking loss function for the image-text pair. ,in, represents the marginal hyperparameter, which is used to control the minimum score difference between positive and negative samples. Represents the relevant positive sample pairs of images and texts, represents a negative sample pair consisting of text and an irrelevant image, represents a negative sample pair consisting of an image and irrelevant text, This solution represents a difficult negative sample image-text pair mining solution. This loss function can help the model improve its ability to distinguish whether image-text pairs match or not, thereby enhancing image-text alignment and multimodal semantic consistency learning.
[0155] In one embodiment, to achieve efficient deployment and continuous optimization of the model in a real-world community property environment, a model inference and data closure phase can also be set up. This phase focuses on edge model deployment, real-time multimodal input processing, natural language retrieval and inference responses, and a feedback-driven model update mechanism. The goal is to build a closed-loop image-text semantic system that integrates real-time intelligent recognition, multi-role interaction, and self-evolution capabilities.
[0156] First, after model training and compression optimization, it is deployed on edge computing devices (such as property servers, smart camera nodes, and security terminals) within the community property environment to support low-latency, high-efficiency on-site image and text reasoning capabilities. Edge devices continuously collect image data from community camera streams and automatically extract and store image frames locally using a predefined frame extraction strategy (including timed and event-triggered frame extraction). This provides a continuous image input source for the model and builds an image index database containing timestamps, location information, and event tags.
[0157] At the same time, in the daily management of the community, a natural language query interface is supported for multi-role users, such as property owners, property security personnel, front desk staff, etc., who can use voice or text input to ask questions based on actual scenario needs. This natural language input form usually contains multimodal clues such as semantic goals, time constraints, and spatial limitations, such as "Where did I see an elderly person riding an electric bike yesterday evening" or "Which building is near a full trash can?" This type of query request is input as user query content, and image-text semantic parsing and cross-modal retrieval are performed according to the aforementioned process: text feature extraction and knowledge graph enhancement processing are performed on the natural language query to construct a joint feature representation with semantic representation and structural knowledge association; the query feature vector is used to perform cross-modal matching calculations with image features in the image library stored at the edge to obtain a matching score; finally, the images are ranked according to the matching score and the most relevant image results are returned to respond to the user request.
[0158] This image-text matching reasoning process not only integrates the semantics of images and text, but also introduces the background semantics and relational reasoning capabilities provided by the knowledge graph, thereby achieving highly accurate content recall under natural language conditions with ambiguous expressions, uncertain targets, and strong ambiguity. Furthermore, during the operation of the system, by analyzing user search behavior, query styles, and feedback results, the system can dynamically collect data from multiple angles to build a closed-loop mechanism: automatically record the interaction log between each user search request and its returned results, and analyze whether the user accepts, modifies, or retries the search results; if the user manually confirms, annotates, or corrects the returned results, the feedback is automatically added to the model training sample pool as a new supervisory signal. In addition, it can also automatically identify query styles or image modalities that the current model performs poorly in specific types of scenarios, and trigger sample expansion mechanisms based on their failed recall samples, such as supplementing similar image frames or annotating typical text patterns.
[0159] Through the automatic collection and aggregation of the above-mentioned feedback data, training sample batches are periodically constructed in the background, and incremental update training is performed on the server side or distributed model training platform, so that the model can continuously improve its generalization ability, semantic alignment robustness, and retrieval and reasoning accuracy in real property scenarios. Ultimately, a self-circulating mechanism of "real-world problems drive data accumulation - data feeds back to model training - model feedback to solve real-world problems" is achieved. Through the implementation of this step, a data flywheel closed-loop system characterized by event-driven, interactive feedback, and intelligent optimization can be formed in the actual property management environment, effectively improving the practicality, interpretability, and continuous evolution capabilities of the multimodal semantic model, and providing strong technical support for the precise management and efficient response of smart communities.
[0160] In summary, the present invention first constructs a property graphic and text data system with scene diversity and multi-dimensional expression to truly restore the daily management context of the community; secondly, it introduces an external knowledge graph to achieve the bridging and enhancement of graphic and text semantics through unified modeling of heterogeneous information; then, by designing a relational learning module, it deeply models the potential semantic structure between graphics and texts; further, it introduces a multi-head cross-attention mechanism to achieve deep interaction and joint representation between graphics and texts; finally, the model is deployed on an edge computing platform, and a continuous optimization mechanism driven by user interaction and data feedback is established to construct an intelligent semantic closed-loop system with real-time response ability, knowledge transfer ability, and scene adaptation ability.
[0161] The beneficial effects of the present invention include:
[0162] Enhanced semantic understanding: Through knowledge graph-assisted modeling and semantic relation learning, the model's ability to understand complex expressions, ambiguous references, and polysemous contexts in images and texts has been significantly improved;
[0163] Improved graphic and text matching accuracy: Through a bidirectional attention mechanism and semantic alignment design, deep and structured interaction and fusion between graphic and text modalities have been achieved, effectively narrowing the graphic and text semantic gap. Of course, on the basis of maintaining accuracy, the ability to handle diverse expressions, implicit relationship understanding, and graphic and text matching generalization in the property scenario has also been improved;
[0164] Strong robustness of intelligent response: When facing different user roles, language expression methods, query intentions, and changes in scenario conditions, the system can still maintain stable and accurate response effects;
[0165] High efficiency in edge deployment: The model can be deployed on community edge devices after quantization and compression, and has the on-site inference ability of low latency, high efficiency, and scalability, meeting the actual application needs of the property;
[0166] Form a data closed-loop mechanism: The system supports automatically collecting semantic feedback and scenario data from natural user interactions, driving the continuous optimization and evolution of the model, and constructing a self-enhancing cycle of "data - model - application" in intelligent property management.
[0167] Combined Figure 2 as shown Figure 2 is a schematic block diagram of a retrieval device based on a knowledge graph provided by an embodiment of the present invention. The retrieval device 200 based on the knowledge graph includes:
[0168] A node retrieval unit 201, configured to obtain a knowledge graph, output a node relationship vector by using a preset relational learning module, and retrieve nodes related to the user query content from the knowledge graph, and generate an external knowledge representation result;
[0169] A feature fusion unit 202 is configured to perform feature fusion on the user query content and the external knowledge representation result using a correlation matrix to obtain a fused feature vector;
[0170] A semantic alignment unit 203 is configured to perform bidirectional semantic alignment of the image and text on the fused feature vector based on an attention mechanism to generate a joint feature vector for cross-modal retrieval;
[0171] The feature matching unit 204 is configured to perform a matching calculation based on the joint feature vector to obtain a matching score representing the semantic relevance between the user query content and the candidate graphic and text data for use in cross-modal retrieval.
[0172] In this embodiment, the node retrieval unit 201 obtains a knowledge graph, uses a preset relationship learning module to output a node relationship vector, and retrieves nodes related to the user query content from the knowledge graph, and generates an external knowledge representation result; the feature fusion unit 202 uses a correlation matrix to fuse the user query content with the external knowledge representation result to obtain a fused feature vector; the semantic alignment unit 203 performs bidirectional semantic alignment of the image and text on the fused feature vector based on the attention mechanism to generate a joint feature vector for cross-modal retrieval; the feature matching unit 204 performs matching calculation based on the joint feature vector to obtain a matching score for characterizing the degree of semantic relevance between the user query content and the candidate image and text data for cross-modal retrieval.
[0173] In one embodiment, the knowledge graph-based retrieval device 200 further includes:
[0174] A data construction unit is used to construct an image-text pairing dataset under a target scene; wherein the image-text pairing dataset includes a plurality of image training data with scene semantics and text description training data corresponding to the image training data;
[0175] A feature extraction unit is used to extract features from the image training data and the text description training data respectively to obtain an image training feature vector and a text training feature vector;
[0176] A node extraction unit, configured to extract nodes and semantic representations related to the target scenario from the knowledge graph to obtain a node semantic set;
[0177] A data fusion unit, configured to fuse the image training feature vector, the text training feature vector, and the node semantic set based on a node fusion mechanism to generate a fused node vector for representing multimodal information;
[0178] A vector construction unit, configured to input the fused node vectors into an original relationship learning module to construct the node relationship vectors among multiple nodes;
[0179] A function construction unit, configured to construct a loss function and use the loss function to train the prediction ability of the original relationship learning module to obtain the preset relationship learning module.
[0180] In one embodiment, the data fusion unit includes:
[0181] A feature generation unit, configured to respectively input the image training feature vectors and text training feature vectors into corresponding projection matrices according to the node semantic sets to generate an image-side feature representation and a text-side feature representation;
[0182] A factor generation unit, configured to respectively input the image-side feature representation and the text-side feature representation into corresponding multi-layer perceptrons to generate an image-side gating factor and a text-side gating factor;
[0183] A factor adjustment unit, configured to adjust the image-side feature representation by using the image-side gating factor and adjust the text-side feature representation by using the text-side gating factor to respectively generate an image-gated feature and a text-gated feature;
[0184] A feature splicing unit, configured to splice the image-gated feature and the text-gated feature and input the spliced result into a hybrid projection matrix to generate the fused node vectors.
[0185] In one embodiment, the node retrieval unit 201 includes:
[0186] A first retrieval unit, configured to, when the user query content is an image, input the image into a preset visual feature extraction model to extract image feature vectors of multiple regions in the image, and at the same time input the image into the knowledge graph to retrieve multiple first nodes related to the image to obtain corresponding first node indexes;
[0187] A first extraction unit, configured to extract the fused node vectors corresponding to the first nodes based on the first node indexes;
[0188] A first modeling unit, configured to input the image feature vectors and the fused node vectors into the preset relationship learning module for relationship modeling and output a first node relationship vector between the image and the first nodes;
[0189] a second retrieval unit configured to, when the user query content is text, input the text into a preset text feature extraction model to extract text feature vectors at each token level in the text, and simultaneously input the text into the knowledge graph to retrieve a plurality of second nodes related to the text and obtain corresponding second node indexes;
[0190] a second extraction unit, configured to extract the fused node vector corresponding to the second node based on the second node index;
[0191] A second modeling unit, configured to input the text feature vector and the fusion node vector into the preset relationship learning module for relationship modeling, and output the second node relationship vector between the text and the second node;
[0192] The vector integration unit is used to integrate the first node relationship vector and the second node relationship vector to obtain the external knowledge representation result.
[0193] In one embodiment, the feature fusion unit 202 includes:
[0194] A vector input unit, configured to input a plurality of vectors in the user query content as a plurality of heterogeneous entities; wherein the vectors include an image feature vector, a text feature vector, a fusion node vector, and a node relationship vector;
[0195] a vector embedding unit, configured to embed the vector into a unified representation space and calculate semantic relevance scores between heterogeneous entities to construct a relevance matrix having relationships between multiple entities;
[0196] a vector setting unit, configured to use the image feature vector or the text feature vector as a query vector based on the correlation matrix;
[0197] a vector mapping unit, configured to map the fused node vector and the node relationship vector to a preset weight matrix based on the correlation matrix to obtain corresponding key vectors and value vectors;
[0198] A score calculation unit is used to calculate the attention score based on the query vector, key vector and value vector to generate the fused feature vector.
[0199] In one embodiment, the semantic alignment unit 203 includes:
[0200] A first execution unit is configured to perform a first attention calculation based on the attention mechanism using the fused image feature vector as a first query vector and the fused text feature vector as a first key vector and a first value vector to obtain a first alignment vector;
[0201] a second execution unit, configured to perform a second attention calculation based on the attention mechanism using the fused text feature vector as a second query vector and the fused image feature vector as a second key vector and a second value vector to obtain a second alignment vector;
[0202] a third execution unit, configured to perform residual connection and regularization processing on the first alignment vectors respectively to generate image feature representation;
[0203] a fourth execution unit, configured to perform residual connection and regularization processing on the second alignment vectors to generate text feature representation;
[0204] a fifth execution unit, configured to input the image feature representation and the text feature representation into a feedforward neural network to generate a joint feature representation after semantic fusion of image and text;
[0205] a sixth execution unit, configured to repeatedly execute the first execution unit to the sixth execution unit at least twice to enhance a deep semantic alignment effect between the image and the text;
[0206] The seventh execution unit is configured to integrate the joint feature representation to obtain the joint feature vector.
[0207] In one embodiment, the feature matching unit 204 includes:
[0208] A linear mapping unit, configured to perform linear mapping on the image local features and the text local features respectively to obtain corresponding image mapping features and text mapping features;
[0209] A weight calculation unit, configured to normalize the image mapping features and the text mapping features respectively to obtain corresponding image attention weights and text attention weights;
[0210] A first multiplication unit is used to multiply the image attention weight and the image mapping feature to generate an image global feature;
[0211] A second multiplication unit is used to multiply the text attention weight by the text mapping feature to generate a text global feature;
[0212] A feature adding unit is used to add the image global feature and the text global feature to generate the matching score.
[0213] Since the embodiments of the apparatus part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the apparatus part, and they will not be repeated here.
[0214] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, the steps provided in the above embodiments can be implemented. The storage medium may include various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0215] An embodiment of the present invention further provides a computer device, which may include a memory and a processor. When the processor calls the computer program stored in the memory, the steps provided in the above embodiments can be implemented. Of course, the computer device may further include various network interfaces, power supplies, and other components.
[0216] The embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part. It should be noted that for those of ordinary skill in the art in the technical field of the present application, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
[0217] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, article or device comprising the element.
Claims
1. A cross-modal information retrieval method based on a knowledge graph, characterized in that Including: Obtain a knowledge graph, use a preset relationship learning module to output node relationship vectors, retrieve nodes related to the user query content from the knowledge graph, and generate an external knowledge representation result; Use a correlation matrix to perform feature fusion on the user query content and the external knowledge representation result to obtain a fused feature vector; Perform image-text bidirectional semantic alignment on the fused feature vector based on an attention mechanism to generate a joint feature vector for cross-modal retrieval; Perform matching calculations based on the joint feature vector to obtain a matching score for characterizing the semantic correlation degree between the user query content and candidate image-text data for cross-modal retrieval; The step of using a preset relationship learning module to output node relationship vectors, retrieve nodes related to the user query content from the knowledge graph, and generate an external knowledge representation result includes: when the user query content is an image, input the image into a preset visual feature extraction model to extract image feature vectors of multiple regions in the image, and at the same time input the image into the knowledge graph to retrieve multiple first nodes related to the image to obtain corresponding first node indexes; extract fused node vectors corresponding to the first nodes based on the first node indexes; Input the image feature vectors and the fused node vectors into the preset relationship learning module for relationship modeling to output a first node relationship vector between the image and the first nodes; when the user query content is text, input the text into a preset text feature extraction model to extract text feature vectors at each Token level in the text, and at the same time input the text into the knowledge graph to retrieve multiple second nodes related to the text to obtain corresponding second node indexes; extract fused node vectors corresponding to the second nodes based on the second node indexes; input the text feature vectors and the fused node vectors into the preset relationship learning module for relationship modeling to output a second node relationship vector between the text and the second nodes; integrate the first node relationship vector and the second node relationship vector to obtain the external knowledge representation result.
2. The cross-modal information retrieval method based on a knowledge graph according to claim 1, wherein Before the step of obtaining a knowledge graph, using a preset relationship learning module to output node relationship vectors, retrieving nodes related to the user query content from the knowledge graph, and generating an external knowledge representation result, it further includes: Construct an image-text pairing dataset in the target scenario; wherein, the image-text pairing dataset includes multiple image training data with scene semantics and corresponding text description training data for the image training data; Respectively perform feature extraction on the image training data and the text description training data to obtain image training feature vectors and text training feature vectors; Extract nodes and semantic representations related to the target scenario from the knowledge graph to obtain a node semantic set; Based on a node fusion mechanism, fuse the image training feature vectors, the text training feature vectors, and the node semantic set to generate a fused node vector for characterizing multi-modal information; Input the fused node vectors into the original relationship learning module to construct the node relationship vectors among multiple nodes; Construct a loss function and use the loss function to train the prediction ability of the original relationship learning module to obtain the preset relationship learning module.
3. The cross-modal information retrieval method based on a knowledge graph according to claim 1, wherein The feature fusion of the user query content and the external knowledge representation result using the correlation matrix to obtain a fused feature vector includes: Take multiple vectors in the user query content as multiple heterogeneous entities for input; where the vectors include image feature vectors, text feature vectors, fused node vectors, and node relationship vectors; Embed the vectors into a unified representation space and calculate the semantic correlation scores among the heterogeneous entities to construct a correlation matrix with the relationships among multiple entities; Based on the correlation matrix, use the image feature vector or the text feature vector as a query vector; Based on the correlation matrix, map the fused node vector and the node relationship vector to a preset weight matrix respectively to obtain corresponding key vectors and value vectors; Calculate attention scores based on the query vector, key vector, and value vector to generate the fused feature vector.
4. The cross-modal information retrieval method based on a knowledge graph according to claim 1, characterized in that The fused feature vector includes the fused image feature vector and the fused text feature vector. The bidirectional semantic alignment of the fused feature vector based on the attention mechanism to generate a joint feature vector for cross-modal retrieval includes the following steps: Step 1: Based on the attention mechanism, use the fused image feature vector as the first query vector, the fused text feature vector as the first key vector and the first value vector, and perform the first attention calculation to obtain the first alignment vector; Step 2: Based on the attention mechanism, use the fused text feature vector as the second query vector, the fused image feature vector as the second key vector and the second value vector, and perform the second attention calculation to obtain the second alignment vector; Step 3: Perform residual connection and regularization processing on the first alignment vector respectively to generate an image feature representation; Step 4: Perform residual connection and regularization processing on the second alignment vector respectively to generate a text feature representation; Step 5: Input the image feature representation and the text feature representation into a feed-forward neural network to generate a joint feature representation after the semantic fusion of the image and text; Step 6: Repeat steps 1 to 5 at least twice to enhance the deep semantic alignment effect between the image and the text; Step 7: Integrate the joint feature representation to obtain the joint feature vector.
5. The cross-modal information retrieval method based on a knowledge graph according to claim 1, wherein The joint feature vector includes image local features and text local features. The matching calculation based on the joint feature vector to obtain a matching score for characterizing the semantic correlation degree between the user query content and the candidate image-text data includes: Perform linear mapping on the image local features and the text local features respectively to obtain corresponding image mapping features and text mapping features; Perform normalization processing on the image mapping feature and the text mapping feature respectively to obtain corresponding image attention weights and text attention weights; Multiply the image attention weight with the image mapping feature to generate an image global feature; Multiply the text attention weight with the text mapping feature to generate a text global feature; Add the image global feature and the text global feature to generate the matching score.
6. The cross-modal information retrieval method based on a knowledge graph according to claim 2, wherein The method of fusing the image training feature vector, the text training feature vector and the node semantic set based on the node fusion mechanism to generate a fusion node vector for representing multimodal information includes: Input the image training feature vector and the text training feature vector into corresponding projection matrices respectively according to the node semantic set to generate an image-side feature representation and a text-side feature representation; Input the image-side feature representation and the text-side feature representation into corresponding multi-layer perceptrons respectively to generate an image-side gating factor and a text-side gating factor; Adjust the image-side feature representation by using the image-side gating factor, and adjust the text-side feature representation by using the text-side gating factor to generate an image gated feature and a text gated feature respectively; Concatenate the image gated feature and the text gated feature and input them into a hybrid projection matrix to generate the fusion node vector.
7. A retrieval device based on a knowledge graph, characterized in that, It includes: A node retrieval unit for obtaining a knowledge graph, outputting a node relationship vector by using a preset relationship learning module, and retrieving nodes related to the user query content from the knowledge graph and generating an external knowledge representation result; A feature fusion unit for fusing the user query content and the external knowledge representation result by using a correlation matrix to obtain a fusion feature vector; A semantic alignment unit for performing two-way image-text semantic alignment on the fusion feature vector based on the attention mechanism to generate a joint feature vector for cross-modal retrieval; A feature matching unit for performing a matching calculation based on the joint feature vector to obtain a matching score for characterizing the semantic correlation degree between the user query content and the candidate image-text data for cross-modal retrieval; The node retrieval unit is specifically configured to, when the user query content is an image, input the image into a preset visual feature extraction model to extract image feature vectors of multiple regions in the image, and at the same time input the image into the knowledge graph to retrieve multiple first nodes related to the image to obtain corresponding first node indexes; extract the fusion node vectors corresponding to the first nodes based on the first node indexes; Input the image feature vector and the fusion node vector into the preset relationship learning module for relationship modeling, and output the first node relationship vector between the image and the first node; when the user query content is text, input the text into the preset text feature extraction model to extract the text feature vectors at the Token level in the text, and at the same time input the text into the knowledge graph to retrieve multiple second nodes related to the text, obtain the corresponding second node indexes; extract the fusion node vectors corresponding to the second nodes based on the second node indexes; input the text feature vectors and the fusion node vectors into the preset relationship learning module for relationship modeling, and output the second node relationship vector between the text and the second node; integrate the first node relationship vector and the second node relationship vector to obtain the external knowledge representation result.
8. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the knowledge graph-based retrieval method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, it implements the knowledge graph-based retrieval method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-modal knowledge graph entity alignment method and system based on inter-modal interaction
CN116932770A
Information retrieval method and device, electronic equipment and storage medium
CN119513334A