Retrieval method and device based on knowledge graph, computer equipment and storage medium

Through the cross-modal information retrieval method based on knowledge graph, the problem of low accuracy of image and text retrieval in the community property management system is solved, efficient semantic alignment and information matching of images and text are achieved, and the accuracy of information retrieval and model adaptability are improved.

CN120196767AActive Publication Date: 2025-06-24SHENZHEN ALL THINGS CLOUD TECH CO LTD +1

Patent Information

Application Number
CN202510671809.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-24
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

In the existing community property management system, image recognition methods cannot process intention information described in natural language, and traditional text retrieval technology is difficult to perceive image content, resulting in low accuracy of information retrieval.

Method used

A cross-modal information retrieval method based on knowledge graph is adopted, and the node relationship vectors of the knowledge graph are obtained, feature fusion and bidirectional semantic alignment of the graph and text are performed, joint feature vectors are generated, and matching calculations are performed to improve the retrieval accuracy.

Benefits of technology

It significantly improves the information retrieval accuracy of the community property management system, enhances the ability to understand complex scenarios, and improves the accuracy of graphics and text matching and the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196767A_ABST
    Figure CN120196767A_ABST
Patent Text Reader

Abstract

The invention discloses a retrieval method and device based on a knowledge graph, computer equipment and a storage medium. The method comprises the steps that the knowledge graph is acquired, a node relation vector is output through a preset relation learning module, nodes related to user query content are retrieved from the knowledge graph, and an external knowledge representation result is generated; performing feature fusion on the user query content and the external knowledge representation result by using the correlation matrix to obtain a fused feature vector; performing image-text bidirectional semantic alignment on the fused feature vector based on an attention mechanism to generate a joint feature vector; and performing matching calculation based on the joint feature vector to obtain a matching score for representing the semantic correlation degree. According to the method, the user query content and the generated external knowledge representation result are fused through the correlation matrix, image-text semantics are aligned through an attention mechanism, the joint feature vector is generated, the matching score is calculated, and cross-modal retrieval is realized by using the matching score, so that the information retrieval accuracy of the community property management system is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a retrieval method, device, computer device and storage medium based on a knowledge graph. Background Art

[0002] With the development of intelligent property management, community property management systems have gradually introduced means such as video surveillance, image acquisition and text repair reports in order to improve service efficiency and management level. However, existing image recognition methods are mostly based on single-modal computer vision models and cannot process the intention information described by users in natural language; traditional text retrieval technologies also have difficulty in perceiving the complex scene content contained in images, and thus cannot achieve deep semantic reasoning based on multi-modal input.

[0003] In recent years, although multi-modal learning technology has made certain progress, its practical application in community property management systems still faces problems such as insufficient graphic and text semantic alignment, weak external knowledge dependence and insufficient model generalization ability, resulting in low information retrieval accuracy in community property management systems. Summary of the Invention

[0004] Embodiments of the present invention provide a retrieval method, device, computer device and storage medium based on a knowledge graph, aiming to solve the problem of low information retrieval accuracy in existing community property management systems.

[0005] In a first aspect, an embodiment of the present invention provides a cross-modal information retrieval method based on a knowledge graph, including:

[0006] Obtain a knowledge graph, use a preset relationship learning module to output node relationship vectors, and retrieve nodes related to the user query content from the knowledge graph, and generate an external knowledge representation result;

[0007] Use a correlation matrix to perform feature fusion on the user query content and the external knowledge representation result to obtain a fusion feature vector;

[0008] Based on an attention mechanism, perform graphic and text bidirectional semantic alignment on the fusion feature vector to generate a joint feature vector for cross-modal retrieval;

[0009] Based on the joint feature vector, perform a matching calculation to obtain a matching score for characterizing the semantic correlation degree between the user query content and the candidate graphic and text data for cross-modal retrieval.

[0010] In a second aspect, an embodiment of the present invention provides a retrieval device based on a knowledge graph, including:

[0011] A node retrieval unit, configured to obtain a knowledge graph, output node relationship vectors by using a preset relationship learning module, retrieve nodes related to user query content from the knowledge graph, and generate an external knowledge representation result;

[0012] A feature fusion unit, configured to perform feature fusion on the user query content and the external knowledge representation result by using a correlation matrix to obtain a fused feature vector;

[0013] A semantic alignment unit, configured to perform two-way text-image semantic alignment on the fused feature vector based on an attention mechanism to generate a joint feature vector for cross-modal retrieval;

[0014] A feature matching unit, configured to perform matching calculation based on the joint feature vector to obtain a matching score for characterizing the semantic correlation degree between the user query content and candidate text-image data for cross-modal retrieval.

[0015] In a third aspect, an embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the cross-modal information retrieval method based on a knowledge graph in the first aspect is implemented.

[0016] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, where a computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the cross-modal information retrieval method based on a knowledge graph in the first aspect is implemented.

[0017] An embodiment of the present invention provides a cross-modal information retrieval method based on a knowledge graph, including obtaining a knowledge graph, outputting node relationship vectors by using a preset relationship learning module, retrieving nodes related to user query content from the knowledge graph, and generating an external knowledge representation result; performing feature fusion on the user query content and the external knowledge representation result by using a correlation matrix to obtain a fused feature vector; performing two-way text-image semantic alignment on the fused feature vector based on an attention mechanism to generate a joint feature vector for cross-modal retrieval; performing matching calculation based on the joint feature vector to obtain a matching score for characterizing the semantic correlation degree between the user query content and candidate text-image data. The present invention fuses the user query content and the generated external knowledge representation result through a correlation matrix, then aligns the text-image semantics through an attention mechanism, generates a joint feature vector and calculates a matching score, thereby realizing cross-modal retrieval. In this way, the information retrieval accuracy of the community property management system is greatly improved.

[0018] An embodiment of the present invention further provides a cross-modal information retrieval device, a computer device, and a storage medium based on a knowledge graph, which also have the above beneficial effects. Brief Description of the Drawings

[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 It is a schematic flowchart of a cross-modal information retrieval method based on a knowledge graph provided by an embodiment of the present invention; Figure 2 It is a schematic block diagram of a cross-modal information retrieval device based on a knowledge graph provided by an embodiment of the present invention. Detailed Embodiments

[0021] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0022] It should be understood that when used in this specification and the appended claims, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0023] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0024] It should be further understood that the term " / and / " as used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0025] Please refer to the following Figure 1 , Figure 1 It is a schematic flowchart of a cross-modal information retrieval method based on a knowledge graph provided by an embodiment of the present invention, specifically including: steps S101 to S104.

[0026] S101. Obtain a knowledge graph, use a preset relationship learning module to output node relationship vectors, retrieve nodes related to the user's query content from the knowledge graph, and generate an external knowledge representation result; S102. Use a correlation matrix to perform feature fusion on the user's query content and the external knowledge representation result to obtain a fused feature vector; S104. Based on the attention mechanism, perform two-way semantic alignment of text and images on the fused feature vector to generate a joint feature vector for cross-modal retrieval; S104. Perform matching calculations based on the joint feature vector to obtain a matching score for characterizing the semantic correlation degree between the user's query content and the candidate text-image data for cross-modal retrieval.

[0027] In one embodiment, before step S101, the following steps are further included:

[0028] Construct an image-text pairing data set in the target scenario; wherein, the image-text pairing data set includes multiple image training data with scene semantics and corresponding text description training data for the image training data;

[0029] Extract features from the image training data and the text description training data respectively to obtain image training feature vectors and text training feature vectors;

[0030] Extract nodes and semantic representations related to the target scenario from the knowledge graph to obtain a node semantic set;

[0031] Based on a node fusion mechanism, fuse the image training feature vectors, text training feature vectors, and the node semantic set to generate a fusion node vector for characterizing multi-modal information;

[0032] Input the fusion node vector into the original relationship learning module to construct the node relationship vectors between multiple nodes;

[0033] Construct a loss function, and use the loss function to train the prediction ability of the original relationship learning module to obtain the preset relationship learning module.

[0034] In this embodiment, to construct the data resources required for training the cross-modal knowledge retrieval model applicable to the community property scenario, it is necessary to first execute the construction process of the image-text pairing data set. The construction process includes three stages: acquisition of image training data, generation of text description training data, and manual verification and revision. Finally, an image-text pairing data set is generated as the basic support for model training and feature learning.

[0035] In the acquisition stage of image training data, image training data is collected based on cameras installed in different functional areas within the property community. To ensure the comprehensiveness and diversity of the data, multiple target scene locations are covered, including public areas, around the swimming pool, children's playground, garage entrances and exits, etc. Further, environmental changes under different weather conditions (such as sunny days, cloudy and rainy days, nights, etc.) and time periods (such as daytime, evening, late at night, etc.) are considered to enhance the adaptability and generalization ability of the data. During the process of acquiring image training data, a frame extraction strategy is set according to the surveillance video sources within the community, and image frames are extracted at a predetermined time interval. Video segments that trigger specific events (such as target movement, abnormal behavior, etc.) are preferentially collected, thus ensuring the full capture and coverage of community event-related targets.

[0036] For the collected image training data, according to 22 types of target detection tasks in community property management, the image training data is initially labeled and scene-classified. The specific classification tasks include: garbage in public areas, employees leaving their posts in the command center, trash cans overflowing, children playing by the pool, congestion at entrances and exits, people lingering at garage entrances and exits, motor vehicles staying put, electric vehicles parked randomly, intrusion into the community perimeter, motor vehicles parked illegally, trash beside the trash can, electric vehicles entering the lobby, entering the pool during non-business hours, electric vehicles entering the elevator, the pool security guard leaving their post, people falling, pets walking alone, children walking alone, face recognition, detection of open flames, playing with mobile phones on duty, and sundries piled up in the fire passage.

[0037] In the generation stage of text description training data, based on the actual business requirements of community property management and the goal of semantic understanding of scene events, multi-angle semantic question-and-answer tasks are constructed for each image. For example, 5 standard questions can be generated for each image, including dimensions such as environmental state description, personnel behavior discrimination, traffic condition assessment, implicit risk inference, and scene applicability judgment. Examples of the questions are as follows: (1) What is the environmental situation of this image? Please describe it in detail; (2) Are there any people in the image? Is there any obvious social interaction between them? (3) Is there any traffic flow or potential traffic hazards in the picture? (4) Does the phenomenon reflected in this picture have potential risks in the context of community management? (5) Which specific scenario in property management does this image apply to? Does it have a warning or reference significance? Based on the above set of questions, using a pre-trained multi-modal large language model (such as Qwen2-VL-7b), each image and the corresponding questions are input to generate preliminary text answers, forming the machine's understanding result of the image semantic information. This process can automatically complete the preliminary semantic annotation of a large number of images, improving the efficiency of dataset construction.

[0038] In the manual verification and revision stage, to ensure the accuracy and usability of the final text description, after the large model generation is completed, manual review and revision can be arranged for the 5 text answers corresponding to each image item by item. The content of manual verification includes the coherence of semantic logic, the accuracy of event and object descriptions, the standardization of grammatical structures, and the supplementation of background details such as time and space. During the verification process, if it is found that the content generated by the model is ambiguous, incomplete, or inconsistent with the image, the manual will modify or rewrite it according to the facts to ensure that the text and image semantics strictly correspond and are trainable. Finally, each image corresponds to 5 manually revised text descriptions, forming the complete semantic annotation of the image.

[0039] Based on the above process, the finally formed image-text paired dataset is as follows: The image part includes T static image frames collected from real community monitoring devices, covering 22 types of typical scenario events; the text part correspondingly generates 5T semantic description statements, covering information such as event status, environmental information, object attributes, and risk metaphors, fully reflecting the multi-modal semantic complementarity and task-driven relevance.

[0040] Furthermore, feature extraction is performed on the image training data and text description data respectively. For the image modality, a convolutional neural network based on ResNet or the Faster-RCNN structure can be used to extract regional-level semantic embedding features to obtain the image training feature vector; for the text modality, each description statement is Token-level encoded through a pre-trained language model (such as BERT) to obtain the text training feature vector. Then, semantic node and structural relationship information related to the target scenario is extracted from the knowledge graph to form a node semantic set.

[0041] To solve the problems of semantic heterogeneity and cross-modal alignment difficulties between the image and text modalities in the community property scenario, this embodiment integrates image information, text information, and knowledge graph structured semantics through a unified modeling method, thereby providing a more generalizable and inferable feature basis for subsequent cross-modal retrieval. Specifically, this unified representation method uses the Visualsem knowledge graph as the semantic enhancement carrier to comprehensively fuse and vectorize model the text annotations, image instances, and semantic nodes contained in the graph. Visualsem is an open knowledge graph for multi-language and multi-modal information processing, covering approximately 1.3 million lexical annotations and 938,000 image resources, forming a multi-modal triple set consisting of 89,896 unique semantic nodes and approximately 1.48 million structural relationships between them , representing the node connected to the node through the relationship .

[0042] Since Visualsem covers a wide range of content, including a large amount of general knowledge information not related to community property management, in order to improve the scene adaptability and accuracy of the model, it is necessary to perform pre-screening processing on this knowledge graph. The specific processing process includes: first, constructing a vocabulary list in the field of property management, and through means of word vector semantic matching and image content analysis, screening out a set of nodes related to the property management scene from the original knowledge graph, denoted as M target nodes, and the total number of annotation statements and images associated with these M nodes. Subsequently, for each target node, collect its associated text annotations and explanatory images, and after manually confirming the semantic validity, retain them as the original materials for cross-modal fusion modeling. Among them, the total number of valid original materials is annotation statements and images. Then, perform feature extraction operations on the text and image data of each node in Visualsem to construct a unified cross-modal representation for subsequent model learning. All nodes in the entire knowledge graph are interlinked with each other in a similar way. Taking the node "Air quality index" in the Visualsem knowledge graph as an example, for the M target nodes selected in the previous step, the specific data processing method is as follows:

[0043] First, for any th node among the M nodes, use the LSTM and ResNet networks to perform feature extraction on the vocabulary annotations and images related to this node respectively, and obtain the corresponding several text and image original high-level semantic representation vectors and respectively, where represents the th annotation related to the th node, represents the th image related to the th node; then, for all text vectors and image vectors , use the average aggregation method to obtain and respectively, to represent the final text feature and the final image feature of the th node. The formulas for text and image average aggregation are as follows: .

[0044] Furthermore, a node fusion mechanism is constructed to unify the modeling of the image training feature set, the text training feature set, and the node semantic set. It includes: using LSTM and ResNet to extract the original semantic vectors from the vocabulary annotations and images related to each node respectively, and forming the node text features and image features through average aggregation; then, through the GraphSAGE graph neural network structure combined with the adjacency structure, perform structure-aware feature aggregation on the node's own information to obtain a node vector containing multi-modal features and structural semantics; subsequently, through the node gating unit, fuse the image training feature set, the text training feature set, and the node semantic vector into a fused node vector.

[0045] Specifically, in the node feature aggregation stage, an improved graph neural network structure - GraphSAGE (Graph Sample and Aggregate) can be used to perform feature aggregation and update on the semantic nodes in the Visualsem knowledge graph. Different from the traditional GCN full-graph convolutional calculation, GraphSAGE takes each node as the center during the calculation process, samples its neighbor nodes to construct a subgraph, and generates a context-aware expression of the target node through local aggregation operations, effectively improving the computational efficiency on large-scale graphs and the dynamics of node expressions. For any target node A in the Visualsem knowledge graph, an Aggregator function is used to learn its local neighbor aggregation feature information. Each Aggregator function aggregates information from different hops (here taking 1) or different search depths of this node, and its role is to convert a set of vectors into a vector. The specific formula of the Aggregator function is as follows: ; Among them, represents the total number of 1-hop nodes of the target node, represents the weight matrix, represents the initial feature vectors of the 1-hop nodes , is the activation function, represents the feature vector of the target node finally generated through the neighbor aggregation operation.

[0046] After completing the structure aggregation, the target node A not only contains its own feature information but also integrates the context semantic features of the adjacent nodes, and has a preliminary graph structure perception ability.

[0047] Further, the fused node vectors are input into the original relationship learning module to construct the node relationship vectors between multiple said nodes. The original relationship learning module constructs a set of node pairs with actual semantic dependencies based on a predefined head and tail node screening strategy, and models the potential semantic relationships between each pair of node pairs. The original relationship learning module uses a unified scale feature mapping and ReLU activation structure to generate relationship vectors. The N generated node pair relationship vectors form a node relationship matrix, which has the ability to recognize 13 predefined semantic relationships.

[0048] The main processes of the original relationship learning module include three stages: feature mapping and scale unification, head and tail node screening and pairing, and relationship vector fusion and non-linear activation. The specific implementation is as follows:

[0049] In the stage of feature mapping and scale unification, let the node features with multimodal fusion information have a dimension of (1, ). Multiply all M with the weight matrix W of dimension ( , ). The purpose is to uniformly map node features from different sources into a standardized scale space, ensuring that information from images and texts can be effectively fused and processed in subsequent relationship learning processes, thereby eliminating possible differences between the original feature spaces. After mapping, the dimension of the M node features is (1, ).

[0050] In the stage of head and tail node screening and matching, after the node features are uniformly mapped to the same scale space, it is necessary to screen and pair the head and tail node concepts. The screened node pairs are those that have clear semantic correspondence relationships in the VisualSem knowledge graph, rather than random pairing. In this way, the interference of irrelevant nodes is avoided, ensuring that the relationship learning module can focus on real semantic relationships rather than irrelevant hypothetical connections. Based on these learned correct relationships, the knowledge understanding breadth of the subsequent model in retrieval and reasoning tasks can be broadened, and by extending and associating relationships with the user's input, deeper possible intentions can be captured. After screening and pairing, there are a total of N pairs of head and tail node pairs (2N M). The 13 inherent relationship types between nodes existing in the VisualSem knowledge graph are: is-a; has-part; related-to; used-for; used-to; subject-of; receives-action; made-of; has-property; gloss-related; synonym; part-of and located-at.

[0051] In the relational vector fusion and non-linear activation stage, linear transformation is performed on the selected N head and tail nodes, mapping them into high-dimensional vectors of dimension, and adding and taking the average of the two paired vectors. Through this operation, the information of the head and tail nodes can be effectively fused to form a comprehensive feature vector to represent their relationship. Subsequently, the ReLU function is used to non-linearly activate this comprehensive feature vector. The ReLU function can further enhance the relationship representation ability of the model by suppressing negative values and keeping positive values unchanged, making the relational vectors between different nodes have stronger discrimination and discriminative power. Finally, N relational vectors between the head and tail nodes predicted by the model are obtained, which will be used for subsequent loss function calculation and gradient update, aiming to cultivate the model's ability to accurately associate external inputs and Visualsem internal knowledge through continuous training.

[0052] The overall process is represented by the following formula: ; where, and represent the features of any pair of paired head and tail nodes among M nodes, W is the weight matrix, represents the linear transformation, represents the relational matrix with the final output dimension of (N, ) (i.e., N node relational vectors).

[0053] Furthermore, to optimize the semantic modeling performance of the original relational learning module, a multi-dimensional loss function is further constructed for training optimization. The training objective of the original relational learning module is to cultivate an external generalization ability of the model under the constraint of known relational Groundtruth, so that when external information is input to the model, it can still accurately and efficiently establish a relational link with the model's internal knowledge.

[0054] First, vector mapping needs to be performed on the 13 inherent relational types in Visualsem as the relational Groundtruth in the loss function. One-Hot encoding is used to uniquely represent each relationship (13-dimensional vector), and then it is mapped into dimensional relational embeddings through a trainable MLP. The specific formula is as follows: ; where, represents the One-Hot encoding of a certain relationship, is a trainable weight matrix responsible for learning the embedding of the One-Hot relationship, is the final relation embedding, and is also the relation groundtruth in the loss function below. The following formula is the correlation loss , contrast loss and triplet loss The specific content forms of the three functions are: ; in, Represents the relationship vector between a pair of nodes predicted by the model, Represents the true relationship vector (Groundtruth) between the node pairs obtained by One-Hot encoding, Indicates that any non The relationship vector of type, represents a constant, represents the sigmoid activation function, ( represents element-wise multiplication) is a score function for a triple.

[0055] Among the three loss functions mentioned above, It uses the conventional cross entropy loss function to narrow the distance between the predicted relationship and the true relationship vector; Based on the cosine similarity consideration, the cosine similarity between the predicted relationship vector and the corresponding true relationship vector is made as high as possible (close to 1), while the cosine similarity with other irrelevant true relationship vectors is made as low as possible, and the gap is made as close to the edge value as possible. ; The role of is to force the model to obtain the same tail entity through the predicted relationship of the head node as that obtained through the real relationship.

[0056] By minimizing the above three loss functions, the model of the original relationship learning module will dynamically adjust the weights and feature representations during the training process to achieve the best ability to predict the true relationship between nodes.

[0057] Furthermore, for the N head and tail node relationship vectors (i.e., predicted relationship matrix) predicted by the model of the original relationship learning module, they are multiplied with the 13 categories of real relationship vectors (composed of the relationship matrix) generated by One-Hot encoding, thereby obtaining an N×13 similarity matrix, in which each row represents the similarity score between a predicted relationship and 13 real relationship categories. Based on the idea that the more similar the vectors, the larger the dot product value, the similarity score of each row is softmax normalized, in which the subscript of the maximum value corresponds to the sequence number of the specific category of the predicted relationship, ranging from 1 to 13.

[0058] Furthermore, in order to verify the effectiveness and robustness of the model of the original relationship learning module in the node relationship modeling task, after completing the construction of the original relationship learning module and the optimization of multiple loss functions, two evaluation indicators commonly used in knowledge graph link prediction can also be set: Mean Reciprocal Rank (MRR) and Hits@n, to intuitively reflect the accuracy and ranking ability of the model in predicting the relationship between head and tail entities. Specifically, the MRR indicator measures the average degree to which the model ranks the real head and tail node relationship in the forefront among all the prediction results. The specific calculation method is as follows: ;

[0059] Among them, N represents the total number of predicted relationships between head and tail entity pairs, It indicates the ranking of the true head-tail relationship in the candidate list output by the model in the i-th sample; the value range of MRR is (0,1]. The larger the value, the more the model tends to rank the correct relationship prediction in the front position, and has a stronger relationship recognition and sorting ability.

[0060] The Hits@n metric focuses on whether the model can predict the top n results (n 13) to accurately identify the true head-tail relationship, we can select n=1, n=3 or n=10 to observe the prediction accuracy at different granularities. The higher the Hits@n, the better the model can identify the correct relationship within a limited candidate range. The specific formula of Hits@n is as follows: ;

[0061] Through the above two indicators, the performance of the model in the task of predicting semantic relations between nodes can be systematically evaluated from the two dimensions of accuracy and sorting ability. With the help of continuous correction and supervision of existing relations in the knowledge graph, the model will gradually learn how to rely on the semantic features of the two nodes themselves to infer the most appropriate relationship type between them. This ability will also be transferred to nodes outside the graph to achieve cross-domain semantic connection and information completion, laying the foundation for subsequent diverse intelligent retrieval and semantic extension understanding. At this point, the original relationship learning module that has been trained and optimized is obtained, which is the preset relationship learning module.

[0062] In one embodiment, the node fusion mechanism is used to fuse the image training feature vector, the text training feature vector and the node semantic set to generate a fused node vector for representing multimodal information, including:

[0063] According to the node semantic set, the image training feature vector and the text training feature vector are respectively input into corresponding projection matrices to generate image-side feature representation and text-side feature representation;

[0064] Input the image - side feature representation and the text - side feature representation into the corresponding multi - layer perceptron respectively to generate an image - side gating factor and a text - side gating factor;

[0065] Use the image - side gating factor to adjust the image - side feature representation, and use the text - side gating factor to adjust the text - side feature representation, and generate an image - gated feature and a text - gated feature respectively;

[0066] Concatenate the image - gated feature and the text - gated feature and input them into a hybrid projection matrix to generate the fused node vector.

[0067] In this embodiment, there is a node gating unit in the node fusion mechanism. Its role is to integrate the text training feature vector and the image training feature vector related to the node into the feature of the node semantic set itself, so as to realize the unified representation of the cross - modal data of the knowledge graph. The specific formula of the node gating unit is as follows: ; Among them, represents the node gating function, represents concatenating vectors, and are the projection matrices about the text training feature vector and the image training feature vector respectively, and represent the multi - layer perceptrons of the text - side feature representation and the image - side feature representation respectively, and are the text - side gating factor and the image - side gating factor respectively, and are the node features integrating the information of the text training feature vector and the image training feature vector. Finally, and are concatenated and then multiplied by the hybrid projection matrix to obtain the node feature with multi - modal data fusion information (i.e., the fused node vector).

[0068] In step S101, first, obtain a pre-constructed knowledge graph, which includes node entities, attribute information, and various semantic relationships related to the community property scenario. The knowledge graph is preferably a multi-modal knowledge graph that integrates image and text modal information, such as the Visualsem knowledge graph, whose nodes have multi-language annotations, illustrative images, and semantic relationship definitions. Use a preset relationship learning module to output node relationship vectors, perform semantic parsing on the query content input by the user, and retrieve target nodes related thereto in the knowledge graph. The preset relationship learning module introduces a graph neural network (such as GraphSAGE) to integrate the text features, image features, and structural features of the nodes, and then generates an external knowledge representation result expressing the query semantics to support subsequent cross-modal semantic modeling.

[0069] In one embodiment, step S101 includes: When the user query content is an image, input the image into a preset visual feature extraction model to extract image feature vectors of multiple regions in the image. At the same time, input the image into the knowledge graph to retrieve multiple first nodes related to the image, and obtain corresponding first node indexes; Extract the fusion node vectors corresponding to the first nodes based on the first node indexes; Input the image feature vectors and the fusion node vectors into the preset relationship learning module for relationship modeling, and output the first node relationship vector between the image and the first nodes; When the user query content is text, input the text into a preset text feature extraction model to extract text feature vectors at the Token level in the text. At the same time, input the text into the knowledge graph to retrieve multiple second nodes related to the text, and obtain corresponding second node indexes; Extract the fusion node vectors corresponding to the second nodes based on the second node indexes; Input the text feature vectors and the fusion node vectors into the preset relationship learning module for relationship modeling, and output the second node relationship vector between the text and the second nodes; Integrate the first node relationship vector and the second node relationship vector to obtain the external knowledge representation result.

[0070] In this embodiment, the image path and the text path are processed respectively based on the user query content to achieve external knowledge enhancement and semantic relationship modeling under the condition of text and image dual modalities. When the user query content is an image, the following operations are performed: Input the image into a preset visual feature extraction model, which is preferably an Faster-RCNN structure and includes a Region Proposal Network (RPN) and an ROI feature extraction module. Generate a number of candidate regions (Regions of Interest) from the image, and calculate the semantic embedding vector for each region to obtain a set of region-level image feature vectors. Each vector corresponds to a local semantic region node (i.e., the first node) in the image, capturing fine-grained spatial semantic information. Then input the image into a knowledge graph, which is a multimodal semantic graph Visualsem, and realize semantic association with the nodes in the knowledge graph through its attached image-text retrieval ability. Use a CLIP model (such as RN50x16) to perform semantic encoding on the input image and the property scene image nodes selected from the knowledge graph. By calculating the cosine similarity of the feature vectors between the images, retrieve several graph nodes most relevant to the input image to obtain the corresponding first node indices. Subsequently, based on the first node indices, extract the corresponding fusion node vectors from a pre-constructed set of fusion node vectors. Input the image feature vectors and the fusion node vectors into a preset relationship learning module. The preset relationship learning module learns the possible explicit semantic relationships (such as "has-part", "used-for", "located-at", etc.) between the image features and the knowledge nodes based on the aforementioned trained structure, and outputs the first node relationship vectors between the image and each retrieved node as the external knowledge representation on the image modality side.

[0071] In the case where the user query content is text, perform the following operations: Input the text into a preset text feature extraction model, preferably the pre-trained language model BERT, to extract the Token-level semantic representation of the input text, obtaining a set of text feature vectors. Each vector corresponds to a semantic unit (such as a word or phrase) and is used to represent the semantic granularity in the text. Input the text into the knowledge graph and perform knowledge enhancement operations similar to those for the image path. The specific steps include using the Sentence-BERT model to perform semantic encoding on the input text and the set of text annotations related to the property domain in the knowledge graph; then, calculating the cosine similarity between the text semantics and screening out several annotations (i.e., the second nodes) with the highest relevance to the input text; finally, determining the corresponding node indices according to the annotation mapping relationship to obtain the second node index set. Extract the corresponding fused node vectors based on the second node indices. Subsequently, input the Token-level feature vectors of the text and the fused node vectors into a preset relationship learning module. Through this preset relationship learning module, learn the semantic relationship connection between the input text content and the graph nodes, and output the second node relationship vector between the text modality and the second nodes, representing the possible semantic relationships between the text semantic units and the graph knowledge entities, such as "subject-of", "has-property", "related-to", etc.

[0072] Finally, merge and integrate the first node relationship vector obtained from the image path and the second node relationship vector obtained from the text path to generate the external knowledge representation result corresponding to the user query content (here there can be two, namely a retrieved node vector and a relationship vector between this retrieved node vector). Through the above processing steps, it is possible to fully mine and model the knowledge structure and semantic context associated with the input content in the preprocessing stage of the text-image query, effectively improving the understanding and reasoning capabilities of fuzzy expressions, implicit relationships, and context semantics in the subsequent text-image matching process, and is particularly suitable for diverse and complex community property intelligent service scenarios.

[0073] In step S102, based on the correlation matrix construction mechanism (the core goal is to jointly embed the input stream (region / token embedding), the knowledge graph retrieval node (retrieved KG node), and the relationship between the two (relation embedding), a total of 3, into a unified representation space), fuse the user query content with the generated external knowledge representation result to obtain the fused multi-modal feature vector (i.e., the fused feature vector). This process takes into account the diversity of semantic expression and the perception ability of external knowledge in terms of structure, improving the understanding accuracy of the model for complex scenarios.

[0074] In one embodiment, step S102 includes: Take multiple vectors in the user query content as multiple heterogeneous entities as input; wherein, the vectors include image feature vectors, text feature vectors, fusion node vectors, and node relationship vectors; Embed the vectors into a unified representation space, and calculate the semantic correlation scores between heterogeneous entities to construct a correlation matrix with the relationships between multiple entities; Based on the correlation matrix, use the image feature vector or the text feature vector as the query vector; Based on the correlation matrix, map the fusion node vector and the node relationship vector to a preset weight matrix respectively to obtain the corresponding key vector and value vector; Calculate the attention scores based on the query vector, key vector, and value vector to generate the fusion feature vector.

[0075] In this embodiment, multiple semantic representation vectors obtained from the user query content are organized as multiple heterogeneous entities as input. This input set includes: multiple region-level image feature vectors extracted in the image modality (such as ROI embeddings extracted by Faster-RCNN), multiple Token-level text feature vectors extracted in the text modality (such as TokenEmbedding output by BERT), fusion node vectors obtained through knowledge graph retrieval (with graph-text semantic and structural information), and node relationship vectors obtained through a preset relationship learning module (representing the semantic relationship between the input features and the knowledge graph). All the above vectors participate in the subsequent unified modeling process as independent semantic entities.

[0076] Furthermore, the system embeds these vectors into a unified representation space. To eliminate the inconsistency of feature dimensions and semantic scales between modalities, a unified embedding transformation matrix is preset, and linear mapping is performed on various vectors to project them into the same embedding space. Then, based on the above embedding representation, calculate the semantic correlation scores between any entities to construct a correlation matrix between entities. For all input entities, calculate the similarity between each pair (such as dot product or cosine similarity) to fill the values in the matrix, and finally form a correlation matrix. After constructing the correlation matrix, perform a multi-modal attention interaction fusion operation based on this correlation matrix. The fusion of external enhanced knowledge and the original input features based on the attention mechanism is divided into an image stream and a text stream for processing. Here, the processing method of the image stream is taken as an example to introduce, and the processing method of the text stream is exactly the same as that of the image stream, as follows: The size of the correlation matrix of the image stream is × , where , let the region-level image features extracted from the CNN Backbone be respectively represented as , ,…… , retrieved from the knowledge graph node features are respectively represented as , ,…… , and these two are predicted by a preset relationship learning module relationship features are respectively represented as , ,…… .

[0077] Based on the attention mechanism, the features are respectively constructed into query, key, and value, referring to the following formula: ; ; ; Among them, the query is only obtained by multiplying the region-level features of the input image with the weight matrix , that is, the query vector only focuses on the original input features. For the key and the value , in addition to focusing on the original features, they also perform weight transformation on the node features and relationship features (a total of features) (multiplying with and ).

[0078] Furthermore, the information fusion process based on the attention mechanism is carried out according to the following formula: ; ; The finally output …… is the region-level image feature (i.e., the fusion feature vector) that integrates external enhanced information (i.e., the external knowledge representation result). Compared with the original input, it not only retains its own representation ability but also integrates richer external semantic-related knowledge, which provides more comprehensive support and guidance for complex and diverse retrieval and question answering in the subsequent community property scenario.

[0079] The correlation matrix of the text stream has a size of q×q, where , let the token-level features extracted from Bert be respectively represented as , ,…… , after similar operations to the image stream, the final output feature representation is …… 。

[0080] In step S103, a text-image bidirectional semantic alignment mechanism is further adopted. The Co-Attention mutual attention structure can be used to perform text-image feature interaction and joint encoding on the fused feature vector. In this mechanism, the image-modal features are used as the query (Query), and the text-modal features are used as the key (Key) and value (Value), as well as the reverse alignment process, to achieve semantic guidance and feature perception between text and image features, and enhance the semantic alignment effect between text and image. Under the action of the multi-layer stacked interaction module, a joint feature vector representing the semantic consistency between the query content and the candidate text-image samples is finally generated, providing high-level semantic support for subsequent matching calculations.

[0081] In one embodiment, step S103 includes the following steps: Step 1: Based on the attention mechanism, use the fused image feature vector as the first query vector, the fused text feature vector as the first key vector and the first value vector, and perform the first attention calculation to obtain the first alignment vector; Step 2: Based on the attention mechanism, use the fused text feature vector as the second query vector, the fused image feature vector as the second key vector and the second value vector, and perform the second attention calculation to obtain the second alignment vector; Step 3: Perform residual connection and regularization processing on the first alignment vector respectively to generate the image feature representation; Step 4: Perform residual connection and regularization processing on the second alignment vector respectively to generate the text feature representation; Step 5: Input the image feature representation and the text feature representation into a feed-forward neural network to generate the joint feature representation after text-image semantic fusion; Step 6: Repeat steps 1 to 5 at least twice to enhance the deep semantic alignment effect between image and text; Step 7: Integrate the joint feature representations to obtain the joint feature vector.

[0082] In this embodiment, Step 1: First, based on the multi-modal attention mechanism, using the fused image feature vector as the first query vector (Query), and the fused text feature vector as the first key vector (Key) and the first value vector (Value), an image-to-text interaction path is constructed. By performing the first attention calculation through this path, the image features can perceive and capture the semantic information expressed in the text, and then generate the first alignment vector on the image side. In a specific implementation, this attention mechanism adopts a multi-head structure and models in parallel in different semantic sub-spaces to enhance the semantic coverage ability and context interaction ability; Step 2: Perform reverse attention path processing, using the fused text feature vector as the second query vector (Query), and the fused image feature vector as the second key vector (Key) and the second value vector (Value), to construct a text-to-image interaction path. By performing the second attention calculation, the text features can focus on the image region information that matches its semantics, and thus output the second alignment vector on the text side; Step 3: For the first alignment vector on the image side, adopt a residual connection mechanism to perform element-wise addition processing on this alignment vector and the original image input features to retain the low-level representation information in the original image features. Subsequently, perform layer normalization processing (Layer Normalization) on the concatenated result to generate the feature representation on the image side. This processing step improves the model stability and avoids feature dissipation, forming a preliminary image representation after image-text interaction; Step 4: For the second alignment vector on the text side, similarly perform residual connection and regularization operations, connect the alignment vector with the initial text embedding vector, and then perform normalization to obtain the text feature representation with the ability to perceive visual information; Step 5: Input the above image feature representation and text feature representation into a feed-forward neural network (FeedForward Network, FFN). This network consists of two fully connected structures, internally integrating an activation function (such as ReLU) and a projection mechanism to further extract the non-linear semantic combinations between cross-modalities. After the feed-forward output undergoes another residual connection and normalization processing, the joint feature representation after image-text semantic fusion at the current layer is output; Step 6: In order to enhance the deep semantic alignment effect between the image and text modalities and improve the hierarchical nature of context understanding, the above Steps 1 to 5 are constructed into a complete image-text interaction sub-module, and this module is stacked multiple times. Preferably, the number of stacks is 5 layers, then the above sub-module is repeatedly executed at least five times, gradually realizing the gradual penetration between image-text semantics layer by layer, so that the final image features contain text semantics and the text features fuse image context, thus forming a high-level cross-modal joint representation; Step 7: Finally, integrate the joint feature representations of text and images output from each interaction layer (e.g., through concatenation, weighted summation, or adaptive aggregation strategies) to generate a set of high-dimensional joint feature vectors, which serve as the basic input for cross-modal similarity calculation and retrieval matching in subsequent steps. This joint feature vector not only retains the original modal information but also integrates the deep semantics from the other modality, and realizes fine-grained context perception through external knowledge graph relationships and attention mechanisms, improving the matching accuracy and generalization ability between text and images.

[0083] Specifically, the calculation can be performed with reference to the following formula: Let the input be , , taking the visual stream as an example (similarly for the text stream): ; ; ; Among them, , , , , and all represent weight coefficients, and represent bias coefficients, h represents the number of attention heads, and the output of the multi-head attention is represented as x after residual connection and regularization. x is input into a feed-forward neural network FFN, and then after residual connection and regularization, is obtained. Similarly, the text stream will obtain .

[0084] For example, stack 5 layers of the above interaction modules to gradually enhance the deep fusion ability of text and image features. Finally, the outputs of the image and text streams are and respectively. At this time, the text semantics have been encoded in the image feature , and the visual context is also included in the text feature , forming a joint feature vector.

[0085] In step S104, perform matching calculation based on the joint feature vector, adopt a weighted aggregation strategy to globally encode the text and image features, and generate a matching score. This matching score is used to measure the semantic correlation degree between the user's query content and each candidate text and image sample, and accordingly complete the cross-modal retrieval task. Preferably, a ranking loss function can be introduced during the matching calculation process to strengthen the semantic discrimination ability between positive and negative samples, thereby improving the efficiency of intelligent question answering, event location, and patrol response in the community property application scenario.

[0086] In one embodiment, step S104 includes: Perform linear mapping on the local image features and local text features respectively to obtain corresponding image mapping features and text mapping features; Perform normalization processing on the image mapping features and text mapping features respectively to obtain corresponding image attention weights and text attention weights; Multiply the image attention weights by the image mapping features to generate image global features; Multiply the text attention weights by the text mapping features to generate text global features; Add the image global features and the text global features to generate the matching score.

[0087] In this embodiment, the matching score can be generated according to the following formula: ; ; ; ; Let and be linearly mapped through weight matrices and respectively, and then after softmax, attention weight factors and are obtained respectively. Multiply the attention weight factors by and respectively to obtain the overall image features and text overall features aggregating all sub-features; then, add and and send the sum into an MLP, and normalize the output to [0, 1] through a sigmoid activation function to represent the probability of image-text matching (i.e., the matching score), where 1 represents image-text matching and 0 represents non-matching of image and text; Finally, create a ranking loss function for the image-text pair based on this matching score , where represents the margin hyperparameter used to control the minimum score gap between positive and negative samples, represents the relevant positive image-text pair, represents the negative sample pair composed of text and an irrelevant image, represents the negative sample pair composed of an image and irrelevant text, It represents a mining solution for difficult negative sample image-text pairs. Through this loss function, the model can be promoted to improve its discrimination ability for whether the image-text pair matches, and strengthen the learning of image-text alignment effect and multi-modal semantic consistency.

[0088] In one embodiment, to achieve the efficient deployment and continuous optimization of the model in the real community property environment, the model inference and data closed-loop stage can also be set. This stage focuses on the edge deployment of the model, real-time multi-modal input processing, natural language retrieval inference response, and feedback-driven model update mechanism, aiming to build a graphic-semantic closed-loop system integrating real-time intelligent recognition, multi-role interaction, and self-evolution capabilities.

[0089] First, after the model training is completed and compressed and optimized, it is deployed on edge computing devices (such as property servers, intelligent camera nodes, security terminals, etc.) in the community property environment to support low-latency and high-efficiency on-site graphic-text inference capabilities. The edge device continuously collects image data from the community camera stream, automatically extracts image frames using a set frame extraction strategy (including timed frame extraction and event-triggered frame extraction) and performs local structured storage, providing a sustainable image input source for the model, and building an image index database containing time stamps, location information, and event tags.

[0090] At the same time, during the daily management of the community, a natural language query interface for multi-role users is supported. For example, owners, property security personnel, front desk staff, etc. can input questions and query instructions based on the actual scenario needs through voice or text input. This natural language input form usually contains multi-modal clues such as semantic targets, time constraints, and spatial limitations, such as "Where was the old man riding an electric bike seen yesterday evening?" or "Which building has a full trash can nearby?" etc. Such query requests are used as user query content input, and graphic-semantic parsing and cross-modal retrieval are performed according to the above process: text feature extraction and knowledge graph enhancement processing are performed on the natural language query to construct a joint feature representation with semantic representation and structural knowledge association; the query feature vector is used to perform cross-modal matching calculation with the image features in the edge-stored image library to obtain a matching score; finally, the images are ranked according to the matching scores and the most relevant image results are returned to respond to user requests.

[0091] The image-text matching reasoning process not only integrates the semantics of images and texts, but also introduces the background semantics and relational reasoning capabilities provided by the knowledge graph, enabling high-accuracy content recall under natural language conditions with vague expressions, uncertain targets, and strong polysemy. Further, during the operation of the system, by analyzing the user's retrieval behavior, query styles, and feedback results, the system can dynamically collect data from multiple perspectives to build a closed-loop mechanism: automatically record the interaction logs of each user retrieval request and its returned results, and analyze whether the user accepts, modifies, or retries the retrieval results; if the user manually confirms, annotates, or corrects the returned results, this feedback is automatically supplemented as a new supervision signal to the model training sample pool. In addition, it can automatically identify query styles or image modalities where the current model performs poorly in specific types of scenarios, and trigger a sample augmentation mechanism based on its recall failure samples, such as supplementing similar image frames or annotating typical text patterns.

[0092] Through the automatic collection and aggregation of the above feedback data, training sample batches are periodically constructed in the background, and incremental update training is performed on the server side or distributed model training platform, enabling the model to continuously improve its generalization ability, semantic alignment robustness, and retrieval reasoning accuracy in real estate scenarios. Finally, a self-loop mechanism of "real-world problem-driven data accumulation - data feeding back model training - model feedback real-world problem solving" is achieved. Through the implementation of this step, a data flywheel closed-loop system characterized by event-driven, interactive feedback, and intelligent optimization can be formed in the actual environment of property management, effectively improving the practicality, interpretability, and continuous evolution ability of the multi-modal semantic model, and providing strong technical support for the precise management and efficient response of smart communities.

[0093] In summary, the present invention first constructs a property image-text data system with scene diversity and multi-dimensional expression, truly restoring the daily management context of the community; secondly, introduces an external knowledge graph to bridge and enhance the image-text semantics through unified modeling of heterogeneous information; then, designs a relational learning module to deeply model the potential semantic structure between images and texts; further introduces a multi-head cross-attention mechanism to achieve deep interaction and joint representation between images and texts; finally, deploys the model on an edge computing platform and establishes a continuous optimization mechanism driven by user interaction and data feedback to construct an intelligent semantic closed-loop system with real-time response ability, knowledge transfer ability, and scene adaptation ability.

[0094] The beneficial effects of the present invention include: Enhanced semantic understanding: Through knowledge graph-assisted modeling and semantic relationship learning, the model's ability to understand complex expressions, vague references, and polysemous contexts in images and texts is significantly improved; Improved Image-Text Matching Accuracy: Through a bidirectional attention mechanism and semantic alignment design, deep and structured interaction and fusion between image and text modalities are achieved, effectively narrowing the semantic gap between images and text. Of course, on the basis of maintaining accuracy, the system also enhances its capabilities in handling diverse expressions, understanding implicit relationships, and generalizing image-text matching in property management scenarios; Strong Robustness of Intelligent Response: In the face of changes in different user roles, language expressions, query intents, and scenario conditions, the system can still maintain stable and accurate response effects; High Efficiency in Edge Deployment: The model can be deployed on community edge devices after quantization and compression, with low-latency, high-efficiency, and scalable on-site inference capabilities, meeting the actual application requirements of property management; Forming a Data Closed-Loop Mechanism: The system supports automatically collecting semantic feedback and scenario data from natural user interactions, driving the continuous optimization and evolution of the model, and constructing a self-enhancing cycle of "data - model - application" in intelligent property management.

[0095] Combined with Figure 2 as shown in Figure 2 FIG. 13 is a schematic block diagram of a retrieval device based on a knowledge graph provided by an embodiment of the present invention. The retrieval device 200 based on the knowledge graph includes: A node retrieval unit 201, configured to obtain a knowledge graph, output a node relationship vector by using a preset relationship learning module, retrieve nodes related to user query content from the knowledge graph, and generate an external knowledge representation result; A feature fusion unit 202, configured to fuse the features of the user query content and the external knowledge representation result by using a correlation matrix to obtain a fused feature vector; A semantic alignment unit 203, configured to perform bidirectional image-text semantic alignment on the fused feature vector based on an attention mechanism to generate a joint feature vector for cross-modal retrieval; A feature matching unit 204, configured to perform a matching calculation based on the joint feature vector to obtain a matching score for characterizing the semantic correlation degree between the user query content and candidate image-text data for cross-modal retrieval.

[0096] In this embodiment, the node retrieval unit 201 obtains a knowledge graph, outputs node relationship vectors by using a preset relationship learning module, retrieves nodes related to the user query content from the knowledge graph, and generates an external knowledge representation result; the feature fusion unit 202 performs feature fusion on the user query content and the external knowledge representation result by using a correlation matrix to obtain a fused feature vector; the semantic alignment unit 203 performs two-way image-text semantic alignment on the fused feature vector based on an attention mechanism to generate a joint feature vector for cross-modal retrieval; the feature matching unit 204 performs matching calculations based on the joint feature vector to obtain a matching score for characterizing the semantic relevance degree between the user query content and the candidate image-text data for cross-modal retrieval.

[0097] In one embodiment, the knowledge graph-based retrieval device 200 further includes: A data construction unit, configured to construct an image-text pairing data set in a target scenario; wherein, the image-text pairing data set includes a plurality of image training data with scene semantics and text description training data corresponding to the image training data; A feature extraction unit, configured to respectively extract features from the image training data and the text description training data to obtain an image training feature vector and a text training feature vector; A node extraction unit, configured to extract nodes and semantic representations related to the target scenario from the knowledge graph to obtain a node semantic set; A data fusion unit, configured to fuse the image training feature vector, the text training feature vector, and the node semantic set based on a node fusion mechanism to generate a fused node vector for characterizing multi-modal information; A vector construction unit, configured to input the fused node vector into an original relationship learning module to construct the node relationship vectors between multiple nodes; A function construction unit, configured to construct a loss function and use the loss function to train the prediction ability of the original relationship learning module to obtain the preset relationship learning module.

[0098] In one embodiment, the data fusion unit includes: A feature generation unit, configured to respectively input the image training feature vector and the text training feature vector into corresponding projection matrices according to the node semantic set to generate an image-side feature representation and a text-side feature representation; A factor generation unit, configured to respectively input the image-side feature representation and the text-side feature representation into corresponding multi-layer perceptrons to generate an image-side gating factor and a text-side gating factor; A factor adjustment unit, configured to adjust the image - side feature representation by using the image - side gating factor, and adjust the text - side feature representation by using the text - side gating factor, respectively generating an image - gated feature and a text - gated feature; A feature splicing unit, configured to splice the image - gated feature and the text - gated feature, and input them into a hybrid projection matrix to generate the fused node vector.

[0099] In one embodiment, the node retrieval unit 201 includes: A first retrieval unit, configured to, when the user query content is an image, input the image into a preset visual feature extraction model to extract image feature vectors of multiple regions in the image, and at the same time input the image into the knowledge graph to retrieve multiple first nodes related to the image, obtaining corresponding first - node indexes; A first extraction unit, configured to extract the fused node vector corresponding to the first node based on the first - node index; A first modeling unit, configured to input the image feature vector and the fused node vector into the preset relationship learning module for relationship modeling, and output a first - node relationship vector between the image and the first node; A second retrieval unit, configured to, when the user query content is text, input the text into a preset text feature extraction model to extract text feature vectors at each Token level in the text, and at the same time input the text into the knowledge graph to retrieve multiple second nodes related to the text, obtaining corresponding second - node indexes; A second extraction unit, configured to extract the fused node vector corresponding to the second node based on the second - node index; A second modeling unit, configured to input the text feature vector and the fused node vector into the preset relationship learning module for relationship modeling, and output a second - node relationship vector between the text and the second node; A vector integration unit, configured to integrate the first - node relationship vector and the second - node relationship vector to obtain the external knowledge representation result.

[0100] In one embodiment, the feature fusion unit 202 includes: A vector input unit, configured to input multiple vectors in the user query content as multiple heterogeneous entities; wherein the vectors include image feature vectors, text feature vectors, fused node vectors, and node relationship vectors; A vector embedding unit, configured to embed the vectors into a unified representation space, and calculate the semantic correlation scores between heterogeneous entities to construct a correlation matrix with the relationships between multiple entities; A vector setting unit, configured to use the image feature vector or the text feature vector as a query vector based on the correlation matrix; A vector mapping unit, configured to map the fusion node vector and the node relationship vector to a preset weight matrix respectively based on the correlation matrix to obtain corresponding key vectors and value vectors; A score calculation unit, configured to calculate an attention score based on the query vector, the key vector, and the value vector to generate the fusion feature vector.

[0101] In one embodiment, the semantic alignment unit 203 includes: A first execution unit, configured to use the fused image feature vector as a first query vector, the fused text feature vector as a first key vector and a first value vector based on the attention mechanism, and perform a first attention calculation to obtain a first alignment vector; A second execution unit, configured to use the fused text feature vector as a second query vector, the fused image feature vector as a second key vector and a second value vector based on the attention mechanism, and perform a second attention calculation to obtain a second alignment vector; A third execution unit, configured to perform residual connection and regularization processing on the first alignment vector respectively to generate an image feature representation; A fourth execution unit, configured to perform residual connection and regularization processing on the second alignment vector respectively to generate a text feature representation; A fifth execution unit, configured to input the image feature representation and the text feature representation into a feed-forward neural network to generate a joint feature representation after graph-text semantic fusion; A sixth execution unit, configured to repeatedly execute the first execution unit to the sixth execution unit at least twice to enhance the deep semantic alignment effect between the image and the text; A seventh execution unit, configured to integrate the joint feature representation to obtain the joint feature vector.

[0102] In one embodiment, the feature matching unit 204 includes: A linear mapping unit, configured to perform linear mapping on the image local feature and the text local feature respectively to obtain corresponding image mapping features and text mapping features; A weight calculation unit, configured to perform normalization processing on the image mapping feature and the text mapping feature respectively to obtain corresponding image attention weights and text attention weights; A first multiplication unit, configured to multiply the image attention weight by the image mapping feature to generate an image global feature; A second multiplication unit, configured to multiply the text attention weight by the text mapping feature to generate a text global feature; A feature addition unit for adding the global image feature and the global text feature to generate the matching score.

[0103] Since the embodiments in the apparatus part correspond to the embodiments in the method part, please refer to the description of the embodiments in the method part for the embodiments in the apparatus part, which will not be elaborated here.

[0104] An embodiment of the present invention also provides a computer-readable storage medium with a computer program stored thereon. When the computer program is executed, the steps provided in the above embodiments can be implemented. The storage medium may include various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0105] An embodiment of the present invention also provides a computer device, which may include a memory and a processor. When the processor calls the computer program stored in the memory, the steps provided in the above embodiments can be implemented. Of course, the computer device may also include various network interfaces, power supplies, and other components.

[0106] The embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple. For the relevant parts, please refer to the description in the method part. It should be noted that for those of ordinary skill in the art of the present technology, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

[0107] It should also be noted that in this specification, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "comprises" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

Claims

1. A cross-modal information retrieval method based on a knowledge graph, characterized in that Including: Obtain a knowledge graph, use a preset relationship learning module to output node relationship vectors, retrieve nodes related to the user query content from the knowledge graph, and generate an external knowledge representation result; Use a correlation matrix to perform feature fusion on the user query content and the external knowledge representation result to obtain a fused feature vector; Based on an attention mechanism, perform text-image two-way semantic alignment on the fused feature vector to generate a joint feature vector for cross-modal retrieval; Based on the joint feature vector, perform a matching calculation to obtain a matching score for characterizing the semantic correlation degree between the user query content and the candidate text-image data for cross-modal retrieval.

2. The cross-modal information retrieval method based on a knowledge graph according to claim 1, wherein Before the step of obtaining the knowledge graph, using a preset relationship learning module to output node relationship vectors, retrieving nodes related to the user query content from the knowledge graph, and generating an external knowledge representation result, further includes: Construct an image-text pairing dataset in a target scenario; wherein, the image-text pairing dataset includes a plurality of image training data with scene semantics and corresponding text description training data for the image training data; Extract features from the image training data and the text description training data respectively to obtain an image training feature vector and a text training feature vector; Extract nodes and semantic representations related to the target scenario from the knowledge graph to obtain a node semantic set; Based on a node fusion mechanism, fuse the image training feature vector, the text training feature vector, and the node semantic set to generate a fused node vector for characterizing multi-modal information; Input the fused node vector into the original relationship learning module to construct the node relationship vectors between the multiple nodes; Construct a loss function, and use the loss function to train the prediction ability of the original relationship learning module to obtain the preset relationship learning module.

3. The cross-modal information retrieval method based on a knowledge graph according to claim 2, wherein The step of using a preset relationship learning module to output node relationship vectors, retrieving nodes related to the user query content from the knowledge graph, and generating an external knowledge representation result includes: When the user query content is an image, input the image into a preset visual feature extraction model to extract image feature vectors of multiple regions in the image, and at the same time input the image into the knowledge graph to retrieve a plurality of first nodes related to the image to obtain corresponding first node indexes; Extract the fused node vectors corresponding to the first nodes based on the first node indexes; Input the image feature vectors and the fused node vectors into the preset relationship learning module for relationship modeling, and output the first node relationship vectors between the image and the first nodes; When the user query content is text, input the text into a preset text feature extraction model to extract text feature vectors at the Token level in the text, and at the same time input the text into the knowledge graph to retrieve a plurality of second nodes related to the text to obtain corresponding second node indexes; Extract the fused node vectors corresponding to the second nodes based on the second node indexes; Input the text feature vector and the fusion node vector into the preset relationship learning module for relationship modeling, and output the second node relationship vector between the text and the second node; Integrate the first node relationship vector and the second node relationship vector to obtain the external knowledge representation result.

4. The cross-modal information retrieval method based on a knowledge graph according to claim 1, wherein The feature fusion of the user query content and the external knowledge representation result using the correlation matrix to obtain a fusion feature vector includes: Take multiple vectors in the user query content as multiple heterogeneous entities; wherein, the vectors include image feature vectors, text feature vectors, fusion node vectors, and node relationship vectors; Embed the vectors into a unified representation space, and calculate the semantic correlation scores between heterogeneous entities to construct a correlation matrix with relationships between multiple entities; Based on the correlation matrix, use the image feature vector or the text feature vector as a query vector; Based on the correlation matrix, map the fusion node vector and the node relationship vector to a preset weight matrix respectively to obtain corresponding key vectors and value vectors; Calculate attention scores based on the query vector, key vector, and value vector to generate the fusion feature vector.

5. The cross-modal information retrieval method based on a knowledge graph according to claim 1, wherein The fusion feature vector includes a fused image feature vector and a fused text feature vector. The method for generating a joint feature vector for cross-modal retrieval by performing two-way semantic alignment of images and texts on the fusion feature vector based on the attention mechanism includes the following steps: Step 1: Based on the attention mechanism, use the fused image feature vector as the first query vector, the fused text feature vector as the first key vector and the first value vector, and perform the first attention calculation to obtain the first alignment vector; Step 2: Based on the attention mechanism, use the fused text feature vector as the second query vector, the fused image feature vector as the second key vector and the second value vector, and perform the second attention calculation to obtain the second alignment vector; Step 3: Perform residual connection and regularization processing on the first alignment vector respectively to generate an image feature representation; Step 4: Perform residual connection and regularization processing on the second alignment vector respectively to generate a text feature representation; Step 5: Input the image feature representation and the text feature representation into a feed-forward neural network to generate a joint feature representation after semantic fusion of images and texts; Step 6: Repeat steps 1 to 5 at least twice to enhance the deep semantic alignment effect between images and texts; Step 7: Integrate the joint feature representations to obtain the joint feature vector.

6. The cross-modal information retrieval method based on a knowledge graph according to claim 1, wherein The joint feature vector includes image local features and text local features. The method for obtaining a matching score for characterizing the semantic correlation degree between the user query content and the candidate image-text data based on the joint feature vector includes: Perform linear mapping on the image local features and the text local features respectively to obtain corresponding image mapping features and text mapping features; Perform normalization processing on the image mapping feature and the text mapping feature respectively to obtain corresponding image attention weights and text attention weights; Multiply the image attention weight by the image mapping feature to generate an image global feature; Multiply the text attention weight by the text mapping feature to generate a text global feature; Add the image global feature and the text global feature to generate the matching score.

7. The cross-modal information retrieval method based on a knowledge graph according to claim 2, wherein The method of fusing the image training feature vector, the text training feature vector and the node semantic set based on the node fusion mechanism to generate a fusion node vector for representing multimodal information includes: Input the image training feature vector and the text training feature vector into corresponding projection matrices respectively according to the node semantic set to generate an image-side feature representation and a text-side feature representation; Input the image-side feature representation and the text-side feature representation into corresponding multi-layer perceptrons respectively to generate an image-side gating factor and a text-side gating factor; Adjust the image-side feature representation by using the image-side gating factor, and adjust the text-side feature representation by using the text-side gating factor to generate an image gated feature and a text gated feature respectively; Concatenate the image gated feature and the text gated feature and input them into a hybrid projection matrix to generate the fusion node vector.

8. A retrieval device based on a knowledge graph, characterized in that, It includes: A node retrieval unit, configured to obtain a knowledge graph, output a node relationship vector by using a preset relationship learning module, and retrieve nodes related to the user query content from the knowledge graph and generate an external knowledge representation result; A feature fusion unit, configured to fuse the user query content and the external knowledge representation result by using a correlation matrix to obtain a fused feature vector; A semantic alignment unit, configured to perform two-way image-text semantic alignment on the fused feature vector based on an attention mechanism to generate a joint feature vector for cross-modal retrieval; A feature matching unit, configured to perform a matching calculation based on the joint feature vector to obtain a matching score for characterizing the semantic correlation degree between the user query content and the candidate image-text data for cross-modal retrieval.

9. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the knowledge graph-based retrieval method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements the knowledge graph-based retrieval method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-modal knowledge graph entity alignment method and system based on inter-modal interaction

    CN116932770A

  • Information retrieval method and device, electronic equipment and storage medium

    CN119513334A

  • Drug repositioning method and system fusing multi-source knowledge graph

    WO2024138803A1

Cited By

  • Semantic enhancement and dynamic completion method and system for power market data graph

    CN120409495A

  • A semantic enhancement and dynamic completion method and system for power market data graph

    CN120409495B

  • Multi-modal knowledge graph construction method in cross-media retrieval

    CN120611774A

  • Fabric cross-modal image-text retrieval method based on knowledge graph and storage medium

    CN120763355A

  • Smart community metadata interaction method and system based on edge computing framework

    CN120803756A