Target key part identification method based on knowledge graph reasoning

By constructing a dynamic knowledge graph and feature fusion network, the problem of insufficient semantic understanding and reasoning ability in the identification of key parts of the target is solved, and high-precision recognition and rapid adaptation in complex scenarios are achieved.

CN120599210APending Publication Date: 2025-09-05BEIJING INST OF CONTROL & ELECTRONICS TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510340808.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing technologies have problems in identifying key parts of targets, such as insufficient semantic understanding, limited reasoning capabilities, high data requirements, and insufficient ability to dynamically update knowledge graphs. In particular, the performance decreases significantly in complex scenarios and small sample scenarios.

Method used

A knowledge graph-based method is adopted to construct a dynamic knowledge graph through graph neural network and attention mechanism, feature fusion is performed by combining visual features and semantic features, and key parts of the target are identified using multimodal fusion network and cross-modal attention mechanism.

Benefits of technology

It significantly improves recognition accuracy in complex scenarios, reduces dependence on large-scale labeled data, enhances the ability to reason and locate key parts of the target, and can quickly adapt to scene changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599210A_ABST
    Figure CN120599210A_ABST
Patent Text Reader

Abstract

The invention discloses a target key part identification method based on knowledge graph reasoning. The method comprises the following steps: obtaining visual features from an image; constructing a knowledge graph, and embedding entities and relationships in the knowledge graph into a continuous vector space based on a graph neural network to obtain knowledge reasoning features; performing feature fusion on the knowledge reasoning features and the visual features to obtain fusion features; and obtaining a detection result by using a detector based on the fusion feature. According to the target key part identification method based on knowledge graph reasoning disclosed by the invention, the identification precision in a complex scene is remarkably improved, and the reasoning and positioning capabilities of the target key part are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for identifying key parts of a target based on knowledge graph reasoning, and belongs to the technical field of image recognition. Background Art

[0002] Identifying key parts of an object is an important research area in computer vision, with applications in a variety of scenarios, including medical image analysis, industrial inspection, and intelligent security. Traditional methods rely primarily on deep learning models, using feature extraction and classification networks to identify key parts. However, these methods have significant shortcomings, such as:

[0003] 1. Insufficient semantic understanding: Traditional methods tend to perform object detection and location directly based on pixel-level or region-level features, lacking a deep semantic understanding of the object and its key parts.

[0004] 2. Limited reasoning capabilities: Existing methods often lack reasoning in their detection results and are unable to incorporate contextual information or prior knowledge to compensate for missing visual features.

[0005] 3. High data requirements: Deep learning models are highly dependent on large-scale labeled data, and their performance degrades significantly when faced with small samples or data-scarce scenarios.

[0006] Knowledge-based object detection methods have attracted much attention in recent years. By introducing knowledge graphs (KGs) and related reasoning techniques, they establish connections between visual data and semantic knowledge, significantly improving the performance of object detection and location localization. The main developments are as follows:

[0007] 1. Object detection based on semantic relationships: Most methods attempt to model the semantic relationships between objects as knowledge graphs. For example, Chen et al. proposed the Scene Graph Generation technique, which converts objects in a scene and their semantic relationships into a structured graph model for object detection and region segmentation.

[0008] 2. Knowledge-guided feature enhancement: Some studies use knowledge graphs to guide the feature extraction process of deep learning networks. For example, Wu et al. proposed a method to enhance target feature representation using a graph convolutional network (GCN). By combining category semantic embedding with visual features, they improved the detection performance of small samples and blurred objects.

[0009] 3. Context-aware semantic reasoning: Zhou et al. enhanced scene understanding capabilities by building a dynamic knowledge graph and combining graph neural networks with attention mechanisms. In particular, in complex scenes, the reasoning process captures the semantic relationship between the object and the context, significantly improving the robustness of object detection.

[0010] 4. Application in cross-modal tasks: In image-text cross-modal tasks, knowledge graphs are used to bridge image content and text descriptions. For example, Anderson et al. proposed the Bottom-Up and Top-Down Attention mechanism, which combines knowledge graphs with visual features for visual question answering and image annotation tasks.

[0011] Although the above methods have made some progress in the field of target detection, research on the recognition of key parts of targets is still insufficient, especially:

[0012] 1. Fine-grained key part recognition: Most existing methods focus on overall target detection while ignoring the precise positioning of key parts within the target.

[0013] 2. Insufficient reasoning complexity: Although some studies have attempted to introduce knowledge reasoning modules, the ability to reason about complex structural targets in a fine-grained manner is still limited.

[0014] 3. Insufficient dynamic updating capabilities of knowledge graphs: Existing methods are usually based on static knowledge graphs and are difficult to adapt to new scenarios or new goals.

[0015] Therefore, it is necessary to conduct more in-depth research on existing target detection methods to solve the above problems. Summary of the Invention

[0016] To overcome the above problems, we conducted in-depth research and proposed a method for identifying key parts of a target based on knowledge graph reasoning, which includes the following steps:

[0017] S1. Obtain visual features from images;

[0018] S2. Build a knowledge graph and embed the entities and relationships in the knowledge graph into a continuous vector space based on a graph neural network to obtain knowledge reasoning features.

[0019] S3, fusing the knowledge reasoning features and the visual features to obtain fused features;

[0020] S4. Based on the fusion features, a detector is used to obtain the detection results.

[0021] In a preferred embodiment, in the knowledge graph, entity appearance and entity relationships are obtained through visual features, and entity names and entity attributes are obtained through semantic features.

[0022] In a preferred embodiment, a language model is used to generate embedded representations of the target and key parts from text descriptions and tags to obtain the semantic features.

[0023] In a preferred embodiment, an attention mechanism is provided in the graph neural network, through which edge weights are dynamically adjusted according to the semantic and contextual importance between nodes.

[0024] In a preferred embodiment, the graph neural network outputs a correlation weight matrix between the target and the part as a knowledge reasoning feature.

[0025] In a preferred embodiment, a dynamic knowledge update mechanism is set up to continuously optimize the structure and content of the knowledge graph through online learning methods.

[0026] In a preferred embodiment, in S3, a multimodal fusion network is used to fuse knowledge reasoning features and visual features.

[0027] In a preferred embodiment, a cross-modal attention mechanism is used to align knowledge reasoning features and visual features.

[0028] The present invention also provides an electronic device, comprising:

[0029] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any one of the above methods.

[0030] The present invention also provides a computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute any one of the above methods.

[0031] The beneficial effects of the present invention include:

[0032] (1) Strong adaptability to complex scenarios: By modeling semantic associations and contextual relationships through knowledge graphs, the recognition accuracy in complex scenarios is significantly improved;

[0033] (2) Low data dependency: The introduction of knowledge graphs reduces the reliance on large-scale annotated data and is particularly suitable for scenarios with sparse data.

[0034] (3) Enhanced reasoning ability: Combining graph neural networks and attention mechanisms improves the ability to reason and locate key parts of the target;

[0035] (4) Dynamic update capability: By introducing a dynamic update mechanism, the model can quickly adapt to scene changes and new task requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 The figure is a flow chart of a target key part identification method based on knowledge graph reasoning according to a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0037] The present invention will be described in further detail below with reference to the accompanying drawings and examples, through which the features and advantages of the present invention will become more clearly understood.

[0038] The word "exemplary" is used exclusively herein to mean "serving as an example, example, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.

[0039] According to the present invention, a target key part identification method based on knowledge graph reasoning is provided. Figure 1 As shown, the following steps are included:

[0040] S1. Obtain visual features from images;

[0041] S2. Build a knowledge graph and embed the entities and relationships in the knowledge graph into a continuous vector space based on a graph neural network to obtain knowledge reasoning features.

[0042] S3, fusing the knowledge reasoning features and the visual features to obtain fused features;

[0043] S4. Based on the fusion features, a detector is used to obtain the detection results.

[0044] In the present invention, there is no limitation on the order of steps S1 and S2. S1 can be performed first and then S2, or S2 can be performed first and then S1, or S1 and S2 can be performed in parallel.

[0045] In S1, the visual features are extracted from the target image.

[0046] Preferably, a convolutional neural network is used to extract multi-scale features from the target image, including position, texture, and color information, and the multi-scale features are expressed in the form of a tensor to obtain the visual features.

[0047] In the present invention, the specific structure of the convolutional neural network is not limited, and those skilled in the art can freely choose according to actual needs, such as using ResNet, EfficientNet, etc.

[0048] In S2, the knowledge graph is a structured knowledge representation method that represents entities and the relationships between them in the form of a graph.

[0049] Preferably, the knowledge graph includes:

[0050] Entity: represents the target and its different parts;

[0051] Relationship: describes various associations between entities, including belonging (parts of the target belong to the target) and adjacency (spatial relationships between target parts).

[0052] Attributes: Detailed information between each entity, including the relative position and function of the entities.

[0053] By expressing image information in the form of knowledge graphs, we can achieve a deeper understanding of image content. Knowledge graphs can reveal implicit relationships between entities in images and promote overall task association analysis. In addition, this expression method can provide us with an intuitive way to explore and analyze image datasets and enhance our understanding of image data.

[0054] According to the present invention, the structured information in the knowledge graph is obtained based on the target images and their descriptions and labels in the training set.

[0055] In the knowledge graph, entity appearance and entity relationships are obtained through visual features, and entity names and entity attributes are obtained through semantic features.

[0056] Furthermore, the semantic features are extracted from text descriptions and tags.

[0057] Preferably, a language model is used to generate embedded representations of the target and key parts from text descriptions and labels to obtain the semantic features.

[0058] In the present invention, the specific structure of the language model is not limited, and those skilled in the art can freely choose it according to actual needs, such as using BERT, GPT, etc.

[0059] According to the present invention, in the knowledge graph, the target and its parts are used as nodes of the knowledge graph, the node attributes include visual features and semantic feature embedding, and the nodes are semantic relationship edges, such as "the wrist belongs to the arm" or "the elbow joint connects the upper arm and forearm".

[0060] Graph Neural Networks (GNN) is a type of neural network that specializes in processing graph-structured data. Unlike traditional neural networks that process data with regular shapes (such as images or text), GNNs can operate directly on graphs, processing the features of nodes in the graph and the connections between them.

[0061] In a preferred embodiment, the knowledge graph is embedded using a graph convolutional network (GCN) or a graph attention network (GAT) in a graph neural network.

[0062] In a preferred embodiment, an attention mechanism is provided in the graph neural network, which dynamically adjusts the edge weights according to the semantic and contextual importance between nodes, filters irrelevant information, and enhances the significance of key nodes.

[0063] In the present invention, the specific structure of the attention mechanism is not limited, and those skilled in the art may adopt any existing known attention mechanism, such as multi-head attention or self-attention mechanism.

[0064] Furthermore, the graph neural network outputs a correlation weight matrix between the target and the part as a knowledge reasoning feature, which represents the importance of the part.

[0065] In a preferred embodiment, a dynamic knowledge update mechanism is also provided to continuously optimize the structure and content of the knowledge graph through online learning methods.

[0066] Specifically, a new target image and its description or label are used to construct structured information. If the structured information is not a node or relationship in the knowledge graph, the knowledge graph is updated to achieve dynamic knowledge update.

[0067] The updating of the knowledge graph includes:

[0068] Node and relationship update: If the newly obtained structured information contains new nodes, add them to the knowledge graph and initialize the knowledge reasoning features; if the weights of the node relationships in the newly obtained structured information are different from those in the knowledge graph, adjust the weight values ​​in the knowledge graph or add new edges;

[0069] Redundant node removal: Low-confidence or redundant nodes and edges are removed by setting thresholds to maintain the efficiency and accuracy of the graph.

[0070] In S3, a multimodal fusion network is used to fuse knowledge reasoning features and visual features.

[0071] Multimodal fusion networks are widely used in image and text generation and understanding, speech and vision, autonomous driving, human-computer interaction and many other aspects.

[0072] In the present invention, the specific structure of the multimodal fusion network is not limited, and those skilled in the art can freely select it according to actual needs, for example, using a feature pyramid network (FPT).

[0073] In the present invention, by fusing knowledge reasoning features and visual features, the key part recognition accuracy can be further improved, and the reasoning and positioning capabilities of the target key parts can be enhanced.

[0074] Preferably, a cross-modal attention mechanism is used to align knowledge reasoning features with visual features. By introducing the attention mechanism, the recognition effect of key areas is enhanced.

[0075] In S4, the detector may adopt any existing detection network, such as a classification network or a regression network.

[0076] Furthermore, the output of the detector includes the coordinates, categories and confidence levels of key parts.

[0077] According to the present invention, based on the fusion features, key parts can be detected, thereby improving the positioning accuracy and robustness of the key parts.

[0078] According to the present invention, similar to traditional visual recognition, in the training phase, steps S1 to S4 are executed to train the model parameters; in the process of recognizing a new image, steps S1, S3, and S4 are executed to obtain the final recognition result.

[0079] Various embodiments of the methods described above in the present invention may be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0080] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0081] Example

[0082] Example 1

[0083] The target recognition experiment using a certain dataset includes the following steps:

[0084] S1. Obtain visual features from images;

[0085] S2. Build a knowledge graph and embed the entities and relationships in the knowledge graph into a continuous vector space based on a graph neural network to obtain knowledge reasoning features.

[0086] S3, fusing the knowledge reasoning features and the visual features to obtain fused features;

[0087] S4. Based on the fusion features, a detector is used to obtain the detection results.

[0088] In S1, a convolutional neural network ResNet is used to extract multi-scale features from the target image, including position, texture, and color information, and the multi-scale feature tensor form is expressed to obtain the visual features.

[0089] In S2, the knowledge graph includes:

[0090] Entity: represents the target and its different parts;

[0091] Relationship: describes various associations between entities, including belonging (parts of the target belong to the target) and adjacency (spatial relationships between target parts).

[0092] Attributes: Detailed information between each entity, including the relative position and function of the entities.

[0093] In the knowledge graph, entity appearance and entity relationships are obtained through visual features, and entity names and entity attributes are obtained through semantic features.

[0094] The language model GPT is used to generate embedded representations of targets and key parts from text descriptions and labels to obtain the semantic features.

[0095] The graph convolutional network (GCN) in the graph neural network is used to embed the knowledge graph. A multi-head attention mechanism is set up in the graph neural network, and a dynamic knowledge update mechanism is also set up. Through online learning methods, the structure and content of the knowledge graph are continuously optimized.

[0096] In S3, a multimodal fusion network is used to fuse knowledge reasoning features and visual features, and the multimodal fusion network adopts a feature pyramid network (FPT).

[0097] In S4, the detector adopts a classification network, and the output includes the coordinates, categories and confidence levels of key parts.

[0098] Comparative Example

[0099] Comparative Example 1

[0100] The same data set as in Example 1 was used to conduct the same experiment as in Example 1, except that S2 was embedded using the TransE model. The TransE model is described in Bordes A, Usunier N, Garcia-Duran A, et al. Translating Embeddings for Modeling Multi-relational Data[J]. Curran Associates Inc. 2013.

[0101] The results of Example 1 and Comparative Example 1 are shown in Table 1.

[0102]

[0103]

[0104] As can be seen from Table 1, the method in Example 1 significantly improves the recognition accuracy in complex scenarios, especially the ability to reason and locate key parts of the target.

[0105] The present invention has been described above with reference to preferred embodiments, but these embodiments are merely exemplary and serve only as illustrations. On this basis, various replacements and improvements can be made to the present invention, all of which fall within the scope of protection of the present invention.

Claims

1. A method for identifying key parts of a target based on knowledge graph reasoning, characterized in that: The following steps are involved: Obtain visual features from images; Build a knowledge graph and embed the entities and relationships in the knowledge graph into a continuous vector space based on a graph neural network to obtain knowledge reasoning features; Fuse knowledge reasoning features and visual features to obtain fused features; Based on the fused features, a detector is used to obtain the detection results.

2. The target key part identification method based on knowledge graph reasoning according to claim 1 is characterized in that: In the knowledge graph, entity appearance and entity relationships are obtained through visual features, and entity names and entity attributes are obtained through semantic features.

3. The target key part identification method based on knowledge graph reasoning according to claim 1 is characterized in that: A language model is used to generate embedded representations of targets and key parts from text descriptions and labels to obtain the semantic features.

4. The target key part identification method based on knowledge graph reasoning according to claim 1 is characterized in that: An attention mechanism is set up in the graph neural network, which dynamically adjusts the edge weights according to the semantic and contextual importance between nodes.

5. The target key part identification method based on knowledge graph reasoning according to claim 1 is characterized in that: The graph neural network outputs a correlation weight matrix between the target and the part as a knowledge reasoning feature.

6. The target key part identification method based on knowledge graph reasoning according to claim 1 is characterized in that: Set up a dynamic knowledge update mechanism and continuously optimize the structure and content of the knowledge graph through online learning methods.

7. The target key part identification method based on knowledge graph reasoning according to claim 1 is characterized in that: A multimodal fusion network is used to fuse knowledge reasoning features and visual features.

8. The target key part identification method based on knowledge graph reasoning according to claim 7 is characterized in that: A cross-modal attention mechanism is used to align knowledge reasoning features and visual features.

9. An electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1 to 8.

10. A computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.