Target detection method, system and device based on knowledge graph and medium

Through a knowledge graph-based method, combined with multi-level visual feature extraction and knowledge graph information fusion, the problem of low object detection efficiency and poor accuracy in complex scenarios is solved, and more efficient and accurate object detection is achieved.

CN120388219APending Publication Date: 2025-07-29GUANGZHOU POWER SUPPLY BUREAU GUANGDONG POWER GRID CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510460146.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Existing deep learning models cannot accurately extract local features in complex scenarios, ignoring the semantic relationship between objects, resulting in low object detection efficiency and poor accuracy.

Method used

Using a knowledge graph-based method, multi-level visual feature extraction and knowledge graph information fusion are improved to improve the accuracy and efficiency of object detection.

Benefits of technology

Improve the accuracy and efficiency of object detection in complex environments and enhance understanding of the relationship between objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388219A_ABST
    Figure CN120388219A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection method, system and device based on a knowledge graph and a medium, and belongs to the technical field of computer vision, and the method comprises the steps: carrying out the multi-level visual feature extraction of an obtained to-be-recognized image, so as to extract a plurality of candidate regions from an obtained visual feature map, selecting a plurality of target detection areas from the plurality of candidate areas; obtaining a knowledge graph corresponding to the to-be-recognized image, and extracting graph information of the knowledge graph; positioning a region visual feature of each target detection region from the visual feature map, and fusing the region visual features with the map information to obtain a vector representation corresponding to each target detection region; and performing target category prediction on each target detection area according to the vector representation and a preset target detection model to obtain a target detection result of the to-be-recognized image. Therefore, by implementing the method and the device, the problem that target detection cannot be performed timely and accurately in the prior art can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular, to an object detection method, system, device and medium based on a knowledge graph. Background Art

[0002] In the current field of visual detection, significant progress has been made in object detection and image classification using deep learning models. Among them, technologies based on convolutional neural networks and Faster R-CNN (Region-based Convolutional Neural Network) can effectively extract features from images for image classification and object detection, and are widely used in multiple fields such as autonomous driving, security monitoring, and medical image analysis.

[0003] However, when using existing deep learning models for object detection, they usually rely on a large amount of labeled data to train the model so that the model learns local features in the image for detection. As a result, when using the model to process complex scenes, blurred images, object occlusion, or difficult-to-distinguish objects, the local features cannot be accurately extracted, thereby resulting in low detection accuracy.

[0004] In addition, most existing visual detection models only rely on visual information in the image for reasoning, lack an understanding of the relationships between objects in multi-object detection, and ignore the context information and semantic relationships between objects. They can usually only detect and classify single objects and cannot perform object detection in complex environments, reducing the efficiency and accuracy of object detection. Summary of the Invention

[0005] The present invention provides an object detection method, system, device and medium based on a knowledge graph, which can perform object detection on complex environments in an image according to the visual features of the image and the information of the knowledge graph, improve the efficiency and accuracy of object detection, and thus solve the problem in the prior art that object detection cannot be performed in complex environments, resulting in low efficiency and poor accuracy of object detection.

[0006] To achieve the above object, in a first aspect, the present invention discloses an object detection method based on a knowledge graph, including:

[0007] Performing multi-level visual feature extraction on the acquired image to be recognized to obtain a visual feature map corresponding to the image to be recognized;

[0008] Extracting a plurality of candidate regions from the visual feature map, and calculating an interest score corresponding to each candidate region, so as to select a plurality of object detection regions from the plurality of candidate regions according to the interest score;

[0009] Obtain the knowledge graph corresponding to the image to be recognized, and extract the graph information of the knowledge graph;

[0010] Locate the regional visual features of each target detection region in the visual feature map, and fuse the regional visual features with the graph information to obtain the vector representation corresponding to each target detection region;

[0011] Perform target category prediction on each target detection region according to the vector representation and a preset target detection model to obtain the target detection result of the image to be recognized.

[0012] A target detection method based on a knowledge graph disclosed by the present invention first extracts multi-level visual features of an image to improve the accuracy of the finally generated visual feature map, providing a more accurate feature basis for subsequent target detection. Then, multiple regions of interest are extracted on the visual feature map to perform target detection on the regions of interest, thereby improving the efficiency and accuracy of target detection. Secondly, obtain the knowledge graph corresponding to the current image to be recognized, and extract the graph information of the knowledge graph to fuse the visual features of each region of interest with the graph information of the knowledge graph, and then enhance the understanding of the relationships between various objects in the image according to the semantic information in the knowledge graph, thereby improving the accuracy of target detection in a complex environment.

[0013] As a preferred example, the extracting multi-level visual features from the obtained image to be recognized to obtain the visual feature map corresponding to the image to be recognized includes:

[0014] Input the obtained image to be recognized into a pre-constructed convolutional neural network, and perform multi-layer convolution and pooling operations on the image to be recognized through the convolutional neural network to obtain the visual feature map corresponding to the image to be recognized.

[0015] In the above solution, visual features of the image are extracted through a convolutional neural network to improve the efficiency and accuracy of feature extraction. Among them, during the feature extraction process, performing multi-layer convolution and pooling operations on the input image can effectively extract different levels of visual features from the image. Furthermore, when performing target detection through the visual feature map containing different levels of visual features, the efficiency and accuracy of target detection can be improved.

[0016] As a preferred example, the extracting multiple candidate regions from the visual feature map and calculating the interest score corresponding to each candidate region to select several target detection regions from the multiple candidate regions according to the interest score includes:

[0017] Input the visual feature map into a pre - constructed region proposal network to perform object detection on the visual feature map through a sliding window mechanism and anchor boxes preset in the region proposal network, and obtain multiple candidate regions;

[0018] For any of the candidate regions, calculate the object classification score and the intersection - over - union score of the candidate region, and use the sum of the object classification score and the intersection - over - union score as the interest score of the candidate region;

[0019] When the interest score is greater than or equal to a preset score threshold, determine that the candidate region is an object detection region.

[0020] In the above solution, in order to improve the efficiency and accuracy of object detection, a region proposal network is used to extract regions of interest from the generated visual feature map for object detection, rather than directly segmenting the original image, so as to optimize the accuracy and computational efficiency of object detection.

[0021] As a preferred example, calculating the interest score corresponding to each candidate region includes:

[0022] For any of the candidate regions;

[0023] Calculate the object classification score corresponding to the candidate region through a preset binary classification application network; wherein, determine whether the candidate region is the background through the object classification score;

[0024] Calculate the intersection area between the candidate region and the anchor box and the union area between the candidate region and the anchor box, and use the ratio of the intersection area to the union area as the intersection - over - union score.

[0025] In the above solution, in order to improve the efficiency of object detection, calculate the object classification score corresponding to the candidate region to determine whether the candidate region is the background, thereby narrowing the scope of object detection. And calculate the intersection area and union area between the candidate region and the anchor box to score the confidence of the candidate region, and then include the candidate regions with high confidence as regions of interest to improve the accuracy of object detection.

[0026] As a preferred example, obtaining the knowledge graph corresponding to the image to be recognized and extracting the graph information of the knowledge graph includes:

[0027] Match the large - scale knowledge base corresponding to the image to be recognized;

[0028] Construct the knowledge graph corresponding to the image to be recognized based on predefined entities, attributes, and relationships between entities; wherein, the entity is the target category;

[0029] Input the knowledge graph into a pre-constructed graph convolutional network to extract the semantic features of each node in the knowledge graph through the graph convolutional network, and use the semantic features as the graph information of the knowledge graph.

[0030] In the above solution, by matching a large-scale knowledge base corresponding to the image to be recognized and constructing a knowledge graph corresponding to the image to be recognized, accurate judgment basis can be provided for object detection by using the semantic information contained in the knowledge graph. Among them, when extracting the semantic features of each node in the knowledge graph, using a graph convolutional neural network can improve the accuracy and efficiency of feature extraction, and thus improve the efficiency and accuracy of object detection.

[0031] As a preferred example, obtaining the knowledge graph corresponding to the image to be recognized further includes:

[0032] Design a dynamic update mechanism corresponding to the knowledge graph; wherein, the dynamic update mechanism includes new entity addition, relationship update, and conflict resolution;

[0033] Obtain the object detection result and update the knowledge graph in real time based on the dynamic update mechanism.

[0034] In the above solution, by designing a dynamic update mechanism for the knowledge graph, the changes in the historical image data corresponding to the image data to be recognized can be responded to in a timely manner, thereby improving the accuracy of the knowledge graph and improving the accuracy of object detection.

[0035] As a preferred example, locating the regional visual features of each object detection region in the visual feature map and fusing the regional visual features with the graph information to obtain the vector representation corresponding to each object detection region respectively includes:

[0036] For any object detection region, extract the regional visual features of the object detection region from the visual feature map according to the bounding box corresponding to the object detection region;

[0037] Map the regional visual features and the semantic features to a fusion space according to a preset feature alignment mechanism;

[0038] Calculate the visual weight of the regional visual features and the semantic weight of the semantic features after the mapping operation according to a preset attention mechanism;

[0039] Fuse the regional visual features and the semantic weights based on the visual weight, the semantic weight, and a preset cross-modal attention mechanism to generate the vector representation corresponding to each object detection region respectively.

[0040] In the above solution, after extracting the visual features of the image and the semantic features of the knowledge graph, this solution deeply fuses the features of the knowledge graph with the visual features of the image through an attention mechanism to improve the accuracy of object detection and reasoning.

[0041] In a second aspect, the present invention discloses an object detection system based on a knowledge graph, including a visual feature module, a graph screening module, a semantic feature module, a feature fusion module, and an object detection module;

[0042] The visual feature module is used to perform multi-level visual feature extraction on the acquired image to be recognized, and obtain a visual feature map corresponding to the image to be recognized;

[0043] The graph screening module is used to extract multiple candidate regions from the visual feature map, and calculate the interest score corresponding to each candidate region, so as to select several object detection regions from the multiple candidate regions according to the interest score;

[0044] The semantic feature module is used to obtain the knowledge graph corresponding to the image to be recognized, and extract the graph information of the knowledge graph;

[0045] The feature fusion module is used to locate the regional visual feature where each object detection region is located in the visual feature map, and fuse the regional visual feature with the graph information to obtain a vector representation corresponding to each object detection region;

[0046] The object detection module is used to perform object category prediction on each object detection region according to the vector representation and a preset object detection model, and obtain the object detection result of the image to be recognized.

[0047] An object detection system based on a knowledge graph disclosed by the present invention first extracts multi-level visual features of an image to improve the accuracy of the finally generated visual feature map, provides a more accurate feature basis for subsequent object detection, then extracts multiple regions of interest on the visual feature map to perform object detection on the regions of interest, thereby improving the efficiency and accuracy of object detection. Secondly, it obtains the knowledge graph corresponding to the current image to be recognized, and extracts the graph information of the knowledge graph, so as to fuse the visual features of each region of interest with the graph information of the knowledge graph, and then enhance the understanding of the relationships between various objects in the image according to the semantic information in the knowledge graph, thereby improving the accuracy of object detection in a complex environment.

[0048] In a third aspect, the present invention discloses a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, a target detection method based on a knowledge graph as described in the first aspect is implemented.

[0049] In a fourth aspect, the present invention discloses a computer-readable storage medium, including: a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute a target detection method based on a knowledge graph as described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] To more clearly illustrate the technical solutions of the present application, the drawings required for implementation will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0051] Figure 1 It is a flowchart of a target detection method based on a knowledge graph provided by an embodiment of the present invention;

[0052] Figure 2 It is a schematic structural diagram of a target detection system based on a knowledge graph provided by an embodiment of the present invention;

[0053] Figure 3 It is a flowchart of a target detection method based on a knowledge graph provided by another embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the present application; the terms "including" and "having" and any variations thereof in the specification and claims of the present application and the above drawings are intended to cover non-exclusive inclusion.

[0056] In the description of the embodiments of the present application, technical terms such as "first" and "second" are only used to distinguish different objects, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity, specific order, or primary-secondary relationship of the indicated technical features. In the description of the embodiments of the present application, the meaning of "a plurality" is more than two, unless otherwise clearly and specifically defined.

[0057] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0058] In the description of the embodiments of the present application, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this text generally represents an "or" relationship between the associated objects before and after.

[0059] In the description of the embodiments of the present application, the term "a plurality" refers to more than two (including two). Similarly, "a plurality of groups" refers to more than two groups (including two groups), and "a plurality of pieces" refers to more than two pieces (including two pieces).

[0060] In the description of the embodiments of the present application, unless otherwise clearly specified and limited, technical terms such as "installation", "connection", "coupling", "fixation", etc. should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can also be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present application can be understood according to specific circumstances.

[0061] Embodiment 1

[0062] See Figure 1 , to solve the problem that it is impossible to detect targets in a complex environment in a timely and accurate manner in the prior art, this embodiment provides a target detection method based on a knowledge graph to improve the efficiency and accuracy of target detection. Among them, the method includes:

[0063] Step 101: Extract multi-level visual features from the obtained image to be recognized to obtain a visual feature map corresponding to the image to be recognized.

[0064] In this embodiment, this step mainly includes: inputting the acquired image to be recognized into a pre-constructed convolutional neural network, and performing multi-layer convolution and pooling operations on the image to be recognized through the convolutional neural network to obtain a visual feature map corresponding to the image to be recognized.

[0065] The above step extracts visual features from the image through a convolutional neural network to improve the efficiency and accuracy of feature extraction. Among them, in the process of feature extraction, performing multi-layer convolution and pooling operations on the input image can effectively extract visual features at different levels from the image. Furthermore, when performing object detection through the visual feature map containing visual features at different levels, the efficiency and accuracy of object detection can be improved.

[0066] Step 102: Extract multiple candidate regions from the visual feature map, and calculate the interest score corresponding to each candidate region, so as to select several object detection regions from multiple candidate regions according to the interest score.

[0067] In this embodiment, this step mainly includes: inputting the visual feature map into a pre-constructed region proposal network, and performing object detection on the visual feature map through a sliding window mechanism and anchor boxes preset in the region proposal network to obtain multiple candidate regions; for any candidate region, calculate the object classification score and the intersection ratio score of the candidate region, and take the sum of the object classification score and the intersection ratio score as the interest score of the candidate region; when the interest score is greater than or equal to a preset score threshold, determine that the candidate region is an object detection region.

[0068] Among them, for any candidate region; calculate the object classification score corresponding to the candidate region through a preset binary classification application network; among them, determine whether the candidate region is a background through the object classification score; calculate the intersection area between the candidate region and the anchor box and the union area between the candidate region and the anchor box, and take the ratio of the intersection area to the union area as the intersection ratio score.

[0069] In the above step, in order to improve the efficiency and accuracy of object detection, a region proposal network is used to extract regions of interest from the generated visual feature map for object detection, rather than directly segmenting the original image, so as to optimize the accuracy and computational efficiency of object detection. Among them, in order to improve the efficiency of object detection, calculate the object classification score corresponding to the candidate region to determine whether the candidate region is a background, thereby narrowing the scope of object detection. And calculate the intersection area and union area between the candidate region and the anchor box to score the confidence of the candidate region, and then include the candidate region with high confidence as the region of interest to improve the accuracy of object detection.

[0070] Step 103: Obtain the knowledge graph corresponding to the image to be recognized, and extract the graph information of the knowledge graph.

[0071] In this embodiment, this step mainly includes: matching the large-scale knowledge base corresponding to the image to be recognized; constructing the knowledge graph corresponding to the image to be recognized based on predefined entities, attributes, and relationships between entities; where the entity is the target category; inputting the knowledge graph into a pre-constructed graph convolutional network to extract the semantic features of each node in the knowledge graph through the graph convolutional network, and using the semantic features as the graph information of the knowledge graph.

[0072] Among them, design a dynamic update mechanism corresponding to the knowledge graph; where the dynamic update mechanism includes new entity addition, relationship update, and conflict resolution; obtain the target detection result, and update the knowledge graph in real time based on the dynamic update mechanism.

[0073] In the above steps, by matching the large-scale knowledge base corresponding to the image to be recognized and constructing the knowledge graph corresponding to the image to be recognized, the semantic information contained in the knowledge graph is used to provide an accurate judgment basis for target detection. Among them, when extracting the semantic features of each node in the knowledge graph, using a graph convolutional neural network can improve the accuracy and efficiency of feature extraction, thereby improving the efficiency and accuracy of target detection. And by designing a dynamic update mechanism for the knowledge graph, it can respond in a timely manner to the changes in the historical image data corresponding to the image data to be recognized, thereby improving the accuracy of the knowledge graph to improve the accuracy of target detection.

[0074] Step 104: Locate the regional visual features of each target detection region in the visual feature map, and fuse the regional visual features and the graph information to obtain a vector representation corresponding to each target detection region.

[0075] In this embodiment, this step is mainly: for any target detection region, extract the regional visual features of the target detection region from the visual feature map according to the bounding box corresponding to the target detection region; map the regional visual features and the semantic features to a fusion space according to a preset feature alignment mechanism; calculate the visual weight of the regional visual features and the semantic weight of the semantic features after the mapping operation according to a preset attention mechanism; fuse the regional visual features and the semantic weight based on the visual weight, the semantic weight, and a preset cross-modal attention mechanism to generate a vector representation corresponding to each target detection region.

[0076] After extracting the visual features of the image and the semantic features of the knowledge graph in the above steps, the present solution deeply fuses the features of the knowledge graph with the visual features of the image through an attention mechanism to improve the accuracy of object detection and reasoning.

[0077] Step 105: Perform object category prediction on each of the object detection regions according to the vector representation and a preset object detection model to obtain the object detection result of the image to be recognized.

[0078] As Figure 2 shown, on the basis of the above method item embodiments, corresponding device item embodiments are provided. Among them, the present embodiment provides an object detection system based on a knowledge graph, including a visual feature module 201, a graph screening module 202, a semantic feature module 203, a feature fusion module 204, and an object detection module 205.

[0079] The visual feature module 201 is configured to perform multi-level visual feature extraction on the acquired image to be recognized to obtain a visual feature map corresponding to the image to be recognized.

[0080] The graph screening module 202 is configured to extract a plurality of candidate regions from the visual feature map and calculate an interest score corresponding to each of the candidate regions, so as to select several object detection regions from the plurality of candidate regions according to the interest score.

[0081] The semantic feature module 203 is configured to obtain the knowledge graph corresponding to the image to be recognized and extract the graph information of the knowledge graph.

[0082] The feature fusion module 204 is configured to locate the regional visual feature where each of the object detection regions is located in the visual feature map, and fuse the regional visual feature with the graph information to obtain a vector representation corresponding to each of the object detection regions respectively.

[0083] The object detection module 205 is configured to perform object category prediction on each of the object detection regions according to the vector representation and a preset object detection model to obtain the object detection result of the image to be recognized.

[0084] It can be understood that the above device item embodiments correspond to the method item embodiments of the present invention, and can implement an object detection method based on a knowledge graph provided by any one of the above method item embodiments of the present invention.

[0085] It should be noted that the device embodiments described above are merely illustrative. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationships between the modules indicate that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0086] Based on the foregoing embodiment of a knowledge graph-based target detection method as Figure 1 shown, this embodiment provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, a knowledge graph-based target detection method of this embodiment is implemented.

[0087] Exemplarily, in this embodiment, the computer program can be divided into one or more modules. The one or more modules are stored in the memory and executed by the processor to complete this embodiment. The one or more module elements can be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program in the terminal device.

[0088] The terminal device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor and a memory.

[0089] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the terminal device, and connects various parts of the entire terminal device through various interfaces and lines.

[0090] Based on the above method embodiments, another embodiment of the present invention provides a computer-readable storage medium, including a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute a target detection method based on a knowledge graph according to any one of the above method embodiments of the present invention.

[0091] Among them, the modules / units integrated in the device / terminal device, if implemented in the form of software functional units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on such an understanding, all or part of the processes in the method of this embodiment can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0092] A target detection method, system, device and medium based on a knowledge graph provided in this embodiment first extracts multi-level visual features of an image to improve the accuracy of the finally generated visual feature map, and provides a more accurate feature basis for subsequent target detection. Then, multiple regions of interest are extracted from the visual feature map to perform target detection on the regions of interest, thereby improving the efficiency and accuracy of target detection. Secondly, a knowledge graph corresponding to the current image to be recognized is obtained, and the graph information of the knowledge graph is extracted to fuse the visual features of each region of interest with the graph information of the knowledge graph. Then, according to the semantic information in the knowledge graph, the understanding of the relationships between various objects in the image is enhanced, so as to improve the accuracy of target detection in a complex environment.

[0093] Embodiment 2

[0094] When using deep learning models for object detection in the prior art, most of them only rely on visual information in images for inference, ignoring context information and semantic relationships between objects. In complex scenarios, the relationships between objects are crucial for understanding object detection, but existing models cannot effectively extract and utilize this information, limiting their inference ability and accuracy in these scenarios. At the same time, most traditional visual detection methods only rely on local features of images for inference and cannot effectively identify the relationships between various objects in the image. In an environment with multiple objects and complex backgrounds, existing models often have difficulty correctly inferring the interactions between objects, thus limiting their effectiveness in practical applications.

[0095] Referring to Figure 3 , this embodiment provides an object detection method based on a knowledge graph to solve the technical problems of low efficiency and low accuracy of object detection caused by the inability of existing deep learning models to infer context information in images. Specifically, the object detection method includes:

[0096] Step 301: Obtain the original data corresponding to the image to be recognized to construct the knowledge graph corresponding to the image to be recognized.

[0097] In this embodiment, a large-scale knowledge base corresponding to the image to be recognized is matched, and the knowledge graph corresponding to the image to be recognized is constructed based on predefined entities, attributes, and relationships between entities. Among them, a dynamic update mechanism corresponding to the knowledge graph is designed; the dynamic update mechanism includes new entity addition, relationship update, and conflict resolution; the object detection result corresponding to each image to be recognized is obtained, and the knowledge graph is updated in real time based on the dynamic update mechanism.

[0098] Specifically, the construction of the knowledge graph involves multiple levels of objects, attributes, behaviors, and the relationships between them. Thus, in an implementation manner of this embodiment, when the object detection is applied to an image corresponding to a power grid transmission line, the process of constructing the knowledge graph of the image including the power grid transmission route is as follows:

[0099] First, define the nodes of the knowledge graph, that is, entities. Taking the power grid transmission route as an example, first find different data sources in the circuit field where the power grid transmission route is located, including but not limited to a power equipment database (for providing equipment parameters and maintenance records), remote sensing images and sensor data (for constructing spatial relationships), and manually labeled data (for verifying and optimizing the model), etc., to find data related to the power grid transmission route, and obtain entities from the data. The entities include but are not limited to: transmission lines, that is, high-voltage lines for power transmission; utility poles, that is, the basic structures supporting the transmission lines; transformers, that is, power equipment for voltage conversion; and distribution equipment, such as circuit breakers, cable joints, etc.

[0100] Secondly, define the relationships in the knowledge graph, i.e., the edges between two nodes, according to the retrieved data. Taking the power grid transmission line as an example, the relationships in the knowledge graph mainly include the following categories: (1) Physical connection relationship: The transmission line connects to the utility pole, the transformer is installed on the utility pole, and the transformer is connected to the power distribution equipment through the transmission line; (2) State influence relationship: For example, line overload may cause the transformer to heat up, and transformer aging may affect the normal operation of the power distribution equipment; (3) Behavioral relationship: For example, the transformer needs to be maintained regularly, and line faults will trigger grid alarms, etc.

[0101] In this embodiment, the knowledge graph is constructed to perform reasoning in the visual detection task, such as reducing false detections: If the model detects a suspected transformer but fails to detect the corresponding utility pole, it can be inferred based on the knowledge graph that there may be a false detection; enhancing the recognition ability: By utilizing the existing equipment distribution information in the knowledge graph, the model can automatically adjust the detection strategy to improve the detection accuracy; anomaly analysis: For example, if the knowledge graph shows that a certain area usually contains 5 utility poles, but only 3 are found in the detection result, it may indicate that there is an abnormal situation in this area.

[0102] The knowledge graph constructed by the above method can effectively improve the accuracy and robustness of object detection, enhance the reasoning ability of the visual detection model, and thus improve its adaptability and intelligent level in complex environments.

[0103] Step 302: Input the image to be recognized into a pre-constructed convolutional neural network so that the convolutional neural network outputs a visual feature map corresponding to the image to be recognized.

[0104] In this embodiment, the obtained image to be recognized is input into a pre-constructed convolutional neural network to perform multi-layer convolution and pooling operations on the image to be recognized through the convolutional neural network, and a visual feature map corresponding to the image to be recognized is obtained.

[0105] To quickly and accurately extract visual features from the image to improve the efficiency and accuracy of object detection. In one implementation manner of this embodiment, a deep convolutional neural network (Convolutional Neural Network) is used to extract visual features.

[0106] Specifically, through multi-layer convolution and pooling operations, the deep convolutional neural network can effectively extract visual features at different levels from the image and generate a visual feature map containing the visual features at different levels. Among them, the convolutional layer of the deep convolutional neural network extracts local features of the image through filters (convolution kernels) and generates a feature map of the image. The convolution operation formula of the deep convolutional neural network is as follows:

[0107]

[0108] Wherein, I is the input image, W i is the convolution kernel, * is the convolution operation, and b i is the bias term. This formula represents the visual features generated by the image through the convolution operation.

[0109] Step 303: Extract a plurality of regions of interest from the visual feature map according to the region proposal network, so as to obtain the region visual features corresponding to each of the regions of interest from the visual feature map.

[0110] In this embodiment, the visual feature map is input into a pre-constructed region proposal network to perform object detection on the visual feature map through a sliding window mechanism and anchor boxes preset in the region proposal network, and obtain a plurality of candidate regions; for any of the candidate regions, calculate the object classification score and the intersection-over-union score of the candidate region, and use the sum of the object classification score and the intersection-over-union score as the interest score of the candidate region; when the interest score is greater than or equal to a preset score threshold, determine that the candidate region is an object detection region, that is, a region of interest.

[0111] In an implementation manner of this embodiment, in order to improve the efficiency and accuracy of object detection, a region proposal network (Region Proposal Network, RPN) is used to extract regions of interest from the visual feature map to perform object detection on the regions of interest, thereby improving the accuracy and computational efficiency of object detection.

[0112] In the implementation manner, the region proposal network adopts a sliding window mechanism on the visual feature map, combines anchor boxes (Anchor Boxes) to generate a plurality of candidate regions, and screens the candidate regions by calculating the objectness score of each candidate region, and finally obtains a plurality of regions of interest, that is, object detection regions.

[0113] Preferably, when using the Region Proposal Network (RPN) to screen regions of interest, the RPN is first trained to improve the accuracy of candidate region selection. During the training process, a preset AnchorBox scoring and regression loss function is used to improve the accuracy of the bounding boxes generated by the RPN for each candidate region. The RPN generates candidate boxes through multi-scale Anchor Boxes and performs Bounding Box Regression on them to optimize the detection accuracy. During training, the Smooth L1 Loss in the regression loss function is used to calculate the deviation between the predicted box and the ground truth box. The formula for the Smooth L1 Loss function is as follows:

[0114]

[0115] Where: N reg represents the number of positive sample regions; t i is the predicted bounding box parameter; is the bounding box parameter of the ground truth target; The Smooth L1 Loss function is defined as follows:

[0116]

[0117] This method avoids the sensitivity of the L2 loss to outliers and improves the regression stability of the bounding boxes.

[0118] After obtaining multiple candidate regions using the trained RPN, the criteria for screening target detection regions from the candidate regions are to remove redundant boxes and retain the optimal detection region, i.e., the target detection region, based on object classification scoring, Intersection over Union (IoU) calculation, Non-Maximum Suppression (NMS), etc.

[0119] Among them, the method for calculating the object classification score is that in the RPN structure, the score of the candidate region is calculated by a binary classification neural network to determine whether the region contains an object (foreground) or background (irrelevant region). The object classification loss function is defined as:

[0120]

[0121] Where: N cls represents the number of candidate regions; y i represents the true category of the i-th candidate region (1 represents foreground, 0 represents background); represents the foreground probability predicted by the RPN.

[0122] Next, the RPN calculates the IoU value (i.e., the intersection over union) between the candidate region and the ground truth target, and determines whether the region has a high target confidence based on this. The IoU calculation formula is as follows:

[0123]

[0124] where: A intersection is the intersection area between the candidate region and the ground truth target; A union is the union area between the candidate region and the ground truth target. When the IoU is higher than the set threshold (e.g., 0.7), the candidate region is considered a region of interest with high confidence, otherwise it is discarded or its priority is reduced.

[0125] The region screening mechanism based on the above IoU and target classification scores helps to reduce false detections and missed detections, and improve the adaptability of the model in complex environments. In addition, the RPN relies on the high-level semantic features extracted by the convolutional neural network to ensure that the candidate regions have strong target representation capabilities, providing high-quality inputs for subsequent classification and bounding box regression.

[0126] Step 304: Input the knowledge graph into a pre-constructed graph convolutional neural network so that the graph convolutional neural network outputs the semantic features corresponding to the knowledge graph.

[0127] In this embodiment, after the knowledge graph is constructed, in order to combine the information in the knowledge graph to improve the accuracy of target detection, a pre-constructed graph convolutional network is used to extract the features of the knowledge graph.

[0128] Among them, in the process of constructing the graph convolutional network, based on each node v in the knowledge graph i represents a specific entity, such as a transmission line, a utility pole, a transformer, etc. In order to construct an effective graph convolutional network (GCN), it is first necessary to define the initial node feature matrix H (0) :

[0129]

[0130] where, The construction methods of include: Feature based on attributes: Using structured data such as device power, voltage, type, etc. as initial features. Feature based on embedding: Adopting an NLP pre-trained model to convert text descriptions into high-dimensional vectors. Feature based on vision: If the node contains image data, features can be extracted through a CNN.

[0131] Next, construct the adjacency matrix: In the knowledge graph, the connection relationship between nodes is represented by the adjacency matrix A:

[0132]

[0133] To stabilize the calculation, a normalized adjacency matrix is introduced:

[0134]

[0135] where D ii = ∑ j A ij is the degree matrix. The normalization process can prevent gradient explosion and improve the training stability of GCN.

[0136] For GCN iterative feature propagation, GCN updates node features at each layer, and the calculation formula is as follows:

[0137]

[0138] where: H (l) is the node feature of the l-th layer; W (l) is the learnable weight matrix; σ(·) is the activation function. Through multi-layer propagation, GCN aggregates neighborhood information to the central node, realizes cross-layer information fusion, and enhances the inference ability of visual detection.

[0139] The graph convolutional network constructed through the above construction process can enhance the representation of nodes in the graph spectrum by propagating the information of adjacent nodes, thereby optimizing visual features. The graph convolutional network plays a key role in the fusion of image features and knowledge graphs. Specifically, the graph convolutional network can capture the relationships between entities in the graph spectrum, propagate the information of nodes in the graph spectrum in an iterative manner, and combine visual features to form a richer feature representation.

[0140] When using the pre-constructed graph convolutional neural network to perform convolutional operations on the knowledge graph to output the semantic features of the knowledge graph, the graph convolutional operation formula in the graph convolutional neural network is:

[0141]

[0142] In the formula: H (l) represents the node feature matrix of the l-th layer, that is, the feature representation of nodes in the current layer; is the normalized adjacency matrix, representing the connection relationship between nodes; W (l) is the learned weight matrix, representing the feature transformation parameters learned by the model during training; σ is the activation function. Through the propagation of the multi-layer graph convolutional network, the representation of nodes gradually contains more context information, which helps to better fuse the knowledge graph and visual features and provides stronger inference ability for object detection.

[0143] Step 305: Through a preset attention mechanism, fuse the region visual features corresponding to each of the regions of interest with the semantic features respectively to obtain the fusion features corresponding to each of the regions of interest.

[0144] In this embodiment, the regional visual feature and the semantic feature are mapped to a fusion space according to a preset feature alignment mechanism; according to a preset attention mechanism, the visual weight of the regional visual feature and the semantic weight of the semantic feature after the mapping operation are calculated; based on the visual weight, the semantic weight, and a preset cross-modal attention mechanism, the regional visual feature and the semantic weight are fused to generate a vector representation corresponding to each of the target detection regions.

[0145] In an implementation manner of this embodiment, during the fusion of the semantic feature and the visual feature of the knowledge graph, the attention mechanism further weights the similarity between different graphs and visual features. By calculating the similarity between the graph and the visual feature, the model can adjust the fusion weight of the feature according to different importance levels, so that important information obtains a greater weight, thereby enhancing the accuracy of reasoning.

[0146] Among them, the attention mechanism is used to perform weighted fusion on the visual feature and the knowledge graph information to improve the reasoning ability of target detection. The calculation process is as follows:

[0147] Construction of the query vector. The query vector Q is formed by the weighted combination of the visual feature and the knowledge graph feature:

[0148] Q = λW Q F CNN +(1 - λ)W Q H GCN

[0149] where: F CNN is the target feature extracted by the convolutional neural network; H GCN is the knowledge graph embedding feature extracted by the graph convolutional network; W Q is a learnable projection matrix; λ controls the fusion ratio of the two features.

[0150] Calculation of the attention weight. The attention mechanism uses Softmax normalization to calculate the similarity between the query vector Q and all key vectors K i . The weight calculation formula is as follows:

[0151]

[0152] where: score(Q, K i ) is used to measure the similarity between Q and K i , and dot product attention can be adopted:

[0153] score(Q, K i ) = Q T K i

[0154] Or additive attention:

[0155] score(Q, K i ) = V T tanh(W Q Q + W K K i );

[0156] Wherein, K i represents the adjacent node features in the knowledge graph or the candidate region vector in the visual features.

[0157] Attention feature fusion, the calculated attention weights are used to weighted-fuse the knowledge graph features and visual features:

[0158]

[0159] Wherein, H fusion is input into the object detection network as the finally fused feature.

[0160] When calculating the attention weights, the index i takes values in the range of all possible candidate key vectors, that is:

[0161] In the knowledge graph scenario, n represents the number of neighbor nodes of the target node, that is:

[0162]

[0163] Wherein, represents the set of neighbor nodes of the target node v.

[0164] In the visual feature scenario, n represents the number of candidate regions in object detection, that is:

[0165]

[0166] Wherein, represents the set of Region Proposals (RoIs).

[0167] For the deep fusion of visual features and knowledge graph features, in order to achieve more effective feature fusion, the present invention further introduces Cross-Modal Attention, enabling visual features and knowledge graph features to complement each other. Under this mechanism, visual features and knowledge graph features are fused through feature alignment and weighted summation:

[0168]

[0169] Wherein: F' CNN = W f F CNN: Through the linear transformation matrix W f Enable the visual feature F extracted by the convolutional neural network (CNN) CNN Adapt to the fusion space. H′ GCN = W g H GCN : Through the linear transformation matrix W g Enable the knowledge graph feature H extracted by the graph convolutional network (GCN) GCN Adapt to the fusion space. α i : Attention weight, indicating the similarity between the query vector Q and the candidate key vector K i The contribution degree of different information sources is determined. Through weighted summation, the fused feature contains both visual information and semantic information of the graph, thereby improving the accuracy of object detection.

[0170] Step 306: Based on the fused feature and a preset object detection model, predict the object type of the image to be recognized, and obtain the object detection result of the image to be recognized.

[0171] Specifically, object detection is performed using a deep learning model through the fused feature. The loss function of this model includes classification loss and regression loss, ensuring accurate classification and localization of the object. The classification loss usually uses cross-entropy loss, and the regression loss uses smooth L1 loss. The formula for the classification loss is:[[]]

[0172]

[0173] In the formula, y i is the true class label of the object, and p i is the predicted class probability.

[0174] The formula for the regression loss is:[[]]

[0175]

[0176] In the formula, b i is the true bounding box,[[]] is the predicted bounding box, and smooth L1 is the smooth L1 loss function.

[0177] Among them, when the deep learning model performs object detection, the fusion feature that combines the semantic features of the knowledge graph can enhance the effect of visual object detection. Among them, after combining the attention mechanism, the knowledge graph features can provide effective prior information in the visual object detection task. For example: Object association enhancement: The knowledge graph can provide structural information about the objects. For example, in the power grid inspection scenario, transformers are usually installed on utility poles. If the visual detection model has a low confidence in the transformer, the detection result can be enhanced through knowledge graph relationship reasoning. Weak supervision detection: For some fuzzy, occluded, or low-resolution targets, the knowledge graph can complement the deficiencies of visual features and improve the reliability of object detection. Spatial constraint optimization: The knowledge graph can provide the spatial relationships between objects, enabling the object detection network to reduce false detections and missed detections in specific scenarios. The reasoning module uses the reasoning rules in the knowledge graph for post-processing.

[0178] Through the relationship information in the graph, the deep learning model can further infer the relationships between the targets and correct or supplement the detection boxes. During the reasoning process, by analyzing the relationships between the nodes, the accuracy of the detection result is further improved.

[0179] Step 307: Evaluate and optimize the object detection model according to the object detection result.

[0180] Specifically, by evaluating the performance of the model in complex scenarios, the improvement of the model's reasoning ability by the knowledge graph is verified. Common evaluation metrics include Average Precision (AP) and Mean Average Precision (mAP).

[0181] The formula for calculating the average precision is as follows:

[0182]

[0183] In the formula, P(r) represents the precision corresponding to the recall rate r. This integral represents averaging the precisions at different recall rates to evaluate the performance of the model at different recall values.

[0184] The formula for calculating the mean average precision is as follows:

[0185]

[0186] In the formula, AP i is the average precision of the i-th category, and n is the number of categories. By comparing with the traditional visual detection model, the improvement of the reasoning ability can be clearly seen.

[0187] In one implementation manner of this embodiment, taking object detection in autonomous driving as an application scenario, when performing object detection in combination with the object detection method as shown in Figure 3 firstly, data collection and preprocessing are carried out. Among them, the autonomous driving vehicle uses multi-modal sensors to collect environmental information, including: camera: collecting RGB images and extracting visual information; lidar (LiDAR): constructing 3D point cloud data to provide spatial information. Millimeter-wave radar: detecting the speed and distance of vehicles, pedestrians, and obstacles.

[0188] Then the visual data is processed by a deep convolutional neural network (CNN) to extract preliminary object features F CNN , and at the same time, the knowledge graph extracts semantic information from road rules, traffic signs, and vehicle behavior patterns to construct knowledge graph features H GCN . The cross-modal attention mechanism is used to calculate the correlation between visual features and knowledge graph features:

[0189]

[0190] where: F′ CNN and H′ GCN are linearly transformed to map different features to the same representation space. α i is calculated by Softmax normalization to measure the importance of visual features and knowledge graph features.

[0191] Combined with traffic rules and pedestrian behavior patterns in the knowledge graph for pedestrian detection to improve detection accuracy, or use vehicle types and driving rules provided by the knowledge graph for constraints for vehicle recognition, or combine visual features and the knowledge graph to enhance the detection ability of occluded objects (such as objects behind the leading vehicle and large trucks).

[0192] According to the detection results, the autonomous driving decision-making system is enabled to perform corresponding operations; such as predicting pedestrian behavior and decelerating or braking in advance; detecting obstacles and reasonably planning the driving path to avoid collisions; combining traffic regulations to ensure the vehicle drives in compliance and improve driving safety.

[0193] By implementing a knowledge-graph-based object detection method provided in this embodiment, the following beneficial effects are obtained:

[0194] (1) By deeply integrating image features with knowledge graph information, the problems of insufficient reasoning ability and lack of understanding of object relationships in traditional visual detection methods in complex scenarios are solved. In the implementation process of the method provided in this embodiment, the visual feature extraction module takes a convolutional neural network as the core, combines the context relationships of objects in the image, accurately extracts image features through deep learning algorithms, and converts visual information into structured data with semantic meaning. The knowledge graph fusion module efficiently combines the semantic information contained in the graph with image features through graph convolutional networks and attention mechanisms, enabling the visual model to not only process local visual features but also perform object detection based on semantic reasoning.

[0195] (2) Based on the topological relationships and inference rules of the knowledge graph, the relationships between objects are inferred and deduced, and the positions and categories of the detection boxes are automatically corrected to ensure the accuracy and rationality of the detection results. In the object detection and reasoning part, the detection results are converted into standardized object information, and through a multi-level optimization mechanism, the detection process can adapt to the requirements of various complex scenarios, while improving the intelligence and accuracy of the detection.

[0196] (3) The inference rules in the graph are used to gradually verify the detection results, and the conflicts and errors in the detection process are eliminated through dynamic optimization algorithms, further improving the reliability and robustness of the model. This method not only improves the efficiency and accuracy of object detection but also provides technical support for intelligent reasoning in the field of computer vision.

[0197] (4) Through the above technical means, the method provided in this embodiment significantly improves the reasoning ability of the visual detection model in complex scenarios, can adapt to diverse object detection requirements, reduces the dependence on manual intervention, and effectively reduces the error rate in traditional detection methods. This technical solution has broad prospects in the application of visual detection and provides important technical support for the intelligent upgrade of the field of computer vision.

[0198] The above is the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. A target detection method based on a knowledge graph, characterized in that, Including: Performing multi-level visual feature extraction on the acquired image to be recognized to obtain a visual feature map corresponding to the image to be recognized; Extracting a plurality of candidate regions from the visual feature map and calculating an interest score corresponding to each candidate region, so as to select a plurality of target detection regions from the plurality of candidate regions according to the interest score; Obtaining a knowledge graph corresponding to the image to be recognized and extracting graph information of the knowledge graph; Locating region visual features where each target detection region is located in the visual feature map, and fusing the region visual features and the graph information to obtain a vector representation corresponding to each target detection region; Performing target category prediction on each target detection region according to the vector representation and a preset target detection model to obtain a target detection result of the image to be recognized.

2. The object detection method based on a knowledge graph according to claim 1, characterized in that, The performing multi-level visual feature extraction on the acquired image to be recognized to obtain a visual feature map corresponding to the image to be recognized includes: Inputting the acquired image to be recognized into a pre-constructed convolutional neural network to perform multi-layer convolution and pooling operations on the image to be recognized through the convolutional neural network to obtain a visual feature map corresponding to the image to be recognized.

3. The object detection method based on a knowledge graph according to claim 1, wherein The extracting a plurality of candidate regions from the visual feature map and calculating an interest score corresponding to each candidate region, so as to select a plurality of target detection regions from the plurality of candidate regions according to the interest score includes: Inputting the visual feature map into a pre-constructed region proposal network to perform target detection on the visual feature map through a sliding window mechanism and anchor boxes preset in the region proposal network to obtain a plurality of candidate regions; For any candidate region, calculating a target classification score and an intersection ratio score of the candidate region, and taking the sum of the target classification score and the intersection ratio score as the interest score of the candidate region; When the interest score is greater than or equal to a preset score threshold, determining that the candidate region is a target detection region.

4. A method for object detection based on a knowledge graph according to claim 3, characterized in that, The calculating an interest score corresponding to each candidate region includes: For any candidate region; Calculating a target classification score corresponding to the candidate region through a preset binary classification application network; wherein, determining whether the candidate region is a background through the target classification score; Calculating the intersection area between the candidate region and the anchor box and the union area between the candidate region and the anchor box, and taking the ratio of the intersection area to the union area as the intersection ratio score.

5. A method for object detection based on a knowledge graph according to claim 1, characterized in that, The obtaining a knowledge graph corresponding to the image to be recognized and extracting graph information of the knowledge graph includes: Matching a large-scale knowledge base corresponding to the image to be recognized; Constructing a knowledge graph corresponding to the image to be recognized based on predefined entities, attributes, and relationships between entities; wherein, the entity is a target category; Inputting the knowledge graph into a pre-constructed graph convolutional network to extract semantic features of each node in the knowledge graph through the graph convolutional network, and taking the semantic features as the graph information of the knowledge graph.

6. The object detection method based on a knowledge graph according to claim 5, wherein The obtaining of the knowledge graph corresponding to the image to be recognized further includes: Designing a dynamic update mechanism corresponding to the knowledge graph; wherein, the dynamic update mechanism includes new entity addition, relationship update, and conflict resolution; Obtaining a target detection result and updating the knowledge graph in real time based on the dynamic update mechanism.

7. The object detection method based on a knowledge graph according to claim 5, wherein, The locating of the regional visual features where each of the target detection regions is located from the visual feature map and the fusing of the regional visual features and the graph information to obtain the vector representation corresponding to each of the target detection regions includes: For any of the target detection regions, extracting the regional visual features of the target detection region from the visual feature map according to the bounding box corresponding to the target detection region; Mapping the regional visual features and the semantic features to a fusion space according to a preset feature alignment mechanism; Calculating the visual weight of the regional visual features and the semantic weight of the semantic features after the mapping operation according to a preset attention mechanism; Fusing the regional visual features and the semantic weight based on the visual weight, the semantic weight, and a preset cross-modal attention mechanism to generate the vector representation corresponding to each of the target detection regions.

8. A target detection system based on a knowledge graph, characterized in that, Including a visual feature module, a graph screening module, a semantic feature module, a feature fusion module, and a target detection module; The visual feature module is configured to perform multi-level visual feature extraction on the obtained image to be recognized to obtain the visual feature map corresponding to the image to be recognized; The graph screening module is configured to extract a plurality of candidate regions from the visual feature map and calculate the interest score corresponding to each of the candidate regions, so as to select a plurality of target detection regions from the plurality of candidate regions according to the interest score; The semantic feature module is configured to obtain the knowledge graph corresponding to the image to be recognized and extract the graph information of the knowledge graph; The feature fusion module is configured to locate the regional visual features where each of the target detection regions is located from the visual feature map and fuse the regional visual features and the graph information to obtain the vector representation corresponding to each of the target detection regions; The target detection module is configured to perform target category prediction on each of the target detection regions according to the vector representation and a preset target detection model to obtain the target detection result of the image to be recognized.

9. A terminal device, characterized in that, Including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements a knowledge graph-based target detection method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, Including: A stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute a knowledge graph-based target detection method according to any one of claims 1-7.

Citation Information

Cited By

  • Urban traffic event semantic recognition method based on knowledge graph

    CN121071561A

  • Knowledge graph driven bearing motor fault visual positioning method and system

    CN121169865A

  • Target detection method, system and device based on hierarchical collaborative reasoning

    CN121330440A

  • Target detection method, device and equipment for aerial image of unmanned aerial vehicle, and medium

    CN121661533A

  • A target detection method, device and equipment for unmanned aerial vehicle aerial image and medium

    CN121661533B