Image target detection method based on graph neural network
The graph neural network-based image detection method improves detection precision by leveraging convolutional networks, region proposal networks, self-attention, and cross-attention layers to enhance feature representation and positioning accuracy.
Patent Information
- Application Number
- CN202510416914.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-03
AI Technical Summary
In the prior art, the detection accuracy of the image object detection model is relatively low.
Using a graph neural network-based method, image feature vectors are extracted through convolutional neural networks, and candidate regions and their position coordinate matrices are generated using the region proposal network, and feature fusion is combined with the self-attention layer and the cross-attention layer to improve detection accuracy.
It improves the accuracy and reliability of target detection, reduces false detection and missed detection, and enhances feature representation capabilities.
Smart Images

Figure CN120318579A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image detection, and relates to but is not limited to an image target detection method based on a graph neural network. Background Art
[0002] Object detection is one of the most basic problems in the field of computer vision, and its purpose is to detect which objects are in a picture or video and determine the specific positions of the objects. Although the image target detection model constructed by deep learning technology has been widely applied in scenarios such as intelligent driving and smart cities. However, in the related art, the target detection model still has the problem of low detection accuracy.
[0003] Therefore, how to quickly improve the image detection accuracy has become an urgent problem to be solved. Summary of the Invention
[0004] In view of this, an embodiment of the present invention provides an image target detection method based on a graph neural network, which at least solves the problem of low image detection accuracy in the related art.
[0005] According to the first aspect of the embodiments of the present invention, an image target detection method based on a graph neural network is provided, including: Inputting the image to be detected into a convolutional neural network model to obtain the extracted image feature vector; Inputting the image feature vector into a region proposal network to obtain the first candidate region feature and the corresponding position coordinate matrix; and obtaining the initial candidate region set through the first candidate region feature and the position coordinate matrix; Obtaining the category label set corresponding to the image to be detected, and based on the category label set, the candidate region set and the self-attention layer, obtaining the transformed second candidate region feature and the transformed label feature; Using the cross-attention layer to perform feature fusion on the second candidate region feature and the label feature to obtain the target feature; and inputting the target feature and the candidate region set into a fully connected layer to obtain the final positions corresponding to the candidate region set and the categories of the objects included in each of the candidate region sets.
[0006] According to the second aspect of the embodiments of the present invention, an electronic device is provided, including: a processor, a memory, a communication interface and a communication bus, and the processor, the memory and the communication interface complete communication with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform the operations corresponding to the method described in the first aspect.
[0007] According to a third aspect of an embodiment of the present invention, there is provided a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in the first aspect is implemented.
[0008] In the solution provided by the embodiment of the present invention, the image to be detected is input into a convolutional neural network model to obtain an extracted image feature vector; the image feature vector is input into a region proposal network to obtain a first candidate region feature and a corresponding position coordinate matrix; and an initial candidate region set is obtained through the first candidate region feature and the position coordinate matrix; a category label set corresponding to the image to be detected is obtained, and based on the category label set, the candidate region set, and a self-attention layer, a transformed second candidate region feature and a transformed label feature are obtained; a cross-attention layer is used to perform feature fusion on the second candidate region feature and the label feature to obtain a target feature; and the target feature and the candidate region set are input into a fully connected layer to obtain the final position corresponding to the candidate region set and the category of the object included in each candidate region set. In this process, by extracting the image feature vector through a convolutional neural network, high-level semantic information in the image can be captured. Using the region proposal network to generate candidate regions and their position coordinate matrices from the image feature vector helps to accurately locate potential target positions, reduce false detections and missed detections. Introducing a self-attention layer can better capture the correlation between different candidate regions and adjust the candidate region features according to the category label set, thereby improving the classification accuracy. The cross-attention layer can further fuse the second candidate region feature and the label feature to achieve a more accurate feature representation. By processing the target feature and the candidate region set obtained through the above steps by the fully connected layer, the exact position of each candidate region and the category of the object it contains can be output, making the final detection result more accurate and reliable. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings, where: Figure 1 is a flowchart showing a method for image target detection based on a graph neural network provided by an embodiment of the present invention Figure 1 ; Figure 2 is a diagram showing the inference process of a self-attention layer provided by an embodiment of the present invention; Figure 3 is a diagram showing the inference process of a cross-attention layer provided by an embodiment of the present invention; Figure 4Flow schematic of an image object detection method provided by an embodiment of the present invention Figure 2 ; Figure 5 Flow schematic of an image object detection method provided by an embodiment of the present invention Figure 3 ; Figure 6 Schematic diagram of the visualization comparison results of the baseline model (Baseline) and the model after adding the IRR framework (Baseline+IRR) on the VOC 2007 test set provided by an embodiment of the present invention; Figure 7 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0010] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0011] In the following description, reference is made to "some embodiments", which describe subsets of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0012] It should be noted that the terms "first\second\third" involved in the embodiments of the present invention are only used to distinguish similar objects, and do not represent a specific order for the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present invention described here can be implemented in an order other than that illustrated or described here.
[0013] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used here have the same meaning as that generally understood by those of ordinary skill in the art in the field to which the embodiments of the present invention belong. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.
[0014] Figure 1Flow schematic of an image target detection method based on a graph neural network provided by an embodiment of the present invention Figure 1 The image target detection method based on a graph neural network provided by an embodiment of the present invention can be executed by an electronic device, and the electronic device can be a computer, a server, etc.
[0015] As Figure 1 shown, the image target detection method based on a graph neural network includes: S101. Input the image to be detected into a convolutional neural network model to obtain the extracted image feature vector.
[0016] In an embodiment of the present invention, the image to be detected contains multiple objects. The "objects" in the image generally refer to specific entities or parts of interest in the image, which can be any things with clear boundaries and features such as people, animals, vehicles, buildings, etc. Inputting the image to be detected into the convolutional neural network model, the convolutional neural network can effectively extract multi-level feature representations from the image to be detected through a series of convolutional operations, pooling operations, and non-linear activation functions, and obtain the extracted image feature vector.
[0017] S102. Input the image feature vector into a region proposal network to obtain the first candidate region feature and the position coordinate matrix of the candidate region; and obtain the initial candidate region set through the first candidate region feature and the position coordinate matrix.
[0018] In an embodiment of the present invention, a candidate region refers to a rectangular region in the detected image that may contain an object of interest. These regions are not the final determined object positions, but potential object positions pre-generated based on certain algorithms or strategies. The candidate regions do not necessarily accurately correspond to the actual object bounding boxes, but they provide a starting point, enabling subsequent object detection algorithms to focus more on these regions for accurate classification and bounding box adjustment. The role of the Region Proposal Network (RPN) is to generate candidate regions on the image and output the prediction values of each candidate region, and these prediction values are used to indicate whether the candidate region contains an object on the image. Specifically, the RPN usually outputs two main pieces of information: The position of the candidate region: These are the coordinates of the regions on the image that may contain the target object (usually the coordinates of the bounding box, including the coordinates of the upper left corner and the lower right corner).
[0019] Prediction values (confidence levels): These values represent the probability that each candidate box contains the target object. For example, a prediction value close to 1 indicates that the candidate region is very likely to contain a target object, while a prediction value close to 0 indicates that the candidate box is very likely not to contain the target object.
[0020] Among them, the image feature vector is input into the region proposal network to obtain the first candidate region features and the position coordinate matrix; the initial candidate region set is obtained through the candidate region features and the position coordinate matrix.
[0021] S103. Obtain the category label set corresponding to the image to be detected, and based on the category label set, the candidate region set, and the self-attention layer, obtain the transformed second candidate region features and the transformed label features.
[0022] In the embodiments of the present invention, the category labels (labels) generally refer to the category information related to the target objects in the image. These category labels are used to indicate the category to which each target object in the image belongs. For example, in an object detection task, if the image contains different types of objects (such as cars, pedestrians, bicycles, etc.), each object will have a corresponding category label. In an object detection task, the category label usually refers to the category of each target object in the image. For example, if there is a car and a pedestrian in an image, the category labels of these two target objects may be "car" and "pedestrian". These category labels are used to guide the model to learn how to distinguish different categories of targets during the training process. When training an object detection model, the category labels are used to calculate the loss function. The output of the model (predicted category and position) will be compared with the true labels and bounding boxes to calculate the classification loss and the regression loss. By minimizing these losses, the model can learn how to more accurately identify and locate the target objects in the image. In the inference (testing) stage, the model will output the predicted category and confidence of each candidate box, and these predicted categories are generated based on the label information learned by the model during the training stage. The category label set is a set formed by organizing all possible category labels and is used for the training and inference of the model. For example, if there are 20 different categories in a dataset, the label set will contain the labels of these 20 categories.
[0023] Further, after obtaining the category label set corresponding to the image to be detected, the category label set and the candidate region set are input into the self-attention layer to obtain the transformed second candidate region features and the transformed label features.
[0024] S104. Use the cross-attention layer to perform feature fusion on the second candidate region features and the label features to obtain the target features; and input the target features and the candidate region set into the fully connected layer to obtain the final positions corresponding to the candidate region set and the categories of the objects included in each candidate region set.
[0025] In an embodiment of the present invention, the cross-attention layer is used to learn the interdependence between the region set and the label set. In the cross-attention layer, the feature representation of the candidate region set interacts with the feature representation of the class label set to capture the global dependence between the candidate regions and the class labels. For example, the cross-attention layer learns the relationship between each candidate region and each class label, thereby updating the feature representation of each candidate region. The second candidate region feature and the label feature are input into the cross-attention layer for feature fusion to obtain the target feature, and the target feature and the candidate region set are input into the fully connected layer for feature integration to obtain the final positions corresponding to the candidate region set and the classes of the objects contained in each candidate region set.
[0026] It can be understood that in an embodiment of the present invention, the image to be detected is input into the convolutional neural network model to obtain the extracted image feature vector; the image feature vector is input into the region proposal network to obtain the first candidate region feature and the corresponding position coordinate matrix; and the initial candidate region set is obtained through the first candidate region feature and the position coordinate matrix; the class label set corresponding to the image to be detected is obtained, and based on the class label set, the candidate region set, and the self-attention layer, the transformed second candidate region feature and the transformed label feature are obtained; the cross-attention layer is used to perform feature fusion on the second candidate region feature and the label feature to obtain the target feature; and the target feature and the candidate region set are input into the fully connected layer to obtain the final positions corresponding to the candidate region set and the classes of the objects contained in each candidate region set. In this process, the convolutional neural network is used to extract the image feature vector, which can capture the high-level semantic information in the image. The region proposal network is used to generate candidate regions and their position coordinate matrices from the image feature vector, which helps to accurately locate potential target positions, reduce false detections and missed detections. Introducing the self-attention layer can better capture the correlation between different candidate regions and adjust the candidate region features according to the class label set, thereby improving the classification accuracy. The cross-attention layer can further fuse the second candidate region feature and the label feature to achieve a more accurate feature representation. Processing the target feature and the candidate region set obtained through the above steps by the fully connected layer can output the accurate positions of each candidate region and the classes of the objects they contain, making the final detection result more accurate and reliable.
[0027] In some embodiments of the present invention, S103 can be implemented through S1031 to S1033, and the following steps are used for illustration.
[0028] S1031. Process the class label set using the global vector word representation algorithm to obtain the corresponding class word vectors.
[0029] S1032. Convert the category word vectors to the feature representation space corresponding to the image feature vectors using a conversion function to obtain the converted category word vectors.
[0030] S1033. Based on the converted category word vectors, the candidate region set, and the self-attention layer, obtain the transformed second candidate region features and the transformed label features.
[0031] In some embodiments of the present invention, the category labels are converted using the Global Vectors for Word Representation (GloVe) algorithm. Assuming there are 3 categories, each category is directly input into the GloVe algorithm to obtain the corresponding category word vectors, and a learnable conversion function is used to adjust these word vectors to adapt them to specific task requirements and ensure that they are in the same representation space as other types of data features (such as image feature vectors), thus obtaining the converted category word vectors.
[0032] Furthermore, the converted category word vectors and the candidate region set are input into the self-attention layer to obtain the transformed second candidate region features and the transformed label features.
[0033] In some embodiments of the present invention, S1033 can be implemented through S301 to S302, and the following steps are used for illustration.
[0034] S301. Input the category word vectors into the self-attention layer to obtain the first query vector, the first key vector, and the first value vector; and based on the first query vector, the first key vector, and the first value vector, obtain the transformed second candidate region features.
[0035] In some embodiments of the present invention, inputting the category word vectors into the self-attention layer to obtain the first query vector, the first key vector, and the first value vector through linear transformation, and the specific transformation is as follows: Among them, is the first query vector, is the first key vector, is the first value vector, , , are transformation matrices respectively, F is the category word vector, and i is the i-th element of the input category word vector.
[0036] Furthermore, based on the first query vector, the first key vector, and the first value vector, calculate the attention scores, apply the softmax function, and perform weighted summation to obtain the transformed second candidate region features. Among them, the calculation of the attention scores and the application of the softmax function are as follows: In the above formula, is the correlation coefficient matrix, and d is the dimension of the first key vector.
[0037] Among them, weighted summation is performed, that is, the learned correlation coefficient matrix is used to aggregate the feature information in the set, and is used to update the feature representation of the region set: In the above formula, is the transformed second candidate region feature.
[0038] S302. Input the candidate region set into the self-attention layer to obtain a second query vector, a second key vector, and a second value vector, and obtain the transformed label feature based on the second query vector, the second key vector, and the second value vector.
[0039] In some embodiments of the present invention, the candidate region set is input into the self-attention layer to obtain a second query vector, a second key vector, and a second value vector. Attention score calculation, application of the softmax function, and weighted summation are performed based on the second query vector, the second key vector, and the second value vector to obtain the transformed label feature. S302 is similar to S301.
[0040] As Figure 2 shown, Figure 2 is an inference process of a self-attention layer provided by an embodiment of the present invention. N1 represents the number of heads of the multi-head attention mechanism used in the self-attention module. In the feed-forward network (FFN), the information in multiple feature spaces is aggregated and calculated through residuals to obtain the final enhanced region feature, as follows: In the above formula, F represents the input feature set input to the self-attention module or the cross-attention module (the specific meaning has been described above), represents the updated feature set after being processed by the attention mechanism, τ(.) is an update function, which is responsible for updating the feature representation according to the original input feature F and the feature processed by the attention mechanism. This function may involve a multi-layer perceptron (MLP, Multi-Layer Perceptron) or other types of neural networks for learning non-linear combinations of features, is a learnable weight parameter matrix for converting the collected feature information into the same dimensional space as the original input feature set F. ← represents that the updated feature set is assigned to F, which means that the original input feature set is replaced by the updated feature set.
[0041] The above formula describes the process of updating the feature representation by combining the original input features and the features processed by the attention mechanism in the self-attention module or cross-attention module. This process involves the aggregation, transformation, and update of features, aiming to enhance the representational ability of features by capturing the interdependencies within and between sets.
[0042] In some embodiments of the present invention, the feature fusion of the second candidate region features and the label features using the cross-attention layer in S104 to obtain the target features can be achieved through S1041 to S1043, which will be described through the following steps.
[0043] S1041: Use the cross-attention layer to perform feature fusion on the second candidate region features and the label features to obtain the fused features.
[0044] As Figure 3 shown, Figure 3 is a diagram of the inference process of a cross-attention layer provided by an embodiment of the present invention. In Figure 3 , the key vectors and value vectors are determined through the second candidate region features ( ) and the label features . At the same time, the query vector is obtained through the second candidate region features. Attention score calculation, application of the softmax function, and weighted summation are performed using the query vector, key vectors, and value vectors. Since there are multiple heads, each head will generate its own context vector. To integrate the information from these different perspectives, the outputs of each head are concatenated together, and then in the feed-forward neural network (FFN), residual networks are used for feature fusion to obtain the fused features.
[0045] S1042: When the fused features meet the preset conditions or the number of iterations reaches the preset iteration conditions, use the fused features as the target features.
[0046] In some embodiments of the present invention, the second candidate region features and the label features are input into the cross-attention layer to perform operations such as transformation of the query vector, key vectors, and value vectors, attention score calculation, application of the softmax function, and weighted summation for feature fusion to obtain the fused features. When the fused features meet the preset conditions or the number of iterations reaches the maximum number of iterations, use the fused features as the target features.
[0047] S1043: When the fused features do not meet the preset conditions or the number of iterations does not reach the preset iteration conditions, input the fused features into the self-attention layer again until the target features are obtained.
[0048] In some embodiments of the present invention, when the fused features do not meet the preset conditions or the number of iterations does not reach the maximum number of iterations, the fused features are fed back to the self-attention layer for feature update and processing operations, and the features output by the self-attention layer are used as the input of the cross-attention layer again until the features output by the cross-attention layer meet the preset conditions or the number of iterations reaches the preset iteration condition, and the target features are obtained.
[0049] In some embodiments of the present invention, S1041 can be implemented by S1041A, and the following steps are used for illustration.
[0050] S1041A: Input the second candidate region feature and the label feature into the cross-attention layer to obtain a third query vector, a third key vector, and a third value vector; and perform feature fusion through the third query vector, the third key vector, and the third value vector to obtain the fused features.
[0051] In some embodiments of the present invention, the second candidate region feature and the label feature are input into the cross-attention layer. By jointly determining the third query vector, the third key vector, and the third value vector from the second candidate region feature and the label feature, attention score calculation, application of the softmax function, and weighted summation are performed based on the third query vector, the third key vector, and the third value vector to obtain the fused features.
[0052] In some embodiments of the present invention, in S104, inputting the target feature and the candidate region set into the fully connected layer to obtain the final position corresponding to the candidate region set and the category of the object included in each of the candidate region sets can be implemented by S201, and the following steps are used for illustration.
[0053] S401: Add the target feature and the candidate region set to obtain the added feature, and input the added feature into the fully connected layer to obtain the final position and category.
[0054] In some embodiments of the present invention, the target feature and the candidate region set are added to obtain the added feature, and the added feature is fed into several fully connected layers, which will finally branch into two outputs: one is the classification score (used to determine which category the object belongs to), and the other is the bounding box regression (used to refine the position and size of the object).
[0055] As Figure 4 shown, Figure 4 is the process schematic of an image target detection method based on a graph neural network provided by an embodiment of the present invention Figure 2 In Figure 4In the method, an image to be detected is input into a convolutional neural network model to obtain an extracted image feature vector. The image feature vector is input into a region proposal network to obtain a first candidate region feature and a corresponding position coordinate matrix. An initial candidate region set is obtained through the first candidate region feature and the position coordinate matrix. Input category labels, such as cls_a, cls_b, and cls_c. These category labels are converted into vector representations through a GloVe model. The category label vectors and the candidate region set are input into a self-attention layer for separate processing to obtain a transformed second candidate region feature and a transformed label feature. The second candidate region feature and the transformed label feature are input into a cross-attention layer for feature fusion. When the fused feature meets a preset condition or the number of iterations reaches a preset iteration condition, the fused feature is used as the target feature. When the fused feature does not meet the preset condition or the number of iterations does not reach the preset iteration condition, the fused feature is input into the self-attention layer until the target feature is obtained. The target feature and the candidate set are added together and input into a fully connected layer to obtain the final positions corresponding to the candidate region set and the categories of the objects included in each candidate region set.
[0056] As Figure 5 shown, Figure 5 is a schematic flowchart of a method for image object detection based on a graph neural network provided by an embodiment of the present invention Figure 3 , that is Figure 5 is an IRR framework integrating a self-attention layer and a cross-attention layer. In Figure 5 , an image is segmented into multiple candidate regions (such as the 4 green circles shown in the figure). These regions may be potential object positions. Each region extracts features through a certain method (such as a convolutional neural network). A pre-trained word embedding (such as GloVe) is used to initialize the category labels to form an initial category feature representation (such as the blue circles in the figure). The candidate region set and the category label set are input into the self-attention layer to obtain a transformed second candidate region feature and a transformed label feature. The transformed second candidate region feature and the transformed label feature are input into the cross-attention layer for feature fusion. The model iterates multiple times to gradually optimize the relationship between the region proposals and the category labels until a convergence condition is reached to obtain the target feature. Finally, based on the target feature, the final object classification result is obtained, that is, the specific category corresponding to each region and the precise adjustment of the object position to ensure that the object bounding box is more accurate.
[0057] In an embodiment of the present invention, the following is Example 1: Taking the typical samples of Pascal VOC as an example, several typical feature extraction methods based on Bayesian networks were selected to test the recognition performance of adding an interactive relationship reasoning framework on different types of detection methods. Table 1 below shows the comparison experimental results between the detection model with the interactive relationship reasoning framework added and the baseline model on the VOC 2007 test set.
[0058] Table 1 Performance comparison between baseline + IRR and baseline on the Pascal VOC test set Among them, Faster R-CNN is the Faster Region Convolutional Neural Network, Faster R-CNN+IRR is the Faster Region Convolutional Neural Network + Interactive Relationship Reasoning Framework, Faster R-CNN is the Faster Region Convolutional Neural Network, Faster R-CNN+IRR is the Faster Region Convolutional Neural Network + Interactive Relationship Reasoning Framework, ResNet50-FPN is the 50-layer Residual Network - Pyramid Network, ResNet101-FPN is the 101-layer Residual Network - Pyramid Network, the baseline model is the model without introducing the IRR framework, and Backbone is the backbone network.
[0059] After adding the IRR framework, the mAP indicators of the baseline model increased by 1.6% and 1.1% respectively. The results show that an image object detection method based on graph neural network of the present invention can enhance the classification and localization capabilities of the algorithm. Thus, it is verified that a reliable relationship structure between the learning region and the label and inside the learning region is beneficial to enhancing the representation ability of features and improving the performance of the detection algorithm.
[0060] For the MS-COCO2017 dataset, the present invention selected several mainstream region-based object detection algorithms as baseline models, including Faster R-CNN, Mask R-CNN, DCNv2withmask, and DCNv2withcascademask. The detection performance of the latter two models is stronger after adding DCNv2 to MaskR-CNN and Cascade MaskR-CNN. ResNet101-FPN was selected as the backbone network for these detection algorithms to conduct experiments. It can be seen from Table 1 that the detection algorithms after adding the IRR framework are basically better than the Baseline model. In terms of evaluation indicators, the AP, Faster R-CNN, and Mask R-CNN detection algorithms increased by 1.7% and 1.5% respectively. It also increased from 43.5% and 47.3% to 44.4% and 47.9% on two more advanced DCNv2-based models.
[0061] In addition, in order to further verify the effectiveness of the IRR framework, the present invention also conducted a visual comparison experiment, such asFigure 6 As shown in the figure. In the experiment, ResNet50-FPN was selected as the backbone network and Faster R-CNN detector was selected as the baseline model. The IRR framework can effectively enhance the detection ability of the detection algorithm. For example, Figure 6 In the first image, based on the baseline model, the model with the IRR framework improves the classification confidence value of the "motorcycle" label by learning the relational structure of the objects in the image, allowing the detection algorithm to filter the detection box according to the confidence threshold. The detection box marked as "motorcycle" is effectively retained, thereby improving the detection ability of the algorithm. For the second image, the IRR framework effectively reduces the redundant detection boxes marked as "boat" through the attention mechanism, thereby achieving correct detection by the detection algorithm. Figure 6 (Figure 9) shows the improvement effect of the IRR (Interactive Relationship Reasoning) framework in the target detection algorithm. Figure 6 It contains the visual comparison results of the baseline model (Baseline) and the model (Baseline+IRR) after adding the IRR framework on the VOC 2007 test set. The first figure shows the improvement of the classification confidence value of the "motorbike" label after adding the IRR framework on the basis of the baseline model (Baseline). The baseline model may not accurately identify or have low confidence detection frames. After adding the IRR framework, by learning the relationship structure of the objects in the image, the model can more accurately identify the "motorcycle" label and filter the detection frame according to the confidence threshold, effectively retaining the detection frame marked as "motorcycle", thereby improving the detection ability of the algorithm. The second figure shows the ability of the IRR framework to effectively reduce redundant detection frames through the attention mechanism. In the baseline model, there may be multiple redundant detection frames marked as "boat". The IRR framework reduces these redundant detection frames by aggregating the feature information of the region set, and realizes the correct detection of the detection algorithm. Figure 6 The visualization results in Figure 2 verify the actual effect of the IRR framework in the target detection algorithm, especially in improving detection accuracy and reducing false detection. The IRR framework learns the correlation coefficient matrix within the candidate region, within the label, and between the region and the label through the self-attention module and the cross-attention module, thereby effectively capturing the global long-range dependency in the image, enhancing the feature representation capability, and thus improving the performance of the detection algorithm.
[0062] The present invention improves the traditional O(n3) algorithm into a more efficient O(n2) algorithm solution, providing improved performance without sacrificing space complexity.
[0063] On this basis, two models are used for object detection respectively, and they are tested respectively to verify the effectiveness of their performance in object detection. First, only an independent attention model is introduced in object detection to examine its role in object detection. On this basis, the IRR model containing only the interactive attention model is used to examine the interactive attention in object recognition. Finally, a comprehensive IRR analysis method is proposed, and on this basis, the influence of the combined detector on the detection effect is tested. Table 2 shows the comparison of ablation performance tested by the volatile organic compound test device in 2007. Baseline: baseline model, Baseline+IRR: baseline+IRR model, Baseline+IRR (including self-attention module): baseline+IRR (including self-attention module), Baseline+IRR (including cross-attention module): baseline+IRR (including cross-attention module), Baseline+IRR (including self-attention and cross-attention modules): baseline+IRR (including self-attention and cross-attention modules). Among them, Baseline+IRR (including self-attention and cross-attention modules) is an image object detection method based on graph neural network proposed in the embodiment of the present invention.
[0064] Table 2 Comparison of ablation performance of different module IRR frameworks in VOC 2007 test set Existing detection algorithms based on graph convolutional neural networks are all based on an artificially constructed fixed graph structure. The existence of unreliable edge-building relationships in the graph will lead to the degradation of the relationship reasoning ability on the graph, thus affecting the performance of the detection algorithm. First, the framework initializes the candidate region features and prior label word vectors extracted as two independent sets, and then designs a self-attention module to learn the interdependence within the set and achieve self-reconstruction within the set. The cross-attention module is implemented to achieve the mutual reconstruction between the region set and the label set. Finally, the similarity of the two updated sets is used to calculate the predicted classification value, and the updated region set is fused with the original features to predict the regression value of the position.
[0065] As Figure 7 shown, it is a schematic structural diagram of an electronic device according to an embodiment of the present invention. The specific embodiments of the present invention do not limit the specific implementation of the electronic device.
[0066] As Figure 7 shown, the electronic device may include: a processor 502, a communication interface 504, a memory 506, and a communication bus 508.
[0067] Wherein: The processor 502, the communication interface 504, and the memory 506 communicate with each other via the communication bus 508.
[0068] The communication interface 504 is used to communicate with other electronic devices or servers.
[0069] The processor 502 is used to execute the program 510, and specifically can execute the relevant steps in the above method embodiments.
[0070] Specifically, the program 510 may include program code, and the program code includes computer operation instructions.
[0071] The processor 502 may be a central processing unit (CPU), or a specific integrated circuit (ASIC) (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.
[0072] The memory 506 is used to store the program 510. The memory 506 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.
[0073] The program 510 is specifically used to cause the processor 502 to execute the operations corresponding to the methods described in the above method embodiments.
[0074] For the specific implementation of each step in the program 510, reference may be made to the corresponding steps and descriptions in the corresponding units in the above method embodiments, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding process descriptions in the foregoing method embodiments, which will not be repeated here.
[0075] It should be noted that according to the needs of implementation, each component / step described in the embodiments of the present invention can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present invention.
[0076] The method according to the embodiments of the present invention can be implemented in hardware, firmware, or can be implemented as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or can be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and to be stored in a local recording medium, so that the method described herein can be stored as such software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or an FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as a RAM, a ROM, a flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown herein.
[0077] Those of ordinary skill in the art can realize that the units and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present invention.
[0078] The above embodiments are only used to illustrate the embodiments of the present invention, rather than to limit the embodiments of the present invention. Those of ordinary skill in the relevant technical field can also make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present invention. The patent protection scope of the embodiments of the present invention shall be defined by the claims.
Claims
1. An image object detection method based on graph neural network, characterized in that, Including: Input the image to be detected into a convolutional neural network model to obtain the extracted image feature vector; Input the image feature vector into the region proposal network to obtain the first candidate region feature and the corresponding position coordinate matrix; and obtain the initial candidate region set through the first candidate region feature and the position coordinate matrix; Obtain the class label set corresponding to the image to be detected, and based on the class label set, the candidate region set, and the self-attention layer, obtain the transformed second candidate region feature and the transformed label feature; Use the cross-attention layer to perform feature fusion on the second candidate region feature and the label feature to obtain the target feature; And input the target feature and the candidate region set into the fully connected layer to obtain the final positions corresponding to the candidate region set and the classes of the objects included in each of the candidate region sets.
2. The method according to claim 1, wherein The obtaining the transformed second candidate region feature and the transformed label feature based on the class label set, the candidate region set, and the self-attention layer includes: Process the class label set using the global vector word representation algorithm to obtain the corresponding class word vector; Use the conversion function to convert the class word vector into the feature representation space corresponding to the image feature vector to obtain the transformed class word vector; Based on the transformed class word vector, the candidate region set, and the self-attention layer, obtain the transformed second candidate region feature and the transformed label feature.
3. The method according to claim 2, wherein The obtaining the transformed second candidate region feature and the transformed label feature based on the transformed class word vector, the candidate region set, and the self-attention layer includes: Input the class word vector into the self-attention layer to obtain the first query vector, the first key vector, and the first value vector; and obtain the transformed second candidate region feature based on the first query vector, the first key vector, and the first value vector; Input the candidate region set into the self-attention layer to obtain the second query vector, the second key vector, and the second value vector, and obtain the transformed label feature based on the second query vector, the second key vector, and the second value vector.
4. The method according to claim 1, characterized in that, The using the cross-attention layer to perform feature fusion on the second candidate region feature and the label feature to obtain the target feature includes: Use the cross-attention layer to perform feature fusion on the second candidate region feature and the label feature to obtain the fused feature; When the fused feature meets the preset condition or the number of iterations reaches the preset iteration condition, use the fused feature as the target feature; When the fused feature does not meet the preset condition or the number of iterations does not reach the preset iteration condition, input the fused feature into the self-attention layer again until the target feature is obtained.
5. The method according to claim 4, wherein The using the cross-attention layer to perform feature fusion on the second candidate region feature and the label feature to obtain the fused feature includes: Input the second candidate region feature and the label feature into the cross-attention layer to obtain a third query vector, a third key vector, and a third value vector; and perform feature fusion through the third query vector, the third key vector, and the third value vector to obtain the fused feature.
6. The method according to claim 1, wherein The step of inputting the target feature and the candidate region set into the fully connected layer to obtain the final positions corresponding to the candidate region set and the categories of the objects included in each of the candidate region sets includes: Add the target feature and the candidate region set to obtain an added feature, and input the added feature into the fully connected layer to obtain the final positions and the categories.
Citation Information
Patent Citations
Object detection model training method and device
CN111709471A
Universal image target detection method and device based on self-attention mechanism
CN113902926A
Target detection method and device, computer equipment and computer readable storage medium
CN114596548A
Remote sensing image aggregation type target identification method and device based on graph attention network
CN115908908A
Clinic-driven multi-label classification framework for medical images
US20240331137A1