Self-training weak supervision object detection method guided by perception graph

The perception graph-guided self-training method addresses boundary box regression uncertainty in weakly supervised object detection by refining feature representations and aligning classification and localization tasks, improving model accuracy and robustness.

CN120318497APending Publication Date: 2025-07-15HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510469989.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the existing weakly supervised object detection methods, the imbalance between classification tasks and positioning tasks leads to uncertainty in bounding box regression, affecting the accuracy of the detection model.

Method used

Using a self-training method guided by perceptual graphs, a training framework is constructed, including pre-trained classifiers, full connection layer, self-training modules and graph guidance modules, a graph convolutional network is used to pass messages on the candidate box relationship diagram, integrate context information, optimize the gap between classification and positioning tasks, and generate accurate task-aware feature representations and location information.

Benefits of technology

It effectively alleviates the uncertainty of bounding box regression in weakly supervised object detection, improves the accuracy and stability of the model, reduces the possibility that the model converges to the local optimal solution, and improves the accuracy of object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318497A_ABST
    Figure CN120318497A_ABST
Patent Text Reader

Abstract

The invention discloses a self-training weak supervision object detection method guided by a perception graph, solves the problem of uncertainty of bounding box regression under a weak supervision self-training framework, and belongs to the field of computer vision and deep learning. The method comprises the steps that a training framework is constructed, wherein the training framework comprises a pre-training classifier, a full connection layer, a self-training module and a graph guiding module; inputting the image and the candidate frame into a pre-training classifier, generating a group of candidate frame feature vectors by an output feature block through two full connection layers, then inputting the candidate frame feature vectors into a self-training module, and generating a pseudo tag of an image guiding module by the self-training module; the graph guidance module obtains accurate task perception features and position information in candidate frame features, clusters redundant candidate frame areas under distance constraint, allocates different weights to each pair of candidate frame areas to construct a graph, and the graph convolution network performs information transmission on the graph to integrate context information from neighbor nodes. The trained self-training module is used for object detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for weakly supervised object detection in the case of scarce annotation information of training data, belonging to the fields of computer vision and deep learning. Background Art

[0002] Currently, it is time-consuming and expensive, and sometimes even infeasible, to obtain a large amount of accurately labeled bounding box data for object detection. Weakly supervised object detection is a method that trains an object detector only relying on image-level labels, where the image-level label indicates whether the object to be detected exists in the current image without annotating the specific position of the object to be detected. Current weakly supervised object detection has achieved remarkable success and wide application in the fields of specific model aircraft detection and defect detection. Weakly supervised detection models usually regard the task as a multi-instance learning task or a latent variable learning task, and train the detector to find the positive candidate bounding box regions as real instances among many candidate box regions, which is achieved by designing appropriate classification losses or complex loss functions. However, weakly supervised detection models usually converge to a local optimal solution, which is far from satisfactory.

[0003] Based on the above, the limitations of existing weakly supervised object detection methods can be summarized as the uncertainty of bounding box regression due to the imbalance between the classification task and the localization task under the weakly supervised self-training module, which in turn affects the accuracy of the detection model. Summary of the Invention

[0004] Aiming at the problem of the uncertainty of bounding box regression due to the imbalance between the classification task and the localization task under the weakly supervised self-training module, the present invention provides a self-training weakly supervised object detection method guided by a perception graph.

[0005] A weakly supervised object detection method guided by a perception graph according to the present invention includes:

[0006] S1. Construct a training framework, including a pre-trained classifier, a fully connected layer, a self-training module, and a graph-guided module;

[0007] According to the input image, a plurality of candidate bounding boxes are obtained, the input image and the candidate bounding boxes are input into the pre-trained classifier to obtain candidate bounding box feature blocks, and the candidate bounding box feature blocks generate a set of candidate bounding box feature vectors through two fully connected layers;

[0008] Each candidate bounding box feature vector is input into the self-training module for training and generates a pseudo-label for the graph-guided module;

[0009] The graph-guided module includes an encoding network, a graph convolutional network, a fully connected layer, and a softmax;

[0010] The encoding network encodes the features of each candidate box in the candidate box features to obtain a feature encoding matrix, clusters the candidate box features based on two distance constraints of spatial correlation and semantic similarity, constructs a candidate box and candidate box relationship graph G(V,E) according to the clustering result, where the candidate boxes are vertices V and the edges between vertices are E, generates a corresponding adjacency matrix according to the candidate box and candidate box relationship graph G(V,E), inputs the feature encoding matrix and the adjacency matrix into the graph convolutional network at the same time, uses the graph convolutional network to transmit messages between adjacent and contextually semantically related candidate boxes, so that the feature information can be aggregated from the context candidate boxes, and the candidate box features output by the integration of the graph convolutional network are then successively passed through a fully connected layer and a softmax layer to output the prediction result;

[0011] S2. Use the self-training module in the trained training framework to perform object detection.

[0012] Preferably, the element A in the adjacency matrix A ij is:

[0013]

[0014] where T sp is the threshold, IoU represents the intersection over union, p i represents candidate box i, p j represents candidate box j; T cos is a hyperparameter, x i represents the feature vector of candidate box i, x j represents the feature vector of candidate box j.

[0015] Preferably, the graph convolutional network includes:

[0016]

[0017] where F (0) is the input feature encoding matrix, W (0) is the weight of the first layer of graph convolution, W (1) is the weight of the second layer of graph convolution; D is the degree matrix of A, and the element D ij =∑ j A ij ; σ represents the non-linear activation layer; F (1) represents the output after passing through the first layer of graph convolution and the non-linear activation layer in sequence, and Q is the output of the graph convolutional network.

[0018] Preferably, the self-training module includes a teacher module and a student module, and the student module includes K + 1 parallel layers;

[0019] Each candidate box feature vector is simultaneously input into the K+1 layer of the teacher module and the student module. The label of the teacher module is the image-level label, and the pseudo-label of the first layer of the student module is obtained from the prediction result of the teacher module; the pseudo-label of the subsequent layer is obtained from the prediction result of the previous layer in the K+1 layer of the student module, and the pseudo-label obtained from the prediction result of the Kth layer is also used as the pseudo-label of the graph guidance module.

[0020] During the inference process, the final output comes from the prediction result of the K+1 layer.

[0021] Preferably, the K+1 layer includes a classification branch and a localization branch. Both the classification branch and the localization branch are fully connected layers and softmax layers connected in sequence, and the classification branch and the localization branch output classification information and position information respectively.

[0022] Preferably, the teacher module includes two branches. Each branch is a fully connected layer and a softmax layer connected in sequence. The fully connected layers in the two branches do not share parameters, and the output of the softmax layers in the two branches is obtained by element-wise multiplication to get the prediction result of the teacher module.

[0023] Preferably, the first layer to the Kth layer all include a fully connected layer and a softmax layer connected in sequence.

[0024] Preferably, the loss function Loss basic of the self-training module gcn and the loss function Loss

[0025] of the graph guidance module basic are added together as the composite loss function, and the training framework is jointly optimized in an end-to-end manner by minimizing the composite loss function.

[0026] Loss basic = Loss w + Loss r

[0027] The loss function Loss w of the teacher module is:

[0028]

[0029] where c is the category, C is the set of image categories, y c = 1 or y c = 0 indicates whether the category c exists, and τ(c) represents the prediction score of the category c;

[0030] The loss function of the first layer to the Kth layer of the student module is:

[0031]

[0032] Among them, ∣R∣ represents the total number of candidate boxes, represents whether the candidate box r is of category c, represents yes, represents no; represents the classification confidence score corresponding to the candidate box r predicted by the k-th layer;

[0033] The loss function of the (K + 1)-th layer of the student module is:

[0034]

[0035] Among them, L smooth-L1 represents the smooth L1 loss for localization, represents the loss weight of the (K + 1)-th layer, represents the pseudo-label obtained from the prediction result of the K-th layer, is the pseudo-label corresponding pseudo-box, X c,r and B c,r are respectively the classification confidence score and the bounding box of the candidate box r predicted by the (K + 1)-th layer for category c, and λ is used to balance the weights of the classification and localization losses of the (K + 1)-th layer;

[0036] The total loss function Loss of the student module r is:

[0037]

[0038] Among them, λ r is used to balance the weights from the first layer to the K-th layer of the student module;

[0039] The loss function Loss of the graph-guided module gcn is:

[0040]

[0041] Among them, represents the prediction result output by the graph-guided module, w c,r represents the loss weight of the graph-guided module, represents the pseudo-label obtained from the prediction result of the K-th layer of the student module.

[0042] Preferably, the pre-trained classifier is the ImageNet classifier.

[0043] Advantages of the present invention: The present invention uses a graph guidance module to solve the problem of bounding box regression uncertainty in the self-training module of the weakly supervised object detection method, thereby alleviating the problem that the model is prone to converge to a local optimal solution during weakly supervised training. The present invention belongs to the research on object detectors for scarce annotations, and to a certain extent, promotes the implementation of object detection technology in artificial intelligence deep learning, which conforms to the development trend of contemporary intelligent manufacturing.

[0044] In the present invention, the VOC2007 and VOC2012 datasets are used to train the model. VOC2007 contains 9,963 images and there are 20 types of objects to be detected. There are 5,011 training sets in the dataset to train the model proposed by the present invention and 4,952 test images to evaluate the accuracy of the model. VOC2012 contains 22,531 images and there are 20 types of objects to be detected. There are 11,540 training sets in the dataset to train the model proposed by the present invention and 10,991 test images to evaluate the accuracy of the model. The experimental results follow the evaluation metrics given officially, that is, the accuracy mAP of the model is evaluated on the test set, and the correct localization performance CorLoc of the model is evaluated on the training set. Description of the Drawings

[0045] Figure 1 is a schematic diagram of the principle of the present invention;

[0046] Figure 2 is a schematic diagram of the principle of the graph guidance module;

[0047] Figure 3 is a comparison of spatial attention maps with and without the graph guidance module;

[0048] Figure 4 is a performance analysis of different parameters in the graph guidance module, where (a) is the parameter Tsp of the spatial adjacency matrix, and (b) is the parameter Tcos of the context adjacency matrix;

[0049] Figure 5 is a comparison of qualitative detection results between this embodiment and the baseline method. Detailed Embodiment

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0051] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0052] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but it is not intended to limit the present invention.

[0053] This embodiment reduces the influence generated during iteration in the training process in two steps. Specifically, in the first step, weakly supervised object detection is modeled as a training framework in which, through a pre-trained classifier, a fully connected layer, and a self-training module, pseudo-labels are generated; in the second step, a graph-guided module is introduced to obtain accurate task-aware feature representations and location information. Under distance constraints, redundant candidate box regions are clustered, and different weights are assigned to each pair of candidate box regions to construct a graph. Then, a graph convolutional network (GCN) is applied to perform information transmission on the graph to integrate context information from neighboring nodes. The entire framework is trained in an end-to-end manner. The graph-guided module proposed in this embodiment effectively improves the feature distribution. Compared with other methods based on the same self-training module, it is further analyzed that the noise in the model comes from the imbalance between classification and localization, and corresponding measures are taken to eliminate them. Specifically, the weakly supervised object detection method guided by a perceptual graph in this embodiment includes:

[0054] Construct a training framework, including a pre-trained classifier, a fully connected layer, a self-training module, and a graph-guided module;

[0055] According to the input image, multiple candidate boxes are obtained. The input image and the candidate boxes are input into the pre-trained classifier to obtain candidate box feature blocks. The candidate box feature blocks generate a set of candidate box feature vectors through two fully connected layers; each candidate box feature vector is input into the self-training module for training. A large number of improperly processed candidate box features as inputs will also bring considerable noise and uncertainty to the model. For this reason, this embodiment proposes a graph-guided module to utilize these candidate boxes. At the same time, the graph-guided module can integrate the accurate context information of each candidate box, refine the detection boundary, and narrow the gap between the classification task and the localization task in the self-training module. In the method proposed in this embodiment, clustering-based candidate boxes are used to construct the dependency relationship between candidate boxes on the graph, and a graph convolutional network is applied to transmit information on the graph. Then, an instance-level classifier based on the graph convolutional network learns to absorb the information in the candidate box graph. Finally, the graph-guided module is supervised by the instance-level pseudo-labels generated by the self-training module. In the proposed graph-guided weak supervision module, the relational graph is defined as G=(V, E, A), where the candidate boxes are used as vertices V={v i}, the edges between vertices are E={e ij =(v i , v j )} representing nodes and edges respectively, and the edges are weighted according to appropriate weights, and A is the adjacency matrix. Specifically, the graph-guided module includes an encoding network, a graph convolutional network, a fully connected layer, and a softmax;

[0056] The encoding network performs feature encoding on each candidate box in the candidate box features to obtain a feature encoding matrix, clusters the candidate box features based on two distance constraints of spatial correlation and semantic similarity, constructs a candidate box - candidate box relationship graph G(V, E) according to the clustering result, generates a corresponding adjacency matrix A according to the candidate box - candidate box relationship graph G(V, E), inputs the feature encoding matrix and the adjacency matrix A into the graph convolutional network at the same time, uses the graph convolutional network to transmit messages between adjacent and context - semantically related candidate boxes, so that feature information can be aggregated from the context candidate boxes, and the candidate box features output by the integration of the graph convolutional network are then successively passed through a fully - connected layer and a softmax layer to output a prediction result;

[0057] The graph - guiding module of this embodiment involves two innovation points: First, a candidate box - candidate box relationship graph G(V, E) is introduced, redundant candidate box regions are clustered under distance constraints, and different weights are assigned to each pair of candidate box regions to construct a graph. Second, a graph convolutional network is applied to perform information transmission on the graph to integrate context information from neighbor nodes for obtaining accurate task - aware feature representations and location information.

[0058] During the training process, the pre - trained classifier is continuously optimized through the graph - guiding module, and then the fully - connected layer and the self - training module are optimized, eliminating the regression uncertainty of weakly - supervised object detection.

[0059] The self - training module of this embodiment includes a teacher module and a student module, and the student module includes K + 1 parallel layers;

[0060] Each candidate box feature vector is input into the K + 1 layers of the teacher module and the student module at the same time. The label of the teacher module is the image - level label, and the prediction result of the teacher module is used as the pseudo - label of the first layer of the student module; the prediction result of the previous layer in the K + 1 layers of the student module is used as the pseudo - label of the next layer, and the prediction result of the Kth layer is the pseudo - label of the graph - guiding module;

[0061] During the inference process, the final output is the prediction result from the K + 1th layer.

[0062] As Figure 1 shown, the teacher module includes two branches, each branch is a fully - connected layer and a softmax layer connected in sequence, the fully - connected layers in the two branches do not share parameters, and the outputs of the softmax layers in the two branches are obtained by element - wise multiplication operation to get the prediction result of the teacher module. Specifically:

[0063] The teacher module is used to generate pseudo - labels, and then the student module is trained under the supervision of the pseudo - labels. Given an image I, and its image - level label y = [y1, y2, …, y c ∈RC×1 , where C represents the number of image categories, and y c = 1 or y c = 0 indicates the presence or absence of class c. Approximately 2000 candidate bounding boxes are generated for the picture through the selective search method where ∣R∣ is the total number of candidate bounding boxes. The image and these candidate bounding boxes are input into a pre-trained classifier, which includes a convolutional layer and a spatial pyramid pooling layer, to generate a convolutional feature map of a fixed size for each candidate bounding box, forming a candidate bounding box feature block. The candidate bounding box feature block generates a set of candidate bounding box feature vectors through two fully connected layers. These candidate bounding box feature vectors are simultaneously input into different branches, including the layers of the teacher module and the student module. First, the teacher module projects the learned features into two different score matrices through two fully connected layers with non-shared parameters, namely x cls and x det ∈R C×|R| , which are used for class c and region r in the image respectively. These two matrices are normalized in different directions through the softmax layer, namely the class direction and the candidate bounding box direction, and are interpreted as performing classification and localization:

[0064]

[0065] Then, the score of each candidate bounding box r for the specified class c is generated through an element-wise multiplication operation x R = σ(x cls ) ⊙ σ(x det ).

[0066] Finally, the predicted score of class c is obtained by summing over all candidate bounding boxes The image module is trained under the supervision of image-level labels

[0067] The image-level loss function Loss of the teacher module w is defined as follows:

[0068]

[0069] When constructing the student module, the pseudo-labels are generated from the prediction results of the teacher module to guide the online training of the student module. The student module has K + 1 layers, and the pseudo-labels generated by the previous layer can be used to supervise the next layer. During the self-training process, the supervision signal of each layer is determined by the prediction results of the previous layer, and the pseudo-labels are continuously updated and optimized during the step-by-step iteration of each layer of the student module, thereby improving the representation ability of the entire module. Layers 1 to K each include a fully connected layer and a softmax layer connected in sequence. Specifically, taking the k-th layer of the student module as an example, assuming that the image module is class c, the k - 1-th layer of the student module outputs a set of candidate bounding box scores Select the candidate bounding box with the highest score And label it as class c, that is Since different candidate boxes usually overlap with each other, and candidate boxes with a high degree of spatial overlap should belong to the same class, the candidate boxes and their adjacent candidate boxes are labeled as class c for refinement in the k-th layer of the student module in the following. That is, if the IoU (Intersection over Union) of candidate box j and is greater than a certain threshold I t , then candidate box j is labeled as class c Otherwise, it is labeled as background Here, the threshold 0.5 is determined through experiments. At the same time, if there is no object of class c in the image, all In this way, the pseudo-labeling of all candidate boxes is completed.

[0070] Loss functions of the 1st to K-th layers of the student module are:

[0071]

[0072] where, |R| represents the total number of candidate boxes, indicates whether candidate box r is class c, indicates yes, indicates no; represents the classification confidence score corresponding to candidate box r predicted by the k-th layer;

[0073] In this embodiment, the position refinement branch of the (K + 1)-th layer is integrated into the self-training module as the last branch of the student module. This branch consists of two fully connected layers parallel to other layers of the student module and performs classification and regression tasks. Specifically, the (K + 1)-th layer includes a classification branch and a localization branch. Both the classification branch and the localization branch are composed of a fully connected layer and a softmax layer connected in sequence, and the classification branch and the localization branch output classification information and position information respectively.

[0074] Loss function of the (K + 1)-th layer of the student module is:

[0075]

[0076] where, L smootj-L1 represents the smooth L1 loss for localization, represents the loss weight of the (K + 1)-th layer, represents the pseudo-label obtained from the prediction result of the K-th layer, is the pseudo-label corresponding pseudo-box, X c,r and B c,rThey are respectively the classification confidence score and bounding box of the candidate box r predicted by the (K + 1)-th layer for the category c, and λ is the weight used to balance the classification and localization losses of the (K + 1)-th layer;

[0077] The total loss function of the student module is defined as follows:

[0078]

[0079] Among them, λ r is used to balance the weights from the first layer to the K-th layer of the student module.

[0080] The principle of the graph guidance module is as Figure 2 shown. Intuitively, nodes that are adjacent in space and semantically related in context usually depict the same object. Based on this view, this embodiment imposes double constraints to cluster candidate boxes, namely spatial correlation and semantic similarity, and constructs a graph by assigning corresponding connection weights to candidate boxes in the same cluster. Specifically, a relational graph G(V, E) is constructed. The adjacency matrix A is used to model this relationship. For the construction of edges, two distance constraints are used, namely the spatial distance d s and the cosine similarity d cos .

[0081] The encoding network performs feature encoding on each candidate box in the candidate box features. Feature encoding: maps the features of each candidate box to an encoding space and feeds it as input into the graph convolutional network. However, the fully connected layer destroys the spatial information of the image features. To retain as much information as possible, in this embodiment, through the global average pooling (GAP) layer of the RoI pooling output feature F ∈ R H ×W×D to generate F ∈ R 1×1×D . The feature encoding of the candidate box is obtained by reducing the dimension F ∈ R 1×D . In other words, each vector encodes the feature representation of a candidate box. The feature encoding matrix of the candidate box is defined as F ∈ R |R|×D .

[0082] If the candidate boxes are clustered according to the spatial distance, then the closer the distance, the more likely the candidate boxes are from the same cluster. Therefore, higher weights should be assigned to them when constructing edges.

[0083] The spatial adjacency matrix is expressed as which imposes spatial constraints on the graph and is defined as follows:

[0084]

[0085] Among them, represents the spatial distance between candidate boxes, T spis a certain threshold. Then, the Intersection over Union (IoU) is selected as the metric to measure the spatial distance between candidate boxes. This metric takes into account the positions of the candidate boxes and can be expressed as:

[0086]

[0087] In addition, candidate boxes with high context semantic relevance also tend to cluster together. This embodiment aggregates more context information to obtain more position feature representations. To this end, different weights are assigned to the corresponding edges to design the adjacency matrix. Specifically, this embodiment uses semantic similarity (i.e., cosine similarity ) to estimate the weights of the edges, which is defined as:

[0088]

[0089] where x i represents the feature vector of each candidate box. Therefore, the context adjacency matrix is as follows:

[0090]

[0091] where T cos is a hyperparameter referring to the semantic similarity threshold. Finally, the spatial adjacency matrix and the context adjacency matrix are combined to calculate the final adjacency matrix

[0092] After constructing the candidate box relationship graph through clustering, a graph convolutional network is applied to operate on the graph to reconstruct the image feature representation. GCN is an effective way of message passing that guides the deep model to learn the neighbor information of the nodes, i.e., the context information of each candidate box in this embodiment. The graph convolution is formulated as:

[0093] F (l+1) = AF (l) W (l)

[0094] where A represents the graph adjacency metric, F (l) is the feature representation of the l-th layer, and A and F are fed into the model as inputs. W is a learnable parameter matrix, and F (l+1) is the hidden or output feature of the l-th layer. In this embodiment, it is most appropriate to select to reconstruct the candidate box features through two layers of graph convolution (GCN). Therefore, this graph convolutional network adds a non-linear activation layer, such as ReLU, after the first layer of graph convolution. In addition, to eliminate the influence of the magnitude difference between features, normalization of the adjacency matrix is introduced. The final expressions of the first layer of graph convolution and the non-linear activation layer are as follows:

[0095]

[0096] Among them, F (0) is the input feature encoding matrix, and W (0) is the weight of the first-layer graph convolution, and the element D in the degree matrix ij = ∑ j A ij ; σ represents the non-linear activation layer;

[0097] The output of the second-layer graph convolution and the non-linear activation layer is:

[0098]

[0099] W (1) is the weight of the second-layer graph convolution, and Q is the output of the graph convolutional network.

[0100] A fully-connected layer is inserted after the graph convolutional network, followed by a softmax operation, which can be expressed as:

[0101]

[0102] The loss function Loss of the graph guidance module gcn is:

[0103]

[0104] Among them, represents the prediction result output by the graph guidance module, and w c,r represents the loss weight of the graph guidance module, represents the pseudo label obtained from the prediction result of the K-th layer of the student module. During the inference process, the final output comes from the last layer (the K + 1-th layer) of the student module.

[0105] Prepare training samples. In this embodiment, the VOC2007 and VOC2012 datasets are selected to verify the effectiveness of this embodiment. VOC2007 contains 9,963 images and there are 20 types of objects to be detected. There are 5,011 training sets in the dataset to train the model proposed in this embodiment and 4,952 test images to evaluate the accuracy of the model. VOC2012 contains 22,531 images and there are 20 types of objects to be detected. There are 11,540 training sets in the dataset to train the model proposed in this embodiment and 10,991 test images to evaluate the accuracy of the model. The experimental results follow the evaluation metrics given by the official, that is, the accuracy mAP of the model is evaluated on the test set, and the correct localization performance CorLoc of the model is evaluated on the training set.

[0106] This embodiment uses the pre-trained VGG16 classification network on the ImageNet dataset as the backbone network of the framework. About 2000 candidate boxes are generated on each image by the selective search method. In the basic multi-instance detector, the refinement branch is trained 3 times. In the self-training module, the student module has 4 branches, including a weighted position refinement branch. For the first three branches of the student module, this embodiment sets the weight coefficient of each student layer to λr = [3, 1, 1]. In the graph-guided module, this embodiment uses a two-layer GCN. During the training process, the entire model is trained for 75k iterations, and the learning rate changes from 0.001 in the first 30k iterations to 0.0001 in the last 45k iterations. This embodiment sets the mini-batch size to 2, the momentum to 0.9, and the weight decay to 0.0005. For training and test data augmentation, the shortest side of the image is adjusted to one of six scales {480, 576, 688, 864, 1000, and 1200}, the longest side is set to not exceed = 2000, and each image is horizontally flipped. This embodiment uses the deep learning framework PyTorch and NVIDIA RTX 3090 GPU to train the entire model.

[0107] Experiments prove that the "weakly supervised object detection method guided by perceptual maps" of this embodiment alleviates the problem of prominent discriminant regions in the weakly supervised object detection process and overcomes the bounding box regression uncertainty in the self-training module. A proposed graph-guided weakly supervised learning module is used to reorganize rich candidate boxes, guide the entire model to uniformly focus on sensitive positions for classification and localization, and obtain accurate feature representations and experimental results.

[0108] The effectiveness of the method proposed in this embodiment is trained and tested on the general VOC2007 and VOC2012 datasets respectively to verify the effectiveness of the method proposed in this embodiment. In addition, the visualization results on the VOC2007 and VOC2012 test sets are also shown.

[0109] (1) Influence of the graph-guided module

[0110] The experimental results in Table 1 are used to analyze the influence of the graph-guided module on the accuracy mAP of the VOC2007 test dataset. First, analyze the advantages of the entire graph-guided module. As shown in the second and third rows of Table 1, the graph-guided module proposed in this embodiment increases the accuracy from 54.3% to 55.2%, which verifies the effectiveness of the proposed graph-guided module. The main reason is that this module reduces the uncertainty brought by redundant candidate boxes and helps the model learn a richer task-aware feature representation. It breaks the barrier between the classification and localization tasks, enabling the overall model method to focus on the target object. In addition, a qualitative experiment was conducted to visualize the feature maps with / without the graph-guided module, asFigure 3 As shown. The first column shows the original image, the second column shows the spatial attention map without the graph guidance module, and the third column shows the result with the graph guidance module. Obviously, the proposed graph guidance module enables the model to have a broader vision. This further proves that the method proposed in this embodiment alleviates the problem that the model converges to the local optimal solution, enabling the model to not only focus on the significant regions with rich classification information but also on the boundaries with rich position information.

[0111] (2) Influence of the candidate box - candidate box relationship graph:

[0112] Analyze the influence of the candidate box - candidate box relationship graph, where the graph adjacency matrix A is removed and the graph convolutional layer is replaced with a convolutional layer. As shown in Table 2, experiments with ("w graph") and without ("w / o graph") the candidate box - candidate box relationship graph are designed here. In addition, the influence of the graph convolutional layer on the model is also analyzed, denoted as "GCN1", "GCN2", and "GCN3". For example, "GCN2" means using two layers of graph convolutional layers to construct the graph guidance module. As shown in Table 2, the first row is the result without the candidate box - candidate box relationship graph. Compared with the first row and the second row in Table 2 (without and with the whole graph), the model with the relationship graph significantly improves the accuracy mAP, from 49.5% to 52.5%, which proves the effectiveness of the candidate box - candidate box relationship graph. Compared with the second row, the third row and the fourth row show the results of using different numbers of layers of GCN. It can be seen that the experimental results corresponding to GCN2 and GCN3 achieve accuracy performances of 55.2% and 52.0% respectively, so the number of convolutional layers involved in the graph guidance module is set to 2.

[0113] (3) Influence of the parameters in the graph guidance module

[0114] Analyze the influence of each parameter in the graph guidance module. The spatial distance threshold Tsp is a parameter for the edges used in candidate box clustering and relationship graph construction. Figure 4 In (a), the analysis results of different Tsp values on the PASCAL VOC2007 test dataset are shown while fixing other parameters. It can be observed that when Tsp = 0.5, the model achieves the highest accuracy of 55.2%, which indicates that too many or too few neighbors will reduce the performance of the model, mainly because too many neighbors will introduce more noise into the model, and too few neighbors will prevent the model from learning more context information. Figure 4 In (b), the detection performance of different T cos values during graph construction is shown, and it represents the cosine similarity threshold during graph construction. The point a on the horizontal axis represents the mean of cos ij minus the standard deviation. As the parameter increases, the performance of the model gradually decreases. When T cosThe model performs best when =0. Therefore, in subsequent experiments, the parameter T cos is fixed at 0.

[0115] (4) Comparison with other baseline models and visualization results

[0116] Experiments were conducted on the PASCAL VOC0712 dataset, and the method proposed in this embodiment was compared with other baseline methods for weakly supervised object detection. The results are shown in Tables 3, 4, and 5. Table 3 reports the detection performance of various methods on the VOC2007 test set. In terms of the accuracy mAP metric, the method of this embodiment is 14.0%, 7.9%, 7.5%, and 1.3% higher than OICR, MELM, CBASH, and SLV respectively, and achieves the highest performance (55.2%). Compared with the self-paced method that can also alleviate the impact of uncertainty in weakly supervised object detection, the method of this embodiment is 9.3% and 7.6% higher respectively. Obviously, the single model without any post-processing not only outperforms the single model but also significantly exceeds the ensemble model or the supervised retraining model. In addition, the proposed perception graph-guided self-training model in this embodiment performs well on both rigid and non-rigid objects. In particular, for "bird", "chair", and "person", the accuracies of 61.3%, 34.6%, and 27.9% are obtained respectively. Table 4 shows the performance of different methods on the VOC2012 test dataset. The proposed single model is about 2% higher than other advanced weakly supervised detection methods and only 0.1% lower than the method of Ren et al. The improvement obtained on the VOC2007 test dataset further verifies the effectiveness of the proposed method. Table 5 shows the localization accuracy results obtained on the VOC2007 and 2012 training validation datasets. The method achieves an average localization accuracy of 70.9% on the VOC2007 training validation dataset, leading other advanced methods by a large margin. For example, the method of this embodiment exceeds methods such as OICR, ZigZag, C-SPCL, and CBASH by 10.3%, 9.7%, 8.5%, and 5.1% respectively in terms of the accuracy mAP. Compared with the SLV method, this method is only 0.1% lower. On the VOC2012 training validation dataset, this method leads other advanced baseline models, and the localization performance Corloc achieves 70.9%.

[0117] The method proposed in this embodiment has significant advantages over other methods, which benefits from the proposed graph-guided module for improving the feature distribution. Compared with the method based on the same self-training module, the advantage is that it not only analyzes the uncertainty of the model and takes corresponding measures to eliminate them. Some detection results such as Figure 5 shown, compared with the baseline results, more accurate bounding boxes are generated, which depends on the ability of the graph-guided module to learn context information and the ability to eliminate uncertainty.

[0118] Effectiveness Analysis of Table 1 Figure Guidance Module

[0119] Method mAP (%) Baseline Model 47.4 Baseline Model + Self-training Module 54.3 Baseline Model + Self-training Module + Graph-guided Module 55.2

[0120] Ablation Study of Table 2 Candidate Box - Candidate Box Relationship Diagram

[0121]

[0122] Accuracy Comparison on VOC2007 Test Set in Table 3

[0123]

[0124]

[0125] Correct Localization Performance Comparison on VOC2012 Training Set in Table 4

[0126]

[0127]

[0128] Correct Localization Performance Comparison on VOC2007 and VOC2012 Training Sets in Table 5

[0129] Method 2007VOC 2012VOC Yang 68.0 69.5 MELM 61.4 - SLV 71.0 69.2 C-SPCI 62.4 - MIST 68.8 70.9 CBASH 65.8 68.3 Ours 70.9 70.9

[0130] Although the present invention has been described herein with reference to particular embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. Accordingly, it should be understood that numerous modifications may be made to the exemplary embodiments, and other arrangements may be designed, provided they do not depart from the spirit and scope of the present invention as defined by the appended claims. It should be understood that different dependent claims and the features described herein may be combined in a manner different from that described in the original claims. It should also be understood that the features described in connection with a single embodiment may be used in other described embodiments.

Claims

1. A self-training weak-supervised object detection method guided by a perception map, characterized in that Including: S1. Construct a training framework, including a pre-trained classifier, a fully connected layer, a self-training module, and a graph-guided module; Based on the input image, obtain multiple candidate boxes, input the input image and the candidate boxes into the pre-trained classifier to obtain candidate box feature blocks, and generate a set of candidate box feature vectors through two fully connected layers for the candidate box feature blocks; Input each candidate box feature vector into the self-training module for training and generate pseudo-labels for the graph-guided module; The graph-guided module includes an encoding network, a graph convolutional network, a fully connected layer, and a softmax; The encoding network performs feature encoding on each candidate box in the candidate box features to obtain a feature encoding matrix, clusters the candidate box features based on two distance constraints of spatial correlation and semantic similarity, constructs a candidate box-candidate box relationship graph G(V, E) according to the clustering result, where the candidate boxes are vertices V, and the edges between vertices are E, generate a corresponding adjacency matrix according to the candidate box-candidate box relationship graph G(V, E), input the feature encoding matrix and the adjacency matrix into the graph convolutional network at the same time, use the graph convolutional network to transmit messages between adjacent and context-semantically related candidate boxes, so that feature information can be aggregated from the context candidate boxes, and the candidate box features integrated and output by the graph convolutional network are successively passed through a fully connected layer and a softmax layer to output the prediction result; S2. Use the self-training module in the trained training framework for object detection.

2. The self-training weak supervised object detection method guided by a perception map according to claim 1, wherein The element A in the adjacency matrix A ij is as follows: Among them, T sp is the threshold, IoU represents the intersection over union, p i represents the candidate box i, p j represents the candidate box j; T cos is a hyperparameter, x i represents the feature vector of the candidate box i, x j represents the feature vector of the candidate box j.

3. The self-training weak supervised object detection method guided by a perception graph according to claim 2, characterized in that The graph convolutional network includes: Among them, F (0) is the input feature encoding matrix, W (0) is the weight of the first-layer graph convolution, W (1) is the weight of the second-layer graph convolution; D is the degree matrix of A, and the element D ij in the degree matrix is ∑ j A ij ; σ represents the non-linear activation layer; F (1) represents the output passing through the first-layer graph convolution and the non-linear activation layer in sequence, and Q is the output of the graph convolutional network.

4. The self-training weakly supervised object detection method guided by a perception map according to claim 1, characterized in that, The self-training module includes a teacher module and a student module, and the student module includes K + 1 parallel layers; Input each candidate box feature vector into the K + 1 layers of the teacher module and the student module at the same time. The label of the teacher module is the image-level label, and the pseudo-label of the first layer of the student module is obtained from the prediction result of the teacher module; the pseudo-label of the next layer is obtained from the prediction result of the previous layer in the K + 1 layers of the student module, and the pseudo-label obtained from the prediction result of the Kth layer is also used as the pseudo-label of the graph-guided module at the same time; During the inference process, the final output is the prediction result from the K + 1th layer.

5. The self-training weakly supervised object detection method guided by a perception map according to claim 4, wherein The K + 1th layer includes a classification branch and a localization branch, and both the classification branch and the localization branch are a fully connected layer and a softmax layer connected in sequence, and the classification branch and the localization branch output classification information and position information respectively.

6. The self-training weak supervised object detection method guided by a perception map according to claim 4, wherein The teacher module includes two branches, each branch is a fully connected layer and a softmax layer connected in sequence, the fully connected layers in the two branches do not share parameters, and the output of the softmax layers in the two branches is obtained by element-wise multiplication to get the prediction result of the teacher module.

7. The self-training weak-supervised object detection method guided by a perception map according to claim 4, wherein Layers from the 1st layer to the Kth layer all include a fully connected layer and a softmax layer connected in sequence.

8. The self-training weak-supervised object detection method guided by a perception map according to claim 4, wherein Add the loss function Loss of the self-training module basic and the loss function Loss of the graph-guided module gcn to obtain a composite loss function, and jointly optimize the training framework in an end-to-end manner by minimizing the composite loss function.

9. The self-training weak supervision object detection method guided by a perception map according to claim 5, wherein The loss function Loss of the self-training module basic is as follows: Loss basic = Loss w + Loss r Loss function of the teacher module w is as follows: where c is the category, C is the set of image categories, and y c = 1 or y c = 0 indicates whether the category c exists, and τ(c) represents the predicted score of the category c; Loss function from the 1st layer to the Kth layer of the student module is as follows: where k = 1, 2, …, K, |R| represents the total number of candidate boxes, and R represents the set of candidate boxes. indicates whether the candidate box r is of class c. indicates yes. indicates no. represents the classification confidence score corresponding to the candidate box r predicted at the k-th layer. The loss function of the (K + 1)-th layer of the student module is as follows: Among them, L smooth-L1 represents the smooth L1 loss for localization, represents the loss weight of the (K + 1)-th layer, represents the pseudo-label obtained from the prediction result of the K-th layer, is the pseudo-label corresponding pseudo-box, X c,r and B c,r are respectively the classification confidence score and the bounding box of the candidate box r predicted by the (K + 1)-th layer for the category c, and λ is used to balance the weights of the classification and localization losses of the (K + 1)-th layer; Total loss function Loss of the student module r is as follows: Among them, λ r is used to balance the weights from the first layer to the Kth layer of the student module; Loss function of the graph guidance module gcn is as follows: Among them, represents the prediction result output by the graph guidance module, w c,r represents the loss weight of the graph guidance module, represents the pseudo-label obtained from the prediction result of the K-th layer of the student module.

10. The self-training weakly supervised object detection method guided by a perception map according to claim 1, characterized in that, The pre-trained classifier is an ImageNet classifier.