A Weakly Supervised Object Localization Method Based on Spatial Awareness Tokens

By introducing space-aware tokens and joint network training, the optimization conflicts of positioning and classification tasks in Transformer weakly supervised object positioning method are solved, and efficient positioning and classification results are achieved, which are suitable for various Transformer backbones.

CN117315219BActive Publication Date: 2025-07-25UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311383335.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-24
Publication Date
2025-07-25
Estimated Expiration
2043-10-24

AI Technical Summary

Technical Problem

The existing weakly supervised object positioning method based on Transformer has optimization conflicts between positioning and classification tasks, resulting in limited performance of the network on both tasks.

Method used

Introduce spatially perceptual tokens specific to positioning tasks, and by constructing a classification-positioning joint network, use spatially perceptual tokens to carry and capture positioning knowledge, generate positioning maps, and combine cross-entropy classification losses, batch area losses and regularization losses for training to avoid optimization conflicts between positioning and classification tasks.

Benefits of technology

Decoupling of positioning and classification tasks is realized, positioning performance is improved, training data and parameters demands are reduced, and the accuracy and stability of positioning results are improved. It is suitable for various Transformer backbones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117315219B_ABST
    Figure CN117315219B_ABST
Patent Text Reader

Abstract

The present invention discloses a weakly supervised object localization method based on spatial perception tokens. The steps include: 1. Serialization processing of images; 2. Constructing a classification-localization joint network based on modules such as spatial perception Transformer; 3. Concatenating class tokens and spatial perception tokens for the processed image sequence as input, generating a localization map after passing through the spatial perception Transformer block, and outputting a class prediction result at the end of the entire network; 4. Averaging the localization maps generated by all spatial perception Transformer blocks as the final localization result; 5. Constructing classification loss, batch area loss, and regularization loss to train the model. The present invention can effectively achieve the weakly supervised object localization task by introducing additional spatial tokens to aggregate information and generate a localization result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and specifically relates to a weakly supervised object localization method based on spatial perception tokens. Background Art

[0002] The purpose of object localization is to identify and locate objects in a given image. However, achieving this task in a fully supervised manner requires accurate bounding box annotations, which limits its practical use. For this reason, weakly supervised object localization (WSOL) emerged. Its purpose is to achieve the object localization task only with image class labels. Since it does not require expensive bounding box labels or pixel-level labels, weakly supervised object localization significantly reduces the manual annotation cost and attracts more and more researchers. Zhou et al. first proposed CAM to achieve the weakly supervised object localization task. By synthesizing deep feature maps and fully connected weights, a class activation map (CAM) is generated to achieve the localization task. Gao et al. extended the idea of CAM to the Transformer backbone, and extracted the attention map and deep feature map in the classification network as the localization map. Subsequent works such as LCTR and SCM further improved the localization quality by enhancing the connectivity and local consistency of the localization map.

[0003] Existing Transformer-based methods all synthesize the feature maps learned from the classification task as the localization map. However, this simple approach leads to an optimization conflict between the localization and classification tasks, thereby reducing the performance of the model on both tasks. On the one hand, allowing the classification feature map to learn more low-discriminative object regions will reduce the classification performance of the network. On the other hand, the learning of the localization map is limited by the characteristics of the classification feature map, and it is difficult to generate balanced and comprehensive responses on the object. Therefore, this potential optimization conflict will limit the performance of the network on the localization and classification tasks. Summary of the Invention

[0004] The present invention is to solve the above-mentioned deficiencies of the existing technologies, and proposes a weakly supervised object localization method based on spatial perception tokens, aiming to solve the optimization problem between the localization and classification tasks. By introducing spatial perception tokens specific to the localization task to carry and capture the localization knowledge of the network, and establishing a classification-localization joint network, weakly supervised object localization is achieved.

[0005] To achieve the above object of the invention, the following technical solutions are adopted:

[0006] The characteristics of a weakly supervised object localization method based on spatial perception tokens of the present invention are as follows: the method is carried out according to the following steps:

[0007] Step 1: Serialization processing of the image:

[0008] Step 1.1: Obtain a set of images with category labels in batch B. Denote the b-th image in the image set as where h, w, and c represent the height, width, and number of channels of the image respectively. Let the image I b have a category label of C is the number of categories in the dataset;

[0009] Step 1.2: Split the image I b into a series of non-overlapping square patches. After linear processing, obtain a sequence of image patches where D represents the number of channels, N represents the length of the sequence, and N = h × w / P 2 , and P is the side length of the patch;

[0010] Step 1.3: Construct a set of category tokens to be learned and spatial perception tokens

[0011] Step 1.4: Concatenate the image patch sequence X b with the category token x cls and the spatial perception token x spa to generate an initial sequence F b ={X b ,x cls ,x spa} and use it as the input to the network, and

[0012] Step 2: Build a cascaded network consisting of L1 Transformer blocks, L2 spatial perception Transformer blocks, one convolutional layer, and one global average pooling layer in a linear concatenation form to construct a classification-localization joint network for generating the classification prediction result and localization prediction result of the image I b :

[0013] Step 2.1: Input F b into the classification-localization joint network. After being processed by L1 Transformer blocks, obtain a feature sequence and input it into the cascaded L2 spatial perception Transformer blocks;

[0014] Step 2.2: Use as the input to the l-th spatial perception Transformer block. After being processed by the l-th spatial perception Transformer block, generate the sequence output by the l-th spatial perception Transformer block and the localization map Thus, the feature sequence output by the L2 spatial perception Transformer block and the localization map

[0015] Step 2.3: The feature sequence After passing through a convolutional layer and a global average pooling layer, the image I b The prediction logic of the category

[0016] Step 2.4: Average all the localization maps in the L2 spatial perception Transformer blocks to obtain the final localization map M b ;

[0017] Step 3: Based on the image I b The prediction logic of the category and the localization map M b , construct a cross-entropy classification loss L cls and two spatial losses, including: the batch area loss L ba and the regularization loss L norm ;

[0018] Step 3.1: Construct the cross-entropy classification loss L cls :

[0019]

[0020] In Equation (4), y b,m represents the true category y b of the image I b on the label of category m; and respectively represent the activation values of the prediction logic on categories m and c;

[0021] Step 3.2: Construct the batch area loss L ba :

[0022]

[0023] In Equation (5), M b (i,j) represents the value of the localization map M b at the pixel point (i,j); λ represents a hyperparameter;

[0024] Step 3.3: Construct the regularization loss L norm :

[0025]

[0026] Step 3.4: Construct the total loss function \(L\) according to Equation (7):

[0027] \(L = L_{}\) cls + \(L_{}\) ba + \(L_{}\) norm (7)

[0028] Step 4: Use the gradient descent method to train the classification - localization joint network, and calculate the total loss function \(L\) to update the network parameters until the total loss function \(L\) converges, so as to obtain a trained classification - localization joint model for realizing the classification and localization prediction of any input image.

[0029] The characteristics of the weakly supervised object localization method based on spatial - aware tokens described in the present invention also lie in that any \(l\) - th spatial - aware Transformer block in the step 2.2 sequentially includes: a first Layer - Norm layer, a spatial query attention module SQA, a skip connection layer, a second Layer - Norm layer, an MLP layer, and a skip connection layer; wherein, the spatial query attention module SQA includes: a localization map generation mechanism layer based on spatial - aware tokens, a localization - guided cross - attention calculation layer;

[0030] Step 2.2.1. After the first Layer - Norm layer of the \(l\) - th spatial - aware Transformer block normalizes the received feature sequence it generates a sequence and inputs it into the spatial query attention module SQA of the \(l\) - th spatial - aware Transformer block;

[0031] Step 2.2.2. The localization map generation mechanism layer in the \(l\) - th spatial - aware Transformer block maps the \(l\) - th sequence to a query matrix a key matrix a value matrix by a fully - connected layer, and according to the spatial - aware token \(x_{}\) spa , it obtains the corresponding query vector from the query matrix Thus, the similarity matrix of the \(l\) - th spatial - aware Transformer block is calculated according to Equation (1)

[0032]

[0033] In Equation (1), is a balance factor, represents the transpose of the matrix ;

[0034] Step 2.2.3. The localization map generation mechanism layer in the l-th spatial perception Transformer block processes the similarity matrix using the Sigmoid activation function to obtain the l-th foreground activation probability map

[0035] Deform the part of X b corresponding to in the foreground activation probability map to obtain the localization map

[0036] Step 2.2.4. The localization-guided cross-attention calculation layer in the l-th spatial perception Transformer block obtains the spatial query attention result of the l-th spatial perception Transformer block according to Equation (2)

[0037]

[0038] In Equation (2), * represents dot product;

[0039] Step 2.2.5. After being processed by the skip connection layer, Layer-Norm layer, MLP layer, and skip connection layer in sequence, the feature sequence output by the l-th spatial perception Transformer block is obtained

[0040] An electronic device according to the present invention includes a memory and a processor, characterized in that the memory is used to store a program for supporting the processor to execute the weakly supervised object localization method, and the processor is configured to execute the program stored in the memory.

[0041] A computer-readable storage medium according to the present invention, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, it executes the steps of the weakly supervised object localization method.

[0042] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0043] 1. The present invention introduces task-specific tokens for localization in the input space to carry and converge localization knowledge, so as to decouple the generation of localization results and classification results. Therefore, the proposed method only needs to fine-tune a small number of parameters related to localization or only needs a small amount of data to obtain competitive results, having data efficiency and fine-tuning efficiency. The method proposed by the present invention far exceeds the prior art when using less than 0.1% of the training data and fine-tuning less than 20% of the parameters, indicating that the proposed method has superior localization performance.

[0044] 2. The present invention uses additional spatial perception tokens for generating localization maps instead of directly synthesizing classification feature maps, thus avoiding potential optimization conflicts between localization and classification tasks, which significantly improves the performance of the proposed method in localization and classification tasks.

[0045] 3. The present invention introduces two spatial loss functions to establish pixel-level supervision of the localization map, making each pixel have an obvious bias in distinguishing between background and foreground. Therefore, the visualization results are sharper and more reliable, and the generated localization results are insensitive to thresholds, with stability and consistency in localization accuracy in various environments.

[0046] 4. The spatial perception tokens and spatial query attention module proposed by the present invention can be adapted to various types and scales of Transformer backbones and can bring consistent improvements. Therefore, the proposed method has universality. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a framework diagram of a weakly supervised object localization based on spatial perception tokens of the present invention;

[0048] Figure 2 It is a structural diagram of a spatial perception Transformer block of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] In this embodiment, a weakly supervised object localization method based on spatial perception tokens, as Figure 1 shown, is carried out according to the following steps:

[0050] Step 1. Serialization processing of the image:

[0051] Step 1.1. Obtain a set of image collections with class labels in batch B, and denote the b-th image in the image collection as where h, w, and c respectively represent the height, width, and number of channels of the image. Let the image I b have a class label of C is the number of classes in the dataset. On the CUB-200 and ImageNet datasets, C is 200 and 1000 respectively;

[0052] Step 1.2. Split the image I b into a series of non-overlapping square patches, and after linear processing, obtain a set of image patch sequences where D represents the number of channels, N represents the length of the sequence, and N = h × w / P 2 , and P is the side length of the patch;

[0053] Step 1.3. Construct a set of learnable class tokens in the input space and spatial perception token

[0054] Step 1.4: Concatenate the image patch sequence X b with the class token x cls and the spatial perception token x spa to generate an initial sequence F b ={X b , x cls , x spa} and use it as the input of the network, and

[0055] Step 2: As Figure 1 shown, construct cascaded L1 Transformer blocks, L2 spatial perception Transformer blocks, a convolutional layer, and a global average pooling layer in a linear concatenation form to construct a classification-localization joint network for generating the classification prediction result and the localization prediction result of the image I b . In this example, L1 is taken as 9 and L2 is taken as 3:

[0056] Step 2.1: Input F b into the classification-localization joint network, and after being processed by L1 Transformer blocks, obtain a feature sequence and input it into the cascaded L2 spatial perception Transformer blocks;

[0057] Step 2.2: Take as the input of the l-th spatial perception Transformer block, and after being processed by the l-th spatial perception Transformer block, generate the sequence output by the l-th spatial perception Transformer block and the localization map Thus, the feature sequence output by the L2-th spatial perception Transformer block and the localization map

[0058] In this embodiment, as Figure 2 shown, any l-th spatial perception Transformer block sequentially includes: a first Layer-Norm layer, a spatial query attention module SQA, a skip connection layer, a second Layer-Norm layer, an MLP layer, and a skip connection layer; among them, the spatial query attention module SQA includes: a localization map generation mechanism layer based on the spatial perception token, a localization-guided cross-attention calculation layer;

[0059] Step 2.2.1. After the first Layer-Norm layer of the l-th spatial perception Transformer block normalizes the received feature sequence it generates a sequence and inputs it into the spatial query attention module SQA of the l-th spatial perception Transformer block;

[0060] Step 2.2.2. The localization map generation mechanism layer in the l-th spatial perception Transformer block maps the l-th sequence to a query matrix a key matrix and a value matrix through a fully connected layer. And according to the spatial perception token x spa , it obtains the corresponding query vector from the query matrix Thus, the similarity matrix of the l-th spatial perception Transformer block is calculated according to Equation (1)

[0061]

[0062] In Equation (1), is the balance factor, represents the transpose of matrix ;

[0063] Step 2.2.3. The localization map generation mechanism layer in the l-th spatial perception Transformer block processes the similarity matrix using the Sigmoid activation function to obtain the l-th foreground activation probability map

[0064] After deforming the corresponding part of X b in the foreground activation probability map , the localization map

[0065] is obtained. Step 2.2.4. The localization-guided cross-attention calculation layer in the l-th spatial perception Transformer block obtains the spatial query attention result of the l-th spatial perception Transformer block according to Equation (2)

[0066]

[0067] In Equation (2), * represents dot product;

[0068] Step 2.2.5. After being processed by the skip connection layer, Layer-Norm layer, MLP layer, and skip connection layer in sequence, the feature sequence output by the $l$-th spatial perception Transformer block is obtained.

[0069] Step 2.3: Feature sequence After passing through a convolutional layer and a global average pooling layer, the image $I$ is obtained. b Prediction logic for categories

[0070] Step 2.4: Average all the localization maps in the $L_2$ spatial perception Transformer blocks to obtain the final localization map $M$. b ;

[0071] Step 3: Based on the image $I$ b Prediction logic for categories and the localization map $M$ b , construct a cross-entropy classification loss $L$ cls and two spatial losses, including: batch area loss $M$ ba and regularization loss $L$ norm ;

[0072] Step 3.1: Construct the cross-entropy classification loss $L$ according to Equation (4) cls :

[0073]

[0074] In Equation (4), $y$ b,m represents the true category $y$ b of the image $I$ b on the label of category $m$; and respectively represent the activation values of the prediction logic on categories $m$ and $c$;

[0075] Step 3.2: Construct the batch area loss $L$ according to Equation (5) ba :

[0076]

[0077] In Equation (5), $M$ b (i,j) represents the value of the localization map $M$ b at the pixel point (i,j); $\lambda$ represents a hyperparameter and can be appropriately adjusted according to different datasets;

[0078] Step 3.3: Construct the regularization loss $L$ according to Equation (6) norm :

[0079]

[0080] Step 3.4: Construct the total loss function \(L\) according to Equation (7):

[0081] \(L = L\) cls + \(L\) ba + \(L\) norm (7)

[0082] Step 4: Use the gradient descent method to train the classification - localization joint network, and calculate the total loss function \(L\) to update the network parameters until the total loss function \(L\) converges, so as to obtain a trained classification - localization joint model for realizing the classification and localization prediction of any input image.

[0083] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above - mentioned method, and the processor is configured to execute the program stored in the memory.

[0084] In this embodiment, a computer - readable storage medium stores a computer program, and when the computer program is run by a processor, it executes the steps of the above - mentioned method.

[0085] Embodiment:

[0086] To verify the effectiveness of the method of the present invention, this embodiment selects the CUB - 200 - 2011 and ImageNet datasets, and uses the Top - 1 localization accuracy, Top - 1 classification accuracy, Top - 5 localization accuracy, and GT category localization accuracy as quantitative evaluation criteria.

[0087] In this embodiment, fifteen methods are selected for comparison with the method of the present invention. The selected methods are CAM, ORNet, BAS, Kim et al., CREAM, GCNet, SPA, FAM, BagCAMs, SPOL, DA - WSOL, ISIC, TS - CAM, LCTR, and SCM; according to the experimental results, the results are shown in Tables 1, 2, 3, and 4 as follows:

[0088] Table 1 Localization results of the method of the present invention and the fifteen selected comparison methods on two datasets

[0089]

[0090]

[0091] Table 2 Localization and classification results of the method of the present invention and two selected comparison methods on the CUB - 200 dataset

[0092]

[0093] Table 3 Localization results of the method of the present invention and TS-CAM under different settings of the number of training samples for each category

[0094]

[0095]

[0096] Table 4 Classification and localization results of the method of the present invention and three methods on the ImageNet dataset under different frozen parameter ratios

[0097]

[0098] The experimental results show that the method of the present invention has better effects compared with the other fifteen methods, thus proving the feasibility of the method proposed by the present invention.

Claims

1. A weakly-supervised object localization method based on spatially-aware tokens, characterized in that, The steps are as follows: Step 1: Serialization processing of the image: Step 1.1: Obtain a set of images with class labels in batch B, and denote the b-th image in the image set as where h, w, and c represent the height, width, and number of channels of the image, respectively. Let the image I b have a class label of C is the number of classes in the dataset; Step 1.

2. Split the image I b into a series of non - overlapping square patches. After linear processing, a sequence of image patches is obtained where D represents the number of channels, N represents the length of the sequence, and N = h×w / P 2 , and P is the side length of the patch; Step 1.3, construct a set of category tokens to be learned in the input space and spatial perception tokens Step 1.

4. Concatenate the image patch sequence X b with the class token x cls and the spatial perception token x spa to generate the initial sequence F b = {X b , x cls , x spa} and use it as the input of the network, and Step 2: Construct a classification-localization joint network in the form of linear cascading with L1 Transformer blocks, L2 spatial-aware Transformer blocks, one convolutional layer, and one global average pooling layer to generate the classification prediction result and the localization prediction result of image I b : Step 2.1: Input F b into the classification-localization joint network. After being processed by L1 Transformer blocks, a feature sequence is obtained and input into L2 cascaded spatial-aware Transformer blocks; Step 2.2: Take as the input of the l-th spatial perception Transformer block. After being processed by the l-th spatial perception Transformer block, generate the sequence output by the l-th spatial perception Transformer block and the positioning map Thus, the feature sequence output by the L2-th spatial perception Transformer block Step 2.3: Feature sequence After passing through a convolutional layer and a global average pooling layer, image I is obtained b Prediction logic for categories Step 2.4: Average all the localization maps in the L2 spatial perception Transformer blocks to obtain the final localization map M b ; Step 3: Based on the image I b Prediction logic for the category and the localization map M b , construct a cross-entropy classification loss L cls and two spatial losses, including: batch area loss L ba and regularization loss L norm ; Step 3.1: Construct the cross-entropy classification loss L according to Equation (4) cls :[[]]END]] In formula (4), y b,m Represents image I b The true category y b The label on category m; and Represents prediction logic Activation values for category m and category c; Step 3.2: Construct the batch area loss L according to Equation (5) ba :[[]]END]] In formula (5), M b (i, j) represents the value of the positioning map M b at the pixel point (i, j); λ represents a hyperparameter; Step 3.3: Construct the regularization loss L according to Equation (6) norm :[[]]END]] Step 3.4: Construct the total loss function L according to Equation (7): L = L cls + L ba + L norm (7) Step 4: Use the gradient descent method to train the classification-localization joint network, and calculate the total loss function L to update the network parameters until the total loss function L converges, so as to obtain a trained classification-localization joint model for realizing the classification and localization prediction of any input image.

2. The weakly supervised object localization method based on spatially-aware tokens according to claim 1, wherein Any l-th spatial perception Transformer block in the step 2.2 sequentially includes: a first Layer-Norm layer, a spatial query attention module SQA, a skip connection layer, a second Layer-Norm layer, an MLP layer, and a skip connection layer; wherein, the spatial query attention module SQA includes: a localization map generation mechanism layer based on spatial perception tokens, a localization-guided cross-attention calculation layer; Step 2.2.1: After the first Layer-Norm layer of the l-th spatial perception Transformer block performs normalization on the received feature sequence it generates a sequence and inputs it into the spatial query attention module SQA of the l-th spatial perception Transformer block; Step 2.2.

2. The localization map generation mechanism layer in the l-th spatial perception Transformer block maps the l-th sequence through a fully connected layer to a query matrix a key matrix and a value matrix and, based on the spatial perception token x spa , obtains the corresponding query vector from the query matrix Thereby, the similarity matrix of the l-th spatial perception Transformer block is calculated according to Equation (1) In formula (1), is the balance factor, represents the matrix transpose; Step 2.2.

3. The localization map generation mechanism layer in the l-th spatial perception Transformer block processes the similarity matrix using the Sigmoid activation function to obtain the l-th foreground activation probability map For X b After deforming the corresponding part in the foreground activation probability map a positioning map is obtained Step 2.2.

4. The location-guided cross-attention calculation layer in the l-th spatial perception Transformer block obtains the spatial query attention result of the l-th spatial perception Transformer block according to Equation (2) In Equation (2), * represents dot product; Step 2.2.5, After being processed successively by a skip connection layer, a Layer-Norm layer, an MLP layer, and a skip connection layer, the feature sequence output by the l-th spatial perception Transformer block is obtained 3. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program for supporting the processor to execute the weakly supervised object localization method described in Claim 1 or 2, and the processor is configured to execute the program stored in the memory.

4. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, it executes the steps of the weakly supervised object localization method described in Claim 1 or 2.

Citation Information

Patent Citations

  • Pollen image classification method based on cross attention distillation Transformer

    CN113887610A

  • Vision Transform network-based weak supervision instance segmentation method and system, and medium

    CN115359254A