An object localization system and method based on self-supervised Transformer

By adopting a target positioning system based on self-supervised Transformer in smart factories, the problems of target positioning complexity and dependence on labeled data in smart factories are solved, and higher positioning accuracy and coverage are achieved.

CN120088466BActive Publication Date: 2025-06-27NINGDE SKEQI INTELLIGENT EQUIP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510566998.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-27
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

Automation systems in smart factories face the problems of target positioning complexity and dependence on large amounts of labeled data. Especially in dynamically changing factory environments, existing methods are difficult to effectively distinguish between objects of interest and background noise.

Method used

A target positioning system based on self-supervised Transformer is adopted, which includes a coding module, a classifier, a discriminator, a pixel-level pseudo-label generator and a positioning module. The target positioning is achieved by generating attention maps, candidate recognition areas, pixel-level pseudo-labels and positioning maps, combined with a conditional random field loss function.

Benefits of technology

Improve the accuracy and coverage of target positioning, can handle background noise and object boundaries more effectively, and reduce dependence on labeled data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088466B_ABST
    Figure CN120088466B_ABST
Patent Text Reader

Abstract

The present invention discloses an object localization system and method based on self-supervised Transformer, belonging to the technical field of object localization. It includes an encoding module, a classifier, a discriminator, a pixel-level pseudo-label generator, and a localization module. Compared with the CNN-based CAM method, the present invention can cover multiple objects in the image based on random sampling through the pixel-level pseudo-label generator, thus providing richer information for the localization task. Compared with the Transformer-based method, the present invention can identify the most discriminative regions in each attention map by using a pre-trained external classifier. This method can ensure that the selected regions of interest (ROIs) can accurately cover the objects in the image, while excluding background noise, greatly improving the accuracy of localization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target localization, and more specifically, to a target localization system and method based on self-supervised Transformer. Background Art

[0002] In recent years, intelligent manufacturing technology has developed rapidly. In these highly automated environments, machine vision systems play a crucial role. They can not only improve production efficiency but also ensure quality control and safety supervision. In this context, weakly supervised object localization (WSOL) technology provides an effective method to identify and locate various objects on the production line, so as to achieve precise automated operation and monitoring. However, the automated systems in intelligent factories face the same challenges as deep learning models: the dependence on a large amount of labeled data. On the production line, obtaining accurate labels for each component or product is not only time-consuming and laborious but also costly. In addition, the dynamic changes and diverse product types in the factory environment further increase the complexity of object localization.

[0003] In this context, the method of self-supervised learning provides an innovative solution for intelligent factories. Self-supervised Transformer is a potential method that can generate good attention maps for different objects in an image without any supervision. However, these maps are usually class-agnostic because the model is unsupervised training, which causes the model to be unable to distinguish the objects of interest from background noise objects. This problem has become a challenge for researchers.

[0004] In the research of object localization, researchers have proposed weakly supervised object localization methods based on CNN. These methods usually rely on class activation mapping (CAM) to generate localization maps. However, these methods usually only focus on the significant regions shared among classes, and due to the small common activation regions between class samples, the localization coverage is limited. With the development of Transformer, researchers have designed weakly supervised object localization methods based on Transformer. These methods utilize the long-range dependencies of Transformer to generate effective localization maps, but they perform poorly in dealing with background noise and object boundaries.

[0005] Therefore, in view of the above technical problems, it is necessary to provide a target localization system and method based on self-supervised Transformer. Summary of the Invention

[0006] The purpose of the present invention is to provide a target localization system and method based on self-supervised Transformer to solve the above problems.

[0007] To achieve the above object, the technical solution provided by an embodiment of the present invention is as follows:

[0008] A target localization system based on self-supervised Transformer, the target localization system includes:

[0009] An encoding module for generating an embedding vector and an attention map corresponding to an image through a Vision Transformer model;

[0010] A classifier, based on the generated attention map, inputs the attention map, and performs classification training on the attention map through standard cross-entropy;

[0011] A discriminator for scoring a set of attention maps through an external classifier to generate candidate recognition regions;

[0012] A pixel-level pseudo-label generator for generating pixel-level pseudo-labels for the foreground and background based on the generated candidate recognition regions;

[0013] A localization module for inputting the embedding vector through a decoder, outputting a background map and a foreground map, and finally outputting a localization map.

[0014] Further, the classifier uses the attention map obtained by the encoding module as input and outputs , where ;

[0015] where f c is a function of the attention map , P r represents the output picture, the embedding vector E, the attention map , X represents the picture, and k represents the position in the output .

[0016] Further, the discriminator further includes

[0017] training an external classifier using a training data set and image class labels ;

[0018] traversing all the attention maps in the set to calculate the bounding boxes of all connected regions in the map;

[0019] scoring each bounding box through the obtained external classifier to obtain a scoring set , and the scoring set only shows the content within the box to the model ;

[0020] Sort the set and retain the top K numerically highest scores, and these K scores will be used as the output of the discriminator; K is greater than 0 and less than or equal to 3, thereby obtaining a set of candidate recognition regions.

[0021] Furthermore, through random position sampling, sample the first n pixels within the bounding box, where n is greater than 0 and less than or equal to 6;

[0022] For each sampled pixel, its position will be encoded in the image where these pixels will be encoded as pseudo-labels where the background label is 0, the foreground label is 1, and the positions of unknown labels are encoded as unknown; for the positions of the pixels, use the cross-entropy formula for training:

[0023] ;

[0024] where, is the pseudo-label, is the activation map.

[0025] Furthermore, the positioning module also includes

[0026] using the embedding vector as the input and outputting where are the background map and the foreground map respectively, represents the decoder, and finally outputs the positioning map,

[0027] To ensure the alignment of the activation map of the output S with the object boundary, use the conditional random field CRF loss, which takes into account both the proximity and color similarity between pixels;

[0028] ;

[0029] where, represents an affinity matrix, captures the color similarity and proximity between pixels in the image and and uses a Gaussian kernel to calculate ;

[0030] The positioning module introduces foreground and background pixels marked by a pixel-level pseudo-label generator to better perform object localization. Therefore, the final training loss function is as follows;

[0031] ;

[0032] where, , is the weighting coefficient, , the validation set is used in training to search .

[0033] On the other hand, a target localization method based on self-supervised Transformer includes

[0034] S1, generating an embedding vector and an attention map corresponding to the image through the Vision Transformer model;

[0035] S2, based on the generated attention map, inputting the attention map and classifying and training the attention map through standard cross-entropy;

[0036] S3, scoring a group of attention maps through an external classifier to generate candidate recognition regions;

[0037] S4, based on the generated candidate recognition regions, generating pixel-level pseudo-labels for the foreground and background through the candidate recognition regions;

[0038] S5, through the decoder, inputting the embedding vector, outputting the background map and the foreground map, and finally outputting the localization map.

[0039] Furthermore, in S2, it also includes

[0040] The classifier uses the attention map obtained by the encoding module as input and outputs , where ; where f c is the function of the attention map , P r represents the output picture, the embedding vector E, the attention map , X represents the picture, and k represents the position in the output .

[0041] Furthermore, in S3, it also includes

[0042] training an external classifier using the training dataset and image class labels ;

[0043] traversing all the attention maps in the set and calculating the bounding boxes of all connected regions in the map;

[0044] scoring each bounding box through the obtained external classifier to obtain a scoring set , and the scoring set only shows the content inside the box to the model ;

[0045] Sort the set and retain the top K highest-scoring values. These K values will be used as the output of the discriminator. K is greater than 0 and less than or equal to 3, thereby obtaining a set of candidate recognition regions.

[0046] Furthermore, in S4, it also includes sampling the first n pixels within the bounding box through random position sampling, where n is greater than 0 and less than or equal to 6.

[0047] The position of each sampled pixel will be encoded in the image where these pixels will be encoded as pseudo-labels where the background label is 0, the foreground label is 1, and the positions of unknown labels are encoded as unknown. For the position of the pixels, the cross-entropy formula is used for training:

[0048] ;

[0049] where, is the pseudo-label, is the activation map.

[0050] Furthermore, in S5, it also includes

[0051] using the embedding vector as the input and outputting where are the background map and the foreground map respectively, represents the decoder, and finally outputs the localization map,

[0052] To ensure the alignment of the activation map of the output S with the object boundary, a conditional random field (CRF) loss is adopted. This loss takes into account both the proximity and color similarity between pixels.

[0053] ;

[0054] where, represents an affinity matrix, captures the color similarity and proximity between pixels in the image and and is calculated using a Gaussian kernel ;

[0055] The localization module introduces foreground and background pixels marked by a pixel-level pseudo-label generator to better perform object localization. Therefore, the final training loss function is as follows:

[0056] ;

[0057] where, , is a weighting coefficient, , and the validation set is used to search during training .

[0058] Compared with the prior art, the advantages of the present invention are as follows:

[0059] Compared with the CNN-based CAM method of the present invention, the present invention can cover multiple objects in the image based on random sampling through a pixel-level pseudo-label generator, thereby providing richer information for the localization task. Compared with the Transformer-based method, the present invention can identify the most discriminative regions in each attention map by using a pre-trained external classifier. This method ensures that the selected regions of interest (ROIs) can accurately cover the objects in the image while excluding background noise, improving the accuracy of localization. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 is a flowchart of the object localization system based on self-supervised Transformer of the present invention;

[0061] Figure 2 is a flowchart of the object localization method based on self-supervised Transformer of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention; obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Embodiment 1:

[0063] Please refer to Figure 1 , an object localization system based on self-supervised Transformer, the object localization system includes: an encoding module for generating an embedding vector and an attention map corresponding to an image through a Vision Transformer model;

[0064] a classifier, based on the generated attention map, inputting the attention map, and performing classification training on the attention map through standard cross-entropy;

[0065] a discriminator for scoring a group of attention maps through an external classifier to generate candidate recognition regions;

[0066] a pixel-level pseudo-label generator for generating pixel-level pseudo-labels for foreground and background based on the generated candidate recognition regions;

[0067] A positioning module, which is used to input an embedded vector through a decoder, output a background map and a foreground map, and finally output a positioning map.

[0068] Specifically, an encoding module. This module uses an existing Vision Transformer model to generate an embedded representation of the corresponding image. The Vision Transformer is a publicly available model. Specifically, the encoding module uses the existing Vision Transformer model to generate an embedded representation of the corresponding image. Specifically, the input image is converted into an embedded vector and an attention map . In the present invention, the encoding part of the model is retained and frozen, and no longer trained for further use;

[0069] A classifier, that is, an external classifier. This module classifies the image based on the embedded representation of the encoding module. The classifier uses the attention map obtained by the encoding module as input and outputs , where , where f c is the attention map function, P r represents the output image, the embedded vector E, the attention map , X represents the image, and k represents the position in the output . The parameters of this module are trained for classification using standard cross-entropy ;

[0070] A discriminator, which generates candidate recognition regions based on the scores calculated by the external classifier for each image. The role of the discriminator is to generate discriminative suggestions from a set of attention maps to better localize objects in the model. The self-supervised Transformer will generate multiple attention maps rich in localization information and decompose the scene into multiple objects or parts of objects through the attention maps, but the attention maps and the localized objects are not associated with semantic meanings. To utilize these rich maps, the present invention designs a discriminator to associate the regions with semantics for obtaining reliable and discriminative candidate recognition regions (ROIs). The inputs of the discriminator are the image and its corresponding category , and the specific process is as follows:

[0071] Step 1: Use the training dataset and image category labels to train an external classifier , and this external classifier will be used as a scoring function to measure the possibility that a certain part of the image is discriminative;

[0072] Step 2: Traverse all the attention maps in the set and calculate the bounding boxes of all connected regions in the map. According to the common assumption in the class activation map, strong activations in the map are considered potential foregrounds, while weak activations are considered backgrounds, thus obtaining the bounding boxes;

[0073] Step 3: Use the external classifier obtained in Step 1 to score each bounding box in Step 2, obtaining a score set , where a higher value indicates a greater likelihood that the box contains discriminative parts associated with the image class label, and only show the content within the box to the model and achieve image perturbation by blurring the regions outside the bounding box, thereby suppressing the information outside the box;

[0074] Step 4: Sort the set and retain the top K with the highest scores. These K will be used as the output of the discriminator. Sort the set and retain the top K numerical values with the highest scores, and the K will be used as the output of the discriminator; the K is greater than 0 and less than or equal to 3, thereby obtaining a set of candidate recognition regions.

[0075] Through the above steps, a set of candidate recognition regions (ROIs) with rich and diverse information can be obtained, which has more advantages than the ROIs based on the class activation map because the method based on the class activation map is limited to a small discriminative region;

[0076] A pixel-level pseudo-label generator generates pixel-level pseudo-labels for the foreground and background based on the candidate recognition regions. The pixel-level pseudo-label generator uses a set of candidate recognition regions generated by the discriminator to generate pixel-level pseudo-labels for the foreground and background, providing support for the training of the localization module, and only using a few pixels at a time. The localization module transfers the knowledge learned from the pseudo-labels to other similar regions to fill in the blanks in the rest of the image;

[0077] The sampling of the foreground pixels is performed within the bounding box. Using random position sampling to avoid overfitting to the bounding box, considering the first n pixels within the bounding box for sampling helps to explore different parts of the region within the bounding box while focusing on potential pixels. Assuming the object is continuous, a polynomial distribution is used to select a pixel;

[0078] The sampling of the background pixels is performed relative to all the bounding boxes to ensure that the sampled background pixels are outside all the bounding boxes. Similar to the foreground pixels, by sorting all the activations of the attention map and selecting the lowest n pixels for sampling. Assuming the background region is uniformly distributed in the image, a uniform sampling is used to select a pixel;

[0079] Each of the sampled background and foreground pixels has its position encoded in the image, where these pixels will be encoded as pseudo-labels Among them, the background label is 0, the foreground label is 1, and the positions of unknown labels are encoded as unknown. For the pixels at the positions, the cross-entropy formula is used for training:

[0080] ;

[0081] Among them, is the pseudo-label, is the activation map. In actual training, for each image, several pixels are sampled from the foreground and background, and it is ensured that the sampled pixels are balanced.

[0082] The localization module is used to perform object localization. The localization module is a decoder for performing object localization, which uses the embedding vector as the input and outputs where are the background map and the foreground map respectively, represents the decoder, and finally outputs the localization map;

[0083] To ensure the alignment of the activation map of the output S with the object boundary, the conditional random field CRF loss is adopted, and this loss takes into account both the proximity and color similarity between pixels;

[0084] ;

[0085] Among them, represents an affinity matrix, captures the color similarity and proximity between the pixels in the image and and is calculated using a Gaussian kernel ;

[0086] The localization module introduces foreground and background pixels marked by the pixel-level pseudo-label generator to better perform object localization. Therefore, the final training loss function is as follows;

[0087] ;

[0088] Among them, , is the weighting coefficient, , and the validation set is used in training to search for .

[0089] Specifically, define the search space: determine the range of hyperparameters to be adjusted;

[0090] Training and evaluation loop: Train the model on the training set for each set of hyperparameters (or model architectures).

[0091] Calculate the performance metrics on the validation set based on the training loss function;

[0092] Select the best configuration based on the calculated performance metrics obtained: retain the hyperparameter combination with the best performance on the validation set.

[0093] Final evaluation: Conduct a final evaluation of the selected model using the test set to obtain .

[0094] Compared with the CNN-based CAM method, the present invention can cover multiple objects in the image based on random sampling through a pixel-level pseudo-label generator, thereby providing richer information for the localization task. Compared with the Transformer-based method, the present invention can identify the most discriminative regions in each attention map by using a pre-trained external classifier. This method ensures that the selected regions of interest (ROIs) can accurately cover the objects in the image while excluding background noise, improving the accuracy of localization.

[0095] See Figure 2 , a self-supervised Transformer-based object localization method, including,

[0096] S1, Generate the embedding vector and attention map corresponding to the image through the Vision Transformer model;

[0097] S2, Based on the generated attention map, input the attention map and perform classification training on the attention map through the standard cross-entropy;

[0098] S3, Score a group of attention maps through an external classifier to generate candidate recognition regions;

[0099] S4, Based on the generated candidate recognition regions, generate pixel-level pseudo-labels for the foreground and background through the candidate recognition regions;

[0100] S5, Through the decoder, input the embedding vector, output the background map and the foreground map, and finally output the localization map.

[0101] Furthermore, in S2, it also includes,

[0102] The classifier uses the attention map obtained by the encoding module as the input and outputs , where .

[0103] Furthermore, in S3, it also includes,

[0104] Train an external classifier using a training dataset and image class labels ;

[0105] Traverse all attention maps in the set and calculate the bounding boxes of all connected regions in the map;

[0106] Use the obtained external classifier to score each bounding box to obtain a score set , the score set only shows the content within the box to the model ;

[0107] Sort the set and retain the top K highest-scoring values, where the K values will be the output of the discriminator; K is greater than 0 and less than or equal to 3, thus obtaining a set of candidate recognition regions.

[0108] Furthermore, in S4, it also includes sampling the first n pixels within the bounding box through random position sampling, where n is greater than 0 and less than or equal to 6;

[0109] For each sampled pixel, its position will be encoded in the image where these pixels will be encoded as pseudo-labels in which the background label is 0, the foreground label is 1, and the positions of unknown labels are encoded as unknown; for the positions of the pixels, use the cross-entropy formula for training:

[0110] ;

[0111] where, is the pseudo-label, is the activation map.

[0112] Furthermore, in S5, it also includes,

[0113] using the embedding vector as input and outputting , where are the background map and the foreground map respectively, represents the decoder, and finally outputs the localization map,

[0114] To ensure the alignment of the activation map of the output S with the object boundary, use the conditional random field CRF loss, which takes into account both the proximity and color similarity between pixels;

[0115] ;

[0116] where, represents an affinity matrix, captures the image Medium pixel and the color similarity and proximity between them are calculated using a Gaussian kernel ;

[0117] The localization module introduces foreground and background pixels marked by a pixel-level pseudo-label generator to better perform object localization. Therefore, the final training loss function is as follows;

[0118] ;

[0119] wherein , is the weighting coefficient, , and the validation set is used to search for during training.

[0120] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed within the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.

[0121] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other implementation manners understandable by those skilled in the art.

Claims

1. A target localization system based on a self-supervised Transformer, characterized by: The target positioning system comprises: The encoding module is used to generate the embedding vector and attention map of the corresponding image through the Vision Transformer model; The classifier, based on the generated attention map, inputs the attention map and trains the classification of the attention map through standard cross entropy; Discriminator, used to score a set of attention maps through an external classifier to generate candidate recognition regions; train an external classifier using the training dataset and image category labels ; Traverse the collection All attention maps in the graph are used to calculate the bounding boxes of all connected regions in the graph; the obtained external classifier Score each bounding box and get a score set , score set Model only Display the content in the box; sort the set and retain the first K values ​​with the highest scores, which will be used as the output of the discriminator; K is greater than 0 and less than or equal to 3, thereby obtaining a set of candidate recognition areas; A pixel-level pseudo-label generator is used to generate pixel-level pseudo-labels for the foreground and background based on the generated candidate recognition areas; the first n pixels in the bounding box are sampled by random position sampling, where n is greater than 0 and less than or equal to 6; the position of each sampled pixel will be encoded in the image where these pixels will be encoded as pseudo labels In which the background label is 0, the foreground label is 1, and the position of the unknown label is encoded as unknown; for the position Pixels are trained using the cross entropy formula: ; in, is a pseudo label, is the activation map, image ; The positioning module is used to input the embedding vector through the decoder, output the background image and foreground image, and finally output the positioning map through the embedding vector As input, output ,in They are background image and foreground image respectively. Represents the decoder, and finally outputs the positioning map, In order to ensure that the activation map of the output S is aligned with the object boundary, the conditional random field CRF loss is adopted, which takes into account both the proximity and color similarity between pixels; ; in, represents an affinity matrix, Captured the image Medium Pixels and The color similarity and proximity between them are calculated using a Gaussian kernel ; The localization module introduces foreground and background pixels marked by a pixel-level pseudo-label generator to better perform target localization. Therefore, the final training loss function is as follows; ; in, , is the weighting coefficient, , use the validation set to search for .

2. The target positioning system based on self-supervised Transformer according to claim 1, characterized in that: The classifier uses the attention map obtained by the encoding module As input, output ,in , Among them, f c For the attention map Function, Pr represents the output image, embedding vector E, attention map , X represents the image, k represents the output The position in.

3. A target localization method based on self-supervised Transformer, characterized in that: include, S1, generates the embedding vector and attention map of the corresponding image through the Vision Transformer model; S2, based on the generated attention map, input the attention map and perform classification training on the attention map through standard cross entropy; S3, score a set of attention maps through an external classifier to generate candidate recognition regions; train an external classifier using the training dataset and image category labels ; Traverse the collection All attention maps in the graph are used to calculate the bounding boxes of all connected regions in the graph; the obtained external classifier Score each bounding box and get a score set , score set Model only Display the contents of the frame; Sort the set, retain the first K values ​​with the highest scores, and the K values ​​will be used as the output of the discriminator; the K is greater than 0 and less than or equal to 3, so as to obtain a set of candidate recognition areas; S4, based on the generated candidate recognition regions, generating pixel-level pseudo labels for the foreground and background through the candidate recognition regions; The first n pixels in the bounding box are sampled by random position sampling, where n is greater than 0 and less than or equal to 6; the position of each sampled pixel will be encoded in the image where these pixels will be encoded as pseudo labels In which the background label is 0, the foreground label is 1, and the position of the unknown label is encoded as unknown; for the position Pixels are trained using the cross entropy formula: ; in, is a pseudo label, is the activation map, image ; S5, through the decoder, input embedding vector, output background image and foreground image, and finally output positioning map, through embedding vector As input, output ,in They are background image and foreground image respectively. Represents the decoder, and finally outputs the positioning map, In order to ensure that the activation map of the output S is aligned with the object boundary, the conditional random field CRF loss is adopted, which takes into account both the proximity and color similarity between pixels; ; in, represents an affinity matrix, Captured the image Medium Pixels and The color similarity and proximity between them are calculated using a Gaussian kernel ; The foreground and background pixels marked by the pixel-level pseudo-label generator are introduced to perform target localization. Therefore, the loss function of the final training is as follows; ; in, , is the weighting coefficient, , use the validation set to search for .

4. The target localization method based on self-supervised Transformer according to claim 3, characterized in that: Also included in S2 is: The classifier uses the attention map obtained by the encoding module As input, output ,in , Among them, f c For the attention map Function, Pr represents the output image, embedding vector E, attention map , X represents the image, k represents the output The position in.

Citation Information

Patent Citations

  • Semantic enhancement feature fusion self-supervised transform-based whole-scene SAR mariculture multi-target extraction method

    CN117036934A

  • Medical image segmentation method and system based on basic model assisted semi-supervised learning

    CN119445120A