Visual converter model evaluation method based on accumulative attention map and Jaccard similarity coefficient

Through the cumulative attention map and JACAD similarity coefficient evaluation method, the problem of high computational complexity of vision transformer models on resource-constrained devices is solved, providing a more comprehensive basis for model selection, and improving the interpretability and reliability of model selection.

CN120375005AActive Publication Date: 2025-07-25杭州智元研究院有限公司
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510850585.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-25
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

The existing vision converter models have high computational complexity on resource-constrained devices and are difficult to select the best model through simple final performance indicators, making it difficult to make choices in practical applications.

Method used

The evaluation method of cumulative attention map and Jaccard similarity coefficient is used to record the attention map of the visual transformer model and calculate the Jaccard similarity coefficient, compare the differences in attention mechanisms of different models, and select the optimal model.

Benefits of technology

An evaluation method other than the final performance indicators is provided, which enhances the interpretability and reliability of model selection, can significantly reflect the attention mechanism differences between models, and helps to select a better visual transformer model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375005A_ABST
    Figure CN120375005A_ABST
Patent Text Reader

Abstract

The invention discloses a visual converter model evaluation method based on an accumulative attention map and a Jaccard similarity coefficient. The method comprises the following steps of performing initialization setting; training a visual converter model which is used as a reference and is based on complete attention; training a plurality of alternative visual converter models based on grouping attention; recording an attention map of a key layer in the visual converter model in the forward propagation process of a small amount of calibration data; determining an accumulated attention map of a key layer in the visual converter model; determining a Jaccard similarity coefficient; and selecting an optimal visual converter model according to the Jaccard similarity coefficient. According to the method, the difference index can be improved in each group attention model with tiny final effect difference, and a more excellent model with higher interpretability can be found more favorably.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of visual perception, and particularly relates to a method for evaluating a vision transformer model based on a cumulative attention map and a Jaccard similarity coefficient. Background Art

[0002] In the field of visual perception, vision backbones based on the Transformer have achieved remarkable success in recent years. Since the Vision Transformer (ViT) was proposed, the Transformer architecture has successfully introduced the self-attention mechanism in natural language processing into computer vision tasks. In the architecture of the vision transformer, an image is first divided into a series of image patches of a fixed size. After information extraction, these image patches form image tokens, and then these image tokens are regarded as sequence data and input into the Transformer, thus realizing the modeling of global image information. This innovation has enabled the Transformer to perform excellently in tasks such as image classification, object detection, and semantic segmentation.

[0003] However, the huge computational consumption and memory occupancy have hindered its further development. Specifically, the self-attention mechanism of the Transformer needs to calculate the relationships between all image patches, and its computational complexity is the square of the number of image tokens. When processing high-resolution images, the number of image tokens will be very large, resulting in an exponential increase in the demand for computational and storage resources. This poses a huge challenge in practical applications, especially on resource-constrained devices such as mobile devices and embedded systems.

[0004] Therefore, many researchers have introduced grouped attention to replace full attention and proposed many vision transformer models based on grouped attention. The core idea of these models is to divide all image tokens into several groups and calculate self-attention only within each group, thereby reducing the computational complexity and significantly reducing the overhead. Grouped attention not only reduces the computational complexity but also can capture local features, similar to the local receptive fields in convolutional neural networks. This method has shown a good balance between performance and efficiency in practical applications. Among the methods proposed by these models, although other improvements have also been introduced simultaneously, such as new position representation methods (such as relative position encoding, coordinate embedding), additional convolutional layers for feature extraction, adjustment of normalization layers, and optimization of activation functions, etc., the key difference between different models lies in how to group the image tokens. For example, the Swin Transformer introduces a sliding window mechanism to limit the self-attention calculation within a local window and gradually fuse global information through cross-window connections. The CSwin Transformer divides the feature map into multiple cross-shaped windows and calculates self-attention separately in the horizontal and vertical directions. The cross-shaped design enables the model to capture longer-range dependencies in each layer.

[0005] However, the final model performances among different models are very similar. For example, in the image classification task on the ImageNet dataset (a large-scale benchmark dataset for computer vision tasks) under the same other conditions, the difference in the top-1 classification accuracy of different grouped attention models is only about 1%. Such a small performance difference may not be significant in practical applications and may even not be statistically significant. This makes it difficult to choose which grouped attention method to use in actual use.

[0006] The reason for this phenomenon may be that these models are very similar in overall architecture and only differ in the way of grouping image tokens. In addition, factors such as training strategies, hyperparameter settings, and data preprocessing may have a greater impact on model performance than the grouping strategy itself. Therefore, simply relying on final metrics such as top-1 classification accuracy to evaluate the model is not sufficient to comprehensively reflect the actual performance and application value of the model. Summary of the Invention

[0007] The object of the present invention is to provide a method for evaluating a vision transformer model based on cumulative attention maps and Jaccard similarity coefficients to solve the above technical problems.

[0008] To achieve the object of the present invention, the present invention provides a method for evaluating a vision transformer model based on cumulative attention maps and Jaccard similarity coefficients, including the following steps:

[0009] Step 1: Select a specific visual perception task, dataset, and model architecture settings to complete the initialization settings;

[0010] Step 2: Utilize the visual perception task, dataset, and model architecture settings to train a reference vision transformer model based on full attention;

[0011] Step 3: Utilize the visual perception task, dataset, and model architecture settings to train multiple alternative vision transformer models based on grouped attention;

[0012] Step 4: Select a small amount of calibration data from the dataset, input the small amount of calibration data into the vision transformer model based on full attention and the vision transformer models based on grouped attention, and record the attention maps of the key layers in the vision transformer models during the forward propagation of the small amount of calibration data;

[0013] Step 5: Determine the cumulative attention maps of the key layers in the vision transformer models according to the attention maps of the key layers in the vision transformer models;

[0014] Step 6: Compare the cumulative attention maps of each alternative vision transformer model based on grouped attention with the cumulative attention maps at the corresponding positions of the vision transformer model to determine the Jaccard similarity coefficient;

[0015] Step 7: Select the optimal vision transformer model according to the Jaccard similarity coefficient.

[0016] Compared with the prior art, the significant progress of the present invention lies in: 1) The present invention proposes a method for evaluating vision transformer models based on cumulative attention maps and Jaccard similarity coefficients, providing a new perspective for model selection in addition to the final performance indicators. This method not only considers the overall performance of the model but also conducts a detailed analysis of the model from the perspective of the attention mechanism, thus providing a more comprehensive basis for model selection; 2) The present invention utilizes the characteristics of the attention mechanism in the vision transformer model to intuitively display the attention distribution inside the model through cumulative attention maps. This evaluation method based on attention maps makes the behavior of the model more transparent and easy to understand. In addition, the Jaccard similarity coefficient, as an intuitive metric standard, further enhances the interpretability of the evaluation results; 3) In existing vision transformer models, when only relying on final performance indicators such as accuracy and recall for comparison, the differences between models may be very small and difficult to distinguish between good and bad. However, the present invention can significantly reflect the differences in the attention mechanism between different models by calculating the Jaccard similarity coefficient between cumulative attention maps. This difference is more obvious than traditional performance indicators, thus providing a more intuitive and reliable basis for model selection.

[0017] To more clearly illustrate the functional characteristics and structural parameters of the present invention, the following further explains in conjunction with the accompanying drawings and specific embodiments. Description of the Drawings

[0018] The accompanying drawings described herein are used to provide a further understanding of the present invention and form a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0019] Figure 1 is the flowchart of the present invention;

[0020] Figure 2 is the line graph of the Jaccard similarity coefficient of the cumulative attention maps of three vision transformers based on grouped attention and the vision transformer model based on full attention in the embodiments of the present invention. Specific Embodiments

[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0022] A method for evaluating a vision transformer model based on cumulative attention maps and Jaccard similarity coefficients of the present invention, in combination with Figure 1 , includes the following steps:

[0023] Step 1: Select specific visual perception tasks, datasets, and model architecture settings to complete the initialization settings;

[0024] The specific initialization settings include:

[0025] Step 1-1: Select specific visual perception tasks and datasets as the basis for subsequent comparison; in this embodiment, the image classification task and the ImageNet dataset are selected.

[0026] Step 1-2: Select the model architecture settings. Since various forms of comparison will be performed on each model later, the architecture design of the models needs to be unified. In the vision transformer model based on full attention and the vision transformer model based on grouped attention, the number of stages and model blocks are equal; the number of image tokens within the model blocks at corresponding positions is equal; in the vision transformer model based on grouped attention, the number of groups within the model blocks at corresponding positions is equal. Referring to the common designs in various vision transformers, in this embodiment, the image size for training and inference is 224 × 224, the number of image tokens in the four stages are 56×56, 28×28, 14×14, and 7×7 respectively, and the group size of grouped attention is 49. Considering representativeness and diversity, three classic vision transformer models are selected in this embodiment, namely the Sliding Window Transformer, the Shuffle Transformer, and the Cross-Domain Transformer.

[0027] Step 2: Use the visual perception task, dataset, and model architecture settings to train a reference vision transformer model based on full attention. According to the initialization settings, use the visual perception task, dataset, and model architecture settings to select appropriate training strategies and hyperparameter settings, and train the vision transformer model based on full attention until convergence.

[0028] Step 3: Use the visual perception task, dataset, and model architecture settings to train multiple alternative vision transformer models based on grouped attention. According to the initialization settings, use the visual perception task, dataset, and model architecture settings to maintain the same training settings as the vision transformer model based on full attention, and train the multiple alternative vision transformer models based on grouped attention until convergence. Unify the experimental conditions to fairly evaluate the impact of the grouped attention mechanism on the model performance. In this embodiment, the Sliding Window Transformer, the Shuffle Transformer, and the Cross-Domain Transformer are trained until convergence.

[0029] Step 4: Select a small amount of calibration data from the dataset, input the small amount of calibration data into the vision transformer model based on full attention and the vision transformer models based on grouped attention, and record the attention maps of the key layers in the vision transformer models during the forward propagation of the small amount of calibration data.

[0030] Step 4-1: Select a small amount of calibration data from the dataset. The calibration data is randomly selected from the training set data or validation set data in the dataset. In this embodiment, 512 pieces of data are selected.

[0031] Step 4-2: Input the small amount of calibration data into the Vision Transformer model. During the forward propagation of the small amount of calibration data, record the attention maps of the key layers in the Vision Transformer model. The attention maps exist in the calculation process of self-attention, which shows how the model distributes attention when processing input data, that is, the degree of attention of each image token to other image tokens. Specifically, in the self-attention mechanism, when inputting image tokens, query (Q), key (K), and value (V) vectors are generated. The attention maps assign weights to the calculation of the value V vectors. The attention map of the layer is as follows:

[0032] ;

[0033] Among them, is the dimension of the key vector, used for scaling to prevent numerical values from being too large; the element of the attention map represents the degree of attention of the th image token to the th image token in the

[0034] layer. The larger the value, the more important it is. In this embodiment, the layers in the third stage are selected. There are 18 layers in total in the third stage.

[0035] Step 5: Determine the cumulative attention map of the key layers in the Vision Transformer model according to the attention maps of the key layers in the Vision Transformer model.

[0036] The attention map in the previous step is an incomplete representation of the mutual attention relationship between each image token in the Vision Transformer model. To represent the mutual attention relationship between each image token in the Vision Transformer model more accurately, the calculation of the cumulative attention map should be adopted. Based on the attention map, in the layer, the cumulative attention of the th image token to the th image token is as follows, where k represents the index of all possible intermediate image tokens in the current

[0037] ;

[0038] In addition, the residual connection transmits the information of the image tokens from the previous layer to the corresponding position in the next layer. Therefore, it is necessary to add the identity matrix I to the cumulative attention map of the current layer. Finally, the entire cumulative attention map is calculated by matrix multiplication:

[0039] ;

[0040] Among them, is the cumulative attention map of the (l - 1)-th layer.

[0041] Step 6: Compare the cumulative attention maps of each alternative group attention-based vision transformer model with the corresponding position cumulative attention map of the vision transformer model to determine the Jaccard similarity coefficient;

[0042] The Jaccard similarity coefficient is defined as the ratio of the size of the intersection of two sets to the size of the union, and its value range is between 0 and 1. The larger the value, the more similar the two sets are; the first attention map and the second attention map The similarity between them is expressed as:

[0043] ;

[0044] Between the full-attention-based vision transformer model and the group attention-based vision transformer model, calculate the Jaccard similarity coefficient of their corresponding layer cumulative attention maps. The specific steps are as follows:

[0045] Obtain the attention map: Extract the attention maps of the corresponding layers from two different vision transformer models;

[0046] Calculate the cumulative attention map: For each layer, summarize or average the results of all attention heads of this layer to obtain the cumulative attention map of this layer;

[0047] Calculate the intersection and union: For each pair of cumulative attention maps (from the same layer of two models), calculate their intersection and union;

[0048] Calculate the Jaccard similarity coefficient: Use the above formula , calculate the Jaccard similarity coefficient of each pair of cumulative attention maps.

[0049] Step 7: Select the optimal vision transformer model according to the Jaccard similarity coefficient: Compare the Jaccard similarity coefficients of each group attention-based vision transformer model with the full-attention-based vision transformer model to reveal which vision transformer model's attention mechanism is closer to the full vision transformer model; since such a comparison is carried out among each vision transformer model, no threshold is set, and the evaluation is based on the comparison results among each vision transformer model. Combining Figure 2 , in the comparison of the similarity coefficients of each layer in the third stage, the similarity coefficients of the cross-domain transformer and the shuffle transformer are much higher than those of the sliding window transformer, which also shows that the attention forms of the cross-domain transformer and the shuffle transformer are closer to the full-attention-based vision transformer model, so they are more applicable vision transformer models under this evaluation criterion.

[0050] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.

[0051] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for evaluating a vision transformer model based on cumulative attention maps and Jaccard similarity coefficients, characterized in that The following steps are involved: Step 1: Select a specific visual perception task, data set, model architecture setting, and complete the initialization settings; Step 2: Using the visual perception task, dataset, and model architecture settings, train a reference full attention-based visual transformer model; Step 3: using the visual perception task, data set, and model architecture settings, train multiple alternative group-attention-based visual transformer models; Step 4: Select a small amount of calibration data from the data set, input the small amount of calibration data into the visual transformer model based on complete attention and the visual transformer model based on grouped attention, and record the attention map of the key layer in the visual transformer model during the forward propagation of the small amount of calibration data; Step 5: determining a cumulative attention map of the key layer in the visual transformer model according to the attention map of the key layer in the visual transformer model; Step 6, comparing the cumulative attention map of each candidate group-attention-based visual transformer model with the corresponding position cumulative attention map of the visual transformer model to determine the Jaccard similarity coefficient; Step 7: Select the optimal visual transformer model according to the Jaccard similarity coefficient.

2. The method for evaluating a vision transformer model based on a cumulative attention map and a Jaccard similarity coefficient according to claim 1, wherein, The initialization settings of step 1 specifically include: Step 1-1: Select specific visual perception tasks and datasets; Step 1-2, select the model architecture settings and unify the model architecture design; in the full attention-based visual transformer model and the group attention-based visual transformer model, the number of stages and model blocks are equal; the number of image tags in the corresponding position model blocks is equal; in the group attention-based visual transformer model, the number of groups in the corresponding position model blocks is equal.

3. The evaluation method of a vision transformer model based on cumulative attention maps and Jaccard similarity coefficients according to claim 1, characterized in that, The step 2 specifically includes: according to the initialization settings, using the visual perception task, data set, and model architecture settings, selecting appropriate training strategies and hyperparameter settings, and training the full attention-based visual transformer model until convergence.

4. A method for evaluating a vision transformer model based on cumulative attention maps and Jaccard similarity coefficients according to claim 3, characterized in that, The step 3 specifically comprises: according to the initialization settings, utilizing the visual perception task, data set, and model architecture settings, maintaining the same training settings as the full attention-based visual transformer, and training multiple alternative grouped attention-based visual transformer models until convergence.

5. A method for evaluating a vision transformer model based on a cumulative attention map and Jaccard similarity coefficient according to claim 1, wherein The step 4 specifically includes: Step 4-1, selecting a small amount of calibration data in the data set, wherein the calibration data is selected in small amounts from the training set data or the validation set data in the data set; Step 4-2, input the small amount of calibration data into the vision transformer model, and record the attention maps of the key layers in the vision transformer model during the forward propagation of the small amount of calibration data, and the attention maps exist in the calculation process of self-attention; specifically, in the self-attention mechanism, query Q, key K, and value V vectors are generated when inputting image tokens, and the attention maps assign weights to the calculation of the value V vectors. The attention map of the layer is: ; Among them, is the dimension of the key vector, used for scaling to prevent excessive numerical values; the elements of the attention map represent that in the th layer, the th image token pays attention to the th image token, and the larger the value, the more important it indicates.

6. A method for evaluating a vision transformer model based on a cumulative attention map and the Jaccard similarity coefficient according to claim 5, characterized in that, The step 5 specifically includes: Based on the attention map, in the -th layer, for the -th image token, the cumulative attention to the -th image token is as follows, where k represents all possible intermediate image token indices in the current l-th layer: ; Cumulative attention map of the current layer Add the identity matrix I, and finally the entire cumulative attention map is calculated using matrix multiplication: ; Among them, is the cumulative attention map of the (l - 1)-th layer.

7. The evaluation method of a vision transformer model based on cumulative attention maps and Jaccard similarity coefficients according to claim 6, characterized in that The step 6 is specifically as follows: The Jaccard similarity coefficient is defined as the ratio of the size of the intersection of two sets to the size of the union, with a numerical range between 0 and 1. The larger the value, the more similar the two sets are; the first attention map and the second attention map The similarity between them is expressed as: ; Between the visual transformer model based on full attention and the visual transformer model based on grouped attention, the Jaccard similarity coefficient of the cumulative attention graphs of their corresponding layers is calculated, and the specific steps are as follows: Obtaining Attention Maps: Extracting attention maps of corresponding layers from two different visual transformer models; Calculate the cumulative attention map: For each layer, summarize or average the results of all attention heads of the layer to get the cumulative attention map of the layer; Calculate intersection and union: For each pair of accumulated attention maps, calculate their intersection and union; Calculate the Jaccard similarity coefficient: Use the above formula , and calculate the Jaccard similarity coefficient for each pair of cumulative attention maps.

8. A method for evaluating a vision transformer model based on cumulative attention maps and Jaccard similarity coefficients according to claim 1, wherein, Specifically, step 7 is as follows: Compare the Jaccard similarity coefficients of each vision transformer model based on grouped attention with the vision transformer model based on full attention, and determine that the attention mechanism of one vision transformer model is closer to that of the full vision transformer model; Since such a comparison is carried out among the vision transformer models, no threshold is set, and the results are evaluated through the mutual comparison of the vision transformer models.

Citation Information

Patent Citations

  • Cone-beam X-ray luminescence tomography method based on grouping attention residual network

    CN113288188A

  • Memory Grounded Conversational Reasoning and Question Answering for Assistant Systems

    US20200410012A1

  • Leveraging Redundancy in Attention with Reuse Transformers

    US20230112862A1

  • Multi-granular clustering-based solution for key-value cache compression

    US20250094712A1

  • Efficient token pruning in transformer-based neural networks

    US20250124105A1