A Light Swin Image Classification Method Based on Sparse Attention InterWindow Block

Through the Light Swin Transformer model based on the sparse attention InterWindow block, the problem of high computational complexity of the image classification model is solved, efficient image information extraction and accurate classification are achieved, and the lightweight and accuracy of the model is improved.

CN118447313BActive Publication Date: 2025-07-22AOBO (JIANGSU) ROBOT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410560817.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-08
Publication Date
2025-07-22
Estimated Expiration
2044-05-08

AI Technical Summary

Technical Problem

The existing image classification model has high problems in computing complexity and parameter quantity, and the traditional Transformer architecture has problems of fixed scale and high computing cost in image classification, making it difficult to efficiently extract image information.

Method used

The Light Swin Transformer model based on the sparse attention InterWindow block is adopted. Through phased feature extraction and sparse attention control, the calculation complexity is reduced and the parameter utilization is improved. Combined with the feature fusion of Swin Transformer's local window attention and the feature fusion of InterWindow blocks, the accuracy of image classification is enhanced.

Benefits of technology

A lightweight image classification model is realized, which improves the accuracy and inference speed of image classification, and has better robustness and parameter efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118447313B_ABST
    Figure CN118447313B_ABST
Patent Text Reader

Abstract

Light Swin Image Classification Method Based on Sparse Attention InterWindow Block. The steps include: First, perform image augmentation processing such as Mixup, horizontal flipping, and random cropping on the image dataset, and label them accordingly. Then, use the Swin Transformer pre-trained model for training to obtain the low-dimensional hierarchical feature representation of the image. Next, use the sparse attention InterWindow block for training to strengthen the spatial feature representation of the image. Finally, use the classifier to process the extracted features to obtain the final image classification result. The Light Swin model proposed by the present invention combines the advantages of the CNN architecture and the Transformer architecture, realizing lightweight and efficient feature extraction. At the same time, in the pre-training stage, the present invention uses the l ∞ -norm to control the sparsity of the attention weights of the model, enabling the model to self-regulate the attention distribution and improving the accuracy and prediction speed of image classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention is applicable to the field of computer vision. Specifically, it relates to a sparsification method based on the Swin Transformer model and an attention module. Its superiority lies in using the sparse attention InterWindow block to enhance the extraction of image spatial features, thereby improving the accuracy of image classification. At the same time, it improves the parameter utilization rate of the model, realizes the lightweight of the model, and the achieved effect is relatively obvious. Background Art

[0002] Currently, image classification is a popular research direction in the field of computer vision. In this field, the core goal of each model establishment is to extract image information more effectively. By analyzing the image content and extracting key features, accurate classification can be achieved. Therefore, how to design a model that can efficiently extract image information has become the key to image classification research. The convolutional neural network (CNN) constructs a powerful feature representation by performing feature mapping on image pixel units and has become the main backbone network for various visual tasks. With the in-depth research of the convolutional neural network, the architecture of CNN has gradually increased, the connections are more extensive, and the convolutional form has become more complex.

[0003] In the field of natural language processing (Natural Language Processing), Transformer introduces an attention mechanism to perform dynamic interaction on each token or pixel unit, mimicking the human attention mechanism. This successful experience has promoted the application of the Transformer architecture in the field of computer vision, especially in image classification. However, the traditional Vision Transformer (VIT) has problems of fixed scale and high computational complexity in image classification. To solve these problems, Swin Transformer adopts a hierarchical mapping method to achieve dense prediction. This use of local window-based attention avoids the use of attention over the entire image and reduces the computational cost of Transformer on images. On the other hand, MobileViT effectively encodes local and global information using MobileViT blocks, which helps to learn better representations with fewer parameters and a simple training method.

[0004] SparseFormer proposes a preliminary framework for sparse visual recognition, which is mainly based on the strategy of operating on pixels in the latent space. SparseFormer further introduces the Rol adjustment technique, whose design concept is derived from an in-depth simulation of the human visual perception process. When humans initially quickly browse an image, they often cannot immediately understand all the content and need to adjust and focus their line of sight multiple times to effectively capture key information. In addition, ideally, the attention weight distribution should be sparse, and sparsity requires that the attention weights of the entire image input pixels be close to zero. Summary of the Invention

[0005] The present invention proposes a lightweight Swin Transformer model based on a sparse attention InterWindow block. By using the InterWindow Block, the model's ability to extract global features from deep feature maps is improved, while the computational complexity is reduced, the efficiency of using model parameters is increased, and the model inference speed is improved; sparse attention is introduced, and the l ∞ norm enables the model to self-regulate the sparsity degree of attention weights during the training process to find the optimal attention distribution.

[0006] Specific content of the invention:

[0007] 1. A Light Swin image classification method based on a sparse attention InterWindow block, characterized by comprising the following steps:

[0008] Step 1) Perform image enhancement on the image classification dataset;

[0009] Step 2) Preprocess the image based on the Swin Transformer pre-trained model, including the following steps:

[0010] Step 2.1) Use the input layer to input the image data obtained in step 1) into the SwinTransformer layer of the model. The shape of the input image data is:

[0011]

[0012] where represents the input image data, B is the batch size of the input image, H and W are the height and width of the input image, and 3 represents that the input image is RGB three-channel image data;

[0013] Step 2.2) Use the Swin Transformer block to perform feature extraction on the input image data in stages. In the first stage, the input feature map is scaled into Size, where C is the set initial dimension; starting from the second stage, the height and width of the input features are halved and the number of feature dimensions is doubled in each stage, and the shape of the finally output feature map is;

[0014]

[0015] where represents the output image data;

[0016] Step 3) Input the data obtained in Step 2) into the InterWindow block, including the following steps:

[0017] Step 3.1) The InterWindow block first divides the feature map into non-overlapping windows, performs an attention calculation within the windows, and adds the result to the shortcut branch. The shape of the output image data remains

[0018] Step 3.2) Flatten the image data into a sequence according to the relative positions of the non-overlapping windows. The length of the sequence is the number of small windows, set the height and width of the small windows to M, the sequence length is The number of sequences is M 2 ones, perform an attention calculation on the sequence data and refold the sequence data back, with the shape of

[0019] Step 3.3) Concatenate the data obtained in Step 3.2) and the original data along the dimension direction. After concatenation, the shape is Perform feature fusion through a convolutional layer. After fusion, the shape is

[0020] Step 3.4) Downsample the data after fusion in Step 3.3) again to Repeat the operations in Step 3.1) - Step 3.3) to obtain the output:

[0021]

[0022] where X represents the output image data;

[0023] Step 4) Linearly flatten the data obtained in Step 3) into the shape of Finally, obtain the output result through a linear layer:

[0024] p = (B, N)

[0025] where p represents the probability of the model predicting the category, and N is the number of output categories;

[0026] Step 5) For the predicted data in each batch, through l∞ The norm regularizes the attention weights and controls the sparsity of its distribution through the hyperparameter η

[0027]

[0028] where loss represents the loss function for model training, N represents the number of predicted categories, W-MSA-weights represents the weight matrix for attention calculation in step 3.1), IW-MSA-weights represents the weight matrix for attention calculation in step 3.2), and y i is the one-hot encoded true label, and p i represents the model prediction probability;

[0029] Step 6) Set the learning rate and the number of iterations, and train on the dataset to obtain the trained model;

[0030] Step 7) Input the predicted image into the model, set the batch size B to 1 at the same time, and sort the output result p=(1,N) in descending order in the second dimension to obtain the category with the largest prediction as the image category;

[0031] 2. The specific steps for image enhancement of the image dataset are as follows:

[0032] Step 1.1) Use the Mixup method to perform image enhancement operations on the image data: The calculation formulas for the image and label after the Mixup operation:

[0033]

[0034] where x i and x j represent the original image data in the dataset, y i and y j correspond to the one-hot labels of the images, λ is the probability value following the Beta distribution, with a range of 0-1, and represent the newly generated image data and the corresponding image labels after the Mixup operation;

[0035] Step 1.2) Perform random cropping, horizontal flipping, and normalization preprocessing operations on the input data of the dataset images; By randomly selecting the image data in the dataset, an image classification dataset after image enhancement is obtained.

[0036] The present invention has the following advantages and beneficial effects:

[0037] Combined with the feature extraction ability of the Swin Transformer pre-trained model for local parts of the model, and the spatial feature extraction ability complemented by the InterWinodw block. By applying the InterWindow block in the high-dimensional space, not only the computational complexity and the number of model parameters are reduced, but also the model's ability to capture global information is greatly enhanced, improving the accuracy of image classification;

[0038] At the same time, in order to control the distribution of attention weights, the concept of sparse attention is utilized. By using the l ∞ norm to control the distribution of attention weights, the sparsity of the attention weight distribution is achieved, enabling the model to accurately focus on the key parts in the image and further improving the performance of the model. Therefore, the image classification model of the present invention has better robustness and better accuracy than traditional CNN and Transformer models, and also has a faster inference speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is the overall flowchart of the Light Swin image classification model of the present invention;

[0040] Figure 2 is the structural diagram of the sparse attention InterWindow block of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] Next, in combination with the drawings and embodiments, the advantages and objectives of the present invention will be further described in detail. It should be understood that the description herein is only for explaining the present invention and is not used to limit the present invention.

[0042] The Light Swin image classification method based on the sparse attention InterWindow block proposed by the present invention first performs image enhancement processing on the image, then uses the Swin Transformer pre-trained model for training to obtain low-level image feature representations, and then uses the sparse attention InterWindow block for training to obtain the spatial feature representations of the image. At the same time, restrictions are imposed on the attention layer during training to make its weight distribution show sparsity. Finally, the image classification result is obtained using the classifier. Attached Figure 1 is the overall flowchart of the Light Swin image classification model of the present invention. Attached Figure 2 is the structural diagram of the sparse attention InterWindow block of the present invention, including the following steps:

[0043] Step 1) Perform image enhancement on the image classification dataset;

[0044] Step 2) Preprocess the image based on the Swin Transformer pre-trained model, including the following steps:

[0045] Step 2.1) Use the input layer to input the image data obtained in Step 1) into the SwinTransformer layer of the model. The shape of the input image data is:

[0046]

[0047] where represents the input image data, B is the batch size of the input image, H and W are the height and width of the input image, and 3 indicates that the input image is RGB three-channel image data;

[0048] Step 2.2) Use the Swin Transformer block to perform feature extraction on the input image data in stages. In the first stage, the input feature map is scaled to size, where C is the set initial dimension; starting from the second stage, every time a stage is experienced, the height and width of the input features are halved, and the number of feature dimensions doubles. The shape of the finally output feature map is;

[0049]

[0050] where represents the output image data;

[0051] Step 3) Input the data obtained in Step 2) into the InterWindow block, including the following steps:

[0052] Step 3.1) The InterWindow block first divides the feature map into non-overlapping windows, performs an attention calculation within the windows, and adds the result to the shortcut branch. The shape of the output image data remains

[0053] Step 3.2) Flatten the image data into a sequence according to the relative positions of the non-overlapping windows. The length of the sequence is the number of small windows. Set the height and width of the small windows to M, and the sequence length is The number of sequences is M 2 Perform an attention calculation on the sequence data and refold the sequence data back, with the shape of

[0054] Step 3.3) Concatenate the data obtained in Step 3.2) and the original data along the dimension direction. After concatenation, the shape is Perform feature fusion through the convolutional layer. After fusion, the shape is

[0055] Step 3.4) Downsample the data after fusion in Step 3.3) again to Repeat the operations in steps 3.1) - 3.3) to obtain the output:

[0056]

[0057] where X represents the output image data;

[0058] Step 4) Linearly flatten the data obtained in step 3) into the shape of, and finally obtain the output result through the linear layer:

[0059] p = (B, N)

[0060] where p represents the probability of the predicted class by the model, and N is the number of output classes;

[0061] Step 5) For the predicted data in each batch, through the l ∞ norm, regularize the attention weights, and control the sparsity of its distribution through the hyperparameter η

[0062]

[0063] where loss represents the loss function of model training, N represents the number of predicted classes, W-MSA-weights represents the weight matrix for attention calculation in step 3.1), IW-MSA-weights represents the weight matrix for attention calculation in step 3.2), y i is the one-hot encoded true label, and p i represents the model predicted probability;

[0064] Step 6) Set the learning rate and the number of iterations, and train on the dataset to obtain the trained model;

[0065] Step 7) Input the predicted image into the model, and at the same time set the batch number B to 1. Sort the output result p = (1, N) in the second dimension from largest to smallest to obtain the class with the largest prediction as the image class;

[0066] 2. The specific steps for image enhancement of the image dataset are as follows:

[0067] Step 1.1) Use the Mixup method to perform image enhancement operations on the image data: The calculation formulas for the image and label after Mixup operation:

[0068]

[0069] where x i and x j represent the original image data in the dataset, and y i and y jOne-hot label of the corresponding image, λ is a probability value following the Beta distribution, with a range of 0-1, and are represented as the newly generated image data and the corresponding image labels after the Mixup operation;

[0070] Step 1.2) Perform random cropping, horizontal flipping, and normalization preprocessing operations on the input data of the dataset images; by randomly selecting the image data in the dataset, obtain the image classification dataset after image enhancement.

Claims

1. A Light Swin image classification method based on a sparse attention InterWindow block, characterized in that It includes the following steps: Step 1) Perform image enhancement on the image classification dataset; Step 2) Preprocess the image based on the Swin Transformer pre-trained model, including the following steps: Step 2.1) Use the input layer to input the image data obtained in Step 1) into the Swin Transformer layer of the model. The shape of the input image data is: Among them represents the input image data, B is the batch size of the input image, H and W are the height and width of the input image, and 3 indicates that the input image is RGB three-channel image data; Step 2.2) Use the Swin Transformer block to perform feature extraction on the input image data in stages. In the first stage, the input feature map is scaled to size, where C is the set initial dimension; starting from the second stage, the height and width of the input features are halved and the number of feature dimensions is doubled after each stage. The shape of the finally output feature map is: Among them represents the output image data; Step 3) Input the data obtained in Step 2) into the InterWindow block, including the following steps: Step 3.1) The InterWindow block first divides the feature map into non-overlapping windows, performs an attention calculation within the windows, and adds the result to the shortcut branch. The output image data still has the shape of Step 3.2) Flatten the image data into a sequence according to the relative positions of non-overlapping windows. The length of the sequence is the number of small windows. Set the height and width of the small window to M, and the sequence length is The number of sequences is M 2 pieces. Perform an attention calculation on the sequence data and refold the sequence data back to a shape of Step 3.3) Combine the data obtained in Step 3.2) and the original data along the dimension direction. After combination, the shape is Perform feature fusion through a convolutional layer. After fusion, the shape is Step 3.4) Downsample the data after fusion in Step 3.3) again to Repeat the operations in Step 3.1) - Step 3.3) to obtain the output: where X represents the output image data; Step 4) Flatten the data obtained in step 3) linearly into shape, and finally obtain the output result through a linear layer: p = (B, N) where p represents the probability of the model predicting the class, and N is the number of output classes; Step 5) For the predicted data in each batch, through the l ∞ -norm, regularize the attention weights and control the sparsity of its distribution through the hyperparameter η where loss represents the loss function for model training, N represents the number of predicted categories, W-MSA-weights represents the weight matrix for attention calculation in step 3.1), IW-MSA-weights represents the weight matrix for attention calculation in step 3.2), y i is the one-hot encoded true label, and p i represents the model prediction probability; Step 6) Set the learning rate and the number of iterations, and train on the dataset to obtain the trained model; Step 7) Input the predicted image into the model, and at the same time set the batch size B to 1. Sort the output result p = (1, N) in the second dimension from largest to smallest to obtain the class with the largest prediction as the image class.

2. The Light Swin image classification method based on the sparse attention InterWindow block according to claim 1, wherein The specific steps for performing image enhancement on the image dataset are as follows: Step 1.1) Use the Mixup method to perform image enhancement operations on the image data: Calculation formulas for the image and label after Mixup operation: where x i and x j represent the original image data in the dataset, y i and y j correspond to the one-hot labels of the images, λ is a probability value following a Beta distribution, ranging from 0 to 1, and represent the newly generated image data and the corresponding image labels after the Mixup operation; Step 1.2) Perform random cropping, horizontal flipping, and normalization preprocessing operations on the input data of the dataset images; By randomly selecting the image data in the dataset, obtain the image classification dataset after image enhancement.

Citation Information

Patent Citations

  • Picture feature extraction method and device, target re-identification method and device and electronic equipment

    CN111310518A

  • Image target detection method based on multi-source information fusion

    CN117830788A