An Image Recognition Method Based on a Lightweight Attention Mechanism with Linear K-Value
By using a linear K value lightweight attention mechanism in the self-attention mechanism, the problem of excessive calculation of the existing attention mechanism is solved, and a more stable recognition speed and saving of calculation amount is achieved.
Patent Information
- Application Number
- CN202311221334.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-09-21
AI Technical Summary
The existing attention mechanism is too much computational in the field of computer vision, resulting in unstable recognition speed.
An image recognition method based on a linear K value lightweight attention mechanism is proposed. By allowing the K value to pass through a linear layer in the self-attention mechanism, the process of multiplying the Q value and the K value is reduced, thereby reducing the calculation amount.
On the basis of ensuring similar effects to the existing attention mechanism, half of the calculation amount is saved and the recognition speed is more stable.
Smart Images

Figure CN117173413B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and pattern recognition, and in particular to an image recognition method based on a linear K-value lightweight attention mechanism. Background Art
[0002] The attention mechanism (Self-Attention) enables the model to adjust its attention and weights according to different parts of the input in order to more accurately understand and process information. It is mainly applied to neural network structures such as Transformer and was initially proposed for the field of natural language processing. However, in recent years, more and more research has shown that the Transformer model can also achieve good performance in the field of computer vision. Vision Transformer (ViT) is a computer vision model based on Transformer. ViT uses a pre-training technique similar to that in the field of natural language processing. By pre-training the model with a large amount of unlabeled data, good results have been achieved in various computer vision tasks, demonstrating the application prospects of Transformer in the field of computer vision. The self-attention mechanism is the most basic module of Transformer. In the self-attention mechanism, the input vector first passes through a linear layer to generate three vectors Q, K, and V, and then these three vectors are used for the calculation of the self-attention mechanism. The self-attention mechanism can capture context information by learning the relationships between different positions in the input sequence. Each input position can interact with other positions and participate in the calculation in different ways.
[0003] The most widely used attention mechanism of Transformer in the visual field is the self-attention mechanism used in ViT. An important reason for its failure to be widely applied in practice is its huge computational complexity of O(N 2 )), where N is the number of pixels in the photo. In the visual field where a photo often has tens of thousands of pixel values, such a computational complexity is extremely terrifying. To cope with the huge computational complexity of Transformer, many improved models have been proposed. For example, Swin Transformer restricts the calculation of attention within a window through W-MSA to reduce the computational complexity, and then uses SW-MSA to move the window to achieve information exchange between windows. In this way, Swin Transformer successfully reduces the calculation of attention from O(N 2) It has been reduced to O(N) and achieved quite good classification accuracy. Subsequently, it has also been used in semantic segmentation, object detection, etc. This attention mechanism has significantly reduced the computational complexity compared to the self-attention mechanism. However, it still takes a long time to perform image classification and object detection. The present invention can further reduce its computational complexity on this basis. Efficient attention proposed a Factorized Attention Mechanism, which formed a new self-attention mechanism by changing the calculation formula of the attention mechanism, effectively reducing the computational complexity. Subsequently, it was improved and used by the Co-Scale Conv-Attentional Image Transformers (CoaT) model, where Softmax is applied element-wise to the tokens in the sequence, and the projected channel C′ = C, achieving a time complexity of O(NC 2 ) However, through experiments on it, this Factorized Attention Mechanism has problems such as slow convergence speed, poor recognition effect, and unstable recognition speed in some models. Summary of the Invention
[0004] The problem to be solved by the present invention is to provide an image recognition method based on a linear K-value lightweight attention mechanism in view of the deficiencies of the above-mentioned prior art, so as to solve the problem of excessive computational complexity of the current attention mechanism and make the recognition speed more stable.
[0005] To achieve the above object of the present invention, the present invention provides an image recognition method based on a linear K-value lightweight attention mechanism, and the method includes the following steps:
[0006] Step 1: Obtain a data set, store the pictures in the data set into folders according to categories respectively, set the folder names as category names, and then divide the data set into a training set and a test set according to a set ratio; the data set includes several pictures;
[0007] Step 2: Preprocess the training set and the test set, and convert the pictures in the training set and the test set into tensors;
[0008] Step 2.1: Randomly crop the pictures in the training set into pictures of a set size, horizontally flip the randomly cropped pictures, and then convert the pictures into tensors;
[0009] Step 2.2: Convert the pictures in the test set into a set size, then crop the pictures in the center, and convert the cropped pictures into tensors; the size of the cropped pictures is the same as the size of the randomly cropped pictures in the training set.
[0010] Step 3: Standardize the tensors in the training set and the test set to obtain the standardized tensors;
[0011] Step 4: Split the standardized tensors into several patches, and downsample the above-mentioned several patches through convolution. After that, obtain the input vector through a flatten layer; the convolution kernel size and the convolution stride of the convolution are both set to the dimension of the patch;
[0012] Step 5: Design a lightweight attention mechanism based on linear K values;
[0013] The lightweight attention mechanism based on linear K values converts the input vector into a key vector and a value vector through a linear layer. The key vector is divided by the square root of the current key vector dimension and then passes through a linear layer to obtain an attention matrix. Then, position encoding is added to the attention matrix, and the attention matrix with position encoding is multiplied by the value vector to obtain the output vector;
[0014] The formula for obtaining an attention matrix by dividing the key vector by the square root of the current dimension and then passing through a linear layer is:
[0015]
[0016] where Atten is the attention matrix; is the key vector matrix, n is the number of key vectors, d k is the dimension of the key vector; Linear represents a linear transformation;
[0017] The formula for multiplying the attention matrix with position encoding by the value vector to obtain the output vector is:
[0018] Atte(X) = Softmax(Atten * *V) (2)
[0019] where Atte(X) is the output vector; Atten * is the attention matrix with position encoding; is the value vector matrix, d v is the dimension of the value vector; Softmax represents the Softmax function.
[0020] Step 6: Build a Transformers model based on the lightweight attention mechanism with linear K values;
[0021] Step 6.1: Build a multi-head attention mechanism Multi-HeadAttention according to the lightweight attention mechanism with linear K values;
[0022] Step 6.2: Build an encoder Encoder block structure;
[0023] Step 6.2.1: Add the Multi-Head Attention mechanism.
[0024] Step 6.2.2: Add an Add&Norm layer after the Multi-Head Attention mechanism.
[0025] Step 6.2.3: Add a multi-layer perceptron (MLP) after the Add&Norm layer.
[0026] Step 6.3: Concatenate several Encoder block structures to obtain a Transformers model based on the linear K-value lightweight attention mechanism; the output vector of the previous Encoder block structure in the Transformers model based on the linear K-value lightweight attention mechanism is used as the input vector of the next Encoder block structure.
[0027] Step 7: Use the training set to train the Transformers model based on the linear K-value lightweight attention mechanism to obtain a trained Transformers model based on the linear K-value lightweight attention mechanism.
[0028] Step 8: Input the test set into the Transformers model based on the linear K-value lightweight attention mechanism to obtain an output vector.
[0029] Step 9: Input the output vector into a linear layer to obtain the recognition result.
[0030] The technical solution adopted by the present invention has the following technical effects compared with the prior art:
[0031] An image recognition method based on the linear K-value lightweight attention mechanism proposed by the present invention, when performing self-attention mechanism calculation, allows the K value to pass through a linear layer, replacing the process of multiplying the Q value and the K value in the existing attention mechanism. On the basis of ensuring similar effects to the existing attention mechanism, it saves nearly half of the computational amount and makes the recognition speed more stable. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a flowchart of an image recognition method based on the linear K-value lightweight attention mechanism in an embodiment of the present invention;
[0033] Figure 2 It is a schematic diagram of the linear K-value lightweight attention mechanism in an embodiment of the present invention;
[0034] Figure 3This is the structural diagram of the Transformers model that establishes a linear K-value lightweight attention mechanism in the embodiments of the present invention. Specific embodiments
[0035] In order to make the technical solutions and their advantages of the present invention clearer, the technical solutions in this application will be described below in conjunction with the accompanying drawings. The specific real-time examples described here are only used to explain the present invention and are not used to limit the scope of the present invention.
[0036] An image recognition method based on a linear K-value lightweight attention mechanism in this embodiment is as Figure 1 shown, and includes the following steps:
[0037] Step 1: Obtain a data set, store the pictures in the data set into folders according to categories respectively, set the folder name as the category name, and then divide the data set into a training set and a test set according to a set ratio; the data set includes several pictures;
[0038] In this embodiment, the ImageNet data set is used; the training set and the test set are divided according to a ratio of 8:2;
[0039] Step 2: Preprocess the training set and the test set, and convert the pictures in the training set and the test set into tensors;
[0040] Step 2.1: Randomly crop the pictures in the training set into pictures of a set size, horizontally flip the randomly cropped pictures, and then convert the pictures into tensors;
[0041] In this embodiment, the pictures in the training set are randomly cropped into pictures of 224*224; the batch size is set to 8;
[0042] Step 2.2: Convert the pictures in the test set into a set size, then crop the pictures in the center, and convert the cropped pictures into tensors; the size of the cropped pictures is the same as the size of the randomly cropped pictures in the training set;
[0043] In this embodiment, the pictures in the test set are first converted into pictures of 256*256, and then cropped in the center of the pictures to be cropped into pictures of 224*224; the batch size is set to 8;
[0044] Step 3: Standardize the tensors in the training set and the test set to obtain standardized tensors;
[0045] Step 4: Split the standardized tensor into several patches, and downsample the above-mentioned several patches through convolution. After that, obtain the input vector through a flatten layer; the convolution kernel size and convolution stride of the convolution are both set to the dimension of the patch;
[0046] Step 5: Design a lightweight attention mechanism based on linear K values;
[0047] As Figure 2 shown, the lightweight attention mechanism based on linear K values converts the input vector into a key vector and a value vector through a linear layer. The key vector is divided by the square root of the current key vector dimension and then passes through a linear layer to obtain an attention matrix. Then, positional encoding is added to the attention matrix, and the attention matrix with positional encoding is multiplied by the value vector to obtain the output vector;
[0048] The formula for obtaining an attention matrix by dividing the key vector by the square root of the current dimension and then passing through a linear layer is:
[0049]
[0050] where Atten is the attention matrix; is the key vector matrix, n is the number of key vectors, d k is the dimension of the key vector; Linear represents a linear transformation; dividing by the square root of the current dimension is used to prevent extreme values. Extreme values will become more extreme after passing through Softmax, which will lead to gradient disappearance.
[0051] The formula for multiplying the attention matrix with positional encoding by the value vector to obtain the output vector is:
[0052] Atte(X) = Softmax(Atten * *V) (2)
[0053] where Atte(X) is the output vector; Atten * is the attention matrix with positional encoding; is the value vector matrix, d v is the dimension of the value vector. Usually, d K = d v ; Softmax represents the Softmax function, which is used for normalization and ensures the non-negativity of the values inside the matrix.
[0054] Step 6: Build a Transformers model based on the lightweight attention mechanism with linear K values;
[0055] Step 6.1: Establish a multi-head attention mechanism Multi-HeadAttention based on the linear K-value lightweight attention mechanism;
[0056] The multi-head attention mechanism Multi-Head Attention is composed of several linear K-value lightweight attention mechanisms;
[0057] Step 6.2: Establish an encoder Encoder block structure;
[0058] Step 6.2.1: Add the multi-head attention mechanism Multi-Head Attention;
[0059] Step 6.2.2: Add an Add&Norm layer after the multi-head attention mechanism Multi-Head Attention;
[0060] Step 6.2.3: Add a multi-layer perceptron MLP after the Add&Norm layer;
[0061] Step 6.2.4: Add an Add&Norm layer after the multi-layer perceptron MLP to obtain the encoder Encoder block structure;
[0062] Step 6.3: Connect several Encoder block structures in series to obtain a Transformers model based on the linear K-value lightweight attention mechanism;
[0063] As Figure 3 shown, in the Transformers model based on the linear K-value lightweight attention mechanism, the output vector of the previous Encoder block structure is used as the input vector of the next Encoder block structure;
[0064] Step 7: Train the Transformers model based on the linear K-value lightweight attention mechanism using the training set to obtain a trained Transformers model based on the linear K-value lightweight attention mechanism;
[0065] Step 8: Input the test set into the Transformers model based on the linear K-value lightweight attention mechanism to obtain an output vector;
[0066] Step 9: Input the output vector into a linear layer to obtain the recognition result.
Claims
1. An image recognition method based on a lightweight attention mechanism with linear K value, characterized in that, It includes the following steps: Step 1: Obtain a dataset, store the pictures in the dataset into folders according to categories respectively, set the folder names as category names, and then divide the dataset into a training set and a test set according to a set ratio; the dataset includes several pictures; Step 2: Preprocess the training set and the test set, and convert the pictures in the training set and the test set into tensors; Step 3: Normalize the tensors in the training set and the test set to obtain normalized tensors; Step 4: Cut the normalized tensors into several patches, perform downsampling on the above-mentioned several patches through convolution, and then obtain an input vector through a faltten layer; the convolution kernel size and the convolution stride of the convolution are both set to the dimension of the patch; Step 5: Design a lightweight attention mechanism based on linear K values; The lightweight attention mechanism based on linear K values converts the input vector into a key vector and a value vector through a linear layer. After the key vector is divided by the square root of the current key vector dimension, a linear layer is used to obtain an attention matrix. Then, position encoding is added to the attention matrix, and the attention matrix with position encoding is multiplied by the value vector to obtain an output vector; The formula for obtaining an attention matrix by dividing the key vector by the square root of the current dimension and then passing through a linear layer is: Among them, Atten is the attention matrix; is the key vector matrix, n is the number of key vectors, and d k is the dimension of the key vector; Linear represents a linear transformation; The formula for multiplying the attention matrix with position encoding by the value vector to obtain an output vector is: Atte(X) = Softmax(Atten * *V) (2) Among them, Atte(X) is the output vector; Atten * is the attention matrix with positional encoding added; is the value vector matrix, d v is the dimension of the value vector; Softmax represents the Softmax function; Step 6: Build a Transformers model based on the lightweight attention mechanism with linear K values; Step 7: Use the training set to train the Transformers model based on the lightweight attention mechanism with linear K values to obtain a trained Transformers model based on the lightweight attention mechanism with linear K values; Step 8: Input the test set into the Transformers model based on the lightweight attention mechanism with linear K values to obtain an output vector; Step 9: Input the output vector into a linear layer to obtain a recognition result.
2. The image recognition method based on a lightweight attention mechanism with linear K value according to claim 1, characterized in that Step 2 includes the following specific steps: Step 2.1: Randomly crop the pictures in the training set into pictures of a set size, horizontally flip the randomly cropped pictures, and then convert the pictures into tensors; Step 2.2: Convert the pictures in the test set into a set size, then crop the pictures in the center, and convert the cropped pictures into tensors; the size of the cropped pictures is the same as the size of the randomly cropped pictures in the training set.
3. A method for image recognition based on a linear K-value lightweight attention mechanism according to claim 1, characterized in that, Step 6 includes the following specific steps: Step 6.1: Build a multi-head attention mechanism Multi-HeadAttention according to the lightweight attention mechanism based on linear K values; Step 6.2: Build an encoder Encoder block structure; Step 6.3: Connect several Encoder block structures in series to obtain a Transformers model based on a lightweight attention mechanism with linear K value; in the Transformers model based on the lightweight attention mechanism with linear K value, the output vector of the previous Encoder block structure is used as the input vector of the next Encoder block structure.
4. The image recognition method based on the linear K - value lightweight attention mechanism according to claim 3, wherein, The specific steps of Step 6.2 are as follows: Step 6.2.1: Add the multi-head attention mechanism Multi-HeadAttention; Step 6.2.2: Add an Add&Norm layer after the multi-head attention mechanism Multi-HeadAttention; Step 6.2.3: Add a multi-layer perceptron MLP after the Add&Norm layer.