An image classification method based on Transformer neural network
Through the image classification method based on Transformer neural network, combined with pre-training, fine-tuning, image chunking and sliding window multi-head attention mechanism, the problem of traditional convolutional neural networks being limited in receptive field and high computational volume during image classification is solved, achieving higher accuracy and real-timeness.
Patent Information
- Application Number
- CN202210464982.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-04-29
AI Technical Summary
Traditional convolutional neural networks have limited receptive fields when classifying images, and cannot effectively capture long-distance features. The multi-head attention mechanism is computationally large, affecting real-time.
The image classification method based on Transformer neural network is adopted to extract common and deep features through pre-training and fine-tuning, combining image chunking, linear transformation and self-attention mechanism, and alternate calculations are further used to reduce the calculation amount.
It improves the accuracy of image classification and the real-time nature of the model, and solves the problems of restricted receptive field and large computing volume of traditional convolutional neural networks.
Smart Images

Figure CN114743022B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image classification, and particularly to an image classification method based on a Transformer neural network. Background Art
[0002] In traditional convolutional neural network algorithms, since the actual receptive field is much smaller than the theoretical receptive field, the network cannot obtain the global view of the entire image during detection. Therefore, there is a certain degree of deviation in identifying the category of the image. When the Transformer neural network detects an image, it can obtain the global view of the image and has a stronger ability to capture long-distance features. Therefore, it can better classify the image. Secondly, image classification in some scenarios requires high real-time performance. In the previous attention mechanisms, global self-attention calculation needs to be performed on the feature map, resulting in a huge amount of computation.
[0003] Liu Bing et al. used a Transformer neural network that performs self-attention calculation globally in "A Hyperspectral Image Classification Method Based on Deep Transformer". Compared with traditional convolutional neural networks, it improved the feature expression ability of the model, but at the same time brought a huge amount of computation and did not well improve the prediction speed of the model. Summary of the Invention
[0004] Aiming at the deficiencies of existing algorithms, the present invention solves the problems that the accuracy of traditional convolutional neural networks in classifying images is not high enough, and the existing multi-head attention mechanism has a large amount of computation.
[0005] The technical solution adopted by the present invention is: an image classification method based on a Transformer neural network includes the following steps:
[0006] S1. Pre-training: Use a large dataset to pre-train the Transformer neural network, extract common features from it, reduce the burden of the model's learning for specific tasks, and enable the model to obtain the ability to express shallow and extensive features.
[0007] S2. Fine-tuning: Use a self-made dataset to fine-tune the Transformer neural network, so that the parameters of the pre-trained model can adapt to classifying specific tasks, and thus enable the model to obtain the ability to express deep and specific features;
[0008] S3. Image Blocking: Process the collected image in blocks;
[0009] Further, a RGB three-channel image with a width and height of H×W is cut into small blocks with a width and height of 4×4 each, and the number of blocks obtained is Then, each small block is flattened and concatenated along the channel dimension. Each small block becomes a tensor of shape 1×1×48, and the entire image becomes a tensor of shape after the chunking operation;
[0010] S4. Linear transformation: Perform a linear transformation on each pixel of the tensor along the channel dimension, and then perform a LayerNorm operation on the transformed tensor along the channel dimension;
[0011] Furthermore, first, perform a linear transformation on each pixel along the channel dimension to adapt to the input of the Transformer neural network;
[0012] Second, perform a LayerNorm operation on the transformed tensor along the channel dimension, as shown in Equation 1:
[0013]
[0014] where E(x) and Var(x) are the mean and variance of the input elements of each layer, respectively; ∈ is a specified extremely small value, and γ and β are the weights and biases of each pixel;
[0015] S5. Feature extraction: Feed the processed image into the trained Transformer neural network for feature extraction;
[0016] Furthermore, first, perform multiple feature extractions on the image through the Transformer neural network, and use the obtained feature maps containing rich target information as detection sources;
[0017] After multiple downsamplings and multiple attention calculations on the three-dimensional data, it becomes data with a dimension of ;
[0018] The Transformer neural network captures features by calculating the correlation of each pixel based on the self-attention mechanism, obtaining a global view of the entire image. The implementation principle of the self-attention mechanism is shown in Equation 2:
[0019]
[0020] where Q (Query), K (Key), and V (Value) are three learnable matrices from the same input, d k is the dimension of a Query and Key vector, and softmax(·) is a normalization function;
[0021] Furthermore, the sliding window multi-head attention mechanism is used, and when dividing the feature map, a certain degree of offset is performed. During the process of feature extraction from the image, window attention and sliding window multi-head attention are alternately used;
[0022] Furthermore, alternately using window attention and sliding window multi-head attention includes:
[0023] The output x obtained from the previous layer l-1 , undergoes LayerNorm operation; then, based on the window multi-head attention mechanism, the correlation between each pixel point is calculated to obtain x l , and a residual connection is added between x l-1 and x l , and then passes through the next LN normalization layer; then the output is sent to a multi-layer perceptron; the multi-layer perceptron consists of a fully connected layer, a GELU activation function, and a Dropout layer; finally, the output x l is sent to the next module containing the sliding window multi-head attention mechanism;
[0024] S6, Detection: Detect the final features extracted in step S5, and classify the image into the class with the highest score;
[0025] Advantages of the present invention:
[0026] 1. The neural network constructed based on the Transformer architecture solves the problems of limited receptive field and inability to capture long-distance features well in traditional convolutional neural networks;
[0027] 2. The original multi-head attention mechanism calculates self-attention based on the global scale, which requires consuming a large amount of computing resources. The present invention changes the global self-attention calculation to window-based self-attention calculation. By alternately using the window multi-head attention and sliding window multi-head attention mechanisms, the computational amount is reduced, thereby improving the training and prediction speed of the neural network. Description of the Drawings
[0028] Figure 1 is the flowchart of the image classification method based on the Transformer neural network of the present invention;
[0029] Figure 2 is the structure diagram of the Transformer neural network of the present invention;
[0030] Figure 3 is the attention mechanism module of the present invention;
[0031] Figure 4 is the structure diagram of the detection module of the present invention;
[0032] Figure 5 It is a schematic diagram when the improved multi-head attention mechanism of the present invention performs a segmentation operation. Detailed implementation manners
[0033] The present invention will be further described below in conjunction with the drawings and embodiments. This figure is a simplified schematic diagram, which only illustrates the basic structure of the present invention in a schematic manner. Therefore, it only shows the components related to the present invention.
[0034] As Figure 1 shown, an image classification method based on a Transformer neural network includes the following steps:
[0035] S1. Pre-training: Use a large dataset to pre-train the Transformer neural network, extract as many common features as possible from it, thereby reducing the burden on the model for learning specific tasks, and enabling the model to obtain the ability to express shallow and extensive features;
[0036] In this embodiment, the ImageNet-22K dataset contains 22,000 different categories and a total of about 15 million images.
[0037] S2. Fine-tuning: For a specific classification task, use a self-made dataset to fine-tune the Transformer neural network, so that the parameters of the pre-trained model can be adapted to classify specific tasks, thereby enabling the model to obtain the ability to express deep and specific features;
[0038] In this embodiment, the CIFAR-100 dataset contains 100 categories, and each category has 600 32×32 color images;
[0039] S3. Image chunking: Perform chunking on the collected images;
[0040] A RGB three-channel image with a width and height of H×W is sliced into small blocks, each block having a width and height of 4×4. Then the number of blocks obtained is Then each small block is flattened and concatenated in the channel dimension. Each small block becomes a tensor with a shape of 1×1×48. After the chunking operation on the entire image, it becomes a tensor with a shape of In this embodiment, the size of the image is 384×384×3.
[0041] S4. Linear transformation: Perform a linear transformation on each pixel of the tensor obtained after the process of step S3 in the channel dimension, and perform a LayerNorm operation on the transformed tensor in the channel dimension;
[0042] First, perform a linear transformation on each pixel in the channel dimension to adapt to the input of the Transformer neural network;
[0043] Secondly, perform LayerNorm operation on the transformed tensor in the channel dimension, as shown in Equation 1:
[0044]
[0045] where E(x) and Var(x) are the mean and variance of the input elements of each layer respectively; ∈ is a specified extremely small value, whose function is to avoid the situation of denominator being zero; γ and β are the weights and biases of each pixel point, used for affine transformation, and the value of ∈ in this embodiment is 0.00001.
[0046] S5. Feature extraction: Send the image processed in step S4 into the trained Transformer neural network for feature extraction;
[0047] The Transformer neural network will perform feature extraction on the picture multiple times, and use the finally obtained feature map containing rich target information as the detection source, as Figure 2 is the overall structure of the Transformer neural network. After each downsampling, the resolution of the input image data is halved and the number of channels is doubled; After the three-dimensional data goes through multiple attention calculations and downsamplings, it becomes a dimension of data.
[0048] The Transformer neural network captures features by calculating the correlation of each pixel based on the self-attention mechanism, so as to obtain the global view of the entire picture; the implementation principle of the self-attention mechanism is shown in Equation 2:
[0049]
[0050] where Q (Query), K (Key), and V (Value) are three learnable matrices from the same input, d k is the dimension of a Query and Key vector, and softmax(·) is the normalization function.
[0051] In the present invention, for the picture input data, an extended window multi-head attention mechanism based on the self-attention mechanism is used. The feature map extracted in each stage is divided into multiple windows, and the self-attention mechanism is calculated within each window. Compared with the original calculation amount of the global-based self-attention mechanism, it is reduced by about three-fifths; but it also brings the problem that the information between each window cannot interact, resulting in a reduction in the receptive field of the network.
[0052] Therefore, the present invention uses a sliding window multi-head attention mechanism. When performing the partitioning operation on the feature map, a certain degree of offset is carried out, expanding the receptive field range. During the process of feature extraction from the picture, these two attention mechanisms are alternately used, which not only reduces the computational amount but also improves the prediction accuracy of the network.
[0053] Such as Figure 3 For Figure 2 the specific structure of the attention mechanism module, the output x obtained from the previous layer l-1 , first undergoes a LayerNorm operation to ensure the stability of the data feature distribution; then, based on the window multi-head attention mechanism, the correlation between each pixel point is calculated to obtain x l , and a residual connection is added between x l-1 and x l ; subsequently, it passes through the next LN normalization layer; then the output is sent to a multi-layer perceptron, which is composed of a fully connected layer, a GELU activation function, and a Dropout layer; finally, the output x l is sent to a similar module containing a sliding window multi-head attention mechanism in the next layer;
[0054] S6, Detection: Detect the final features extracted in step S5, and classify the picture into the category with the highest score;
[0055] Such as Figure 4 For the structural diagram of the detection module, first pass through a LayerNorm layer to accelerate the convergence speed of the model; then pass through an average pooling layer to calculate the average value of the image area as the value after pooling of this area; then pass through a fully connected layer to map the learned "feature representation" to the sample label space, that is, classify the target image, and finally output the classification result through the output layer.
[0056] Compare the Transformer neural network proposed in the present invention with the traditional convolutional neural network to verify the effectiveness of the neural network constructed based on the Transformer architecture in the image classification task;
[0057] This embodiment uses the CIFAR-100 dataset, which contains 100 categories, each category has 600 color images with a resolution of 32×32, and is divided into 50,000 training set images and 10,000 test set images. The learning rate is set to 0.001, the batch size is set to 64, the number of Epochs is set to 30, and each epoch iterates through the training set once; the comparison results of the training accuracy are shown in Table 1:
[0058] Table 1 Comparison of accuracies in the image classification task
[0059]
[0060] As can be seen from Table 1, the accuracy of the Transformer neural network structure constructed by the present invention is superior to that of the other two convolutional neural networks, which shows the superiority and effectiveness of the algorithm of the present invention and has greater application value.
[0061] As Figure 5 shown, it is a schematic diagram of the improved multi-head attention mechanism of the present invention when performing the splitting operation; among them, the square diagram composed of black lines is the feature map, and the orange lines are the splitting lines. In the window multi-head attention, the splitting will be performed according to the layout of the orange lines in the left figure, and a feature map will be divided into nine small windows of equal size. Then, local self-attention calculations will be performed within each window. Although the amount of calculation is reduced, it also brings the problem that information cannot be exchanged between each window, resulting in a reduction in the receptive field of the network.
[0062] Therefore, after each module using window multi-head attention, a module using sliding window multi-head attention follows. As Figure 5 shown in the right figure of, when the sliding window multi-head attention divides the feature map, a certain degree of offset is performed, and the window sizes after division are different. Then, self-attention calculations are performed within the nine windows respectively;
[0063] Table 2 Comparison of Three Kinds of Multi-Head Attention Mechanisms
[0064]
[0065] It can be clearly seen from Table 2 that the amounts of calculation of both window multi-head attention and sliding window multi-head attention are much smaller than that of multi-head attention. Although neither window multi-head attention nor sliding window multi-head attention can obtain the global view of the image when used alone, in the present invention, these two kinds of attention are alternately used in the network structure, so that the network can obtain the global view of the entire image, while reducing the amount of calculation without reducing the learning and expression capabilities of the network.
[0066] Taking the above ideal embodiments according to the present invention as an inspiration, through the above description, relevant staff can completely make various changes and modifications within the scope not deviating from the technical idea of this invention. The technical scope of this invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.
Claims
1. An image classification method based on the Transformer neural network, characterized in that, it includes the following steps: S1. Pretraining: Use a large dataset to pre-train the Transformer neural network and extract common features from it; S2. Fine-tuning: Use a self-made dataset to fine-tune the Transformer neural network so that the parameters of the pre-trained model can be adapted to classify specific tasks; S3. Image chunking: Process the collected images in chunks; S4. Linear transformation: Perform a linear transformation on each pixel of the tensor in the channel dimension and perform a LayerNorm operation on the transformed tensor in the channel dimension; The step S4 includes: First, perform a linear transformation on each pixel in the channel dimension Second, perform a LayerNorm operation on the transformed tensor in the channel dimension, as shown in Equation 1: where E(x) and Var(x) are the mean and variance of the input elements of each layer respectively; ∈ is a specified extremely small value; γ and β are the weights and biases of each pixel; S5. Feature extraction: Feed the processed image into the trained Transformer neural network for feature extraction; The step S5 includes: First, perform multiple feature extractions on the picture through the Transformer neural network and use the obtained feature map containing target information as the detection source; Secondly, the three-dimensional data undergoes multiple downsamplings and multiple attention calculations and becomes data with a dimension of . Finally, calculate the correlation of each pixel based on the self-attention mechanism to obtain the global view of the entire picture; The self-attention mechanism is to use a sliding window multi-head attention mechanism and perform a certain degree of offset when dividing the feature map; during the process of feature extraction of the picture, window multi-head attention and sliding window multi-head attention are alternately used; The alternate use of window multi-head attention and sliding window multi-head attention includes: First, the output x obtained from the previous layer l-1 , undergoes the LayerNorm operation; Then, calculate the correlation between each pixel based on the window multi-head attention mechanism to obtain x l , and add a residual connection between x l-1 and x l . Subsequently, pass through the next LN normalization layer, and then send the output to a multi-layer perceptron; Finally, output x l is fed into the next module with a sliding window multi-head attention mechanism; S6. Detection: Detect the final features extracted in step S5 and classify the picture into the class with the highest score.
2. The image classification method based on the Transformer neural network according to claim 1, characterized in that: The multi-layer perceptron is composed of a fully connected layer, a GELU activation function, and a Dropout layer.
Citation Information
Patent Citations
Transform-based key feature enhanced gastric cancer image recognition method
CN114119585A