Transform neural network-based image classification method
Through the image classification method of dynamic adaptive chunking and cross-scale feature interaction, the problem of imbalance between efficiency and accuracy in image classification is solved, the model's sensitivity to edges and textures is enhanced, and the efficient image classification effect is achieved.
Patent Information
- Application Number
- CN202510493850.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-19
- Publication Date
- 2025-07-25
AI Technical Summary
Existing image classification methods cannot dynamically block according to the local complexity of different image areas, and lack explicit interactions with features of different scales, resulting in an imbalance in efficiency and accuracy.
By using the Sobel operator to calculate the gradient amplitude component of the image, dynamically adaptively block, and combining the pre-trained ViT model to expand the linear projection layer weight, the multi-head attention mechanism or cross-scale cross-attention mechanism is used for feature processing to generate highly discriminant image-level representations.
The model's sensitivity to edges and textures is enhanced, the calculation quantity and accuracy is dynamically balanced, and the pre-trained model is effectively utilized, which can achieve efficient fusion of global semantics and local details, and improve the performance of image classification.
Smart Images

Figure CN120375080A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image classification, relates to Transformer neural network technology, and specifically is an image classification method based on Transformer neural network. Background Art
[0002] Image classification, as one of the core tasks of computer vision, has wide applications in fields such as intelligent security, medical image diagnosis, and remote sensing data analysis. With the development of deep learning, vision models based on Transformer (such as ViT) have shown excellent performance in image classification tasks due to their powerful global modeling ability.
[0003] Traditional ViT divides an image into equal-sized patches through fixed partitioning, and after linear projection, inputs them into the Transformer encoder. However, its fixed partitioning strategy is difficult to balance the local complexity of different image regions. Using fine-grained partitioning in texture-smooth regions will increase computational redundancy, while using coarse-grained partitioning in regions with rich details may lose key local features, resulting in an imbalance between efficiency and accuracy.
[0004] In addition, existing methods mostly rely on pixel values as input features, ignoring structural information such as image edges and textures, as well as local structural changes that characterize images. In the application of pre-trained models, the linear projection layer of ViT is usually designed for three-channel RGB images. When introducing additional features (such as gradient magnitude), directly expanding the channels may lead to training instability caused by random initialization of parameters, or overfitting due to fine-tuning with small samples. There is an urgent need for a parameter expansion strategy that can utilize the advantages of pre-trained weights and adapt to new features.
[0005] On the other hand, traditional fixed-partitioning models lack explicit interaction of features at different scales and are difficult to balance global semantics and local details.
[0006] In view of the above technical problems, the present application provides a solution. Summary of the Invention
[0007] The purpose of the present invention is to provide an image classification method based on Transformer neural network to solve the problems that existing image classification methods cannot perform dynamic partitioning according to the local complexity of different image regions and lack explicit interaction of features at different scales.
[0008] The technical problem that the present invention needs to solve is: how to provide an image classification method based on Transformer neural network that can perform dynamic partitioning according to the local complexity of different image regions and has explicit interaction of features at different scales.
[0009] The object of the present invention can be achieved by the following technical solutions:
[0010] An image classification method based on the Transformer neural network, comprising the following steps:
[0011] Obtain the input image, and use the Sobel operator to calculate the gradient magnitude component of each pixel in the input image respectively;
[0012] Calculate the gradient variation value of the input image according to the gradient magnitude component, and perform dynamic adaptive partitioning on the input image according to the gradient variation value;
[0013] Flatten the partition into a vector sequence, expand the weights of the linear projection layer according to the pre-trained ViT model, and linearly project the vector sequence into d-dimensional embedding vectors through the linear projection layer;
[0014] Add positional encoding to the embedding vectors to generate feature vectors;
[0015] According to the type of the dynamic adaptive partition, use the multi-head attention mechanism or the cross-scale cross-attention mechanism of the Transformer to process the feature vectors, and output the class probability distribution of the input image.
[0016] Further, the specific process of using the Sobel operator to calculate the gradient magnitude component of each pixel includes:
[0017] Each pixel takes its 3*3 neighborhood to form a pixel region;
[0018] Align the three RGB components of the pixel region with the centers of two groups of preset convolution kernels respectively, and perform element-wise multiplication and then summation to obtain the horizontal gradient component and the vertical gradient component;
[0019] Perform a square root operation on the sum of the squares of the horizontal gradient component and the vertical gradient component to obtain the gradient magnitude component of each pixel on the three RGB channels. The calculation formula is:
[0020] where G x,R 、G x,G 、G x,B are the horizontal gradient components, and G y,R 、G y,G 、G y,B are the vertical gradient components;
[0021] Normalize the gradient magnitude component.
[0022] Further, the process of performing dynamic adaptive partitioning on the input image according to the gradient variation value includes:
[0023] Calculate the gradient magnitude TF of each pixel in the input image;
[0024] Calculate the gradient variation value of the input image according to the gradient magnitude TF of the pixel, and compare the gradient variation value with a preset gradient variation threshold:
[0025] If the gradient variation value is greater than or equal to the gradient variation threshold, use hybrid block division;
[0026] If the gradient variation value is less than the gradient variation threshold, use average block division.
[0027] Further, the specific process of the average block division includes:
[0028] Calculate the gradient mean value of the input image according to the gradient magnitude TF of the pixel, and compare the gradient mean value with a preset gradient average threshold: if the gradient mean value is greater than or equal to the gradient average threshold, use the average block division of small blocks; if the gradient mean value is less than the gradient average threshold, use the average block division of large blocks
[0029] The specific process of hybrid block division includes: evenly divide the input image into several large blocks, calculate the gradient mean value of each large block and compare it with the gradient average threshold: if the gradient mean value is less than the gradient average threshold, no processing is required; if the gradient mean value is greater than or equal to the gradient average threshold, divide the large block into small blocks evenly.
[0030] Further, the specific process of the extended linear projection layer weights includes: load the pre-trained ViT model, extend the input channels of the pre-trained ViT model to six channels, use the projection weights of the pre-trained ViT model for the first three channels, and initialize the last three channels through staged training.
[0031] Further, the specific process of the staged training includes: in the first stage, freeze the projection weights corresponding to the RGB three channels of the pre-trained model, and only train the relevant parameters of the gradient magnitude component channel; in the second stage, copy the projection weights corresponding to the RGB three channels to the gradient channels and adjust them with a low learning rate.
[0032] Further, the calculation formula for mapping the vector sequence of each block to a d-dimensional embedding vector is:
[0033] F small = Linear(F small , d), F large = Linear(F large , d), where Fsmall is the small block feature corresponding to the vector sequence of the small block, and Flarge is the large block feature corresponding to the vector sequence of the large block.
[0034] Furthermore, the specific process of adding positional encoding to the embedding vectors to generate input feature vectors includes:
[0035] For each block center point (xi, yi), calculate its normalized coordinates;
[0036] Add 2D positional encoding to the embedding vector corresponding to each block. The calculation formula is:
[0037] Epatch = Linear(Flatten(Xpatch)) + PE2D(xi, yi);
[0038] The calculation formula of 2D positional encoding is:
[0039] PE2D(xi, yi) = Concat(PE(xi), PE(yi)) ∈ Rd, where PE(·) is the sine function encoding of ViT.
[0040] Furthermore, the specific process of processing the feature vectors using the multi-head attention mechanism includes:
[0041] Splice the randomly initialized CLS token to the head of the feature vector to form an input sequence;
[0042] The CLS token obtains the output vector of the CLS token by interacting with the input sequence, and inputs the output vector of the CLS token into the fully connected layer to output the class probability distribution.
[0043] Furthermore, the specific process of processing the feature vectors using the cross-scale cross-attention mechanism includes:
[0044] Obtain the feature vector Flarge of the large block and the feature vector Fsmall of the small block;
[0045] Use the feature vector Flarge as the query, and map the feature vector Fsmall as the key-value pair to the query Q, key K, and value V spaces respectively;
[0046] Calculate the correlation between the feature vector Flarge and the feature vector Fsmall. The calculation formula is:
[0047] The Softmax function normalizes the weights into a probability distribution;
[0048] Generate the fused feature vector Ffused by weighting the feature vector Fsmall of the small block. The specific calculation formula is:
[0049]
[0050] Dynamically adjust the fused feature vector Ffused and the feature vector Flarge of the original large block, and generate a weight matrix G:
[0051] Where W g ∈R 2d×d is a learnable parameter, and σ is the Sigmoid function;
[0052] Dynamically mix the feature vector Ffused and the feature vector Flarge of the original large block in combination with the weight matrix G. The specific calculation formula is:
[0053] ⊙ represents element-wise multiplication;
[0054] For the mixed features Perform global average pooling to obtain a global feature vector f ∈ R d ; Input the global feature vector f into a multi-layer perceptron classifier to output class probabilities. The specific calculation formula is:
[0055] Logits = MLP(f) ∈ R C , where C is the number of classes.
[0056] The present invention has the following beneficial effects:
[0057] 1. Extract the gradient information of the RGB components in the input image through the Sobel operator, enhance the sensitivity of the model to edges and textures, and make up for the deficiency of pure pixel features;
[0058] 2. Dynamically select the block strategy and block size according to the gradient variation value and mean. For images with different local complexities, large blocks are used in smooth areas to reduce the calculation amount, and small blocks are used in areas with rich details to capture local features, greatly balancing efficiency and accuracy;
[0059] 3. Expand the channels on the basis of the pre-trained ViT, design a six-channel linear projection layer (RGB + gradient magnitude), and adopt a staged training strategy of freezing the original weights to fine-tune at a low learning rate when designing the weights of the three channels of the gradient magnitude, effectively utilizing the pre-trained ViT basis and avoiding overfitting under small data;
[0060] 4. Generate position encoding based on the block center coordinates, retain the spatial position information while being compatible with different block sizes, and enhance the model's understanding of the image structure;
[0061] 5. Complete multi-scale feature fusion through cross-scale attention guidance, feature splicing, and dynamic selection of gated fusion features and original features, realize the efficient fusion of global semantics and local details, and finally generate a highly discriminative image-level representation, significantly improving the performance of the fine-grained classification task. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0063] Figure 1 It is the method logic block diagram of the present invention;
[0064] Figure 2 It is the overall method flowchart of the present invention;
[0065] Figure 3 It is the method flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0066] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0067] As Figures 1-3 shown, an image classification method based on the Transformer neural network includes the following steps:
[0068] S1: Obtain the input image, and use the Sobel operator to calculate the gradient magnitude component of each pixel in the input image respectively;
[0069] Specifically, in some embodiments, each pixel takes its 3*3 neighborhood to form a 3-row and 3-column pixel region centered on the target pixel, and the gradient magnitude component of each pixel is independently calculated on three RGB channels; the calculation process of the gradient magnitude component includes: using the Sobel operator to calculate the three RGB channels of the 3*3 pixel region respectively, aligning the three RGB channels of the pixel region with the centers of two groups of convolution kernels respectively, and performing element-wise multiplication and then summing to obtain the horizontal gradient component G x,R and the vertical gradient component G y,R and the horizontal gradient component G x,G and the vertical gradient component G y,G and the horizontal gradient component G x,B and the vertical gradient component G y,B ; taking the square root of the sum of the squares of the horizontal gradient components and the vertical gradient components on the three RGB channels to obtain the gradient magnitude component of each pixel on the three RGB channels, and the specific formula is:
[0070]
[0071] Further, in some embodiments, RGB is a commonly used color space. Each pixel of the input image consists of three channels, namely R, G, and B; R represents the red channel, G represents the green channel, and B represents the blue channel. Described in 24-bit color, R occupies 8 bits with values from 0 to 255, G occupies 8 bits with values from 0 to 255, and B occupies 8 bits with values from 0 to 255;
[0072] Preferably, in some embodiments, the Sobel operator refers to two groups of 3×3 convolution kernels, specifically, the horizontal direction convolution kernel [-1, 0, 1; -2, 0, 2; -1, 0, 1] and the vertical direction convolution kernel [-1, -2, -1; 0, 0, 0; 1, 2, 1] can be adopted;
[0073] Further, in some embodiments, when the pixel region is located at the image boundary, that is, there are blanks in the 3*3 neighborhood of the target pixel, the blank pixel region is filled with 0 to keep the output size consistent;
[0074] Specifically, in some embodiments, the gradient magnitude components of the three RGB channels are respectively normalized to the range of [0, 1]; the normalization calculation formula is:
[0075] Xnorm = (X - Xmin) / (Xmax - Xmin), where Xmax is the maximum value of the gradient magnitude component, Xmin is the minimum value of the gradient magnitude component, and Xnorm is the value of the normalized gradient magnitude component;
[0076] Extracting the gradient information of the RGB components in the input image through the Sobel operator enhances the sensitivity of the model to edges and textures and makes up for the deficiency of pure pixel features;
[0077] S2: Calculate the gradient variation value of the input image according to the gradient magnitude component, and perform dynamic adaptive block division on the input image according to the gradient variation value;
[0078] Specifically, in some embodiments, the block division types of dynamic adaptive block division include average block division and hybrid block division;
[0079] Further, in some embodiments, the selection process of the block division method includes: calculating the sum of the gradient magnitude components of each pixel on the three RGB channels to obtain the gradient magnitude TF of each pixel in the input image;
[0080] Calculating the variance of the gradient magnitude TF of the input image to obtain the gradient variation value, and comparing the gradient variation value with a preset gradient variation threshold:
[0081] If the gradient variation value is greater than or equal to the gradient variation threshold, hybrid block division is adopted;
[0082] If the gradient variation value is less than the gradient variation threshold, average block division is adopted; further, the average value of the gradient magnitude TF is calculated to obtain the gradient mean value, and the gradient mean value is compared with the preset gradient average threshold: if the gradient mean value is greater than or equal to the gradient average threshold, small blocks are used for average block division, and if the gradient mean value is less than the gradient average threshold, large blocks are used for average block division;
[0083] Further, in some embodiments, average block division evenly divides the input image into N blocks of size P*P (P takes the value of P1 or P2); hybrid block division divides the input image into N l large blocks of size P1*P1 and N s small blocks of size P2*P2, and N = N l + N s ;
[0084] Further, in some embodiments, the specific process of hybrid block division includes: first, the input image is evenly divided into several large blocks of size P1*P1, and the gradient mean value of each large block is calculated. If the gradient mean value is less than the gradient average threshold, the large block is retained. If the gradient mean value is greater than or equal to the gradient average threshold, the large block is evenly divided into small blocks of size P2*P2;
[0085] Preferably, in some embodiments, the small block is an 8×8 block and the large block is a 16×16 block;
[0086] The block division strategy and block size are dynamically selected according to the gradient variation value and mean value. For images with different local complexities, large blocks are used in smooth regions to reduce the computational amount, and small blocks are used in regions with rich details to capture local features, greatly balancing efficiency and accuracy;
[0087] S3: Flatten the divided blocks into a vector sequence, expand the weights of the linear projection layer according to the pre-trained ViT model, and linearly project the vector sequence into d-dimensional embedding vectors through the linear projection layer;
[0088] Specifically, in some embodiments, the divided blocks are flattened into a vector sequence Xpatch in the order from left to right and from top to bottom. The dimension of the vector sequence Xpatch is P*P*C, where P*P is the resolution of the block, that is, the number of pixels in the block; C is the number of channels, representing the dimension of the image data. Each pixel has three channels to store the RGB components representing colors and three channels to store the gradient magnitude components corresponding to RGB, so the number of channels C is 6;
[0089] Specifically, in some embodiments, the input channels of the linear projection layer of the pre-trained ViT model are three channels, and the current input channels are six channels, which need to be expanded. The specific process includes: selecting the pre-trained ViT model with the original input of three channels as the basic pre-trained model; loading the parameters of the pre-trained ViT model and expanding the input channels: the first three channels (RGB) directly use the weight matrix of the pre-trained ViT model, and the last three channels (gradient) are initialized by staged training.
[0090] Further, in some embodiments, the staged training is specifically as follows: in the first stage, freeze the projection weights corresponding to the three RGB channels and only train the relevant parameters of the gradient channels; in the second stage, copy the original RGB three-channel weights to the gradient channels and fine-tune the entire model using a low learning rate (1e-5).
[0091] Specifically, in some embodiments, the vector sequence Xpatch of each patch is mapped to a d-dimensional embedding vector through the trained linear projection layer.
[0092] Further, in some embodiments, the vector sequence of the small patch corresponds to the feature vector Fsmall of the small patch, the dimension of the feature vector Fsmall of the small patch is d2 (P2 * P2 * 6), the vector sequence of the large patch corresponds to the feature vector Flarge of the large patch, and the dimension of the feature vector Flarge of the large patch is d1 (P1 * P1 * 6). They are unified to the same dimension d through the linear projection layer:
[0093] F small =Linear(F small ,d), F large =Linear(F large ,d);
[0094] On the basis of the pre-trained ViT model, expand the channels and design a six-channel linear projection layer (RGB + gradient). When designing the weights of the three gradient channels, adopt a staged training strategy from freezing the original weights to fine-tuning with a low learning rate, effectively utilizing the basis of the pre-trained ViT model and avoiding overfitting under small data.
[0095] S4: Add positional encoding to the embedding vector and generate the feature vector Epatch.
[0096] Specifically, in some embodiments, for the center point (xi, yi) of each patch, calculate its normalized coordinates:
[0097] where S is the patch size, and W and H are the width and height of the input image.
[0098] Specifically, in some embodiments, 2D positional encoding is added to the embedding vectors corresponding to each patch to generate a feature vector Epatch, and the calculation formula is:
[0099] Epatch = Linear(Flatten(Xpatch)) + PE2D(xi, yi);
[0100] Preferably, in some embodiments, the calculation formula for 2D positional encoding using two-dimensional sine positional encoding is:
[0101] PE2D(xi, yi) = Concat(PE(xi), PE(yi)) ∈ Rd, where PE(·) is the sine function encoding of the traditional ViT, and the specific calculation formula is:
[0102] PE(pos, 2k) = sin(100002k / dpos), PE(pos, 2k + 1) = cos(100002k / dpos);
[0103] Generating positional encoding based on the patch center coordinates, while retaining spatial position information and being compatible with different patch sizes, enhances the model's understanding of the image structure;
[0104] S5: According to the method type of dynamic adaptive patching, use the multi-head attention mechanism or cross-scale cross-attention mechanism of Transformer to process the feature vector Epatch, and output the class probability distribution;
[0105] Obtain the method type adopted when performing dynamic adaptive patching. If average patching is adopted, use the multi-head attention mechanism of Transformer to process the feature vector Epatch. If hybrid patching is adopted, use the cross-scale cross-attention mechanism of Transformer to process the feature vector Epatch;
[0106] Specifically, in some embodiments, the specific process of using the multi-head attention mechanism of Transformer to process the feature vector Epatch includes: forming an embedding sequence of the input image from the feature vectors Epatch of all patches, and splicing the CLS token to the head of the embedding sequence of the input image to form an input sequence:
[0107] Input = [CLS, Epatch1, Epatch2,..., Epatch N ;
[0108] Furthermore, in some embodiments, the CLS token is a learnable vector initialized randomly; in each layer of the Transformer encoder, the CLS token interacts with the feature vectors Epatch of all patches of the input image through multi-head self-attention: the CLS token serves as the "query" (Q), attends to the "key-value" (K-V) pairs of all patches, and through the attention weights, the CLS token integrates the semantic information of different patches. After passing through all the Transformer layers, the output vector of the CLS token is extracted as the global representation of the input image; the output vector of the CLS token is input into a fully connected layer to output the class probability distribution.
[0109] Specifically, in some embodiments, the specific process of processing the feature vector Epatch using the cross-scale cross-attention mechanism of the Transformer includes:
[0110] The first sub-step: Obtain the feature vector Epatch of the large patch and label it as the feature vector
[0111] Obtain the feature vector Epatch of the small patch and label it as the feature vector
[0112] Feature vector As the query (Query), the feature vector
[0113] As the key-value pair (Key-Value), map the feature vector and the feature vector to the query (Q), key (K), and value (V) spaces respectively, and the calculation formula is:
[0114] Q = F large W q K = F small W k V = F small W v ;
[0115] where W q W k W v ∈R d×d is a learnable parameter matrix that maps the input features to a unified attention space;
[0116] The second sub-step: Calculate the correlation between the feature vector Flarge and the feature vector Fsmall through scaled dot-product attention, and the calculation formula is:
[0117] Used to scale the dot product result to prevent gradient vanishing; the Softmax function normalizes the weights into a probability distribution;
[0118] The third sub-step: Use the attention weights to weight the feature vector Fsmall to generate the fused feature vector Ffused, so that each large block feature incorporates the related local detail information. The specific calculation formula is:
[0119]
[0120] The fourth sub-step: Dynamically adjust the contributions of the fused feature vector Ffused and the original feature, suppress the noise region, enhance the key details, and input the fused feature into the gating network to generate a weight matrix:
[0121] where W g ∈R 2d×d is a learnable parameter, σ is the Sigmoid function σ(x) = 1 / (1 + e^(-x)), which maps continuous values to probability weight output values in the range of [0, 1]; when the gradient amplitude of the input feature is large, the output of the σ function is close to 1, retaining the complete information of the feature channel; otherwise, the noise feature is attenuated;
[0122] The fifth sub-step: Dynamically mix the fused feature and the original feature through the gating weight. The specific calculation formula is:
[0123] ⊙ represents element-wise multiplication. The gating mechanism enables the model to autonomously decide whether to rely on the fused feature or retain the original global feature;
[0124] Exemplarily, if a large block corresponds to a smooth background area, the gating weight G approaches 0, and the final feature mainly retains the original large block information; if it corresponds to an edge area, G approaches 1, and more of the fused detail features are adopted;
[0125] The sixth sub-step: Perform global average pooling on the fused feature vector to obtain the feature vector f ∈ R d ; The specific operation process of global average pooling includes: calculating the average value of all spatial positions of each channel of the fused feature vector, and finally outputting a global feature vector f with a length of the number of channels C;
[0126] Input the global feature vector f into a multi-layer perceptron classifier to output the class probability. The specific calculation formula is:
[0127] Logits = MLP(f) ∈ R C , where C is the number of classes;
[0128] Multi-scale feature fusion is completed by cross-scale attention guidance, feature splicing, and dynamic selection of gated fusion features and original features, achieving efficient fusion of global semantics and local details, and finally generating highly discriminative image-level representations, significantly improving the performance of fine-grained classification tasks.
[0129] The above content is only an example and illustration of the structure of the present invention. Those skilled in the art of the present technology can make various modifications or supplements to the described specific embodiments or use similar methods to replace them, as long as they do not deviate from the structure of the invention or exceed the scope defined by the claims of the present invention, they should fall within the protection scope of the present invention.
[0130] In the description of this specification, the description with reference to terms such as "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0131] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the present invention to the specific implementation manners. Obviously, according to the content of this specification, many modifications and changes can be made. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. The present invention is only subject to the claims and their full scope and equivalents.
Claims
1. An image classification method based on the Transformer neural network, characterized in that, It includes the following steps: Obtain an input image, and use the Sobel operator to calculate the gradient magnitude component of each pixel in the input image respectively; Calculate the gradient variation value of the input image according to the gradient magnitude component, and perform dynamic adaptive block division on the input image according to the gradient variation value; Flatten the divided blocks into a vector sequence, expand the weights of the linear projection layer according to the pre-trained ViT model, and linearly project the vector sequence into d-dimensional embedding vectors through the linear projection layer; Add position encoding to the embedding vectors and generate feature vectors; According to the type of the dynamic adaptive block division, use the multi-head attention mechanism or the cross-scale cross-attention mechanism of the Transformer to process the feature vectors, and output the class probability distribution of the input image.
2. The image classification method based on the Transformer neural network according to claim 1, wherein, The specific process of calculating the gradient magnitude component of each pixel using the Sobel operator includes: Each pixel takes its 3*3 neighborhood to form a pixel region; Align the three RGB components of the pixel region with the centers of two groups of preset convolution kernels respectively, perform element-wise multiplication and then sum to obtain the horizontal gradient component and the vertical gradient component; Perform a square root operation on the sum of the squares of the horizontal gradient component and the vertical gradient component to obtain the gradient magnitude component of each pixel on the three RGB channels. The calculation formula is: Among them, G x,R , G x,G , G x,B are horizontal gradient components, and G y,R , G y,G , G y,B are vertical gradient components; Normalize the gradient magnitude component.
3. A method for image classification based on a Transformer neural network according to claim 1, characterized in that, The process of performing dynamic adaptive block division on the input image according to the gradient variation value includes: Calculate the gradient magnitude TF of each pixel in the input image; Calculate the gradient variation value of the input image according to the gradient magnitude TF of the pixel, and compare the gradient variation value with a preset gradient variation threshold: If the gradient variation value is greater than or equal to the gradient variation threshold, adopt hybrid block division; If the gradient variation value is less than the gradient variation threshold, adopt average block division.
4. A method for image classification based on a Transformer neural network according to claim 3, characterized in that, The specific process of the average block division includes: Calculate the gradient mean of the input image according to the gradient magnitude TF of the pixel, and compare the gradient mean with a preset gradient average threshold: if the gradient mean is greater than or equal to the gradient average threshold, adopt the average block division with small blocks, if the gradient mean is less than the gradient average threshold, adopt the average block division with large blocks The specific process of hybrid block division includes: uniformly divide the input image into several large blocks, calculate the gradient mean of each large block and compare it with the gradient average threshold: if the gradient mean is less than the gradient average threshold, no processing is required, if the gradient mean is greater than or equal to the gradient average threshold, divide the large block into small blocks evenly.
5. A method for image classification based on a Transformer neural network according to claim 1, characterized in that, The specific process of expanding the weights of the linear projection layer includes: load the pre-trained ViT model, expand the input channels of the pre-trained ViT model to six channels, use the projection weights of the pre-trained ViT model for the first three channels, and initialize the last three channels through staged training.
6. The image classification method based on the Transformer neural network according to claim 5, wherein The specific process of the staged training includes: in the first stage, freeze the projection weights corresponding to the three RGB channels of the pre-trained model, and only train the relevant parameters of the gradient magnitude component channel; in the second stage, copy the projection weights corresponding to the three RGB channels to the gradient channels and adjust them with a low learning rate.
7. A method for image classification based on the Transformer neural network according to claim 1, characterized in that, The calculation formula for mapping the vector sequence of each block into d-dimensional embedding vectors is: F small = Linear(F small , d), F large = Linear(F large , d), where Fsmall is the small-block feature corresponding to the vector sequence of small blocks, and Flarge is the large-block feature corresponding to the vector sequence of large blocks.
8. A method for image classification based on a Transformer neural network according to claim 1, characterized in that, The specific process of adding positional encoding to the embedded vector to generate the input feature vector includes: For each block center point (xi, yi), calculate its normalized coordinates. Add 2D positional encoding to the embedded vector corresponding to each block. The calculation formula is: Epatch = Linear(Flatten(Xpatch)) + PE2D(xi, yi); The calculation formula for 2D positional encoding is: PE2D(xi, yi) = Concat(PE(xi), PE(yi)) ∈ Rd, where PE(·) is the sine function encoding of ViT.
9. A method for image classification based on a Transformer neural network according to claim 1, characterized in that, The specific process of processing the feature vector using the multi-head attention mechanism includes: Concatenate the randomly initialized CLS token to the head of the feature vector to form the input sequence. The CLS token obtains the output vector of the CLS token by interacting with the input sequence, and the output vector of the CLS token is input into the fully connected layer to output the class probability distribution.
10. A method for image classification based on a Transformer neural network according to claim 1, characterized in that, The specific process of processing the feature vector by the cross-scale cross-attention mechanism includes: Obtain the feature vector Flarge of the large block and the feature vector Fsmall of the small block. Map the feature vector Flarge as the query and the feature vector Fsmall as the key-value pair to the query Q, key K, and value V spaces respectively. Calculate the correlation between the feature vector Flarge and the feature vector Fsmall. The calculation formula is: The Softmax function normalizes the weights into a probability distribution; Generate the fused feature vector Ffused by weighting the feature vector Fsmall of the small block. The specific calculation formula is: Dynamically adjust the fused feature vector Ffused and the original feature vector Flarge of the large block to generate the weight matrix G: where W g ∈R 2d×d is a learnable parameter and σ is the Sigmoid function; Dynamically mix the feature vector Ffused and the original feature vector Flarge of the large block in combination with the weight matrix G. The specific calculation formula is: ⊙ represents element-wise multiplication; For the mixed features perform global average pooling to obtain a global feature vector f ∈ R d ; input the global feature vector f into a multi-layer perceptron classifier to output class probabilities. The specific calculation formula is as follows: Logits = MLP(f) ∈ ℝ C , where C is the number of classes.
Citation Information
Cited By
Cross-modal image-text analysis method for machine vision
CN121210958A
Machine vision-oriented cross-modal graphic-text analysis method
CN121210958B