CNN-Transform staged fusion-based image classification method and system
By adopting a phased fusion method based on CNN-Transformer in the image classification task, using a multi-level feature interaction mechanism and an adaptive gating fusion module, the problem of waste of resources and inconsistent optimization goals of existing hybrid models is solved, and high-precision image classification is achieved, especially in complex scenarios to improve robustness.
Patent Information
- Application Number
- CN202510226187.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing CNN-Transformer hybrid model has problems of wasting resources and inconsistent optimization goals in image classification tasks, resulting in local optimal traps and the dynamic interaction between local and global features is not fully considered.
Using a phased fusion method based on CNN-Transformer, multi-level local features are extracted and global semantic features are generated through a multi-level feature interaction mechanism and an adaptive gating fusion module, adaptive weighted fusion is realized, and shallow and deep fusion features are fused through cross-level jump connections.
It significantly improves image classification accuracy, enhances the complementarity between local details and global context features, and improves classification robustness, especially in complex backgrounds or small-target scenarios.
Smart Images

Figure CN120147730A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly relates to an image classification method and system based on staged fusion of CNN-Transformer. Background Art
[0002] In recent years, convolutional neural networks (CNNs) and Transformer models have made remarkable progress in the field of computer vision and become the mainstream methods for image classification tasks. CNNs extract local features through convolutional kernels, with the advantages of translational invariance and parameter sharing, and are widely used in image classification tasks. Typical CNN models include ResNet, EfficientNet, and MobileNet, etc. Transformer was initially used in natural language processing (NLP) and was later introduced into the field of computer vision through Vision Transformer (ViT). Transformer captures global dependencies through the self-attention mechanism, breaking through the locality limitation of CNNs. In order to combine the local feature extraction ability of CNNs and the global modeling ability of Transformer, researchers have proposed various hybrid models, such as ConvNeXt, BoTNet, and MobileViT, etc. Existing hybrid models mostly adopt fixed-ratio fusion or simple concatenation, without fully considering the dynamic interaction between local and global features; the independent calculations of CNN and Transformer modules lead to resource waste; the optimization objectives of CNN and Transformer are inconsistent, and it is easy to fall into local optima during joint training. Summary of the Invention
[0003] Technical Objective: Aiming at the existing deficiencies, the present invention discloses an image classification method and system based on staged fusion of CNN-Transformer. Through a multi-level feature interaction mechanism and an adaptive gating fusion module, while reducing the computational complexity, the image classification accuracy is significantly improved.
[0004] Technical Solution: To achieve the above technical objective, the present invention adopts the following technical solutions:
[0005] An image classification method based on staged fusion of CNN-Transformer, comprising the following steps:
[0006] The input image is preprocessed and then input into the CNN backbone network to extract multi-level local features;
[0007] Non-overlapping image patches are divided on the middle and high-level feature maps of the CNN backbone network, and after adding learnable position encoding, they are input into the Transformer module to generate global semantic features;
[0008] The local features of the CNN backbone network and the global features of the Transformer module are adaptively weighted and fused through a dynamic gating fusion module;
[0009] The shallow features of the CNN backbone network and the deep features of the Transformer module are fused through cross-level skip connections to output the classification results.
[0010] The present invention also provides an image classification system based on phased fusion of CNN-Transformer, including:
[0011] A data preprocessing module for image size normalization, enhancement, and standardization;
[0012] A phased fusion module for implementing an image classification method based on phased fusion of CNN-Transformer as described above;
[0013] A deployment optimization module that supports model quantization, operator fusion, and deployment on edge devices.
[0014] Advantageous effects: An image classification method and system based on phased fusion of CNN-Transformer provided by the present invention have the following advantageous effects:
[0015] 1. The present invention extracts multi-level local features through the CNN backbone network, combines the global semantic features generated by the Transformer module, and realizes adaptive weighted fusion through the dynamic gating fusion module. This module uses a channel attention branch and a spatial attention branch to jointly calculate the fusion weights, significantly enhancing the complementarity of local details and global context features. The dynamic weight allocation can be adaptively adjusted according to the input content, avoiding information redundancy or loss, thereby improving the classification accuracy.
[0016] 2. The present invention solves the problem of size mismatch between shallow features and deep features in traditional methods by cross-level fusing the CNN shallow features and the Transformer deep features, using a combination of adaptive pooling and bilinear upsampling. This design not only retains the low-level details but also enhances the expression ability of high-level semantics. In complex background or small target scenarios, the classification robustness is improved by about 15%.
[0017] 3. The present invention introduces learnable position encoding in the Transformer module to replace the fixed position encoding of traditional ViT. This design enables the model to automatically learn the spatial position relationship according to the task requirements. Especially when dealing with irregular objects or rotated and scaled images, the ability to model spatial information is significantly improved. Description of the Drawings
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art.
[0019] Figure 1 This is the overall flowchart of the image classification method based on stage-by-stage fusion of CNN-Transformer of the present invention;
[0020] Figure 2 This is the overall architecture diagram of the stage-by-stage fusion of CNN-Transformer of the present invention;
[0021] Figure 3 This is the detailed diagram of the dynamic gating fusion module of the present invention;
[0022] Figure 4 This is the schematic diagram of the cross-level skip connection of the present invention;
[0023] Figure 5 This is the overall block diagram of the image classification system based on stage-by-stage fusion of CNN-Transformer of the present invention. Detailed implementation manners
[0024] The following will more clearly and completely illustrate the present invention by way of a preferred embodiment in conjunction with the accompanying drawings, but the present invention is not limited to the scope of the described embodiments.
[0025] As Figure 1 and Figure 2 shown, an image classification method based on stage-by-stage fusion of CNN-Transformer includes the following steps:
[0026] S1. After preprocessing, the input image is input into the CNN backbone network to extract multi-level local features;
[0027] The preprocessing of the input image includes size normalization, data augmentation, and standardization. Size normalization is to scale the size of the input image to a fixed size such as 224×224. Data augmentation applies random horizontal flipping, rotation (±15°), and color jitter (brightness and contrast adjustment ±20%). Standardization is normalized according to the ImageNet mean ([0.485, 0.456, 0.406]) and standard deviation ([0.229, 0.224, 0.225]).
[0028] A depthwise separable convolution is used to construct the CNN backbone network, such as the improved MobileNetV3. The feature levels extract shallow features (such as the 2nd layer, size 112×112) to capture edges and textures, middle-level features (such as the 4th layer, 56×56) to capture local semantics, and deep features (such as the 6th layer, 28×28) to capture high-level abstractions. The calculation formula for a single layer is:
[0029] F out = ReLU6(BN(DConv 3×3 (F in )))
[0030] where DConv 3×3 is depthwise separable convolution, BN is batch normalization, and ReLU6 is an activation function.
[0031] S2. Divide non-overlapping image patches on the middle and high-level feature maps of the CNN backbone network, add learnable position encoding, and then input them into the Transformer module to generate global semantic features;
[0032] On the middle and high-level feature maps (such as 56×56) of the CNN backbone network, divide 16×16 image patches. Each patch is flattened into a vector, and then add learnable position encoding and input it into the Transformer module. Dynamically generate position encoding through a fully connected layer and add it to the patch vector. The generation formula of the learnable position encoding is:
[0033] P (i,j) = E·Embedding(i×W + j)
[0034] where i and j are the spatial coordinates of the feature map, W is the width of the feature map, is the learnable embedding matrix, D is the feature dimension, L is the maximum position encoding length, and Embedding represents a function that maps position indices to integer encodings.
[0035] Compared with the fixed sine encoding of ViT, the learnable encoding can adapt to task-specific spatial relationships. For example, for irregular lesions in medical images, the matrix E can be optimized through backpropagation and initialized with a random Gaussian distribution.
[0036] In one embodiment, the Transformer module adopts a linear attention mechanism, and its calculation process is:
[0037]
[0038] where represent the query matrix, key matrix, and value matrix respectively, n is the length of the input sequence, d is the feature dimension, 1 represents a vector of all 1s with dimension n×1, and diag represents extracting the diagonal elements of the matrix to generate a vector. The linear attention mechanism can replace the standard multi-head attention, reducing memory occupancy.
[0039] Diagonal approximation can be used to avoid calculating the complete attention matrix, saving GPU memory (measured to reduce by 40%)
[0040] S3. As Figure 3As shown, the local features of the CNN backbone network and the global features of the Transformer module are adaptively weighted and fused through the dynamic gating fusion module;
[0041] The local features F of the CNN backbone network cnn and the global features F of the Transformer module trans are concatenated along the channel dimension, and the channel attention branch generates channel dimension weights. The calculation formula is:
[0042]
[0043] where W c represents the channel dimension weight, which is used to weight the importance of each channel. Sigmoid represents the activation function, and MLP represents the multi-layer perceptron. There are two fully connected layers, and the dimension of the middle layer is 1 / 4 of the input channel number. For example, if the input is 512 dimensions, the middle layer is 128 dimensions. GAP represents global average pooling, which compresses the spatial dimension. represents channel concatenation, and F cnn and F trans are the local features of the CNN backbone network and the global features of the Transformer module respectively;
[0044] The spatial dimension weights are generated through the spatial attention branch. The calculation formula is:
[0045]
[0046] where W s represents the spatial dimension weight, and Conv represents the convolution formula;
[0047] The final fusion weights are calculated by combining the channel dimension weights and the spatial dimension weights. The calculation formula is:
[0048] G cnn =W c ·W s
[0049] G trans =1 - G cnn
[0050] where G cnn and G trans represent the dynamically generated weight matrices, and the dimension is the same as that of F cnn ;
[0051] S4. The shallow features of the CNN backbone network and the deep features of the Transformer module are fused through cross-level skip connections, and the classification results are output.
[0052] The cross-level skip connections include the following operations:
[0053] As Figure 4 shown, after adaptively pooling the shallow features (such as 112×112) of the CNN backbone network to the target size (such as 28×28), they are concatenated with the deep features of the Transformer module;
[0054] After bilinearly upsampling the deep features (such as 28×28) of the Transformer module to the shallow size (such as 112×112), they are added to the shallow features of the CNN backbone network.
[0055] The fusion method is to concatenate the shallow features after pooling with the deep features to enhance detail retention, and add the upsampled deep features to the shallow features to improve semantic information transmission.
[0056] In one embodiment, the target size of the adaptive pooling is 1 / 2 of the deep feature size output by the Transformer module.
[0057] As Figure 5 shown, the present invention also provides an image classification system based on CNN-Transformer staged fusion, including:
[0058] A data preprocessing module for image size normalization, enhancement and standardization. Image enhancement adopts CutMix (random cropping and mixing regions) and MixUp (linear mixing of images) strategies to improve the generalization ability of small samples. Standardization supports dynamic mean and standard deviation calculation (applicable to non-ImageNet datasets);
[0059] A staged fusion module for implementing an image classification method based on CNN-Transformer staged fusion as described above. It uses end-to-end training, uses a cross-entropy loss function, combines an AdamW optimizer (initial learning rate 3e-4, weight decay 0.05), and the training strategy adopts a cosine annealing learning rate scheduler with a maximum of 300 epochs;
[0060] A deployment optimization module that supports model quantization, operator fusion and deployment on edge devices. Model quantization is divided into post-training quantization (PTQ) and quantization-aware training (QAT). PTQ converts FP32 weights to INT8, and the calibration set uses 500 images. QAT simulates quantization errors during training, and the accuracy loss is controlled within 1%. Operator fusion combines Conv-BN-ReLU into a single operator, and the inference speed is increased by 20%. Edge deployment supports TensorRT engine optimization, and realizes 45FPS real-time inference on Jetson Nano.
[0061] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An image classification method based on CNN-Transformer staged fusion, characterized in that: The following steps are involved: After preprocessing, the input image is input into the CNN backbone network to extract multi-level local features; Divide non-overlapping blocks on the mid- and high-level feature maps of the CNN backbone network, add learnable position encodings and input them into the Transformer module to generate global semantic features; The local features of the CNN backbone network and the global features of the Transformer module are adaptively weighted and fused through the dynamic gated fusion module; The shallow features of the CNN backbone network are fused with the deep features of the Transformer module through cross-level skip connections to output the classification results.
2. The image classification method based on CNN-Transformer staged fusion according to claim 1, characterized in that: The generation formula of the learnable positional encoding is: P (i,j) =E·Embedding(i×W+j) Among them, i and j are the spatial coordinates of the feature map, W is the width of the feature map, is a learnable embedding matrix, D is the feature dimension, L is the maximum position encoding length, and Embedding represents a function that maps position indexes to integer encodings.
3. The image classification method based on CNN-Transformer staged fusion according to claim 1, characterized in that: The Transformer module uses a linear attention mechanism, and its calculation process is: in, They represent query matrix, key matrix, and value matrix respectively. n is the input sequence length, d is the feature dimension, 1 represents a full 1 vector, the dimension is n×1, and diag represents extracting the diagonal elements of the matrix to generate a vector. The linear attention mechanism can replace the standard multi-head attention and reduce memory usage.
4. The image classification method based on CNN-Transformer staged fusion according to claim 1, characterized in that: The dynamic gated fusion module generates weights by the following steps: Concatenate the local features of the CNN backbone network with the global features of the Transformer module; The channel dimension weight is generated by the channel attention branch, and the calculation formula is: Among them, W c represents the channel dimension weight, Sigmoid represents the activation function, MLP represents the multi-layer perceptron, GAP represents the global average pooling, represents channel splicing, F cnn and F trans They are the local features of the CNN backbone network and the global features of the Transformer module; The spatial dimension weight is generated by the spatial attention branch, and the calculation formula is: Among them, W s represents the spatial dimension weight, Conv represents the convolution formula; The final fusion weight is calculated by combining the channel dimension weight and the spatial dimension weight. The calculation formula is: G cnn =W c ·W s G trans =1-G cnn Among them, G cnn and G trans Represents a dynamically generated weight matrix with the same dimensions as F cnn same.
5. The image classification method based on CNN-Transformer staged fusion according to claim 4 is characterized in that: In the channel attention branch of the dynamic gated fusion module, the MLP contains two fully connected layers, and the dimension of the middle layer is one quarter of the number of input channels.
6. The image classification method based on CNN-Transformer staged fusion according to claim 1, characterized in that: The fusion formula is: F out =G cnn ⊙F cnn +G trans ⊙F trans Among them, F out The fusion formula of the local features of the CNN backbone network and the global features of the Transformer module, F cnn and F trans are the local features of the CNN backbone network and the global features of the Transformer module, G cnn and G trans represents the dynamically generated weight matrix, and ⊙ represents the element-by-element multiplication operation.
7. The image classification method based on CNN-Transformer staged fusion according to claim 1, characterized in that: The cross-level skip connection includes the following operations: Adaptively pooling the shallow features of the CNN backbone network to the target size, and then concatenating them with the deep features of the Transformer module; The deep features of the Transformer module are bilinearly upsampled and then added to the shallow features of the CNN backbone network.
8. The image classification method based on CNN-Transformer staged fusion according to claim 7, characterized in that: The target size of the adaptive pooling is 1 / 2 of the deep feature size output by the Transformer module.
9. The image classification method based on CNN-Transformer staged fusion according to claim 1, characterized in that: The CNN backbone network is a deep separable convolutional structure, and the single-layer calculation is: F out =ReLU6(BN(DConv 3×3 (F in ))) Among them, DConv 3×3 is a depth-wise separable convolution, BN is batch normalization, and ReLU6 is the activation function.
10. An image classification system based on CNN-Transformer staged fusion, characterized in that: include: Data preprocessing module for image size normalization, enhancement and standardization; A staged fusion module, used to implement an image classification method based on CNN-Transformer staged fusion as described in any one of claims 1 to 9; Deployment optimization module to support model quantization, operator fusion and deployment on edge devices.
Citation Information
Cited By
Group behavior identification method and system based on cross-feature interaction Transform
CN120388335A
Image interpretable classification method and device, computer equipment and storage medium
CN120783103A
Image segmentation method fusing state modeling and convolution perception mechanism
CN121010761A
Confusion expression recognition method and device fusing lightweight double-branch attention, equipment and storage medium
CN121074966A
Attention guidance-based airport runway line detection method
CN121170742A