Convolution and sparse vision Transform fused efficient feature learning method and system

By designing the convolution and sparse self-attention fusion module (CATB) and building a CATFormer network, the problem of limited local information capture capabilities and high computational complexity in feature learning is solved, efficient global feature learning is achieved, and the performance of image processing tasks is improved.

CN120337996APending Publication Date: 2025-07-18UNIV OF JINAN
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510482870.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing convolutional neural networks and visual Transformers have problems with limited local information capture capabilities and high computational complexity in feature learning, making it difficult to effectively integrate global information.

Method used

Convolution and sparse self-attention fusion module (CATB) is designed, and the self-attention mechanism is sparse by sparse query vectors and combined with convolutional diffusion characteristics. Multiple CATB modules are stacked to form a CATFormer network, forming a feature pyramid structure, reducing the computational complexity and efficiently capturing global features.

Benefits of technology

While reducing the amount of calculation, CATFormer can effectively capture the global information of the image and transmit global features, significantly improving the expressive ability of feature representation, and is suitable for tasks such as image classification, object detection and semantic segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337996A_ABST
    Figure CN120337996A_ABST
Patent Text Reader

Abstract

The invention provides an efficient feature learning method and system fusing convolution and sparse vision Transform, and relates to the field of computer vision. According to the method, a convolution and sparse self-attention fusion module (CATB) is designed as a main construction block of a neural network CATFormer, uniform sampling and interaction are carried out in an image processing task according to the spatial position of a feature map, the sparsification of a self-attention mechanism is realized, the self-attention calculation complexity is reduced, and meanwhile, the self-attention fusion efficiency is improved. Global information of an image captured by a neural network is realized, and the global information is effectively transmitted by using the diffusion characteristic of convolution, so that efficient feature learning is realized; a plurality of CATB modules are stacked to form a neural network CATFormer, and a sampling step size # imgabs0 # is set in the CATB to realize sparse sampling, so that a token with global representativeness is extracted, and the obtained neural network model can be used for trunks of visual tasks such as image classification, target detection, instance segmentation, semantic segmentation and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and specifically to an efficient feature learning method and system that combines convolution and sparse vision transformers. Background Art

[0002] Convolutional neural networks (CNNs) and vision transformers (ViTs) are two mainstream feature learning methods, showing good prospects in multiple research and application fields. CNNs capture local features through the sliding window of convolutional kernels, which limits their ability to capture global information. ViTs can capture global information well, but have a high computational complexity. Local ViTs reduce the computational complexity through non-overlapping windows, and this window limitation usually requires relying on deeper network layers to effectively fuse features through cross-window interactions. To address these issues, many studies have focused on optimizing the network structure to reduce the computational complexity, such as designing efficient convolutional operations and exploring the sparse attention mechanism of ViTs. Convolutional operations use large-kernel convolutions to increase the receptive field and use position offsets to break through the limitations of small convolutional kernel windows. Similarly, ViTs have adopted non-overlapping sub-windows and sparse attention mechanisms to improve efficiency without sacrificing the advantage of the global receptive field. Summary of the Invention

[0003] The present invention provides an efficient feature learning method that combines convolution and sparse vision transformers, including the following steps: S1. Design a convolution and sparse self-attention fusion module (CATB) as the main building block of the neural network. In image processing tasks, efficiently capture global features while reducing the computational amount by sparsifying the query vector, and use the diffusion characteristics of convolution to effectively transmit global features, enabling the neural network to capture the long-range and short-range dependencies of features and perform efficient feature learning; S2. Stack multiple CATB network modules to form a convolution and sparse vision transformer fusion network CATFormer. CATFormer includes four stages, each stage is composed of CATB. In CATB, sample tokens according to the stride to sparsify the query vector, efficiently capture global features while reducing the computational amount, and then use the diffusion ability of convolution to effectively transmit global features. Perform downsampling operations on the image between stages to form a feature pyramid structure.

[0004] Preferably, the convolution and sparse self-attention fusion module (CATB) in S1 is composed of a convolutional position embedding (CPE), a sparse multi-head self-attention (SMHSA), a local-global information fusion convolution (LGC), and a feed-forward network (FFN). Given an input image , where C is the embedding dimension, H is the height of the input image, and W is the width of the input image. When input into the CATB, it can be represented by the following process: , , , .

[0005] Preferably, for the CATB in S1, the CATB uniformly samples tokens globally according to the spatial positions in the input feature map, so as to extract tokens with global representativeness for sparse self-attention calculation. Specifically, the SMHSA first performs uniform sampling on the input feature map, selects tokens at intervals of stride in the length and width directions. The final number of sampled tokens is . These sampled tokens are linearly mapped to generate query vectors . At the same time, the input feature map is downsampled by convolution and linearly mapped to generate key-value vectors , , the sampled , , generate global feature information after being calculated by the self-attention mechanism, and finally restore the tokens containing global information to the original sampling positions to complete the global enhancement process of the features.

[0006] Preferably, the CATB in S1 utilizes the diffusion ability of convolution to effectively transfer global features to surrounding regions, strengthens the interaction between local features and global context, and significantly improves the expression ability of the overall feature representation. Specifically, the feature map output by the SMHSA contains both global information tokens and local information tokens. In the LGC, the interaction between global and local information is promoted through hierarchical convolution operations, and further information fusion between channels is completed through point convolution to further optimize the feature expression;

[0007] Preferably, the neural network CATFormer in S2 includes four stages. Each stage consists of a CATB, and image downsampling operations are performed between stages to form a feature pyramid structure. Each stage of the CATFormer adopts a CATB module, samples tokens according to stride to sparsify the query vectors, efficiently capture global features while reducing the computational amount, and then utilize the diffusion ability of convolution to effectively transfer global features.

[0008] The present invention also provides an efficient feature learning system integrating convolution and sparse vision transformers, which is characterized by including: a convolution and sparse vision transformer fusion module (CATB), which samples tokens according to the spatial position of the feature map and interacts to sparsify the self-attention mechanism, reducing the computational complexity of self-attention while ensuring that the neural network can effectively capture the global information of the image. In addition, by utilizing the diffusion characteristics of convolution, it can effectively transmit global information, thus achieving efficient feature learning; a stacking module, where multiple CATB modules are stacked to form a four-stage neural network, and different strides are set in each stage Sampling tokens sparsifies the query vector, efficiently capturing global features while reducing the computational amount, and then using the diffusion ability of convolution to effectively transmit the global features, which is an efficient feature learning system.

[0009] Compared with the prior art, the present invention has the following technical effects: CATFormer is a neural network architecture that fuses a sparse vision transformer that sparsifies query vectors with convolution. By sparsifying query vectors, it efficiently captures global features while reducing the computational amount, and uses the diffusion characteristics of convolution to effectively transmit global features, enabling the neural network to capture the long-range and short-range dependencies of features and perform efficient feature learning. Each stage of CATFormer is composed of CATB, and different strides are set in SMHSA in CATB Sampling tokens sparsifies the query vector, efficiently capturing global features while reducing the computational amount. In CATB, the diffusion ability of convolution is used to effectively transmit global features to the surrounding areas, strengthening the interaction between local features and global context and significantly improving the expressive ability of the overall feature representation. Description of the Drawings

[0010] Figure 1 It is the neural network architecture diagram provided by the embodiment of the present invention; Figure 2 It is the configuration of the neural network family provided by the embodiment of the present invention; Figure 3 It is the classification performance comparison result diagram of the method of the embodiment of the present invention and several different methods on the ImageNet-1K dataset; Figure 4 It is the object detection and instance segmentation performance comparison result diagram of the method of the embodiment of the present invention and several different methods on the COCO val2017 dataset; Figure 5 It is the semantic segmentation performance comparison result diagram of the method of the embodiment of the present invention and several different methods on the ADE20K dataset. Detailed Embodiment

[0011] The present invention aims to provide an efficient feature learning method and system that integrates convolution and sparse vision transformers, designs a convolution and sparse self-attention fusion module (CATB) as the main building block of the neural network CATFormer, uniformly samples and interacts according to the spatial position of the feature map in image processing tasks, realizes the sparsification of the self-attention mechanism, reduces the computational complexity of self-attention while enabling the neural network to capture the global information of the image, and utilizes the diffusion characteristics of convolution to effectively transmit the global information, thereby achieving efficient feature learning; stacking multiple CATB modules constitutes the neural network CATFormer, setting the sampling step r in CATB to achieve sparse sampling, thereby extracting tokens with global representativeness, and the obtained neural network model can be used as the backbone for visual tasks such as image classification, object detection, instance segmentation, and semantic segmentation.

[0012] Please refer to Figure 1 As shown, an efficient feature learning neural network method and system that integrates convolution and sparse vision transformers in the embodiments of the present application: S1. Design a convolution and sparse self-attention fusion module (CATB) as the main building block. In image processing tasks, while reducing the computational amount by sparsifying the query vector, efficiently capture global features, and utilize the diffusion characteristics of convolution to effectively transmit global features, enabling the neural network to capture the long-range and short-range dependencies of features and perform efficient feature learning; S2. Stacking multiple CATB network modules constitutes a convolution and sparse vision transformer fusion network CATFormer. CATFormer includes four stages, each stage is composed of CATB. In CATB, sample tokens according to the stride to achieve the sparsification of the query vector, efficiently capture global features while reducing the computational amount, and then utilize the diffusion ability of convolution to effectively transmit global features. Downsampling operations of the image are performed between stages to form a feature pyramid structure.

[0013] Furthermore, for the convolution and sparse self-attention fusion module (CATB) in S1, CATB is composed of convolution position embedding (CPE), sparse multi-head self-attention (SMHSA), local-global information fusion convolution (LGC), and feed-forward network (FFN). Given an input image , where C is the embedding dimension, H is the height of the input image, and W is the width of the input image, input it into CATB, which can be represented by the following process: , , , . First, through with zero padding Depth convolution encodes positional information into , which is a learnable positional embedding. After positional encoding, is added to the original value through a residual connection to obtain a feature map . Then, SMHSA is used to perform discrete self-attention calculation on the obtained feature map for discrete sampling tokens to obtain a global feature representation. After that, the tokens containing global information are restored to the original sampling positions to ensure the consistency of the input and output dimensions, resulting in a new feature representation . N represents the number of all tokens, and C is the embedding dimension. is effectively transmitted to the surrounding regions through LGC to strengthen the interaction between local features and global context, obtaining a new feature representation . is input into FFN to achieve cross-channel interaction. FFN consists of two linear layers, and a GELU function is added between the two linear layers. Finally, the output of FFN is added to the original value through a residual connection to obtain . is the output of CATB.

[0014] The CATB in S1 uniformly samples tokens globally according to the spatial positions in the input feature map, thereby extracting tokens with global representativeness for sparse self-attention calculation. Specifically, SMHSA first uniformly samples the input feature map , selects tokens at intervals of stride in the length and width directions, which is set to different values according to different stages, [4, 2, 1, 1] for the first to fourth stages respectively. The final number of sampled tokens is . These sampled tokens are linearly mapped to generate query vectors . At the same time, the input feature map is downsampled by convolution and linearly mapped to generate key-value vectors . . The sampled , , generate global feature information after being calculated by the self-attention mechanism. Finally, the tokens containing global information are restored to the original sampling positions to obtain a feature representation containing local and global information. In LGC, the global feature is first effectively transmitted to the surrounding regions through hierarchical convolution to strengthen the interaction between local features and global context, and then information interaction between channels is achieved through point convolution, significantly improving the expressive ability of the overall feature representation.

[0015] Furthermore, the neural network CATFormer in S2, where CATFormer consists of four stages, each stage is composed of CATB, and downsampling operations of images are performed between stages to form a feature pyramid structure. We instantiate WGViT with five different model sizes by setting different numbers of blocks, as Figure 2 shown. To obtain hierarchical feature maps, we use downsampling layers between each stage to adjust the different scales of the feature maps. Figure 2 The Blocks in correspond to the number of blocks in the four stages of CATFormer. The number of blocks in each stage is the number of blocks of CATB. Each stage of CATFormer adopts the CATB module. According to the stride

[0016] This example provides an efficient feature learning system that fuses convolution and sparse vision Transformer. It is characterized by a convolution and sparse vision Transformer fusion module (CATB) that uniformly samples tokens according to the spatial position of the feature map and interacts, realizing the sparsification of the self-attention mechanism. While reducing the computational complexity of self-attention, it ensures that the neural network can effectively capture the global information of the image. In addition, using the diffusion characteristics of convolution, it can effectively transmit global information, thus realizing efficient feature learning; a stacking module, where multiple CATB modules are stacked to form a four-stage neural network. Different strides are set in each stage to sample tokens to sparsify the query vector, efficiently capture global features while reducing the computational amount, and then use the diffusion ability of convolution to effectively transmit global features, which is an efficient feature learning system that fuses convolution and sparse vision Transformer. The obtained neural network model can be used as the backbone for visual tasks such as image classification, object detection, instance segmentation, and semantic segmentation.

[0017] Furthermore, as Figure 3 , Figure 4 and Figure 5 shown, Figure 3 shows the image classification results of CATFormer on the ImageNet-1K dataset. Figure 4 shows the results of CATFormer as the backbone network for object detection and instance segmentation tasks on the COCO dataset. Figure 5shows its performance in the semantic segmentation task on the ADE20K dataset. It can be seen that while maintaining a low number of parameters and computational complexity, CATFormer can still achieve excellent performance in a variety of downstream tasks, fully demonstrating its good balance between performance and efficiency.

[0018] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the inventive concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.

Claims

1. An efficient feature learning method that fuses convolution and sparse vision transformers, characterized in that, Including the following steps: S1. Design a Convolutional and Sparse Self-Attention Fusion Module (CATB) as the main building block of the neural network. In image processing tasks, by sparsifying the query vector, it can efficiently capture global features while reducing the computational cost, and utilize the diffusion property of convolution to effectively transmit global features, enabling the neural network to capture long-range and short-range dependencies of features; S11. For the Convolution and Sparse Self-Attention Fusion Module (CATB), CATB consists of Convolutional Position Embedding (CPE), Sparse Multi-Head Self-Attention (SMHSA), Local-Global Information Fusion Convolution (LGC), and Feed-Forward Network (FFN). Given an input image , where C is the embedding dimension, H is the height of the input image, and W is the width of the input image. Inputting it into CATB can be represented by the following process: , , , ; S12. CATB uniformly samples tokens globally according to the spatial positions in the input feature map, so as to extract globally representative tokens for sparse self-attention calculation. Specifically, SMHSA first performs uniform sampling on the input feature map by taking tokens at intervals with a stride in the length and width directions. Finally, the number of tokens obtained by sampling is . These sampled tokens are linearly mapped to generate query vectors . At the same time, the input feature map is downsampled by convolution and linearly mapped to generate key-value vectors . , , , generate global feature information after being calculated by the self-attention mechanism. Finally, the tokens containing global information are restored to the original sampling positions to complete the global enhancement process of the features; S13. CATB utilizes the diffusion property of convolution to propagate global features to the surrounding regions, realizing the interaction between local features and global context. Specifically, the feature map output by SMHSA contains both global information tokens and local information tokens. In LGC, first, hierarchical convolution operations are used to promote the interaction between global and local information, and then pointwise convolution is used to complete the information fusion between channels; S2. Stacking multiple CATB network modules constitutes the convolutional and sparse vision transformer fusion network CATFormer. CATFormer consists of four stages, each stage is composed of CATB. In CATB, tokens are sampled according to the stride to sparsify the query vector, efficiently capture global features while reducing the computational cost, and then utilize the diffusion ability of convolution to effectively transmit the global features. Image downsampling operations are performed between stages to form a feature pyramid structure.

2. An efficient feature learning method integrating convolution and sparse vision Transformer according to claim 1, characterized in that, The neural network CATFormer in S2, where CATFormer consists of four stages, each stage is composed of CATB, and downsampling operations of images are performed between stages to form a feature pyramid structure. Each stage of CATFormer is a CATB module. According to the stride Sample tokens to capture global features, and then utilize the diffusion ability of convolution to achieve effective transmission of global features.

3. An efficient feature learning system integrating convolution and sparse vision transformers, characterized in that, Including: Convolutional and Sparse Vision Transformer Fusion Module (CATB), which samples tokens according to the spatial position of the feature map and interacts, realizing the sparsification of the self-attention mechanism. While reducing the computational complexity of self-attention, it ensures that the neural network can effectively capture the global information of the image. In addition, by utilizing the diffusion property of convolution, it can effectively transmit global information, thus achieving efficient feature learning; Stacked module, multiple CATB modules are stacked to form a four-stage neural network, and different strides are set for each stage Sparse sampling tokens are used to capture global features, and then the diffusion ability of convolution is utilized to achieve effective transmission of global features, which is an efficient feature learning system for implementing any one of the methods described in claims 1-2.