Transformer Dense Prediction With Resolution-Preserving Patch Tokens
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing dense prediction techniques in computer vision suffer from loss of feature resolution and granularity due to down-sampling, which is hard to recover in downstream decoders, leading to inaccurate predictions.
Innovation Solution
An encoder-decoder architecture using vision transformers maintains constant feature dimensionality and global receptive field, leveraging a bag-of-words representation reassembled into image-like feature representations, and combining them using convolutional decoders to achieve fine-grained and globally coherent predictions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If down-sampling is used in fully-convolutional deep networks, then computational complexity is reduced, but feature resolution and granularity are lost
Solution Approach 1:
The patent segments the image into non-overlapping patches and processes each patch independently through the transformer encoder, maintaining original resolution while reducing computational complexity through localized processing. This segmentation approach allows the model to handle high-resolution images without the need for aggressive down-sampling.
Solution Approach 2:
The patent transforms the spatial dimension problem by projecting image patches into a token sequence dimension, where transformer self-attention mechanisms operate. This dimensional transformation allows global context to be captured without losing spatial resolution, as the attention mechanism processes all positions simultaneously rather than through hierarchical down-sampling.
2Adaptability or versatility
If down-sampling is applied in encoder stages, then global context is captured, but spatial details are lost and hard to recover in decoder
Solution Approach 1:
The transformer encoder serves multiple functions simultaneously: it captures global context through self-attention mechanisms while preserving spatial resolution through patch-based processing. The same encoder structure handles both contextual understanding and spatial feature extraction, eliminating the need for separate down-sampling operations that would lose detail.
Solution Approach 2:
The patent introduces patch embedding tokens as intermediaries between the input image and the transformer processing. These tokens maintain a one-to-one correspondence with image patches, serving as a bridge that preserves spatial information while enabling global context capture through transformer attention mechanisms.
3Adaptability or versatility
If transformer blocks are used instead of convolutions, then global receptive field is achieved, but computational requirements increase
Solution Approach 1:
By segmenting the image into patches and processing them as tokens, the patent reduces the computational burden of transformer self-attention from O(N²) where N is the number of pixels, to O(M²) where M is the number of patches. This segmentation dramatically reduces computational requirements while maintaining global receptive field through attention mechanisms.
Data Source
AI summary
Methods, apparatus, systems and articles of manufacture disclosed herein perform dense prediction of an input image using transformers at an encoder stage and at a reassembly stage of an image processing system. A disclosed apparatus includes an encoder with an embedder to convert an input image to a plurality of tokens representing features extracted from the input image. The tokens are embedded with a learnable position embedding. The encoder also includes one or more transformers configured in a sequence of stages to relate the tokens to each other. The apparatus further includes a decoder that includes one or more of reassemblers to assemble the tokens into feature representations, one or more of fusion blocks to combine the feature representations to generate a final feature representation, and an output head to generate a dense prediction based on the final feature representation and based on an output task.


