Transformer Dense Prediction With Resolution-Preserving Patch Tokens

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing dense prediction techniques in computer vision suffer from loss of feature resolution and granularity due to down-sampling, which is hard to recover in downstream decoders, leading to inaccurate predictions.

Innovation Solution

An encoder-decoder architecture using vision transformers maintains constant feature dimensionality and global receptive field, leveraging a bag-of-words representation reassembled into image-like feature representations, and combining them using convolutional decoders to achieve fine-grained and globally coherent predictions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If down-sampling is used in fully-convolutional deep networks, then computational complexity is reduced, but feature resolution and granularity are lost

Engineering Contradiction:
Improvecomputational complexityVSAvoidfeature resolution
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent segments the image into non-overlapping patches and processes each patch independently through the transformer encoder, maintaining original resolution while reducing computational complexity through localized processing. This segmentation approach allows the model to handle high-resolution images without the need for aggressive down-sampling.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the spatial dimension problem by projecting image patches into a token sequence dimension, where transformer self-attention mechanisms operate. This dimensional transformation allows global context to be captured without losing spatial resolution, as the attention mechanism processes all positions simultaneously rather than through hierarchical down-sampling.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If down-sampling is applied in encoder stages, then global context is captured, but spatial details are lost and hard to recover in decoder

Engineering Contradiction:
Improveglobal context captureVSAvoidspatial detail recovery
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The transformer encoder serves multiple functions simultaneously: it captures global context through self-attention mechanisms while preserving spatial resolution through patch-based processing. The same encoder structure handles both contextual understanding and spatial feature extraction, eliminating the need for separate down-sampling operations that would lose detail.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces patch embedding tokens as intermediaries between the input image and the transformer processing. These tokens maintain a one-to-one correspondence with image patches, serving as a bridge that preserves spatial information while enabling global context capture through transformer attention mechanisms.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If transformer blocks are used instead of convolutions, then global receptive field is achieved, but computational requirements increase

Engineering Contradiction:
Improveglobal receptive fieldVSAvoidcomputational requirements
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

By segmenting the image into patches and processing them as tokens, the patent reduces the computational burden of transformer self-attention from O(N²) where N is the number of pixels, to O(M²) where M is the number of patches. This segmentation dramatically reduces computational requirements while maintaining global receptive field through attention mechanisms.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12380714B2Methods and apparatus to perform dense prediction using transformer blocks
Publication Date: 2025.08.05 INTEL CORP
  • US12380714B2 patent drawing
  • US12380714B2 patent drawing
  • US12380714B2 patent drawing

AI summary

Methods, apparatus, systems and articles of manufacture disclosed herein perform dense prediction of an input image using transformers at an encoder stage and at a reassembly stage of an image processing system. A disclosed apparatus includes an encoder with an embedder to convert an input image to a plurality of tokens representing features extracted from the input image. The tokens are embedded with a learnable position embedding. The encoder also includes one or more transformers configured in a sequence of stages to relate the tokens to each other. The apparatus further includes a decoder that includes one or more of reassemblers to assemble the tokens into feature representations, one or more of fusion blocks to combine the feature representations to generate a final feature representation, and an output head to generate a dense prediction based on the final feature representation and based on an output task.