Face Super-Resolution via Cascaded Encoder and Deformable Attention
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current face super-resolution (FSR) technologies require complex network structures, increasing memory and computation costs, and often need additional face priori annotations, which can be cumbersome and costly.
Innovation Solution
A method and electronic device using a cascaded encoder structure with cross-attention and transformable attention models to aggregate multi-level image features, generating super-resolution images and key point coordinates without requiring additional facial prior annotations, by integrating face priori information through a deformable attention mechanism.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If complex network structures are used for face super-resolution, then FSR performance is improved, but memory and computation costs increase
Solution Approach 1:
The network is divided into multiple encoder modules that process different levels of feature maps from the multi-level feature pyramid. Each encoder module independently processes specific feature levels and aggregates results, breaking down the complex super-resolution task into manageable segments that reduce overall computational burden while maintaining performance.
Solution Approach 2:
The patent introduces a multi-level feature pyramid dimension, extracting feature maps at multiple scales and levels. This dimensional expansion allows the network to capture both fine-grained and coarse-grained facial features simultaneously, improving FSR performance without requiring a single overly complex network structure.
2Manufacturing precision
If additional face priori annotations are required, then FSR performance is improved, but annotation costs and complexity increase
Solution Approach 1:
The network generates its own face priori information internally through the multi-level feature pyramid extraction and encoder-based feature aggregation. Instead of relying on external annotated face priors, the system self-generates the necessary structural information from the input low-resolution image, eliminating the need for costly additional annotations while maintaining FSR performance.
3Manufacturing precision
If complex network structures are used, then super-resolution quality is improved, but training and operation costs increase
Solution Approach 1:
By segmenting the processing task across multiple encoder modules that operate on different feature levels, the patent reduces the computational complexity of each individual processing step. This segmentation allows for more efficient training and operation with lower energy costs while maintaining high super-resolution quality through the aggregation of features from multiple levels.
Solution Approach 2:
The network processes only the essential feature levels and regions that contribute most to super-resolution quality, rather than uniformly processing all possible features. This selective partial processing reduces overall computational energy requirements while maintaining the quality necessary for effective face super-resolution.
Data Source
AI summary
An electronic device and a processor-implemented method with image processing are provided. The processor-implemented method comprises generating an initial image feature matrix of a face image based on a multi-level feature map of the face image; generating an initial face priori feature matrix of the face image based on a final level feature map of the multi-level feature map; and generating a super-resolution image of the face image and/or key point coordinates of the face image by using one or more encoders, based on the initial image feature matrix and the initial face priori feature matrix, wherein, when the one or more encoders are plural encoders, they are cascaded.


