Facial Image Object Removal Using Attention Maps

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current facial image beautification techniques require large amounts of paired data for model training, making it difficult to collect and increasing training costs.

Innovation Solution

A method and apparatus for image processing that uses a model trained on unpaired data, specifically generating an attention map of the object to be removed, allowing for the removal of predetermined objects from facial images without the need for paired data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If paired data is used for model training, then the accuracy of object removal is improved, but the data collection difficulty and training cost increase

Engineering Contradiction:
Improveobject removal accuracyVSAvoiddata collection ease
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent introduces an attention map as an intermediary element that guides the object removal process. The attention map highlights the target object's location and characteristics, enabling the model to focus on relevant features without requiring paired training data. This intermediary mechanism bridges the gap between unpaired data and accurate object removal, resolving the contradiction by maintaining high accuracy while eliminating the need for difficult paired data collection

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the training paradigm from paired data to unpaired data, fundamentally altering the data parameter. By using unpaired data with attention maps, the model learns to identify and remove objects based on attention guidance rather than direct pixel-to-pixel mapping. This parameter change enables training with easily collected unpaired data while maintaining removal accuracy through the attention mechanism

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If paired data is used for model training, then the object removal quality is improved, but the training cost increases

Engineering Contradiction:
Improveobject removal qualityVSAvoidtraining data quantity
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The attention map serves as a mediator that enables high-quality object removal without requiring large quantities of paired training data. By providing spatial and feature guidance through the attention map, the model can achieve precise object removal using unpaired data, thus reducing the training data quantity requirement while maintaining removal quality

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent uses the attention map to copy or transfer the object's spatial and feature information to guide the removal process. Instead of requiring numerous paired examples, the attention map captures essential object characteristics that can be applied during inference, reducing the need for extensive paired training data while preserving removal quality

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20240290135A1Method, electornic device, and storage medium for image processing
Publication Date: 2024.08.29 DOUYIN VISION CO LTD
  • US20240290135A1 patent drawing
  • US20240290135A1 patent drawing
  • US20240290135A1 patent drawing

AI summary

Embodiments of the disclosure provide a method, apparatus, electronic device (700), and storage medium for image processing. The method includes: inputting a to-be-processed facial image to a predetermined model (S110); and outputting, by the predetermined model, a target facial image (S120) with a predetermined object removed from the to-be-processed facial image; wherein the predetermined model is trained and generated based on an attention map (a) of the predetermined object. Since the predetermined model is trained based on the attention map (a) of the predetermined object, it is able to first generate the attention map (a) of the predetermined object based on unpaired data training, and then train to remove the predetermined object from the facial image with the attention map (a) of the predetermined object.