Text-Guided Image Editing With Feature Fusion and Detail Preservation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current image processing technologies struggle to accurately generate target images based on user input text descriptions due to difficulties in determining the specific elements within initial images, leading to potential loss of image information during editing.

Innovation Solution

A method involving a multimodal model to fuse text and image features, using a feature conversion model to align dimensions, and a diffusion model to generate target images with aligned features, ensuring accurate representation of user-specified visual effects while preserving initial image details.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If image processing techniques are used to generate target images based on text descriptions, then user experience is improved, but image information may be lost during editing

Engineering Contradiction:
Improveuser experienceVSAvoidimage information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent segments the image processing into multiple independent modules: a first processing module that preserves original image information, a second processing module that applies visual effects based on text descriptions, and a fusion module that combines results. This segmentation allows each module to perform its function optimally without interfering with information preservation in other modules.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediate feature representations as mediators between the text description and the image processing. The text description is converted to text features, which then guide the visual effect application without directly modifying the original image pixels. This intermediary approach prevents direct information loss while still achieving the desired visual effects.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If visual effects are applied to initial images based on text input, then image editing accuracy is improved, but image quality may deteriorate

Engineering Contradiction:
Improveimage editing accuracyVSAvoidimage quality
Core Design Contradiction:
Measurement precisionVSManufacturing precision

Solution Approach 1:

The patent performs preliminary processing to extract and preserve key image features before applying visual effects. The first processing module captures essential image information in advance, creating a reference that guides the subsequent effect application. This preliminary action ensures that even as visual effects are applied, the original image quality characteristics are maintained as a foundation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a composite representation by fusing multiple feature types: original image features, text description features, and processed visual effect features. This composite approach combines the precision of text-guided editing with the quality preservation of original image characteristics, achieving both high editing accuracy and maintained image quality.

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS20260065525A1Image processing
Publication Date: 2026.03.05 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20260065525A1 patent drawing
  • US20260065525A1 patent drawing
  • US20260065525A1 patent drawing

AI summary

A method, apparatus, device, and computer-readable storage medium for image processing are provided. The method includes receiving a text input for an initial image, the text input describing a visual effect for the initial image. A fusion feature for the text input and the initial image is generated based on the text input and the initial image. A target image corresponding to the initial image is generated based on a first image feature of the initial image and the fusion feature, the target image having a visual element related to the visual effect. The fusion of text and image can better express the desired visual effect.