Composite Image Parsing for Accurate Layer Element Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for decomposing composite images to extract image elements are inaccurate due to complex superposition and interdependence relationships between layers, leading to poor extraction results.

Innovation Solution

An image processing method that utilizes an image parsing model to obtain structured data, including position and appearance information of image elements, by segmenting and reconstructing composite images using a combination of visual encoders, a large language model, and vector quantization techniques to accurately extract and edit image elements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing decomposition methods are used to extract image elements from composite images, then the extraction process can be performed, but the extraction accuracy is poor due to complex superposition and interdependence relationships between layers

Engineering Contradiction:
Improveextraction accuracyVSAvoidcomplexity of decomposition method
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the composite image into multiple independent layers through layer separation technology. Each layer contains specific image elements (background layer, main image layer, decorative graphic layer, text layer) that can be independently processed. This segmentation resolves the complexity of handling superposed elements by treating each layer separately, thereby improving extraction accuracy without requiring overly complex decomposition algorithms.

Inventive Principle:
Principle #1Segmentation

2Ease of operation

If layer separation technology is applied to understand the layered structure of composite images, then image editing and material archiving can be improved, but the processing complexity increases

Engineering Contradiction:
Improveimage editing capabilityVSAvoidprocessing system complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent introduces structured data as an intermediary representation between the composite image and the editing operations. The image parsing model converts the complex layered image structure into organized structured data that captures positional and appearance information. This intermediary representation simplifies subsequent editing operations while managing the complexity of layer separation through automated parsing rather than manual processing.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20260105663A1Image processing method, apparatus, electronic device and storage medium
Publication Date: 2026.04.16 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20260105663A1 patent drawing
  • US20260105663A1 patent drawing
  • US20260105663A1 patent drawing

AI summary

The embodiments of the present disclosure provide an image processing method, an apparatus, an electronic device and a storage medium by obtaining an image to be processed, wherein the image to be processed comprises at least two image elements; obtaining structured data corresponding to the image elements by calling an image parsing model to process the image to be processed, wherein the structured data comprises position information and appearance information of the image elements, the position information represents position features of the image elements in the image to be processed, and the appearance information represents appearance features of the image elements in the image to be processed; and generating an output image containing at least one image element in response to an editing instruction for the structured data.