Composite Image Parsing for Accurate Layer Element Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for decomposing composite images to extract image elements are inaccurate due to complex superposition and interdependence relationships between layers, leading to poor extraction results.
Innovation Solution
An image processing method that utilizes an image parsing model to obtain structured data, including position and appearance information of image elements, by segmenting and reconstructing composite images using a combination of visual encoders, a large language model, and vector quantization techniques to accurately extract and edit image elements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing decomposition methods are used to extract image elements from composite images, then the extraction process can be performed, but the extraction accuracy is poor due to complex superposition and interdependence relationships between layers
Solution Approach 1:
The patent applies segmentation by dividing the composite image into multiple independent layers through layer separation technology. Each layer contains specific image elements (background layer, main image layer, decorative graphic layer, text layer) that can be independently processed. This segmentation resolves the complexity of handling superposed elements by treating each layer separately, thereby improving extraction accuracy without requiring overly complex decomposition algorithms.
2Ease of operation
If layer separation technology is applied to understand the layered structure of composite images, then image editing and material archiving can be improved, but the processing complexity increases
Solution Approach 1:
The patent introduces structured data as an intermediary representation between the composite image and the editing operations. The image parsing model converts the complex layered image structure into organized structured data that captures positional and appearance information. This intermediary representation simplifies subsequent editing operations while managing the complexity of layer separation through automated parsing rather than manual processing.
Data Source
AI summary
The embodiments of the present disclosure provide an image processing method, an apparatus, an electronic device and a storage medium by obtaining an image to be processed, wherein the image to be processed comprises at least two image elements; obtaining structured data corresponding to the image elements by calling an image parsing model to process the image to be processed, wherein the structured data comprises position information and appearance information of the image elements, the position information represents position features of the image elements in the image to be processed, and the appearance information represents appearance features of the image elements in the image to be processed; and generating an output image containing at least one image element in response to an editing instruction for the structured data.


