Generative Image Editing With Selective Feature Trimming
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional generative neural networks for digital image generation and editing require significant computing resources and processing time, especially for high-resolution images and iterative editing processes, due to generating whole images instead of targeted image portions.
Innovation Solution
A local refinement generative system that utilizes a transformer-based encoder neural network to encode global context information into tokens, trims the latent feature vector to relevant image subsets, and uses a generative decoder neural network to generate and blend digital image data only for masked portions, reducing resource usage and processing time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If generative neural networks generate whole images for digital image editing, then comprehensive image content is produced, but computing resources and processing time increase significantly
Solution Approach 1:
The patent divides the image editing task into two segments: (1) generating a coarse whole image at low resolution to establish overall structure and context, and (2) refining only specific local regions at high resolution to achieve detailed accuracy. This segmentation allows the system to maintain comprehensive image content while significantly reducing computing resources by avoiding full high-resolution generation.
Solution Approach 2:
The patent applies partial action by generating only the necessary portions of the image at high resolution. Instead of processing the entire image at full detail, the system performs excessive action on coarse generation (ensuring complete coverage) and partial action on refinement (only where needed), optimizing the balance between completeness and efficiency.
2Manufacturing precision
If generative neural networks process high-resolution images, then image quality improves, but processing time increases due to iterative editing processes
Solution Approach 1:
The patent segments the resolution processing into two stages: coarse-resolution whole-image generation followed by fine-resolution local refinement. This allows high image quality to be achieved in refined regions without processing the entire high-resolution image iteratively, thereby reducing processing time while maintaining quality where it matters.
Solution Approach 2:
The patent performs preliminary coarse-generation of the entire image at low resolution before conducting detailed refinement. This preliminary action establishes the overall image structure and context, so that subsequent high-resolution processing can focus only on specific regions, reducing the total time required for high-quality output.
3Manufacturing precision
If generative neural networks generate digital image content for entire images, then complete image coverage is achieved, but resource consumption increases
Solution Approach 1:
The patent segments resource allocation by dedicating minimal resources to whole-image coarse generation and concentrating computational resources only on local refinement regions. This ensures complete image content coverage through the coarse pass while achieving high-quality detailed content only where required, significantly reducing overall computing resource consumption.
Solution Approach 2:
The patent applies excessive action in the coarse generation phase (processing entire image at low resolution to ensure complete coverage) and partial action in the refinement phase (processing only selected regions at high resolution). This strategy achieves complete image content while optimizing resource usage by avoiding unnecessary high-resolution processing in all regions.
Data Source
AI summary
Methods, systems, and non-transitory computer readable storage media are disclosed for modifying digital images via a generative neural network with local refinement. The disclosed system generates, utilizing an encoder neural network, a latent feature vector of a digital image by encoding global context information of the digital image into the latent feature vector. The disclosed system also determines a modified latent feature vector by trimming the latent feature vector to a feature subset corresponding to a masked portion of the digital image. Additionally, the disclosed system generates, utilizing a generative decoder neural network on the modified latent feature vector, digital image data corresponding to the masked portion of the digital image. The disclosed system also generates a modified digital image including the digital image data corresponding to the masked portion combined with additional portions of the digital image.


