Image Stylization Using Pose and Facial Attribute Guidance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Stylized images generated by existing methods, such as Pix2pix and CycleGAN, suffer from single image style, unrealistic local distortion, and poor generalization, while diffusion models struggle with low relevance to original images and require manual input of text prompts.
Innovation Solution
An image processing method that recognizes and analyzes original images to determine theme elements and facial attributes, performs pose estimation, adds noise, and generates images based on text prompt information, facial attributes, and pose information to improve relevance and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If diffusion models are used for image generation, then image quality and style diversity are improved, but relevance to original images deteriorates
Solution Approach 1:
The patent employs feedback mechanisms by extracting facial attribute information and pose information from the original image, then using these extracted features as guidance during the diffusion generation process. This feedback loop ensures that the generated image maintains structural and semantic consistency with the original while achieving style transformation, thus resolving the contradiction between style diversity and relevance.
Solution Approach 2:
The patent changes key parameters by extracting and preserving specific attributes (facial attributes, pose information) from the original image while allowing other parameters (style, color, texture) to vary through the diffusion process. This selective parameter preservation enables the generated image to maintain relevance in terms of identity and structure while achieving style diversity.
2Measurement precision
If manual text prompt input is required for image generation, then control over generated image content is improved, but operation complexity and time consumption increase
Solution Approach 1:
The patent implements self-service by automatically extracting theme elements, facial attribute information, and pose information from the input image without requiring manual text prompt input. The system serves itself by generating appropriate control parameters (text prompts, attribute constraints) based on the image content, thereby maintaining control precision while dramatically simplifying operation.
Solution Approach 2:
The patent performs preliminary actions by pre-extracting and analyzing image features (theme elements, facial attributes, pose information) before the generation process. This preliminary analysis prepares control parameters in advance, eliminating the need for manual prompt input during operation while ensuring precise control over the generated image content.
3Productivity
If stylization is performed using traditional methods, then processing speed is improved, but image quality and realism deteriorate
Solution Approach 1:
The patent applies partial action by using a simplified diffusion process that focuses only on style transformation while preserving the original image's structural content. Instead of performing full image generation, the method applies diffusion selectively to style parameters, maintaining processing speed while improving image quality and realism in the stylized output.
Data Source
AI summary
An image processing method, an image processing device, an electronic device, and a computer-readable storage medium are provided. The image processing method includes: recognizing and analyzing an original image to obtain theme elements of the original image and facial attribute information of a target object in the original image, and determining text prompt information based on a predetermined target style, the theme elements and the facial attribute information; performing pose estimation on the original image to obtain pose information of the target object in the original image; performing noise adding processing on the original image to obtain a target noise image; generating image noise based on the target noise image, the text prompt information, the facial attribute information, and the pose information; and generating a target image of the predetermined target style based on the image noise and the target noise image.


