Face Video Decoder Using Generative Synthesis for Image Quality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding technologies face challenges in improving coding efficiency, enhancing image quality, reducing processing amounts, and circuit scales, and appropriately selecting elements or operations such as filters, blocks, motion vectors, and reference pictures.
Innovation Solution
A decoder configuration that includes memory and circuitry to decode fundamental images, geometric information, and background information, using a generative model to synthesize face videos, reducing background distortion and enhancing image quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If conventional video coding technologies are used, then coding efficiency can be maintained at current levels, but image quality enhancement and processing amount reduction are limited
Solution Approach 1:
The patent segments the video processing into distinct components: fundamental image extraction, geometric information processing, and background information handling. This segmentation allows each component to be optimized independently, improving overall image quality while maintaining coding efficiency through specialized processing pipelines for each segment.
Solution Approach 2:
The patent employs parameter changes by utilizing generative models that transform input parameters (fundamental image, geometric information, background information) into enhanced output images. The generative model adjusts various image parameters such as resolution, quality metrics, and visual characteristics to achieve superior image quality without proportionally increasing processing amounts.
2Manufacturing precision
If advanced video coding schemes are implemented, then image quality and coding efficiency can be improved, but processing amounts and circuit scales increase
Solution Approach 1:
The patent applies preliminary action by pre-processing and organizing video data into structured components (fundamental images, geometric information, background information) before main processing. This preliminary organization reduces the complexity of subsequent processing steps, allowing advanced image quality enhancement with reduced overall processing amounts.
Solution Approach 2:
The patent uses copying mechanisms where fundamental images and background information are reused across multiple frames and processing operations. Instead of processing complete video data repeatedly, the system copies and applies base components across different temporal instances, significantly reducing processing amounts while maintaining high image quality through the generative model.
3Manufacturing precision
If generative models are used to synthesize face videos, then background distortion is reduced and image quality is enhanced, but device complexity increases
Solution Approach 1:
The patent implements universality by designing a multi-functional processing system where the generative model serves multiple purposes: background distortion reduction, image quality enhancement, and synthesized video generation. This single multi-functional component replaces what would otherwise require multiple separate processing circuits, thereby managing device complexity while achieving superior image quality.
Solution Approach 2:
The generative model acts as an intermediary between the input data (fundamental image, geometric information, background information) and the final synthesized video output. This intermediary layer processes and reconciles multiple input streams, reducing background distortion and enhancing image quality without requiring direct complex interactions between all input components, thus managing circuit scale effectively.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A decoder (200) includes memory (252) and circuitry (251) coupled to the memory (252). In operation, the circuitry (251): decodes, from one or more streams, (i) a fundamental image that is an image including a face, (ii) geometric information indicating geometric attributes of a subject and corresponding to each of frames of a captured video by a camera, and (iii) background information regarding a background image (S401); and generates a synthesized face video using a generative model from the fundamental image, the geometric information, and the background information (S402). The synthesized face video is a video including the face and synthesized with the background image.