Generative Model Facial Image Enhancement via Multi-Frame Loss Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Facial recognition in videos is hindered by external factors like noise and light, resulting in poor-quality and misidentified facial images, which complicates subsequent processes.

Innovation Solution

A method and apparatus that utilize a pre-trained generative model to enhance facial image quality by acquiring multiple frames from a video, updating model parameters based on a loss function calculated from the probability and similarity between generated and standard facial images, using machine learning techniques and models like Long-Short Term Memory Networks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If multiple frames are processed to generate a single facial image, then the quality and authenticity of the generated image is improved, but the processing time and computational complexity increases

Engineering Contradiction:
Improveimage qualityVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-training the generative model offline using large datasets before actual video processing. The pre-trained model contains learned facial features and patterns that enable rapid generation of high-quality images during video processing, reducing real-time computational burden while maintaining image quality

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates multiple copies of facial information from different video frames and processes them through the generative model to produce a single enhanced facial image. By copying and comparing facial features across multiple frames, the system reconstructs a high-quality facial image that captures the essential characteristics of the subject

Inventive Principle:
Principle #26Copying

2Reliability

If a pre-trained generative model is used to generate facial images, then the authenticity of generated images is improved, but the device complexity increases

Engineering Contradiction:
Improveimage authenticityVSAvoidmodel complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system introduces a discriminative model as an intermediary component that works in conjunction with the generative model. The discriminative model evaluates the authenticity of generated images by comparing them against real facial images, providing feedback that guides the generative model to produce more authentic results. This intermediary mechanism enhances reliability while distributing computational complexity across two specialized models rather than one complex model

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If facial images with poor quality are used for recognition, then the processing speed is maintained, but the recognition accuracy deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidrecognition accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system copies facial information from multiple video frames and combines them to create a single high-quality facial image for recognition. By aggregating information from multiple sources (frames), the system reconstructs a clearer, more accurate representation of the subject's face that maintains or improves recognition accuracy while still enabling efficient processing through the pre-trained model

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11978245B2Method and apparatus for generating image
Publication Date: 2024.05.07 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US11978245B2 patent drawing
  • US11978245B2 patent drawing
  • US11978245B2 patent drawing

AI summary

The present disclosure discloses a method and apparatus for generating an image. A specific embodiment of the method comprises: acquiring at least two frames of facial images extracted from a target video; and inputting the at least two frames of facial images into a pre-trained generative model to generate a single facial image. The generative model updates a model parameter using a loss function in a training process, and the loss function is determined based on a probability of the single facial generative image being a real facial image and a similarity between the single facial generative image and a standard facial image. According to this embodiment, authenticity of the single facial image generated by the generative model may be enhanced, and then a quality of a facial image obtained based on the video is improved.