Videophone Image Processing for Authentic Communication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional cellular videophones lack the ability to modify or substitute undesirable facial images during video calls, leading to a loss of enhanced communication through facial expressions and lip movements.
Innovation Solution
A videophone image processing system that allows users to select and transmit a preferred or avatar image, incorporating their actual facial features and background, using object-based video segmentation and processing to maintain lip movements and expressions, while replacing the original image.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a user's actual facial image is transmitted during video calls, then communication authenticity is improved, but undesirable appearance aspects reduce communication effectiveness
Solution Approach 1:
The system creates a virtual copy (avatar) of the user's face that replicates their facial features, expressions, and lip movements. This avatar serves as a substitute for the actual facial image, maintaining the authenticity of communication while eliminating undesirable appearance aspects. The avatar is generated by capturing facial geometry and texture data, then rendering it in real-time with controlled lighting and background.
Solution Approach 2:
The avatar acts as an intermediary between the user's actual appearance and the communication partner. Instead of transmitting the raw facial image directly, the system processes it through the avatar intermediary, which filters out undesirable aspects while preserving essential communication elements like expressions and lip movements.
2Object-affected harmful factors
If an avatar image is used to replace the actual facial image, then appearance quality is improved, but communication authenticity may be reduced
Solution Approach 1:
The system applies different quality levels to different parts of the facial representation. Critical communication elements such as facial expressions, eye movements, and lip synchronisation are rendered with high fidelity to maintain authenticity, while less critical aspects can be stylized or simplified. This selective application of quality ensures both appearance enhancement and communication authenticity.
Solution Approach 2:
The avatar is designed to be dynamically responsive, capturing and reproducing real-time facial expressions and lip movements. This dynamic behavior maintains communication authenticity by ensuring the avatar reflects the user's actual emotional state and speech patterns, rather than being a static or pre-recorded image.
3Productivity
If real-time facial video processing is performed, then communication effectiveness is improved, but processing complexity increases
Solution Approach 1:
The facial processing system is divided into separate functional modules: facial feature detection, geometry extraction, texture mapping, avatar rendering, and lip-synchronisation. This segmentation allows each module to be optimized independently and processed in parallel, reducing overall processing complexity while maintaining real-time performance.
Solution Approach 2:
The system performs preliminary processing by capturing and storing facial geometry and texture data in advance, creating a base model that can be quickly rendered in real-time. This pre-processing reduces the computational burden during actual video calls, as the system only needs to animate and render the pre-prepared model rather than processing raw facial data from scratch.
Data Source
AI summary
Herein described is a system and method for modifying facial video transmitted from a first videophone to a second videophone during a videophone conversation. A videophone comprises a videophone image processing system (VIPS) that stores one or more preferred images. The one or more preferred images may comprise an image of a person presented in an attractive appearance. The one or more preferred images may comprise one or more avatars. Additionally, the VIPS may be used to incorporate one or more facial features of the person into a preferred image or avatar. Furthermore, a replacement background may be incorporated into the preferred image or avatar. The VIPS transmits a preferred image of a first speaker of a first videophone to a second speaker of a second videophone by capturing an actual image of the first speaker and substituting at least a portion of said actual image with a stored image.


