Lip Motion Enhancement for Hearing-Impaired Communication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional technologies for facilitating calls with hearing-impaired individuals fail to effectively convey the nuances of utterance content and lip motion, leading to inadequate recognition of intended messages.
Innovation Solution
A communication device with a display control system that captures and enhances lip motion data, comparing recognition rates from voice and lip motion recognition to generate and display enhanced lip motion images, ensuring better content recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If voice recognition is used to convey utterance content, then communication efficiency is improved, but nuance and emotional tone are lost
Solution Approach 1:
The system segments the communication information into multiple components: voice recognition results for efficiency, lip motion data for nuance, and synthesized speech for emotional tone. Each component handles a different aspect of the utterance, allowing simultaneous achievement of communication efficiency and information preservation.
Solution Approach 2:
The system creates a composite communication output by combining three different information sources: text from voice recognition, visual lip motion patterns, and synthesized speech with emotional tone. This composite approach preserves both efficiency and nuance that would be lost in single-mode communication.
2Loss of information
If lip motion is displayed as-is, then visual information is preserved, but recognition accuracy is insufficient when lip motion is small
Solution Approach 1:
The system dynamically adjusts lip motion display based on recognition confidence levels. When lip motion is small and recognition is uncertain, the system enhances the visual display of lip patterns to improve detectability without distorting the underlying visual information when motion is sufficient.
Solution Approach 2:
The system changes display parameters of lip motion visualization based on detected motion magnitude and recognition accuracy. When lip motion is small, the system adjusts contrast, amplification, or highlighting parameters to enhance visibility and recognition accuracy while preserving authentic visual information.
3Quantity of substance
If voice data alone is transmitted, then communication bandwidth is reduced, but hearing-impaired individuals cannot understand the message
Solution Approach 1:
The system creates a universal communication solution that serves both hearing-impaired and hearing-normal individuals. By incorporating visual lip motion patterns alongside voice data, the system ensures message understanding for hearing-impaired users while maintaining compatibility and utility for hearing-normal users, achieving multi-functionality.
Solution Approach 2:
The system introduces visual lip motion patterns as an intermediary channel that bridges the gap for hearing-impaired individuals. This intermediate visual representation of speech patterns allows message understanding without requiring direct audio perception, while serving as a supplement for all users.
Data Source
AI summary
An disclosure includes: moving image acquisition unit configured to acquire moving image data obtained through moving image capturing of at least a mouth part of an utterer; a lip detection unit configured to detect a lip part from the moving image data and detect motion of the lip part; a moving image processing unit configured to generate a moving image enhanced to increase the motion of the lip part detected by the lip detection unit; and a display control unit configured to control a display panel to display the moving image generated by the moving image processing unit.


