Electronic Device Lip Reading Voice Recognition Noise
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Intelligent agent services face challenges in accurately recognizing user voice commands due to noise interference, which affects the performance of voice recognition systems.
Innovation Solution
The electronic device employs lip reading technology to improve the accuracy of intelligent agent services by analyzing image information and combining it with voice recognition, allowing users to correct unclear voice inputs based on lip shape.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If voice recognition is used to recognize user speech, then the system can process user commands, but recognition accuracy deteriorates in noisy environments
Solution Approach 1:
The patent combines voice recognition with lip reading technology to recognize user speech. The processor integrates results from both voice input and lip movement analysis to determine user commands, thereby maintaining high recognition accuracy in noisy environments where voice-only systems fail.
Solution Approach 2:
Lip reading serves as an intermediary method to complement voice recognition. When voice recognition accuracy is compromised by noise, the system uses lip movement analysis from image data as an alternative pathway to recognize user speech, mediating the harmful effect of noise interference.
2Reliability
If lip reading technology is added to improve recognition accuracy, then speech recognition reliability improves, but device complexity increases
Solution Approach 1:
The camera module, originally designed for general imaging purposes, is utilized for lip reading analysis. This multi-functional use of existing hardware avoids adding dedicated complex devices, as the same camera serves both regular imaging and lip movement detection functions.
Solution Approach 2:
The system merges voice recognition and lip reading processing into a unified speech recognition framework. The processor integrates both modalities within a single system architecture, managing multiple recognition pathways without requiring entirely separate complex systems.
3Reliability
If image information is acquired for lip reading, then speech recognition accuracy improves, but energy consumption increases
Solution Approach 1:
The system dynamically adjusts when to activate lip reading based on noise conditions. The processor activates image acquisition and lip reading processing selectively when voice recognition is likely to be compromised by noise, rather than continuously, thereby reducing overall energy consumption while maintaining accuracy when needed.
Solution Approach 2:
The system applies lip reading partially rather than continuously. By using lip reading only in specific conditions (noisy environments) rather than always, the system achieves the necessary speech recognition accuracy improvement without the excessive energy cost of constant image processing.
Data Source
Figure 1
Figure 2A
Figure 2B
AI summary
An electronic device includes: a camera; a microphone; a display; a memory; and a processor configured to receive an input for activating an intelligent agent service from a user while at least one application is executed, identify context information of the electronic device, control to acquire image information of the user through the camera, based on the identified context information, detect movement of a user's lips included in the acquired image information to recognize a speech of the user, and perform a function corresponding to the recognized speech.