Facial Feature Detection for Speech Activity Start and End
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice-controlled applications face challenges in determining the start and end of speech in noisy environments, leading to unintended noise processing and unnatural user interactions, particularly in environments with background noise or multiple speakers.
Innovation Solution
Facial feature detection is used to determine the user's intention to interact with a device by detecting the presence of a face looking at the screen and mouth movement, allowing for hands-free operation and natural command inputs without the need for trigger words or button presses.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual button press is used to control speech start and stop, then speech processing accuracy is improved, but ease of operation deteriorates
Solution Approach 1:
The patent replaces the mechanical button-pressing system with an automated facial feature detection system. The camera captures facial images, and software algorithms detect mouth movement patterns to automatically determine speech start and end points, eliminating the need for manual mechanical input while maintaining accurate speech processing.
Solution Approach 2:
The system enables self-service operation by automatically detecting when the user is speaking through facial feature analysis. The device monitors its own operational state using the camera and algorithms, eliminating the need for external manual control while maintaining precise speech processing.
2Measurement precision
If trigger words are used to signal speech commands, then speech start detection is improved, but naturalness of interaction deteriorates
Solution Approach 1:
The patent replaces the linguistic trigger-word system with a visual facial detection system. Instead of requiring users to speak specific trigger words, the camera-based system directly detects mouth movement patterns that indicate speech initiation, enabling more natural interactions without artificial linguistic constraints.
3Device complexity
If audio-only processing is used, then device complexity is reduced, but reliability deteriorates in noisy environments
Solution Approach 1:
The patent merges audio processing with visual facial feature detection. The system combines data from both the microphone (audio signals) and camera (facial images), using facial mouth movement detection to validate and enhance speech detection reliability, particularly in noisy environments where audio alone may be insufficient.
Data Source
AI summary
Disclosed herein are systems, methods, and non-transitory computer-readable storage media for processing audio. A system configured to practice the method monitors, via a processor of a computing device, an image feed of a user interacting with the computing device and identifies an audio start event in the image feed based on face detection of the user looking at the computing device or a specific region of the computing device. The image feed can be a video stream. The audio start event can be based on a head size, orientation or distance from the computing device, eye position or direction, device orientation, mouth movement, and/or other user features. Then the system initiates processing of a received audio signal based on the audio start event. The system can also identify an audio end event in the image feed and end processing of the received audio signal based on the end event.


