Video-Guided Voice Recognition for Hands-Free Speech Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice recognition technologies require separate user gestures to initiate and end voice recording, which can be inconvenient in situations like driving or carrying burdens, and struggle with accurately distinguishing speech from noise in various environments.
Innovation Solution
A method and apparatus that utilize video recognition to detect speech start and end times without requiring user gestures, using a combination of video and audio data analysis to convert the device into a voice recognition mode and generate relevant audio or video data for accurate voice command recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If voice recognition requires separate user gestures to initiate and end recording, then the system can accurately control recording timing, but the ease of operation deteriorates in situations like driving or carrying burdens
Solution Approach 1:
The system uses the user's own video image to automatically detect speech start and end times through mouth movement analysis, eliminating the need for separate gesture operations. The device serves itself by using its camera to capture the user's face and automatically triggering voice recognition based on detected mouth movements, making the system both easy to operate and reliable.
2Measurement precision
If the system uses only audio data for voice recognition, then the device complexity is reduced, but the measurement precision of speech start and end times deteriorates in noisy environments
Solution Approach 1:
The system merges video data from the camera with audio data from the microphone to detect speech start and end times. By combining these two data sources, the system achieves higher measurement precision in determining when speech begins and ends, while the video component helps distinguish actual speech from background noise in challenging acoustic environments.
3Loss of information
If the system records all audio data continuously, then no speech data is missed, but the loss of information increases due to including noise and silent sections
Solution Approach 1:
The system performs preliminary analysis of video data to detect mouth movements before processing audio data. By预先 detecting when the user's mouth moves (indicating speech), the system can selectively record only the relevant audio segments, avoiding the inclusion of noise and silent sections while ensuring no actual speech is missed.
Data Source
AI summary
Provided are a method and an apparatus for performing exact start and end recognition of voice based on video recognition. The method includes determining whether a speech starts based on at least one of first video and audio data before conversion into a voice recognition mode, converting into the voice recognition mode and generating second audio data including a voice command, when it is determined that speech starts, and determining whether the speech is terminated based on at least one of second video and audio data after conversion into the voice recognition mode.


