Depth-Camera Speech Recognition for Noisy In-Vehicle Voice Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In-vehicle speech recognition systems face challenges due to environmental noise changes, affecting accuracy and potentially causing navigation errors and increased driving risk.
Innovation Solution
A method that combines audio and video features by using a depth camera to capture facial depth images, extracting mouth shape features, and fusing them with voice features to improve recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If voice interaction is used in in-vehicle environment, then convenience and safety are improved, but speech recognition accuracy deteriorates due to environmental noise
Solution Approach 1:
The patent combines audio features extracted from voice signals with visual features extracted from facial depth images to form fused features for speech recognition. This multi-modal fusion approach leverages both auditory and visual information to overcome the limitations of single-modality recognition in noisy in-vehicle environments, thereby maintaining high recognition accuracy while preserving the convenience of voice interaction.
2Reliability
If voice interaction is used in in-vehicle environment, then safety is improved by freeing hands, but speech recognition accuracy deteriorates due to environmental noise
Solution Approach 1:
The system fuses audio and visual modalities to enhance speech recognition reliability. By combining voice signals with facial depth image features, the system creates a more robust recognition mechanism that is less susceptible to environmental noise interference, thereby maintaining high safety standards while enabling hands-free operation.
Solution Approach 2:
The patent introduces facial depth images as an intermediary visual modality to supplement voice recognition. The depth information from facial images serves as an additional channel that helps disambiguate voice commands in noisy environments, acting as a mediator that enhances the reliability of speech recognition without requiring physical contact or visual line-of-sight to the driver.
3Device complexity
If audio features alone are used for speech recognition, then system complexity is reduced, but recognition accuracy deteriorates in noisy environments
Solution Approach 1:
The patent integrates audio feature extraction with visual feature extraction from facial depth images. The system processes both modalities simultaneously and fuses their features to achieve superior recognition accuracy in noisy in-vehicle environments, demonstrating that the increased complexity is justified by the significant improvement in recognition reliability.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Embodiments of this application provide an electronic device, a speech recognition method therefor, and a medium, and relate to a speech recognition technology in the field of artificial intelligence (Artificial Intelligence, AI). The speech recognition method in this application includes: obtaining a facial depth image and a to-be-recognized voice of a user, where the facial depth image is an image collected by using a depth camera; recognizing a mouth shape feature from the facial depth image, and recognizing a voice feature from a to-be-recognized audio; and fusing the voice feature and the mouth shape feature into an audio-video feature, and recognizing, based on the audio-video feature, a voice uttered by the user. According to the method, because the mouth shape feature extracted from the facial depth image is not affected by light of an environment, the mouth shape feature can more accurately reflect a mouth shape change obtained when the user utters the voice. The mouth shape feature extracted from the facial depth image and the voice feature are fused, so that speech recognition accuracy can be improved.