Vehicle Audio Signal Processing for Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition systems in vehicles face challenges in accurately distinguishing intended voice commands from various background audio sources, leading to false positives and reduced recognition rates due to noise and echo interference.
Innovation Solution
A system and method that process audio signals captured from microphones in vehicles by identifying and reducing unwanted audio sources, using reference signals from mixed audio outputs to improve speech recognition, with separate processing for human perception and machine recognition, employing noise reduction and echo cancellation techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If audio is captured continuously from the microphone to enable speech recognition, then the system can respond to voice commands, but false positive recognition results occur due to background noise and audio sources
Solution Approach 1:
The patent extracts and removes known audio sources (navigation prompts, music, chimes, text-to-speech output) from the microphone signal before processing. By identifying these specific audio sources and selectively removing them, the system prevents false positives while maintaining continuous listening capability for genuine voice commands.
Solution Approach 2:
The patent introduces an intermediary processing stage between audio capture and speech recognition. This intermediate processor analyzes the microphone signal, identifies known audio sources, and removes them before the signal reaches the speech recognition system, thereby improving recognition accuracy without sacrificing responsiveness.
2Reliability
If echo cancellation and noise reduction are applied to process the audio signal, then recognition rates improve, but the processing complexity increases
Solution Approach 1:
The patent applies preliminary action by removing known audio sources from the microphone signal before the speech recognition processing occurs. This pre-processing step simplifies the subsequent recognition task by eliminating predictable interference sources, thereby improving recognition rates without requiring overly complex real-time processing during the recognition phase.
3Adaptability or versatility
If multiple audio sources are present in the vehicle environment, then the audio system provides rich functionality, but the speech recognition system experiences false positives and reduced accuracy
Solution Approach 1:
The patent selectively extracts and removes specific known audio sources (navigation prompts, music, chimes, text-to-speech) from the mixed microphone signal. This extraction process preserves the functionality of these audio sources while eliminating their interference with speech recognition, allowing the system to maintain rich audio functionality without sacrificing detection accuracy.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A system and method for processing an audio signal captured from a microphone may reproduce a known audio signal with an audio transducer into an acoustic space. The known audio signal may include content from one or more audio sources. A microphone audio signal may be captured from the acoustic space where the microphone audio signal comprises the known audio signal and one or more unknown audio signals. Processing control information may be accessed. The known audio signal may be reduced in the microphone audio signal responsive to the processing control information where the processing control information indicates one or more characteristics of a downstream audio processor that processes the microphone audio signal.