Speech Processing Device Dereverberation via Adaptive Model Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition technologies face challenges in reverberation environments due to the computational load and processing delay associated with long reverberation times, and the inability to accurately select acoustic models that account for the positional relationship between sound sources and sound collection units, leading to decreased speech recognition accuracy.
Innovation Solution
A speech processing device that includes a reverberation characteristic selection unit to correlate correction data with adaptive acoustic models trained for specific reverberation characteristics, and a dereverberation unit to remove reverberation components based on selected correction data, considering the distance between the sound source and sound collection unit, thereby improving speech recognition accuracy without measuring reverberation characteristics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If inverse filter processing is used to estimate impulse response and reconstruct sound source signal, then speech recognition accuracy is improved, but computational load excessively increases and processing delay becomes remarkable
Solution Approach 1:
The patent pre-calculates and stores impulse response data for multiple predetermined positions before actual speech recognition. During processing, the system simply retrieves the pre-computed impulse response corresponding to the estimated speaker position, avoiding the need to perform complex inverse filter processing in real-time. This preliminary computation approach significantly reduces computational load and processing delay while maintaining speech recognition accuracy.
2Productivity
If acoustic models are selected based on reverberation time, then speech recognition is performed, but the positional relationship between sound source and sound collection unit is not considered, leading to decreased speech recognition accuracy
Solution Approach 1:
The patent creates separate impulse response data for multiple predetermined positions in the reverberation space, with each position having its own characteristic impulse response. The system then determines the speaker's position and selects the corresponding locally-appropriate impulse response data, rather than using a single global acoustic model. This local quality approach accounts for the positional relationship between sound source and sound collection unit, improving speech recognition accuracy.
3Device complexity
If reverberation time is used as the sole characteristic for acoustic model selection, then processing is simplified, but the ratio of reverberation component intensity to direct sound intensity varies depending on distance, causing inaccurate model selection
Solution Approach 1:
The patent transitions from selecting acoustic models based on a single parameter (reverberation time) to selecting based on spatial position coordinates. By introducing the spatial dimension and creating impulse response data for multiple predetermined positions, the system can accurately represent the varying relationship between direct sound and reverberation components at different distances. This dimensional change enables more accurate acoustic model selection while maintaining reasonable processing complexity.
Data Source
AI summary
A speech processing device includes a reverberation characteristic selection unit configured to correlate correction data indicating a contribution of a reverberation component based on a corresponding reverberation characteristic with an adaptive acoustic model which is trained using reverbed speech to which a reverberation based on the corresponding reverberation characteristic is added for each of reverberation characteristics, to calculate likelihoods based on the adaptive acoustic models for a recorded speech, and to select correction data corresponding to the adaptive acoustic model having the calculated highest likelihood, and a dereverberation unit configured to remove the reverberation component from the speech based on the correction data.


