Speech Recognition Dereverberation Using Segment-Specific Weighting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition technologies face challenges in accurately processing speeches recorded in reverberant environments due to the overlap of reverberation components from previous speech, leading to lower recognition rates, as they do not adequately consider the varying energy levels between words.
Innovation Solution
A speech processing device and method that includes a speech recognition unit, a reverberation influence storage unit, and a reverberation reduction unit, which selectively removes reverberation components based on the degree of influence specific to each recognition segment or word pair, using dereverberation parameters and power spectral density to improve accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If reverberation components are removed using a uniform weighting coefficient based on reverberation time, then the reverberation reduction process is simple, but speech recognition accuracy cannot be sufficiently improved because differences between words are not considered
Solution Approach 1:
The patent applies local quality by transitioning from a uniform weighting coefficient applied to all speech segments to word-specific or recognition segment-specific weighting coefficients. Each word or recognition segment is assigned its own weighting coefficient based on its individual characteristics (such as power spectral density), allowing differential treatment of different parts of the speech signal to improve recognition accuracy while maintaining reasonable processing complexity
Solution Approach 2:
The patent implements dynamics by making the weighting coefficient adaptive rather than static. The weighting coefficient is dynamically determined based on the power spectral density of each recognition segment and the estimated reverberation characteristics, allowing the system to adapt to varying speech conditions and reverberation levels across different words and time periods
2Object-affected harmful factors
If all energy based on reverberations is removed from currently-observed sound energy, then reverberation influence is reduced, but speech recognition accuracy is not sufficiently improved due to uniform treatment of all words
Solution Approach 1:
The patent addresses this contradiction by applying local quality through word-specific weighting coefficients. Instead of uniformly removing all reverberation energy, the system selectively removes reverberation components from each word or recognition segment based on its specific characteristics. This localized approach ensures that reverberation reduction is tailored to each word's acoustic properties, thereby sufficiently improving speech recognition accuracy
Solution Approach 2:
The patent employs parameter changes by using power spectral density as a key parameter to determine the weighting coefficient for each recognition segment. By changing the parameter from a fixed reverberation-time-based coefficient to a dynamic coefficient derived from actual speech characteristics (power spectral density), the system achieves both effective reverberation reduction and improved recognition accuracy
Data Source
AI summary
A speech processing device includes a speech recognition unit configured to sequentially recognize recognition segments from an input speech, a reverberation influence storage unit configured to store a degree of reverberation influence indicating an influence of a reverberation based on a preceding speech to a subsequent speech subsequent to the preceding speech and a recognition segment group including a plurality of recognition segments in correlation with each other, a reverberation influence selection unit configured to select the degree of reverberation influence corresponding to the recognition segment group which includes the plurality of recognition segments recognized by the speech recognition unit from the reverberation influence storage unit, and a reverberation reduction unit configured to remove a reverberation component weighted with the degree of reverberation influence from the speech from which at least a part of recognition segments of the recognition segment group is recognized.


