Speech Processing Device with Position-Based Feature Adjustment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Speech recognition rates are lowered due to changes in sound source position under reverberation conditions, as the positional relationship between the sound source and recording point affects the degree of reverberation, leading to decreased accuracy in speech recognition processing.
Innovation Solution
A speech processing device that includes a sound source localization unit, a reverberation suppression unit, a feature quantity calculation unit, and a feature quantity adjustment unit, which determines the sound source position, suppresses reverberation components, calculates adjusted feature quantities, and performs speech recognition using these adjusted quantities to maintain recognition accuracy despite changes in sound source position.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech recognition processing is performed for recorded speech under reverberation conditions, then speech recognition can be performed in indoor environments, but speech recognition rate becomes lower than original sound
Solution Approach 1:
The invention extracts and removes the reverberation component from the recorded speech signal. The reverberation suppression unit separates the direct sound from the reverberation component and eliminates the reverberation, keeping only the clean direct sound for speech recognition processing. This extraction approach resolves the contradiction by removing the harmful reverberation while preserving the useful speech content.
Solution Approach 2:
The invention uses the captured reverberation component to generate position information about the sound source. By analyzing the characteristics of the reverberation, the system determines the positional relationship between the sound source and recording point, then uses this position information to adjust the acoustic model. This converts the harmful reverberation into a useful source of positional information that improves recognition accuracy.
2Adaptability or versatility
If the position of sound source or recording point changes, then mobility and flexibility are improved, but speech recognition rate is lowered due to different degrees of reverberation
Solution Approach 1:
The invention dynamically adapts the acoustic model based on the determined sound source position. The system adjusts the acoustic model parameters according to the positional relationship between sound source and recording point, allowing the recognition system to optimize for each specific configuration. This dynamic adaptation resolves the contradiction by making the system flexible to position changes while maintaining high recognition rates.
Solution Approach 2:
The invention changes the parameters of the acoustic model based on the sound source position information. By adjusting acoustic model parameters according to the determined position, the system compensates for variations in reverberation characteristics that occur with different positional relationships. This parameter adjustment maintains speech recognition accuracy despite changes in sound source or recording point position.
3Device complexity
If reverberation suppression is performed without position information, then processing complexity is reduced, but speech recognition accuracy decreases due to position-dependent reverberation characteristics
Solution Approach 1:
The invention performs preliminary determination of sound source position information from the reverberation component before the speech recognition process. By extracting position information in advance and using it to adjust the acoustic model beforehand, the system prepares the optimal recognition parameters for the specific situation. This preliminary action resolves the contradiction by enabling accurate position-based adaptation without adding significant complexity to the main recognition process.
Data Source
AI summary
A speech processing device includes a sound source localization unit configured to determine a sound source position from acquired speech, a reverberation suppression unit configured to suppress a reverberation component of the speech to generate dereverberated speech, a feature quantity calculation unit configured to calculate a feature quantity of the dereverberated speech, a feature quantity adjustment unit configured to multiply the feature quantity by an adjustment factor corresponding to the sound source position to calculate an adjusted feature quantity, and a speech recognition unit configured to perform speech recognition using the adjusted feature quantity.


