Speech Processing Apparatus Direction of Arrival Correction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speaker clustering methods using microphone arrays struggle with accurate recognition of speakers due to inaccuracies in direction of arrival estimation and speaker positioning, leading to poor clustering performance.
Innovation Solution
A speech processing apparatus that includes an acquisition unit, separation unit, calculation unit, estimation unit, correction unit, and clustering unit, which separates speech into sections, calculates and corrects similarity based on direction of arrival, and clusters sections with similar acoustic features to improve recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If direction of arrival estimation is used for speaker clustering, then speaker recognition can be performed, but recognition accuracy deteriorates due to estimation inaccuracies
Solution Approach 1:
The patent introduces a correction unit that acts as an intermediary between the direction of arrival estimation unit and the clustering unit. This correction unit refines the estimated direction of arrival values by correcting estimation errors, thereby improving the accuracy of speaker clustering without requiring a complete redesign of the estimation system. The correction unit processes the estimated directions and produces corrected directions that better represent the true speaker positions.
Solution Approach 2:
The system implements a feedback mechanism where the clustering results and similarity scores are used to refine and correct the direction of arrival estimates. The correction unit utilizes information from the clustering process to adjust the estimated directions, creating a feedback loop that continuously improves estimation accuracy based on actual clustering performance.
2Device complexity
If speaker clustering is performed based on acoustic features alone, then processing is simple, but clustering accuracy deteriorates due to insufficient spatial information
Solution Approach 1:
The patent merges two different types of information: acoustic features (similarity scores from speech content) and spatial features (corrected direction of arrival estimates). By combining these complementary information sources, the system achieves more accurate speaker clustering than would be possible using either feature type alone, while maintaining reasonable processing complexity through efficient integration methods.
Solution Approach 2:
The system transitions from one-dimensional acoustic feature analysis to two-dimensional analysis by incorporating the spatial dimension (direction of arrival) as an additional feature space. This dimensional expansion allows the clustering algorithm to differentiate between speakers more effectively by considering both what they say (acoustic) and where they are located (spatial).
Data Source
AI summary
In a speech processing apparatus, an acquisition unit is configured to acquire a speech. A separation unit is configured to separate the speech into a plurality of sections in accordance with a prescribed rule. A calculation unit is configured to calculate a degree of similarity in each combination of the sections. An estimation unit is configured to estimate, with respect to the each section, a direction of arrival of the speech. A correction unit is configured to group the sections whose directions of arrival are mutually similar into a same group and correct the degree of similarity with respect to the combination of the sections in the same group. A clustering unit is configured to cluster the sections by using the corrected degree of similarity.


