Joint Speaker Location and Identification Neural Network
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speaker recognition and location systems face challenges in accurately determining speaker changes and spatial positions, especially when speakers have similar vocal characteristics or are close to each other, as they either lack access to spatial information or vocal characteristics.
Innovation Solution
A joint speaker location/speaker identification neural network is trained using both magnitude and phase information features from multi-channel audio signals, enabling the system to exploit both types of information for more accurate speaker identification and location results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If speaker recognition systems use only magnitude features, then the system complexity is reduced, but the measurement precision of speaker identification deteriorates
Solution Approach 1:
The patent combines magnitude features and phase information features into a unified neural network model for joint speaker identification and location. This merging allows the system to utilize both types of features simultaneously, improving measurement precision while managing complexity through integrated architecture design.
Solution Approach 2:
The neural network is designed to perform multiple functions: speaker identification, speaker location, and phase information processing. This multi-functionality allows the system to extract both magnitude and phase features through a single unified model, improving overall system efficiency and accuracy.
2Speed
If speaker recognition systems use only spatial information, then the processing speed is improved, but the measurement precision of speaker identification deteriorates
Solution Approach 1:
The system merges spatial information (phase features) with spectral information (magnitude features) in a joint neural network model. This combination allows the system to maintain fast processing through efficient feature extraction while improving identification accuracy by utilizing complementary information from both magnitude and phase domains.
3Device complexity
If separate systems are used for speaker identification and location, then the device complexity is reduced, but the measurement precision of both functions deteriorates
Solution Approach 1:
The patent implements a joint neural network that simultaneously performs speaker identification and location estimation by processing both magnitude and phase features. This unified approach improves measurement precision for both functions by allowing the system to leverage correlations between speaker characteristics and spatial position, while the modular neural network architecture manages device complexity effectively.
Solution Approach 2:
The neural network is designed as a universal model that handles multiple tasks: speaker identification, speaker location, and feature extraction. This multi-functional design improves overall system performance by sharing computational resources and learning joint representations, while maintaining manageable device complexity through efficient architecture design.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Computing devices and methods utilizing a joint speaker location/speaker identification neural network are provided. In one example a computing device receives a multi-channel audio signal of an utterance spoken by a user. Magnitude and phase information features are extracted from the signal and inputted into a joint speaker location/speaker identification neural network that is trained via utterances from a plurality of persons. A user embedding comprising speaker identification characteristics and location characteristics is received from the neural network and compared to a plurality of enrollment embeddings extracted from the plurality of utterances that are each associated with an identity of a corresponding person. Based at least on the comparisons, the user is matched to an identity of one of the persons, and the identity of the person is outputted.