Joint Speaker Location and Identification Neural Network

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speaker recognition and location systems face challenges in accurately determining speaker changes and spatial positions, especially when speakers have similar vocal characteristics or are close to each other, as they either lack access to spatial information or vocal characteristics.

Innovation Solution

A joint speaker location/speaker identification neural network is trained using both magnitude and phase information features from multi-channel audio signals, enabling the system to exploit both types of information for more accurate speaker identification and location results.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If speaker recognition systems use only magnitude features, then the system complexity is reduced, but the measurement precision of speaker identification deteriorates

Engineering Contradiction:
Improvesystem complexityVSAvoidspeaker identification accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent combines magnitude features and phase information features into a unified neural network model for joint speaker identification and location. This merging allows the system to utilize both types of features simultaneously, improving measurement precision while managing complexity through integrated architecture design.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The neural network is designed to perform multiple functions: speaker identification, speaker location, and phase information processing. This multi-functionality allows the system to extract both magnitude and phase features through a single unified model, improving overall system efficiency and accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Speed

If speaker recognition systems use only spatial information, then the processing speed is improved, but the measurement precision of speaker identification deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidspeaker identification accuracy
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The system merges spatial information (phase features) with spectral information (magnitude features) in a joint neural network model. This combination allows the system to maintain fast processing through efficient feature extraction while improving identification accuracy by utilizing complementary information from both magnitude and phase domains.

Inventive Principle:
Principle #5Merging (Combining)

3Device complexity

If separate systems are used for speaker identification and location, then the device complexity is reduced, but the measurement precision of both functions deteriorates

Engineering Contradiction:
Improvedevice complexityVSAvoidspeaker identification and location accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent implements a joint neural network that simultaneously performs speaker identification and location estimation by processing both magnitude and phase features. This unified approach improves measurement precision for both functions by allowing the system to leverage correlations between speaker characteristics and spatial position, while the modular neural network architecture manages device complexity effectively.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The neural network is designed as a universal model that handles multiple tasks: speaker identification, speaker location, and feature extraction. This multi-functional design improves overall system performance by sharing computational resources and learning joint representations, while maintaining manageable device complexity through efficient architecture design.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3791393B1Speaker recognition/location using neural network
Publication Date: 2023.01.25 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3791393B1 patent drawingFigure 1
  • EP3791393B1 patent drawingFigure 2
  • EP3791393B1 patent drawingFigure 3

AI summary

Computing devices and methods utilizing a joint speaker location/speaker identification neural network are provided. In one example a computing device receives a multi-channel audio signal of an utterance spoken by a user. Magnitude and phase information features are extracted from the signal and inputted into a joint speaker location/speaker identification neural network that is trained via utterances from a plurality of persons. A user embedding comprising speaker identification characteristics and location characteristics is received from the neural network and compared to a plurality of enrollment embeddings extracted from the plurality of utterances that are each associated with an identity of a corresponding person. Based at least on the comparisons, the user is matched to an identity of one of the persons, and the identity of the person is outputted.