Audio Fingerprint Speaker Identification System

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In multi-party communications sessions, identifying the current speaker can be challenging, especially when multiple speakers share a single client device or use telephones, as visual indicators based on account information are insufficient and may lead to confusion about who is speaking.

Innovation Solution

A computer system that generates an audio fingerprint of the current speaker and performs automated speaker recognition by comparing it against stored fingerprints, providing metadata to client devices to identify the speaker and enhance user interfaces with real-time speaker information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If visual indicators based on account information are used to identify speakers, then the system is simple to operate, but speaker identification becomes inaccurate when multiple speakers share a single client device

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces visual indicators based on account information with audio fingerprinting technology. The system generates unique audio fingerprints from speakers' voice characteristics and compares them against a database of stored fingerprints to automatically identify speakers. This substitutes the mechanical/account-based identification system with an acoustic-based automated recognition system, achieving accurate speaker identification even when multiple speakers share a single client device.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If automated speaker recognition is implemented, then speaker identification accuracy improves, but processing time and computational resources increase

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-generating and storing audio fingerprints in a database before actual speaker identification is needed. When a speaker speaks, the system only needs to compare the current audio fingerprint against the pre-stored ones, significantly reducing processing time. This preliminary preparation of fingerprint data enables fast real-time speaker recognition without requiring complex on-the-fly analysis.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If audio fingerprinting is used for speaker recognition, then speaker identification becomes accurate, but the system requires more computational resources

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidcomputational energy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system extracts only the essential acoustic features from audio signals to create compact audio fingerprints, rather than processing entire audio streams. By extracting and storing only the distinctive fingerprint characteristics in a database, the system reduces computational energy requirements for speaker identification while maintaining high accuracy. This extraction of essential features minimizes the computational burden during real-time recognition.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentEP3271917B1Communicating metadata that identifies a current speaker
Publication Date: 2021.07.14 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3271917B1 patent drawingFigure 1
  • EP3271917B1 patent drawingFigure 2A
  • EP3271917B1 patent drawingFigure 2B

AI summary

A computer system may communicate metadata that identifies a current speaker. The computer system may receive audio data that represents speech of the current speaker, generate an audio fingerprint of the current speaker based on the audio data, and perform automated speaker recognition by comparing the audio fingerprint of the current speaker against stored audio fingerprints contained in a speaker fingerprint repository. The computer system may communicate data indicating that the current speaker is unrecognized to a client device of an observer and receive tagging information that identifies the current speaker from the client device of the observer. The computer system may store the audio fingerprint of the current speaker and metadata that identifies the current speaker in the speaker fingerprint repository and communicate the metadata that identifies the current speaker to at least one of the client device of the observer or a client device of a different observer.