Heterogeneous Audio Processing Architecture for Speaker Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speaker recognition technologies are not optimally suited for energy-efficient operation in various environments and use scenarios, particularly in devices with limited power sources like mobile phones, and struggle with unobtrusive training and efficient identification in real-time conditions.

Innovation Solution

A heterogeneous architecture with a first processing unit handling frequently performed, low-power audio processing tasks and a second unit handling less frequent, more power-intensive tasks, along with unobtrusive training methods using audio segments from everyday conversations and channel compensation techniques, enables efficient speaker identification and model generation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speaker recognition technology processes audio signals in real-time with comprehensive analysis, then identification accuracy is improved, but power consumption increases

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidpower consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The audio processing pipeline is segmented into multiple stages: audio strength testing, speech detection, frame quality assessment, and speaker model matching. Each stage filters and processes only necessary portions of the audio signal, avoiding comprehensive analysis at every step. This segmentation allows the system to maintain high identification accuracy while reducing overall power consumption by eliminating redundant processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs partial processing by applying admission tests and quality thresholds that filter out low-quality speech frames before they undergo full speaker model matching. By performing only the necessary portion of processing on each audio frame (skipping full analysis for frames that fail quality checks), the system maintains accuracy for valid speech while reducing power consumption from unnecessary processing of invalid or low-quality frames.

Inventive Principle:
Principle #16Partial or excessive action

2Reliability

If the system continuously monitors and processes all audio signals for speaker identification, then identification reliability is improved, but device battery life deteriorates

Engineering Contradiction:
Improveidentification reliabilityVSAvoidbattery life
Core Design Contradiction:
ReliabilityVSDuration of action of moving object

Solution Approach 1:

The system implements periodic processing by continuously monitoring audio strength and speech presence at regular intervals, then performing full speaker model matching only when speech is detected and quality thresholds are met. This periodic approach ensures reliable speaker identification when needed while allowing the system to enter lower-power states during intervals without speech, thereby extending battery life.

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The system performs preliminary audio strength testing and speech detection before committing to full speaker model matching. These preliminary actions filter out non-speech or low-quality audio segments, ensuring that computationally intensive processing is only applied to promising candidates. This maintains identification reliability by pre-screening inputs while conserving battery power by avoiding unnecessary full processing of unsuitable audio segments.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If the system applies rigorous quality tests and admission criteria to speech frames, then speaker model accuracy is improved, but processing time increases

Engineering Contradiction:
Improvespeaker model accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The quality assessment process is segmented into multiple independent tests: audio strength testing, speech detection, frame quality assessment, and admission testing. Each test operates on different aspects of the audio signal and can be performed in parallel or selectively. This segmentation allows the system to apply rigorous quality criteria for accurate speaker modeling while optimizing processing time by skipping tests for frames that fail early criteria or processing tests concurrently where possible.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8731936B2Energy-efficient unobtrusive identification of a speaker
Publication Date: 2014.05.20 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8731936B2 patent drawing
  • US8731936B2 patent drawing
  • US8731936B2 patent drawing

AI summary

Functionality is described herein for recognizing speakers in an energy-efficient manner. The functionality employs a heterogeneous architecture that comprises at least a first processing unit and a second processing unit. The first processing unit handles a first set of audio processing tasks (associated with the detection of speech) while the second processing unit handles a second set of audio processing tasks (associated with the identification of speakers), where the first set of tasks consumes less power than the second set of tasks. The functionality also provides unobtrusive techniques for collecting audio segments for training purposes. The functionality also encompasses new applications which may be invoked in response to the recognition of speakers.