Temporal Convolutional Network for Latent Speaker State Embeddings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional automated speech recognition (ASR) systems fail to accurately capture latent speaker states, such as emotions and prosodic cues, due to their reliance on text modeling, language dependency, and high computational costs, leading to loss of information and inefficiencies in data usage.

Innovation Solution

A fused deep neural network architecture that processes raw audio waveforms to generate identity embeddings, using temporal convolutional networks trained with triplet loss functions and similarity learning techniques, allowing for end-to-end processing without spectral transformation or language transcription, and enabling efficient detection of speaker attributes like emotion and language.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional ASR systems use text modeling and spectral transformations to process speech, then speech recognition accuracy is improved, but computational cost and data processing time increase significantly

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidcomputational cost
Core Design Contradiction:
Measurement precisionVSUse of energy by stationary object

Solution Approach 1:

The patent extracts and processes only the most relevant acoustic features directly from raw waveforms using a streamlined neural network architecture, eliminating the need for complex spectral transformations and extensive feature engineering while maintaining recognition accuracy

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces traditional mechanical signal processing pipelines (spectral transformation, feature extraction, language modeling) with a direct end-to-end neural network approach that processes raw audio waveforms, significantly reducing computational overhead

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If conventional ASR systems perform extensive spectral transformation and language transcription, then speech understanding is improved, but information loss increases

Engineering Contradiction:
Improvespeech understanding accuracyVSAvoidacoustic information loss
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

Instead of transforming speech into spectral representations and then back to text, the patent inverts the approach by directly mapping raw waveforms to semantic meaning through neural networks, preserving acoustic information throughout the processing pipeline

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The patent introduces latent speaker state embeddings as intermediate representations that capture both acoustic and semantic information, serving as a bridge between raw audio and textual output while preserving information from both domains

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If conventional ASR systems use language-dependent modeling, then recognition accuracy for trained languages is improved, but adaptability to new languages and dialects deteriorates

Engineering Contradiction:
Improverecognition accuracyVSAvoidlanguage adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent implements a universal speaker state representation that captures fundamental acoustic patterns applicable across languages and dialects, allowing the same model architecture to adapt to multiple languages without retraining the core feature extraction components

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent performs preliminary extraction of language-invariant speaker states from raw audio before language-specific processing, enabling the system to adapt to new languages more efficiently by leveraging pre-learned acoustic patterns

Inventive Principle:
Principle #10Preliminary action

4Measurement precision

If conventional ASR systems process complete audio streams with complex models, then recognition accuracy is improved, but processing speed and real-time performance deteriorate

Engineering Contradiction:
Improverecognition accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSSpeed

Solution Approach 1:

The patent segments the audio processing into discrete frames with overlapping windows, allowing parallel processing of multiple segments through the neural network and enabling real-time processing with maintained accuracy

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent processes only the most informative acoustic features and speaker states at each time step rather than analyzing the complete audio spectrum, reducing computational load while maintaining recognition accuracy through selective feature processing

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11646037B2Sample-efficient representation learning for real-time latent speaker state characterization
Publication Date: 2023.05.09 UNITY TECH SF
  • US11646037B2 patent drawing
  • US11646037B2 patent drawing
  • US11646037B2 patent drawing

AI summary

Systems, methods, and non-transitory computer-readable media can provide audio waveform data that corresponds to a voice sample to a temporal convolutional network for evaluation. The temporal convolutional network can pre-process the audio waveform data and can output an identity embedding associated with the audio waveform data. The identity embedding associated with the voice sample can be obtained from the temporal convolutional network. Information describing a speaker associated with the voice sample can be determined based at least in part on the identity embedding.