Speaker Separation via Real-Time Latent State Characterization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional approaches to automated speech recognition (ASR) and speaker diarization fail to accurately capture latent speaker states and non-verbal acoustic cues due to their reliance on text modeling, spectral representations, and high computational costs.

Innovation Solution

The implementation of a fused deep neural network architecture that processes raw audio waveforms to generate identity embeddings, enabling real-time speaker diarization without the need for spectral transformation or language transcription.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional ASR approaches use text modeling and spectral representations, then speech recognition capability is achieved, but ability to capture latent speaker states and non-verbal acoustic cues deteriorates

Engineering Contradiction:
Improvecapture of latent speaker statesVSAvoidnon-verbal acoustic cues
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent extracts and processes raw audio waveforms directly without converting to spectral representations or text, thereby preserving latent speaker states and non-verbal acoustic cues that would be lost in conventional text-modeling approaches

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs multiple functions simultaneously: speaker diarization, latent speaker state characterization, and speech recognition, all while processing raw audio to preserve information that would be lost in specialized conventional approaches

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If conventional speaker diarization uses spectral representations and text transcription, then speaker separation is achieved, but real-time processing capability deteriorates due to computational costs

Engineering Contradiction:
Improvereal-time processing speedVSAvoidcomputational cost
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent replaces computationally intensive spectral transformation and text transcription mechanisms with a direct neural network processing approach on raw audio waveforms, reducing computational overhead while maintaining real-time capability

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs preliminary processing by generating identity embeddings from raw audio in real-time, enabling subsequent speaker diarization decisions to be made quickly without requiring heavy post-processing computational resources

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12315516B2Speaker separation based on real-time latent speaker state characterization
Publication Date: 2025.05.27 UNITY TECH SF
  • US12315516B2 patent drawing
  • US12315516B2 patent drawing
  • US12315516B2 patent drawing

AI summary

Systems, methods, and non-transitory computer-readable media can obtain a stream of audio waveform data that represents speech involving a plurality of speakers. As the stream of audio waveform data is obtained, a plurality of audio chunks can be determined. An audio chunk can be associated with one or more identity embeddings. The stream of audio waveform data can be segmented into a plurality of segments based on the plurality of audio chunks and respective identity embeddings associated with the plurality of audio chunks. A segment can be associated with a speaker included in the plurality of speakers. Information describing the plurality of segments associated with the stream of audio waveform data can be provided.