Teleconference Diarization Using Multi-Modal Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional diarization systems struggle to accurately identify speakers in audio and video teleconferences, especially when there are overlaps, short utterances, muttering, or other audio artifacts, leading to incomplete or inaccurate speaker identification.

Innovation Solution

A method involving a computer system that processes audio, video, and metadata components of a recorded teleconference to parse speech segments, tag speaker identification, and use neural networks for accurate speaker labeling, incorporating transcription data and visual features to improve diarization accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional diarization systems are used to identify speakers in teleconferences, then the system complexity remains low, but the speaker identification accuracy deteriorates in complex scenarios with overlaps, short utterances, and audio artifacts

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The audio stream is divided into multiple speech segments based on voice activity detection and speaker change detection. Each segment is then independently analyzed and labeled with speaker identity, allowing the system to handle complex overlapping scenarios by processing smaller, manageable units rather than the entire audio stream at once.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transitions from analyzing only audio features to incorporating multiple dimensions including video feeds, transcription data, and metadata. This multi-dimensional approach provides additional cues for speaker identification, improving accuracy in challenging scenarios where audio alone is insufficient.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If multiple data sources (audio, video, transcription, metadata) are integrated to improve speaker identification, then the accuracy improves, but the data processing complexity increases

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The integration of multiple data sources is performed in a segmented manner, where each speech segment is processed independently with all data sources aligned to that segment. This reduces the overall processing complexity by breaking down the large-scale integration problem into smaller, manageable tasks.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system creates a universal diarization framework that can handle multiple types of input data (audio, video, transcription, metadata) through a common processing pipeline. This multi-functional approach allows the same system to process diverse data sources without requiring separate specialized processors for each type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of information

If speech segments are parsed and labeled with speaker information to improve diarization accuracy, then the information completeness improves, but the processing time increases

Engineering Contradiction:
Improveinformation completenessVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system performs preliminary voice activity detection and speaker change detection to identify potential speech segments before detailed speaker identification is performed. This preliminary segmentation prepares the data structure in advance, allowing faster processing during the actual speaker identification phase.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system replaces traditional mechanical speaker identification methods with neural network-based approaches. The neural networks are trained offline to learn speaker characteristics, enabling fast online identification without complex real-time mechanical processing of all audio features.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11978457B2Method for uniquely identifying participants in a recorded streaming teleconference
Publication Date: 2024.05.07 GONG IO INC
  • US11978457B2 patent drawing
  • US11978457B2 patent drawing
  • US11978457B2 patent drawing

AI summary

Methods for uniquely identifying respective participants in a teleconference involving obtaining components of the teleconference including an audio component, a video component, teleconference metadata, and transcription data, parsing components into plural speech segments, tagging respective speech segments with speaker identification information, and diarizing the teleconference so as to label respective speech segments.