Multimodal Diarization for Accurate Meeting Transcripts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current automatic transcription systems for conversations lack accuracy and functionality, particularly in multi-speaker scenarios, and do not effectively incorporate metadata to enhance transcript usability and collaboration.

Innovation Solution

The system employs a multimodal diarization model to identify and label speakers, incorporates metadata such as speaker diarization, timestamp markers, and hyperlinks, and enables collaborative editing, using domain-specific language models to improve transcription accuracy and user interaction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automatic transcription systems are used for multi-speaker conversations, then transcription speed and productivity are improved, but transcription accuracy and speaker identification deteriorate

Engineering Contradiction:
Improvetranscription speedVSAvoidtranscription accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The audio stream is segmented into multiple channels, with each channel dedicated to a specific speaker. The system performs speaker diarization to identify and separate different speakers' audio, then transcribes each channel independently. This segmentation allows the system to maintain high transcription accuracy for each speaker while processing multi-speaker conversations efficiently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer between audio input and transcription output that includes speaker identification, audio channel separation, and metadata generation. This intermediary layer processes the audio to create structured, speaker-labeled segments before transcription, thereby improving overall transcription accuracy without compromising productivity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If metadata is incorporated into transcripts, then transcript usability and information completeness are improved, but system complexity and processing requirements worsen

Engineering Contradiction:
Improveinformation completenessVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by generating metadata (speaker labels, timestamps, audio channel information) during the transcription process itself, rather than as a separate post-processing step. This preliminary generation of metadata integrates information enrichment into the core transcription workflow, reducing overall system complexity while maintaining information completeness.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If speaker diarization and audio channel separation are implemented, then speaker identification accuracy is improved, but computational requirements and processing time worsen

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system implements self-service by using the audio signal itself to perform speaker identification and channel separation. The audio characteristics (voice patterns, spectral features) are analyzed directly from the input signal to automatically segment and label speaker channels without requiring external reference data or manual intervention, thereby reducing processing overhead while maintaining high identification accuracy.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250014582A1Method and system for conversation transcription with metadata
Publication Date: 2025.01.09 SOUNDHOUND AI IP LLC
  • US20250014582A1 patent drawing
  • US20250014582A1 patent drawing
  • US20250014582A1 patent drawing

AI summary

Methods and systems for enabling an efficient review of meeting content via a metadata-enriched, speaker-attributed and multiuser-editable transcript are disclosed. By incorporating speaker diarization and other metadata, the system can provide a structured and effective way to review and/or edit the transcript by one or more editors. One type of metadata can be image or video data to represent the meeting content. Furthermore, the present subject matter utilizes a multimodal diarization model to identify and label different speakers. The system can synchronize various sources of data, e.g., audio channel data, voice feature vectors, acoustic beamforming, image identification, and extrinsic data, to implement speaker diarization.