Multi-Speaker Speech Transcription via Visual Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech transcription systems struggle to efficiently transcribe interactions between multiple people in real-time, often requiring specific commands or instructions and lacking natural user interaction.

Innovation Solution

A system that combines speech recognition, speaker identification, and visual pattern recognition using AI/ML models to transcribe speech in real-time, identifying speakers based on image data and generating additional data such as calendar invitations, task lists, and notifications, without the need for specific commands.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech recognition systems use specific commands or instructions, then transcription accuracy improves, but user interaction naturalness deteriorates

Engineering Contradiction:
Improvetranscription accuracyVSAvoiduser interaction naturalness
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system automatically performs speaker identification and transcription without requiring users to issue commands. The speech recognition system serves itself by autonomously capturing, processing, and transcribing speech segments, eliminating the need for users to initiate transcription processes with specific commands while maintaining high accuracy through automated speaker verification

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary speaker identification and segmentation of speech before full transcription occurs. By pre-identifying speakers and separating their speech segments, the system prepares the data structure needed for accurate transcription without requiring users to provide commands during the actual speech interaction

Inventive Principle:
Principle #10Preliminary action

2Loss of information

If the system transcribes all speech segments from multiple speakers, then completeness of transcription improves, but system complexity increases

Engineering Contradiction:
Improvetranscription completenessVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system divides the transcription task into separate segments for each speaker. By identifying speakers first and then transcribing their individual speech segments separately, the system maintains complete transcription coverage while managing complexity through modular processing of speaker-specific audio streams rather than attempting to process all speech as a single complex stream

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The speech recognition system is designed to handle multiple speakers universally using the same transcription pipeline. Rather than requiring different processing paths for different numbers of speakers, the system uses a universal approach where speaker identification automatically adapts the processing to handle any number of speakers, reducing overall system complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11749285B2Speech transcription using multiple data sources
Publication Date: 2023.09.05 META PLATFORMS TECHNOLOGIES LLC
  • US11749285B2 patent drawing
  • US11749285B2 patent drawing
  • US11749285B2 patent drawing

AI summary

This disclosure describes transcribing speech using audio, image, and other data. A system is described that includes an audio capture system configured to capture audio data associated with a plurality of speakers, an image capture system configured to capture images of one or more of the plurality of speakers, and a speech processing engine. The speech processing engine may be configured to recognize a plurality of speech segments in the audio data, identify, for each speech segment of the plurality of speech segments and based on the images, a speaker associated with the speech segment, transcribe each of the plurality of speech segments to produce a transcription of the plurality of speech segments including, for each speech segment in the plurality of speech segments, an indication of the speaker associated with the speech segment, and analyze the transcription to produce additional data derived from the transcription.