Speaker Thumbnail Selection in Diarized Transcripts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video editing is tedious, challenging, and often beyond the skill level of many users due to its time-based nature and lack of intuitive interaction modalities.

Innovation Solution

The use of transcript interactions for video segment selection and editing, including face-aware speaker diarization, music-aware speaker diarization, and transcript paragraph segmentation, to facilitate text-based editing of audio and video assets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional time-based video editing is used, then video frames can be selected and edited, but the process becomes tedious and challenging for users

Engineering Contradiction:
Improveease of video editingVSAvoidtime required for editing
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent replaces the mechanical time-based editing system with a text-based interaction system. Instead of manually dragging and dropping video frames on a timeline, users interact with transcribed text to select and edit video segments. The system automatically maps text selections to corresponding video timecodes and performs editing operations, substituting manual mechanical operations with automated text-driven control.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces text transcription as an intermediary layer between the user and the video content. Users interact with the text transcript rather than directly with video frames, and the system uses this text selection as a mediator to automatically identify and edit the corresponding video segments. This intermediary simplifies user interaction while maintaining precise control over video editing.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If speaker diarization is implemented to identify different speakers, then speaker information can be displayed, but the complexity of the system increases

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments the video content by identifying and separating different speaker segments through diarization. The system divides the continuous video and audio stream into discrete speaker segments, each associated with a specific speaker identity. This segmentation enables the system to track and display speaker information throughout the video without requiring complex real-time analysis of the entire video stream.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs speaker diarization and speaker identification as preliminary actions during video processing, before the user interacts with the content. The system pre-analyzes the audio and video streams to identify speakers, generate transcripts with speaker labels, and create speaker segment information in advance. This preliminary processing reduces the complexity of real-time operations and enables efficient speaker information display during user interaction.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12300272B2Speaker thumbnail selection and speaker visualization in diarized transcripts for text-based video
Publication Date: 2025.05.13 ADOBE INC
  • US12300272B2 patent drawing
  • US12300272B2 patent drawing
  • US12300272B2 patent drawing

AI summary

Embodiments of the present invention provide systems, methods, and computer storage media for selection of the best image of a particular speaker's face in a video, and visualization in a diarized transcript. In an example embodiment, candidate images of a face of a detected speaker are extracted from frames of a video identified by a detected face track for the face, and a representative image of the detected speaker's face is selected from the candidate images based on image quality, facial emotion (e.g., using an emotion classifier that generates a happiness score), a size factor (e.g., favoring larger images), and/or penalizing images that appear towards the beginning or end of a face track. As such, each segment of the transcript is presented with the representative image of the speaker who spoke that segment and/or input is accepted changing the representative image associated with each speaker.