Visual-Text Search Interface for Transcript-Based Video Editing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video editing tools are tedious and challenging for many users due to the need for fine-grained interactions with video frames, and existing speaker diarization techniques often result in over-segmentation and inaccuracies, especially when dealing with music and small faces.

Innovation Solution

A hybrid speaker diarization technique that combines audio and visual cues to accurately identify speakers, along with transcript segmentation and interaction modalities that allow users to select and edit video segments through transcript interactions, including face-aware and music-aware diarization to improve video editing efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional video editing tools are used with fine-grained interactions with video frames, then editing precision is improved, but user ease of operation deteriorates

Engineering Contradiction:
Improveediting precisionVSAvoiduser ease of operation
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent introduces a transcript as an intermediary representation between the video content and user interactions. Instead of requiring users to directly interact with video frames, the system transcribes audio to text and allows users to select and edit video segments by interacting with the transcript, thereby simplifying the operation interface while maintaining editing precision

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical interaction model (direct frame selection) with a semantic interaction model (text-based selection). By substituting the traditional time-based selection mechanism with text-based selection, the system enables users to locate and edit specific video segments more easily through semantic search and transcript interaction

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If audio-only speaker diarization is used, then processing speed is improved, but measurement precision deteriorates due to over-segmentation and inaccuracy

Engineering Contradiction:
Improveprocessing speedVSAvoidspeaker identification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent merges audio-based speaker diarization with visual-based speaker identification by combining audio features (voice characteristics) with visual features (face recognition). This fusion allows the system to leverage the speed advantage of audio processing while compensating for its inaccuracies through visual verification, thereby improving overall speaker identification accuracy without significant speed penalty

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a composite speaker diarization system that integrates multiple data sources (audio and visual streams) to form a more robust and accurate speaker identification model. By treating audio and visual information as complementary materials, the system achieves both speed and accuracy through synergistic processing

Inventive Principle:
Principle #40Composite materials

3Measurement precision

If hybrid audio-visual speaker diarization is applied, then speaker identification accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the speaker diarization process into distinct stages: audio-based initial diarization, visual-based refinement, and integration. By dividing the complex task into manageable segments, the system can process audio and visual data separately before combining results, thereby reducing overall system complexity while maintaining high accuracy

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary audio-based speaker diarization before incorporating visual information. This preliminary action establishes a baseline speaker segmentation that can then be refined using visual data, allowing the system to build complexity incrementally rather than requiring all processing to occur simultaneously, thus managing system complexity more effectively

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12367238B2Visual and text search interface for text-based video editing
Publication Date: 2025.07.22 ADOBE INC
  • US12367238B2 patent drawing
  • US12367238B2 patent drawing
  • US12367238B2 patent drawing

AI summary

Embodiments of the present invention provide systems, methods, and computer storage media for a visual and text search interface used to navigate a video transcript. In an example embodiment, a freeform text query triggers a visual search for frames of a loaded video that match the freeform text query (e.g., frame embeddings that match a corresponding embedding of the freeform query), and triggers a text search for matching words from a corresponding transcript or from tags of detected features from the loaded video. Visual search results are displayed (e.g., in a row of tiles that can be scrolled to the left and right), and textual search results are displayed (e.g., in a row of tiles that can be scrolled up and down). Selecting (e.g., clicking or tapping on) a search result tile navigates a transcript interface to a corresponding portion of the transcript.