Audio Clarification via Text Overlay for Video Playback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face inefficiencies in clarifying unclear or missed auditory information in video content, such as unclear dialogue or song lyrics, due to the need to rewind videos or search online, which can be distracting and ineffective, especially when intrinsic characteristics like accents are involved.

Innovation Solution

A system and method that allows users to search for and receive immediate clarification of words or lyrics heard in audio or video content through a combination of audio and video content identification, matching, and technical processing, providing coherent and intelligible textual information, including context and sentiment, even for live or previously unrecorded content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the user rewinds the video to re-hear the verbal content, then the user can hear the content again, but the user's time expended to watch the video increases

Engineering Contradiction:
Improveclarity of verbal contentVSAvoidtime expended to watch video
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary action by automatically generating and displaying text transcripts of verbal content as the video plays, so that when users miss or hear content unclearly, the text is already available for immediate reference without requiring video rewinding

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces text transcript as an intermediary medium between the audio verbal content and the user's understanding. Instead of directly replaying audio, the system mediates through text representation, allowing users to clarify missed content efficiently

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If the user searches online for information to clarify the verbal content, then the user can find information about the content, but the user is distracted from the video itself

Engineering Contradiction:
Improveinformation about verbal contentVSAvoiduser distraction from video
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent merges the video playback interface with the text transcript display, combining two functions (video viewing and text reference) into a single integrated interface. This eliminates the need for users to switch to separate search tools or websites, keeping them engaged with the video while providing access to textual clarification

Inventive Principle:
Principle #5Merging (Combining)

3Reliability

If the video content contains intrinsic characteristics such as accents, then the verbal content may be unclear to some users, but rewinding does not cure the problem

Engineering Contradiction:
Improveclarity of verbal contentVSAvoidaccommodation of different user needs
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system changes the parameter of information representation from audio-only to text-based, providing an alternative modal representation that accommodates users with different needs (e.g., those who struggle with certain accents, language learners, or users who prefer reading). This parameter change makes the content adaptable to diverse user characteristics

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP4123481A1Clarifying audible verbal information in video content
Publication Date: 2023.01.25 GOOGLE LLC
  • EP4123481A1 patent drawingFigure 1A
  • EP4123481A1 patent drawingFigure 1B
  • EP4123481A1 patent drawingFigure 2

AI summary

A method at a server includes: receiving a user request to clarify audible verbal information associated with a media content item playing in proximity to a client device, where the user request includes an audio sample of the media content item and a user query, and the audio sample corresponds to a portion of the media content item proximate in time to issuance of the user query; in response to the user request: identifying the media content item and a first playback position in the media content corresponding to the audio sample; in accordance with the first playback position and identity of the media content item, obtaining textual information corresponding to the user query for a respective portion of the media content item; and transmitting to the client device at least a portion of the textual information.