Voice-Driven Graphic Overlay Timing for Video Images

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing electronic devices lack an efficient method to automatically convert voice into text and apply relevant graphic data to video content based on keywords, requiring manual user intervention for selecting and timing graphic data application.

Innovation Solution

An electronic device with a processor that analyzes voice signals from video content, extracts keywords, determines corresponding graphic data, and applies it to selected images at appropriate times, optionally with user input or automatic selection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual user intervention is used for selecting and timing graphic data application, then user control over graphic data selection is maintained, but user convenience deteriorates and time consumption increases

Engineering Contradiction:
Improveuser convenienceVSAvoidmanual operation complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system performs automatic keyword extraction from voice signals, automatic graphic data determination based on keywords, and automatic timing identification without requiring manual user intervention for these tasks, allowing the system to serve itself

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system extracts keywords from voice signals and determines corresponding graphic data in advance before the user needs to apply them, preparing the graphic data and timing information beforehand to eliminate manual search and selection steps

Inventive Principle:
Principle #10Preliminary action

2Productivity

If automatic keyword extraction and graphic data determination is implemented, then productivity increases and time consumption decreases, but system complexity increases

Engineering Contradiction:
Improvegraphic data application efficiencyVSAvoidprocessing system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system divides the complex task of graphic data application into separate modular functions: voice signal processing, keyword extraction, graphic data determination, and timing identification, where each module handles a specific aspect independently

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Keywords serve as an intermediary between voice signals and graphic data, translating spoken content into selectable graphic data categories, while timing information acts as an intermediary to automatically determine when to apply graphic data based on voice output timing

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of time

If voice-based automatic graphic data selection is implemented, then user convenience improves and manual search time is eliminated, but voice recognition accuracy requirements increase

Engineering Contradiction:
Improvemanual search timeVSAvoidvoice recognition accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

Instead of requiring perfect voice recognition of the entire sentence, the system extracts only key words or phrases that are sufficient to determine appropriate graphic data, accepting partial recognition results that meet the minimum threshold for accurate graphic data selection

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP3906553B1Electronic device for providing graphic data based on voice and operating method thereof
Publication Date: 2026.04.08 SAMSUNG ELECTRONICS CO LTD
  • EP3906553B1 patent drawingFigure 1
  • EP3906553B1 patent drawingFigure 2
  • EP3906553B1 patent drawingFigure 3

AI summary

An electronic device for providing graphic data based on a voice, and an operation method therefor are provided. The electronic device includes a display, and a processor, and the processor is configured to obtain at least one keyword from a voice signal related to a plurality of images, determine at least one graphic data corresponding to the at least one keyword, select at least one of the plurality of images, based on a point in time at which a voice corresponding to a keyword that corresponds to the determined graphic data is output, and perform control so as to apply the determined graphic data to the at least one selected image.