Speech Segmentation for Precise Media Navigation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face difficulty in navigating media content using Fast Forward/REW functions, as these often skip critical segments of the storyline, making it hard to follow the narrative.

Innovation Solution

A system comprising a server that detects speech data from media content, divides it into segments based on speakers and breaks, and a media content playing device that receives these segments and control signals to skip forward or rewind to specific points within the narrative.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If Fast Forward/REW functions skip segments of media content by time units, then navigation speed is improved, but storyline continuity deteriorates

Engineering Contradiction:
Improvenavigation speedVSAvoidstoryline continuity
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The patent segments the media content into speech data segments based on speaker changes and breaks in the audio stream. Each segment is identified with metadata including start time, end time, and speaker information. This segmentation allows the system to navigate between meaningful narrative units rather than arbitrary time intervals, maintaining storyline continuity while enabling efficient navigation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system detects speech data from the media content and uses this feedback to dynamically identify and segment the content. The speech detection algorithm analyzes audio characteristics to determine when speakers change or when breaks occur, providing feedback that guides the segmentation process. This feedback mechanism ensures that segments are defined by actual narrative structure rather than fixed time units.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If media content is divided into speech data segments based on speakers and breaks, then navigation precision is improved, but system complexity increases

Engineering Contradiction:
Improvenavigation precisionVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts speech data from the media content stream using speech recognition technology. The system isolates and processes only the audio components that contain speech information, separating this data from the video and other audio tracks. This extraction process enables precise identification of speech segments without requiring complex analysis of the entire media stream, reducing overall system complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system introduces an intermediary layer of speech data segments that mediates between the raw media content and the user navigation requests. Rather than directly analyzing the complete media stream for navigation decisions, the system uses these intermediate speech segments as reference points. This intermediary structure simplifies the navigation process while maintaining precision in locating specific narrative moments.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9830933B2Media content playing scheme
Publication Date: 2017.11.28 KT CORP
  • US9830933B2 patent drawing
  • US9830933B2 patent drawing
  • US9830933B2 patent drawing

AI summary

A system may include a server configured to detect speech data from media content and to divide the detected speech data into one or more speech data segments in accordance with at least a respective speaker and a break in the detected speech data; and a media content playing device configured to receive the speech data segments from the server, to receive, from an input device, a control signal to play the media content, and to skip forward or rewind to play the media content starting at the identified starting point corresponding to a first one of the respective speech data segments.