Scene Break Detection via Multimedia Embedding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for detecting scene breaks in media content are inefficient and prone to variability, as they rely heavily on human intervention, which can lead to inconsistent placement of cue points for targeted media content.

Innovation Solution

The system segments media content into units by detecting unit boundaries and encodes these units into a multimedia representation in an embedding space across different media modalities. A sequence classifier then identifies scene boundaries based on these representations, improving accuracy and reducing human error.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If human intervention is used to detect scene breaks, then the process is simple to implement, but the accuracy and consistency of scene boundary identification deteriorates due to variability and human error

Engineering Contradiction:
Improveease of implementationVSAvoidscene break detection accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent replaces the mechanical human decision-making process with an automated machine learning system. The system uses feature encoders to extract visual, audio, and text features, processes them through a sequence classifier trained to identify scene boundaries, and outputs automated scene break detections. This substitution eliminates human variability while maintaining implementation feasibility through standardized computational processes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by allowing the media content to be automatically analyzed without human intervention. The sequence classifier independently processes the encoded features and identifies scene boundaries based on learned patterns from training data, making the system autonomous and eliminating the need for manual scene break detection.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If automated feature encoding and sequence classification are applied, then the accuracy of scene boundary identification improves, but the device complexity and computational resources required increase

Engineering Contradiction:
Improvescene break detection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the complex scene break detection task into distinct modular components: feature encoding (extracting visual, audio, and text features), sequence classification (identifying scene boundaries), and post-processing (refining detections). Each module can be independently developed, trained, and optimized, reducing overall system complexity while maintaining high accuracy through specialized processing at each stage.

Inventive Principle:
Principle #1Segmentation

3Reliability

If multiple media modalities are encoded into embedding space, then the relevance of targeted media content insertion improves, but the computational energy and processing time required increase

Engineering Contradiction:
Improverelevance of content insertionVSAvoidcomputational energy consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary encoding of media content into embedding space during an offline preprocessing stage. Feature encoders extract and encode visual, audio, and text features from the entire media content before insertion points are needed. This preliminary action stores the encoded representations for rapid retrieval and comparison during targeted content insertion, reducing real-time computational energy requirements while maintaining high relevance through multi-modal feature integration.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250142183A1Scene break detection
Publication Date: 2025.05.01 ROKU INC
  • US20250142183A1 patent drawing
  • US20250142183A1 patent drawing
  • US20250142183A1 patent drawing

AI summary

Disclosed herein are system, apparatus, article of manufacture, method and/or computer program product embodiments, and/or combinations and sub-combinations thereof, for identifying scene breaks in media content. An example method comprises segmenting media content into a sequence of units by detecting unit boundaries. One or more feature encoders are applied to generate in an embedding space a multimedia representation of features of each unit in the sequence across different media modalities. A sequence classifier is applied to identify whether a unit boundary is a scene boundary based on the multimedia representation of units in the embedding space in at least a subset of the sequence of units.