Cue Point Aware Content Stitching for Gap-Free Ad Insertion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for stitching secondary audio and video content, such as advertisements, into primary content like television shows or movies, face challenges due to cue point agnostic encoding, which results in overlapping or gaps between content segments, especially when using dynamic durations.

Innovation Solution

Implement cue point aware encoding (CAE) to align video and audio fragments accurately, ensuring the accumulated duration of audio segments is equal to or longer than video segments, but not by more than one frame duration, and adjust cue points to the precision of milliseconds.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If cue point agnostic encoding is used with fixed fragment durations, then encoding simplicity is maintained, but content alignment precision deteriorates causing overlaps and gaps

Engineering Contradiction:
Improveencoding simplicityVSAvoidcontent alignment precision
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent changes the encoding approach from cue point agnostic to cue point aware encoding, transforming how video and audio fragments are encoded to track and align cue points accurately. This parameter change in the encoding method enables precise content alignment while maintaining manageable encoding processes through systematic cue point tracking.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If dynamic fragment durations are used to adapt to content variations, then content adaptability improves, but alignment precision deteriorates due to timing mismatches

Engineering Contradiction:
Improvecontent adaptabilityVSAvoidalignment precision
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent applies preliminary action by pre-identifying and marking cue points in the content before fragmentation. These cue points serve as predetermined reference markers that guide subsequent alignment operations, ensuring that even with dynamic fragment durations, the content segments can be accurately reassembled without overlaps or gaps.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback mechanisms through cue point tracking that monitors the alignment status of video and audio fragments throughout the encoding and playback process. This feedback system allows the player to detect and correct timing deviations, ensuring accurate synchronization even when fragment durations vary dynamically.

Inventive Principle:
Principle #23Feedback

3Ease of operation

If secondary content is inserted at fixed time locations, then insertion simplicity is maintained, but content continuity deteriorates due to overlaps and gaps

Engineering Contradiction:
Improveinsertion simplicityVSAvoidcontent continuity
Core Design Contradiction:
Ease of operationVSStability of the object's composition

Solution Approach 1:

The patent transforms the static approach of fixed-time insertion into a dynamic system where insertion points are determined by cue point locations rather than fixed timestamps. This dynamic approach allows the system to adapt insertion timing to the actual content structure, maintaining continuity while still providing systematic and manageable insertion operations through cue point guidance.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12483736B1Systems and methods for audio and video content stitching
Publication Date: 2025.11.25 AMAZON TECH INC
  • US12483736B1 patent drawing
  • US12483736B1 patent drawing
  • US12483736B1 patent drawing

AI summary

Systems and methods for audio and video content stitching are provided. an example method may include receiving a first video fragment and a second video fragment included within primary video content, the first video fragment and the second video fragment including one or more video frames, wherein the first video fragment is a first time duration, wherein the second video fragment is a second time duration, and wherein the first time duration and the second time duration are different. The example method may also include receiving a third video fragment associated with secondary video content. The example method may also include determining a first cue point within the primary video content, wherein the first cue point is located at a first time between a starting time of the first video fragment and a second time corresponding to a sum of the first time and a duration of the first time duration. The example method may also include adding the secondary video content to the primary video content at the first cue point. The example method may also include presenting the secondary video content at the first cue point within the primary video content.