Directed Content Insertion via Audio Tag Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The insertion of directed content into video assets is computationally intensive due to the analysis of large amounts of image data, requiring significant computing resources and time, which is inefficient compared to analyzing audio data using speech recognition techniques.

Innovation Solution

A computing system analyzes audio data to generate a time series of tags representing keywords or keyphrases, which are then used to create a mapping between time and directed content assets, allowing for the insertion of contextually meaningful content into video assets by associating specific content with specific timestamps.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If image analysis of video frames is used to identify video attributes, then contextually meaningful directed content can be inserted into the video, but the computational resources and time required increase significantly

Engineering Contradiction:
Improvevideo attribute identification accuracyVSAvoidcontent insertion efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent extracts and analyzes only the audio component from video content, separating it from the full video analysis process. This extraction of the audio track allows for efficient speech recognition-based attribute identification without the computational burden of analyzing all video frames, thereby improving productivity while maintaining adequate measurement precision for content insertion decisions

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces the mechanical image analysis process with a speech recognition system that processes audio data. Instead of using computationally intensive computer vision algorithms to analyze video frames, the system substitutes this with more efficient audio processing and natural language processing techniques, achieving the same goal of identifying video attributes with significantly reduced computational resources

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of information

If thousands to tens of thousands of video frames are analyzed, then comprehensive video context is obtained, but computing resources (compute time, data storage, communication bandwidth) are heavily consumed

Engineering Contradiction:
Improvevideo context information completenessVSAvoidcomputing resources consumed
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The patent extracts the audio track as a separate data stream from the video content, isolating the speech information that carries contextual meaning. This extraction allows the system to obtain comprehensive video context through audio analysis alone, without needing to process the vast quantity of visual frame data, thereby reducing the quantity of computing resources consumed while preventing loss of essential information

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of analyzing video frames to extract audio information (the conventional approach), the patent inverts the process by directly analyzing the audio track that is already embedded in or synchronized with the video. This inversion recognizes that speech recognition on audio data provides the same contextual information more efficiently, reducing resource consumption while maintaining information completeness

Inventive Principle:
Principle #13The other way round (Inversion)

3Productivity

If speech recognition techniques are applied to audio data, then the complexity and computational resources needed are dramatically reduced, but this approach must still generate contextually accurate tags for directed content insertion

Engineering Contradiction:
Improvemetadata generation speedVSAvoidtag accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces speech recognition technology as an intermediary between the audio data and the video attribute identification process. This intermediary converts spoken language in the audio track into text tags that capture the contextual meaning of the video content, achieving high tag accuracy while maintaining the productivity benefits of audio-based analysis instead of frame-by-frame image processing

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11490133B1Insertion of directed content into a video asset
Publication Date: 2022.11.01 AMAZON TECH INC
  • US11490133B1 patent drawing
  • US11490133B1 patent drawing
  • US11490133B1 patent drawing

AI summary

Technologies are provided for insertion of directed content into video assets based on analysis of audio data corresponding to the video assets. Some embodiments include a computing system that can receive audio data corresponding to a video asset. The computing system can generate, using the audio data, a time series of tags corresponding to speech conveyed by the video asset. The computing system can then generate a time-asset mapping between time and directed content using the time series and a correlation policy. The directed content can include digital media intended for a defined audience, for example. The time-asset mapping associates groups of directed content assets to respective specific times in the time series. The computing system can insert, using the mapping, a defined directed content asset from a group of directed content assets identified in the mapping.