Directed Content Insertion via Audio Tag Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The insertion of directed content into video assets is computationally intensive due to the analysis of large amounts of image data, requiring significant computing resources and time, which is inefficient compared to analyzing audio data using speech recognition techniques.
Innovation Solution
A computing system analyzes audio data to generate a time series of tags representing keywords or keyphrases, which are then used to create a mapping between time and directed content assets, allowing for the insertion of contextually meaningful content into video assets by associating specific content with specific timestamps.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If image analysis of video frames is used to identify video attributes, then contextually meaningful directed content can be inserted into the video, but the computational resources and time required increase significantly
Solution Approach 1:
The patent extracts and analyzes only the audio component from video content, separating it from the full video analysis process. This extraction of the audio track allows for efficient speech recognition-based attribute identification without the computational burden of analyzing all video frames, thereby improving productivity while maintaining adequate measurement precision for content insertion decisions
Solution Approach 2:
The patent replaces the mechanical image analysis process with a speech recognition system that processes audio data. Instead of using computationally intensive computer vision algorithms to analyze video frames, the system substitutes this with more efficient audio processing and natural language processing techniques, achieving the same goal of identifying video attributes with significantly reduced computational resources
2Loss of information
If thousands to tens of thousands of video frames are analyzed, then comprehensive video context is obtained, but computing resources (compute time, data storage, communication bandwidth) are heavily consumed
Solution Approach 1:
The patent extracts the audio track as a separate data stream from the video content, isolating the speech information that carries contextual meaning. This extraction allows the system to obtain comprehensive video context through audio analysis alone, without needing to process the vast quantity of visual frame data, thereby reducing the quantity of computing resources consumed while preventing loss of essential information
Solution Approach 2:
Instead of analyzing video frames to extract audio information (the conventional approach), the patent inverts the process by directly analyzing the audio track that is already embedded in or synchronized with the video. This inversion recognizes that speech recognition on audio data provides the same contextual information more efficiently, reducing resource consumption while maintaining information completeness
3Productivity
If speech recognition techniques are applied to audio data, then the complexity and computational resources needed are dramatically reduced, but this approach must still generate contextually accurate tags for directed content insertion
Solution Approach 1:
The patent introduces speech recognition technology as an intermediary between the audio data and the video attribute identification process. This intermediary converts spoken language in the audio track into text tags that capture the contextual meaning of the video content, achieving high tag accuracy while maintaining the productivity benefits of audio-based analysis instead of frame-by-frame image processing
Data Source
AI summary
Technologies are provided for insertion of directed content into video assets based on analysis of audio data corresponding to the video assets. Some embodiments include a computing system that can receive audio data corresponding to a video asset. The computing system can generate, using the audio data, a time series of tags corresponding to speech conveyed by the video asset. The computing system can then generate a time-asset mapping between time and directed content using the time series and a correlation policy. The directed content can include digital media intended for a defined audience, for example. The time-asset mapping associates groups of directed content assets to respective specific times in the time series. The computing system can insert, using the mapping, a defined directed content asset from a group of directed content assets identified in the mapping.


