Multimodal Content Channel Programming for Contextual Ad Placement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing online advertising methods are inefficient and disruptive, lacking relevance to the content being consumed, leading to decreased user engagement and revenue loss for publishers.
Innovation Solution
A system utilizing multimodal metadata extraction and deep learning to understand video content on a scene-by-scene basis, enabling contextual advertising and content modification through enriched metadata, allowing for relevant ad placement and dynamic content customization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If traditional online advertising methods are used, then advertisements can be delivered to users, but user engagement decreases and the ads become disruptive
Solution Approach 1:
The patent applies local quality by making advertisements contextually relevant to specific content segments rather than uniformly displaying ads throughout. The system analyzes video content at the scene level and places ads only where contextually appropriate, transforming ads from disruptive elements to integrated content enhancements.
Solution Approach 2:
The patent introduces an intermediary system that includes a processor and database for storing and analyzing content metadata. This intermediary layer mediates between the content delivery system and advertising system, enabling contextual matching without direct ad insertion that would disrupt content flow.
2Measurement precision
If advertisements are targeted based on user data collection, then ad relevance increases, but user privacy concerns increase
Solution Approach 1:
The patent extracts advertising targeting capabilities from direct user data collection and relocates them to content metadata analysis. Instead of tracking user behavior across websites, the system extracts contextual information from content itself (scene descriptions, objects, actions) to determine appropriate ad placements, eliminating the need for invasive user profiling.
Solution Approach 2:
The patent creates a copy of user interests and preferences through content metadata rather than directly collecting user data. The metadata serves as a proxy that reflects what content users engage with, providing targeting information without requiring actual user tracking or personal data collection.
3Measurement precision
If contextual advertising is implemented, then ad relevance to content improves, but system complexity increases
Solution Approach 1:
The patent applies preliminary action by pre-processing and storing content metadata in a database before ad placement is needed. The system analyzes and indexes content characteristics in advance, so when ad placement is required, the relevant contextual information is already prepared and readily accessible, reducing real-time computational complexity.
Solution Approach 2:
The patent segments content into discrete analyzable units (scenes) with associated metadata, making the complex task of contextual advertising manageable. By breaking content into scenes with standardized metadata fields (objects, actions, locations), the system can process and match ads systematically without overwhelming complexity.
4Measurement precision
If deep learning models are used for content analysis, then understanding of video content improves, but computational resources required increase
Solution Approach 1:
The patent applies preliminary action by pre-computing and storing deep learning embeddings and metadata for content analysis before ad placement decisions are needed. The computationally intensive deep learning models process content in advance, converting videos into structured metadata representations that can be quickly queried without requiring heavy computational resources at ad placement time.
Solution Approach 2:
The patent creates a compressed representation (embedding) of complex video content through deep learning, storing this simplified copy in the database. Instead of processing the full video content again during ad placement, the system queries and compares these pre-computed embeddings, significantly reducing real-time computational resource requirements while maintaining content understanding accuracy.
Data Source
AI summary
A system for programming content channels based on multimodal extraction of metadata from content, wherein the metadata is extracted by processing one or more scenes in the content to extract metadata corresponding to multiple extraction modes, and an embedding model for each extraction mode wherein an aggregated embedding model responsive to the extracted metadata for each mode formulates an aggregated embedding. The system includes a consumer information database containing information regarding consumer consumption of content including at least an identification of content consumed. Programming logic based at least in part on comparing an embedding of metadata extracted by multimodal extraction of content consumed by the consumer to the aggregated embeddings including a vector comparison processor for determining the distance between the embedding generated based upon content consumed and the aggregated embedding is provided. The programming logic may use the distance between the embedding generated based on content consumed and the aggregated embedding to generate a program guide tailored to viewer preferences.


