Multimodal Content Channel Programming for Contextual Ad Placement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing online advertising methods are inefficient and disruptive, lacking relevance to the content being consumed, leading to decreased user engagement and revenue loss for publishers.

Innovation Solution

A system utilizing multimodal metadata extraction and deep learning to understand video content on a scene-by-scene basis, enabling contextual advertising and content modification through enriched metadata, allowing for relevant ad placement and dynamic content customization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If traditional online advertising methods are used, then advertisements can be delivered to users, but user engagement decreases and the ads become disruptive

Engineering Contradiction:
Improveuser engagementVSAvoidad disruption
Core Design Contradiction:
Ease of operationVSObject-affected harmful factors

Solution Approach 1:

The patent applies local quality by making advertisements contextually relevant to specific content segments rather than uniformly displaying ads throughout. The system analyzes video content at the scene level and places ads only where contextually appropriate, transforming ads from disruptive elements to integrated content enhancements.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent introduces an intermediary system that includes a processor and database for storing and analyzing content metadata. This intermediary layer mediates between the content delivery system and advertising system, enabling contextual matching without direct ad insertion that would disrupt content flow.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If advertisements are targeted based on user data collection, then ad relevance increases, but user privacy concerns increase

Engineering Contradiction:
Improvead targeting precisionVSAvoiduser privacy
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent extracts advertising targeting capabilities from direct user data collection and relocates them to content metadata analysis. Instead of tracking user behavior across websites, the system extracts contextual information from content itself (scene descriptions, objects, actions) to determine appropriate ad placements, eliminating the need for invasive user profiling.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates a copy of user interests and preferences through content metadata rather than directly collecting user data. The metadata serves as a proxy that reflects what content users engage with, providing targeting information without requiring actual user tracking or personal data collection.

Inventive Principle:
Principle #26Copying

3Measurement precision

If contextual advertising is implemented, then ad relevance to content improves, but system complexity increases

Engineering Contradiction:
Improvecontextual relevanceVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-processing and storing content metadata in a database before ad placement is needed. The system analyzes and indexes content characteristics in advance, so when ad placement is required, the relevant contextual information is already prepared and readily accessible, reducing real-time computational complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments content into discrete analyzable units (scenes) with associated metadata, making the complex task of contextual advertising manageable. By breaking content into scenes with standardized metadata fields (objects, actions, locations), the system can process and match ads systematically without overwhelming complexity.

Inventive Principle:
Principle #1Segmentation

4Measurement precision

If deep learning models are used for content analysis, then understanding of video content improves, but computational resources required increase

Engineering Contradiction:
Improvecontent understandingVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent applies preliminary action by pre-computing and storing deep learning embeddings and metadata for content analysis before ad placement decisions are needed. The computationally intensive deep learning models process content in advance, converting videos into structured metadata representations that can be quickly queried without requiring heavy computational resources at ad placement time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a compressed representation (embedding) of complex video content through deep learning, storing this simplified copy in the database. Instead of processing the full video content again during ad placement, the system queries and compares these pre-computed embeddings, significantly reducing real-time computational resource requirements while maintaining content understanding accuracy.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250267325A1System for programming content channels
Publication Date: 2025.08.21 ANOKI INC
  • US20250267325A1 patent drawing
  • US20250267325A1 patent drawing
  • US20250267325A1 patent drawing

AI summary

A system for programming content channels based on multimodal extraction of metadata from content, wherein the metadata is extracted by processing one or more scenes in the content to extract metadata corresponding to multiple extraction modes, and an embedding model for each extraction mode wherein an aggregated embedding model responsive to the extracted metadata for each mode formulates an aggregated embedding. The system includes a consumer information database containing information regarding consumer consumption of content including at least an identification of content consumed. Programming logic based at least in part on comparing an embedding of metadata extracted by multimodal extraction of content consumed by the consumer to the aggregated embeddings including a vector comparison processor for determining the distance between the embedding generated based upon content consumed and the aggregated embedding is provided. The programming logic may use the distance between the embedding generated based on content consumed and the aggregated embedding to generate a program guide tailored to viewer preferences.