Contextual Ad Stitching Through Multimodal Video Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing online advertising methods are inefficient and disruptive, leading to ad blocking and inadequate revenue replacement, with a lack of relevance in content-based targeting, and there is a need for more effective contextual advertising systems.

Innovation Solution

A system that utilizes multimodal metadata extraction and deep learning to understand video content on a scene-by-scene basis, enabling contextual advertising and seamless integration of ads based on rich metadata, using generative AI to modify content and provide dynamically customized creatives.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional display advertising is used, then advertisers can reach audiences, but the ads are disruptive and lead to ad blocking

Engineering Contradiction:
Improvead delivery effectivenessVSAvoiduser disruption and ad blocking
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent merges advertising content with entertainment content by seamlessly integrating ads into video streams. Ads are not displayed as separate elements but are woven into the content flow, making them indistinguishable from regular video content. This eliminates the disruptive nature of traditional ads while maintaining advertising effectiveness.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system uses an intermediary processing layer that includes AI-driven content analysis, ad placement algorithms, and video processing modules. This intermediary layer analyzes content context, selects appropriate ad placements, and modifies video streams to integrate ads naturally, acting as a mediator between advertisers and content consumers.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If contextual advertising based on content is implemented, then ad relevance improves, but the system complexity increases

Engineering Contradiction:
Improvecontextual ad targetingVSAvoidsystem architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system segments the video content into discrete units (scenes, shots, or frames) and analyzes each segment independently to determine appropriate ad placements. This segmentation allows the complex contextual analysis to be broken down into manageable processing steps, reducing overall system complexity while maintaining high adaptability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary content analysis and indexing before ad placement is needed. By pre-processing video content to extract metadata, identify contexts, and build knowledge graphs, the system reduces the computational burden during real-time ad selection and placement, thereby managing complexity while maintaining versatility.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If deep learning and multimodal metadata extraction are used to understand video content, then contextual understanding improves, but computational resources and processing time increase

Engineering Contradiction:
Improvecontent understanding accuracyVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

Instead of uniformly processing the entire video stream with high computational intensity, the system applies local quality analysis by focusing deep learning models only on specific segments or frames where ad placements are most effective. This selective processing maintains high content understanding accuracy for critical moments while reducing overall computational resource consumption.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system uses partial action by extracting only the most relevant metadata and contextual information needed for ad placement decisions, rather than completely analyzing every aspect of the video content. This approach achieves sufficient content understanding accuracy for effective advertising while significantly reducing computational resource requirements.

Inventive Principle:
Principle #16Partial or excessive action

4Productivity

If seamless ad integration is achieved, then user engagement improves, but the difficulty of detecting and measuring ad effectiveness increases

Engineering Contradiction:
Improveuser engagementVSAvoidad effectiveness measurement
Core Design Contradiction:
ProductivityVSDifficulty of detecting and measuring

Solution Approach 1:

The system incorporates feedback mechanisms that track user interactions with integrated ads, measure engagement metrics, and provide real-time data back to the advertising platform. This feedback loop enables continuous optimization of ad placements while providing measurable insights into ad effectiveness despite the seamless integration.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system uses subtle visual differentiation techniques, such as slight color or brightness changes in integrated ads compared to surrounding content, to maintain seamlessness while providing detectable cues for measurement. These subtle changes are imperceptible to users but can be detected by analysis algorithms to measure ad effectiveness.

Inventive Principle:
Principle #32Color changes

Data Source

PatentUS20250265620A1System for seamlessly stitching fully contextual ads to content for immersive advertising
Publication Date: 2025.08.21 ANOKI INC
  • US20250265620A1 patent drawing
  • US20250265620A1 patent drawing
  • US20250265620A1 patent drawing

AI summary

A system for contextual modification of content based on multimodal extraction of metadata from the content, wherein the metadata is extracted by processing one or more scenes in the content to extract metadata corresponding to multiple extraction modes, and an embedding model for each extraction mode wherein an aggregated embedding model responsive to the extracted metadata for each mode formulates an aggregated embedding. A process controller may include an embedding extractor responsive to a control input. The control input may specify one or more features appearing in the content defining a content modification opportunity. The embedding extractor may include an embedding model coordinated with the embedding model for one or more of the embedding modes to generate an opportunity embedding in the form of a vector. A vector comparison processor determines the distance between the opportunity embedding and the aggregated embedding, wherein the embeddings are in the form of vectors. The process controller is responsive to the vector comparison processor to generate edit control instructions indicating a modification of the content upon detection of the content modification opportunity. A content editor is responsive to the edit control instructions to modify the content. The content editor uses generative AI techniques to modify the content by replacing an element appearing within the content during the content modification opportunity with an element correlated with the element appearing in said content.