Scene-Level Video Search Using Multimodal Embeddings for Contextual Ads
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing online advertising methods are inefficient and disruptive, leading to ad blocking and revenue loss, while lacking in targeted and contextual ad delivery.
Innovation Solution
A system utilizing AI techniques to process video content on a scene-by-scene and frame-by-frame basis, extracting multimodal metadata for indexing and searching, allowing for free-form, contextual, and detailed video searches.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional online advertising methods are used, then ad delivery can be implemented, but ad effectiveness is reduced due to disruption and ad blocking
Solution Approach 1:
The video content is segmented into individual scenes and frames, allowing ads to be placed at specific contextual moments rather than interrupting the entire video. This scene-by-scene indexing enables precise control over ad placement timing and context, reducing user disruption while maintaining delivery effectiveness
Solution Approach 2:
Different parts of the video content receive different treatments based on their contextual relevance. Ads are placed selectively in scenes that match the ad content, creating locally optimized ad experiences rather than uniform ad delivery. This allows high-value ad placements in relevant scenes while avoiding disruption in unrelated portions
2Adaptability or versatility
If contextual advertising is implemented, then ad targeting improves, but system complexity increases due to content analysis requirements
Solution Approach 1:
The system performs preliminary analysis of video content during ingestion, creating scene indexes and metadata beforehand. This pre-processing stores contextual information about each scene (objects, actions, settings) in an accessible format, so that when ads need to be matched, the system queries pre-existing indexes rather than analyzing raw video in real-time, significantly reducing operational complexity
Solution Approach 2:
The patent introduces an intermediary layer between raw video content and ad delivery - a scene indexing system that translates video content into structured metadata and embeddings. This intermediary representation simplifies subsequent ad matching operations by providing a standardized interface between content and advertising systems, reducing the complexity of direct video-ad analysis
3Measurement precision
If detailed content indexing is performed, then search precision improves, but processing time and computational resources increase
Solution Approach 1:
The system creates compressed representations (embeddings) of video scenes that capture essential contextual information in a simplified numerical form. These embeddings are mathematical approximations that preserve semantic meaning while occupying minimal storage and enabling rapid comparison operations, achieving high search precision without processing the full complexity of original video data
Solution Approach 2:
The patent transforms video content from its original complex format into different parameter spaces - converting visual and audio data into numerical embeddings that can be efficiently searched and compared. This parameter transformation allows precise content matching through mathematical operations on simplified representations, dramatically reducing processing time while maintaining search accuracy
Data Source
AI summary
A system for contextual searching of content based on a content query. The content is related to multimodal metadata extracted from the content. A search vector is created using a compatible metadata extractor and the distance between said search vector and an embedding is indicative matching the content to the search content. The search content may be broken out by scene and the embedding for each scene may be established independently to serve as search terms, which may be combined trough logical combinations.


