Contextual Advertising Using Multimodal Video Metadata and Generative AI
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing online advertising methods are inefficient and disruptive, leading to ad blocking and revenue loss, and lack relevance to the content being consumed, resulting in lower engagement and monetization.
Innovation Solution
A system that utilizes multimodal metadata extraction and generative AI to understand video content on a scene-by-scene basis, enabling contextual advertising by dynamically selecting and generating ads based on rich metadata, including viewer information, to enhance relevance and viewer engagement.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If traditional online advertising methods are used, then ad delivery is achieved, but advertising relevance to content and viewer engagement deteriorates
Solution Approach 1:
The system extracts metadata from video content in advance and creates contextual profiles before ad delivery. This preliminary analysis of content themes, objects, and scenes enables the system to pre-identify suitable advertising opportunities, ensuring high relevance when ads are ultimately delivered without compromising viewer engagement
Solution Approach 2:
The system applies different advertising strategies to different segments of video content based on local characteristics. By analyzing specific scenes, objects, and themes within the video, the system tailors ad selection to match the particular context of each segment, thereby improving advertising relevance and maintaining viewer engagement throughout the viewing experience
2Adaptability or versatility
If programmatic advertising with cookies is used, then user targeting is achieved, but viewer disruption and ad blocking increase
Solution Approach 1:
The system introduces contextual content metadata as an intermediary between user targeting and ad delivery. Instead of directly targeting users based on cookies, the system uses content context as a mediator to select ads that are relevant to what the viewer is watching, thereby maintaining targeting effectiveness while reducing disruption and ad blocking
Solution Approach 2:
The system converts the potential harm of disruptive advertising into benefit by using content context to guide ad placement. Ads are positioned within the video content flow in a way that complements rather than interrupts the viewing experience, transforming what would be disruptive interruptions into relevant, contextually-appropriate content extensions
3Productivity
If display advertising is used, then visual ad delivery is achieved, but advertising effectiveness and ROI deteriorate
Solution Approach 1:
The system changes the parameters of ad delivery by incorporating multiple content dimensions including themes, objects, scenes, and viewer context. This multi-parameter approach to ad selection and customization significantly improves advertising effectiveness and ROI while maintaining efficient ad delivery through automated processing
Data Source
AI summary
A system for contextual modification of media content based on multimodal extraction of metadata from the media, wherein the metadata is extracted by processing one or more scenes in the media to extract metadata corresponding to multiple extraction modes, and an embedding model for each extraction mode wherein an aggregated embedding model responsive to the extracted metadata for each mode formulates an aggregated embedding. A process controller may include an embedding extractor responsive to a control input. The control input may specify one or more features appearing in the content defining a media modification opportunity. The embedding extractor may include an embedding model coordinated with the embedding model for one or more of the embedding modes to generate an opportunity embedding in the form of a vector. A vector comparison processor determines the distance between the opportunity embedding and the aggregated embedding, wherein the embeddings are in the form of vectors. The process controller is responsive to the vector comparison processor to generate edit control instructions indicating a modification of the media upon detection of the media modification opportunity. A media content editor is responsive to the edit control instructions to modify the media. The media content editor uses generative AI techniques to modify the media content.


