Contextual Ad Serving via Real-Time Audio Transcription

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing ad serving systems for audio and video content are limited by requiring search queries or metadata that may not accurately match the content, leading to irrelevant advertisements, and they fail to continuously provide contextually relevant ads without user input.

Innovation Solution

A system that continuously monitors media streams, uses voice recognition to transcribe audio or video content into text, analyzes this text to identify relevant advertisements, and displays them as clickable links next to the media playback, without requiring user input or predefined keywords.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If ad serving systems use metadata or search queries to match advertisements with audio/video content, then advertisements can be displayed without user input, but the relevance of advertisements to actual content deteriorates because metadata is limited and may not accurately reflect content

Engineering Contradiction:
Improveautomatic ad serving without user inputVSAvoidadvertising relevance to content
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent replaces manual metadata creation and keyword-based matching with automatic speech recognition technology that transcribes audio content into text. This substitution enables the system to directly analyze the actual spoken content rather than relying on imperfect metadata, thereby improving advertising relevance while maintaining automatic operation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If ad serving systems continuously monitor and analyze media streams in real-time, then advertising relevance to content is improved, but system complexity and processing requirements worsen

Engineering Contradiction:
Improveadvertising relevance to contentVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements preliminary speech recognition processing that transcribes audio content into text and stores it for later retrieval. By performing the complex speech recognition and text analysis in advance rather than in real-time during ad serving, the system reduces processing complexity while maintaining high advertising relevance when ads are displayed.

Inventive Principle:
Principle #10Preliminary action

3Ease of manufacture

If ad serving systems rely on hand-entered metadata to describe content, then implementation is simpler, but the quantity and accuracy of content information worsens leading to failed content matching

Engineering Contradiction:
Improvesystem implementation simplicityVSAvoidcontent information accuracy
Core Design Contradiction:
Ease of manufactureVSLoss of information

Solution Approach 1:

The patent enables the system to automatically generate its own content descriptions through speech recognition technology that transcribes audio into text without requiring manual metadata entry. This self-service approach eliminates the need for hand-entered metadata while providing accurate, comprehensive content information for advertising matching.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS9087331B2Contextual advertising for video and audio media
Publication Date: 2015.07.21 TVEYES
  • US9087331B2 patent drawing
  • US9087331B2 patent drawing
  • US9087331B2 patent drawing

AI summary

A system and method for serving contextually relevant advertisements is provided, including monitoring a media stream to indentify an audio or video asset, extracting and storing corresponding text from the asset, retrieving stored text when the asset is selected by a user, analyzing the text and identifying relevant advertisements that are then displayed to a user as a clickable text next to the playing video or audio asset. In further embodiments, the method may include the steps of retrieving a variable length portion of the text corresponding to the portion of the asset being played by the user, analyzing the portion of the text to identify advertisements relevant to the corresponding portion of the asset, displaying the advertisements during the playback of the portion of the asset, and then repeating the steps until the playback of the whole asset is completed.