Audio Media Identification via Speech-to-Text Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Radio listeners face difficulties in identifying and purchasing media objects, such as e-books, during audio broadcasts due to the need to manually recall identifying information, which can be challenging, especially in situations like moving vehicles, and existing technologies do not provide on-demand access to the desired media.

Innovation Solution

A media-identifying device with software that captures audio segments, converts them to text, and transmits the data over wireless or short-range connections to identify media objects, allowing for on-demand purchase and download of corresponding media items, such as e-books, using a network-connected device.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If radio listeners manually recall identifying information to purchase media objects, then they can potentially purchase the media, but the process becomes difficult and impractical in situations like moving vehicles

Engineering Contradiction:
Improveease of media identification and purchaseVSAvoidrecall accuracy of identifying information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent introduces a media identification device as an intermediary between the radio broadcast and the listener. This device captures audio segments from the radio, converts speech to text, and automatically transmits the text to identify the media object, eliminating the need for manual recall of identifying information by the listener.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system enables self-service by automatically capturing, processing, and transmitting the media identifying information without requiring active listener intervention. The device autonomously performs speech-to-text conversion and sends the text for media identification, allowing listeners to obtain media objects even when unable to manually record or recall information.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If listeners want on-demand access to media objects from radio broadcasts, then convenience is improved, but existing technologies do not provide this capability

Engineering Contradiction:
Improveon-demand media access capabilityVSAvoidsystem complexity for media identification
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The media identification device performs multiple functions: capturing audio segments from various sources (radio, other audio devices), converting speech to text, transmitting text for media identification, and enabling purchase and download of media objects. This multi-functional approach provides on-demand media access while consolidating complexity into a single device rather than requiring multiple separate systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If the device captures and transmits audio segments for media identification, then identification accuracy is improved, but data transmission requirements increase

Engineering Contradiction:
Improvemedia identification accuracyVSAvoiddata transmission volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system extracts only the essential identifying information from the audio broadcast by converting speech to text and transmitting only the text segment rather than the entire audio segment. This extraction approach maintains identification accuracy while significantly reducing the quantity of data that needs to be transmitted compared to sending raw audio segments.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS8571863B1Apparatus and methods for identifying a media object from an audio play out
Publication Date: 2013.10.29 INTELLECTUAL VENTURES FUND 79 LLC
  • US8571863B1 patent drawing
  • US8571863B1 patent drawing
  • US8571863B1 patent drawing

AI summary

In one example, a device captures a segment of audio played out over an audio source in response to a control signal from a user interface of the device. The device causes any speech of the captured segment to be converted into a text segment using a local or remote speech-to-text component. The device causes the text segment to be provided to a remote network device for analysis. In response to the providing, the device receives back a tag that identifies an electronic media object. The electronic media object may correspond to, for example, an electronic book, and the device may use the tag to purchase and download the book. The device may be configured to convert text of the downloaded electronic book to synthesized speech, and then cause the synthesized speech to be played out over a local or remote speaker.