Text-to-Speech Converter for Broadcast Video Description

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Providing video-descriptive services for visually impaired individuals is cost- and bandwidth-prohibitive due to the significant amount of space or bandwidth required for an additional audio track.

Innovation Solution

A system and method that generates a data signal with text description data corresponding to a video signal, converts it to a first audio signal, and plays it as an audible signal at a display device, reducing bandwidth requirements by using a text-to-speech converter within the receiving device.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If an additional audio track is provided for video-descriptive service, then the service quality for visually impaired users is improved, but the bandwidth requirement increases significantly

Engineering Contradiction:
Improveservice qualityVSAvoidbandwidth
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential descriptive information from the video content and transmits it as compact text data rather than full audio. The receiving device then synthesizes speech from this extracted text, significantly reducing the bandwidth required while maintaining the core functionality of describing video content for visually impaired users.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of transmitting audio data and converting it to speech at the broadcast side, the patent inverts the process by transmitting text data and performing text-to-speech conversion at the receiving device. This inversion dramatically reduces the amount of data that needs to be transmitted over the broadcast network.

Inventive Principle:
Principle #13The other way round (Inversion)

2Adaptability or versatility

If an additional audio track is provided for video-descriptive service, then the accessibility for visually impaired users is improved, but the cost becomes prohibitive

Engineering Contradiction:
ImproveaccessibilityVSAvoidcost
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The patent extracts only the essential descriptive information from the video content and transmits it as compact text data rather than full audio. The receiving device then synthesizes speech from this extracted text, significantly reducing the bandwidth required while maintaining the core functionality of describing video content for visually impaired users.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the data format parameter from audio (64 kbps) to text, which requires significantly fewer bits to transmit. This parameter change in the data representation fundamentally reduces both bandwidth requirements and associated costs while preserving accessibility functionality.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If text-to-speech conversion is performed at the receiving device, then bandwidth efficiency is improved, but device complexity increases

Engineering Contradiction:
ImprovebandwidthVSAvoiddevice complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent enables the receiving device to perform text-to-speech conversion using its own built-in capabilities. This self-service approach allows the device to generate the audio output locally from the transmitted text data, eliminating the need for separate external text-to-speech hardware while reducing bandwidth requirements.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS8804035B1Method and system for communicating descriptive data in a television broadcast system
Publication Date: 2014.08.12 DIRECTV LLC
  • US8804035B1 patent drawing
  • US8804035B1 patent drawing
  • US8804035B1 patent drawing

AI summary

A method and system for communicating text descriptive data has a receiving device that receives a data signal having text description data corresponding to a description of a video signal. A text-to-speech converter associated with the receiving device converts the text description data to a first audio signal. A display device in communication with the text-to-speech converter converts the first audio signal associated with the receiving device to an audible signal.