AI Audio Description Generation for Live Events

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating audio descriptions of live events are either delayed or require human intervention, making them time-consuming and costly, and often include non-relevant details.

Innovation Solution

Automatically generating audio descriptions using machine learning models that analyze video frames to identify relevant visual elements and produce semantic representations, which are then converted into audio descriptions in real time, without the need for human captioning.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If human captioning is used to generate audio descriptions, then accuracy and relevance of descriptions improve, but time consumption and cost increase

Engineering Contradiction:
Improveaccuracy of audio descriptionsVSAvoidtime consumption for generation
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent uses machine learning models to create automated copies of the human captioning process. Instead of manually analyzing video frames and generating audio descriptions, the system employs trained models that replicate human descriptive capabilities, producing accurate audio descriptions at machine speed without requiring human intervention for each event.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical human captioning process with an automated machine learning system. The mechanical system of human observers watching videos and writing descriptions is substituted with electronic processing through trained machine learning models that automatically generate audio descriptions from video input, eliminating time consumption and cost while maintaining accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If human captioning is used to generate audio descriptions, then accuracy and relevance of descriptions improve, but cost increases

Engineering Contradiction:
Improveaccuracy of audio descriptionsVSAvoidcost of generation
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent uses machine learning models to create automated copies of the human captioning process. Instead of manually analyzing video frames and generating audio descriptions, the system employs trained models that replicate human descriptive capabilities, producing accurate audio descriptions at machine speed without requiring human intervention for each event.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical human captioning process with an automated machine learning system. The mechanical system of human observers watching videos and writing descriptions is substituted with electronic processing through trained machine learning models that automatically generate audio descriptions from video input, eliminating time consumption and cost while maintaining accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Quantity of substance

If all visual elements are described in audio descriptions, then completeness improves, but relevance and consumer experience worsen due to non-relevant details

Engineering Contradiction:
Improvecompleteness of descriptionsVSAvoidrelevance to consumer experience
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent applies local quality by differentiating the treatment of various visual elements in the video. Instead of uniformly describing all elements, the machine learning models identify and prioritize visually important elements (such as players, ball, goals) while filtering out less relevant background elements. This selective description approach ensures completeness of important events while maintaining relevance to consumer experience.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent uses partial action by selectively describing only the most relevant visual elements rather than attempting to describe everything in the video. The machine learning models identify key objects and actions that matter to the event narrative, providing sufficient description for consumer understanding without including excessive non-relevant details that would degrade the experience.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11736775B1Artificial intelligence audio descriptions for live events
Publication Date: 2023.08.22 AMAZON TECH INC
  • US11736775B1 patent drawing
  • US11736775B1 patent drawing
  • US11736775B1 patent drawing

AI summary

Methods and apparatus are described for generating audio descriptions of live events in near real time. Visual elements are identified in video frames. Semantic representations of the visual elements are determined and used to generate an audio description. The audio description is provided to a client device for playback during the live event as an alternative to the original audio content.