Automated Video Description Generation Using Neural Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Generating textual descriptions for video content is a time-consuming manual process, and not all users can quickly read or access textual information presented alongside video, highlighting the need for automated generation and presentation of such descriptions.

Innovation Solution

The system automatically generates textual descriptions of video content using neural networks and machine learning algorithms, analyzing individual frames and audio segments to identify objects, actions, and sounds, and presents them in both text and audio formats, with the option to refine descriptions using knowledge graph data for accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual generation of textual descriptions is used, then accuracy and quality of descriptions can be maintained, but time consumption and labor effort increase significantly

Engineering Contradiction:
Improvedescription generation speedVSAvoidtime for generating descriptions
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces the manual mechanical process of watching video and typing descriptions with an automated system using neural networks and machine learning algorithms. The system processes video frames, audio segments, and metadata automatically to generate textual descriptions, eliminating the need for human operators to manually view and describe each video segment.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables videos to describe themselves by automatically analyzing their own content (frames, audio, metadata) and generating textual descriptions without external human intervention. The automated generation system processes the video content and produces descriptions independently, making the description generation process self-sufficient.

Inventive Principle:
Principle #25Self-service

2Ease of operation

If textual descriptions are presented alongside video content, then user experience is improved, but accessibility for visually impaired users and reading speed limitations remain issues

Engineering Contradiction:
Improveuser accessibilityVSAvoidaccessibility for different user needs
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The system provides multiple output formats for video descriptions - both textual and audible - making it universally accessible to different user needs. Visually impaired users can access descriptions through audio, while users with reading speed limitations can also benefit from audible presentation, thereby serving multiple user groups with a single system.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces an audible description track as an intermediary medium between the video content and the user. Instead of requiring direct visual reading of text, the system converts textual descriptions into audible format that can be played alongside the video, serving as a mediator that makes content accessible to users who cannot efficiently process visual text.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If automated generation systems are implemented, then productivity increases, but system complexity and computational resources required increase

Engineering Contradiction:
Improvedescription generation efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system divides the complex task of video description generation into separate processing modules: video frame analysis, audio segment analysis, metadata processing, and description generation. Each module handles a specific aspect of the input data independently, allowing for specialized processing and easier maintenance of the overall system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent combines multiple input sources (video frames, audio segments, metadata) into a unified processing pipeline that generates comprehensive textual descriptions. By merging these different data types and processing them together through the neural network system, the approach achieves more accurate and contextually rich descriptions than processing each source separately.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS10999566B1Automated generation and presentation of textual descriptions of video content
Publication Date: 2021.05.04 AMAZON TECH INC
  • US10999566B1 patent drawing
  • US10999566B1 patent drawing
  • US10999566B1 patent drawing

AI summary

Systems, methods, and computer-readable media are disclosed for systems and methods for automated generation of textual descriptions of video content. Example methods may include determining, by one or more computer processors coupled to memory, a first segment of video content, the first segment including a first set of frames and first audio content, determining, using a first neural network, a first action that occurs in the first set of frames, and determining a first sound present in the first audio content. Some methods may include generating a vector representing the first action and the first sound, and generating, using a second neural network and the vector, a first textual description of the first segment, where the first textual description includes words that describe events of the first segment.