Automated Narrative Generation for Digital Media Collections

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face the challenge of manually describing large collections of digital media, such as photos and videos, which can be time-consuming and inefficient when sharing them, as existing systems lack the ability to automatically generate narrative descriptions.

Innovation Solution

A computer-implemented method and system that recognizes objects like faces and landmarks within media content, extracts metadata, and uses parameterized templates to automatically generate a narrative description of the media collection, enabling efficient content summarization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If users manually describe large collections of digital media, then the description can be customized and detailed, but the process becomes time-consuming and inefficient

Engineering Contradiction:
Improvecontent understandingVSAvoiddescription time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system enables automatic generation of narrative descriptions by having the media collection describe itself through object recognition and template-based compilation, eliminating the need for manual user input while preserving content understanding

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of manual description writing with an automated computer-based system that uses object recognition, metadata extraction, and template compilation to generate narratives automatically

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of information

If users review individual images to understand content, then complete understanding is achieved, but the process becomes inefficient for large collections

Engineering Contradiction:
Improvecontent comprehensionVSAvoidreview efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The system extracts essential content information from individual images through object recognition and metadata extraction, then compiles these extracted elements into a cohesive narrative description that provides complete understanding without requiring review of all individual images

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces an intermediary narrative description that mediates between individual images and user understanding, allowing users to comprehend content through the generated narrative rather than directly reviewing each image

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If automatic object recognition is implemented, then narrative generation becomes efficient, but system complexity increases

Engineering Contradiction:
Improvenarrative generation speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments the complex task of narrative generation into distinct modules: object recognition module, metadata extraction module, and narrative compilation module, allowing each to handle specific functions and reducing overall system complexity through functional decomposition

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10303756B2Creating a narrative description of media content and applications thereof
Publication Date: 2019.05.28 GOOGLE LLC
  • US10303756B2 patent drawing
  • US10303756B2 patent drawing
  • US10303756B2 patent drawing

AI summary

This invention relates to creating a narrative description of media content. In an embodiment, a computer-implemented method describes content of a group of images. The group of images includes a first image and a second image. A first object in the first image is recognized to determine a first content data. A second object in the second image is recognized to determine a second content data. Finally, a narrative description of the group of images is determined according to a parameterized template and the first and second content data.