Audio Track Generation for Image Sequences

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for generating audio to accompany digital image collections lack sensitivity to the visual, semantic, and emotive nature of the images, and are unable to effectively handle thematic groupings beyond sequential presentations.

Innovation Solution

A method that analyzes multimedia objects and their metadata to identify recurring thematic patterns, using techniques such as frequent item set mining, face detection, and sentiment analysis, to generate audio tracks that vary in instrumentation, tonality, and tempo, responsive to the visual and semantic content of the images, including the identification of animate and inanimate objects, scenes, and activities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If automated audio generation is applied to digital image collections, then the viewing experience is enhanced with audio components, but the sensitivity to visual, semantic, and emotive nature of images is lost

Engineering Contradiction:
Improveautomated audio generationVSAvoidloss of visual, semantic, and emotive content sensitivity
Core Design Contradiction:
Extent of automationVSLoss of information

Solution Approach 1:

The system performs preliminary analysis of image collections to extract visual, semantic, and emotive features before generating audio. This includes analyzing color palettes, detecting objects and scenes, identifying emotions in facial expressions, and extracting metadata to create a comprehensive representation of the image content that guides subsequent audio generation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces an intermediary processing layer that translates visual and semantic image properties into audio characteristics. This intermediary analysis stage extracts features such as dominant colors, emotional content, scene types, and temporal patterns, then maps these to corresponding audio parameters like tone, tempo, instrumentation, and dynamics

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If sequential presentation methods are used for image collections, then the presentation structure is simple, but thematic groupings and recurring patterns cannot be effectively identified

Engineering Contradiction:
Improvepresentation structureVSAvoidloss of thematic patterns and recurring themes
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The system segments the image collection into thematic groups based on shared visual and semantic characteristics. It identifies recurring patterns such as repeated objects, consistent color schemes, common emotions, and temporal sequences, then organizes images into coherent thematic segments that can be presented with appropriate audio accompaniment for each group

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transitions from simple sequential presentation to multi-dimensional organization by adding thematic and semantic dimensions to the presentation structure. It analyzes images along multiple axes including visual properties, emotional content, temporal relationships, and semantic categories, creating a rich organizational framework that goes beyond linear sequencing

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Adaptability or versatility

If video assets with audio tracks are used, then visual and audio content are combined, but the audio quality is inferior and snippets form only a fraction of the overall rendering

Engineering Contradiction:
Improvecombination of visual and audio contentVSAvoidaudio quality
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The system generates audio tracks autonomously based on the analysis of the image collection itself, without relying on pre-existing video audio. It uses the visual and semantic content of the images to synthesize custom audio that is specifically tailored to match the content, ensuring high audio quality and complete coverage of the presentation

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10699684B2Method for creating audio tracks for accompanying visual imagery
Publication Date: 2020.06.30 KODAK ALARIS LLC
  • US10699684B2 patent drawing
  • US10699684B2 patent drawing
  • US10699684B2 patent drawing

AI summary

Methods of creating one or more audio objects to accompany a sequence of multimedia objects are disclosed. According to one embodiment, the method includes using a processor to analyze the multimedia objects and corresponding recorded metadata to generate derived metadata. The method further receives a selection of one or more analysis tools that are configured to analyze the recorded and derived metadata. Next, a selected subset of multimedia objects are identified and sequenced, which will ultimately be coupled to and accompanied by one or more audio objects. Lastly, an embodiment of the present invention generates an audio track to accompany the selected subset of multimedia objects.