Audio-Visual Slideshow Generation Using Audio Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video summarization methods are inadequate for consumer-quality videos due to diverse content characteristics, challenging conditions like uneven illumination and background noise, and difficulty in assessing user satisfaction, making it hard to identify specific objects or events and generate effective summaries.

Innovation Solution

A method for producing an audio-visual slideshow that automatically analyzes the audio soundtrack to segment and classify audio frames, selects key image frames from the video track, and combines them to form a synchronized audio-visual summary, considering audio diversity, visual diversity, facial quality, and overall image quality, using techniques like Bayesian Information Criterion for audio segmentation and K-means clustering for image selection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If video summarization methods are designed for high quality professional videos, then summarization accuracy is improved, but applicability to consumer-quality videos deteriorates

Engineering Contradiction:
Improvesummarization accuracyVSAvoidapplicability to consumer videos
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent changes the parameter of quality expectation from high-quality professional videos to consumer-quality videos with uncontrolled conditions. The system adapts to varying illumination, noise levels, and camera stability by using robust audio-visual feature extraction that does not require ideal recording conditions.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent creates a universal video summarization system that works across diverse video types (professional and consumer videos). The audio-visual slideshow generation method is designed to handle multiple video quality levels and content types, making it applicable to general consumer video collections rather than specific professional domains.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If object/event detection methods are used for video summarization, then semantic accuracy is improved, but robustness to background noise and clutter deteriorates

Engineering Contradiction:
Improvesemantic accuracyVSAvoidrobustness to noise
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent segments the video into audio segments and visual segments separately, then combines them to form an audio-visual slideshow. This segmentation allows the system to process audio and visual information independently, avoiding the need for complex object detection in noisy visual environments while maintaining semantic accuracy through audio cues.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses audio as an intermediary to bridge the gap between visual content and semantic meaning. Instead of directly detecting objects in cluttered visual scenes, the system extracts semantic information from audio and matches it with representative visual frames, providing robustness to visual noise and clutter.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If audio soundtrack analysis is performed in noisy consumer videos, then audio segment identification is improved, but accuracy deteriorates due to background noise

Engineering Contradiction:
Improveaudio segmentation capabilityVSAvoidaudio segment accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent merges audio analysis with visual analysis to create an audio-visual slideshow. By combining both modalities, the system compensates for weaknesses in individual channels - visual information helps disambiguate audio segments in noisy conditions, while audio provides temporal structure to visual content.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS10134440B2Video summarization using audio and visual cues
Publication Date: 2018.11.20 KODAK ALARIS LLC
  • US10134440B2 patent drawing
  • US10134440B2 patent drawing
  • US10134440B2 patent drawing

AI summary

A method for producing an audio-visual slideshow for a video sequence having an audio soundtrack and a corresponding video track including a time sequence of image frames, comprising: segmenting the audio soundtrack into a plurality of audio segments; subdividing the audio segments into a sequence of audio frames; determining a corresponding audio classification for each audio frame; automatically selecting a subset of the audio segments responsive to the audio classification for the corresponding audio frames; for each of the selected audio segments automatically analyzing the corresponding image frames to select one or more key image frames; merging the selected audio segments to form an audio summary; forming an audio-visual slideshow by combining the selected key frames with the audio summary, wherein the selected key frames are displayed synchronously with their corresponding audio segment; and storing the audio-visual slideshow in a processor-accessible storage memory.