Summary Image Generation Using Music Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image processing technologies fail to effectively generate summary images from long, monotonous sequences of images captured by wearable or action cameras, as they only detect noteworthy sections without setting adoptable sections within the summary image.

Innovation Solution

An information processing method and apparatus that analyzes input images based on music and scene information to set adoptable sections within the summary image, using an image analysis unit and extraction unit to select and position unit images according to music analysis and editing information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If continuous image-capturing is performed for a long time, then more images are captured, but composition becomes monotonous and images are difficult to enjoy

Engineering Contradiction:
Improvenumber of captured imagesVSAvoidenjoyability of images
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent extracts noteworthy sections from the continuous captured images and uses only those sections to generate a summary image. This extraction process removes the monotonous parts while preserving the interesting moments, solving the problem of enjoyingability despite capturing many images.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent divides the continuous image sequence into discrete noteworthy sections based on detection results, and selectively arranges these segmented sections in the summary image. This segmentation allows the system to present only the interesting parts, improving enjoyability.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If technologies detect noteworthy sections from original images, then interesting points are identified, but the sections are adopted in their original states without optimization

Engineering Contradiction:
Improvedetection accuracy of noteworthy sectionsVSAvoidprocessing complexity for summary image generation
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent performs preliminary detection of noteworthy sections before generating the summary image. By pre-identifying and marking the interesting sections, the system simplifies the subsequent summary generation process, making it easier to create optimized summaries without re-processing the entire image set.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary processing step that takes the detected noteworthy sections and transforms them into an optimized summary format. This intermediary process bridges the detection phase and the final output, enabling efficient summary generation with reduced complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If summary images are generated by abbreviating interesting points, then image content is condensed, but the process lacks systematic section adoption positioning

Engineering Contradiction:
Improveefficiency of image summarizationVSAvoidcomplexity of section positioning system
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent employs a dynamic section adoption positioning system that automatically adjusts the selection and arrangement of noteworthy sections based on their detected characteristics. This dynamic approach enables efficient summarization without requiring complex manual positioning, as the system adaptively determines the optimal summary composition.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10984248B2Setting of input images based on input music
Publication Date: 2021.04.20 SONY GROUP CORP
  • US10984248B2 patent drawing
  • US10984248B2 patent drawing
  • US10984248B2 patent drawing

AI summary

An information processing apparatus includes one or more processors configured to analyze content of a plurality of input images, extract one or more unit images from the plurality of input images based on the analysis and set a position of each of the one or more unit images adopted in a summary image based on an input music.