Moving Image Audio Analysis for Scene Search Without Keywords

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video editing systems require users to recognize and input keywords to find specific scenes in a moving image, making it difficult to identify desired scenes when keywords are unknown or forgotten.

Innovation Solution

An information processing apparatus that analyzes voice information in moving images, identifies speech errors and sound types, and displays corresponding scene information for easy selection by the user, allowing for efficient scene identification and editing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If keyword-based video search is used, then video scene identification is possible, but it becomes difficult to identify desired scenes when keywords are unknown or forgotten

Engineering Contradiction:
Improvevideo scene identification accuracyVSAvoiduser operation convenience
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent introduces sound type information (clapping, cheering, applause, etc.) as an intermediary element between the user and the video content. Instead of requiring direct keyword input, the system automatically extracts and presents sound type characteristics that serve as mediators for scene identification, allowing users to find desired scenes without needing to know specific keywords

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs self-service by automatically extracting voice information, analyzing sound types, and generating searchable metadata without user intervention. The video processing system itself generates the search indices (sound type information) that users will later query, eliminating the need for users to manually input keywords or understand the content beforehand

Inventive Principle:
Principle #25Self-service

2Loss of time

If manual keyword recognition is required, then specific scenes can be found, but editing burden and time consumption increase

Engineering Contradiction:
Improvescene search timeVSAvoidediting efficiency
Core Design Contradiction:
Loss of timeVSProductivity

Solution Approach 1:

The system performs preliminary action by automatically extracting voice information and analyzing sound types during video processing, before the user needs to search for scenes. This pre-computation of sound type metadata eliminates the need for time-consuming manual keyword recognition during the editing phase, significantly reducing both search time and editing burden

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250383755A1Information processing apparatus, control method, and recording medium
Publication Date: 2025.12.18 CANON KK
  • US20250383755A1 patent drawing
  • US20250383755A1 patent drawing
  • US20250383755A1 patent drawing

AI summary

An information processing apparatus includes one or more memories storing instructions, and one or more processors in communication with the one or more memories, that upon execution of the stored instructions, configures the one or more processors to acquire voice information included in a moving image, analyze the acquired voice information, based on a result of the analysis, display information in a manner that allows for selection by a user, wherein the display information includes either or both of information indicating a speech error and information indicating a sound type included in the voice information, and in accordance with information being selected by the user, display information regarding a time at which sound corresponding to the selected information is emitted.