Moving Image Audio Analysis for Scene Search Without Keywords
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video editing systems require users to recognize and input keywords to find specific scenes in a moving image, making it difficult to identify desired scenes when keywords are unknown or forgotten.
Innovation Solution
An information processing apparatus that analyzes voice information in moving images, identifies speech errors and sound types, and displays corresponding scene information for easy selection by the user, allowing for efficient scene identification and editing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If keyword-based video search is used, then video scene identification is possible, but it becomes difficult to identify desired scenes when keywords are unknown or forgotten
Solution Approach 1:
The patent introduces sound type information (clapping, cheering, applause, etc.) as an intermediary element between the user and the video content. Instead of requiring direct keyword input, the system automatically extracts and presents sound type characteristics that serve as mediators for scene identification, allowing users to find desired scenes without needing to know specific keywords
Solution Approach 2:
The system performs self-service by automatically extracting voice information, analyzing sound types, and generating searchable metadata without user intervention. The video processing system itself generates the search indices (sound type information) that users will later query, eliminating the need for users to manually input keywords or understand the content beforehand
2Loss of time
If manual keyword recognition is required, then specific scenes can be found, but editing burden and time consumption increase
Solution Approach 1:
The system performs preliminary action by automatically extracting voice information and analyzing sound types during video processing, before the user needs to search for scenes. This pre-computation of sound type metadata eliminates the need for time-consuming manual keyword recognition during the editing phase, significantly reducing both search time and editing burden
Data Source
AI summary
An information processing apparatus includes one or more memories storing instructions, and one or more processors in communication with the one or more memories, that upon execution of the stored instructions, configures the one or more processors to acquire voice information included in a moving image, analyze the acquired voice information, based on a result of the analysis, display information in a manner that allows for selection by a user, wherein the display information includes either or both of information indicating a speech error and information indicating a sound type included in the voice information, and in accordance with information being selected by the user, display information regarding a time at which sound corresponding to the selected information is emitted.


