Audio-Linked Image Layouts with Text for Scene Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in selecting desired scenes for generating layouts from private events, such as personal photos, making it difficult to create meaningful memories.
Innovation Solution
An information processing apparatus that acquires image and audio data, converts audio to text, selects desired images, and generates layout images with associated text, using machine learning and audio recognition techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If audio recognition technique is used to convert audio data to text and generate layout, then text association with images is improved, but difficulty in selecting desired scenes remains
Solution Approach 1:
The patent replaces manual scene selection with automatic scene selection based on audio analysis. The system detects scenes where the specific person is speaking by analyzing audio data, automatically determining which images should be included in the layout without requiring manual intervention.
Solution Approach 2:
The system performs self-service by automatically selecting appropriate scenes for layout generation. Through audio recognition and speaker identification, the system autonomously determines which images contain the specific person speaking, eliminating the need for user intervention in scene selection.
2Ease of operation
If automatic scene selection is implemented, then ease of operation is improved, but selection accuracy may deteriorate
Solution Approach 1:
The system uses audio feedback to improve scene selection accuracy. By continuously analyzing audio data and identifying when the specific person is speaking, the system provides feedback that guides the automatic selection process, ensuring that only relevant scenes are included in the layout.
Solution Approach 2:
The patent segments the audio data into distinct speech events associated with specific persons. By dividing the audio stream and matching it with corresponding images, the system achieves accurate scene selection through systematic segmentation of audio-visual data.
Data Source
AI summary
An information processing apparatus capable of selecting a desired image, and generating a layout image including the selected image and text associated therewith. A network controller acquires image data including an image of a person, and audio data associated with the image data, and an audio data conversion section converts the acquired audio data to text data. The image data is selected on an operation screen, and a layout image generation section generates a layout image in which are arranged specific text data of voice uttered by the person included in the selected image data, in the text data, and the selected image data.


