Image Processing Device for Composite Image Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image processing systems fail to effectively combine frame images with character strings corresponding to voices from moving images, leading to limited variation in composite images and difficulty in selecting the correct voice for individuals in multi-person scenarios.
Innovation Solution
An image processing device that extracts frame images, detects person regions, evaluates and specifies central persons, converts voices to character strings, and generates association information to combine frame images with character strings, allowing for the selection of representative images and strings based on user input and correction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If frame images and character strings corresponding to voices at the time point when frame images are captured are combined, then the composite image can be generated, but character strings corresponding to voices at other time points cannot be combined, resulting in no variation in the composite image
Solution Approach 1:
The system performs preliminary voice extraction and association processing during the image processing stage. Voice data is extracted from the moving image, associated with corresponding frame images, and stored in advance. This preliminary action enables users to access and combine multiple voice-character string pairs with frame images without complex real-time processing, thereby increasing composite image variation while maintaining simplicity.
Solution Approach 2:
The system segments voice data into multiple discrete voice-character string pairs, each associated with specific frame images. Instead of treating voice data as a single unified element, the system divides it into separable units that can be independently selected and combined with different frame images, enabling multiple combination variations without increasing overall system complexity.
2Adaptability or versatility
If frame images and character strings corresponding to voices at time points other than the time point when frame images are captured are combined, then variation in composite images can be achieved, but a considerable effort is necessary for determining which person a voice belongs to or for selecting a desired voice from plural voices
Solution Approach 1:
The system introduces an intermediary association mechanism that links voice data with frame images through intermediate data structures. Voice-character string pairs are associated with frame images via time point information and person identification data, serving as intermediaries that automatically match voices with corresponding persons. This intermediary structure eliminates the need for users to manually determine voice-person associations, reducing selection effort while maintaining composite image variation.
Solution Approach 2:
The system provides feedback by displaying multiple candidate voice-character string pairs with their associated frame images to users. This feedback mechanism allows users to see available options and make informed selections without having to perform complex analysis. The system processes and presents voice options systematically, reducing the cognitive effort required for voice selection while preserving the ability to generate varied composite images.
3Quantity of substance
If multiple voices are extracted from a moving image with multiple persons, then more voice options are available, but it becomes difficult to determine which voice corresponds to which person
Solution Approach 1:
The system applies local quality by associating specific attributes (time point, person identification) with each voice-character string pair. Instead of treating all voice data uniformly, the system attaches localized identification information to each voice segment, enabling precise tracking of which voice belongs to which person. This localized attribution maintains high voice-person association accuracy even when multiple voices are extracted from a moving image with multiple persons.
Solution Approach 2:
The system performs preliminary voice-person association processing by extracting voices, identifying corresponding persons, and creating associated data structures before user interaction. During this preliminary stage, the system analyzes the moving image to link each voice segment with the correct person using techniques such as face recognition, spatial positioning, or temporal correlation. This advance processing ensures accurate voice-person mapping is established before users need to make selections, maintaining precision while providing multiple voice options.
Data Source
AI summary
In the image processing device, a display section displays, based on association information, association information between an icon indicating a central person and a character string corresponding to a voice of the central person. An instruction reception section receives a target person designation instruction for designating a central person corresponding to an icon selected by the user as a combination target person, and then, a combination section combines a frame image at an arbitrary time point when the target person is present with a character image of a character string corresponding to a voice of the target person in an arbitrary time period, according to the target person designation instruction, to generate a composite image.


