Image Processing Apparatus Mixed Voice Character String Combination
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for combining character strings with still or moving images fail to produce meaningful results when the voice within the moving image is a mixed voice, leading to unnatural character strings being generated.
Innovation Solution
An image processing apparatus that selects parts of a moving image, extracts voices during specific times, and combines character strings based on predetermined conditions, ensuring that only appropriate character strings are used when the voice is mixed, using a combination of selection, extraction, and processing units to handle mixed voices effectively.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a character string corresponding to the voice is combined in a moving image, then the image processing function is improved, but unnatural character strings are generated when the voice is a mixed voice
Solution Approach 1:
The system performs preliminary voice quality assessment before combining character strings with images. By evaluating whether the voice is a mixed voice in advance, the system determines whether to proceed with automatic character string combination or to use predetermined character strings, thereby preventing unnatural combinations from occurring in the first place
Solution Approach 2:
The system introduces an intermediate evaluation step that acts as a mediator between voice extraction and character string combination. This intermediary process assesses voice quality and controls the flow to appropriate processing paths, ensuring that character strings are only combined when the voice is suitable
2Productivity
If automatic character string combination is performed, then processing efficiency is improved, but meaningless character strings are generated from mixed voices
Solution Approach 1:
The system performs preliminary voice quality assessment before combining character strings with images. By evaluating whether the voice is a mixed voice in advance, the system determines whether to proceed with automatic character string combination or to use predetermined character strings, thereby preventing unnatural combinations from occurring in the first place
Solution Approach 2:
The system introduces an intermediate evaluation step that acts as a mediator between voice extraction and character string combination. This intermediary process assesses voice quality and controls the flow to appropriate processing paths, ensuring that character strings are only combined when the voice is suitable
Data Source
AI summary
An object of one embodiment of the present disclosure is to provide a product with a high added value to a user by preventing an erroneous character string from being combined in a case where a voice before and after an image selected from within a moving image is a mixed voice. One embodiment of the present disclosure is an image processing apparatus including: a selection unit configured to select, from a moving image including a plurality of frames, a part of the moving image; an extraction unit configured to extract a voice during a predetermined time corresponding to the selected part in the moving image; and a combination unit configured to combine a character string based on a voice extracted by the extraction unit, with the part of the moving image selected by the selection unit or a frame among frames corresponding to the part.


