Speech Image Caption Layout Without Face Overlap
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing information processing systems fail to allow users to understand spoken content of a speaking person while simultaneously seeing their facial expressions due to overlapping display configurations.
Innovation Solution
An information processing system that acquires a speech image including a speaking person and controls the display of spoken content in a specific region that does not overlap with the face, determining the region based on size, visibility, and features to enhance clarity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If spoken content is displayed in a region overlapping with the face of the speaking person, then the display region is maximized, but the facial expression becomes difficult to understand
Solution Approach 1:
The display screen is segmented into multiple regions: a first region for displaying spoken content and a second region for displaying the speech image. The system determines a non-overlapping display region by excluding the face area from the speech image, thereby separating the text display function from the visual expression display function while maximizing the available display area.
2Loss of information
If spoken content is displayed in a region not overlapping with the face, then facial expression visibility is maintained, but the display region area is reduced
Solution Approach 1:
The system transitions from a single-dimension display approach to a two-dimensional region allocation approach. By defining specific coordinate ranges for the first and second regions, the system optimizes the use of available screen space while maintaining non-overlapping display areas, thereby maximizing the display region area without compromising facial expression visibility.
3Loss of information
If the display region size is reduced to avoid overlap, then facial expression visibility is maintained, but the readability of spoken content decreases
Solution Approach 1:
The system dynamically adjusts display parameters including the size, position, and font characteristics of the spoken content based on the determined non-overlapping region. By optimizing these parameters within the available display region, the system maintains high readability of spoken content while preserving facial expression visibility.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An information processing system includes one or plural processors configured to acquire a speech image including a speaking person, acquire display information for displaying spoken content of the speaking person, and perform a control of displaying the display information in a specific region not overlapping with a face of the speaking person in the speech image.