Speech Image Caption Layout Without Face Overlap

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing information processing systems fail to allow users to understand spoken content of a speaking person while simultaneously seeing their facial expressions due to overlapping display configurations.

Innovation Solution

An information processing system that acquires a speech image including a speaking person and controls the display of spoken content in a specific region that does not overlap with the face, determining the region based on size, visibility, and features to enhance clarity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Area of stationary object

If spoken content is displayed in a region overlapping with the face of the speaking person, then the display region is maximized, but the facial expression becomes difficult to understand

Engineering Contradiction:
Improvedisplay region areaVSAvoidfacial expression visibility
Core Design Contradiction:
Area of stationary objectVSLoss of information

Solution Approach 1:

The display screen is segmented into multiple regions: a first region for displaying spoken content and a second region for displaying the speech image. The system determines a non-overlapping display region by excluding the face area from the speech image, thereby separating the text display function from the visual expression display function while maximizing the available display area.

Inventive Principle:
Principle #1Segmentation

2Loss of information

If spoken content is displayed in a region not overlapping with the face, then facial expression visibility is maintained, but the display region area is reduced

Engineering Contradiction:
Improvefacial expression visibilityVSAvoiddisplay region area
Core Design Contradiction:
Loss of informationVSArea of stationary object

Solution Approach 1:

The system transitions from a single-dimension display approach to a two-dimensional region allocation approach. By defining specific coordinate ranges for the first and second regions, the system optimizes the use of available screen space while maintaining non-overlapping display areas, thereby maximizing the display region area without compromising facial expression visibility.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Loss of information

If the display region size is reduced to avoid overlap, then facial expression visibility is maintained, but the readability of spoken content decreases

Engineering Contradiction:
Improvefacial expression visibilityVSAvoidspoken content readability
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The system dynamically adjusts display parameters including the size, position, and font characteristics of the spoken content based on the determined non-overlapping region. By optimizing these parameters within the available display region, the system maintains high readability of spoken content while preserving facial expression visibility.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP4664444A1Information processing system, program, and information processing method
Publication Date: 2025.12.17 FUJIFILM BUSINESS INNOVATION CORP
  • EP4664444A1 patent drawingFigure 1
  • EP4664444A1 patent drawingFigure 2
  • EP4664444A1 patent drawingFigure 3

AI summary

An information processing system includes one or plural processors configured to acquire a speech image including a speaking person, acquire display information for displaying spoken content of the speaking person, and perform a control of displaying the display information in a specific region not overlapping with a face of the speaking person in the speech image.