Speaker Identification System with Dynamic Image Overlay

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speaker identification systems face challenges in displaying speaker information clearly and simply on non-dedicated devices like TVs, particularly in cases of erroneous detection, where users lack a straightforward method to correct errors and the display becomes cluttered, especially when multiple speakers are identified consecutively.

Innovation Solution

A speaker identification method and system that acquires voice signals from speakers, generates speaker voice signals, identifies registered voice signals, and displays associated speaker images on a display device like a TV, allowing for easy correction of erroneous detections by overwriting the registered voice signals and using a database to store and manage voice information with user attributes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If speaker identification is performed continuously to improve accuracy, then identification reliability is improved, but display clutter increases when multiple speakers are identified consecutively

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoiddisplay clutter
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by displaying speaker images before voice recognition results are finalized. This allows the display to prepare for upcoming speaker information without waiting for complete processing, improving responsiveness while maintaining clear display management through timeout mechanisms that automatically remove images after a set period.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The display system dynamically adjusts by showing only the most recent speaker image at any given time, rather than accumulating all speaker images. The display updates in real-time based on current speaker identification, automatically replacing previous images to maintain clarity while continuously tracking multiple speakers through sequential updates.

Inventive Principle:
Principle #15Dynamics

2Productivity

If speaker information is displayed immediately to improve user experience, then responsiveness is improved, but original content on the display may be obstructed

Engineering Contradiction:
Improveresponse speedVSAvoiddisplay area availability
Core Design Contradiction:
ProductivityVSArea of stationary object

Solution Approach 1:

The speaker image is displayed in a specific localized region (corner or overlay area) rather than occupying the entire display. This allows the original content to remain visible in the main display area while speaker information is presented in a dedicated local zone, maintaining both responsiveness and content visibility.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The speaker image display operates periodically with automatic timeout removal. The image appears immediately when a speaker is identified, remains visible for a predetermined period, then automatically disappears. This periodic display pattern ensures responsive information delivery while periodically clearing the display to prevent obstruction.

Inventive Principle:
Principle #19Periodic action

3Loss of information

If voice recognition results are displayed as text to improve information delivery, then information completeness is improved, but ease of understanding decreases for users with reading difficulties

Engineering Contradiction:
Improveinformation completenessVSAvoidease of understanding
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The system creates a visual copy of the speaker representation (speaker image) that conveys identification information without requiring text reading. The speaker image serves as a visual alternative to text-based voice recognition results, providing the same identification function through image recognition rather than reading, thus accommodating users with different processing needs.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS9710219B2Speaker identification method, speaker identification device, and speaker identification system
Publication Date: 2017.07.18 PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
  • US9710219B2 patent drawing
  • US9710219B2 patent drawing
  • US9710219B2 patent drawing

AI summary

The present disclosure is a speaker identification method in a speaker identification system. The system stores registered voice signals and speaker images, the registered voice signals being respectively generated based on voices of speakers, the speaker images being respectively associated with the registered voice signals and respectively representing the speakers. The method includes: acquiring voice of a speaker positioned around a display; generating a speaker voice signal from the voice of the speaker; identifying a registered voice signal corresponding to the speaker voice signal, from the stored registered voice signals; and displaying the speaker image, which is associated with the identified registered voice signal, on the display, at least while the voice of the speaker which forms a basis of generation of the speaker voice signal is being acquired.