Smart Glasses Speaker Selection via Gaze Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for interpreting conversations among multiple people speaking different languages are inefficient due to lack of clear speaker selection, excessive resource consumption, and accuracy degradation from crosstalk in short-range communication networks.

Innovation Solution

An apparatus and method using smart glasses with a camera, gaze-tracking camera, and interpretation target processor to capture and track interlocutors, generate a virtual conversation space, and select a target person for interpretation based on user gaze, allowing for efficient and accurate connection of speakers in a conversation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If N-to-N connection method is used for multiple people, then all languages can be interpreted into all languages, but excessive costs are required for the system

Engineering Contradiction:
Improvelanguage interpretation coverageVSAvoidsystem cost
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent divides the N-to-N interpretation problem into N separate 1-to-N interpretation tasks. Each participant is assigned a dedicated interpretation terminal that interprets only the speech of other participants into their native language, rather than requiring all participants to be connected to all other participants simultaneously. This segmentation reduces the total number of language pairs that need to be supported.

Inventive Principle:
Principle #1Segmentation

2Ease of operation

If intensity measurement of short-range network is used for connecting multiple users, then connections can be established, but accuracy in interlocutor connection is degraded due to crosstalk

Engineering Contradiction:
Improveconnection establishmentVSAvoidinterlocutor connection accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent introduces a camera-based visual identification system as an intermediary between the short-range communication and the interpretation process. The camera captures images of participants, and the system identifies interlocutors based on visual recognition rather than relying solely on signal intensity measurements. This intermediary mechanism eliminates crosstalk issues by providing unambiguous visual identification of the intended interlocutor.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If no clarified method of changing speakers is provided, then people can speak freely, but people may speak at the same time

Engineering Contradiction:
Improvespeaking freedomVSAvoidspeaker turn management
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent implements a feedback mechanism where the system continuously monitors who is speaking and provides visual feedback to participants through the smart glasses display. When a participant finishes speaking, the system highlights the next speaker in the sequence, providing clear visual cues about whose turn it is to speak. This feedback loop helps participants understand speaking turns without requiring explicit verbal announcements.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10796106B2Apparatus and method for selecting speaker by using smart glasses
Publication Date: 2020.10.06 ELECTRONICS & TELECOMM RES INST
  • US10796106B2 patent drawing
  • US10796106B2 patent drawing
  • US10796106B2 patent drawing

AI summary

Provided are an apparatus and method for selecting a speaker by using smart glasses. The apparatus includes a camera configured to capture a front angle video of a user and track guest interpretation interlocutors in the captured video, smart glasses configured to display a virtual space map image including the guest interpretation interlocutors tracked through the camera, a gaze-tracking camera configured to select a target person for interpretation by tracking a gaze of the user so that a guest interpretation interlocutor displayed in the video may be selected, and an interpretation target processor configured to provide an interpretation service in connection with the target person selected through the gaze-tracking camera.